{
  "id": 255338,
  "title": "same pipeline different models break the submission",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/255338",
  "author_name": "",
  "post_date": "2021-07-27T05:27:49.191257600Z",
  "votes": 3,
  "comment_count": 15,
  "views": 0,
  "content": "<p>I have same pipeline but using different models (all lightgbm models), the submission returning an error: <strong>Submission Scoring Error</strong></p>\n<p>Anyone facing the same issue? </p>",
  "messages": [
    {
      "id": "1401215",
      "postDate": "07/27/2021 05:27:49",
      "content": "<p>I have same pipeline but using different models (all lightgbm models), the submission returning an error: <strong>Submission Scoring Error</strong></p>\n<p>Anyone facing the same issue? </p>",
      "rawMarkdown": "I have same pipeline but using different models (all lightgbm models), the submission returning an error: **Submission Scoring Error**\n\n\nAnyone facing the same issue?",
      "votes": null
    },
    {
      "id": "1401237",
      "postDate": "07/27/2021 06:13:13",
      "content": "<p>Any kaggle staff can help? <a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a>  <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> </p>\n<p>Possible to have a check on  my notebook? It's simple notebook for inference only. </p>\n<p><a href=\"https://www.kaggle.com/lsw1993/999-sub\" target=\"_blank\">https://www.kaggle.com/lsw1993/999-sub</a> (Made it visible to public for temporary review)</p>",
      "rawMarkdown": "Any kaggle staff can help? @wcukierski  @juliaelliott \n\nPossible to have a check on  my notebook? It's simple notebook for inference only. \n\nhttps://www.kaggle.com/lsw1993/999-sub (Made it visible to public for temporary review)",
      "votes": null
    },
    {
      "id": "1401241",
      "postDate": "07/27/2021 06:20:34",
      "content": "<p>Feel sucks - wasted all the submissions - but cannot get it working. </p>\n<p>What's going on? It went well on test but the error only occurs during submission…. @@ Really hate this.</p>",
      "rawMarkdown": "Feel sucks - wasted all the submissions - but cannot get it working. \n\nWhat's going on? It went well on test but the error only occurs during submission.... @@ Really hate this.",
      "votes": null
    },
    {
      "id": "1401464",
      "postDate": "07/27/2021 10:48:26",
      "content": "<p><a href=\"https://www.kaggle.com/lsw1993\" target=\"_blank\">@lsw1993</a> <br>\nYou can check your submission with API emulator. Hope this notebook helps:<br>\n<a href=\"https://www.kaggle.com/nyanpn/api-emulator-for-debugging-your-code-locally\" target=\"_blank\">https://www.kaggle.com/nyanpn/api-emulator-for-debugging-your-code-locally</a></p>",
      "rawMarkdown": "lsw1993 \nYou can check your submission with API emulator. Hope this notebook helps:\nhttps://www.kaggle.com/nyanpn/api-emulator-for-debugging-your-code-locally",
      "votes": null
    },
    {
      "id": "1401545",
      "postDate": "07/27/2021 12:02:46",
      "content": "<p><a href=\"https://ibb.co/BcQytH7\" target=\"_blank\">https://ibb.co/BcQytH7</a></p>\n<p>Thanks but it's perfectly fine….. To put into context, it's same models with different parameters.</p>",
      "rawMarkdown": "https://ibb.co/BcQytH7\n\nThanks but it's perfectly fine..... To put into context, it's same models with different parameters.",
      "votes": null
    },
    {
      "id": "1401592",
      "postDate": "07/27/2021 12:27:28",
      "content": "<p><a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a> <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <br>\nMy other queries: </p>\n<ol>\n<li><p>Is there a limit to the Kaggle dataset I am using (or connecting to the submission)?</p></li>\n<li><p>If test had returned smoothly, why my final submission will fail?</p></li>\n<li><p>Is it possible not to count failed submissions? </p></li>\n</ol>",
      "rawMarkdown": "wcukierski @juliaelliott \nMy other queries: \n1. Is there a limit to the Kaggle dataset I am using (or connecting to the submission)?\n\n2. If test had returned smoothly, why my final submission will fail?\n\n3. Is it possible not to count failed submissions?",
      "votes": null
    },
    {
      "id": "1401633",
      "postDate": "07/27/2021 12:58:28",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/lsw1993\" target=\"_blank\">@lsw1993</a>, we unfortunately cannot provide personal debugging support, but <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">https://www.kaggle.com/code-competition-debugging</a> provides answers to some of your questions. On (1), there is no limit to using a dataset you have attached to a notebook. Good luck!</p>",
      "rawMarkdown": "Hey @lsw1993, we unfortunately cannot provide personal debugging support, but https://www.kaggle.com/code-competition-debugging provides answers to some of your questions. On (1), there is no limit to using a dataset you have attached to a notebook. Good luck!",
      "votes": null
    },
    {
      "id": "1401642",
      "postDate": "07/27/2021 13:10:11",
      "content": "<p>(3)? Considering the competition is ending.</p>",
      "rawMarkdown": "(3)? Considering the competition is ending.",
      "votes": null
    },
    {
      "id": "1401644",
      "postDate": "07/27/2021 13:12:16",
      "content": "<p>Made my notebook public at: <a href=\"https://www.kaggle.com/lsw1993/999-sub\" target=\"_blank\">https://www.kaggle.com/lsw1993/999-sub</a></p>",
      "rawMarkdown": "Made my notebook public at: https://www.kaggle.com/lsw1993/999-sub",
      "votes": null
    },
    {
      "id": "1401667",
      "postDate": "07/27/2021 13:31:23",
      "content": "<p>Per the code debugging page:</p>\n<blockquote>\n  <p>Submissions that error also count towards your team’s daily submission limit, otherwise such submissions could be used to mine for hidden information.</p>\n</blockquote>",
      "rawMarkdown": "Per the code debugging page:\n\n> Submissions that error also count towards your team’s daily submission limit, otherwise such submissions could be used to mine for hidden information.",
      "votes": null
    },
    {
      "id": "1401691",
      "postDate": "07/27/2021 13:51:41",
      "content": "<p><a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a> <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a></p>\n<p>Wouldn't it be systemic error/shortage if test is successful but not the final submission? <br>\nFurther more details to be shown would be easier for us to debug moving forwards in my opinion. </p>\n<p>Perhaps a printed line of last run error would be suffice.</p>",
      "rawMarkdown": "wcukierski @juliaelliott\n\nWouldn't it be systemic error/shortage if test is successful but not the final submission? \nFurther more details to be shown would be easier for us to debug moving forwards in my opinion. \n\nPerhaps a printed line of last run error would be suffice.",
      "votes": null
    },
    {
      "id": "1401845",
      "postDate": "07/27/2021 15:43:28",
      "content": "<p><a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">Will</a> - it still amazes me that kaggle staff believes that folks can mine for hidden information - for sure the future is hidden - what the heck could someone learn by mining in this competition.  </p>\n<p>It's fine that this is the company line but for the furtherance of my statistical training please provide a single sample of what could be learned by mining when the final private test has yet to occur and all the data used in the public LB has already been shared.</p>",
      "rawMarkdown": "[Will](https://www.kaggle.com/wcukierski) - it still amazes me that kaggle staff believes that folks can mine for hidden information - for sure the future is hidden - what the heck could someone learn by mining in this competition.  \n\nIt's fine that this is the company line but for the furtherance of my statistical training please provide a single sample of what could be learned by mining when the final private test has yet to occur and all the data used in the public LB has already been shared.",
      "votes": null
    },
    {
      "id": "1401926",
      "postDate": "07/27/2021 16:42:26",
      "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> prior to \"sync rerun\" code competitions, competitions like this would suffer from large failure rates during the 2nd stage (as high as 40% in the worst cases). Participants would adapt to the realities of the stage one dataset but were not incentivized to future proof their models. When new data would arrive and there is a 0 where there had never been a 0 before, code failed, we can't intervene, and lots of unhappy folks were left without a rank.</p>\n<p>The benefit of a truly blind holdout set (whether the data is in the future or not) is that it front loads this \"tough love\" into stage one. For example, let's pretend the holdout set has a funny value that says a player is -92 ft tall. If you can see this fact, you might choose to patch or ignore that row. If you can't, you are pushed to make your code resilient to entire classes of bad inputs, missing values, impossible data events, and so on. The net result of this has been a dramatic reduction in rerun failure rates.</p>\n<p>I don't deny this can be frustrating. I also agree the probing safeguards are less relevant to true predict-the-future competitions. The pain we endure now is to create robust ML models that will run correctly on data they've never seen.</p>",
      "rawMarkdown": "pcjimmmy prior to \"sync rerun\" code competitions, competitions like this would suffer from large failure rates during the 2nd stage (as high as 40% in the worst cases). Participants would adapt to the realities of the stage one dataset but were not incentivized to future proof their models. When new data would arrive and there is a 0 where there had never been a 0 before, code failed, we can't intervene, and lots of unhappy folks were left without a rank.\n\nThe benefit of a truly blind holdout set (whether the data is in the future or not) is that it front loads this \"tough love\" into stage one. For example, let's pretend the holdout set has a funny value that says a player is -92 ft tall. If you can see this fact, you might choose to patch or ignore that row. If you can't, you are pushed to make your code resilient to entire classes of bad inputs, missing values, impossible data events, and so on. The net result of this has been a dramatic reduction in rerun failure rates.\n\nI don't deny this can be frustrating. I also agree the probing safeguards are less relevant to true predict-the-future competitions. The pain we endure now is to create robust ML models that will run correctly on data they've never seen.",
      "votes": null
    },
    {
      "id": "1401940",
      "postDate": "07/27/2021 16:55:46",
      "content": "<p>OK - I guess I get it.  Self driving cars can be made more effective if they crash and no error data is available to the engineer - hence they have to code against all possible errors.  But of course only 5 crashes per day will be permitted :)</p>",
      "rawMarkdown": "OK - I guess I get it.  Self driving cars can be made more effective if they crash and no error data is available to the engineer - hence they have to code against all possible errors.  But of course only 5 crashes per day will be permitted :)",
      "votes": null
    },
    {
      "id": "1401993",
      "postDate": "07/27/2021 17:42:03",
      "content": "<p>Maybe the status can be made verbose, rather than just as such:</p>\n<p><a href=\"https://ibb.co/zR4DYH5\" target=\"_blank\">https://ibb.co/zR4DYH5</a></p>",
      "rawMarkdown": "Maybe the status can be made verbose, rather than just as such:\n\nhttps://ibb.co/zR4DYH5",
      "votes": null
    },
    {
      "id": "1402890",
      "postDate": "07/28/2021 16:00:23",
      "content": "<p>I can't seem to access the notebook</p>",
      "rawMarkdown": "I can't seem to access the notebook",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1401237,
      "author_name": "lsw1993",
      "author_url": "",
      "post_date": "07/27/2021 06:13:13",
      "content": "<p>Any kaggle staff can help? <a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a>  <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> </p>\n<p>Possible to have a check on  my notebook? It's simple notebook for inference only. </p>\n<p><a href=\"https://www.kaggle.com/lsw1993/999-sub\" target=\"_blank\">https://www.kaggle.com/lsw1993/999-sub</a> (Made it visible to public for temporary review)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1401592,
          "author_name": "lsw1993",
          "author_url": "",
          "post_date": "07/27/2021 12:27:28",
          "content": "<p><a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a> <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <br>\nMy other queries: </p>\n<ol>\n<li><p>Is there a limit to the Kaggle dataset I am using (or connecting to the submission)?</p></li>\n<li><p>If test had returned smoothly, why my final submission will fail?</p></li>\n<li><p>Is it possible not to count failed submissions? </p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1401633,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "07/27/2021 12:58:28",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/lsw1993\" target=\"_blank\">@lsw1993</a>, we unfortunately cannot provide personal debugging support, but <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">https://www.kaggle.com/code-competition-debugging</a> provides answers to some of your questions. On (1), there is no limit to using a dataset you have attached to a notebook. Good luck!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1401642,
          "author_name": "lsw1993",
          "author_url": "",
          "post_date": "07/27/2021 13:10:11",
          "content": "<p>(3)? Considering the competition is ending.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1401667,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "07/27/2021 13:31:23",
          "content": "<p>Per the code debugging page:</p>\n<blockquote>\n  <p>Submissions that error also count towards your team’s daily submission limit, otherwise such submissions could be used to mine for hidden information.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1401691,
          "author_name": "lsw1993",
          "author_url": "",
          "post_date": "07/27/2021 13:51:41",
          "content": "<p><a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a> <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a></p>\n<p>Wouldn't it be systemic error/shortage if test is successful but not the final submission? <br>\nFurther more details to be shown would be easier for us to debug moving forwards in my opinion. </p>\n<p>Perhaps a printed line of last run error would be suffice.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1401845,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "07/27/2021 15:43:28",
          "content": "<p><a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">Will</a> - it still amazes me that kaggle staff believes that folks can mine for hidden information - for sure the future is hidden - what the heck could someone learn by mining in this competition.  </p>\n<p>It's fine that this is the company line but for the furtherance of my statistical training please provide a single sample of what could be learned by mining when the final private test has yet to occur and all the data used in the public LB has already been shared.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1401926,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "07/27/2021 16:42:26",
          "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> prior to \"sync rerun\" code competitions, competitions like this would suffer from large failure rates during the 2nd stage (as high as 40% in the worst cases). Participants would adapt to the realities of the stage one dataset but were not incentivized to future proof their models. When new data would arrive and there is a 0 where there had never been a 0 before, code failed, we can't intervene, and lots of unhappy folks were left without a rank.</p>\n<p>The benefit of a truly blind holdout set (whether the data is in the future or not) is that it front loads this \"tough love\" into stage one. For example, let's pretend the holdout set has a funny value that says a player is -92 ft tall. If you can see this fact, you might choose to patch or ignore that row. If you can't, you are pushed to make your code resilient to entire classes of bad inputs, missing values, impossible data events, and so on. The net result of this has been a dramatic reduction in rerun failure rates.</p>\n<p>I don't deny this can be frustrating. I also agree the probing safeguards are less relevant to true predict-the-future competitions. The pain we endure now is to create robust ML models that will run correctly on data they've never seen.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1401940,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "07/27/2021 16:55:46",
          "content": "<p>OK - I guess I get it.  Self driving cars can be made more effective if they crash and no error data is available to the engineer - hence they have to code against all possible errors.  But of course only 5 crashes per day will be permitted :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1401993,
          "author_name": "lsw1993",
          "author_url": "",
          "post_date": "07/27/2021 17:42:03",
          "content": "<p>Maybe the status can be made verbose, rather than just as such:</p>\n<p><a href=\"https://ibb.co/zR4DYH5\" target=\"_blank\">https://ibb.co/zR4DYH5</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1401241,
      "author_name": "lsw1993",
      "author_url": "",
      "post_date": "07/27/2021 06:20:34",
      "content": "<p>Feel sucks - wasted all the submissions - but cannot get it working. </p>\n<p>What's going on? It went well on test but the error only occurs during submission…. @@ Really hate this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1401464,
      "author_name": "nyanpn",
      "author_url": "",
      "post_date": "07/27/2021 10:48:26",
      "content": "<p><a href=\"https://www.kaggle.com/lsw1993\" target=\"_blank\">@lsw1993</a> <br>\nYou can check your submission with API emulator. Hope this notebook helps:<br>\n<a href=\"https://www.kaggle.com/nyanpn/api-emulator-for-debugging-your-code-locally\" target=\"_blank\">https://www.kaggle.com/nyanpn/api-emulator-for-debugging-your-code-locally</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1401545,
          "author_name": "lsw1993",
          "author_url": "",
          "post_date": "07/27/2021 12:02:46",
          "content": "<p><a href=\"https://ibb.co/BcQytH7\" target=\"_blank\">https://ibb.co/BcQytH7</a></p>\n<p>Thanks but it's perfectly fine….. To put into context, it's same models with different parameters.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1401644,
          "author_name": "lsw1993",
          "author_url": "",
          "post_date": "07/27/2021 13:12:16",
          "content": "<p>Made my notebook public at: <a href=\"https://www.kaggle.com/lsw1993/999-sub\" target=\"_blank\">https://www.kaggle.com/lsw1993/999-sub</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1402890,
          "author_name": "redfoongus",
          "author_url": "",
          "post_date": "07/28/2021 16:00:23",
          "content": "<p>I can't seem to access the notebook</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1401215": "I have same pipeline but using different models (all lightgbm models), the submission returning an error: **Submission Scoring Error**\n\n\nAnyone facing the same issue?",
    "1401237": "Any kaggle staff can help? @wcukierski  @juliaelliott \n\nPossible to have a check on  my notebook? It's simple notebook for inference only. \n\nhttps://www.kaggle.com/lsw1993/999-sub (Made it visible to public for temporary review)",
    "1401241": "Feel sucks - wasted all the submissions - but cannot get it working. \n\nWhat's going on? It went well on test but the error only occurs during submission.... @@ Really hate this.",
    "1401464": "lsw1993 \nYou can check your submission with API emulator. Hope this notebook helps:\nhttps://www.kaggle.com/nyanpn/api-emulator-for-debugging-your-code-locally",
    "1401545": "https://ibb.co/BcQytH7\n\nThanks but it's perfectly fine..... To put into context, it's same models with different parameters.",
    "1401592": "wcukierski @juliaelliott \nMy other queries: \n1. Is there a limit to the Kaggle dataset I am using (or connecting to the submission)?\n\n2. If test had returned smoothly, why my final submission will fail?\n\n3. Is it possible not to count failed submissions?",
    "1401633": "Hey @lsw1993, we unfortunately cannot provide personal debugging support, but https://www.kaggle.com/code-competition-debugging provides answers to some of your questions. On (1), there is no limit to using a dataset you have attached to a notebook. Good luck!",
    "1401642": "(3)? Considering the competition is ending.",
    "1401644": "Made my notebook public at: https://www.kaggle.com/lsw1993/999-sub",
    "1401667": "Per the code debugging page:\n\n> Submissions that error also count towards your team’s daily submission limit, otherwise such submissions could be used to mine for hidden information.",
    "1401691": "wcukierski @juliaelliott\n\nWouldn't it be systemic error/shortage if test is successful but not the final submission? \nFurther more details to be shown would be easier for us to debug moving forwards in my opinion. \n\nPerhaps a printed line of last run error would be suffice.",
    "1401845": "[Will](https://www.kaggle.com/wcukierski) - it still amazes me that kaggle staff believes that folks can mine for hidden information - for sure the future is hidden - what the heck could someone learn by mining in this competition.  \n\nIt's fine that this is the company line but for the furtherance of my statistical training please provide a single sample of what could be learned by mining when the final private test has yet to occur and all the data used in the public LB has already been shared.",
    "1401926": "pcjimmmy prior to \"sync rerun\" code competitions, competitions like this would suffer from large failure rates during the 2nd stage (as high as 40% in the worst cases). Participants would adapt to the realities of the stage one dataset but were not incentivized to future proof their models. When new data would arrive and there is a 0 where there had never been a 0 before, code failed, we can't intervene, and lots of unhappy folks were left without a rank.\n\nThe benefit of a truly blind holdout set (whether the data is in the future or not) is that it front loads this \"tough love\" into stage one. For example, let's pretend the holdout set has a funny value that says a player is -92 ft tall. If you can see this fact, you might choose to patch or ignore that row. If you can't, you are pushed to make your code resilient to entire classes of bad inputs, missing values, impossible data events, and so on. The net result of this has been a dramatic reduction in rerun failure rates.\n\nI don't deny this can be frustrating. I also agree the probing safeguards are less relevant to true predict-the-future competitions. The pain we endure now is to create robust ML models that will run correctly on data they've never seen.",
    "1401940": "OK - I guess I get it.  Self driving cars can be made more effective if they crash and no error data is available to the engineer - hence they have to code against all possible errors.  But of course only 5 crashes per day will be permitted :)",
    "1401993": "Maybe the status can be made verbose, rather than just as such:\n\nhttps://ibb.co/zR4DYH5",
    "1402890": "I can't seem to access the notebook"
  },
  "source": "meta"
}