{
  "id": 249673,
  "title": "Yet another scoring error !!! I need help",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/249673",
  "author_name": "",
  "post_date": "2021-06-29T10:43:07.634845500Z",
  "votes": 6,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi kagglers, i have a submission scoring error on one my kernels. Yet all my test were succesful, i checked the number of lines to be predicted, fixed missing data, reduced memory usage, etc. Therefore i have no idea of what is going on behind the hood. Any help will be appreciated. Here is the <a href=\"https://www.kaggle.com/ulrich07/mlb-fe-pipeline-normal-training\" target=\"_blank\">link</a> to the notebook.</p>\n<p>Thanks in advance.</p>\n<p><strong>Partial Solution</strong>: Building on <a href=\"https://www.kaggle.com/mlconsult/1-38-lb-lightgbm-with-target-statistics\" target=\"_blank\">this notebook</a>, i had a partial solution with <a href=\"https://www.kaggle.com/ulrich07/mlb-debug-ann\" target=\"_blank\">this</a>. </p>\n<p><strong>Special thanks</strong> to <a href=\"https://www.kaggle.com/assign\" target=\"_blank\">@assign</a> for pointing out the code block creating this error <a href=\"https://www.kaggle.com/assign/debug-others-mlb-fe-pipeline-normal-training\" target=\"_blank\">here</a>.</p>",
  "messages": [
    {
      "id": "1369423",
      "postDate": "06/29/2021 10:43:07",
      "content": "<p>Hi kagglers, i have a submission scoring error on one my kernels. Yet all my test were succesful, i checked the number of lines to be predicted, fixed missing data, reduced memory usage, etc. Therefore i have no idea of what is going on behind the hood. Any help will be appreciated. Here is the <a href=\"https://www.kaggle.com/ulrich07/mlb-fe-pipeline-normal-training\" target=\"_blank\">link</a> to the notebook.</p>\n<p>Thanks in advance.</p>\n<p><strong>Partial Solution</strong>: Building on <a href=\"https://www.kaggle.com/mlconsult/1-38-lb-lightgbm-with-target-statistics\" target=\"_blank\">this notebook</a>, i had a partial solution with <a href=\"https://www.kaggle.com/ulrich07/mlb-debug-ann\" target=\"_blank\">this</a>. </p>\n<p><strong>Special thanks</strong> to <a href=\"https://www.kaggle.com/assign\" target=\"_blank\">@assign</a> for pointing out the code block creating this error <a href=\"https://www.kaggle.com/assign/debug-others-mlb-fe-pipeline-normal-training\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "Hi kagglers, i have a submission scoring error on one my kernels. Yet all my test were succesful, i checked the number of lines to be predicted, fixed missing data, reduced memory usage, etc. Therefore i have no idea of what is going on behind the hood. Any help will be appreciated. Here is the [link](https://www.kaggle.com/ulrich07/mlb-fe-pipeline-normal-training) to the notebook.\n\nThanks in advance.\n\n**Partial Solution**: Building on [this notebook](https://www.kaggle.com/mlconsult/1-38-lb-lightgbm-with-target-statistics), i had a partial solution with [this](https://www.kaggle.com/ulrich07/mlb-debug-ann). \n\n**Special thanks** to @assign for pointing out the code block creating this error [here](https://www.kaggle.com/assign/debug-others-mlb-fe-pipeline-normal-training).",
      "votes": null
    },
    {
      "id": "1369605",
      "postDate": "06/29/2021 13:22:25",
      "content": "<p>Is there any progress?<br>\nOr any doubts about them?</p>",
      "rawMarkdown": "Is there any progress?\nOr any doubts about them?",
      "votes": null
    },
    {
      "id": "1369613",
      "postDate": "06/29/2021 13:24:56",
      "content": "<p>I still have no idea. It is frustrating, i tried many things but still get \"Submission Scoring Error\". My today 5 submissions are over.</p>",
      "rawMarkdown": "I still have no idea. It is frustrating, i tried many things but still get \"Submission Scoring Error\". My today 5 submissions are over.",
      "votes": null
    },
    {
      "id": "1369659",
      "postDate": "06/29/2021 13:45:43",
      "content": "<p>one thing I observed was error event before finish <strong>iter_test</strong></p>",
      "rawMarkdown": "one thing I observed was error event before finish **iter_test**",
      "votes": null
    },
    {
      "id": "1369664",
      "postDate": "06/29/2021 13:51:09",
      "content": "<p><a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> sorry to hear about the error. Your notebook is doing quite a lot of processing without much in the way of error handling, which can unfortunately result in the silliest little issue failing the whole notebook.</p>\n<p>Check out the tips at the bottom of the <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">code competition debugging</a>. If I were you, I would (a) start by breaking your notebook into more modular steps that can be more readily tested on the unseen test set (b) wrap your processing functions in error handling that allows them to \"fail safe\" (c) make your goal to get <em>any</em> score from the whole pipeline (rip out anything that isn't mandatory, substitute fancy methods with dumb-but-safe alternatives). Careful about hard-coding assumptions of what the data is allowed to look like (categorical levels, missing values, etc).</p>\n<p>Good luck!</p>",
      "rawMarkdown": "ulrich07 sorry to hear about the error. Your notebook is doing quite a lot of processing without much in the way of error handling, which can unfortunately result in the silliest little issue failing the whole notebook.\n\nCheck out the tips at the bottom of the [code competition debugging](https://www.kaggle.com/code-competition-debugging). If I were you, I would (a) start by breaking your notebook into more modular steps that can be more readily tested on the unseen test set (b) wrap your processing functions in error handling that allows them to \"fail safe\" (c) make your goal to get _any_ score from the whole pipeline (rip out anything that isn't mandatory, substitute fancy methods with dumb-but-safe alternatives). Careful about hard-coding assumptions of what the data is allowed to look like (categorical levels, missing values, etc).\n\nGood luck!",
      "votes": null
    },
    {
      "id": "1369730",
      "postDate": "06/29/2021 14:34:58",
      "content": "<p>Yes,  i know it fails before the end of the loop. It looks like it enconters a configuration in the hidden dataset not accounted for in my existing functions. I will follow <a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a> advice to fix it.</p>",
      "rawMarkdown": "Yes,  i know it fails before the end of the loop. It looks like it enconters a configuration in the hidden dataset not accounted for in my existing functions. I will follow @wcukierski advice to fix it.",
      "votes": null
    },
    {
      "id": "1370185",
      "postDate": "06/30/2021 01:46:21",
      "content": "<p>in prepare_test, <strong>PART ONE</strong></p>\n<pre><code>    if hid_df[colname].iloc[0] ==  hid_df[colname].iloc[0]:\n        pscores = pd.DataFrame(eval(hid_df[colname].iloc[0]))\n        pscores = pscores.rename(columns=ren_player)\n        pscores[\"index\"] = (pd.to_datetime(pscores[\"gameDate\"]) - START_DATE).dt.days\n        pscores = pscores[pscores_id + pscores_cat + pscores_num]\n</code></pre>\n<p>include bug, others will not make error (I don't check net.predict(Xe) make error or not)</p>\n<p>I don't know which line makes error, but will helpful enough</p>",
      "rawMarkdown": "in prepare_test, **PART ONE**\n```\n\n    if hid_df[colname].iloc[0] ==  hid_df[colname].iloc[0]:\n        pscores = pd.DataFrame(eval(hid_df[colname].iloc[0]))\n        pscores = pscores.rename(columns=ren_player)\n        pscores[\"index\"] = (pd.to_datetime(pscores[\"gameDate\"]) - START_DATE).dt.days\n        pscores = pscores[pscores_id + pscores_cat + pscores_num]\n```\ninclude bug, others will not make error (I don't check net.predict(Xe) make error or not)\n\nI don't know which line makes error, but will helpful enough",
      "votes": null
    },
    {
      "id": "1370370",
      "postDate": "06/30/2021 06:02:02",
      "content": "<p>Many thks <a href=\"https://www.kaggle.com/assign\" target=\"_blank\">@assign</a> </p>",
      "rawMarkdown": "Many thks @assign",
      "votes": null
    },
    {
      "id": "1370646",
      "postDate": "06/30/2021 09:51:20",
      "content": "<p><a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> - posted elsewhere there seems to be an issue with Rosters. While train had rosters entry for all days, that does not seem to be the case in the submitted test versions so some days could be missing rosters and you may want to handle that.  So maybe do a check like pd.isna(…) around the columns with json data or whatever you prefer.</p>\n<p>Also atm there are issues with notebooks ending with -<br>\nrpc error: code = Unknown desc = generic::unknown: retry budget exhausted (10 attempts): oneshotUpload: googleapi: Error 503: { \"error\": { \"code\": 503, \"message\": \"Server is unavailable\" } }</p>\n<p>Nothing wrong in the Notebook but gets stuck at the Writing Notebook… and ends up Failed status even though it produces output like submission files which are fine.  Not sure what that is about. </p>",
      "rawMarkdown": "ulrich07 - posted elsewhere there seems to be an issue with Rosters. While train had rosters entry for all days, that does not seem to be the case in the submitted test versions so some days could be missing rosters and you may want to handle that.  So maybe do a check like pd.isna(...) around the columns with json data or whatever you prefer.\n\nAlso atm there are issues with notebooks ending with -\nrpc error: code = Unknown desc = generic::unknown: retry budget exhausted (10 attempts): oneshotUpload: googleapi: Error 503: { \"error\": { \"code\": 503, \"message\": \"Server is unavailable\" } }\n\nNothing wrong in the Notebook but gets stuck at the Writing Notebook... and ends up Failed status even though it produces output like submission files which are fine.  Not sure what that is about.",
      "votes": null
    },
    {
      "id": "1370978",
      "postDate": "06/30/2021 14:31:38",
      "content": "<p>Thks <a href=\"https://www.kaggle.com/something4kag\" target=\"_blank\">@something4kag</a> . It seems harder than it looks</p>",
      "rawMarkdown": "Thks @something4kag . It seems harder than it looks",
      "votes": null
    },
    {
      "id": "1372829",
      "postDate": "07/02/2021 04:26:36",
      "content": "<p>Just for curiosity. Since we have rosters for all days and the submitted test version has missing rosters, can Kaggle fix this issue from its end? It seems a fixable issue to me</p>",
      "rawMarkdown": "Just for curiosity. Since we have rosters for all days and the submitted test version has missing rosters, can Kaggle fix this issue from its end? It seems a fixable issue to me",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1369605,
      "author_name": "assign",
      "author_url": "",
      "post_date": "06/29/2021 13:22:25",
      "content": "<p>Is there any progress?<br>\nOr any doubts about them?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1369613,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "06/29/2021 13:24:56",
          "content": "<p>I still have no idea. It is frustrating, i tried many things but still get \"Submission Scoring Error\". My today 5 submissions are over.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1369659,
          "author_name": "assign",
          "author_url": "",
          "post_date": "06/29/2021 13:45:43",
          "content": "<p>one thing I observed was error event before finish <strong>iter_test</strong></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1369730,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "06/29/2021 14:34:58",
          "content": "<p>Yes,  i know it fails before the end of the loop. It looks like it enconters a configuration in the hidden dataset not accounted for in my existing functions. I will follow <a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a> advice to fix it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1369664,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "06/29/2021 13:51:09",
      "content": "<p><a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> sorry to hear about the error. Your notebook is doing quite a lot of processing without much in the way of error handling, which can unfortunately result in the silliest little issue failing the whole notebook.</p>\n<p>Check out the tips at the bottom of the <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">code competition debugging</a>. If I were you, I would (a) start by breaking your notebook into more modular steps that can be more readily tested on the unseen test set (b) wrap your processing functions in error handling that allows them to \"fail safe\" (c) make your goal to get <em>any</em> score from the whole pipeline (rip out anything that isn't mandatory, substitute fancy methods with dumb-but-safe alternatives). Careful about hard-coding assumptions of what the data is allowed to look like (categorical levels, missing values, etc).</p>\n<p>Good luck!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1370185,
      "author_name": "assign",
      "author_url": "",
      "post_date": "06/30/2021 01:46:21",
      "content": "<p>in prepare_test, <strong>PART ONE</strong></p>\n<pre><code>    if hid_df[colname].iloc[0] ==  hid_df[colname].iloc[0]:\n        pscores = pd.DataFrame(eval(hid_df[colname].iloc[0]))\n        pscores = pscores.rename(columns=ren_player)\n        pscores[\"index\"] = (pd.to_datetime(pscores[\"gameDate\"]) - START_DATE).dt.days\n        pscores = pscores[pscores_id + pscores_cat + pscores_num]\n</code></pre>\n<p>include bug, others will not make error (I don't check net.predict(Xe) make error or not)</p>\n<p>I don't know which line makes error, but will helpful enough</p>",
      "votes": null,
      "replies": [
        {
          "id": 1370370,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "06/30/2021 06:02:02",
          "content": "<p>Many thks <a href=\"https://www.kaggle.com/assign\" target=\"_blank\">@assign</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1370646,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "06/30/2021 09:51:20",
      "content": "<p><a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> - posted elsewhere there seems to be an issue with Rosters. While train had rosters entry for all days, that does not seem to be the case in the submitted test versions so some days could be missing rosters and you may want to handle that.  So maybe do a check like pd.isna(…) around the columns with json data or whatever you prefer.</p>\n<p>Also atm there are issues with notebooks ending with -<br>\nrpc error: code = Unknown desc = generic::unknown: retry budget exhausted (10 attempts): oneshotUpload: googleapi: Error 503: { \"error\": { \"code\": 503, \"message\": \"Server is unavailable\" } }</p>\n<p>Nothing wrong in the Notebook but gets stuck at the Writing Notebook… and ends up Failed status even though it produces output like submission files which are fine.  Not sure what that is about. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1370978,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "06/30/2021 14:31:38",
          "content": "<p>Thks <a href=\"https://www.kaggle.com/something4kag\" target=\"_blank\">@something4kag</a> . It seems harder than it looks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1372829,
      "author_name": "kaitehtzeng",
      "author_url": "",
      "post_date": "07/02/2021 04:26:36",
      "content": "<p>Just for curiosity. Since we have rosters for all days and the submitted test version has missing rosters, can Kaggle fix this issue from its end? It seems a fixable issue to me</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1369423": "Hi kagglers, i have a submission scoring error on one my kernels. Yet all my test were succesful, i checked the number of lines to be predicted, fixed missing data, reduced memory usage, etc. Therefore i have no idea of what is going on behind the hood. Any help will be appreciated. Here is the [link](https://www.kaggle.com/ulrich07/mlb-fe-pipeline-normal-training) to the notebook.\n\nThanks in advance.\n\n**Partial Solution**: Building on [this notebook](https://www.kaggle.com/mlconsult/1-38-lb-lightgbm-with-target-statistics), i had a partial solution with [this](https://www.kaggle.com/ulrich07/mlb-debug-ann). \n\n**Special thanks** to @assign for pointing out the code block creating this error [here](https://www.kaggle.com/assign/debug-others-mlb-fe-pipeline-normal-training).",
    "1369605": "Is there any progress?\nOr any doubts about them?",
    "1369613": "I still have no idea. It is frustrating, i tried many things but still get \"Submission Scoring Error\". My today 5 submissions are over.",
    "1369659": "one thing I observed was error event before finish **iter_test**",
    "1369664": "ulrich07 sorry to hear about the error. Your notebook is doing quite a lot of processing without much in the way of error handling, which can unfortunately result in the silliest little issue failing the whole notebook.\n\nCheck out the tips at the bottom of the [code competition debugging](https://www.kaggle.com/code-competition-debugging). If I were you, I would (a) start by breaking your notebook into more modular steps that can be more readily tested on the unseen test set (b) wrap your processing functions in error handling that allows them to \"fail safe\" (c) make your goal to get _any_ score from the whole pipeline (rip out anything that isn't mandatory, substitute fancy methods with dumb-but-safe alternatives). Careful about hard-coding assumptions of what the data is allowed to look like (categorical levels, missing values, etc).\n\nGood luck!",
    "1369730": "Yes,  i know it fails before the end of the loop. It looks like it enconters a configuration in the hidden dataset not accounted for in my existing functions. I will follow @wcukierski advice to fix it.",
    "1370185": "in prepare_test, **PART ONE**\n```\n\n    if hid_df[colname].iloc[0] ==  hid_df[colname].iloc[0]:\n        pscores = pd.DataFrame(eval(hid_df[colname].iloc[0]))\n        pscores = pscores.rename(columns=ren_player)\n        pscores[\"index\"] = (pd.to_datetime(pscores[\"gameDate\"]) - START_DATE).dt.days\n        pscores = pscores[pscores_id + pscores_cat + pscores_num]\n```\ninclude bug, others will not make error (I don't check net.predict(Xe) make error or not)\n\nI don't know which line makes error, but will helpful enough",
    "1370370": "Many thks @assign",
    "1370646": "ulrich07 - posted elsewhere there seems to be an issue with Rosters. While train had rosters entry for all days, that does not seem to be the case in the submitted test versions so some days could be missing rosters and you may want to handle that.  So maybe do a check like pd.isna(...) around the columns with json data or whatever you prefer.\n\nAlso atm there are issues with notebooks ending with -\nrpc error: code = Unknown desc = generic::unknown: retry budget exhausted (10 attempts): oneshotUpload: googleapi: Error 503: { \"error\": { \"code\": 503, \"message\": \"Server is unavailable\" } }\n\nNothing wrong in the Notebook but gets stuck at the Writing Notebook... and ends up Failed status even though it produces output like submission files which are fine.  Not sure what that is about.",
    "1370978": "Thks @something4kag . It seems harder than it looks",
    "1372829": "Just for curiosity. Since we have rosters for all days and the submitted test version has missing rosters, can Kaggle fix this issue from its end? It seems a fixable issue to me"
  },
  "source": "meta"
}