{
  "id": 392641,
  "title": "Submission format not accepted",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/392641",
  "author_name": "",
  "post_date": "2023-03-06T09:37:26.744907300Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi Everyone! <br>\nSo my submissions keep getting rejected with the following error-message:<br>\nSubmission Scoring Error. Your notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected. See more debugging tips</p>\n<p>However, I compared the new format with another submission of mine (which got accepted, so it has the right format) and the new format is identical to the one getting accepted. Any Ideas where my mistake is? Both submussions are of type 'pandas.core.frame.DataFrame' with 54 rows and 2 columns (correct and session_id) and both have all entrys either 1 or 0 of type  'numpy.int64'&gt;</p>\n<p>Does this happen to someone else? Today, I already lost 4 submission-entrys because this</p>",
  "messages": [
    {
      "id": "2170791",
      "postDate": "03/06/2023 09:37:26",
      "content": "<p>Hi Everyone! <br>\nSo my submissions keep getting rejected with the following error-message:<br>\nSubmission Scoring Error. Your notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected. See more debugging tips</p>\n<p>However, I compared the new format with another submission of mine (which got accepted, so it has the right format) and the new format is identical to the one getting accepted. Any Ideas where my mistake is? Both submussions are of type 'pandas.core.frame.DataFrame' with 54 rows and 2 columns (correct and session_id) and both have all entrys either 1 or 0 of type  'numpy.int64'&gt;</p>\n<p>Does this happen to someone else? Today, I already lost 4 submission-entrys because this</p>",
      "rawMarkdown": "Hi Everyone! \nSo my submissions keep getting rejected with the following error-message:\nSubmission Scoring Error. Your notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected. See more debugging tips\n\nHowever, I compared the new format with another submission of mine (which got accepted, so it has the right format) and the new format is identical to the one getting accepted. Any Ideas where my mistake is? Both submussions are of type 'pandas.core.frame.DataFrame' with 54 rows and 2 columns (correct and session_id) and both have all entrys either 1 or 0 of type  'numpy.int64'>\n\nDoes this happen to someone else? Today, I already lost 4 submission-entrys because this",
      "votes": null
    },
    {
      "id": "2173009",
      "postDate": "03/08/2023 02:23:27",
      "content": "<p>I had the same error the other day.  Particularly, I was trying (in <code>pandas</code>) to convert a column with <code>NaN</code>, which can be in only <code>int64</code> or <code>float</code>, to <code>int32</code>. That issue has the following error message</p>\n<pre><code>IntCastingNaNError: Cannot convert non-finite values (NA or inf) to integer\n</code></pre>\n<p>Even though the root cause was actually in my code, not in the submission file, I think because the notebook did not have this error message in commit and only triggered it when rerunning the notebook for scoring, Kaggle falsely recognizes as a <code>Submission Scoring Error</code>.</p>\n<p>To check if the issue is in your code, keep the feature engineering and prediction parts in your code, but don't update the <code>sample_submission.csv</code>. Instead, set the <code>correct</code> column in <code>sample_submission.csv</code> to all 0s and call <code>predict</code>, then submit for scoring. If you still have the <code>Submission Scoring Error</code>, it means the issue is in your code, not the submission file.</p>",
      "rawMarkdown": "I had the same error the other day.  Particularly, I was trying (in `pandas`) to convert a column with `NaN`, which can be in only `int64` or `float`, to `int32`. That issue has the following error message\n```\nIntCastingNaNError: Cannot convert non-finite values (NA or inf) to integer\n```\nEven though the root cause was actually in my code, not in the submission file, I think because the notebook did not have this error message in commit and only triggered it when rerunning the notebook for scoring, Kaggle falsely recognizes as a `Submission Scoring Error`.\n\nTo check if the issue is in your code, keep the feature engineering and prediction parts in your code, but don't update the `sample_submission.csv`. Instead, set the `correct` column in `sample_submission.csv` to all 0s and call `predict`, then submit for scoring. If you still have the `Submission Scoring Error`, it means the issue is in your code, not the submission file.",
      "votes": null
    },
    {
      "id": "2174689",
      "postDate": "03/09/2023 09:55:04",
      "content": "<p>That's a really good idea, thank you for that! FYI, it apparently really is my code. I didn't think that because it shows a complete submission in the output file after running and also no errors. Nevertheless, guess I'll now start looking for the error. But knowing there actually is one makes it more fun hahaha so thanks</p>",
      "rawMarkdown": "That's a really good idea, thank you for that! FYI, it apparently really is my code. I didn't think that because it shows a complete submission in the output file after running and also no errors. Nevertheless, guess I'll now start looking for the error. But knowing there actually is one makes it more fun hahaha so thanks",
      "votes": null
    },
    {
      "id": "2237783",
      "postDate": "04/28/2023 00:51:22",
      "content": "<p>If you still have the Submission Scoring Error, it means the issue is in your code, not the submission file. Does the problem relate to changing another data type?</p>",
      "rawMarkdown": "If you still have the Submission Scoring Error, it means the issue is in your code, not the submission file. Does the problem relate to changing another data type?",
      "votes": null
    },
    {
      "id": "2237807",
      "postDate": "04/28/2023 01:12:13",
      "content": "<p>Hey! You sound like you know how to fix it, could you help me see what's wrong with my simple submission code?</p>\n<p>import numpy as np<br>\nimport pandas as pd<br>\nimport jo_wilder<br>\nenv = jo_wilder.make_env()<br>\niter_test = env.iter_test()<br>\ncounter = 0</p>\n<h1>The API will deliver two dataframes in this specific order,</h1>\n<h1>for every session+level grouping (one group per session for each checkpoint)</h1>\n<p>for (test, sample_submission) in iter_test:<br>\n    if counter == 0:<br>\n        print(sample_submission.head())<br>\n        print(test.head())<br>\n        print(test.shape)</p>\n<pre><code>## users make predictions here using the test data\ntest['correct'] = 1\n\n## env.predict appends the session+level sample_submission to the overall\n## submission\nenv.predict(test)\ncounter += 1\n</code></pre>\n<p>df = pd.read_csv('submission.csv')<br>\ndf.head()</p>",
      "rawMarkdown": "Hey! You sound like you know how to fix it, could you help me see what's wrong with my simple submission code?\n\nimport numpy as np\nimport pandas as pd\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\ncounter = 0\n# The API will deliver two dataframes in this specific order,\n# for every session+level grouping (one group per session for each checkpoint)\nfor (test, sample_submission) in iter_test:\n    if counter == 0:\n        print(sample_submission.head())\n        print(test.head())\n        print(test.shape)\n        \n    ## users make predictions here using the test data\n    test['correct'] = 1\n    \n    ## env.predict appends the session+level sample_submission to the overall\n    ## submission\n    env.predict(test)\n    counter += 1\ndf = pd.read_csv('submission.csv')\ndf.head()",
      "votes": null
    },
    {
      "id": "2304746",
      "postDate": "06/16/2023 07:53:23",
      "content": "<p>hey - i am encountering the same issue. were you able to resolve the problem?</p>",
      "rawMarkdown": "hey - i am encountering the same issue. were you able to resolve the problem?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2173009,
      "author_name": "hoangnguyen719",
      "author_url": "",
      "post_date": "03/08/2023 02:23:27",
      "content": "<p>I had the same error the other day.  Particularly, I was trying (in <code>pandas</code>) to convert a column with <code>NaN</code>, which can be in only <code>int64</code> or <code>float</code>, to <code>int32</code>. That issue has the following error message</p>\n<pre><code>IntCastingNaNError: Cannot convert non-finite values (NA or inf) to integer\n</code></pre>\n<p>Even though the root cause was actually in my code, not in the submission file, I think because the notebook did not have this error message in commit and only triggered it when rerunning the notebook for scoring, Kaggle falsely recognizes as a <code>Submission Scoring Error</code>.</p>\n<p>To check if the issue is in your code, keep the feature engineering and prediction parts in your code, but don't update the <code>sample_submission.csv</code>. Instead, set the <code>correct</code> column in <code>sample_submission.csv</code> to all 0s and call <code>predict</code>, then submit for scoring. If you still have the <code>Submission Scoring Error</code>, it means the issue is in your code, not the submission file.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2174689,
          "author_name": "evolml",
          "author_url": "",
          "post_date": "03/09/2023 09:55:04",
          "content": "<p>That's a really good idea, thank you for that! FYI, it apparently really is my code. I didn't think that because it shows a complete submission in the output file after running and also no errors. Nevertheless, guess I'll now start looking for the error. But knowing there actually is one makes it more fun hahaha so thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2237783,
          "author_name": "yoshitown",
          "author_url": "",
          "post_date": "04/28/2023 00:51:22",
          "content": "<p>If you still have the Submission Scoring Error, it means the issue is in your code, not the submission file. Does the problem relate to changing another data type?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2237807,
              "author_name": "nicholasleong92",
              "author_url": "",
              "post_date": "04/28/2023 01:12:13",
              "content": "<p>Hey! You sound like you know how to fix it, could you help me see what's wrong with my simple submission code?</p>\n<p>import numpy as np<br>\nimport pandas as pd<br>\nimport jo_wilder<br>\nenv = jo_wilder.make_env()<br>\niter_test = env.iter_test()<br>\ncounter = 0</p>\n<h1>The API will deliver two dataframes in this specific order,</h1>\n<h1>for every session+level grouping (one group per session for each checkpoint)</h1>\n<p>for (test, sample_submission) in iter_test:<br>\n    if counter == 0:<br>\n        print(sample_submission.head())<br>\n        print(test.head())<br>\n        print(test.shape)</p>\n<pre><code>## users make predictions here using the test data\ntest['correct'] = 1\n\n## env.predict appends the session+level sample_submission to the overall\n## submission\nenv.predict(test)\ncounter += 1\n</code></pre>\n<p>df = pd.read_csv('submission.csv')<br>\ndf.head()</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2304746,
      "author_name": "vranjes",
      "author_url": "",
      "post_date": "06/16/2023 07:53:23",
      "content": "<p>hey - i am encountering the same issue. were you able to resolve the problem?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2170791": "Hi Everyone! \nSo my submissions keep getting rejected with the following error-message:\nSubmission Scoring Error. Your notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected. See more debugging tips\n\nHowever, I compared the new format with another submission of mine (which got accepted, so it has the right format) and the new format is identical to the one getting accepted. Any Ideas where my mistake is? Both submussions are of type 'pandas.core.frame.DataFrame' with 54 rows and 2 columns (correct and session_id) and both have all entrys either 1 or 0 of type  'numpy.int64'>\n\nDoes this happen to someone else? Today, I already lost 4 submission-entrys because this",
    "2173009": "I had the same error the other day.  Particularly, I was trying (in `pandas`) to convert a column with `NaN`, which can be in only `int64` or `float`, to `int32`. That issue has the following error message\n```\nIntCastingNaNError: Cannot convert non-finite values (NA or inf) to integer\n```\nEven though the root cause was actually in my code, not in the submission file, I think because the notebook did not have this error message in commit and only triggered it when rerunning the notebook for scoring, Kaggle falsely recognizes as a `Submission Scoring Error`.\n\nTo check if the issue is in your code, keep the feature engineering and prediction parts in your code, but don't update the `sample_submission.csv`. Instead, set the `correct` column in `sample_submission.csv` to all 0s and call `predict`, then submit for scoring. If you still have the `Submission Scoring Error`, it means the issue is in your code, not the submission file.",
    "2174689": "That's a really good idea, thank you for that! FYI, it apparently really is my code. I didn't think that because it shows a complete submission in the output file after running and also no errors. Nevertheless, guess I'll now start looking for the error. But knowing there actually is one makes it more fun hahaha so thanks",
    "2237783": "If you still have the Submission Scoring Error, it means the issue is in your code, not the submission file. Does the problem relate to changing another data type?",
    "2237807": "Hey! You sound like you know how to fix it, could you help me see what's wrong with my simple submission code?\n\nimport numpy as np\nimport pandas as pd\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\ncounter = 0\n# The API will deliver two dataframes in this specific order,\n# for every session+level grouping (one group per session for each checkpoint)\nfor (test, sample_submission) in iter_test:\n    if counter == 0:\n        print(sample_submission.head())\n        print(test.head())\n        print(test.shape)\n        \n    ## users make predictions here using the test data\n    test['correct'] = 1\n    \n    ## env.predict appends the session+level sample_submission to the overall\n    ## submission\n    env.predict(test)\n    counter += 1\ndf = pd.read_csv('submission.csv')\ndf.head()",
    "2304746": "hey - i am encountering the same issue. were you able to resolve the problem?"
  },
  "source": "meta"
}