{
  "id": 520940,
  "title": "Submission Scoring Error",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/520940",
  "author_name": "",
  "post_date": "2024-07-18T06:01:15.022740Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>** help**<br>\nI've tried it many times, but I still get this error😭😭😭<br>\nYour notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected. </p>",
  "messages": [
    {
      "id": "2926907",
      "postDate": "07/18/2024 06:01:15",
      "content": "<p>** help**<br>\nI've tried it many times, but I still get this error😭😭😭<br>\nYour notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected. </p>",
      "rawMarkdown": "** help**\nI've tried it many times, but I still get this error😭😭😭\nYour notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected.",
      "votes": null
    },
    {
      "id": "2926915",
      "postDate": "07/18/2024 06:15:50",
      "content": "<p>Here are some things you can check:</p>\n<ul>\n<li>While saving, don't output index. E.g.: <code>submission.to_csv('/kaggle/working/submission.csv', index=False)</code> where <code>submission</code> is your final submission dataframe.</li>\n<li>Assert that all row_ids in sample_submission.csv are present in your final dataframe and that there are no additional row_ids. That is, <code>set(ss_df.row_id.to_list())==set(submission.row_id.to_list())</code> where <code>submission</code> is your final submission dataframe.</li>\n<li>Make sure row_ids are unique, that is, don't have duplicates.</li>\n<li>Round the floats, to say 6 decimals, as exponential notations can cause error.</li>\n<li>I don't think row_ids should be ordered in the same as sample submission, but if that is a problem, use this code: <code>submission = ss_df.merge(submission, on=\"row_id\", validate=\"1:1\")</code>. Afterwards assert that there are no nan values.</li>\n</ul>\n<p>If the \"Submission Scoring Error\" changes to \"Notebook Throw an Exception\" then you know that some problem exists in your code.</p>\n<p>Also see my <a href=\"https://www.kaggle.com/code/coderrkj/rsna-resnet-starter-notebook\" target=\"_blank\">starter notebook</a> where I used a flag to replace the test data with a sample selected from the train set. The reason for the flag <code>(len(sub) &lt;= 25)</code> so that when submitting it for scoring, it will be false and the actual test data will be used.</p>",
      "rawMarkdown": "Here are some things you can check:\n\n- While saving, don't output index. E.g.: `submission.to_csv('/kaggle/working/submission.csv', index=False)` where `submission` is your final submission dataframe.\n- Assert that all row_ids in sample_submission.csv are present in your final dataframe and that there are no additional row_ids. That is, `set(ss_df.row_id.to_list())==set(submission.row_id.to_list())` where `submission` is your final submission dataframe.\n- Make sure row_ids are unique, that is, don't have duplicates.\n- Round the floats, to say 6 decimals, as exponential notations can cause error.\n- I don't think row_ids should be ordered in the same as sample submission, but if that is a problem, use this code: `submission = ss_df.merge(submission, on=\"row_id\", validate=\"1:1\")`. Afterwards assert that there are no nan values.\n\nIf the \"Submission Scoring Error\" changes to \"Notebook Throw an Exception\" then you know that some problem exists in your code.\n\nAlso see my [starter notebook](https://www.kaggle.com/code/coderrkj/rsna-resnet-starter-notebook) where I used a flag to replace the test data with a sample selected from the train set. The reason for the flag `(len(sub) <= 25)` so that when submitting it for scoring, it will be false and the actual test data will be used.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2926915,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "07/18/2024 06:15:50",
      "content": "<p>Here are some things you can check:</p>\n<ul>\n<li>While saving, don't output index. E.g.: <code>submission.to_csv('/kaggle/working/submission.csv', index=False)</code> where <code>submission</code> is your final submission dataframe.</li>\n<li>Assert that all row_ids in sample_submission.csv are present in your final dataframe and that there are no additional row_ids. That is, <code>set(ss_df.row_id.to_list())==set(submission.row_id.to_list())</code> where <code>submission</code> is your final submission dataframe.</li>\n<li>Make sure row_ids are unique, that is, don't have duplicates.</li>\n<li>Round the floats, to say 6 decimals, as exponential notations can cause error.</li>\n<li>I don't think row_ids should be ordered in the same as sample submission, but if that is a problem, use this code: <code>submission = ss_df.merge(submission, on=\"row_id\", validate=\"1:1\")</code>. Afterwards assert that there are no nan values.</li>\n</ul>\n<p>If the \"Submission Scoring Error\" changes to \"Notebook Throw an Exception\" then you know that some problem exists in your code.</p>\n<p>Also see my <a href=\"https://www.kaggle.com/code/coderrkj/rsna-resnet-starter-notebook\" target=\"_blank\">starter notebook</a> where I used a flag to replace the test data with a sample selected from the train set. The reason for the flag <code>(len(sub) &lt;= 25)</code> so that when submitting it for scoring, it will be false and the actual test data will be used.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2926907": "** help**\nI've tried it many times, but I still get this error😭😭😭\nYour notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected.",
    "2926915": "Here are some things you can check:\n\n- While saving, don't output index. E.g.: `submission.to_csv('/kaggle/working/submission.csv', index=False)` where `submission` is your final submission dataframe.\n- Assert that all row_ids in sample_submission.csv are present in your final dataframe and that there are no additional row_ids. That is, `set(ss_df.row_id.to_list())==set(submission.row_id.to_list())` where `submission` is your final submission dataframe.\n- Make sure row_ids are unique, that is, don't have duplicates.\n- Round the floats, to say 6 decimals, as exponential notations can cause error.\n- I don't think row_ids should be ordered in the same as sample submission, but if that is a problem, use this code: `submission = ss_df.merge(submission, on=\"row_id\", validate=\"1:1\")`. Afterwards assert that there are no nan values.\n\nIf the \"Submission Scoring Error\" changes to \"Notebook Throw an Exception\" then you know that some problem exists in your code.\n\nAlso see my [starter notebook](https://www.kaggle.com/code/coderrkj/rsna-resnet-starter-notebook) where I used a flag to replace the test data with a sample selected from the train set. The reason for the flag `(len(sub) <= 25)` so that when submitting it for scoring, it will be false and the actual test data will be used."
  },
  "source": "meta"
}