{
  "id": 514655,
  "title": "Scoring error",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/514655",
  "author_name": "",
  "post_date": "2024-06-25T04:03:48.724435800Z",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I can't manage to get my submission score. I know it has the right columns, right number of rows and, the columns sum to 1, all the 'row_id' values are good. <br>\nI need to know if the error could be coming from the fact that the rows must be in a particular order. Also, if it's not the case, what could cause a scoring error ? <br>\nI leave the head of the submission dataframe (training data). <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14858850%2F6abdea3a5b804b6772c0145d7a57995b%2FScreenshot%202024-06-24%20at%2020.55.43.png?generation=1719288201549913&amp;alt=media\"></p>",
  "messages": [
    {
      "id": "2888751",
      "postDate": "06/25/2024 04:03:48",
      "content": "<p>I can't manage to get my submission score. I know it has the right columns, right number of rows and, the columns sum to 1, all the 'row_id' values are good. <br>\nI need to know if the error could be coming from the fact that the rows must be in a particular order. Also, if it's not the case, what could cause a scoring error ? <br>\nI leave the head of the submission dataframe (training data). <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14858850%2F6abdea3a5b804b6772c0145d7a57995b%2FScreenshot%202024-06-24%20at%2020.55.43.png?generation=1719288201549913&amp;alt=media\"></p>",
      "rawMarkdown": "I can't manage to get my submission score. I know it has the right columns, right number of rows and, the columns sum to 1, all the 'row_id' values are good. \nI need to know if the error could be coming from the fact that the rows must be in a particular order. Also, if it's not the case, what could cause a scoring error ? \nI leave the head of the submission dataframe (training data). \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14858850%2F6abdea3a5b804b6772c0145d7a57995b%2FScreenshot%202024-06-24%20at%2020.55.43.png?generation=1719288201549913&alt=media)",
      "votes": null
    },
    {
      "id": "2889431",
      "postDate": "06/25/2024 13:42:49",
      "content": "<p>Are you saving the DataFrame with or without index? The correct way is like this: <code>my_df.to_csv(\"/kaggle/working/submission.csv\", index=False)</code></p>",
      "rawMarkdown": "Are you saving the DataFrame with or without index? The correct way is like this: `my_df.to_csv(\"/kaggle/working/submission.csv\", index=False)`",
      "votes": null
    },
    {
      "id": "2889689",
      "postDate": "06/25/2024 16:14:51",
      "content": "<p>Thank you for your response. Yes, this is the way I save the dataframe.</p>",
      "rawMarkdown": "Thank you for your response. Yes, this is the way I save the dataframe.",
      "votes": null
    },
    {
      "id": "2890284",
      "postDate": "06/26/2024 04:02:07",
      "content": "<p>Here are some things you can check:</p>\n<ol>\n<li>Assert that all <code>row_ids</code> in <code>sample_submission.csv</code> are present in your final dataframe and that there are no additional <code>row_ids</code>. That is, <code>set(ss_df.row_id.to_list())==set(my_df.row_id.to_list())</code> where <code>my_df</code> is your final submission dataframe.</li>\n<li>Make sure <code>row_ids</code> are unique.</li>\n<li>Round the floats, to say 6 decimals, as exponential notations can cause error.</li>\n<li>I don't think <code>row_ids</code> should be ordered in the same as sample submission, but if that is a problem, use this code: <code>my_df= ss_df.merge(my_df, on=\"row_id\", validate=\"1:1:)</code>. Afterwards assert that there are no nan values.</li>\n</ol>\n<p>If the \"Submission Scoring Error\" changes to \"Notebook Throw an Exception\" then you know that some problem exists in your code.</p>\n<p>Also see my <a href=\"https://www.kaggle.com/code/coderrkj/rsna-resnet-starter-notebook\" target=\"_blank\">starter notebook</a> where I used a flag to replace the test data with a sample selected from the train set. The reason for the flag (<code>len(sub) &lt;= 25</code>) so that when submitting it for scoring, it will be false and the actual test data will be used.</p>",
      "rawMarkdown": "Here are some things you can check:\n1. Assert that all `row_ids` in `sample_submission.csv` are present in your final dataframe and that there are no additional `row_ids`. That is, `set(ss_df.row_id.to_list())==set(my_df.row_id.to_list())` where `my_df` is your final submission dataframe.\n1. Make sure `row_ids` are unique.\n1. Round the floats, to say 6 decimals, as exponential notations can cause error.\n1. I don't think `row_ids` should be ordered in the same as sample submission, but if that is a problem, use this code: `my_df= ss_df.merge(my_df, on=\"row_id\", validate=\"1:1:)`. Afterwards assert that there are no nan values.\n\nIf the \"Submission Scoring Error\" changes to \"Notebook Throw an Exception\" then you know that some problem exists in your code.\n\nAlso see my [starter notebook](https://www.kaggle.com/code/coderrkj/rsna-resnet-starter-notebook) where I used a flag to replace the test data with a sample selected from the train set. The reason for the flag (`len(sub) <= 25`) so that when submitting it for scoring, it will be false and the actual test data will be used.",
      "votes": null
    },
    {
      "id": "2890625",
      "postDate": "06/26/2024 07:40:09",
      "content": "<p>Have you solved the problem? I had the same problem this morning</p>",
      "rawMarkdown": "Have you solved the problem? I had the same problem this morning",
      "votes": null
    },
    {
      "id": "2900398",
      "postDate": "07/02/2024 08:51:16",
      "content": "<p>Could you give me some advice for the related topic (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/516357)\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/516357)</a>. Thanks!</p>",
      "rawMarkdown": "Could you give me some advice for the related topic (https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/516357). Thanks!",
      "votes": null
    },
    {
      "id": "2900399",
      "postDate": "07/02/2024 08:52:01",
      "content": "<p>Have you tackled this problem?</p>",
      "rawMarkdown": "Have you tackled this problem?",
      "votes": null
    },
    {
      "id": "2916921",
      "postDate": "07/11/2024 11:03:33",
      "content": "<p>Thanks for your suggestion, however, it would interest you to know that the sample_submission is not part of the training datasets given. Kindly check</p>",
      "rawMarkdown": "Thanks for your suggestion, however, it would interest you to know that the sample_submission is not part of the training datasets given. Kindly check",
      "votes": null
    },
    {
      "id": "2938923",
      "postDate": "07/28/2024 15:35:42",
      "content": "<p>I am facing same issue, did you got any leads ?<br>\nis it necessary that all rows sum up to 1?<br>\nthere may be rounding errors etc. </p>",
      "rawMarkdown": "I am facing same issue, did you got any leads ?\nis it necessary that all rows sum up to 1?\nthere may be rounding errors etc.",
      "votes": null
    },
    {
      "id": "2938991",
      "postDate": "07/28/2024 16:38:19",
      "content": "<p>Summing up to 1 is not strictly necessary but all values need to be non-negative.<br>\nTrying to check all the 4 points I have listed below in one of the comments.</p>",
      "rawMarkdown": "Summing up to 1 is not strictly necessary but all values need to be non-negative.\nTrying to check all the 4 points I have listed below in one of the comments.",
      "votes": null
    },
    {
      "id": "2939936",
      "postDate": "07/29/2024 16:04:15",
      "content": "<p>I have a doubt.<br>\nDoes kaggle update sample_submission.csv file as well for code competitions with hidden test set at the time of scoring?</p>\n<p>If not then asserting that the row_ids set in sample_submission.csv and submission.csv are same does not make sense.</p>",
      "rawMarkdown": "I have a doubt.\nDoes kaggle update sample_submission.csv file as well for code competitions with hidden test set at the time of scoring?\n\nIf not then asserting that the row_ids set in sample_submission.csv and submission.csv are same does not make sense.",
      "votes": null
    },
    {
      "id": "2940038",
      "postDate": "07/29/2024 17:57:29",
      "content": "<p>It does, at least for this competition. See one of the baseline notebooks initially published at the start of this completion: <a href=\"https://www.kaggle.com/code/krabhijeet/rsna2024-baseline-submission\" target=\"_blank\">https://www.kaggle.com/code/krabhijeet/rsna2024-baseline-submission</a></p>\n<p>They are using <code>sample_submission.csv</code> and nothing else to get a public score around 1.</p>",
      "rawMarkdown": "It does, at least for this competition. See one of the baseline notebooks initially published at the start of this completion: https://www.kaggle.com/code/krabhijeet/rsna2024-baseline-submission\n\nThey are using `sample_submission.csv` and nothing else to get a public score around 1.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2889431,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "06/25/2024 13:42:49",
      "content": "<p>Are you saving the DataFrame with or without index? The correct way is like this: <code>my_df.to_csv(\"/kaggle/working/submission.csv\", index=False)</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 2889689,
          "author_name": "juliengenzling",
          "author_url": "",
          "post_date": "06/25/2024 16:14:51",
          "content": "<p>Thank you for your response. Yes, this is the way I save the dataframe.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2890284,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "06/26/2024 04:02:07",
      "content": "<p>Here are some things you can check:</p>\n<ol>\n<li>Assert that all <code>row_ids</code> in <code>sample_submission.csv</code> are present in your final dataframe and that there are no additional <code>row_ids</code>. That is, <code>set(ss_df.row_id.to_list())==set(my_df.row_id.to_list())</code> where <code>my_df</code> is your final submission dataframe.</li>\n<li>Make sure <code>row_ids</code> are unique.</li>\n<li>Round the floats, to say 6 decimals, as exponential notations can cause error.</li>\n<li>I don't think <code>row_ids</code> should be ordered in the same as sample submission, but if that is a problem, use this code: <code>my_df= ss_df.merge(my_df, on=\"row_id\", validate=\"1:1:)</code>. Afterwards assert that there are no nan values.</li>\n</ol>\n<p>If the \"Submission Scoring Error\" changes to \"Notebook Throw an Exception\" then you know that some problem exists in your code.</p>\n<p>Also see my <a href=\"https://www.kaggle.com/code/coderrkj/rsna-resnet-starter-notebook\" target=\"_blank\">starter notebook</a> where I used a flag to replace the test data with a sample selected from the train set. The reason for the flag (<code>len(sub) &lt;= 25</code>) so that when submitting it for scoring, it will be false and the actual test data will be used.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2900398,
          "author_name": "tuanle98",
          "author_url": "",
          "post_date": "07/02/2024 08:51:16",
          "content": "<p>Could you give me some advice for the related topic (<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/516357)\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/516357)</a>. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2916921,
          "author_name": "samsonolutoberu",
          "author_url": "",
          "post_date": "07/11/2024 11:03:33",
          "content": "<p>Thanks for your suggestion, however, it would interest you to know that the sample_submission is not part of the training datasets given. Kindly check</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2890625,
      "author_name": "",
      "author_url": "",
      "post_date": "06/26/2024 07:40:09",
      "content": "<p>Have you solved the problem? I had the same problem this morning</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2900399,
      "author_name": "tuanle98",
      "author_url": "",
      "post_date": "07/02/2024 08:52:01",
      "content": "<p>Have you tackled this problem?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2938923,
      "author_name": "rohitchaudhari25",
      "author_url": "",
      "post_date": "07/28/2024 15:35:42",
      "content": "<p>I am facing same issue, did you got any leads ?<br>\nis it necessary that all rows sum up to 1?<br>\nthere may be rounding errors etc. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2938991,
          "author_name": "coderrkj",
          "author_url": "",
          "post_date": "07/28/2024 16:38:19",
          "content": "<p>Summing up to 1 is not strictly necessary but all values need to be non-negative.<br>\nTrying to check all the 4 points I have listed below in one of the comments.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2939936,
              "author_name": "rohitchaudhari25",
              "author_url": "",
              "post_date": "07/29/2024 16:04:15",
              "content": "<p>I have a doubt.<br>\nDoes kaggle update sample_submission.csv file as well for code competitions with hidden test set at the time of scoring?</p>\n<p>If not then asserting that the row_ids set in sample_submission.csv and submission.csv are same does not make sense.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2940038,
                  "author_name": "coderrkj",
                  "author_url": "",
                  "post_date": "07/29/2024 17:57:29",
                  "content": "<p>It does, at least for this competition. See one of the baseline notebooks initially published at the start of this completion: <a href=\"https://www.kaggle.com/code/krabhijeet/rsna2024-baseline-submission\" target=\"_blank\">https://www.kaggle.com/code/krabhijeet/rsna2024-baseline-submission</a></p>\n<p>They are using <code>sample_submission.csv</code> and nothing else to get a public score around 1.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2888751": "I can't manage to get my submission score. I know it has the right columns, right number of rows and, the columns sum to 1, all the 'row_id' values are good. \nI need to know if the error could be coming from the fact that the rows must be in a particular order. Also, if it's not the case, what could cause a scoring error ? \nI leave the head of the submission dataframe (training data). \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14858850%2F6abdea3a5b804b6772c0145d7a57995b%2FScreenshot%202024-06-24%20at%2020.55.43.png?generation=1719288201549913&alt=media)",
    "2889431": "Are you saving the DataFrame with or without index? The correct way is like this: `my_df.to_csv(\"/kaggle/working/submission.csv\", index=False)`",
    "2889689": "Thank you for your response. Yes, this is the way I save the dataframe.",
    "2890284": "Here are some things you can check:\n1. Assert that all `row_ids` in `sample_submission.csv` are present in your final dataframe and that there are no additional `row_ids`. That is, `set(ss_df.row_id.to_list())==set(my_df.row_id.to_list())` where `my_df` is your final submission dataframe.\n1. Make sure `row_ids` are unique.\n1. Round the floats, to say 6 decimals, as exponential notations can cause error.\n1. I don't think `row_ids` should be ordered in the same as sample submission, but if that is a problem, use this code: `my_df= ss_df.merge(my_df, on=\"row_id\", validate=\"1:1:)`. Afterwards assert that there are no nan values.\n\nIf the \"Submission Scoring Error\" changes to \"Notebook Throw an Exception\" then you know that some problem exists in your code.\n\nAlso see my [starter notebook](https://www.kaggle.com/code/coderrkj/rsna-resnet-starter-notebook) where I used a flag to replace the test data with a sample selected from the train set. The reason for the flag (`len(sub) <= 25`) so that when submitting it for scoring, it will be false and the actual test data will be used.",
    "2890625": "Have you solved the problem? I had the same problem this morning",
    "2900398": "Could you give me some advice for the related topic (https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/516357). Thanks!",
    "2900399": "Have you tackled this problem?",
    "2916921": "Thanks for your suggestion, however, it would interest you to know that the sample_submission is not part of the training datasets given. Kindly check",
    "2938923": "I am facing same issue, did you got any leads ?\nis it necessary that all rows sum up to 1?\nthere may be rounding errors etc.",
    "2938991": "Summing up to 1 is not strictly necessary but all values need to be non-negative.\nTrying to check all the 4 points I have listed below in one of the comments.",
    "2939936": "I have a doubt.\nDoes kaggle update sample_submission.csv file as well for code competitions with hidden test set at the time of scoring?\n\nIf not then asserting that the row_ids set in sample_submission.csv and submission.csv are same does not make sense.",
    "2940038": "It does, at least for this competition. See one of the baseline notebooks initially published at the start of this completion: https://www.kaggle.com/code/krabhijeet/rsna2024-baseline-submission\n\nThey are using `sample_submission.csv` and nothing else to get a public score around 1."
  },
  "source": "meta"
}