{
  "id": 222370,
  "title": "Submission Scoring Error",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/222370",
  "author_name": "Mikhail Gurevich",
  "post_date": "2021-02-26T17:04:18.666000",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello!</p>\n<p>I'm getting a submission scoring error even with a sample_submission.csv<br>\nI'm starting to believe that I'm doing something really wrong :-)</p>\n<p>I have following:<br>\n1) Private notebook with HPA dataset and my own public dataset <strong><em>hpa-public-test-submission</em></strong>.<br>\n2) <strong><em>hpa-public-test-submission</em></strong> contains a copy of sample_submission.csv renamed to my_empty_submission.csv (<a href=\"https://www.kaggle.com/mgurevich/hpa-public-test-submission\" target=\"_blank\">https://www.kaggle.com/mgurevich/hpa-public-test-submission</a>)</p>\n<p>I load it with: <br>\n<code>submission = pd.read_csv('/kaggle/input/hpa-public-test-submission/my_empty_submission.csv')</code></p>\n<p>And then save it:<br>\n<code>submission.to_csv('submission.csv', index=False)</code></p>\n<p>Then I get <strong>submission scoring error</strong></p>\n<p>My goal is to prepare a pipeline for submitting my submissions on public test data.</p>\n<p>Do we have any means to know why exactly the submission fails?</p>",
  "messages": [
    {
      "id": 1219791,
      "postDate": "2021-02-27T08:43:55.210Z",
      "content": "<p>Had a similar issue yesterday. The problem: You are only saving predictions for the public test dataset and not for the hidden one. You get a submission scoring error because Kaggle doesn't find most of the image id's it's looking for.</p>\n<p>So, how to get predictions for the hidden test set? Well, the robust option is to implement the whole routine of: 1. segmenting the images in the test dataset and 2. predicting class labels for all cells in the test dataset. The reason for this is that when you submit a notebook for scoring, the public test dataset gets swapped out with the pubic test dataset + the hidden test dataset and your notebook is rerun.</p>\n<p>You can also ignore the hidden test dataset by first reading in the default \"sample_submission.csv\" from the \"/kaggle/input/hpa-…/\" folder (Note: this file get's swapped out when scoring with an appropriate sample_submission.csv for public+hidden test data as well!) and then merge your predictions to it based on image ID. That way you only add your predicitons for the public test set and leave the default for the hidden test set. (I.e. your public LB score is representative, but your private LB score will be 0). </p>",
      "rawMarkdown": "Had a similar issue yesterday. The problem: You are only saving predictions for the public test dataset and not for the hidden one. You get a submission scoring error because Kaggle doesn't find most of the image id's it's looking for.\n\nSo, how to get predictions for the hidden test set? Well, the robust option is to implement the whole routine of: 1. segmenting the images in the test dataset and 2. predicting class labels for all cells in the test dataset. The reason for this is that when you submit a notebook for scoring, the public test dataset gets swapped out with the pubic test dataset + the hidden test dataset and your notebook is rerun.\n\nYou can also ignore the hidden test dataset by first reading in the default \"sample_submission.csv\" from the \"/kaggle/input/hpa-.../\" folder (Note: this file get's swapped out when scoring with an appropriate sample_submission.csv for public+hidden test data as well!) and then merge your predictions to it based on image ID. That way you only add your predicitons for the public test set and leave the default for the hidden test set. (I.e. your public LB score is representative, but your private LB score will be 0). ",
      "votes": 4,
      "replies": [
        {
          "id": 1219836,
          "postDate": "2021-02-27T09:24:03.240Z",
          "content": "<p>Thank you very much! Now it works - I didn't take into consideration that sample_submission.csv was being swapped as well.</p>",
          "rawMarkdown": "Thank you very much! Now it works - I didn't take into consideration that sample_submission.csv was being swapped as well.",
          "votes": 1
        },
        {
          "id": 1235637,
          "postDate": "2021-03-12T10:31:37.850Z",
          "content": "<p>How do got this information? Are there some documentation how a submission exactly works?</p>",
          "rawMarkdown": "How do got this information? Are there some documentation how a submission exactly works?"
        }
      ]
    },
    {
      "id": 1219315,
      "postDate": "2021-02-26T17:04:18.667Z",
      "content": "<p>Hello!</p>\n<p>I'm getting a submission scoring error even with a sample_submission.csv<br>\nI'm starting to believe that I'm doing something really wrong :-)</p>\n<p>I have following:<br>\n1) Private notebook with HPA dataset and my own public dataset <strong><em>hpa-public-test-submission</em></strong>.<br>\n2) <strong><em>hpa-public-test-submission</em></strong> contains a copy of sample_submission.csv renamed to my_empty_submission.csv (<a href=\"https://www.kaggle.com/mgurevich/hpa-public-test-submission\" target=\"_blank\">https://www.kaggle.com/mgurevich/hpa-public-test-submission</a>)</p>\n<p>I load it with: <br>\n<code>submission = pd.read_csv('/kaggle/input/hpa-public-test-submission/my_empty_submission.csv')</code></p>\n<p>And then save it:<br>\n<code>submission.to_csv('submission.csv', index=False)</code></p>\n<p>Then I get <strong>submission scoring error</strong></p>\n<p>My goal is to prepare a pipeline for submitting my submissions on public test data.</p>\n<p>Do we have any means to know why exactly the submission fails?</p>",
      "rawMarkdown": "Hello!\n\nI'm getting a submission scoring error even with a sample_submission.csv\nI'm starting to believe that I'm doing something really wrong :-)\n\nI have following:\n1) Private notebook with HPA dataset and my own public dataset ***hpa-public-test-submission***.\n2) ***hpa-public-test-submission*** contains a copy of sample_submission.csv renamed to my_empty_submission.csv (https://www.kaggle.com/mgurevich/hpa-public-test-submission)\n\nI load it with: \n`submission = pd.read_csv('/kaggle/input/hpa-public-test-submission/my_empty_submission.csv')`\n\nAnd then save it:\n`submission.to_csv('submission.csv', index=False)`\n\nThen I get **submission scoring error**\n\nMy goal is to prepare a pipeline for submitting my submissions on public test data.\n\nDo we have any means to know why exactly the submission fails?",
      "votes": 4
    },
    {
      "id": 1266976,
      "postDate": "2021-04-08T08:06:43.777Z",
      "content": "<p>Did you find the reason for the error?</p>",
      "rawMarkdown": "Did you find the reason for the error?"
    },
    {
      "id": 1219348,
      "postDate": "2021-02-26T18:17:44.923Z",
      "content": "<p>Not sure I understand why you would want to do this - the submission that gets scored cannot be a saved file of the public test predictions.  Only submissions that predict the private test images are scored.</p>",
      "rawMarkdown": "Not sure I understand why you would want to do this - the submission that gets scored cannot be a saved file of the public test predictions.  Only submissions that predict the private test images are scored."
    }
  ],
  "comments": [
    {
      "id": 1219791,
      "author_name": "Manuel",
      "author_url": "",
      "post_date": "2021-02-27T08:43:55.210000",
      "content": "<p>Had a similar issue yesterday. The problem: You are only saving predictions for the public test dataset and not for the hidden one. You get a submission scoring error because Kaggle doesn't find most of the image id's it's looking for.</p>\n<p>So, how to get predictions for the hidden test set? Well, the robust option is to implement the whole routine of: 1. segmenting the images in the test dataset and 2. predicting class labels for all cells in the test dataset. The reason for this is that when you submit a notebook for scoring, the public test dataset gets swapped out with the pubic test dataset + the hidden test dataset and your notebook is rerun.</p>\n<p>You can also ignore the hidden test dataset by first reading in the default \"sample_submission.csv\" from the \"/kaggle/input/hpa-…/\" folder (Note: this file get's swapped out when scoring with an appropriate sample_submission.csv for public+hidden test data as well!) and then merge your predictions to it based on image ID. That way you only add your predicitons for the public test set and leave the default for the hidden test set. (I.e. your public LB score is representative, but your private LB score will be 0). </p>",
      "votes": 4,
      "replies": [
        {
          "id": 1219836,
          "author_name": "Mikhail Gurevich",
          "author_url": "",
          "post_date": "2021-02-27T09:24:03.240000",
          "content": "<p>Thank you very much! Now it works - I didn't take into consideration that sample_submission.csv was being swapped as well.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1235637,
          "author_name": "LucaMTB",
          "author_url": "",
          "post_date": "2021-03-12T10:31:37.850000",
          "content": "<p>How do got this information? Are there some documentation how a submission exactly works?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1266976,
      "author_name": "Zaakcii Ru",
      "author_url": "",
      "post_date": "2021-04-08T08:06:43.777000",
      "content": "<p>Did you find the reason for the error?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1219348,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2021-02-26T18:17:44.923000",
      "content": "<p>Not sure I understand why you would want to do this - the submission that gets scored cannot be a saved file of the public test predictions.  Only submissions that predict the private test images are scored.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1219791": "Had a similar issue yesterday. The problem: You are only saving predictions for the public test dataset and not for the hidden one. You get a submission scoring error because Kaggle doesn't find most of the image id's it's looking for.\n\nSo, how to get predictions for the hidden test set? Well, the robust option is to implement the whole routine of: 1. segmenting the images in the test dataset and 2. predicting class labels for all cells in the test dataset. The reason for this is that when you submit a notebook for scoring, the public test dataset gets swapped out with the pubic test dataset + the hidden test dataset and your notebook is rerun.\n\nYou can also ignore the hidden test dataset by first reading in the default \"sample_submission.csv\" from the \"/kaggle/input/hpa-.../\" folder (Note: this file get's swapped out when scoring with an appropriate sample_submission.csv for public+hidden test data as well!) and then merge your predictions to it based on image ID. That way you only add your predicitons for the public test set and leave the default for the hidden test set. (I.e. your public LB score is representative, but your private LB score will be 0). ",
    "1219315": "Hello!\n\nI'm getting a submission scoring error even with a sample_submission.csv\nI'm starting to believe that I'm doing something really wrong :-)\n\nI have following:\n1) Private notebook with HPA dataset and my own public dataset ***hpa-public-test-submission***.\n2) ***hpa-public-test-submission*** contains a copy of sample_submission.csv renamed to my_empty_submission.csv (https://www.kaggle.com/mgurevich/hpa-public-test-submission)\n\nI load it with: \n`submission = pd.read_csv('/kaggle/input/hpa-public-test-submission/my_empty_submission.csv')`\n\nAnd then save it:\n`submission.to_csv('submission.csv', index=False)`\n\nThen I get **submission scoring error**\n\nMy goal is to prepare a pipeline for submitting my submissions on public test data.\n\nDo we have any means to know why exactly the submission fails?",
    "1266976": "Did you find the reason for the error?",
    "1219348": "Not sure I understand why you would want to do this - the submission that gets scored cannot be a saved file of the public test predictions.  Only submissions that predict the private test images are scored."
  }
}