{
  "id": 237419,
  "title": "How does final submission on Wednesday actually work?",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/237419",
  "author_name": "Michael Watson",
  "post_date": "2021-05-08T16:29:12.203000",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I'm new to kaggle and am struggling to understand how final submission actually works. Is it safe to say that if your submission runs on the public leaderboard test set, it will also run on the private test set on Wednesday?</p>\n<p>Most inference / submission notebooks for example, copy the sample_submission.csv file and then append their encoded solution on the back of this. I'm assuming that to determine the final leaderboard new rows appended to the sample_submission.csv file and new images are added to the ../input/hpa-single-cell-image-classification/test  directory. This is such that if you have a successful submission to the public leaderboard everything should run as expected.</p>\n<p>Thanks for your help, like I said I'm a noob and have no clue how this works.</p>",
  "messages": [
    {
      "id": 1298205,
      "postDate": "2021-05-08T16:39:16.363Z",
      "content": "<p>On doomsday, sample_submission.csv and test directory will be replaced. New rows will not be added to the existing csv, instead you will get a completely new dataset. Most inference notebooks publicly released will not work on the private test set because they predict only for the public test data. So, if your submission works on the public test set, it is not guaranteed to work on the private test set. You have to write your final submission notebook such that it takes the images from the test directory and sample_submission.csv and outputs a submission.csv by predicting on the fly. An easy way to think about whether your notebook will run successfully is: If sample_submission.csv and test images are replaced with different data, will your notebook be able to produce predictions in the right format? If yes, then you are on the right path.</p>",
      "rawMarkdown": "On doomsday, sample_submission.csv and test directory will be replaced. New rows will not be added to the existing csv, instead you will get a completely new dataset. Most inference notebooks publicly released will not work on the private test set because they predict only for the public test data. So, if your submission works on the public test set, it is not guaranteed to work on the private test set. You have to write your final submission notebook such that it takes the images from the test directory and sample_submission.csv and outputs a submission.csv by predicting on the fly. An easy way to think about whether your notebook will run successfully is: If sample_submission.csv and test images are replaced with different data, will your notebook be able to produce predictions in the right format? If yes, then you are on the right path.",
      "votes": 1,
      "replies": [
        {
          "id": 1298207,
          "postDate": "2021-05-08T16:44:30.373Z",
          "content": "<p>Thanks a lot. Yep, I think we're alright with our current submissions. Kaggle could potentially be slightly clearer about how this works in their documentation but thanks very much for the quick reply. </p>",
          "rawMarkdown": "Thanks a lot. Yep, I think we're alright with our current submissions. Kaggle could potentially be slightly clearer about how this works in their documentation but thanks very much for the quick reply. "
        },
        {
          "id": 1298210,
          "postDate": "2021-05-08T16:47:09.513Z",
          "content": "<p>You're welcome. As far as I know, there is one place that mentions what happens on the last day. Click on \"Submit Predictions\", and you will see:</p>\n<blockquote>\n  <p>In this competition, we will privately re-run your selected Notebook Version with a hidden test set substituted into the competition dataset. We then extract your chosen Output File from the re-run and use that to determine your score.</p>\n</blockquote>",
          "rawMarkdown": "You're welcome. As far as I know, there is one place that mentions what happens on the last day. Click on \"Submit Predictions\", and you will see:\n\n> In this competition, we will privately re-run your selected Notebook Version with a hidden test set substituted into the competition dataset. We then extract your chosen Output File from the re-run and use that to determine your score.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1298196,
      "postDate": "2021-05-08T16:29:12.203Z",
      "content": "<p>I'm new to kaggle and am struggling to understand how final submission actually works. Is it safe to say that if your submission runs on the public leaderboard test set, it will also run on the private test set on Wednesday?</p>\n<p>Most inference / submission notebooks for example, copy the sample_submission.csv file and then append their encoded solution on the back of this. I'm assuming that to determine the final leaderboard new rows appended to the sample_submission.csv file and new images are added to the ../input/hpa-single-cell-image-classification/test  directory. This is such that if you have a successful submission to the public leaderboard everything should run as expected.</p>\n<p>Thanks for your help, like I said I'm a noob and have no clue how this works.</p>",
      "rawMarkdown": "I'm new to kaggle and am struggling to understand how final submission actually works. Is it safe to say that if your submission runs on the public leaderboard test set, it will also run on the private test set on Wednesday?\n\nMost inference / submission notebooks for example, copy the sample_submission.csv file and then append their encoded solution on the back of this. I'm assuming that to determine the final leaderboard new rows appended to the sample_submission.csv file and new images are added to the ../input/hpa-single-cell-image-classification/test  directory. This is such that if you have a successful submission to the public leaderboard everything should run as expected.\n\nThanks for your help, like I said I'm a noob and have no clue how this works.\n  ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1298205,
      "author_name": "novice03",
      "author_url": "",
      "post_date": "2021-05-08T16:39:16.363000",
      "content": "<p>On doomsday, sample_submission.csv and test directory will be replaced. New rows will not be added to the existing csv, instead you will get a completely new dataset. Most inference notebooks publicly released will not work on the private test set because they predict only for the public test data. So, if your submission works on the public test set, it is not guaranteed to work on the private test set. You have to write your final submission notebook such that it takes the images from the test directory and sample_submission.csv and outputs a submission.csv by predicting on the fly. An easy way to think about whether your notebook will run successfully is: If sample_submission.csv and test images are replaced with different data, will your notebook be able to produce predictions in the right format? If yes, then you are on the right path.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1298207,
          "author_name": "Michael Watson",
          "author_url": "",
          "post_date": "2021-05-08T16:44:30.373000",
          "content": "<p>Thanks a lot. Yep, I think we're alright with our current submissions. Kaggle could potentially be slightly clearer about how this works in their documentation but thanks very much for the quick reply. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1298210,
          "author_name": "novice03",
          "author_url": "",
          "post_date": "2021-05-08T16:47:09.513000",
          "content": "<p>You're welcome. As far as I know, there is one place that mentions what happens on the last day. Click on \"Submit Predictions\", and you will see:</p>\n<blockquote>\n  <p>In this competition, we will privately re-run your selected Notebook Version with a hidden test set substituted into the competition dataset. We then extract your chosen Output File from the re-run and use that to determine your score.</p>\n</blockquote>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1298205": "On doomsday, sample_submission.csv and test directory will be replaced. New rows will not be added to the existing csv, instead you will get a completely new dataset. Most inference notebooks publicly released will not work on the private test set because they predict only for the public test data. So, if your submission works on the public test set, it is not guaranteed to work on the private test set. You have to write your final submission notebook such that it takes the images from the test directory and sample_submission.csv and outputs a submission.csv by predicting on the fly. An easy way to think about whether your notebook will run successfully is: If sample_submission.csv and test images are replaced with different data, will your notebook be able to produce predictions in the right format? If yes, then you are on the right path.",
    "1298196": "I'm new to kaggle and am struggling to understand how final submission actually works. Is it safe to say that if your submission runs on the public leaderboard test set, it will also run on the private test set on Wednesday?\n\nMost inference / submission notebooks for example, copy the sample_submission.csv file and then append their encoded solution on the back of this. I'm assuming that to determine the final leaderboard new rows appended to the sample_submission.csv file and new images are added to the ../input/hpa-single-cell-image-classification/test  directory. This is such that if you have a successful submission to the public leaderboard everything should run as expected.\n\nThanks for your help, like I said I'm a noob and have no clue how this works.\n  "
  }
}