{
  "id": 275983,
  "title": "question about test data",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/275983",
  "author_name": "",
  "post_date": "2021-10-02T14:22:16.440621Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello,</p>\n<p>I'm new to the kaggle competition and not very clear of the submission dataset.<br>\nCould you answer my questions?<br>\nThanks in advance.</p>\n<ol>\n<li><p>After the deadline, would the system run our kernels on the test dataset again or just use the created submission.csv of the kernel?</p></li>\n<li><p>According to leaderboard, the public score is calculated based on 22% of the test data. Does this mean:<br>\n(a)0.22 * 87=19 scans are used for public scoring or<br>\n(b)there are totally 87/0.22=395 scans and 87 of them are used?</p></li>\n<li><p>If (a) is the case, then does it mean as long as all the 87 values are calculated, the private score would also be better than 0.5 if we don't consider overfit?<br>\nIf (b) is the case, then how shall we get the ids when we create the kernels since the submission_sample.csv only contains 87 ids?</p></li>\n</ol>",
  "messages": [
    {
      "id": "1531996",
      "postDate": "10/02/2021 14:22:16",
      "content": "<p>Hello,</p>\n<p>I'm new to the kaggle competition and not very clear of the submission dataset.<br>\nCould you answer my questions?<br>\nThanks in advance.</p>\n<ol>\n<li><p>After the deadline, would the system run our kernels on the test dataset again or just use the created submission.csv of the kernel?</p></li>\n<li><p>According to leaderboard, the public score is calculated based on 22% of the test data. Does this mean:<br>\n(a)0.22 * 87=19 scans are used for public scoring or<br>\n(b)there are totally 87/0.22=395 scans and 87 of them are used?</p></li>\n<li><p>If (a) is the case, then does it mean as long as all the 87 values are calculated, the private score would also be better than 0.5 if we don't consider overfit?<br>\nIf (b) is the case, then how shall we get the ids when we create the kernels since the submission_sample.csv only contains 87 ids?</p></li>\n</ol>",
      "rawMarkdown": "Hello,\n\nI'm new to the kaggle competition and not very clear of the submission dataset.\nCould you answer my questions?\nThanks in advance.\n\n1. After the deadline, would the system run our kernels on the test dataset again or just use the created submission.csv of the kernel?\n\n2. According to leaderboard, the public score is calculated based on 22% of the test data. Does this mean:\n(a)0.22 * 87=19 scans are used for public scoring or\n(b)there are totally 87/0.22=395 scans and 87 of them are used?\n3. If (a) is the case, then does it mean as long as all the 87 values are calculated, the private score would also be better than 0.5 if we don't consider overfit?\nIf (b) is the case, then how shall we get the ids when we create the kernels since the submission_sample.csv only contains 87 ids?",
      "votes": null
    },
    {
      "id": "1532420",
      "postDate": "10/03/2021 01:12:24",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cvplayer\" target=\"_blank\">@cvplayer</a>. </p>\n<ol>\n<li><p>From what I know, the private scores are already calculated at the same time the public scores are calculated when you submit the kernel. But they are just hidden for us as of the moment.  It would take so much time for them to rerun all our codes at the end of the competition.</p></li>\n<li><p>It's case (b). So that the competitors cannot probe the entire test set.</p></li>\n<li><p>When you submit the kernel, the Kaggle system automatically updates the submission_sample.csv to include all the 395 ids. Basically, the entire <code>rsna-miccai</code> directory is updated to include all the 395 ids. That is why, your kernel should be able to <strong>dynamically</strong> read new ids and new images and process them. </p></li>\n</ol>",
      "rawMarkdown": "Hi @cvplayer. \n\n1. From what I know, the private scores are already calculated at the same time the public scores are calculated when you submit the kernel. But they are just hidden for us as of the moment.  It would take so much time for them to rerun all our codes at the end of the competition.\n\n2. It's case (b). So that the competitors cannot probe the entire test set.\n\n3. When you submit the kernel, the Kaggle system automatically updates the submission_sample.csv to include all the 395 ids. Basically, the entire `rsna-miccai` directory is updated to include all the 395 ids. That is why, your kernel should be able to **dynamically** read new ids and new images and process them.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1532420,
      "author_name": "rickandjoe",
      "author_url": "",
      "post_date": "10/03/2021 01:12:24",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cvplayer\" target=\"_blank\">@cvplayer</a>. </p>\n<ol>\n<li><p>From what I know, the private scores are already calculated at the same time the public scores are calculated when you submit the kernel. But they are just hidden for us as of the moment.  It would take so much time for them to rerun all our codes at the end of the competition.</p></li>\n<li><p>It's case (b). So that the competitors cannot probe the entire test set.</p></li>\n<li><p>When you submit the kernel, the Kaggle system automatically updates the submission_sample.csv to include all the 395 ids. Basically, the entire <code>rsna-miccai</code> directory is updated to include all the 395 ids. That is why, your kernel should be able to <strong>dynamically</strong> read new ids and new images and process them. </p></li>\n</ol>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1531996": "Hello,\n\nI'm new to the kaggle competition and not very clear of the submission dataset.\nCould you answer my questions?\nThanks in advance.\n\n1. After the deadline, would the system run our kernels on the test dataset again or just use the created submission.csv of the kernel?\n\n2. According to leaderboard, the public score is calculated based on 22% of the test data. Does this mean:\n(a)0.22 * 87=19 scans are used for public scoring or\n(b)there are totally 87/0.22=395 scans and 87 of them are used?\n3. If (a) is the case, then does it mean as long as all the 87 values are calculated, the private score would also be better than 0.5 if we don't consider overfit?\nIf (b) is the case, then how shall we get the ids when we create the kernels since the submission_sample.csv only contains 87 ids?",
    "1532420": "Hi @cvplayer. \n\n1. From what I know, the private scores are already calculated at the same time the public scores are calculated when you submit the kernel. But they are just hidden for us as of the moment.  It would take so much time for them to rerun all our codes at the end of the competition.\n\n2. It's case (b). So that the competitors cannot probe the entire test set.\n\n3. When you submit the kernel, the Kaggle system automatically updates the submission_sample.csv to include all the 395 ids. Basically, the entire `rsna-miccai` directory is updated to include all the 395 ids. That is why, your kernel should be able to **dynamically** read new ids and new images and process them."
  },
  "source": "meta"
}