{
  "id": 513789,
  "title": "Question - testing set number of subjects",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/513789",
  "author_name": "",
  "post_date": "2024-06-21T14:34:52.553312300Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>We noticed that we had a single study_id in the test dataset, so a single prediction to make (with 75 different labels). As of our understanding, this share of the data is supposed to be a 36% of the full dataset according to general information provided. From that, we deduce that there would be 3 (or 9 if one patient has 25 predictions) different patients in the full testing set, which looks suspicious considering the amount of data in the training set. </p>\n<p>Is it possible to clarify and/or get some clarifications about that aspect? </p>\n<p>Thanks in advance!</p>",
  "messages": [
    {
      "id": "2882774",
      "postDate": "06/21/2024 14:34:52",
      "content": "<p>We noticed that we had a single study_id in the test dataset, so a single prediction to make (with 75 different labels). As of our understanding, this share of the data is supposed to be a 36% of the full dataset according to general information provided. From that, we deduce that there would be 3 (or 9 if one patient has 25 predictions) different patients in the full testing set, which looks suspicious considering the amount of data in the training set. </p>\n<p>Is it possible to clarify and/or get some clarifications about that aspect? </p>\n<p>Thanks in advance!</p>",
      "rawMarkdown": "We noticed that we had a single study_id in the test dataset, so a single prediction to make (with 75 different labels). As of our understanding, this share of the data is supposed to be a 36% of the full dataset according to general information provided. From that, we deduce that there would be 3 (or 9 if one patient has 25 predictions) different patients in the full testing set, which looks suspicious considering the amount of data in the training set. \n\nIs it possible to clarify and/or get some clarifications about that aspect? \n\nThanks in advance!",
      "votes": null
    },
    {
      "id": "2882778",
      "postDate": "06/21/2024 14:41:40",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/maxradx\" target=\"_blank\">@maxradx</a>, what you see in the <code>test_images</code> folder is just a sample. Once you make a submission, your code will run against the true <code>test_images</code> folder. This is done so that the <code>test_images</code> directory is not leaked to participants. </p>",
      "rawMarkdown": "Hi @maxradx, what you see in the `test_images` folder is just a sample. Once you make a submission, your code will run against the true `test_images` folder. This is done so that the `test_images` directory is not leaked to participants.",
      "votes": null
    },
    {
      "id": "2882903",
      "postDate": "06/21/2024 15:59:02",
      "content": "<p>In addition, I want to mention that the 38% (in the leaderboard) is referring to percentage of the actual hidden test cases (not the one example in the sample given) that is used in computation of the score displayed in the public leader board. The actual private leaderboard (which determines prizes) will use the remaining 62%.</p>",
      "rawMarkdown": "In addition, I want to mention that the 38% (in the leaderboard) is referring to percentage of the actual hidden test cases (not the one example in the sample given) that is used in computation of the score displayed in the public leader board. The actual private leaderboard (which determines prizes) will use the remaining 62%.",
      "votes": null
    },
    {
      "id": "2882966",
      "postDate": "06/21/2024 17:09:26",
      "content": "<blockquote>\n  <p>In addition, I want to mention that the 38% (in the leaderboard) is referring to percentage of the actual hidden test cases (not the one example in the sample given) that is used in computation of the score displayed in the public leader board. The actual private leaderboard (which determines prizes) will use the remaining 62%.</p>\n</blockquote>\n<p>Great, this is exactly the information we were missing. Thank you for clarifying!</p>",
      "rawMarkdown": "> In addition, I want to mention that the 38% (in the leaderboard) is referring to percentage of the actual hidden test cases (not the one example in the sample given) that is used in computation of the score displayed in the public leader board. The actual private leaderboard (which determines prizes) will use the remaining 62%.\n\nGreat, this is exactly the information we were missing. Thank you for clarifying!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2882778,
      "author_name": "brendanartley",
      "author_url": "",
      "post_date": "06/21/2024 14:41:40",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/maxradx\" target=\"_blank\">@maxradx</a>, what you see in the <code>test_images</code> folder is just a sample. Once you make a submission, your code will run against the true <code>test_images</code> folder. This is done so that the <code>test_images</code> directory is not leaked to participants. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2882903,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "06/21/2024 15:59:02",
      "content": "<p>In addition, I want to mention that the 38% (in the leaderboard) is referring to percentage of the actual hidden test cases (not the one example in the sample given) that is used in computation of the score displayed in the public leader board. The actual private leaderboard (which determines prizes) will use the remaining 62%.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2882966,
      "author_name": "jcohenadad",
      "author_url": "",
      "post_date": "06/21/2024 17:09:26",
      "content": "<blockquote>\n  <p>In addition, I want to mention that the 38% (in the leaderboard) is referring to percentage of the actual hidden test cases (not the one example in the sample given) that is used in computation of the score displayed in the public leader board. The actual private leaderboard (which determines prizes) will use the remaining 62%.</p>\n</blockquote>\n<p>Great, this is exactly the information we were missing. Thank you for clarifying!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2882774": "We noticed that we had a single study_id in the test dataset, so a single prediction to make (with 75 different labels). As of our understanding, this share of the data is supposed to be a 36% of the full dataset according to general information provided. From that, we deduce that there would be 3 (or 9 if one patient has 25 predictions) different patients in the full testing set, which looks suspicious considering the amount of data in the training set. \n\nIs it possible to clarify and/or get some clarifications about that aspect? \n\nThanks in advance!",
    "2882778": "Hi @maxradx, what you see in the `test_images` folder is just a sample. Once you make a submission, your code will run against the true `test_images` folder. This is done so that the `test_images` directory is not leaked to participants.",
    "2882903": "In addition, I want to mention that the 38% (in the leaderboard) is referring to percentage of the actual hidden test cases (not the one example in the sample given) that is used in computation of the score displayed in the public leader board. The actual private leaderboard (which determines prizes) will use the remaining 62%.",
    "2882966": "> In addition, I want to mention that the 38% (in the leaderboard) is referring to percentage of the actual hidden test cases (not the one example in the sample given) that is used in computation of the score displayed in the public leader board. The actual private leaderboard (which determines prizes) will use the remaining 62%.\n\nGreat, this is exactly the information we were missing. Thank you for clarifying!"
  },
  "source": "meta"
}