{
  "id": 182460,
  "title": "Questions about test data",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/182460",
  "author_name": "",
  "post_date": "2020-09-12T21:53:10.077894700Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi all, I have a few questions about the test data.</p>\n<ol>\n<li><p> <strong>EDIT</strong> This is wrong.</p></li>\n<li><p>I also noticed that the time taken for a submission run suggests that ALL of the 200 test files are in the test dicom directory. Is this true? If so, can I assume that the runtime for a public submission is the same as for a private submission?</p></li>\n</ol>\n<p><strong>EDIT</strong></p>\n<p>Upon receiving a first answer from Ahmed I realised I needed to clarify my question at the expense of brevity. I also made a mistake the first time around. Please forget the above.</p>\n<p>There are 3 ways to think of the test data.<br>\ni) The <strong>sample test set</strong>, which consists of the data available to us in our kernels.<br>\nii) The <strong>public test set</strong>, which consists of 15% of the data in the <strong>hidden test set</strong>.<br>\nii) The <strong>private test set</strong>, which consists of 85% of the data in the <strong>hidden test set</strong>.<br>\n* The <strong>hidden test set</strong> has 200 patients in it.<br>\n** Each test set consists of a <strong>dicom directory</strong> and a <strong>test.csv</strong> file.<br>\nWhile I'm at it I might as well ask, are these sets mutually exclusive?</p>\n<p>So here is my situation and question:</p>\n<p>For a submission I take all files in the <strong>dicom directory</strong> and use them as the patients in my submission.csv. Here's what I observed:</p>\n<ul>\n<li>The time taken for the submission made me conclude I was processing closer to 200 dicom files, instead of the 30 I expected. That is, I concluded that the <strong>public dicom directory</strong> contains ALL dicom folders from the <strong>hidden test set</strong>. That, or it's been accidentally switched with the <strong>private dicom directory</strong>.</li>\n<li>Another indicator to suggest my conclusion is that the first time I tried submitting, I was given a resource exhaustion error. I was caching the preprocessed dicom files on disk in 64-bit format and I would have gone over memory for 200 folders. I changed to 16-bit format and there was no error. 30 dicom folders wouldn't have caused this issue.</li>\n</ul>\n<p>So then my questions here are:</p>\n<ol>\n<li><p>Did I conclude correctly? And if so….</p></li>\n<li><p>How do we then decide which patients to submit? Is it:<br>\na) ONLY ALL patients in the <strong>dicom directory</strong>?<br>\nb) ONLY ALL patients in the <strong>test.csv</strong>?<br>\nc) Union of patients in <strong>test.csv</strong> and <strong>dicom directory</strong>?<br>\nd) Intersection of patients in <strong>test.csv</strong> and <strong>dicom directory</strong>?</p></li>\n<li><p>Does the <strong>private dicom directory</strong> also contain all 200 <strong>hidden test set</strong> patients?</p></li>\n<li><p>Do the <strong>public test.csv</strong> and <strong>private test.csv</strong> also contain all 200 <strong>hidden test set</strong> patients?</p></li>\n<li><p>Are you getting us to evaluate the whole <strong>hidden test set</strong> and then just choosing to show the score for a 15% subset on the public leaderboard? What about the private?</p></li>\n</ol>\n<p>Sorry for the super long question! I just thought it would be better to fully clarify everything in one go.</p>",
  "messages": [
    {
      "id": "1008244",
      "postDate": "09/12/2020 21:53:10",
      "content": "<p>Hi all, I have a few questions about the test data.</p>\n<ol>\n<li><p> <strong>EDIT</strong> This is wrong.</p></li>\n<li><p>I also noticed that the time taken for a submission run suggests that ALL of the 200 test files are in the test dicom directory. Is this true? If so, can I assume that the runtime for a public submission is the same as for a private submission?</p></li>\n</ol>\n<p><strong>EDIT</strong></p>\n<p>Upon receiving a first answer from Ahmed I realised I needed to clarify my question at the expense of brevity. I also made a mistake the first time around. Please forget the above.</p>\n<p>There are 3 ways to think of the test data.<br>\ni) The <strong>sample test set</strong>, which consists of the data available to us in our kernels.<br>\nii) The <strong>public test set</strong>, which consists of 15% of the data in the <strong>hidden test set</strong>.<br>\nii) The <strong>private test set</strong>, which consists of 85% of the data in the <strong>hidden test set</strong>.<br>\n* The <strong>hidden test set</strong> has 200 patients in it.<br>\n** Each test set consists of a <strong>dicom directory</strong> and a <strong>test.csv</strong> file.<br>\nWhile I'm at it I might as well ask, are these sets mutually exclusive?</p>\n<p>So here is my situation and question:</p>\n<p>For a submission I take all files in the <strong>dicom directory</strong> and use them as the patients in my submission.csv. Here's what I observed:</p>\n<ul>\n<li>The time taken for the submission made me conclude I was processing closer to 200 dicom files, instead of the 30 I expected. That is, I concluded that the <strong>public dicom directory</strong> contains ALL dicom folders from the <strong>hidden test set</strong>. That, or it's been accidentally switched with the <strong>private dicom directory</strong>.</li>\n<li>Another indicator to suggest my conclusion is that the first time I tried submitting, I was given a resource exhaustion error. I was caching the preprocessed dicom files on disk in 64-bit format and I would have gone over memory for 200 folders. I changed to 16-bit format and there was no error. 30 dicom folders wouldn't have caused this issue.</li>\n</ul>\n<p>So then my questions here are:</p>\n<ol>\n<li><p>Did I conclude correctly? And if so….</p></li>\n<li><p>How do we then decide which patients to submit? Is it:<br>\na) ONLY ALL patients in the <strong>dicom directory</strong>?<br>\nb) ONLY ALL patients in the <strong>test.csv</strong>?<br>\nc) Union of patients in <strong>test.csv</strong> and <strong>dicom directory</strong>?<br>\nd) Intersection of patients in <strong>test.csv</strong> and <strong>dicom directory</strong>?</p></li>\n<li><p>Does the <strong>private dicom directory</strong> also contain all 200 <strong>hidden test set</strong> patients?</p></li>\n<li><p>Do the <strong>public test.csv</strong> and <strong>private test.csv</strong> also contain all 200 <strong>hidden test set</strong> patients?</p></li>\n<li><p>Are you getting us to evaluate the whole <strong>hidden test set</strong> and then just choosing to show the score for a 15% subset on the public leaderboard? What about the private?</p></li>\n</ol>\n<p>Sorry for the super long question! I just thought it would be better to fully clarify everything in one go.</p>",
      "rawMarkdown": "Hi all, I have a few questions about the test data.\n\n1. ~~I noticed one of the dicom files in the public test directory is not in the test.csv. Is there going to be any discrepancy in the private test data? If so, what takes precedence? Should I submit patient_weeks for patients in test.csv but not in the test dicom directory?~~ **EDIT** This is wrong.\n\n2. I also noticed that the time taken for a submission run suggests that ALL of the 200 test files are in the test dicom directory. Is this true? If so, can I assume that the runtime for a public submission is the same as for a private submission?\n\n**EDIT**\n\nUpon receiving a first answer from Ahmed I realised I needed to clarify my question at the expense of brevity. I also made a mistake the first time around. Please forget the above.\n\nThere are 3 ways to think of the test data.\ni) The **sample test set**, which consists of the data available to us in our kernels.\nii) The **public test set**, which consists of 15% of the data in the **hidden test set**.\nii) The **private test set**, which consists of 85% of the data in the **hidden test set**.\n\\* The **hidden test set** has 200 patients in it.\n\\** Each test set consists of a **dicom directory** and a **test.csv** file.\nWhile I'm at it I might as well ask, are these sets mutually exclusive?\n\nSo here is my situation and question:\n\nFor a submission I take all files in the **dicom directory** and use them as the patients in my submission.csv. Here's what I observed:\n- The time taken for the submission made me conclude I was processing closer to 200 dicom files, instead of the 30 I expected. That is, I concluded that the **public dicom directory** contains ALL dicom folders from the **hidden test set**. That, or it's been accidentally switched with the **private dicom directory**.\n- Another indicator to suggest my conclusion is that the first time I tried submitting, I was given a resource exhaustion error. I was caching the preprocessed dicom files on disk in 64-bit format and I would have gone over memory for 200 folders. I changed to 16-bit format and there was no error. 30 dicom folders wouldn't have caused this issue.\n\nSo then my questions here are:\n\n1. Did I conclude correctly? And if so....\n\n2. How do we then decide which patients to submit? Is it:\n  a) ONLY ALL patients in the **dicom directory**?\n  b) ONLY ALL patients in the **test.csv**?\n  c) Union of patients in **test.csv** and **dicom directory**?\n  d) Intersection of patients in **test.csv** and **dicom directory**?\n\n3. Does the **private dicom directory** also contain all 200 **hidden test set** patients?\n\n4. Do the **public test.csv** and **private test.csv** also contain all 200 **hidden test set** patients?\n\n5. Are you getting us to evaluate the whole **hidden test set** and then just choosing to show the score for a 15% subset on the public leaderboard? What about the private?\n\nSorry for the super long question! I just thought it would be better to fully clarify everything in one go.",
      "votes": null
    },
    {
      "id": "1008257",
      "postDate": "09/12/2020 22:43:39",
      "content": "<p>For the first question, the data in the public test directory is just a placeholder for the real test set (which is hidden from you) to let you know the structure of the hidden data. When you submit your solution, it will be run against the hidden set. I am afraid I am not getting your second question?</p>",
      "rawMarkdown": "For the first question, the data in the public test directory is just a placeholder for the real test set (which is hidden from you) to let you know the structure of the hidden data. When you submit your solution, it will be run against the hidden set. I am afraid I am not getting your second question?",
      "votes": null
    },
    {
      "id": "1008569",
      "postDate": "09/13/2020 08:14:01",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ahmedhshahin\" target=\"_blank\">@ahmedhshahin</a>. Thanks so much for the amazing response time! I'm already aware of what you've stated in your response and I realise my question needs to be written more carefully. I've updated it now. Apologies for the length, but I wanted to be clear.</p>",
      "rawMarkdown": "Hi @ahmedhshahin. Thanks so much for the amazing response time! I'm already aware of what you've stated in your response and I realise my question needs to be written more carefully. I've updated it now. Apologies for the length, but I wanted to be clear.",
      "votes": null
    },
    {
      "id": "1008935",
      "postDate": "09/13/2020 14:10:07",
      "content": "<p>No worries!</p>\n<ul>\n<li>Yes, these sets are mutually exclusive.</li>\n<li>Yes, you are processing the around 200 dicom files as they're in the same folder. But, the score you see on the leaderboard is for 15% only, the score of the rest is kept private. So, we process all the test data but show you the score of a fraction of it.</li>\n<li>The patients in the test.csv and the folder are the same (i.e. their intersection = union)</li>\n<li>question 5: yes, the score of 15% of the patients for the public leaderboard, and the score of 85% of the patients for the private one.</li>\n</ul>\n<p>Does that answer your questions?</p>",
      "rawMarkdown": "No worries!\n- Yes, these sets are mutually exclusive.\n- Yes, you are processing the around 200 dicom files as they're in the same folder. But, the score you see on the leaderboard is for 15% only, the score of the rest is kept private. So, we process all the test data but show you the score of a fraction of it.\n- The patients in the test.csv and the folder are the same (i.e. their intersection = union)\n- question 5: yes, the score of 15% of the patients for the public leaderboard, and the score of 85% of the patients for the private one.\n\nDoes that answer your questions?",
      "votes": null
    },
    {
      "id": "1009096",
      "postDate": "09/13/2020 16:38:15",
      "content": "<p>Yes! And again, thanks for the speed and clarity.</p>",
      "rawMarkdown": "Yes! And again, thanks for the speed and clarity.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1008257,
      "author_name": "ahmedhshahin",
      "author_url": "",
      "post_date": "09/12/2020 22:43:39",
      "content": "<p>For the first question, the data in the public test directory is just a placeholder for the real test set (which is hidden from you) to let you know the structure of the hidden data. When you submit your solution, it will be run against the hidden set. I am afraid I am not getting your second question?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1008569,
          "author_name": "alexandersoare",
          "author_url": "",
          "post_date": "09/13/2020 08:14:01",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ahmedhshahin\" target=\"_blank\">@ahmedhshahin</a>. Thanks so much for the amazing response time! I'm already aware of what you've stated in your response and I realise my question needs to be written more carefully. I've updated it now. Apologies for the length, but I wanted to be clear.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1008935,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "09/13/2020 14:10:07",
          "content": "<p>No worries!</p>\n<ul>\n<li>Yes, these sets are mutually exclusive.</li>\n<li>Yes, you are processing the around 200 dicom files as they're in the same folder. But, the score you see on the leaderboard is for 15% only, the score of the rest is kept private. So, we process all the test data but show you the score of a fraction of it.</li>\n<li>The patients in the test.csv and the folder are the same (i.e. their intersection = union)</li>\n<li>question 5: yes, the score of 15% of the patients for the public leaderboard, and the score of 85% of the patients for the private one.</li>\n</ul>\n<p>Does that answer your questions?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1009096,
          "author_name": "alexandersoare",
          "author_url": "",
          "post_date": "09/13/2020 16:38:15",
          "content": "<p>Yes! And again, thanks for the speed and clarity.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1008244": "Hi all, I have a few questions about the test data.\n\n1. ~~I noticed one of the dicom files in the public test directory is not in the test.csv. Is there going to be any discrepancy in the private test data? If so, what takes precedence? Should I submit patient_weeks for patients in test.csv but not in the test dicom directory?~~ **EDIT** This is wrong.\n\n2. I also noticed that the time taken for a submission run suggests that ALL of the 200 test files are in the test dicom directory. Is this true? If so, can I assume that the runtime for a public submission is the same as for a private submission?\n\n**EDIT**\n\nUpon receiving a first answer from Ahmed I realised I needed to clarify my question at the expense of brevity. I also made a mistake the first time around. Please forget the above.\n\nThere are 3 ways to think of the test data.\ni) The **sample test set**, which consists of the data available to us in our kernels.\nii) The **public test set**, which consists of 15% of the data in the **hidden test set**.\nii) The **private test set**, which consists of 85% of the data in the **hidden test set**.\n\\* The **hidden test set** has 200 patients in it.\n\\** Each test set consists of a **dicom directory** and a **test.csv** file.\nWhile I'm at it I might as well ask, are these sets mutually exclusive?\n\nSo here is my situation and question:\n\nFor a submission I take all files in the **dicom directory** and use them as the patients in my submission.csv. Here's what I observed:\n- The time taken for the submission made me conclude I was processing closer to 200 dicom files, instead of the 30 I expected. That is, I concluded that the **public dicom directory** contains ALL dicom folders from the **hidden test set**. That, or it's been accidentally switched with the **private dicom directory**.\n- Another indicator to suggest my conclusion is that the first time I tried submitting, I was given a resource exhaustion error. I was caching the preprocessed dicom files on disk in 64-bit format and I would have gone over memory for 200 folders. I changed to 16-bit format and there was no error. 30 dicom folders wouldn't have caused this issue.\n\nSo then my questions here are:\n\n1. Did I conclude correctly? And if so....\n\n2. How do we then decide which patients to submit? Is it:\n  a) ONLY ALL patients in the **dicom directory**?\n  b) ONLY ALL patients in the **test.csv**?\n  c) Union of patients in **test.csv** and **dicom directory**?\n  d) Intersection of patients in **test.csv** and **dicom directory**?\n\n3. Does the **private dicom directory** also contain all 200 **hidden test set** patients?\n\n4. Do the **public test.csv** and **private test.csv** also contain all 200 **hidden test set** patients?\n\n5. Are you getting us to evaluate the whole **hidden test set** and then just choosing to show the score for a 15% subset on the public leaderboard? What about the private?\n\nSorry for the super long question! I just thought it would be better to fully clarify everything in one go.",
    "1008257": "For the first question, the data in the public test directory is just a placeholder for the real test set (which is hidden from you) to let you know the structure of the hidden data. When you submit your solution, it will be run against the hidden set. I am afraid I am not getting your second question?",
    "1008569": "Hi @ahmedhshahin. Thanks so much for the amazing response time! I'm already aware of what you've stated in your response and I realise my question needs to be written more carefully. I've updated it now. Apologies for the length, but I wanted to be clear.",
    "1008935": "No worries!\n- Yes, these sets are mutually exclusive.\n- Yes, you are processing the around 200 dicom files as they're in the same folder. But, the score you see on the leaderboard is for 15% only, the score of the rest is kept private. So, we process all the test data but show you the score of a fraction of it.\n- The patients in the test.csv and the folder are the same (i.e. their intersection = union)\n- question 5: yes, the score of 15% of the patients for the public leaderboard, and the score of 85% of the patients for the private one.\n\nDoes that answer your questions?",
    "1009096": "Yes! And again, thanks for the speed and clarity."
  },
  "source": "meta"
}