{
  "id": 240329,
  "title": "Understand the Competition Better - Demystifying Submission Details",
  "url": "/competitions/siim-covid19-detection/discussion/240329",
  "author_name": "",
  "post_date": "2021-05-19T09:57:34.316778900Z",
  "votes": 139,
  "comment_count": 37,
  "views": 0,
  "content": "<p>Hey Everyone,</p>\n<p>I have seen few discussions where people are confused about the submission file. So, let's take it step by step and understand it.</p>\n<p>Submission file looks like this:</p>\n<blockquote>\n  <p>Id,PredictionString<br>\n  2b95d54e4be65_study,negative 1 0 0 1 1<br>\n  2b95d54e4be66_study,typical 1 0 0 1 1<br>\n  2b95d54e4be67_study,indeterminate 1 0 0 1 1 atypical 1 0 0 1 1<br>\n  2b95d54e4be68_image,none 1 0 0 1 1<br>\n  2b95d54e4be69_image,opacity 0.5 100 100 200 200 opacity 0.7 10 10 20 20</p>\n</blockquote>\n<p>IDs with <strong>_study</strong> are at <strong>study level</strong> and with <strong>_image</strong> are at <strong>image level</strong>.</p>\n<h4>Study-Level Labels</h4>\n<p>So, the studies can have more than one label from the following labels:</p>\n<blockquote>\n  <p><strong>'negative', 'typical', 'indeterminate', 'atypical'</strong></p>\n</blockquote>\n<p>And we need to predict from at least one of these labels for each study in the test set. The format for the <code>PredictionString</code> would be <code>negative 1 0 0 1 1</code> for single label and <code>indeterminate 1 0 0 1 1 atypical 1 0 0 1 1</code> for multi-label.</p>\n<p><code>negative</code> being the label or class ID (one of the four labels), followed by <code>1</code> which is a <code>confidence</code> score and followed by <code>0 0 1 1</code> which is a one-pixel bounding box.</p>\n<p>The bounding box will always be <code>0 0 1 1</code> irrespective of the label because we are using same submission file for both classification and object detection tasks and this format is used so that the evaluation metric (mAP) does not get affected by the classification task.</p>\n<blockquote>\n  <p>In this competition, we are making predictions at both a study (multi-image) and image level.</p>\n</blockquote>\n<p>A study can have multiple images, this is based on the above statement and you can also observe this once you join <code>train_study_level.csv</code> and <code>train_image_level.csv</code>.</p>\n<h4>Image-Level Labels</h4>\n<p>So, the images can have multiple objects in them and we must find the bounding boxes of these objects.</p>\n<p>And the format for <code>PredictionString</code> would be <code>opacity 0.5 100 100 200 200</code> for image with single object and <code>opacity 0.5 100 100 200 200 opacity 0.7 10 10 20 20 etc</code> for image with multiple objects.</p>\n<p><code>opacity</code> being the class ID, followed by <code>1</code> which is a <code>confidence</code> score and followed by <code>100 100 200 200</code> which is a bounding box in the format <code>xmin ymin xmax ymax</code>.</p>\n<p>Suppose, if you predict that there are <strong>NO objects</strong> in the image, then the <code>PredictionString</code> would be <code>none 1 0 0 1 1</code>.</p>\n<p><code>none</code> is the class ID for <strong>No Finding</strong>, followed by <code>1</code> which is a <code>confidence</code> score and followed by <code>0 0 1 1</code> which is a one-pixel bounding box.</p>\n<p>And each image has only one label from <strong>'negative', 'typical', 'indeterminate', 'atypical'</strong>. Thanks <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> for bringing this up.</p>\n<h3>Conclusion</h3>\n<p>In this competition, we are making predictions at both study-level and image-level. So, it is a <strong>multi-label classification</strong> at study-level and <strong>object detection</strong> at image-level.</p>\n<p>Even though I say it's a multi-label classification, the <code>train_study_level.csv</code> has only one label per study. 😄</p>\n<p>You would have to predict at study-level and at image-level for every image. And since a study can have multiple images, the final prediction at study-level is combined prediction at study-level for each image.</p>\n<p>Let's say you have a study with two images and predictions for both images at study-level are <code>negative</code> then the final prediction at study-level will be <code>negative 1 0 0 1 1</code> and you have another study with three images and predictions for two images at study-level are <code>negative</code> and for one of them it's <code>indeterminate</code> then the final prediction at study-level will be <code>negative 1 0 0 1 1 indeterminate 1 0 0 1 1</code>.</p>\n<p>We can consider the confidence score to be 1 always for every image at study-level because if you look at the train data, you will see that each image has only one label from <strong>'negative', 'typical', 'indeterminate', 'atypical'</strong>.</p>\n<p>There are still some things that are not clear like how can the predictions at study-level be multi-label etc. These things are nicely pointed out <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240250#1315782\" target=\"_blank\">here</a> by <a href=\"https://www.kaggle.com/kevinpdesai\" target=\"_blank\">@kevinpdesai</a>. I will update this thread once I get clarity based on the response from the host.</p>\n<p>I tried my best to summarize this, so I hope it makes sense to you and was helpful. Let me know if I conveyed it wrong anywhere.</p>\n<p>Happy Kaggling! :))</p>",
  "messages": [
    {
      "id": "1314661",
      "postDate": "05/19/2021 09:57:34",
      "content": "<p>Hey Everyone,</p>\n<p>I have seen few discussions where people are confused about the submission file. So, let's take it step by step and understand it.</p>\n<p>Submission file looks like this:</p>\n<blockquote>\n  <p>Id,PredictionString<br>\n  2b95d54e4be65_study,negative 1 0 0 1 1<br>\n  2b95d54e4be66_study,typical 1 0 0 1 1<br>\n  2b95d54e4be67_study,indeterminate 1 0 0 1 1 atypical 1 0 0 1 1<br>\n  2b95d54e4be68_image,none 1 0 0 1 1<br>\n  2b95d54e4be69_image,opacity 0.5 100 100 200 200 opacity 0.7 10 10 20 20</p>\n</blockquote>\n<p>IDs with <strong>_study</strong> are at <strong>study level</strong> and with <strong>_image</strong> are at <strong>image level</strong>.</p>\n<h4>Study-Level Labels</h4>\n<p>So, the studies can have more than one label from the following labels:</p>\n<blockquote>\n  <p><strong>'negative', 'typical', 'indeterminate', 'atypical'</strong></p>\n</blockquote>\n<p>And we need to predict from at least one of these labels for each study in the test set. The format for the <code>PredictionString</code> would be <code>negative 1 0 0 1 1</code> for single label and <code>indeterminate 1 0 0 1 1 atypical 1 0 0 1 1</code> for multi-label.</p>\n<p><code>negative</code> being the label or class ID (one of the four labels), followed by <code>1</code> which is a <code>confidence</code> score and followed by <code>0 0 1 1</code> which is a one-pixel bounding box.</p>\n<p>The bounding box will always be <code>0 0 1 1</code> irrespective of the label because we are using same submission file for both classification and object detection tasks and this format is used so that the evaluation metric (mAP) does not get affected by the classification task.</p>\n<blockquote>\n  <p>In this competition, we are making predictions at both a study (multi-image) and image level.</p>\n</blockquote>\n<p>A study can have multiple images, this is based on the above statement and you can also observe this once you join <code>train_study_level.csv</code> and <code>train_image_level.csv</code>.</p>\n<h4>Image-Level Labels</h4>\n<p>So, the images can have multiple objects in them and we must find the bounding boxes of these objects.</p>\n<p>And the format for <code>PredictionString</code> would be <code>opacity 0.5 100 100 200 200</code> for image with single object and <code>opacity 0.5 100 100 200 200 opacity 0.7 10 10 20 20 etc</code> for image with multiple objects.</p>\n<p><code>opacity</code> being the class ID, followed by <code>1</code> which is a <code>confidence</code> score and followed by <code>100 100 200 200</code> which is a bounding box in the format <code>xmin ymin xmax ymax</code>.</p>\n<p>Suppose, if you predict that there are <strong>NO objects</strong> in the image, then the <code>PredictionString</code> would be <code>none 1 0 0 1 1</code>.</p>\n<p><code>none</code> is the class ID for <strong>No Finding</strong>, followed by <code>1</code> which is a <code>confidence</code> score and followed by <code>0 0 1 1</code> which is a one-pixel bounding box.</p>\n<p>And each image has only one label from <strong>'negative', 'typical', 'indeterminate', 'atypical'</strong>. Thanks <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> for bringing this up.</p>\n<h3>Conclusion</h3>\n<p>In this competition, we are making predictions at both study-level and image-level. So, it is a <strong>multi-label classification</strong> at study-level and <strong>object detection</strong> at image-level.</p>\n<p>Even though I say it's a multi-label classification, the <code>train_study_level.csv</code> has only one label per study. 😄</p>\n<p>You would have to predict at study-level and at image-level for every image. And since a study can have multiple images, the final prediction at study-level is combined prediction at study-level for each image.</p>\n<p>Let's say you have a study with two images and predictions for both images at study-level are <code>negative</code> then the final prediction at study-level will be <code>negative 1 0 0 1 1</code> and you have another study with three images and predictions for two images at study-level are <code>negative</code> and for one of them it's <code>indeterminate</code> then the final prediction at study-level will be <code>negative 1 0 0 1 1 indeterminate 1 0 0 1 1</code>.</p>\n<p>We can consider the confidence score to be 1 always for every image at study-level because if you look at the train data, you will see that each image has only one label from <strong>'negative', 'typical', 'indeterminate', 'atypical'</strong>.</p>\n<p>There are still some things that are not clear like how can the predictions at study-level be multi-label etc. These things are nicely pointed out <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240250#1315782\" target=\"_blank\">here</a> by <a href=\"https://www.kaggle.com/kevinpdesai\" target=\"_blank\">@kevinpdesai</a>. I will update this thread once I get clarity based on the response from the host.</p>\n<p>I tried my best to summarize this, so I hope it makes sense to you and was helpful. Let me know if I conveyed it wrong anywhere.</p>\n<p>Happy Kaggling! :))</p>",
      "rawMarkdown": "Hey Everyone,\n\nI have seen few discussions where people are confused about the submission file. So, let's take it step by step and understand it.\n\nSubmission file looks like this:\n> Id,PredictionString\n2b95d54e4be65_study,negative 1 0 0 1 1\n2b95d54e4be66_study,typical 1 0 0 1 1\n2b95d54e4be67_study,indeterminate 1 0 0 1 1 atypical 1 0 0 1 1\n2b95d54e4be68_image,none 1 0 0 1 1\n2b95d54e4be69_image,opacity 0.5 100 100 200 200 opacity 0.7 10 10 20 20\n\nIDs with **_study** are at **study level** and with **_image** are at **image level**.\n\n#### Study-Level Labels\nSo, the studies can have more than one label from the following labels:\n> **'negative', 'typical', 'indeterminate', 'atypical'**\n\nAnd we need to predict from at least one of these labels for each study in the test set. The format for the `PredictionString` would be `negative 1 0 0 1 1` for single label and `indeterminate 1 0 0 1 1 atypical 1 0 0 1 1` for multi-label.\n\n`negative` being the label or class ID (one of the four labels), followed by `1` which is a `confidence` score and followed by `0 0 1 1` which is a one-pixel bounding box.\n\nThe bounding box will always be `0 0 1 1` irrespective of the label because we are using same submission file for both classification and object detection tasks and this format is used so that the evaluation metric (mAP) does not get affected by the classification task.\n\n> In this competition, we are making predictions at both a study (multi-image) and image level.\n\nA study can have multiple images, this is based on the above statement and you can also observe this once you join `train_study_level.csv` and `train_image_level.csv`.\n\n#### Image-Level Labels\nSo, the images can have multiple objects in them and we must find the bounding boxes of these objects.\n\nAnd the format for `PredictionString` would be `opacity 0.5 100 100 200 200` for image with single object and `opacity 0.5 100 100 200 200 opacity 0.7 10 10 20 20 etc` for image with multiple objects.\n\n`opacity` being the class ID, followed by `1` which is a `confidence` score and followed by `100 100 200 200` which is a bounding box in the format `xmin ymin xmax ymax`.\n\nSuppose, if you predict that there are **NO objects** in the image, then the `PredictionString` would be `none 1 0 0 1 1`.\n\n`none` is the class ID for **No Finding**, followed by `1` which is a `confidence` score and followed by `0 0 1 1` which is a one-pixel bounding box.\n\nAnd each image has only one label from **'negative', 'typical', 'indeterminate', 'atypical'**. Thanks @awsaf49 for bringing this up.\n\n### Conclusion\nIn this competition, we are making predictions at both study-level and image-level. So, it is a **multi-label classification** at study-level and **object detection** at image-level.\n\nEven though I say it's a multi-label classification, the `train_study_level.csv` has only one label per study. 😄\n\nYou would have to predict at study-level and at image-level for every image. And since a study can have multiple images, the final prediction at study-level is combined prediction at study-level for each image.\n\nLet's say you have a study with two images and predictions for both images at study-level are `negative` then the final prediction at study-level will be `negative 1 0 0 1 1` and you have another study with three images and predictions for two images at study-level are `negative` and for one of them it's `indeterminate` then the final prediction at study-level will be `negative 1 0 0 1 1 indeterminate 1 0 0 1 1`.\n\nWe can consider the confidence score to be 1 always for every image at study-level because if you look at the train data, you will see that each image has only one label from **'negative', 'typical', 'indeterminate', 'atypical'**.\n\nThere are still some things that are not clear like how can the predictions at study-level be multi-label etc. These things are nicely pointed out [here](https://www.kaggle.com/c/siim-covid19-detection/discussion/240250#1315782) by @kevinpdesai. I will update this thread once I get clarity based on the response from the host.\n\nI tried my best to summarize this, so I hope it makes sense to you and was helpful. Let me know if I conveyed it wrong anywhere.\n\nHappy Kaggling! :))",
      "votes": null
    },
    {
      "id": "1314667",
      "postDate": "05/19/2021 10:02:41",
      "content": "<p>I think you should also mention one thing, <strong>each image has only one label from 'negative', 'typical', 'indeterminate', 'atypical'</strong></p>",
      "rawMarkdown": "I think you should also mention one thing, **each image has only one label from 'negative', 'typical', 'indeterminate', 'atypical'**",
      "votes": null
    },
    {
      "id": "1314674",
      "postDate": "05/19/2021 10:06:28",
      "content": "<p>Thanks for mentioning this, I will add it. :)</p>",
      "rawMarkdown": "Thanks for mentioning this, I will add it. :)",
      "votes": null
    },
    {
      "id": "1314890",
      "postDate": "05/19/2021 12:38:28",
      "content": "<p>But after checking the train_study_level.csv, I didn't see any study level sample have mulitple label. </p>\n<p>And what confused me also is that what is the real meaning if a study level sample have mulitple labels. For example if a study level sample contain 9 images, and its real label is 'typical' and 'indeterminate', what is the medical meaning of this, what is the labels' relationship with images ?</p>",
      "rawMarkdown": "But after checking the train_study_level.csv, I didn't see any study level sample have mulitple label. \n\nAnd what confused me also is that what is the real meaning if a study level sample have mulitple labels. For example if a study level sample contain 9 images, and its real label is 'typical' and 'indeterminate', what is the medical meaning of this, what is the labels' relationship with images ?",
      "votes": null
    },
    {
      "id": "1314957",
      "postDate": "05/19/2021 13:20:44",
      "content": "<p>I may have to update this a bit, this is entirely based on the evaluation page and the initial EDA I did. Will investigate further and get back to you.</p>",
      "rawMarkdown": "I may have to update this a bit, this is entirely based on the evaluation page and the initial EDA I did. Will investigate further and get back to you.",
      "votes": null
    },
    {
      "id": "1315065",
      "postDate": "05/19/2021 14:32:21",
      "content": "<p><a href=\"https://www.kaggle.com/mrxuehb\" target=\"_blank\">@mrxuehb</a>, we can have a prediction for a study which is <code>typical 1 0 0 1 1 indeterminate 1 0 0 1 1</code> because a study can have multiple images and prediction for each image at study-level can be different.</p>\n<p>This is just an intuition and I am not sure if this is appropriate in case of the data we have.</p>",
      "rawMarkdown": "mrxuehb, we can have a prediction for a study which is `typical 1 0 0 1 1 indeterminate 1 0 0 1 1` because a study can have multiple images and prediction for each image at study-level can be different.\n\nThis is just an intuition and I am not sure if this is appropriate in case of the data we have.",
      "votes": null
    },
    {
      "id": "1315131",
      "postDate": "05/19/2021 15:29:29",
      "content": "<p>'negative being the label or class ID (one of the four labels), followed by 1 which is a confidence score and followed by 0 0 1 1 which is a one-pixel bounding box.'</p>\n<p>hi, can I know what does it mean by 'one-pixel bounding box'? I tried google search but couldn't find about it. Thanks!</p>",
      "rawMarkdown": "'negative being the label or class ID (one of the four labels), followed by 1 which is a confidence score and followed by 0 0 1 1 which is a one-pixel bounding box.'\n\nhi, can I know what does it mean by 'one-pixel bounding box'? I tried google search but couldn't find about it. Thanks!",
      "votes": null
    },
    {
      "id": "1315150",
      "postDate": "05/19/2021 15:45:22",
      "content": "<p>Since they are using the same submission file for both classification and object detection tasks, they need to make sure that the evaluation metric (mAP) does not get affected. For this, they need a consistent format, so they are using <code>0 0 1 1</code> as default bounding box for classification task.</p>",
      "rawMarkdown": "Since they are using the same submission file for both classification and object detection tasks, they need to make sure that the evaluation metric (mAP) does not get affected. For this, they need a consistent format, so they are using `0 0 1 1` as default bounding box for classification task.",
      "votes": null
    },
    {
      "id": "1315705",
      "postDate": "05/20/2021 04:45:55",
      "content": "<p>Oh, so the 0 0 1 1 are the same for every study level classification right? we can just set it to the value directly?</p>\n<p>For study case with multiple images, when inferencing do we fit all the images in the same study and then average the results to make a final prediction for the study case?</p>\n<p>Thank you.</p>",
      "rawMarkdown": "Oh, so the 0 0 1 1 are the same for every study level classification right? we can just set it to the value directly?\n\nFor study case with multiple images, when inferencing do we fit all the images in the same study and then average the results to make a final prediction for the study case?\n\nThank you.",
      "votes": null
    },
    {
      "id": "1316147",
      "postDate": "05/20/2021 10:39:17",
      "content": "<p>I have a confussion about this study <strong>2b95d54e4be66_study</strong> </p>\n<blockquote>\n  <p>Id,PredictionString<br>\n  2b95d54e4be65_study,negative 1 0 0 1 1<br>\n  <strong>2b95d54e4be66_study</strong>,typical 1 0 0 1 1<br>\n  <strong>2b95d54e4be66_study</strong>,indeterminate 1 0 0 1 1 atypical 1 0 0 1 1<br>\n  2b95d54e4be68_image,none 1 0 0 1 1<br>\n  2b95d54e4be69_image,opacity 0.5 100 100 200 200 opacity 0.7 10 10 20 20</p>\n</blockquote>\n<p>based on what you mentioned above, we must find one line for each study.<br>\nwhy this study has two lines in the submission?</p>",
      "rawMarkdown": "I have a confussion about this study **2b95d54e4be66_study** \n\n> Id,PredictionString\n2b95d54e4be65_study,negative 1 0 0 1 1\n**2b95d54e4be66_study**,typical 1 0 0 1 1\n**2b95d54e4be66_study**,indeterminate 1 0 0 1 1 atypical 1 0 0 1 1\n2b95d54e4be68_image,none 1 0 0 1 1\n2b95d54e4be69_image,opacity 0.5 100 100 200 200 opacity 0.7 10 10 20 20\n\nbased on what you mentioned above, we must find one line for each study.\nwhy this study has two lines in the submission?",
      "votes": null
    },
    {
      "id": "1316164",
      "postDate": "05/20/2021 10:55:23",
      "content": "<blockquote>\n  <p>But after checking the train_study_level.csv, I didn't see any study level sample have mulitple label.</p>\n</blockquote>\n<p>Yes, the train data did not have this but the test data can, that's what it's mentioned in the evaluation page.</p>\n<p><a href=\"https://www.kaggle.com/mrxuehb\" target=\"_blank\">@mrxuehb</a>, I have updated my initial explanation a bit, see if it makes sense now.</p>",
      "rawMarkdown": "> But after checking the train_study_level.csv, I didn't see any study level sample have mulitple label.\n\nYes, the train data did not have this but the test data can, that's what it's mentioned in the evaluation page.\n\n@mrxuehb, I have updated my initial explanation a bit, see if it makes sense now.",
      "votes": null
    },
    {
      "id": "1316173",
      "postDate": "05/20/2021 11:00:31",
      "content": "<p>It's an error I guess (I just copied from evaluation page), it should have been <code>2b95d54e4be67_study</code> 😄. One study will have only one row in the submission file.</p>",
      "rawMarkdown": "It's an error I guess (I just copied from evaluation page), it should have been `2b95d54e4be67_study` 😄. One study will have only one row in the submission file.",
      "votes": null
    },
    {
      "id": "1316180",
      "postDate": "05/20/2021 11:02:28",
      "content": "<p>Oh, so the 0 0 1 1 are the same for every study level classification right? we can just set it to the value directly? - <strong>Yes</strong></p>\n<p>For study case with multiple images, when inferencing do we fit all the images in the same study and then average the results to make a final prediction for the study case? - <strong>There will be no averaging involved since this is a classification task and the final prediction at study-level is combined prediction at study-level for each image.</strong></p>\n<p><strong>Lets say you have a study with two images and predictions for both images at study-level are <code>negative</code> then the final prediction at study-level will be <code>negative 1 0 0 1 1</code> and you have another study with three images and predictions for two images at study-level are <code>negative</code> and for one of them it's <code>indeterminate</code> then the final prediction at study-level will be <code>negative 1 0 0 1 1 indeterminate 1 0 0 1 1</code>.</strong></p>\n<p><strong>The confidence score will always be 1 for every image because if you look at the train data, you will see that each image has only one label from 'negative', 'typical', 'indeterminate', 'atypical'.</strong></p>",
      "rawMarkdown": "Oh, so the 0 0 1 1 are the same for every study level classification right? we can just set it to the value directly? - **Yes**\n\nFor study case with multiple images, when inferencing do we fit all the images in the same study and then average the results to make a final prediction for the study case? - **There will be no averaging involved since this is a classification task and the final prediction at study-level is combined prediction at study-level for each image.**\n\n**Lets say you have a study with two images and predictions for both images at study-level are `negative` then the final prediction at study-level will be `negative 1 0 0 1 1` and you have another study with three images and predictions for two images at study-level are `negative` and for one of them it's `indeterminate` then the final prediction at study-level will be `negative 1 0 0 1 1 indeterminate 1 0 0 1 1`.**\n\n**The confidence score will always be 1 for every image because if you look at the train data, you will see that each image has only one label from 'negative', 'typical', 'indeterminate', 'atypical'.**",
      "votes": null
    },
    {
      "id": "1316413",
      "postDate": "05/20/2021 14:22:03",
      "content": "<p>Hi, thanks for your reply! It helps me to understand this competition. Thank you!</p>",
      "rawMarkdown": "Hi, thanks for your reply! It helps me to understand this competition. Thank you!",
      "votes": null
    },
    {
      "id": "1316506",
      "postDate": "05/20/2021 15:52:34",
      "content": "<p>Thanks for the catch on the error in the Evaluation page example. We've fixed!</p>",
      "rawMarkdown": "Thanks for the catch on the error in the Evaluation page example. We've fixed!",
      "votes": null
    },
    {
      "id": "1316546",
      "postDate": "05/20/2021 16:31:32",
      "content": "<p>I think so, thanks <a href=\"https://www.kaggle.com/hassiahk\" target=\"_blank\">@hassiahk</a> for the clarification and <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> for fixing the error.</p>",
      "rawMarkdown": "I think so, thanks @hassiahk for the clarification and @juliaelliott for fixing the error.",
      "votes": null
    },
    {
      "id": "1316588",
      "postDate": "05/20/2021 17:33:20",
      "content": "<p>For study-level, how is mAP calculated if all conf == 1 ??<br>\nI guess there's gonna be large penalty if our prediction is wrong.</p>",
      "rawMarkdown": "For study-level, how is mAP calculated if all conf == 1 ??\nI guess there's gonna be large penalty if our prediction is wrong.",
      "votes": null
    },
    {
      "id": "1316675",
      "postDate": "05/20/2021 19:14:15",
      "content": "<p>We can consider conf as 1 since for every train image there can only be one label at study-level and this is a classification task. And the other way would be to just consider the probability as conf. I am waiting for official answer from the host.</p>\n<p>I am not sure how much difference there will be in mAP if you take conf as 1 and conf as probability.</p>",
      "rawMarkdown": "We can consider conf as 1 since for every train image there can only be one label at study-level and this is a classification task. And the other way would be to just consider the probability as conf. I am waiting for official answer from the host.\n\nI am not sure how much difference there will be in mAP if you take conf as 1 and conf as probability.",
      "votes": null
    },
    {
      "id": "1316949",
      "postDate": "05/21/2021 03:51:20",
      "content": "<p>This was very helpful thanks for sharing </p>",
      "rawMarkdown": "This was very helpful thanks for sharing",
      "votes": null
    },
    {
      "id": "1317031",
      "postDate": "05/21/2021 05:36:06",
      "content": "<p><a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> <a href=\"https://www.kaggle.com/hassiahk\" target=\"_blank\">@hassiahk</a> but in the Evaluation page it is given<br>\n<code>2b95d54e4be67_study,indeterminate 1 0 0 1 1 atypical 1 0 0 1 1</code><br>\nthis image had 2 labels, is it a mistake?</p>",
      "rawMarkdown": "awsaf49 @hassiahk but in the Evaluation page it is given\n`2b95d54e4be67_study,indeterminate 1 0 0 1 1 atypical 1 0 0 1 1`\nthis image had 2 labels, is it a mistake?",
      "votes": null
    },
    {
      "id": "1317252",
      "postDate": "05/21/2021 08:53:42",
      "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a>, I don't think it's a mistake because it is clearly mentioned that a study in the test set can have more than one label. But this is not the case with train set. So <a href=\"https://www.kaggle.com/kevinpdesai\" target=\"_blank\">@kevinpdesai</a> has asked in this <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240250#1315782\" target=\"_blank\">thread</a>, how can there be multiple labels for a study and I am waiting for a response from the host. So, it is still not clear for me.</p>\n<p>And <code>2b95d54e4be67_study</code> is not an image, it's a study ID and a study can have multiple images. </p>",
      "rawMarkdown": "mrinath, I don't think it's a mistake because it is clearly mentioned that a study in the test set can have more than one label. But this is not the case with train set. So @kevinpdesai has asked in this [thread](https://www.kaggle.com/c/siim-covid19-detection/discussion/240250#1315782), how can there be multiple labels for a study and I am waiting for a response from the host. So, it is still not clear for me.\n\nAnd `2b95d54e4be67_study` is not an image, it's a study ID and a study can have multiple images.",
      "votes": null
    },
    {
      "id": "1317294",
      "postDate": "05/21/2021 09:32:48",
      "content": "<p>Yes, I understand what you are trying to say.<br>\nOne thing I am not getting is, what do images of the same StudyInstanceUID even represent? They have the same images but some of them don't have boxes.</p>",
      "rawMarkdown": "Yes, I understand what you are trying to say.\nOne thing I am not getting is, what do images of the same StudyInstanceUID even represent? They have the same images but some of them don't have boxes.",
      "votes": null
    },
    {
      "id": "1317530",
      "postDate": "05/21/2021 13:16:25",
      "content": "<p>I'm a beginner Plzz help me out while making with me a team plzzzz. </p>",
      "rawMarkdown": "I'm a beginner Plzz help me out while making with me a team plzzzz.",
      "votes": null
    },
    {
      "id": "1317538",
      "postDate": "05/21/2021 13:25:25",
      "content": "<p>I think there is a discussion post for making teams, u can see there.</p>",
      "rawMarkdown": "I think there is a discussion post for making teams, u can see there.",
      "votes": null
    },
    {
      "id": "1317705",
      "postDate": "05/21/2021 16:06:00",
      "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> I agree with you. It is not clear what does a study represent. I have asked the same and other questions in that <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240250#1315782\" target=\"_blank\">thread</a>. Waiting on the host <a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> to respond to the questions.</p>",
      "rawMarkdown": "mrinath I agree with you. It is not clear what does a study represent. I have asked the same and other questions in that [thread](https://www.kaggle.com/c/siim-covid19-detection/discussion/240250#1315782). Waiting on the host @paras42 to respond to the questions.",
      "votes": null
    },
    {
      "id": "1318843",
      "postDate": "05/22/2021 15:58:49",
      "content": "<p>For the  <strong>image level</strong> features the correct output will be  either <code>opacity,confidence score, xmin ,ymin ,xmax ,ymax</code> or <code>none 1 0 0 1 1.</code> (suppose single box) right?<br>\nno need to predict the 'negative', 'typical', 'indeterminate', 'atypical' classes?</p>",
      "rawMarkdown": "For the  **image level** features the correct output will be  either ` opacity,confidence score, xmin ,ymin ,xmax ,ymax` or `none 1 0 0 1 1.` (suppose single box) right?\nno need to predict the 'negative', 'typical', 'indeterminate', 'atypical' classes?",
      "votes": null
    },
    {
      "id": "1318864",
      "postDate": "05/22/2021 16:17:29",
      "content": "<p>Yes, based on the current explanation, the class label prediction is only for study level. However, as we discussed above, it is not clear as to what a study represents. There may be 1-to-1 correlation with the images. In that case, it will implicitly be an image level prediction. Still waiting on the host <a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> for clarification.<br>\nAs for the image level features, there can be multiple boxes and hence there can be multiple such <code>opacity confidence score xmin ymin xmax ymax</code> values to be reported for a single image. If there are no boxes then you report <code>none 1 0 0 1 1</code>.</p>",
      "rawMarkdown": "Yes, based on the current explanation, the class label prediction is only for study level. However, as we discussed above, it is not clear as to what a study represents. There may be 1-to-1 correlation with the images. In that case, it will implicitly be an image level prediction. Still waiting on the host @paras42 for clarification.\nAs for the image level features, there can be multiple boxes and hence there can be multiple such `opacity confidence score xmin ymin xmax ymax` values to be reported for a single image. If there are no boxes then you report `none 1 0 0 1 1`.",
      "votes": null
    },
    {
      "id": "1322807",
      "postDate": "05/25/2021 17:34:59",
      "content": "<p>Just wanted to comment here, because I was confused on that, in my thread <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/241238\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/241238</a> Phil wrote \"all studies should have one study-level label.\"</p>",
      "rawMarkdown": "Just wanted to comment here, because I was confused on that, in my thread https://www.kaggle.com/c/siim-covid19-detection/discussion/241238 Phil wrote \"all studies should have one study-level label.\"",
      "votes": null
    },
    {
      "id": "1325800",
      "postDate": "05/28/2021 02:43:22",
      "content": "<p>Thank you for the attempt to clarify this =)</p>",
      "rawMarkdown": "Thank you for the attempt to clarify this =)",
      "votes": null
    },
    {
      "id": "1335158",
      "postDate": "06/04/2021 03:29:56",
      "content": "<p>We became aware of the duplicates recently, as well as images that are similar to the first image, but taken at a slightly different position, or using different processing.  These represent a minority of the studies in this dataset (less than 5%). In these cases, there are multiple images in a study. Usually, there are 2 images, but sometimes more. Also, in these cases, the prediction for the study level should be the same as the image level for all images.  For example, let's say a patient was imaged twice in one setting. The annotators gave same image level prediction for each of the images (for example \"typical appearance\").  For the remainder of the studies (&gt;95%), there is only 1 image.  I hope this settles some confusion.</p>",
      "rawMarkdown": "We became aware of the duplicates recently, as well as images that are similar to the first image, but taken at a slightly different position, or using different processing.  These represent a minority of the studies in this dataset (less than 5%). In these cases, there are multiple images in a study. Usually, there are 2 images, but sometimes more. Also, in these cases, the prediction for the study level should be the same as the image level for all images.  For example, let's say a patient was imaged twice in one setting. The annotators gave same image level prediction for each of the images (for example \"typical appearance\").  For the remainder of the studies (>95%), there is only 1 image.  I hope this settles some confusion.",
      "votes": null
    },
    {
      "id": "1335329",
      "postDate": "06/04/2021 06:40:52",
      "content": "<p>Thanks for the clarification :)</p>",
      "rawMarkdown": "Thanks for the clarification :)",
      "votes": null
    },
    {
      "id": "1346441",
      "postDate": "06/12/2021 11:59:57",
      "content": "<p>A study may contain multiple images. But we are given only one class at study level. Should we consider that all images inside that study are under the same class?</p>",
      "rawMarkdown": "A study may contain multiple images. But we are given only one class at study level. Should we consider that all images inside that study are under the same class?",
      "votes": null
    },
    {
      "id": "1358444",
      "postDate": "06/20/2021 13:10:19",
      "content": "<p>Thank you for the explanation :)</p>",
      "rawMarkdown": "Thank you for the explanation :)",
      "votes": null
    },
    {
      "id": "1366809",
      "postDate": "06/27/2021 08:18:55",
      "content": "<p>Hello! Do you know if we can submit only the csv file with the predictions generated locally? Or should I do the predictions on the notebook and do the csv file from those?</p>",
      "rawMarkdown": "Hello! Do you know if we can submit only the csv file with the predictions generated locally? Or should I do the predictions on the notebook and do the csv file from those?",
      "votes": null
    },
    {
      "id": "1391136",
      "postDate": "07/17/2021 10:35:26",
      "content": "<p>May I ask for a question here? In train_image_level.csv, what is the relation between \"id\" and \"StudyInstanceUID\"? and how do we name the \"id\" when we have \"StudyInstanceUID\" in test data? Help me!!</p>",
      "rawMarkdown": "May I ask for a question here? In train_image_level.csv, what is the relation between \"id\" and \"StudyInstanceUID\"? and how do we name the \"id\" when we have \"StudyInstanceUID\" in test data? Help me!!",
      "votes": null
    },
    {
      "id": "1391140",
      "postDate": "07/17/2021 10:36:22",
      "content": "<p><a href=\"https://www.kaggle.com/hassiahk\" target=\"_blank\">@hassiahk</a> </p>",
      "rawMarkdown": "hassiahk",
      "votes": null
    },
    {
      "id": "1399417",
      "postDate": "07/25/2021 09:34:32",
      "content": "<p>Thanks for great explanation :)</p>",
      "rawMarkdown": "Thanks for great explanation :)",
      "votes": null
    },
    {
      "id": "1402000",
      "postDate": "07/27/2021 17:55:34",
      "content": "<p>What do the 'negative', 'typical', 'indeterminate', and 'atypical' labels mean? Which one represents the presence or absence of COVID-19 in the X-Ray?</p>",
      "rawMarkdown": "What do the 'negative', 'typical', 'indeterminate', and 'atypical' labels mean? Which one represents the presence or absence of COVID-19 in the X-Ray?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1314667,
      "author_name": "awsaf49",
      "author_url": "",
      "post_date": "05/19/2021 10:02:41",
      "content": "<p>I think you should also mention one thing, <strong>each image has only one label from 'negative', 'typical', 'indeterminate', 'atypical'</strong></p>",
      "votes": null,
      "replies": [
        {
          "id": 1314674,
          "author_name": "hassiahk",
          "author_url": "",
          "post_date": "05/19/2021 10:06:28",
          "content": "<p>Thanks for mentioning this, I will add it. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1317031,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "05/21/2021 05:36:06",
          "content": "<p><a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> <a href=\"https://www.kaggle.com/hassiahk\" target=\"_blank\">@hassiahk</a> but in the Evaluation page it is given<br>\n<code>2b95d54e4be67_study,indeterminate 1 0 0 1 1 atypical 1 0 0 1 1</code><br>\nthis image had 2 labels, is it a mistake?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1317252,
          "author_name": "hassiahk",
          "author_url": "",
          "post_date": "05/21/2021 08:53:42",
          "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a>, I don't think it's a mistake because it is clearly mentioned that a study in the test set can have more than one label. But this is not the case with train set. So <a href=\"https://www.kaggle.com/kevinpdesai\" target=\"_blank\">@kevinpdesai</a> has asked in this <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240250#1315782\" target=\"_blank\">thread</a>, how can there be multiple labels for a study and I am waiting for a response from the host. So, it is still not clear for me.</p>\n<p>And <code>2b95d54e4be67_study</code> is not an image, it's a study ID and a study can have multiple images. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1317294,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "05/21/2021 09:32:48",
          "content": "<p>Yes, I understand what you are trying to say.<br>\nOne thing I am not getting is, what do images of the same StudyInstanceUID even represent? They have the same images but some of them don't have boxes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1317705,
          "author_name": "kevinpdesai",
          "author_url": "",
          "post_date": "05/21/2021 16:06:00",
          "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> I agree with you. It is not clear what does a study represent. I have asked the same and other questions in that <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240250#1315782\" target=\"_blank\">thread</a>. Waiting on the host <a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> to respond to the questions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1314890,
      "author_name": "mrxuehb",
      "author_url": "",
      "post_date": "05/19/2021 12:38:28",
      "content": "<p>But after checking the train_study_level.csv, I didn't see any study level sample have mulitple label. </p>\n<p>And what confused me also is that what is the real meaning if a study level sample have mulitple labels. For example if a study level sample contain 9 images, and its real label is 'typical' and 'indeterminate', what is the medical meaning of this, what is the labels' relationship with images ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1314957,
          "author_name": "hassiahk",
          "author_url": "",
          "post_date": "05/19/2021 13:20:44",
          "content": "<p>I may have to update this a bit, this is entirely based on the evaluation page and the initial EDA I did. Will investigate further and get back to you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1315065,
          "author_name": "hassiahk",
          "author_url": "",
          "post_date": "05/19/2021 14:32:21",
          "content": "<p><a href=\"https://www.kaggle.com/mrxuehb\" target=\"_blank\">@mrxuehb</a>, we can have a prediction for a study which is <code>typical 1 0 0 1 1 indeterminate 1 0 0 1 1</code> because a study can have multiple images and prediction for each image at study-level can be different.</p>\n<p>This is just an intuition and I am not sure if this is appropriate in case of the data we have.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1316164,
          "author_name": "hassiahk",
          "author_url": "",
          "post_date": "05/20/2021 10:55:23",
          "content": "<blockquote>\n  <p>But after checking the train_study_level.csv, I didn't see any study level sample have mulitple label.</p>\n</blockquote>\n<p>Yes, the train data did not have this but the test data can, that's what it's mentioned in the evaluation page.</p>\n<p><a href=\"https://www.kaggle.com/mrxuehb\" target=\"_blank\">@mrxuehb</a>, I have updated my initial explanation a bit, see if it makes sense now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1322807,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "05/25/2021 17:34:59",
          "content": "<p>Just wanted to comment here, because I was confused on that, in my thread <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/241238\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/241238</a> Phil wrote \"all studies should have one study-level label.\"</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1315131,
      "author_name": "leezhixiong",
      "author_url": "",
      "post_date": "05/19/2021 15:29:29",
      "content": "<p>'negative being the label or class ID (one of the four labels), followed by 1 which is a confidence score and followed by 0 0 1 1 which is a one-pixel bounding box.'</p>\n<p>hi, can I know what does it mean by 'one-pixel bounding box'? I tried google search but couldn't find about it. Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1315150,
          "author_name": "hassiahk",
          "author_url": "",
          "post_date": "05/19/2021 15:45:22",
          "content": "<p>Since they are using the same submission file for both classification and object detection tasks, they need to make sure that the evaluation metric (mAP) does not get affected. For this, they need a consistent format, so they are using <code>0 0 1 1</code> as default bounding box for classification task.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1315705,
          "author_name": "leezhixiong",
          "author_url": "",
          "post_date": "05/20/2021 04:45:55",
          "content": "<p>Oh, so the 0 0 1 1 are the same for every study level classification right? we can just set it to the value directly?</p>\n<p>For study case with multiple images, when inferencing do we fit all the images in the same study and then average the results to make a final prediction for the study case?</p>\n<p>Thank you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1316180,
          "author_name": "hassiahk",
          "author_url": "",
          "post_date": "05/20/2021 11:02:28",
          "content": "<p>Oh, so the 0 0 1 1 are the same for every study level classification right? we can just set it to the value directly? - <strong>Yes</strong></p>\n<p>For study case with multiple images, when inferencing do we fit all the images in the same study and then average the results to make a final prediction for the study case? - <strong>There will be no averaging involved since this is a classification task and the final prediction at study-level is combined prediction at study-level for each image.</strong></p>\n<p><strong>Lets say you have a study with two images and predictions for both images at study-level are <code>negative</code> then the final prediction at study-level will be <code>negative 1 0 0 1 1</code> and you have another study with three images and predictions for two images at study-level are <code>negative</code> and for one of them it's <code>indeterminate</code> then the final prediction at study-level will be <code>negative 1 0 0 1 1 indeterminate 1 0 0 1 1</code>.</strong></p>\n<p><strong>The confidence score will always be 1 for every image because if you look at the train data, you will see that each image has only one label from 'negative', 'typical', 'indeterminate', 'atypical'.</strong></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1316413,
          "author_name": "leezhixiong",
          "author_url": "",
          "post_date": "05/20/2021 14:22:03",
          "content": "<p>Hi, thanks for your reply! It helps me to understand this competition. Thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1316588,
          "author_name": "drtausamaru",
          "author_url": "",
          "post_date": "05/20/2021 17:33:20",
          "content": "<p>For study-level, how is mAP calculated if all conf == 1 ??<br>\nI guess there's gonna be large penalty if our prediction is wrong.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1316675,
          "author_name": "hassiahk",
          "author_url": "",
          "post_date": "05/20/2021 19:14:15",
          "content": "<p>We can consider conf as 1 since for every train image there can only be one label at study-level and this is a classification task. And the other way would be to just consider the probability as conf. I am waiting for official answer from the host.</p>\n<p>I am not sure how much difference there will be in mAP if you take conf as 1 and conf as probability.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1316147,
      "author_name": "ahmed3991",
      "author_url": "",
      "post_date": "05/20/2021 10:39:17",
      "content": "<p>I have a confussion about this study <strong>2b95d54e4be66_study</strong> </p>\n<blockquote>\n  <p>Id,PredictionString<br>\n  2b95d54e4be65_study,negative 1 0 0 1 1<br>\n  <strong>2b95d54e4be66_study</strong>,typical 1 0 0 1 1<br>\n  <strong>2b95d54e4be66_study</strong>,indeterminate 1 0 0 1 1 atypical 1 0 0 1 1<br>\n  2b95d54e4be68_image,none 1 0 0 1 1<br>\n  2b95d54e4be69_image,opacity 0.5 100 100 200 200 opacity 0.7 10 10 20 20</p>\n</blockquote>\n<p>based on what you mentioned above, we must find one line for each study.<br>\nwhy this study has two lines in the submission?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1316173,
          "author_name": "hassiahk",
          "author_url": "",
          "post_date": "05/20/2021 11:00:31",
          "content": "<p>It's an error I guess (I just copied from evaluation page), it should have been <code>2b95d54e4be67_study</code> 😄. One study will have only one row in the submission file.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1316506,
          "author_name": "juliaelliott",
          "author_url": "",
          "post_date": "05/20/2021 15:52:34",
          "content": "<p>Thanks for the catch on the error in the Evaluation page example. We've fixed!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1316546,
          "author_name": "ahmed3991",
          "author_url": "",
          "post_date": "05/20/2021 16:31:32",
          "content": "<p>I think so, thanks <a href=\"https://www.kaggle.com/hassiahk\" target=\"_blank\">@hassiahk</a> for the clarification and <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> for fixing the error.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1316949,
      "author_name": "saketkattuboina",
      "author_url": "",
      "post_date": "05/21/2021 03:51:20",
      "content": "<p>This was very helpful thanks for sharing </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1317530,
      "author_name": "sagarhm",
      "author_url": "",
      "post_date": "05/21/2021 13:16:25",
      "content": "<p>I'm a beginner Plzz help me out while making with me a team plzzzz. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1317538,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "05/21/2021 13:25:25",
          "content": "<p>I think there is a discussion post for making teams, u can see there.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1318843,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "05/22/2021 15:58:49",
      "content": "<p>For the  <strong>image level</strong> features the correct output will be  either <code>opacity,confidence score, xmin ,ymin ,xmax ,ymax</code> or <code>none 1 0 0 1 1.</code> (suppose single box) right?<br>\nno need to predict the 'negative', 'typical', 'indeterminate', 'atypical' classes?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1318864,
          "author_name": "kevinpdesai",
          "author_url": "",
          "post_date": "05/22/2021 16:17:29",
          "content": "<p>Yes, based on the current explanation, the class label prediction is only for study level. However, as we discussed above, it is not clear as to what a study represents. There may be 1-to-1 correlation with the images. In that case, it will implicitly be an image level prediction. Still waiting on the host <a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> for clarification.<br>\nAs for the image level features, there can be multiple boxes and hence there can be multiple such <code>opacity confidence score xmin ymin xmax ymax</code> values to be reported for a single image. If there are no boxes then you report <code>none 1 0 0 1 1</code>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1325800,
      "author_name": "kmjr95",
      "author_url": "",
      "post_date": "05/28/2021 02:43:22",
      "content": "<p>Thank you for the attempt to clarify this =)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1335158,
      "author_name": "paras42",
      "author_url": "",
      "post_date": "06/04/2021 03:29:56",
      "content": "<p>We became aware of the duplicates recently, as well as images that are similar to the first image, but taken at a slightly different position, or using different processing.  These represent a minority of the studies in this dataset (less than 5%). In these cases, there are multiple images in a study. Usually, there are 2 images, but sometimes more. Also, in these cases, the prediction for the study level should be the same as the image level for all images.  For example, let's say a patient was imaged twice in one setting. The annotators gave same image level prediction for each of the images (for example \"typical appearance\").  For the remainder of the studies (&gt;95%), there is only 1 image.  I hope this settles some confusion.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1335329,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "06/04/2021 06:40:52",
          "content": "<p>Thanks for the clarification :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1346441,
      "author_name": "ajinkyadeshpande39",
      "author_url": "",
      "post_date": "06/12/2021 11:59:57",
      "content": "<p>A study may contain multiple images. But we are given only one class at study level. Should we consider that all images inside that study are under the same class?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1358444,
      "author_name": "manharsharma007",
      "author_url": "",
      "post_date": "06/20/2021 13:10:19",
      "content": "<p>Thank you for the explanation :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1366809,
      "author_name": "ivgona",
      "author_url": "",
      "post_date": "06/27/2021 08:18:55",
      "content": "<p>Hello! Do you know if we can submit only the csv file with the predictions generated locally? Or should I do the predictions on the notebook and do the csv file from those?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1391136,
      "author_name": "jingyuma19980404",
      "author_url": "",
      "post_date": "07/17/2021 10:35:26",
      "content": "<p>May I ask for a question here? In train_image_level.csv, what is the relation between \"id\" and \"StudyInstanceUID\"? and how do we name the \"id\" when we have \"StudyInstanceUID\" in test data? Help me!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1391140,
      "author_name": "jingyuma19980404",
      "author_url": "",
      "post_date": "07/17/2021 10:36:22",
      "content": "<p><a href=\"https://www.kaggle.com/hassiahk\" target=\"_blank\">@hassiahk</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1399417,
      "author_name": "leesuaa",
      "author_url": "",
      "post_date": "07/25/2021 09:34:32",
      "content": "<p>Thanks for great explanation :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1402000,
      "author_name": "chigozirimifebi",
      "author_url": "",
      "post_date": "07/27/2021 17:55:34",
      "content": "<p>What do the 'negative', 'typical', 'indeterminate', and 'atypical' labels mean? Which one represents the presence or absence of COVID-19 in the X-Ray?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1314661": "Hey Everyone,\n\nI have seen few discussions where people are confused about the submission file. So, let's take it step by step and understand it.\n\nSubmission file looks like this:\n> Id,PredictionString\n2b95d54e4be65_study,negative 1 0 0 1 1\n2b95d54e4be66_study,typical 1 0 0 1 1\n2b95d54e4be67_study,indeterminate 1 0 0 1 1 atypical 1 0 0 1 1\n2b95d54e4be68_image,none 1 0 0 1 1\n2b95d54e4be69_image,opacity 0.5 100 100 200 200 opacity 0.7 10 10 20 20\n\nIDs with **_study** are at **study level** and with **_image** are at **image level**.\n\n#### Study-Level Labels\nSo, the studies can have more than one label from the following labels:\n> **'negative', 'typical', 'indeterminate', 'atypical'**\n\nAnd we need to predict from at least one of these labels for each study in the test set. The format for the `PredictionString` would be `negative 1 0 0 1 1` for single label and `indeterminate 1 0 0 1 1 atypical 1 0 0 1 1` for multi-label.\n\n`negative` being the label or class ID (one of the four labels), followed by `1` which is a `confidence` score and followed by `0 0 1 1` which is a one-pixel bounding box.\n\nThe bounding box will always be `0 0 1 1` irrespective of the label because we are using same submission file for both classification and object detection tasks and this format is used so that the evaluation metric (mAP) does not get affected by the classification task.\n\n> In this competition, we are making predictions at both a study (multi-image) and image level.\n\nA study can have multiple images, this is based on the above statement and you can also observe this once you join `train_study_level.csv` and `train_image_level.csv`.\n\n#### Image-Level Labels\nSo, the images can have multiple objects in them and we must find the bounding boxes of these objects.\n\nAnd the format for `PredictionString` would be `opacity 0.5 100 100 200 200` for image with single object and `opacity 0.5 100 100 200 200 opacity 0.7 10 10 20 20 etc` for image with multiple objects.\n\n`opacity` being the class ID, followed by `1` which is a `confidence` score and followed by `100 100 200 200` which is a bounding box in the format `xmin ymin xmax ymax`.\n\nSuppose, if you predict that there are **NO objects** in the image, then the `PredictionString` would be `none 1 0 0 1 1`.\n\n`none` is the class ID for **No Finding**, followed by `1` which is a `confidence` score and followed by `0 0 1 1` which is a one-pixel bounding box.\n\nAnd each image has only one label from **'negative', 'typical', 'indeterminate', 'atypical'**. Thanks @awsaf49 for bringing this up.\n\n### Conclusion\nIn this competition, we are making predictions at both study-level and image-level. So, it is a **multi-label classification** at study-level and **object detection** at image-level.\n\nEven though I say it's a multi-label classification, the `train_study_level.csv` has only one label per study. 😄\n\nYou would have to predict at study-level and at image-level for every image. And since a study can have multiple images, the final prediction at study-level is combined prediction at study-level for each image.\n\nLet's say you have a study with two images and predictions for both images at study-level are `negative` then the final prediction at study-level will be `negative 1 0 0 1 1` and you have another study with three images and predictions for two images at study-level are `negative` and for one of them it's `indeterminate` then the final prediction at study-level will be `negative 1 0 0 1 1 indeterminate 1 0 0 1 1`.\n\nWe can consider the confidence score to be 1 always for every image at study-level because if you look at the train data, you will see that each image has only one label from **'negative', 'typical', 'indeterminate', 'atypical'**.\n\nThere are still some things that are not clear like how can the predictions at study-level be multi-label etc. These things are nicely pointed out [here](https://www.kaggle.com/c/siim-covid19-detection/discussion/240250#1315782) by @kevinpdesai. I will update this thread once I get clarity based on the response from the host.\n\nI tried my best to summarize this, so I hope it makes sense to you and was helpful. Let me know if I conveyed it wrong anywhere.\n\nHappy Kaggling! :))",
    "1314667": "I think you should also mention one thing, **each image has only one label from 'negative', 'typical', 'indeterminate', 'atypical'**",
    "1314674": "Thanks for mentioning this, I will add it. :)",
    "1314890": "But after checking the train_study_level.csv, I didn't see any study level sample have mulitple label. \n\nAnd what confused me also is that what is the real meaning if a study level sample have mulitple labels. For example if a study level sample contain 9 images, and its real label is 'typical' and 'indeterminate', what is the medical meaning of this, what is the labels' relationship with images ?",
    "1314957": "I may have to update this a bit, this is entirely based on the evaluation page and the initial EDA I did. Will investigate further and get back to you.",
    "1315065": "mrxuehb, we can have a prediction for a study which is `typical 1 0 0 1 1 indeterminate 1 0 0 1 1` because a study can have multiple images and prediction for each image at study-level can be different.\n\nThis is just an intuition and I am not sure if this is appropriate in case of the data we have.",
    "1315131": "'negative being the label or class ID (one of the four labels), followed by 1 which is a confidence score and followed by 0 0 1 1 which is a one-pixel bounding box.'\n\nhi, can I know what does it mean by 'one-pixel bounding box'? I tried google search but couldn't find about it. Thanks!",
    "1315150": "Since they are using the same submission file for both classification and object detection tasks, they need to make sure that the evaluation metric (mAP) does not get affected. For this, they need a consistent format, so they are using `0 0 1 1` as default bounding box for classification task.",
    "1315705": "Oh, so the 0 0 1 1 are the same for every study level classification right? we can just set it to the value directly?\n\nFor study case with multiple images, when inferencing do we fit all the images in the same study and then average the results to make a final prediction for the study case?\n\nThank you.",
    "1316147": "I have a confussion about this study **2b95d54e4be66_study** \n\n> Id,PredictionString\n2b95d54e4be65_study,negative 1 0 0 1 1\n**2b95d54e4be66_study**,typical 1 0 0 1 1\n**2b95d54e4be66_study**,indeterminate 1 0 0 1 1 atypical 1 0 0 1 1\n2b95d54e4be68_image,none 1 0 0 1 1\n2b95d54e4be69_image,opacity 0.5 100 100 200 200 opacity 0.7 10 10 20 20\n\nbased on what you mentioned above, we must find one line for each study.\nwhy this study has two lines in the submission?",
    "1316164": "> But after checking the train_study_level.csv, I didn't see any study level sample have mulitple label.\n\nYes, the train data did not have this but the test data can, that's what it's mentioned in the evaluation page.\n\n@mrxuehb, I have updated my initial explanation a bit, see if it makes sense now.",
    "1316173": "It's an error I guess (I just copied from evaluation page), it should have been `2b95d54e4be67_study` 😄. One study will have only one row in the submission file.",
    "1316180": "Oh, so the 0 0 1 1 are the same for every study level classification right? we can just set it to the value directly? - **Yes**\n\nFor study case with multiple images, when inferencing do we fit all the images in the same study and then average the results to make a final prediction for the study case? - **There will be no averaging involved since this is a classification task and the final prediction at study-level is combined prediction at study-level for each image.**\n\n**Lets say you have a study with two images and predictions for both images at study-level are `negative` then the final prediction at study-level will be `negative 1 0 0 1 1` and you have another study with three images and predictions for two images at study-level are `negative` and for one of them it's `indeterminate` then the final prediction at study-level will be `negative 1 0 0 1 1 indeterminate 1 0 0 1 1`.**\n\n**The confidence score will always be 1 for every image because if you look at the train data, you will see that each image has only one label from 'negative', 'typical', 'indeterminate', 'atypical'.**",
    "1316413": "Hi, thanks for your reply! It helps me to understand this competition. Thank you!",
    "1316506": "Thanks for the catch on the error in the Evaluation page example. We've fixed!",
    "1316546": "I think so, thanks @hassiahk for the clarification and @juliaelliott for fixing the error.",
    "1316588": "For study-level, how is mAP calculated if all conf == 1 ??\nI guess there's gonna be large penalty if our prediction is wrong.",
    "1316675": "We can consider conf as 1 since for every train image there can only be one label at study-level and this is a classification task. And the other way would be to just consider the probability as conf. I am waiting for official answer from the host.\n\nI am not sure how much difference there will be in mAP if you take conf as 1 and conf as probability.",
    "1316949": "This was very helpful thanks for sharing",
    "1317031": "awsaf49 @hassiahk but in the Evaluation page it is given\n`2b95d54e4be67_study,indeterminate 1 0 0 1 1 atypical 1 0 0 1 1`\nthis image had 2 labels, is it a mistake?",
    "1317252": "mrinath, I don't think it's a mistake because it is clearly mentioned that a study in the test set can have more than one label. But this is not the case with train set. So @kevinpdesai has asked in this [thread](https://www.kaggle.com/c/siim-covid19-detection/discussion/240250#1315782), how can there be multiple labels for a study and I am waiting for a response from the host. So, it is still not clear for me.\n\nAnd `2b95d54e4be67_study` is not an image, it's a study ID and a study can have multiple images.",
    "1317294": "Yes, I understand what you are trying to say.\nOne thing I am not getting is, what do images of the same StudyInstanceUID even represent? They have the same images but some of them don't have boxes.",
    "1317530": "I'm a beginner Plzz help me out while making with me a team plzzzz.",
    "1317538": "I think there is a discussion post for making teams, u can see there.",
    "1317705": "mrinath I agree with you. It is not clear what does a study represent. I have asked the same and other questions in that [thread](https://www.kaggle.com/c/siim-covid19-detection/discussion/240250#1315782). Waiting on the host @paras42 to respond to the questions.",
    "1318843": "For the  **image level** features the correct output will be  either ` opacity,confidence score, xmin ,ymin ,xmax ,ymax` or `none 1 0 0 1 1.` (suppose single box) right?\nno need to predict the 'negative', 'typical', 'indeterminate', 'atypical' classes?",
    "1318864": "Yes, based on the current explanation, the class label prediction is only for study level. However, as we discussed above, it is not clear as to what a study represents. There may be 1-to-1 correlation with the images. In that case, it will implicitly be an image level prediction. Still waiting on the host @paras42 for clarification.\nAs for the image level features, there can be multiple boxes and hence there can be multiple such `opacity confidence score xmin ymin xmax ymax` values to be reported for a single image. If there are no boxes then you report `none 1 0 0 1 1`.",
    "1322807": "Just wanted to comment here, because I was confused on that, in my thread https://www.kaggle.com/c/siim-covid19-detection/discussion/241238 Phil wrote \"all studies should have one study-level label.\"",
    "1325800": "Thank you for the attempt to clarify this =)",
    "1335158": "We became aware of the duplicates recently, as well as images that are similar to the first image, but taken at a slightly different position, or using different processing.  These represent a minority of the studies in this dataset (less than 5%). In these cases, there are multiple images in a study. Usually, there are 2 images, but sometimes more. Also, in these cases, the prediction for the study level should be the same as the image level for all images.  For example, let's say a patient was imaged twice in one setting. The annotators gave same image level prediction for each of the images (for example \"typical appearance\").  For the remainder of the studies (>95%), there is only 1 image.  I hope this settles some confusion.",
    "1335329": "Thanks for the clarification :)",
    "1346441": "A study may contain multiple images. But we are given only one class at study level. Should we consider that all images inside that study are under the same class?",
    "1358444": "Thank you for the explanation :)",
    "1366809": "Hello! Do you know if we can submit only the csv file with the predictions generated locally? Or should I do the predictions on the notebook and do the csv file from those?",
    "1391136": "May I ask for a question here? In train_image_level.csv, what is the relation between \"id\" and \"StudyInstanceUID\"? and how do we name the \"id\" when we have \"StudyInstanceUID\" in test data? Help me!!",
    "1391140": "hassiahk",
    "1399417": "Thanks for great explanation :)",
    "1402000": "What do the 'negative', 'typical', 'indeterminate', and 'atypical' labels mean? Which one represents the presence or absence of COVID-19 in the X-Ray?"
  },
  "source": "meta"
}