{
  "id": 240250,
  "title": "Annotation Grading Methodology",
  "url": "/competitions/siim-covid19-detection/discussion/240250",
  "author_name": "ParasLakhani",
  "post_date": "2021-05-19T03:57:20.503000",
  "votes": 104,
  "comment_count": 41,
  "views": 0,
  "content": "<p>In this challenge, the chest radiographs (CXRs) were categorized using a specific grading schema, based on a published paper:  </p>\n<p>Litmanovich DE, Chung M, Kirkbride RR, Kicska G, Kanne JP. Review of chest radiograph findings of COVID-19 pneumonia and suggested reporting language. Journal of thoracic imaging. 2020 Nov 14;35(6):354-60.</p>\n<p><a href=\"https://journals.lww.com/thoracicimaging/Fulltext/2020/11000/Review_of_Chest_Radiograph_Findings_of_COVID_19.4.aspx\" target=\"_blank\">https://journals.lww.com/thoracicimaging/Fulltext/2020/11000/Review_of_Chest_Radiograph_Findings_of_COVID_19.4.aspx</a></p>\n<p>Per the grading schema, chest radiographs are classified into one of four categories, which are mutually exclusive:</p>\n<p><strong>1. Typical Appearance:</strong> Multifocal bilateral, peripheral opacities with rounded morphology, lower lung–predominant distribution </p>\n<p><strong>2. Indeterminate Appearance:</strong> Absence of typical findings AND unilateral, central or upper lung predominant distribution</p>\n<p><strong>3. Atypical Appearance:</strong>   Pneumothorax, pleural effusion, pulmonary edema, lobar consolidation, solitary lung nodule or mass, diffuse tiny nodules, cavity             </p>\n<p><strong>4. Negative for Pneumonia:</strong>  No lung opacities  </p>\n<p>Bounding boxes were placed on lung opacities, whether typical or indeterminate. Bounding boxes were also placed on some atypical findings including solitary lobar consolidation, nodules/masses, and cavities.  Bounding boxes were not placed on pleural effusions, or pneumothoraces.  No bounding boxes were placed for the negative for pneumonia category.</p>\n<p>In cases of multiple adjacent opacities, we opted for one large bounding box, rather than multiple adjacent smaller boxes, to improve consistency in the labeling.  </p>\n<p>Annotators did have access to the COVID status for each patient, but were asked to adhere to the grading system above irrespective of the status. As such, some patients who were COVID negative still had chest radiographs with  typical appearances. Similarly, some patients who were COVID positive had atypical appearances, or were negative for pneumonia (no lung opacities), because the grading system is based off the chest radiographic findings alone.</p>\n<p>The goal in this challenge is to determine the appropriate category for each radiograph, as well as localize the lung opacities with a bounding box prediction. </p>",
  "messages": [
    {
      "id": 1314244,
      "postDate": "2021-05-19T03:57:20.503Z",
      "content": "<p>In this challenge, the chest radiographs (CXRs) were categorized using a specific grading schema, based on a published paper:  </p>\n<p>Litmanovich DE, Chung M, Kirkbride RR, Kicska G, Kanne JP. Review of chest radiograph findings of COVID-19 pneumonia and suggested reporting language. Journal of thoracic imaging. 2020 Nov 14;35(6):354-60.</p>\n<p><a href=\"https://journals.lww.com/thoracicimaging/Fulltext/2020/11000/Review_of_Chest_Radiograph_Findings_of_COVID_19.4.aspx\" target=\"_blank\">https://journals.lww.com/thoracicimaging/Fulltext/2020/11000/Review_of_Chest_Radiograph_Findings_of_COVID_19.4.aspx</a></p>\n<p>Per the grading schema, chest radiographs are classified into one of four categories, which are mutually exclusive:</p>\n<p><strong>1. Typical Appearance:</strong> Multifocal bilateral, peripheral opacities with rounded morphology, lower lung–predominant distribution </p>\n<p><strong>2. Indeterminate Appearance:</strong> Absence of typical findings AND unilateral, central or upper lung predominant distribution</p>\n<p><strong>3. Atypical Appearance:</strong>   Pneumothorax, pleural effusion, pulmonary edema, lobar consolidation, solitary lung nodule or mass, diffuse tiny nodules, cavity             </p>\n<p><strong>4. Negative for Pneumonia:</strong>  No lung opacities  </p>\n<p>Bounding boxes were placed on lung opacities, whether typical or indeterminate. Bounding boxes were also placed on some atypical findings including solitary lobar consolidation, nodules/masses, and cavities.  Bounding boxes were not placed on pleural effusions, or pneumothoraces.  No bounding boxes were placed for the negative for pneumonia category.</p>\n<p>In cases of multiple adjacent opacities, we opted for one large bounding box, rather than multiple adjacent smaller boxes, to improve consistency in the labeling.  </p>\n<p>Annotators did have access to the COVID status for each patient, but were asked to adhere to the grading system above irrespective of the status. As such, some patients who were COVID negative still had chest radiographs with  typical appearances. Similarly, some patients who were COVID positive had atypical appearances, or were negative for pneumonia (no lung opacities), because the grading system is based off the chest radiographic findings alone.</p>\n<p>The goal in this challenge is to determine the appropriate category for each radiograph, as well as localize the lung opacities with a bounding box prediction. </p>",
      "rawMarkdown": "In this challenge, the chest radiographs (CXRs) were categorized using a specific grading schema, based on a published paper:  \n\nLitmanovich DE, Chung M, Kirkbride RR, Kicska G, Kanne JP. Review of chest radiograph findings of COVID-19 pneumonia and suggested reporting language. Journal of thoracic imaging. 2020 Nov 14;35(6):354-60.\n\nhttps://journals.lww.com/thoracicimaging/Fulltext/2020/11000/Review_of_Chest_Radiograph_Findings_of_COVID_19.4.aspx\n\nPer the grading schema, chest radiographs are classified into one of four categories, which are mutually exclusive:\n\n**1. Typical Appearance:** Multifocal bilateral, peripheral opacities with rounded morphology, lower lung–predominant distribution \n\n**2. Indeterminate Appearance:** Absence of typical findings AND unilateral, central or upper lung predominant distribution\n\n**3. Atypical Appearance:**   Pneumothorax, pleural effusion, pulmonary edema, lobar consolidation, solitary lung nodule or mass, diffuse tiny nodules, cavity\t         \n \n**4. Negative for Pneumonia:**  No lung opacities  \n\nBounding boxes were placed on lung opacities, whether typical or indeterminate. Bounding boxes were also placed on some atypical findings including solitary lobar consolidation, nodules/masses, and cavities.  Bounding boxes were not placed on pleural effusions, or pneumothoraces.  No bounding boxes were placed for the negative for pneumonia category.\n\nIn cases of multiple adjacent opacities, we opted for one large bounding box, rather than multiple adjacent smaller boxes, to improve consistency in the labeling.  \n\nAnnotators did have access to the COVID status for each patient, but were asked to adhere to the grading system above irrespective of the status. As such, some patients who were COVID negative still had chest radiographs with  typical appearances. Similarly, some patients who were COVID positive had atypical appearances, or were negative for pneumonia (no lung opacities), because the grading system is based off the chest radiographic findings alone.\n\nThe goal in this challenge is to determine the appropriate category for each radiograph, as well as localize the lung opacities with a bounding box prediction. \n",
      "votes": 104
    },
    {
      "id": 1315782,
      "postDate": "2021-05-20T05:47:27.627Z",
      "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> Thanks for the information. Can you please clarify the following?</p>\n<ol>\n<li>The data overview mentions that each study can consist of multiple images. Can you please explain what does multiple image mean? Does it mean - for a given study there are 1+ chest radiographs taken at different angles but at the same time? Or does it mean chest radiographs taken at different times? If timing is a factor, then the data becomes a time series data.</li>\n<li>Is there an upper-bound on the number of images per study? For e.g., each study can have no more than 3 images.</li>\n<li>Description in the 'Evaluation' page mentions 'Studies in the test set may contain more than one label.' How is this possible? As per the above explanation, it seems that the goal for each study is to classify into 1 of 4 categories.</li>\n<li>In reference to the above point, if a given study can be classified into multiple categories, why would that happen? Will there be 1-to-1 correspondence with the number of images in the study? For example, if there are 3 images in the study, there will be three different predictions? If so then this just becomes an image level labeling problem.</li>\n<li>The sample submission shown in the 'Evaluation' section of 'Overview' shows the following two lines for the same study, but with one line reporting one class and the other line reporting two classes.<br>\n<code>2b95d54e4be66_study,typical 1 0 0 1 1\n2b95d54e4be66_study,indeterminate 1 0 0 1 1 atypical 1 0 0 1 1</code><br>\n(a) What is the reason for splitting that? (b) Will the confidence value for study predictions be always 1? If not, why would it be non-one? (c) How are we supposed to split and report the study labels?</li>\n</ol>",
      "rawMarkdown": "@paras42 Thanks for the information. Can you please clarify the following?\n1. The data overview mentions that each study can consist of multiple images. Can you please explain what does multiple image mean? Does it mean - for a given study there are 1+ chest radiographs taken at different angles but at the same time? Or does it mean chest radiographs taken at different times? If timing is a factor, then the data becomes a time series data.\n2. Is there an upper-bound on the number of images per study? For e.g., each study can have no more than 3 images.\n3. Description in the 'Evaluation' page mentions 'Studies in the test set may contain more than one label.' How is this possible? As per the above explanation, it seems that the goal for each study is to classify into 1 of 4 categories.\n4. In reference to the above point, if a given study can be classified into multiple categories, why would that happen? Will there be 1-to-1 correspondence with the number of images in the study? For example, if there are 3 images in the study, there will be three different predictions? If so then this just becomes an image level labeling problem.\n5. The sample submission shown in the 'Evaluation' section of 'Overview' shows the following two lines for the same study, but with one line reporting one class and the other line reporting two classes.\n`2b95d54e4be66_study,typical 1 0 0 1 1\n2b95d54e4be66_study,indeterminate 1 0 0 1 1 atypical 1 0 0 1 1 `\n(a) What is the reason for splitting that? (b) Will the confidence value for study predictions be always 1? If not, why would it be non-one? (c) How are we supposed to split and report the study labels?",
      "votes": 25,
      "replies": [
        {
          "id": 1316270,
          "postDate": "2021-05-20T12:27:06.690Z",
          "content": "<p><a href=\"https://www.kaggle.com/hassiahk\" target=\"_blank\">@hassiahk</a> thanks for explaining the competition in the other <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240329#1314661\" target=\"_blank\">thread</a>. That has been really helpful for me to understand. However, I still have the above doubts. Do you have answers to any of the above?</p>",
          "rawMarkdown": "@hassiahk thanks for explaining the competition in the other [thread](https://www.kaggle.com/c/siim-covid19-detection/discussion/240329#1314661). That has been really helpful for me to understand. However, I still have the above doubts. Do you have answers to any of the above?"
        },
        {
          "id": 1316303,
          "postDate": "2021-05-20T12:51:22.483Z",
          "content": "<p>I think this is just a mistake in the description and there is no split for the ID <code>2b95d54e4be66_study</code>. The later one in the evaluation description should have been <code>2b95d54e4be67_study</code>.</p>",
          "rawMarkdown": "I think this is just a mistake in the description and there is no split for the ID `2b95d54e4be66_study`. The later one in the evaluation description should have been `2b95d54e4be67_study`.",
          "votes": 1
        },
        {
          "id": 1316391,
          "postDate": "2021-05-20T14:04:29.957Z",
          "content": "<p>Yeah, that is my thinking as well. But what about multiple labels per study? And the corresponding confidence levels? <a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a>  can expand on these!</p>",
          "rawMarkdown": "Yeah, that is my thinking as well. But what about multiple labels per study? And the corresponding confidence levels? @paras42  can expand on these!"
        },
        {
          "id": 1321539,
          "postDate": "2021-05-24T18:30:07.117Z",
          "content": "<p>Kevin. I did a deeper dive and there is a clarification to my original response.</p>\n<p>Most of the studies only have 1 image.</p>\n<p>In some cases, however, there are studies with more than 1 image.  In these cases, patients were imaged more than once on the same date/time (same StudyInstanceUID).  In some cases, there is motion artifact, so the tech re-took the image. In other cases, different image processing is applied (the images look almost identical, but there is subtle change in contrast).  In other cases, there are coverage, image penetration, or other technique issues, presumably resulting in the technologist needing to retake radiographs.</p>\n<p>There is no upper bound but most patients only have 1 image.</p>\n<p>Regarding your question about number of labels per image, the annotators were asked to only provide one label (e.g. Typical, Indeterminate, Atypical, or Negative for Pneumonia).  However, to paraphrase a recent discussion <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a>, with the hybrid image- and study-level labels we can't enforce mutually exclusive labels for studies (that is we can't stop Kaggle competitors them from submitting multiple labels), but any non-germane labels will count against the submission's overall score.</p>",
          "rawMarkdown": "Kevin. I did a deeper dive and there is a clarification to my original response.\n\nMost of the studies only have 1 image.\n\nIn some cases, however, there are studies with more than 1 image.  In these cases, patients were imaged more than once on the same date/time (same StudyInstanceUID).  In some cases, there is motion artifact, so the tech re-took the image. In other cases, different image processing is applied (the images look almost identical, but there is subtle change in contrast).  In other cases, there are coverage, image penetration, or other technique issues, presumably resulting in the technologist needing to retake radiographs.\n\nThere is no upper bound but most patients only have 1 image.\n\nRegarding your question about number of labels per image, the annotators were asked to only provide one label (e.g. Typical, Indeterminate, Atypical, or Negative for Pneumonia).  However, to paraphrase a recent discussion @philculliton, with the hybrid image- and study-level labels we can't enforce mutually exclusive labels for studies (that is we can't stop Kaggle competitors them from submitting multiple labels), but any non-germane labels will count against the submission's overall score.",
          "votes": 6
        },
        {
          "id": 1321562,
          "postDate": "2021-05-24T18:54:23.407Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> for the clarification.<br>\nYou mention that there should be only one image per study. However, this is not true as per the dataset. A lot of users have performed EDA on the dataset. One of which (<a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240878\" target=\"_blank\">thread</a>) shows explicitly in detail that there are multiple images in each study. Can you please confirm?<br>\nAlso, can you please get the dataset as well as the Overview / Evaluation / Data sections to be fixed to reflect what you are saying? Maybe <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> can help and confirm when everything is fixed? Thanks.</p>",
          "rawMarkdown": "Thanks @paras42 for the clarification.\nYou mention that there should be only one image per study. However, this is not true as per the dataset. A lot of users have performed EDA on the dataset. One of which ([thread](https://www.kaggle.com/c/siim-covid19-detection/discussion/240878)) shows explicitly in detail that there are multiple images in each study. Can you please confirm?\nAlso, can you please get the dataset as well as the Overview / Evaluation / Data sections to be fixed to reflect what you are saying? Maybe @juliaelliott can help and confirm when everything is fixed? Thanks.",
          "votes": 1
        },
        {
          "id": 1321702,
          "postDate": "2021-05-24T21:59:51.933Z",
          "content": "<p>Kevin, sorry you are right. I updated my response above.  </p>",
          "rawMarkdown": "Kevin, sorry you are right. I updated my response above.  "
        },
        {
          "id": 1321717,
          "postDate": "2021-05-24T22:16:25.330Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> for the update. This clarifies things. However, based on the explanation, all studies with more than one images can be assumed as noisy studies. Also, there is no extra information obtained by having multiple images, except the fact that one of them will be good images. In order to get the best results from a machine learning model, it is better that we have good data. Especially in this case where it is very clear which of the images is the correct one to be used. Just a comment!</p>\n<p>As for the number of images in a study, is the number of images in the test dataset also going to be random? The reason for this question is that in order to design an effective model, it is important to have the input as a fixed number and  shape / size. Surely there are workarounds, but it would be best to get a fixed upper bound to best exploit the data and let the DL model make better inference.</p>\n<p>Also, I understand that there can be multiple labels for a study, based on the image(s) in the study. However, I am not clear on how the confidence value will be calculated. Will the confidence value for study predictions always be 1?</p>\n<p>Thanks!</p>",
          "rawMarkdown": "Thanks @paras42 for the update. This clarifies things. However, based on the explanation, all studies with more than one images can be assumed as noisy studies. Also, there is no extra information obtained by having multiple images, except the fact that one of them will be good images. In order to get the best results from a machine learning model, it is better that we have good data. Especially in this case where it is very clear which of the images is the correct one to be used. Just a comment!\n\nAs for the number of images in a study, is the number of images in the test dataset also going to be random? The reason for this question is that in order to design an effective model, it is important to have the input as a fixed number and  shape / size. Surely there are workarounds, but it would be best to get a fixed upper bound to best exploit the data and let the DL model make better inference.\n\nAlso, I understand that there can be multiple labels for a study, based on the image(s) in the study. However, I am not clear on how the confidence value will be calculated. Will the confidence value for study predictions always be 1?\n\nThanks!",
          "votes": 2
        },
        {
          "id": 1322093,
          "postDate": "2021-05-25T07:37:54.210Z",
          "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> In the training set, if a study has multiple images, only 1 of the images will has bounding boxes labelled. Is the test set also annotated in this way? If yes, for the image level prediction, the model will need to guess which image in a study is annotated too?</p>",
          "rawMarkdown": "@paras42 In the training set, if a study has multiple images, only 1 of the images will has bounding boxes labelled. Is the test set also annotated in this way? If yes, for the image level prediction, the model will need to guess which image in a study is annotated too?",
          "votes": 7
        },
        {
          "id": 1322940,
          "postDate": "2021-05-25T20:03:22.410Z",
          "content": "<p>We are now aware regarding some studies having multiple images (some are duplicates), some with bounding boxes, and others with not.  We are investigating different approaches to remedy this issue.  We will soon follow-up when we arrive at a solution and let you know.</p>",
          "rawMarkdown": "We are now aware regarding some studies having multiple images (some are duplicates), some with bounding boxes, and others with not.  We are investigating different approaches to remedy this issue.  We will soon follow-up when we arrive at a solution and let you know.",
          "votes": 7,
          "replies": [
            {
              "id": 1323793,
              "postDate": "2021-05-26T12:52:26.097Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 1322949,
          "postDate": "2021-05-25T20:17:24.633Z",
          "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> Thanks for the reply!</p>",
          "rawMarkdown": "@paras42 Thanks for the reply!"
        },
        {
          "id": 1332133,
          "postDate": "2021-06-01T23:56:54.187Z",
          "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> any update on the issue in the study-level data? - Regarding some studies having multiple images (some are duplicates), some with bounding boxes, and others with not.</p>",
          "rawMarkdown": "@paras42 any update on the issue in the study-level data? - Regarding some studies having multiple images (some are duplicates), some with bounding boxes, and others with not."
        },
        {
          "id": 1335165,
          "postDate": "2021-06-04T03:43:57.743Z",
          "content": "<p>We are addressing this issue currently regarding the duplicates and test set. We'll let you know when this process is completed. </p>",
          "rawMarkdown": "We are addressing this issue currently regarding the duplicates and test set. We'll let you know when this process is completed. "
        },
        {
          "id": 1351109,
          "postDate": "2021-06-16T04:38:54.043Z",
          "content": "<p>Regarding how to handle duplicates, and updates regarding the test labels, please see this post:  <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/246597\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/246597</a></p>\n<p>Hope this helps answer the question</p>",
          "rawMarkdown": "Regarding how to handle duplicates, and updates regarding the test labels, please see this post:  https://www.kaggle.com/c/siim-covid19-detection/discussion/246597\n\nHope this helps answer the question"
        }
      ]
    },
    {
      "id": 1324365,
      "postDate": "2021-05-26T20:51:28.880Z",
      "content": "<p>Can you please explain the evaluation better? It is extremely confusing based on your description. </p>\n<ol>\n<li>Does the \"opacity\" class label in the image level object detection carry any meaning or is it just a placeholder?</li>\n<li>How do the study level labels come into play in the evaluation? In the Overview you claim mAP is used for evaluation. That's fine for an object detection task. But how are the classification labels incorporated in the evaluation, if at all? </li>\n<li>What is the reason you want us to predict 1 pixel bounding boxes for study level images? I am guessing they are just there because you had to cram two prediction tasks into the same submission file format? And that specific 1 pixel bounding box you supplied as example is always correct, therefore <code>IoU = 1</code> for all study level predictions ?</li>\n</ol>\n<p>Are the study level labels incorporated into the object detection mAP calculations because we supply bogus bounding boxes for them that are always correct? </p>\n<p>This is a $100,000 competition. I can't believe this is not explained better in the evaluation page. It is not at all obvious for an outsider what is going on here.</p>",
      "rawMarkdown": "Can you please explain the evaluation better? It is extremely confusing based on your description. \n\n1. Does the \"opacity\" class label in the image level object detection carry any meaning or is it just a placeholder?\n2. How do the study level labels come into play in the evaluation? In the Overview you claim mAP is used for evaluation. That's fine for an object detection task. But how are the classification labels incorporated in the evaluation, if at all? \n3. What is the reason you want us to predict 1 pixel bounding boxes for study level images? I am guessing they are just there because you had to cram two prediction tasks into the same submission file format? And that specific 1 pixel bounding box you supplied as example is always correct, therefore `IoU = 1` for all study level predictions ?\n\nAre the study level labels incorporated into the object detection mAP calculations because we supply bogus bounding boxes for them that are always correct? \n\nThis is a $100,000 competition. I can't believe this is not explained better in the evaluation page. It is not at all obvious for an outsider what is going on here.",
      "votes": 22
    },
    {
      "id": 1315420,
      "postDate": "2021-05-19T19:15:59.777Z",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a>. That is very clear.</p>\n<p>Any word on the number of annotators per image for the train and the test datasets ?</p>",
      "rawMarkdown": "Thanks a lot @paras42. That is very clear.\n\nAny word on the number of annotators per image for the train and the test datasets ?",
      "votes": 3,
      "replies": [
        {
          "id": 1321498,
          "postDate": "2021-05-24T18:00:27.770Z",
          "content": "<p>Alex, due to time constraints, there is only 1 annotator for both train and test datasets. However, annotators did have the ability to ask for a 2nd opinion on some cases, as part of the adjudication. </p>",
          "rawMarkdown": "Alex, due to time constraints, there is only 1 annotator for both train and test datasets. However, annotators did have the ability to ask for a 2nd opinion on some cases, as part of the adjudication. ",
          "votes": 1
        },
        {
          "id": 1349580,
          "postDate": "2021-06-14T23:20:21.407Z",
          "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> <br>\nIs the annotator for train &amp; public test &amp; private test the same person?</p>",
          "rawMarkdown": "@paras42 \nIs the annotator for train & public test & private test the same person?"
        },
        {
          "id": 1351125,
          "postDate": "2021-06-16T04:55:01.077Z",
          "content": "<p>There were 22 annotators. We did not know which cases comprised the test and train datasets.  This was done after the fact by the Kaggle team in a random fashion.   </p>",
          "rawMarkdown": "There were 22 annotators. We did not know which cases comprised the test and train datasets.  This was done after the fact by the Kaggle team in a random fashion.   ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1350682,
      "postDate": "2021-06-15T16:44:51.233Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> ! Thank you so much for your responses above! I am still quite unsure of how we are expected to generate one single study-level label if a study contains multiple images (I saw posts by the Kaggle staff recently saying that there should be only one label per study). For instance, it seems to me from looking at the data that images with \"none\" (i.e., no bounding boxes found) are associated with the \"negative\" study-level label. However, in the train_image_level.csv file on the data tab, study \"00f9e183938e\" contains two images: \"6534a837497d_image\" which has label \"none\" and \"74077a8e3b7c_image\" which has opacities and bounding boxes. However, the overall study label for these two images is \"atypical\" ( see the label for 00f9e183938e_study in train_study_level.csv on the data tab). </p>\n<p>I saw in this discussion that you mentioned this: \"Regarding your question about number of labels per image, the annotators were asked to only provide one label (e.g. Typical, Indeterminate, Atypical, or Negative for Pneumonia)\". Does this mean that for the case above, both \"74077a8e3b7c_image\" and \"6534a837497d_image\" would have a ground truth label \"atypical\" (to match the overall study-level label for that images)? In general, if a study has multiple images, can we consider the ground truth label of each image in the study to be the same as the overall study-level ground truth label (for the classification part of the problem)? I'm sorry if this post is kind of long but this would really clarify some of the competition rules for me! Thank you very much!</p>",
      "rawMarkdown": "Hi @paras42 ! Thank you so much for your responses above! I am still quite unsure of how we are expected to generate one single study-level label if a study contains multiple images (I saw posts by the Kaggle staff recently saying that there should be only one label per study). For instance, it seems to me from looking at the data that images with \"none\" (i.e., no bounding boxes found) are associated with the \"negative\" study-level label. However, in the train_image_level.csv file on the data tab, study \"00f9e183938e\" contains two images: \"6534a837497d_image\" which has label \"none\" and \"74077a8e3b7c_image\" which has opacities and bounding boxes. However, the overall study label for these two images is \"atypical\" ( see the label for 00f9e183938e_study in train_study_level.csv on the data tab). \n\nI saw in this discussion that you mentioned this: \"Regarding your question about number of labels per image, the annotators were asked to only provide one label (e.g. Typical, Indeterminate, Atypical, or Negative for Pneumonia)\". Does this mean that for the case above, both \"74077a8e3b7c_image\" and \"6534a837497d_image\" would have a ground truth label \"atypical\" (to match the overall study-level label for that images)? In general, if a study has multiple images, can we consider the ground truth label of each image in the study to be the same as the overall study-level ground truth label (for the classification part of the problem)? I'm sorry if this post is kind of long but this would really clarify some of the competition rules for me! Thank you very much!",
      "votes": 1,
      "replies": [
        {
          "id": 1351079,
          "postDate": "2021-06-16T04:18:51.633Z",
          "content": "<p>AndyWL, Thanks for the question.  Our annotation team updated the labels for the test datasets (private and public) but the train labels remain the same.</p>\n<p>The updated test labels corrects the problem regarding duplicates. Essentially, bounding box information is now provided for duplicates or similar images that are part of the same study.  </p>\n<p>Regarding the train labels, in situations where there are 2 or more images belonging to a study, I would recommend only using the labels for the image with the bounding boxes and disregard the other images. These other images (duplicates or similar images) were likely not looked at by the annotators as we were unaware of them during the annotation process.  For example, if there are 2 images at the study level, use the data for the image with the bounding boxes.  </p>\n<p>I just created a post here that discusses this:</p>\n<p><a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/246597\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/246597</a></p>",
          "rawMarkdown": "AndyWL, Thanks for the question.  Our annotation team updated the labels for the test datasets (private and public) but the train labels remain the same.\n\nThe updated test labels corrects the problem regarding duplicates. Essentially, bounding box information is now provided for duplicates or similar images that are part of the same study.  \n\nRegarding the train labels, in situations where there are 2 or more images belonging to a study, I would recommend only using the labels for the image with the bounding boxes and disregard the other images. These other images (duplicates or similar images) were likely not looked at by the annotators as we were unaware of them during the annotation process.  For example, if there are 2 images at the study level, use the data for the image with the bounding boxes.  \n\nI just created a post here that discusses this:\n\nhttps://www.kaggle.com/c/siim-covid19-detection/discussion/246597\n\n",
          "votes": 3
        }
      ]
    },
    {
      "id": 1355560,
      "postDate": "2021-06-18T11:16:30.610Z",
      "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a>  What is the reason you want us to predict 1 pixel bounding boxes for study level images?Is it okay if we take confidence as 1 for all study level images?</p>",
      "rawMarkdown": "@paras42  What is the reason you want us to predict 1 pixel bounding boxes for study level images?Is it okay if we take confidence as 1 for all study level images?",
      "votes": 2
    },
    {
      "id": 1321503,
      "postDate": "2021-05-24T18:02:28.103Z",
      "content": "<p>As a clarification to the original post, the annotators were not asked to placed bounding boxes on masses or nodules.  The remainder of the above post is correct.</p>",
      "rawMarkdown": "As a clarification to the original post, the annotators were not asked to placed bounding boxes on masses or nodules.  The remainder of the above post is correct.",
      "votes": 2,
      "replies": [
        {
          "id": 1348703,
          "postDate": "2021-06-14T07:34:32.683Z",
          "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a>  <br>\n1) i still have a confusion as to why we create category \"Negative for Pneumonia \"<br>\n2)secondly when  we say lung opacity it means thats a  Covid cause opacity ? <br>\n3) Thirdly, CXR where there are no BBX but still Negative for Pneumonia so are these cases of Non Covid Pneumonia ?<br>\n4) Is every  Covid 19 infection has lung opacity but every lung opacity may not be covid ?</p>\n<p>Please dont ignore providing clarifications, it may help many .</p>",
          "rawMarkdown": "@paras42  \n1) i still have a confusion as to why we create category \"Negative for Pneumonia \"\n2)secondly when  we say lung opacity it means thats a  Covid cause opacity ? \n3) Thirdly, CXR where there are no BBX but still Negative for Pneumonia so are these cases of Non Covid Pneumonia ?\n4) Is every  Covid 19 infection has lung opacity but every lung opacity may not be covid ?\n\nPlease dont ignore providing clarifications, it may help many ."
        },
        {
          "id": 1351122,
          "postDate": "2021-06-16T04:52:00.670Z",
          "content": "<p>Jaideep,</p>\n<p>Hope this helps:</p>\n<p>1) The classes were taken from a published paper, and also because it follows the same format as the publicly available annotated RICORD COVID-19 chest X-ray dataset. Negative for pneumonia is the same as no lung opacites.</p>\n<p>2) lung opacity is a general term to refer to any process that causes the lung to look opaque on a chest radiograph.  This may mean pneumonia, edema, hemorrhage or cancer, so it is not specific.  Regarding pneumonia, it may mean COVID-19 pneumonia or another type of pneumonia.</p>\n<p>3) Negative for pneumonia means a radiologist did not see any lung opacites.  However, it is known that some patients will have a normal or negative chest radiograph, but still test positive for COVID-19 on PCR.</p>\n<p>4) In this dataset, as in the real-world, patient's with COVID-19 may have a normal chest radiograph, but more commonly patients with COVID do have lung opacities.  We also know that the severity of COVID correlates to the severity of opacities on radiography. If patients have lung opacities, it could be COVID, or it could be something else, like pulmonary edema.  However, based on many observations and experience, and as described in the referenced publication, a typical pattern of COVID-19 on radiography are bilateral, peripheral opacities, and is more likely to represent COVID infection than an indeterminate or atypical pattern.</p>",
          "rawMarkdown": "Jaideep,\n\nHope this helps:\n\n1) The classes were taken from a published paper, and also because it follows the same format as the publicly available annotated RICORD COVID-19 chest X-ray dataset. Negative for pneumonia is the same as no lung opacites.\n\n2) lung opacity is a general term to refer to any process that causes the lung to look opaque on a chest radiograph.  This may mean pneumonia, edema, hemorrhage or cancer, so it is not specific.  Regarding pneumonia, it may mean COVID-19 pneumonia or another type of pneumonia.\n\n3) Negative for pneumonia means a radiologist did not see any lung opacites.  However, it is known that some patients will have a normal or negative chest radiograph, but still test positive for COVID-19 on PCR.\n\n4) In this dataset, as in the real-world, patient's with COVID-19 may have a normal chest radiograph, but more commonly patients with COVID do have lung opacities.  We also know that the severity of COVID correlates to the severity of opacities on radiography. If patients have lung opacities, it could be COVID, or it could be something else, like pulmonary edema.  However, based on many observations and experience, and as described in the referenced publication, a typical pattern of COVID-19 on radiography are bilateral, peripheral opacities, and is more likely to represent COVID infection than an indeterminate or atypical pattern.",
          "votes": 3
        },
        {
          "id": 1353954,
          "postDate": "2021-06-17T09:56:19.537Z",
          "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a>  thanks for spending your time. </p>\n<p>I just now look for only few things</p>\n<p>1) what can we say about those CXR which have no bbx but still marked as 0 - Negative for Pneumonia ,is it the case of Non Covid Pneuomonia ? are those ones  belonging to  ** Bounding boxes were not placed on pleural effusions, or pneumothoraces. No bounding boxes were placed for the negative for pneumonia category.**</p>\n<p>2) For study level classification  in case we have multiple images in a study what should be overall label for the given study  in your this <a href=\"thread\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/246597</a>   you have provided info regarding the bounding box availability of extra image in study.     As per competition data page<br>\n<code>2b95d54e4be65_study,negative 1 0 0 1 1</code>  , we are giving confidence for all the classes for a study , if there are multiple images in a study it means there will be more than one row for a study ?</p>",
          "rawMarkdown": "@paras42  thanks for spending your time. \n\nI just now look for only few things\n\n1) what can we say about those CXR which have no bbx but still marked as 0 - Negative for Pneumonia ,is it the case of Non Covid Pneuomonia ? are those ones  belonging to  ** Bounding boxes were not placed on pleural effusions, or pneumothoraces. No bounding boxes were placed for the negative for pneumonia category.**\n\n2) For study level classification  in case we have multiple images in a study what should be overall label for the given study  in your this [https://www.kaggle.com/c/siim-covid19-detection/discussion/246597](thread)   you have provided info regarding the bounding box availability of extra image in study.     As per competition data page\n`2b95d54e4be65_study,negative 1 0 0 1 1 `  , we are giving confidence for all the classes for a study , if there are multiple images in a study it means there will be more than one row for a study ?"
        },
        {
          "id": 1376869,
          "postDate": "2021-07-05T11:54:36.310Z",
          "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a>  <br>\nabout your reply <br>\n\"3) Negative for pneumonia means a radiologist did not see any lung opacites. However, it is known that some patients will have a normal or negative chest radiograph, but still test positive for COVID-19 on PCR.\"</p>\n<p>How are we deciding the whether such instances are Typical/Atypical/Indeterminate when we know CXR dont show  opacities</p>\n<p>below are study level labels for such cases</p>\n<pre><code>Typical Appearance                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              131\nIndeterminate Appearance                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                         42\nAtypical Appearance                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              88\n</code></pre>",
          "rawMarkdown": "@paras42  \nabout your reply \n\"3) Negative for pneumonia means a radiologist did not see any lung opacites. However, it is known that some patients will have a normal or negative chest radiograph, but still test positive for COVID-19 on PCR.\"\n\nHow are we deciding the whether such instances are Typical/Atypical/Indeterminate when we know CXR dont show  opacities\n\nbelow are study level labels for such cases\n```\nTypical Appearance                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              131\nIndeterminate Appearance                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                         42\nAtypical Appearance                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              88\n```"
        }
      ]
    },
    {
      "id": 1379729,
      "postDate": "2021-07-07T14:57:05.643Z",
      "content": "<p>request to clarify the evaluation better, it would be helpful to understand better </p>",
      "rawMarkdown": "request to clarify the evaluation better, it would be helpful to understand better "
    },
    {
      "id": 1366314,
      "postDate": "2021-06-26T16:46:07.680Z",
      "content": "<p>Can we use custom pre-trained models which are trained on private cloud for this competition </p>",
      "rawMarkdown": "Can we use custom pre-trained models which are trained on private cloud for this competition "
    },
    {
      "id": 1341970,
      "postDate": "2021-06-09T05:47:03.253Z",
      "content": "<p>Appreciate <a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> very much!</p>\n<p>Because this is a competition about covid19, I understand that annotating the bounding box of covid19 is difficult and ambiguous. But for the patient-level classification labels, any type of lung appearance do not definitely means covid19 positive or negative. i.e. the covid19 positive may be any type of appearance and so do the covid19 negative cases.<br>\nAs stated in the supported paper, these type of lung findings may indicate different meanings related to covid19, but</p>\n<ol>\n<li><p>I wonder why not directly predicting the covid19 positive/negative for patient level, instead of predicting lung appearance types?</p></li>\n<li><p>Is it possible to provide the covid19 status for each patient so that we can understand the relationship between these labels and covid19 more?</p></li>\n</ol>\n<p>Greatly thanks!</p>",
      "rawMarkdown": "Appreciate @paras42 very much!\n\nBecause this is a competition about covid19, I understand that annotating the bounding box of covid19 is difficult and ambiguous. But for the patient-level classification labels, any type of lung appearance do not definitely means covid19 positive or negative. i.e. the covid19 positive may be any type of appearance and so do the covid19 negative cases.\nAs stated in the supported paper, these type of lung findings may indicate different meanings related to covid19, but\n\n1.    I wonder why not directly predicting the covid19 positive/negative for patient level, instead of predicting lung appearance types?\n\n2.    Is it possible to provide the covid19 status for each patient so that we can understand the relationship between these labels and covid19 more?\n\nGreatly thanks!"
    },
    {
      "id": 1332013,
      "postDate": "2021-06-01T21:06:39.507Z",
      "content": "<blockquote>\n  <p>Bounding boxes were not placed on pleural effusions, or pneumothoraces. No bounding boxes were placed for the negative for pneumonia category.</p>\n</blockquote>\n<p>There is one image ('09bb2eb902cb_study', '4390dc824c58_image') that has no boxes and comes with 'Typical Appearance'. Probably annotators missed it.</p>",
      "rawMarkdown": ">  Bounding boxes were not placed on pleural effusions, or pneumothoraces. No bounding boxes were placed for the negative for pneumonia category.\n\nThere is one image ('09bb2eb902cb_study', '4390dc824c58_image') that has no boxes and comes with 'Typical Appearance'. Probably annotators missed it."
    },
    {
      "id": 1324333,
      "postDate": "2021-05-26T20:13:57.713Z",
      "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> Does the dataset contain both <strong>PA</strong> and <strong>AP</strong> view?</p>",
      "rawMarkdown": "@paras42 Does the dataset contain both **PA** and **AP** view?",
      "replies": [
        {
          "id": 1335159,
          "postDate": "2021-06-04T03:32:06.023Z",
          "content": "<p>Yes it contains both views.</p>",
          "rawMarkdown": "Yes it contains both views.",
          "votes": 2
        },
        {
          "id": 1343363,
          "postDate": "2021-06-10T07:09:09.880Z",
          "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a>  can there be an official announcement  to ack the various cited issues with Train dataset  , many GM/competitors are holding on their joining the competition knowing about the issues so we are unclear if there will be rest of LB once issues are sorted  or it would continue as is</p>",
          "rawMarkdown": "@paras42  can there be an official announcement  to ack the various cited issues with Train dataset  , many GM/competitors are holding on their joining the competition knowing about the issues so we are unclear if there will be rest of LB once issues are sorted  or it would continue as is",
          "votes": 3
        },
        {
          "id": 1351086,
          "postDate": "2021-06-16T04:21:34.673Z",
          "content": "<p>Jaideep, thanks for pointing that out.</p>\n<p>I just created a post here, which hopefully clarify things:</p>\n<p><a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/246597\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/246597</a></p>",
          "rawMarkdown": "Jaideep, thanks for pointing that out.\n\nI just created a post here, which hopefully clarify things:\n\nhttps://www.kaggle.com/c/siim-covid19-detection/discussion/246597"
        }
      ]
    },
    {
      "id": 1340852,
      "postDate": "2021-06-08T09:55:18.947Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1315151,
      "postDate": "2021-05-19T15:46:25.983Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1444525,
      "postDate": "2021-08-04T06:12:52.123Z",
      "content": "<p>thank you very much</p>",
      "rawMarkdown": "thank you very much"
    },
    {
      "id": 1408422,
      "postDate": "2021-08-02T12:54:15.573Z",
      "content": "<p>Thanks a lot, this is very helpful!</p>",
      "rawMarkdown": "Thanks a lot, this is very helpful!\n\n"
    },
    {
      "id": 1315240,
      "postDate": "2021-05-19T16:46:59.477Z",
      "content": "<p>Thanks a lot, this is very helpful!</p>",
      "rawMarkdown": "Thanks a lot, this is very helpful!"
    }
  ],
  "comments": [
    {
      "id": 1315782,
      "author_name": "Kevin Desai",
      "author_url": "",
      "post_date": "2021-05-20T05:47:27.627000",
      "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> Thanks for the information. Can you please clarify the following?</p>\n<ol>\n<li>The data overview mentions that each study can consist of multiple images. Can you please explain what does multiple image mean? Does it mean - for a given study there are 1+ chest radiographs taken at different angles but at the same time? Or does it mean chest radiographs taken at different times? If timing is a factor, then the data becomes a time series data.</li>\n<li>Is there an upper-bound on the number of images per study? For e.g., each study can have no more than 3 images.</li>\n<li>Description in the 'Evaluation' page mentions 'Studies in the test set may contain more than one label.' How is this possible? As per the above explanation, it seems that the goal for each study is to classify into 1 of 4 categories.</li>\n<li>In reference to the above point, if a given study can be classified into multiple categories, why would that happen? Will there be 1-to-1 correspondence with the number of images in the study? For example, if there are 3 images in the study, there will be three different predictions? If so then this just becomes an image level labeling problem.</li>\n<li>The sample submission shown in the 'Evaluation' section of 'Overview' shows the following two lines for the same study, but with one line reporting one class and the other line reporting two classes.<br>\n<code>2b95d54e4be66_study,typical 1 0 0 1 1\n2b95d54e4be66_study,indeterminate 1 0 0 1 1 atypical 1 0 0 1 1</code><br>\n(a) What is the reason for splitting that? (b) Will the confidence value for study predictions be always 1? If not, why would it be non-one? (c) How are we supposed to split and report the study labels?</li>\n</ol>",
      "votes": 25,
      "replies": [
        {
          "id": 1316270,
          "author_name": "Kevin Desai",
          "author_url": "",
          "post_date": "2021-05-20T12:27:06.690000",
          "content": "<p><a href=\"https://www.kaggle.com/hassiahk\" target=\"_blank\">@hassiahk</a> thanks for explaining the competition in the other <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240329#1314661\" target=\"_blank\">thread</a>. That has been really helpful for me to understand. However, I still have the above doubts. Do you have answers to any of the above?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1316303,
          "author_name": "Haswanth Aekula",
          "author_url": "",
          "post_date": "2021-05-20T12:51:22.483000",
          "content": "<p>I think this is just a mistake in the description and there is no split for the ID <code>2b95d54e4be66_study</code>. The later one in the evaluation description should have been <code>2b95d54e4be67_study</code>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1316391,
          "author_name": "Kevin Desai",
          "author_url": "",
          "post_date": "2021-05-20T14:04:29.957000",
          "content": "<p>Yeah, that is my thinking as well. But what about multiple labels per study? And the corresponding confidence levels? <a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a>  can expand on these!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1321539,
          "author_name": "ParasLakhani",
          "author_url": "",
          "post_date": "2021-05-24T18:30:07.117000",
          "content": "<p>Kevin. I did a deeper dive and there is a clarification to my original response.</p>\n<p>Most of the studies only have 1 image.</p>\n<p>In some cases, however, there are studies with more than 1 image.  In these cases, patients were imaged more than once on the same date/time (same StudyInstanceUID).  In some cases, there is motion artifact, so the tech re-took the image. In other cases, different image processing is applied (the images look almost identical, but there is subtle change in contrast).  In other cases, there are coverage, image penetration, or other technique issues, presumably resulting in the technologist needing to retake radiographs.</p>\n<p>There is no upper bound but most patients only have 1 image.</p>\n<p>Regarding your question about number of labels per image, the annotators were asked to only provide one label (e.g. Typical, Indeterminate, Atypical, or Negative for Pneumonia).  However, to paraphrase a recent discussion <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a>, with the hybrid image- and study-level labels we can't enforce mutually exclusive labels for studies (that is we can't stop Kaggle competitors them from submitting multiple labels), but any non-germane labels will count against the submission's overall score.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1321562,
          "author_name": "Kevin Desai",
          "author_url": "",
          "post_date": "2021-05-24T18:54:23.407000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> for the clarification.<br>\nYou mention that there should be only one image per study. However, this is not true as per the dataset. A lot of users have performed EDA on the dataset. One of which (<a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240878\" target=\"_blank\">thread</a>) shows explicitly in detail that there are multiple images in each study. Can you please confirm?<br>\nAlso, can you please get the dataset as well as the Overview / Evaluation / Data sections to be fixed to reflect what you are saying? Maybe <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> can help and confirm when everything is fixed? Thanks.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1321702,
          "author_name": "ParasLakhani",
          "author_url": "",
          "post_date": "2021-05-24T21:59:51.933000",
          "content": "<p>Kevin, sorry you are right. I updated my response above.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1321717,
          "author_name": "Kevin Desai",
          "author_url": "",
          "post_date": "2021-05-24T22:16:25.330000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> for the update. This clarifies things. However, based on the explanation, all studies with more than one images can be assumed as noisy studies. Also, there is no extra information obtained by having multiple images, except the fact that one of them will be good images. In order to get the best results from a machine learning model, it is better that we have good data. Especially in this case where it is very clear which of the images is the correct one to be used. Just a comment!</p>\n<p>As for the number of images in a study, is the number of images in the test dataset also going to be random? The reason for this question is that in order to design an effective model, it is important to have the input as a fixed number and  shape / size. Surely there are workarounds, but it would be best to get a fixed upper bound to best exploit the data and let the DL model make better inference.</p>\n<p>Also, I understand that there can be multiple labels for a study, based on the image(s) in the study. However, I am not clear on how the confidence value will be calculated. Will the confidence value for study predictions always be 1?</p>\n<p>Thanks!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1322093,
          "author_name": "ShinSiang",
          "author_url": "",
          "post_date": "2021-05-25T07:37:54.210000",
          "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> In the training set, if a study has multiple images, only 1 of the images will has bounding boxes labelled. Is the test set also annotated in this way? If yes, for the image level prediction, the model will need to guess which image in a study is annotated too?</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1322940,
          "author_name": "ParasLakhani",
          "author_url": "",
          "post_date": "2021-05-25T20:03:22.410000",
          "content": "<p>We are now aware regarding some studies having multiple images (some are duplicates), some with bounding boxes, and others with not.  We are investigating different approaches to remedy this issue.  We will soon follow-up when we arrive at a solution and let you know.</p>",
          "votes": 7,
          "replies": [
            {
              "id": 1323793,
              "author_name": "",
              "author_url": "",
              "post_date": "2021-05-26T12:52:26.097000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1322949,
          "author_name": "ShinSiang",
          "author_url": "",
          "post_date": "2021-05-25T20:17:24.633000",
          "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> Thanks for the reply!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1332133,
          "author_name": "Kevin Desai",
          "author_url": "",
          "post_date": "2021-06-01T23:56:54.187000",
          "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> any update on the issue in the study-level data? - Regarding some studies having multiple images (some are duplicates), some with bounding boxes, and others with not.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1335165,
          "author_name": "ParasLakhani",
          "author_url": "",
          "post_date": "2021-06-04T03:43:57.743000",
          "content": "<p>We are addressing this issue currently regarding the duplicates and test set. We'll let you know when this process is completed. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1351109,
          "author_name": "ParasLakhani",
          "author_url": "",
          "post_date": "2021-06-16T04:38:54.043000",
          "content": "<p>Regarding how to handle duplicates, and updates regarding the test labels, please see this post:  <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/246597\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/246597</a></p>\n<p>Hope this helps answer the question</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1324365,
      "author_name": "Faton",
      "author_url": "",
      "post_date": "2021-05-26T20:51:28.880000",
      "content": "<p>Can you please explain the evaluation better? It is extremely confusing based on your description. </p>\n<ol>\n<li>Does the \"opacity\" class label in the image level object detection carry any meaning or is it just a placeholder?</li>\n<li>How do the study level labels come into play in the evaluation? In the Overview you claim mAP is used for evaluation. That's fine for an object detection task. But how are the classification labels incorporated in the evaluation, if at all? </li>\n<li>What is the reason you want us to predict 1 pixel bounding boxes for study level images? I am guessing they are just there because you had to cram two prediction tasks into the same submission file format? And that specific 1 pixel bounding box you supplied as example is always correct, therefore <code>IoU = 1</code> for all study level predictions ?</li>\n</ol>\n<p>Are the study level labels incorporated into the object detection mAP calculations because we supply bogus bounding boxes for them that are always correct? </p>\n<p>This is a $100,000 competition. I can't believe this is not explained better in the evaluation page. It is not at all obvious for an outsider what is going on here.</p>",
      "votes": 22,
      "replies": []
    },
    {
      "id": 1315420,
      "author_name": "Alexandre Cadrin-Chênevert",
      "author_url": "",
      "post_date": "2021-05-19T19:15:59.777000",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a>. That is very clear.</p>\n<p>Any word on the number of annotators per image for the train and the test datasets ?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1321498,
          "author_name": "ParasLakhani",
          "author_url": "",
          "post_date": "2021-05-24T18:00:27.770000",
          "content": "<p>Alex, due to time constraints, there is only 1 annotator for both train and test datasets. However, annotators did have the ability to ask for a 2nd opinion on some cases, as part of the adjudication. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1349580,
          "author_name": "YujiAriyasu",
          "author_url": "",
          "post_date": "2021-06-14T23:20:21.407000",
          "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> <br>\nIs the annotator for train &amp; public test &amp; private test the same person?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1351125,
          "author_name": "ParasLakhani",
          "author_url": "",
          "post_date": "2021-06-16T04:55:01.077000",
          "content": "<p>There were 22 annotators. We did not know which cases comprised the test and train datasets.  This was done after the fact by the Kaggle team in a random fashion.   </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1350682,
      "author_name": "AndyWL",
      "author_url": "",
      "post_date": "2021-06-15T16:44:51.233000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> ! Thank you so much for your responses above! I am still quite unsure of how we are expected to generate one single study-level label if a study contains multiple images (I saw posts by the Kaggle staff recently saying that there should be only one label per study). For instance, it seems to me from looking at the data that images with \"none\" (i.e., no bounding boxes found) are associated with the \"negative\" study-level label. However, in the train_image_level.csv file on the data tab, study \"00f9e183938e\" contains two images: \"6534a837497d_image\" which has label \"none\" and \"74077a8e3b7c_image\" which has opacities and bounding boxes. However, the overall study label for these two images is \"atypical\" ( see the label for 00f9e183938e_study in train_study_level.csv on the data tab). </p>\n<p>I saw in this discussion that you mentioned this: \"Regarding your question about number of labels per image, the annotators were asked to only provide one label (e.g. Typical, Indeterminate, Atypical, or Negative for Pneumonia)\". Does this mean that for the case above, both \"74077a8e3b7c_image\" and \"6534a837497d_image\" would have a ground truth label \"atypical\" (to match the overall study-level label for that images)? In general, if a study has multiple images, can we consider the ground truth label of each image in the study to be the same as the overall study-level ground truth label (for the classification part of the problem)? I'm sorry if this post is kind of long but this would really clarify some of the competition rules for me! Thank you very much!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1351079,
          "author_name": "ParasLakhani",
          "author_url": "",
          "post_date": "2021-06-16T04:18:51.633000",
          "content": "<p>AndyWL, Thanks for the question.  Our annotation team updated the labels for the test datasets (private and public) but the train labels remain the same.</p>\n<p>The updated test labels corrects the problem regarding duplicates. Essentially, bounding box information is now provided for duplicates or similar images that are part of the same study.  </p>\n<p>Regarding the train labels, in situations where there are 2 or more images belonging to a study, I would recommend only using the labels for the image with the bounding boxes and disregard the other images. These other images (duplicates or similar images) were likely not looked at by the annotators as we were unaware of them during the annotation process.  For example, if there are 2 images at the study level, use the data for the image with the bounding boxes.  </p>\n<p>I just created a post here that discusses this:</p>\n<p><a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/246597\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/246597</a></p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1355560,
      "author_name": "Anagha zachariah",
      "author_url": "",
      "post_date": "2021-06-18T11:16:30.610000",
      "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a>  What is the reason you want us to predict 1 pixel bounding boxes for study level images?Is it okay if we take confidence as 1 for all study level images?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1321503,
      "author_name": "ParasLakhani",
      "author_url": "",
      "post_date": "2021-05-24T18:02:28.103000",
      "content": "<p>As a clarification to the original post, the annotators were not asked to placed bounding boxes on masses or nodules.  The remainder of the above post is correct.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1348703,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2021-06-14T07:34:32.683000",
          "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a>  <br>\n1) i still have a confusion as to why we create category \"Negative for Pneumonia \"<br>\n2)secondly when  we say lung opacity it means thats a  Covid cause opacity ? <br>\n3) Thirdly, CXR where there are no BBX but still Negative for Pneumonia so are these cases of Non Covid Pneumonia ?<br>\n4) Is every  Covid 19 infection has lung opacity but every lung opacity may not be covid ?</p>\n<p>Please dont ignore providing clarifications, it may help many .</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1351122,
          "author_name": "ParasLakhani",
          "author_url": "",
          "post_date": "2021-06-16T04:52:00.670000",
          "content": "<p>Jaideep,</p>\n<p>Hope this helps:</p>\n<p>1) The classes were taken from a published paper, and also because it follows the same format as the publicly available annotated RICORD COVID-19 chest X-ray dataset. Negative for pneumonia is the same as no lung opacites.</p>\n<p>2) lung opacity is a general term to refer to any process that causes the lung to look opaque on a chest radiograph.  This may mean pneumonia, edema, hemorrhage or cancer, so it is not specific.  Regarding pneumonia, it may mean COVID-19 pneumonia or another type of pneumonia.</p>\n<p>3) Negative for pneumonia means a radiologist did not see any lung opacites.  However, it is known that some patients will have a normal or negative chest radiograph, but still test positive for COVID-19 on PCR.</p>\n<p>4) In this dataset, as in the real-world, patient's with COVID-19 may have a normal chest radiograph, but more commonly patients with COVID do have lung opacities.  We also know that the severity of COVID correlates to the severity of opacities on radiography. If patients have lung opacities, it could be COVID, or it could be something else, like pulmonary edema.  However, based on many observations and experience, and as described in the referenced publication, a typical pattern of COVID-19 on radiography are bilateral, peripheral opacities, and is more likely to represent COVID infection than an indeterminate or atypical pattern.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1353954,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2021-06-17T09:56:19.537000",
          "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a>  thanks for spending your time. </p>\n<p>I just now look for only few things</p>\n<p>1) what can we say about those CXR which have no bbx but still marked as 0 - Negative for Pneumonia ,is it the case of Non Covid Pneuomonia ? are those ones  belonging to  ** Bounding boxes were not placed on pleural effusions, or pneumothoraces. No bounding boxes were placed for the negative for pneumonia category.**</p>\n<p>2) For study level classification  in case we have multiple images in a study what should be overall label for the given study  in your this <a href=\"thread\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/246597</a>   you have provided info regarding the bounding box availability of extra image in study.     As per competition data page<br>\n<code>2b95d54e4be65_study,negative 1 0 0 1 1</code>  , we are giving confidence for all the classes for a study , if there are multiple images in a study it means there will be more than one row for a study ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1376869,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2021-07-05T11:54:36.310000",
          "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a>  <br>\nabout your reply <br>\n\"3) Negative for pneumonia means a radiologist did not see any lung opacites. However, it is known that some patients will have a normal or negative chest radiograph, but still test positive for COVID-19 on PCR.\"</p>\n<p>How are we deciding the whether such instances are Typical/Atypical/Indeterminate when we know CXR dont show  opacities</p>\n<p>below are study level labels for such cases</p>\n<pre><code>Typical Appearance                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              131\nIndeterminate Appearance                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                         42\nAtypical Appearance                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              88\n</code></pre>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1379729,
      "author_name": "Jagtaran Kaushal",
      "author_url": "",
      "post_date": "2021-07-07T14:57:05.643000",
      "content": "<p>request to clarify the evaluation better, it would be helpful to understand better </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1366314,
      "author_name": "Raja Thevar",
      "author_url": "",
      "post_date": "2021-06-26T16:46:07.680000",
      "content": "<p>Can we use custom pre-trained models which are trained on private cloud for this competition </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1341970,
      "author_name": "lazyterence",
      "author_url": "",
      "post_date": "2021-06-09T05:47:03.253000",
      "content": "<p>Appreciate <a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> very much!</p>\n<p>Because this is a competition about covid19, I understand that annotating the bounding box of covid19 is difficult and ambiguous. But for the patient-level classification labels, any type of lung appearance do not definitely means covid19 positive or negative. i.e. the covid19 positive may be any type of appearance and so do the covid19 negative cases.<br>\nAs stated in the supported paper, these type of lung findings may indicate different meanings related to covid19, but</p>\n<ol>\n<li><p>I wonder why not directly predicting the covid19 positive/negative for patient level, instead of predicting lung appearance types?</p></li>\n<li><p>Is it possible to provide the covid19 status for each patient so that we can understand the relationship between these labels and covid19 more?</p></li>\n</ol>\n<p>Greatly thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1332013,
      "author_name": "Sergey Zlobin",
      "author_url": "",
      "post_date": "2021-06-01T21:06:39.507000",
      "content": "<blockquote>\n  <p>Bounding boxes were not placed on pleural effusions, or pneumothoraces. No bounding boxes were placed for the negative for pneumonia category.</p>\n</blockquote>\n<p>There is one image ('09bb2eb902cb_study', '4390dc824c58_image') that has no boxes and comes with 'Typical Appearance'. Probably annotators missed it.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1324333,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2021-05-26T20:13:57.713000",
      "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a> Does the dataset contain both <strong>PA</strong> and <strong>AP</strong> view?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1335159,
          "author_name": "ParasLakhani",
          "author_url": "",
          "post_date": "2021-06-04T03:32:06.023000",
          "content": "<p>Yes it contains both views.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1343363,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2021-06-10T07:09:09.880000",
          "content": "<p><a href=\"https://www.kaggle.com/paras42\" target=\"_blank\">@paras42</a>  can there be an official announcement  to ack the various cited issues with Train dataset  , many GM/competitors are holding on their joining the competition knowing about the issues so we are unclear if there will be rest of LB once issues are sorted  or it would continue as is</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1351086,
          "author_name": "ParasLakhani",
          "author_url": "",
          "post_date": "2021-06-16T04:21:34.673000",
          "content": "<p>Jaideep, thanks for pointing that out.</p>\n<p>I just created a post here, which hopefully clarify things:</p>\n<p><a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/246597\" target=\"_blank\">https://www.kaggle.com/c/siim-covid19-detection/discussion/246597</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1340852,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-06-08T09:55:18.947000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1315151,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-19T15:46:25.983000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1444525,
      "author_name": "Shun Katagiri",
      "author_url": "",
      "post_date": "2021-08-04T06:12:52.123000",
      "content": "<p>thank you very much</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1408422,
      "author_name": "Hyunhp",
      "author_url": "",
      "post_date": "2021-08-02T12:54:15.573000",
      "content": "<p>Thanks a lot, this is very helpful!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1315240,
      "author_name": "Kerem Turgutlu",
      "author_url": "",
      "post_date": "2021-05-19T16:46:59.477000",
      "content": "<p>Thanks a lot, this is very helpful!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1314244": "In this challenge, the chest radiographs (CXRs) were categorized using a specific grading schema, based on a published paper:  \n\nLitmanovich DE, Chung M, Kirkbride RR, Kicska G, Kanne JP. Review of chest radiograph findings of COVID-19 pneumonia and suggested reporting language. Journal of thoracic imaging. 2020 Nov 14;35(6):354-60.\n\nhttps://journals.lww.com/thoracicimaging/Fulltext/2020/11000/Review_of_Chest_Radiograph_Findings_of_COVID_19.4.aspx\n\nPer the grading schema, chest radiographs are classified into one of four categories, which are mutually exclusive:\n\n**1. Typical Appearance:** Multifocal bilateral, peripheral opacities with rounded morphology, lower lung–predominant distribution \n\n**2. Indeterminate Appearance:** Absence of typical findings AND unilateral, central or upper lung predominant distribution\n\n**3. Atypical Appearance:**   Pneumothorax, pleural effusion, pulmonary edema, lobar consolidation, solitary lung nodule or mass, diffuse tiny nodules, cavity\t         \n \n**4. Negative for Pneumonia:**  No lung opacities  \n\nBounding boxes were placed on lung opacities, whether typical or indeterminate. Bounding boxes were also placed on some atypical findings including solitary lobar consolidation, nodules/masses, and cavities.  Bounding boxes were not placed on pleural effusions, or pneumothoraces.  No bounding boxes were placed for the negative for pneumonia category.\n\nIn cases of multiple adjacent opacities, we opted for one large bounding box, rather than multiple adjacent smaller boxes, to improve consistency in the labeling.  \n\nAnnotators did have access to the COVID status for each patient, but were asked to adhere to the grading system above irrespective of the status. As such, some patients who were COVID negative still had chest radiographs with  typical appearances. Similarly, some patients who were COVID positive had atypical appearances, or were negative for pneumonia (no lung opacities), because the grading system is based off the chest radiographic findings alone.\n\nThe goal in this challenge is to determine the appropriate category for each radiograph, as well as localize the lung opacities with a bounding box prediction. \n",
    "1315782": "@paras42 Thanks for the information. Can you please clarify the following?\n1. The data overview mentions that each study can consist of multiple images. Can you please explain what does multiple image mean? Does it mean - for a given study there are 1+ chest radiographs taken at different angles but at the same time? Or does it mean chest radiographs taken at different times? If timing is a factor, then the data becomes a time series data.\n2. Is there an upper-bound on the number of images per study? For e.g., each study can have no more than 3 images.\n3. Description in the 'Evaluation' page mentions 'Studies in the test set may contain more than one label.' How is this possible? As per the above explanation, it seems that the goal for each study is to classify into 1 of 4 categories.\n4. In reference to the above point, if a given study can be classified into multiple categories, why would that happen? Will there be 1-to-1 correspondence with the number of images in the study? For example, if there are 3 images in the study, there will be three different predictions? If so then this just becomes an image level labeling problem.\n5. The sample submission shown in the 'Evaluation' section of 'Overview' shows the following two lines for the same study, but with one line reporting one class and the other line reporting two classes.\n`2b95d54e4be66_study,typical 1 0 0 1 1\n2b95d54e4be66_study,indeterminate 1 0 0 1 1 atypical 1 0 0 1 1 `\n(a) What is the reason for splitting that? (b) Will the confidence value for study predictions be always 1? If not, why would it be non-one? (c) How are we supposed to split and report the study labels?",
    "1324365": "Can you please explain the evaluation better? It is extremely confusing based on your description. \n\n1. Does the \"opacity\" class label in the image level object detection carry any meaning or is it just a placeholder?\n2. How do the study level labels come into play in the evaluation? In the Overview you claim mAP is used for evaluation. That's fine for an object detection task. But how are the classification labels incorporated in the evaluation, if at all? \n3. What is the reason you want us to predict 1 pixel bounding boxes for study level images? I am guessing they are just there because you had to cram two prediction tasks into the same submission file format? And that specific 1 pixel bounding box you supplied as example is always correct, therefore `IoU = 1` for all study level predictions ?\n\nAre the study level labels incorporated into the object detection mAP calculations because we supply bogus bounding boxes for them that are always correct? \n\nThis is a $100,000 competition. I can't believe this is not explained better in the evaluation page. It is not at all obvious for an outsider what is going on here.",
    "1315420": "Thanks a lot @paras42. That is very clear.\n\nAny word on the number of annotators per image for the train and the test datasets ?",
    "1350682": "Hi @paras42 ! Thank you so much for your responses above! I am still quite unsure of how we are expected to generate one single study-level label if a study contains multiple images (I saw posts by the Kaggle staff recently saying that there should be only one label per study). For instance, it seems to me from looking at the data that images with \"none\" (i.e., no bounding boxes found) are associated with the \"negative\" study-level label. However, in the train_image_level.csv file on the data tab, study \"00f9e183938e\" contains two images: \"6534a837497d_image\" which has label \"none\" and \"74077a8e3b7c_image\" which has opacities and bounding boxes. However, the overall study label for these two images is \"atypical\" ( see the label for 00f9e183938e_study in train_study_level.csv on the data tab). \n\nI saw in this discussion that you mentioned this: \"Regarding your question about number of labels per image, the annotators were asked to only provide one label (e.g. Typical, Indeterminate, Atypical, or Negative for Pneumonia)\". Does this mean that for the case above, both \"74077a8e3b7c_image\" and \"6534a837497d_image\" would have a ground truth label \"atypical\" (to match the overall study-level label for that images)? In general, if a study has multiple images, can we consider the ground truth label of each image in the study to be the same as the overall study-level ground truth label (for the classification part of the problem)? I'm sorry if this post is kind of long but this would really clarify some of the competition rules for me! Thank you very much!",
    "1355560": "@paras42  What is the reason you want us to predict 1 pixel bounding boxes for study level images?Is it okay if we take confidence as 1 for all study level images?",
    "1321503": "As a clarification to the original post, the annotators were not asked to placed bounding boxes on masses or nodules.  The remainder of the above post is correct.",
    "1379729": "request to clarify the evaluation better, it would be helpful to understand better ",
    "1366314": "Can we use custom pre-trained models which are trained on private cloud for this competition ",
    "1341970": "Appreciate @paras42 very much!\n\nBecause this is a competition about covid19, I understand that annotating the bounding box of covid19 is difficult and ambiguous. But for the patient-level classification labels, any type of lung appearance do not definitely means covid19 positive or negative. i.e. the covid19 positive may be any type of appearance and so do the covid19 negative cases.\nAs stated in the supported paper, these type of lung findings may indicate different meanings related to covid19, but\n\n1.    I wonder why not directly predicting the covid19 positive/negative for patient level, instead of predicting lung appearance types?\n\n2.    Is it possible to provide the covid19 status for each patient so that we can understand the relationship between these labels and covid19 more?\n\nGreatly thanks!",
    "1332013": ">  Bounding boxes were not placed on pleural effusions, or pneumothoraces. No bounding boxes were placed for the negative for pneumonia category.\n\nThere is one image ('09bb2eb902cb_study', '4390dc824c58_image') that has no boxes and comes with 'Typical Appearance'. Probably annotators missed it.",
    "1324333": "@paras42 Does the dataset contain both **PA** and **AP** view?",
    "1340852": "",
    "1315151": "",
    "1444525": "thank you very much",
    "1408422": "Thanks a lot, this is very helpful!\n\n",
    "1315240": "Thanks a lot, this is very helpful!"
  }
}