{
  "id": 64723,
  "title": "Pneumonia Dataset Annotation Methods",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/64723",
  "author_name": "Anouk Stein, MD",
  "post_date": "2018-09-01T01:59:07.407000",
  "votes": 47,
  "comment_count": 16,
  "views": 0,
  "content": "<h1>Images</h1>\n\n<p>The dataset was composed of a subset of 30,000 exams from the original 112,000 dataset (train + test) from the NIH CXR14 dataset using their original labels which were derived from radiology reports and, therefore with the understanding that they were not always accurate. The 30,000 selected exams were comprised of 15,000 exams with pneumonia-like labels ('Pneumonia', 'Infiltration', and 'Consolidation'), a random selection of 7,500 exams with a 'No Findings' label, and another random selection of 7,500 exams without the pneumonia-like labels and without the 'No Findings' label. Random unique identifiers were generated for each of the 30,000 exams. </p>\n\n<h1>Annotation</h1>\n\n<p>All readers used a commercial web-based annotation system and were blinded to the other readers’ annotations. For training of our evaluators, all participating radiologists first practiced on the same set of 50 warm-up chest x-rays in a blinded fashion, and then were unblinded to the other radiologists annotations for the same 50 chest x-rays, as an initial calibration and to allow for questions (eg, does an x-ray with healed rib fractures count and no other findings count as ‘Normal’ -- [yes it was considered ‘Normal’]). The final list of labels were ‘Opacity (High Probability)’, ‘Opacity (Medium Probability)’, ‘Opacity (Low Probability)’, ‘No Opacity / Not Normal’, and ‘Normal.’  We also had a ‘Question’ label designate questions to be answered by an experienced chest radiologist.  </p>\n\n<p>Six board-certified radiologists annotated all 30,000 chest radiographs distributed evenly to determine whether lung opacities suspicious for pneumonia were present on the image with their corresponding bounding box to specify the location. There was also collaboration with multiple members of the Society of Thoracic Radiology (STR), who were primarily responsible for annotating the test set.  A total of 12 STR experts double annotated approximately 4,500 chest x-rays (each of the 4,500 chest x-rays were annotated by two different STR radiologists).   </p>\n\n<p>The final test set cases were multi-read by 3 radiologists,  2 STR radiologists plus an additional radiologist from the initial group of 6 radiologists.</p>\n\n<h1>Adjudication</h1>\n\n<p>Multiread cases that did not agree and multiread cases with isolated (non-overlapping) bounding boxes were adjudicated by  one of two thoracic radiology specialists with more than 10 years of experience. The specialist saw the annotations of all three readers for the adjudicated cases. Somewhat over 15% of the triple read cases were individually adjudicated. For the remaining bounding boxes, the intersection was used if there was at least a 50% overlap by one of the boxes. While using the intersection removed some pixel data, having multiple readers added positive pixels. By incorporating 1500 of the approximately 4500 triple read cases into the training set, we hope to have averaged out some of the potential discrepancies between single and multi-read cases. The remaining 3000 triple read cases comprised the test set. For weak labels, the majority vote was used. </p>\n\n<p>A bounding box in a multiread case was considered isolated if it did not overlap with the bounding boxes of either of the other two readers. That is, the two other readers did not flag that area of the image as being suspicious for pneumonia. If the adjudicator agreed that the isolated box was valid, the box remained as a positive minority opinion. Otherwise, it was removed.</p>\n\n<p>Initially, the bounding boxes were given a confidence score. For the final dataset, low confidence boxes were removed and high/medium were combined into one category of likely pneumonia.  If the case only had a solitary low probability bounding box, the box was removed and the case was labeled Not Normal/No Lung Opacity.</p>\n\n<h1>Assumptions</h1>\n\n<p>The radiologists were from top academic institutions and all had many years of experience. They were given the following instructions:</p>\n\n<ul>\n<li><p>Lung Opacity (bounding box) - a finding on chest radiograph that in a patient with cough and fever has a high likelihood of being pneumonia </p></li>\n<li><p>With the understanding that in the absence of clinical information, lateral radiograph, and serial exams, we have to make assumptions </p></li>\n<li><p>Include any area that is more opaque than the surrounding area (Fleischner definition) </p></li>\n<li><p>Exclude: obvious mass(es), nodule(s), lobar collapse, linear atelectasis </p></li>\n</ul>\n\n<p>In the cases labeled Not Normal/No Lung Opacity, no lung opacity refers to no opacity suspicious for pneumonia. Other non-pneumonia opacities may be present. Also, some of the not normal cases have subtle abnormalities which require a trained eye to discern. (Which, for now, keeps radiologists around.)</p>\n\n<p>Good luck!</p>\n\n<p>Anouk Stein, MD &amp; George Shih, MD from <a href=\"http://md.ai/\">MD.ai</a></p>",
  "messages": [
    {
      "id": 379805,
      "postDate": "2018-09-01T01:59:07.407Z",
      "content": "<h1>Images</h1>\n\n<p>The dataset was composed of a subset of 30,000 exams from the original 112,000 dataset (train + test) from the NIH CXR14 dataset using their original labels which were derived from radiology reports and, therefore with the understanding that they were not always accurate. The 30,000 selected exams were comprised of 15,000 exams with pneumonia-like labels ('Pneumonia', 'Infiltration', and 'Consolidation'), a random selection of 7,500 exams with a 'No Findings' label, and another random selection of 7,500 exams without the pneumonia-like labels and without the 'No Findings' label. Random unique identifiers were generated for each of the 30,000 exams. </p>\n\n<h1>Annotation</h1>\n\n<p>All readers used a commercial web-based annotation system and were blinded to the other readers’ annotations. For training of our evaluators, all participating radiologists first practiced on the same set of 50 warm-up chest x-rays in a blinded fashion, and then were unblinded to the other radiologists annotations for the same 50 chest x-rays, as an initial calibration and to allow for questions (eg, does an x-ray with healed rib fractures count and no other findings count as ‘Normal’ -- [yes it was considered ‘Normal’]). The final list of labels were ‘Opacity (High Probability)’, ‘Opacity (Medium Probability)’, ‘Opacity (Low Probability)’, ‘No Opacity / Not Normal’, and ‘Normal.’  We also had a ‘Question’ label designate questions to be answered by an experienced chest radiologist.  </p>\n\n<p>Six board-certified radiologists annotated all 30,000 chest radiographs distributed evenly to determine whether lung opacities suspicious for pneumonia were present on the image with their corresponding bounding box to specify the location. There was also collaboration with multiple members of the Society of Thoracic Radiology (STR), who were primarily responsible for annotating the test set.  A total of 12 STR experts double annotated approximately 4,500 chest x-rays (each of the 4,500 chest x-rays were annotated by two different STR radiologists).   </p>\n\n<p>The final test set cases were multi-read by 3 radiologists,  2 STR radiologists plus an additional radiologist from the initial group of 6 radiologists.</p>\n\n<h1>Adjudication</h1>\n\n<p>Multiread cases that did not agree and multiread cases with isolated (non-overlapping) bounding boxes were adjudicated by  one of two thoracic radiology specialists with more than 10 years of experience. The specialist saw the annotations of all three readers for the adjudicated cases. Somewhat over 15% of the triple read cases were individually adjudicated. For the remaining bounding boxes, the intersection was used if there was at least a 50% overlap by one of the boxes. While using the intersection removed some pixel data, having multiple readers added positive pixels. By incorporating 1500 of the approximately 4500 triple read cases into the training set, we hope to have averaged out some of the potential discrepancies between single and multi-read cases. The remaining 3000 triple read cases comprised the test set. For weak labels, the majority vote was used. </p>\n\n<p>A bounding box in a multiread case was considered isolated if it did not overlap with the bounding boxes of either of the other two readers. That is, the two other readers did not flag that area of the image as being suspicious for pneumonia. If the adjudicator agreed that the isolated box was valid, the box remained as a positive minority opinion. Otherwise, it was removed.</p>\n\n<p>Initially, the bounding boxes were given a confidence score. For the final dataset, low confidence boxes were removed and high/medium were combined into one category of likely pneumonia.  If the case only had a solitary low probability bounding box, the box was removed and the case was labeled Not Normal/No Lung Opacity.</p>\n\n<h1>Assumptions</h1>\n\n<p>The radiologists were from top academic institutions and all had many years of experience. They were given the following instructions:</p>\n\n<ul>\n<li><p>Lung Opacity (bounding box) - a finding on chest radiograph that in a patient with cough and fever has a high likelihood of being pneumonia </p></li>\n<li><p>With the understanding that in the absence of clinical information, lateral radiograph, and serial exams, we have to make assumptions </p></li>\n<li><p>Include any area that is more opaque than the surrounding area (Fleischner definition) </p></li>\n<li><p>Exclude: obvious mass(es), nodule(s), lobar collapse, linear atelectasis </p></li>\n</ul>\n\n<p>In the cases labeled Not Normal/No Lung Opacity, no lung opacity refers to no opacity suspicious for pneumonia. Other non-pneumonia opacities may be present. Also, some of the not normal cases have subtle abnormalities which require a trained eye to discern. (Which, for now, keeps radiologists around.)</p>\n\n<p>Good luck!</p>\n\n<p>Anouk Stein, MD &amp; George Shih, MD from <a href=\"http://md.ai/\">MD.ai</a></p>",
      "rawMarkdown": "# Images\n\nThe dataset was composed of a subset of 30,000 exams from the original 112,000 dataset (train + test) from the NIH CXR14 dataset using their original labels which were derived from radiology reports and, therefore with the understanding that they were not always accurate. The 30,000 selected exams were comprised of 15,000 exams with pneumonia-like labels ('Pneumonia', 'Infiltration', and 'Consolidation'), a random selection of 7,500 exams with a 'No Findings' label, and another random selection of 7,500 exams without the pneumonia-like labels and without the 'No Findings' label. Random unique identifiers were generated for each of the 30,000 exams. \n\n# Annotation\n\nAll readers used a commercial web-based annotation system and were blinded to the other readers’ annotations. For training of our evaluators, all participating radiologists first practiced on the same set of 50 warm-up chest x-rays in a blinded fashion, and then were unblinded to the other radiologists annotations for the same 50 chest x-rays, as an initial calibration and to allow for questions (eg, does an x-ray with healed rib fractures count and no other findings count as ‘Normal’ -- [yes it was considered ‘Normal’]). The final list of labels were ‘Opacity (High Probability)’, ‘Opacity (Medium Probability)’, ‘Opacity (Low Probability)’, ‘No Opacity / Not Normal’, and ‘Normal.’  We also had a ‘Question’ label designate questions to be answered by an experienced chest radiologist.  \n\nSix board-certified radiologists annotated all 30,000 chest radiographs distributed evenly to determine whether lung opacities suspicious for pneumonia were present on the image with their corresponding bounding box to specify the location. There was also collaboration with multiple members of the Society of Thoracic Radiology (STR), who were primarily responsible for annotating the test set.  A total of 12 STR experts double annotated approximately 4,500 chest x-rays (each of the 4,500 chest x-rays were annotated by two different STR radiologists).   \n\nThe final test set cases were multi-read by 3 radiologists,  2 STR radiologists plus an additional radiologist from the initial group of 6 radiologists.\n\n# Adjudication\n\nMultiread cases that did not agree and multiread cases with isolated (non-overlapping) bounding boxes were adjudicated by  one of two thoracic radiology specialists with more than 10 years of experience. The specialist saw the annotations of all three readers for the adjudicated cases. Somewhat over 15% of the triple read cases were individually adjudicated. For the remaining bounding boxes, the intersection was used if there was at least a 50% overlap by one of the boxes. While using the intersection removed some pixel data, having multiple readers added positive pixels. By incorporating 1500 of the approximately 4500 triple read cases into the training set, we hope to have averaged out some of the potential discrepancies between single and multi-read cases. The remaining 3000 triple read cases comprised the test set. For weak labels, the majority vote was used. \n\nA bounding box in a multiread case was considered isolated if it did not overlap with the bounding boxes of either of the other two readers. That is, the two other readers did not flag that area of the image as being suspicious for pneumonia. If the adjudicator agreed that the isolated box was valid, the box remained as a positive minority opinion. Otherwise, it was removed.\n\nInitially, the bounding boxes were given a confidence score. For the final dataset, low confidence boxes were removed and high/medium were combined into one category of likely pneumonia.  If the case only had a solitary low probability bounding box, the box was removed and the case was labeled Not Normal/No Lung Opacity.\n\n# Assumptions\n\nThe radiologists were from top academic institutions and all had many years of experience. They were given the following instructions:\n\n- Lung Opacity (bounding box) - a finding on chest radiograph that in a patient with cough and fever has a high likelihood of being pneumonia \n\n- With the understanding that in the absence of clinical information, lateral radiograph, and serial exams, we have to make assumptions \n\n- Include any area that is more opaque than the surrounding area (Fleischner definition) \n\n- Exclude: obvious mass(es), nodule(s), lobar collapse, linear atelectasis \n\nIn the cases labeled Not Normal/No Lung Opacity, no lung opacity refers to no opacity suspicious for pneumonia. Other non-pneumonia opacities may be present. Also, some of the not normal cases have subtle abnormalities which require a trained eye to discern. (Which, for now, keeps radiologists around.)\n\nGood luck!\n\nAnouk Stein, MD &amp; George Shih, MD from [MD.ai][1]\n\n\n  [1]: http://md.ai",
      "votes": 47
    },
    {
      "id": 844865,
      "postDate": "2020-05-12T22:34:15.657Z",
      "content": "<p>Pneumonia Boxes Likelihood Data\nThe attached csv contains data from the original project and includes the likelihood that the annotated opacity is pneumonia as determined by the reading radiologist(s). The categories are high, medium, and low. We are releasing this data refinement in conjunction with the RSNA in the hopes that it may help in developing COVID-19 CXR models.</p>",
      "rawMarkdown": "Pneumonia Boxes Likelihood Data\nThe attached csv contains data from the original project and includes the likelihood that the annotated opacity is pneumonia as determined by the reading radiologist(s). The categories are high, medium, and low. We are releasing this data refinement in conjunction with the RSNA in the hopes that it may help in developing COVID-19 CXR models.",
      "votes": 4,
      "replies": [
        {
          "id": 1970576,
          "postDate": "2022-10-04T06:33:55.727Z",
          "content": "<p><a href=\"https://www.kaggle.com/anoukstein\" target=\"_blank\">@anoukstein</a> I plan on using the dataset from this challenge for an experiment to study how clinicians interact with an AI classifier, and I have a few questions about this likelihood data:</p>\n<ol>\n<li>I assume these 15,000 images are part of the stage 2_train_images dataset? If so, does the annotation of medium and high probability only correspond to Lung Opacity images from the training set (or do they also include Not Normal/No Lung Opacity)?</li>\n<li>Since the probability data does not contain the patientID field, is there some way to correlate which annotation corresponds to which patientID from the train_images data?</li>\n</ol>",
          "rawMarkdown": "@anoukstein I plan on using the dataset from this challenge for an experiment to study how clinicians interact with an AI classifier, and I have a few questions about this likelihood data:\n1. I assume these 15,000 images are part of the stage 2_train_images dataset? If so, does the annotation of medium and high probability only correspond to Lung Opacity images from the training set (or do they also include Not Normal/No Lung Opacity)?\n2. Since the probability data does not contain the patientID field, is there some way to correlate which annotation corresponds to which patientID from the train_images data?"
        },
        {
          "id": 1971863,
          "postDate": "2022-10-04T19:40:06.803Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/anirudhsripada\" target=\"_blank\">@anirudhsripada</a> !</p>\n<ol>\n<li>The 15K images are only those with pneumonia and the probability is the estimated probability of pneumonia. An obvious pneumonia would get a high probability while an opacity that could be pneumonia or could just represent atalectasis or something else would get a lower probability.  </li>\n<li>To match these up to the patient id, you would need to create a metadata file (see this excellent notebook by <a href=\"https://www.kaggle.com/jhoward\" target=\"_blank\">@jhoward</a> for instructions: <a href=\"https://www.kaggle.com/code/jhoward/creating-a-metadata-dataframe-fastai\" target=\"_blank\">https://www.kaggle.com/code/jhoward/creating-a-metadata-dataframe-fastai</a>)</li>\n</ol>\n<p>Good luck with your project!<br>\nBest, Anouk</p>",
          "rawMarkdown": "Hi @anirudhsripada !\n\n1. The 15K images are only those with pneumonia and the probability is the estimated probability of pneumonia. An obvious pneumonia would get a high probability while an opacity that could be pneumonia or could just represent atalectasis or something else would get a lower probability.  \n2. To match these up to the patient id, you would need to create a metadata file (see this excellent notebook by @jhoward for instructions: https://www.kaggle.com/code/jhoward/creating-a-metadata-dataframe-fastai)\n\nGood luck with your project!\nBest, Anouk"
        },
        {
          "id": 1972590,
          "postDate": "2022-10-05T08:22:54.837Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/anoukstein\" target=\"_blank\">@anoukstein</a>, thanks a lot for the reply!</p>",
          "rawMarkdown": "Hi @anoukstein, thanks a lot for the reply!"
        }
      ]
    },
    {
      "id": 411822,
      "postDate": "2018-10-29T04:32:26.023Z",
      "content": "<p>@Anouk, according to your description, 2000 images should have been included in the stage 2 test set. However, the stage 2 test set now comprises 3000 images. Can you please explain why?</p>",
      "rawMarkdown": "@Anouk, according to your description, 2000 images should have been included in the stage 2 test set. However, the stage 2 test set now comprises 3000 images. Can you please explain why?",
      "votes": 1
    },
    {
      "id": 380548,
      "postDate": "2018-09-03T01:00:17.767Z",
      "content": "<p>Are the test images in stage 1 and stage 2 annotated in the same way? (i.e. multi-read)</p>",
      "rawMarkdown": "Are the test images in stage 1 and stage 2 annotated in the same way? (i.e. multi-read)",
      "votes": 2,
      "replies": [
        {
          "id": 380608,
          "postDate": "2018-09-03T04:33:00.173Z",
          "content": "<p>Yes. The test images were randomly selected from the 4500 triple read cases. All the readers volunteered their time and did practice cases before they read cases from the main dataset. We are truly grateful that they contributed their effort toward this competition. It will be exciting to see the final algorithms!</p>",
          "rawMarkdown": "Yes. The test images were randomly selected from the 4500 triple read cases. All the readers volunteered their time and did practice cases before they read cases from the main dataset. We are truly grateful that they contributed their effort toward this competition. It will be exciting to see the final algorithms!",
          "votes": 1
        },
        {
          "id": 380866,
          "postDate": "2018-09-03T15:43:13.613Z",
          "content": "<p>Thanks Anouk for the original post and this reply. Very useful information. That is really a great effort and contribution from all the radiologists who participated to these important annotations.</p>\n\n<p>To summarize, is it possible to confirm the available information ?</p>\n\n<p>So 4500 triple read cases :</p>\n\n<ul>\n<li>1500 in training dataset</li>\n<li>1000 in test dataset stage 1 (current LB stage)</li>\n<li>2000 in test dataset stage 2 (final stage)</li>\n</ul>",
          "rawMarkdown": "Thanks Anouk for the original post and this reply. Very useful information. That is really a great effort and contribution from all the radiologists who participated to these important annotations.\n\nTo summarize, is it possible to confirm the available information ?\n\nSo 4500 triple read cases :\n\n - 1500 in training dataset\n - 1000 in test dataset stage 1 (current LB stage)\n - 2000 in test dataset stage 2 (final stage)",
          "votes": 1
        },
        {
          "id": 380903,
          "postDate": "2018-09-03T17:29:06.803Z",
          "content": "<p>Alex, yes. The above is true, thanks for asking. </p>",
          "rawMarkdown": "Alex, yes. The above is true, thanks for asking. "
        },
        {
          "id": 391730,
          "postDate": "2018-09-22T09:47:29.100Z",
          "content": "<p>Is it possible to get IDs for \"1500 in training dataset\"?</p>",
          "rawMarkdown": "Is it possible to get IDs for \"1500 in training dataset\"?",
          "votes": 8
        },
        {
          "id": 414913,
          "postDate": "2018-11-03T21:58:00.487Z",
          "content": "<p>i'm wondering if that's \"2000 in test dataset stage 2\" is still true. As we know there were 3000 images in 2nd stage test set and if it's in line with 1st stage test set, then it should be all triple read cases.</p>",
          "rawMarkdown": "i'm wondering if that's \"2000 in test dataset stage 2\" is still true. As we know there were 3000 images in 2nd stage test set and if it's in line with 1st stage test set, then it should be all triple read cases.",
          "votes": 2
        }
      ]
    },
    {
      "id": 386509,
      "postDate": "2018-09-13T03:22:41.440Z",
      "content": "<p>Are there repeat examinations of the same patient? E.g. if the same patient had multiple radiographs could both be possibly included and how would we identify these? </p>",
      "rawMarkdown": "Are there repeat examinations of the same patient? E.g. if the same patient had multiple radiographs could both be possibly included and how would we identify these? "
    },
    {
      "id": 380157,
      "postDate": "2018-09-01T21:04:51.557Z",
      "content": "<p>Thank you for sharing! This clarifies some issues, especially about opacities in the Not Normal / No Lung Opacity class. </p>",
      "rawMarkdown": "Thank you for sharing! This clarifies some issues, especially about opacities in the Not Normal / No Lung Opacity class. "
    },
    {
      "id": 391059,
      "postDate": "2018-09-21T07:42:49.167Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 387371,
      "postDate": "2018-09-14T19:58:54.200Z",
      "content": "<p>This is very helpful! Thank you :)</p>",
      "rawMarkdown": "This is very helpful! Thank you :)",
      "votes": 1
    },
    {
      "id": 380744,
      "postDate": "2018-09-03T11:20:08.580Z",
      "content": "<p>Thank you for the info.</p>",
      "rawMarkdown": "Thank you for the info.\n"
    }
  ],
  "comments": [
    {
      "id": 844865,
      "author_name": "Anouk Stein, MD",
      "author_url": "",
      "post_date": "2020-05-12T22:34:15.657000",
      "content": "<p>Pneumonia Boxes Likelihood Data\nThe attached csv contains data from the original project and includes the likelihood that the annotated opacity is pneumonia as determined by the reading radiologist(s). The categories are high, medium, and low. We are releasing this data refinement in conjunction with the RSNA in the hopes that it may help in developing COVID-19 CXR models.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1970576,
          "author_name": "Anirudh Sripada",
          "author_url": "",
          "post_date": "2022-10-04T06:33:55.727000",
          "content": "<p><a href=\"https://www.kaggle.com/anoukstein\" target=\"_blank\">@anoukstein</a> I plan on using the dataset from this challenge for an experiment to study how clinicians interact with an AI classifier, and I have a few questions about this likelihood data:</p>\n<ol>\n<li>I assume these 15,000 images are part of the stage 2_train_images dataset? If so, does the annotation of medium and high probability only correspond to Lung Opacity images from the training set (or do they also include Not Normal/No Lung Opacity)?</li>\n<li>Since the probability data does not contain the patientID field, is there some way to correlate which annotation corresponds to which patientID from the train_images data?</li>\n</ol>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1971863,
          "author_name": "Anouk Stein, MD",
          "author_url": "",
          "post_date": "2022-10-04T19:40:06.803000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/anirudhsripada\" target=\"_blank\">@anirudhsripada</a> !</p>\n<ol>\n<li>The 15K images are only those with pneumonia and the probability is the estimated probability of pneumonia. An obvious pneumonia would get a high probability while an opacity that could be pneumonia or could just represent atalectasis or something else would get a lower probability.  </li>\n<li>To match these up to the patient id, you would need to create a metadata file (see this excellent notebook by <a href=\"https://www.kaggle.com/jhoward\" target=\"_blank\">@jhoward</a> for instructions: <a href=\"https://www.kaggle.com/code/jhoward/creating-a-metadata-dataframe-fastai\" target=\"_blank\">https://www.kaggle.com/code/jhoward/creating-a-metadata-dataframe-fastai</a>)</li>\n</ol>\n<p>Good luck with your project!<br>\nBest, Anouk</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1972590,
          "author_name": "Anirudh Sripada",
          "author_url": "",
          "post_date": "2022-10-05T08:22:54.837000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/anoukstein\" target=\"_blank\">@anoukstein</a>, thanks a lot for the reply!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 411822,
      "author_name": "Giulia Savorgnan",
      "author_url": "",
      "post_date": "2018-10-29T04:32:26.023000",
      "content": "<p>@Anouk, according to your description, 2000 images should have been included in the stage 2 test set. However, the stage 2 test set now comprises 3000 images. Can you please explain why?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 380548,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2018-09-03T01:00:17.767000",
      "content": "<p>Are the test images in stage 1 and stage 2 annotated in the same way? (i.e. multi-read)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 380608,
          "author_name": "Anouk Stein, MD",
          "author_url": "",
          "post_date": "2018-09-03T04:33:00.173000",
          "content": "<p>Yes. The test images were randomly selected from the 4500 triple read cases. All the readers volunteered their time and did practice cases before they read cases from the main dataset. We are truly grateful that they contributed their effort toward this competition. It will be exciting to see the final algorithms!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 380866,
          "author_name": "Alexandre Cadrin-Chênevert",
          "author_url": "",
          "post_date": "2018-09-03T15:43:13.613000",
          "content": "<p>Thanks Anouk for the original post and this reply. Very useful information. That is really a great effort and contribution from all the radiologists who participated to these important annotations.</p>\n\n<p>To summarize, is it possible to confirm the available information ?</p>\n\n<p>So 4500 triple read cases :</p>\n\n<ul>\n<li>1500 in training dataset</li>\n<li>1000 in test dataset stage 1 (current LB stage)</li>\n<li>2000 in test dataset stage 2 (final stage)</li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 380903,
          "author_name": "Anouk Stein, MD",
          "author_url": "",
          "post_date": "2018-09-03T17:29:06.803000",
          "content": "<p>Alex, yes. The above is true, thanks for asking. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 391730,
          "author_name": "ZFTurbo",
          "author_url": "",
          "post_date": "2018-09-22T09:47:29.100000",
          "content": "<p>Is it possible to get IDs for \"1500 in training dataset\"?</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 414913,
          "author_name": "Botond Egri",
          "author_url": "",
          "post_date": "2018-11-03T21:58:00.487000",
          "content": "<p>i'm wondering if that's \"2000 in test dataset stage 2\" is still true. As we know there were 3000 images in 2nd stage test set and if it's in line with 1st stage test set, then it should be all triple read cases.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 386509,
      "author_name": "Jarrel Seah",
      "author_url": "",
      "post_date": "2018-09-13T03:22:41.440000",
      "content": "<p>Are there repeat examinations of the same patient? E.g. if the same patient had multiple radiographs could both be possibly included and how would we identify these? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 380157,
      "author_name": "Guy Zahavi",
      "author_url": "",
      "post_date": "2018-09-01T21:04:51.557000",
      "content": "<p>Thank you for sharing! This clarifies some issues, especially about opacities in the Not Normal / No Lung Opacity class. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 391059,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-09-21T07:42:49.167000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 387371,
      "author_name": "Daniel Souza",
      "author_url": "",
      "post_date": "2018-09-14T19:58:54.200000",
      "content": "<p>This is very helpful! Thank you :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 380744,
      "author_name": "Oluwatobi",
      "author_url": "",
      "post_date": "2018-09-03T11:20:08.580000",
      "content": "<p>Thank you for the info.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "379805": "# Images\n\nThe dataset was composed of a subset of 30,000 exams from the original 112,000 dataset (train + test) from the NIH CXR14 dataset using their original labels which were derived from radiology reports and, therefore with the understanding that they were not always accurate. The 30,000 selected exams were comprised of 15,000 exams with pneumonia-like labels ('Pneumonia', 'Infiltration', and 'Consolidation'), a random selection of 7,500 exams with a 'No Findings' label, and another random selection of 7,500 exams without the pneumonia-like labels and without the 'No Findings' label. Random unique identifiers were generated for each of the 30,000 exams. \n\n# Annotation\n\nAll readers used a commercial web-based annotation system and were blinded to the other readers’ annotations. For training of our evaluators, all participating radiologists first practiced on the same set of 50 warm-up chest x-rays in a blinded fashion, and then were unblinded to the other radiologists annotations for the same 50 chest x-rays, as an initial calibration and to allow for questions (eg, does an x-ray with healed rib fractures count and no other findings count as ‘Normal’ -- [yes it was considered ‘Normal’]). The final list of labels were ‘Opacity (High Probability)’, ‘Opacity (Medium Probability)’, ‘Opacity (Low Probability)’, ‘No Opacity / Not Normal’, and ‘Normal.’  We also had a ‘Question’ label designate questions to be answered by an experienced chest radiologist.  \n\nSix board-certified radiologists annotated all 30,000 chest radiographs distributed evenly to determine whether lung opacities suspicious for pneumonia were present on the image with their corresponding bounding box to specify the location. There was also collaboration with multiple members of the Society of Thoracic Radiology (STR), who were primarily responsible for annotating the test set.  A total of 12 STR experts double annotated approximately 4,500 chest x-rays (each of the 4,500 chest x-rays were annotated by two different STR radiologists).   \n\nThe final test set cases were multi-read by 3 radiologists,  2 STR radiologists plus an additional radiologist from the initial group of 6 radiologists.\n\n# Adjudication\n\nMultiread cases that did not agree and multiread cases with isolated (non-overlapping) bounding boxes were adjudicated by  one of two thoracic radiology specialists with more than 10 years of experience. The specialist saw the annotations of all three readers for the adjudicated cases. Somewhat over 15% of the triple read cases were individually adjudicated. For the remaining bounding boxes, the intersection was used if there was at least a 50% overlap by one of the boxes. While using the intersection removed some pixel data, having multiple readers added positive pixels. By incorporating 1500 of the approximately 4500 triple read cases into the training set, we hope to have averaged out some of the potential discrepancies between single and multi-read cases. The remaining 3000 triple read cases comprised the test set. For weak labels, the majority vote was used. \n\nA bounding box in a multiread case was considered isolated if it did not overlap with the bounding boxes of either of the other two readers. That is, the two other readers did not flag that area of the image as being suspicious for pneumonia. If the adjudicator agreed that the isolated box was valid, the box remained as a positive minority opinion. Otherwise, it was removed.\n\nInitially, the bounding boxes were given a confidence score. For the final dataset, low confidence boxes were removed and high/medium were combined into one category of likely pneumonia.  If the case only had a solitary low probability bounding box, the box was removed and the case was labeled Not Normal/No Lung Opacity.\n\n# Assumptions\n\nThe radiologists were from top academic institutions and all had many years of experience. They were given the following instructions:\n\n- Lung Opacity (bounding box) - a finding on chest radiograph that in a patient with cough and fever has a high likelihood of being pneumonia \n\n- With the understanding that in the absence of clinical information, lateral radiograph, and serial exams, we have to make assumptions \n\n- Include any area that is more opaque than the surrounding area (Fleischner definition) \n\n- Exclude: obvious mass(es), nodule(s), lobar collapse, linear atelectasis \n\nIn the cases labeled Not Normal/No Lung Opacity, no lung opacity refers to no opacity suspicious for pneumonia. Other non-pneumonia opacities may be present. Also, some of the not normal cases have subtle abnormalities which require a trained eye to discern. (Which, for now, keeps radiologists around.)\n\nGood luck!\n\nAnouk Stein, MD &amp; George Shih, MD from [MD.ai][1]\n\n\n  [1]: http://md.ai",
    "844865": "Pneumonia Boxes Likelihood Data\nThe attached csv contains data from the original project and includes the likelihood that the annotated opacity is pneumonia as determined by the reading radiologist(s). The categories are high, medium, and low. We are releasing this data refinement in conjunction with the RSNA in the hopes that it may help in developing COVID-19 CXR models.",
    "411822": "@Anouk, according to your description, 2000 images should have been included in the stage 2 test set. However, the stage 2 test set now comprises 3000 images. Can you please explain why?",
    "380548": "Are the test images in stage 1 and stage 2 annotated in the same way? (i.e. multi-read)",
    "386509": "Are there repeat examinations of the same patient? E.g. if the same patient had multiple radiographs could both be possibly included and how would we identify these? ",
    "380157": "Thank you for sharing! This clarifies some issues, especially about opacities in the Not Normal / No Lung Opacity class. ",
    "391059": "",
    "387371": "This is very helpful! Thank you :)",
    "380744": "Thank you for the info.\n"
  }
}