{
  "id": 64620,
  "title": "Data Leak - ViewPosition (AP/PA) correlated with target",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/64620",
  "author_name": "Giulia Savorgnan",
  "post_date": "2018-08-30T21:03:27.604000",
  "votes": 19,
  "comment_count": 31,
  "views": 0,
  "content": "<p>The \"view position\" of the radiograph image appears to be really important in terms of target split. This information is extracted from the metadata header of the radiograph images. </p>\n\n<p>For ViewPosition=='AP' the Target [0,1] split is 0.62 - 0.37\nFor ViewPosition=='PA' the Target [0,1] split is 0.91 - 0.09</p>\n\n<p>AP means Anterior-Posterior, whereas PA means Posterior-Anterior. This <a href=\"https://www.med-ed.virginia.edu/courses/rad/cxr/technique3chest.html\">webpage</a> explains that \"Whenever possible the patient should be imaged in an upright PA position.  AP views are less useful and should be reserved for very ill patients who cannot stand erect\". One way to interpret this target unbalance is that patients that are imaged in an AP position are those that are more ill, and therefore more likely to have contracted pneumonia. Note that the absolute split between AP and PA images is about 50-50, so the above consideration is extremely significant. </p>\n\n<p>You can find more details at this kernel: <a href=\"https://www.kaggle.com/giuliasavorgnan/start-here-beginner-intro-to-lung-opacity-s1\">https://www.kaggle.com/giuliasavorgnan/start-here-beginner-intro-to-lung-opacity-s1</a></p>",
  "messages": [
    {
      "id": 379091,
      "postDate": "2018-08-30T21:03:27.603Z",
      "content": "<p>The \"view position\" of the radiograph image appears to be really important in terms of target split. This information is extracted from the metadata header of the radiograph images. </p>\n\n<p>For ViewPosition=='AP' the Target [0,1] split is 0.62 - 0.37\nFor ViewPosition=='PA' the Target [0,1] split is 0.91 - 0.09</p>\n\n<p>AP means Anterior-Posterior, whereas PA means Posterior-Anterior. This <a href=\"https://www.med-ed.virginia.edu/courses/rad/cxr/technique3chest.html\">webpage</a> explains that \"Whenever possible the patient should be imaged in an upright PA position.  AP views are less useful and should be reserved for very ill patients who cannot stand erect\". One way to interpret this target unbalance is that patients that are imaged in an AP position are those that are more ill, and therefore more likely to have contracted pneumonia. Note that the absolute split between AP and PA images is about 50-50, so the above consideration is extremely significant. </p>\n\n<p>You can find more details at this kernel: <a href=\"https://www.kaggle.com/giuliasavorgnan/start-here-beginner-intro-to-lung-opacity-s1\">https://www.kaggle.com/giuliasavorgnan/start-here-beginner-intro-to-lung-opacity-s1</a></p>",
      "rawMarkdown": "The \"view position\" of the radiograph image appears to be really important in terms of target split. This information is extracted from the metadata header of the radiograph images. \n\nFor ViewPosition=='AP' the Target [0,1] split is 0.62 - 0.37\nFor ViewPosition=='PA' the Target [0,1] split is 0.91 - 0.09\n\nAP means Anterior-Posterior, whereas PA means Posterior-Anterior. This [webpage](https://www.med-ed.virginia.edu/courses/rad/cxr/technique3chest.html) explains that \"Whenever possible the patient should be imaged in an upright PA position.  AP views are less useful and should be reserved for very ill patients who cannot stand erect\". One way to interpret this target unbalance is that patients that are imaged in an AP position are those that are more ill, and therefore more likely to have contracted pneumonia. Note that the absolute split between AP and PA images is about 50-50, so the above consideration is extremely significant. \n\nYou can find more details at this kernel: https://www.kaggle.com/giuliasavorgnan/start-here-beginner-intro-to-lung-opacity-s1",
      "votes": 19
    },
    {
      "id": 379551,
      "postDate": "2018-08-31T14:56:13.163Z",
      "content": "<p>John Zech has a well-written blog post about this issue:  <a href=\"https://medium.com/@jrzech/what-are-radiological-deep-learning-models-actually-learning-f97a546c5b98\">https://medium.com/@jrzech/what-are-radiological-deep-learning-models-actually-learning-f97a546c5b98</a> It is correct to say that an algorithm that concludes that portable xrays are more likely to show pathology does not add anything to medical knowledge. Mihai is spot on when he emphasizes the bounding boxes over the classification.</p>\n\n<p>I am a radiologist employed by MD.ai</p>",
      "rawMarkdown": "John Zech has a well-written blog post about this issue:  https://medium.com/@jrzech/what-are-radiological-deep-learning-models-actually-learning-f97a546c5b98 It is correct to say that an algorithm that concludes that portable xrays are more likely to show pathology does not add anything to medical knowledge. Mihai is spot on when he emphasizes the bounding boxes over the classification.\n\nI am a radiologist employed by MD.ai",
      "votes": 6,
      "replies": [
        {
          "id": 379761,
          "postDate": "2018-08-31T22:29:49.793Z",
          "content": "<p>as much as I wanted to be true (that boxes are more important than classification) - it is not entirely true. The model metric is very sensitive if you put at least one box on a \"good\" x-ray. so having a very good classifier on top of box detection is very important, as classifier would be a very good proxy for thresholding when to show a box and not. I would even start working at this problem from classification perspective (maybe even use whole NIH dataset for that) and only later focus on detection.</p>\n\n<p>So yeah. knowing if the image is PA/AP will be very handy, as @Giulia has stated!</p>",
          "rawMarkdown": "as much as I wanted to be true (that boxes are more important than classification) - it is not entirely true. The model metric is very sensitive if you put at least one box on a \"good\" x-ray. so having a very good classifier on top of box detection is very important, as classifier would be a very good proxy for thresholding when to show a box and not. I would even start working at this problem from classification perspective (maybe even use whole NIH dataset for that) and only later focus on detection.\n\nSo yeah. knowing if the image is PA/AP will be very handy, as @Giulia has stated!",
          "votes": 9
        },
        {
          "id": 379762,
          "postDate": "2018-08-31T22:33:18.790Z",
          "content": "<p>I see your point. Even though there will be bias, classification will still be very useful. It would help to look at  the false positive rate to make sure it's  similar on PA and AP. Also, maybe redacting the 'Portable' label on the films themselves would make a difference. You'd probably need to make a similar redaction box for some non portable films. It might not make a difference because there are other signs that the film is AP/portable ie. larger heart silhouette, often less lung inflation because it's harder for the patient to take a deep breath.</p>",
          "rawMarkdown": "I see your point. Even though there will be bias, classification will still be very useful. It would help to look at  the false positive rate to make sure it's  similar on PA and AP. Also, maybe redacting the 'Portable' label on the films themselves would make a difference. You'd probably need to make a similar redaction box for some non portable films. It might not make a difference because there are other signs that the film is AP/portable ie. larger heart silhouette, often less lung inflation because it's harder for the patient to take a deep breath.",
          "votes": 4
        },
        {
          "id": 379863,
          "postDate": "2018-09-01T05:04:33.067Z",
          "content": "<p>The classification problem is pretty \"simple\" ...  See <a href=\"https://github.com/Azure/AzureChestXRay\">here</a> some recent results ;-)</p>\n\n<p>The segmentation one is the more complicated one. But with tweaking it can be solve using something like <a href=\"https://github.com/matterport/Mask_RCNN\">this</a>.</p>\n\n<p><a href=\"/raddar\">@raddar</a>. you said it yourself, \"if you put at least one box\"</p>",
          "rawMarkdown": "The classification problem is pretty \"simple\" ...  See [here][1] some recent results ;-)\n\nThe segmentation one is the more complicated one. But with tweaking it can be solve using something like [this][2].\n\n@raddar. you said it yourself, \"if you put at least one box\"\n\n\n  [1]: https://github.com/Azure/AzureChestXRay\n  [2]: https://github.com/matterport/Mask_RCNN"
        }
      ]
    },
    {
      "id": 380890,
      "postDate": "2018-09-03T16:49:08.043Z",
      "content": "<p>How is this a data leak? This information is available to the doctor when he takes the image as well.</p>",
      "rawMarkdown": "How is this a data leak? This information is available to the doctor when he takes the image as well.",
      "votes": 4,
      "replies": [
        {
          "id": 380905,
          "postDate": "2018-09-03T17:30:21.337Z",
          "content": "<p>Good point!</p>",
          "rawMarkdown": "Good point!"
        },
        {
          "id": 380936,
          "postDate": "2018-09-03T18:44:54.400Z",
          "content": "<p>Agreed. I guess from the perspective that the task should only consider the image this could be considered a leak. But this is a feature which is included in the DICOM metadata which has clearly been altered as part of the dataset preparation (fields removed or set to the same value throughout dataset) so it was clearly intended that this feature would be accessible to us (not a leak). That's not to say that it isn't an important feature -- it clearly is.</p>",
          "rawMarkdown": "Agreed. I guess from the perspective that the task should only consider the image this could be considered a leak. But this is a feature which is included in the DICOM metadata which has clearly been altered as part of the dataset preparation (fields removed or set to the same value throughout dataset) so it was clearly intended that this feature would be accessible to us (not a leak). That's not to say that it isn't an important feature -- it clearly is."
        },
        {
          "id": 381014,
          "postDate": "2018-09-03T22:36:55.807Z",
          "content": "<p>It means the data is biased. In its essence there's nothing wrong with the view point, but as Giulia pointed out in her kernel, it reveals the ill/critical state of the patient (who can't even stand up). The latter correlates with harsher illnesses like pneumonia.\nRecommend reading the kernel ;)</p>",
          "rawMarkdown": "It means the data is biased. In its essence there's nothing wrong with the view point, but as Giulia pointed out in her kernel, it reveals the ill/critical state of the patient (who can't even stand up). The latter correlates with harsher illnesses like pneumonia.\nRecommend reading the kernel ;)",
          "votes": 1
        }
      ]
    },
    {
      "id": 379252,
      "postDate": "2018-08-31T04:54:27.647Z",
      "content": "<p>Don't worry, there's actually a huge \"classification\" leak because the data is based on NIH and the test data from this competition is present in the training data from NIH.</p>\n\n<p>Keep in mind what does this competition tries to solve! It's not the classification, it's the bounding boxes!! </p>",
      "rawMarkdown": "Don't worry, there's actually a huge \"classification\" leak because the data is based on NIH and the test data from this competition is present in the training data from NIH.\n\nKeep in mind what does this competition tries to solve! It's not the classification, it's the bounding boxes!! ",
      "votes": 4,
      "replies": [
        {
          "id": 381697,
          "postDate": "2018-09-05T02:52:39.043Z",
          "content": "<p>The images are based on NIH.  The data - the labels and bounding boxes - are newly created by the host team.  There isn't a leak.  It's just based on a common set of images.</p>",
          "rawMarkdown": "The images are based on NIH.  The data - the labels and bounding boxes - are newly created by the host team.  There isn't a leak.  It's just based on a common set of images."
        },
        {
          "id": 381850,
          "postDate": "2018-09-05T09:59:23.297Z",
          "content": "<p>Do you mean you've found all the images form NIH  that are part of the \"test\" set as mislabeled? If not, it means that the presence of pneumonia in the \"test\" set of images is already labeled by NIH set. An that's a leak.  (ex. pID 2118026e-ed2f-4d1e-baa2-0e9e1004f72c in Test and    00006473_001 in NIH db)</p>",
          "rawMarkdown": "Do you mean you've found all the images form NIH  that are part of the \"test\" set as mislabeled? If not, it means that the presence of pneumonia in the \"test\" set of images is already labeled by NIH set. An that's a leak.  (ex. pID 2118026e-ed2f-4d1e-baa2-0e9e1004f72c in Test and \t00006473_001 in NIH db)",
          "votes": 1
        },
        {
          "id": 381895,
          "postDate": "2018-09-05T11:29:38.067Z",
          "content": "<p>This data set does not use the NIH labels.  The criteria for whether or not something is labeled with \"pneumonia\" in this particular set is multi-tiered and most certainly does <em>not</em> match the criteria used in the labels you're looking at.</p>\n\n<p>How many radiologists provided the labels you're looking at?  How did they agree on a final diagnosis?  Doctor-doctor agreement for this problem is not exceedingly high, especially if low-grade cases are included.  Were they in your set?  Do you know what severity was assigned to the labels you're looking at?  The host here spent a great deal of effort with multiple passes, multiple doctors, and adjudication to ensure very high-quality labels.  If you're looking at one radiologist's labels (or worse, labels created without a radiologist) you're not looking at remotely the same set of labels used here.  Which is part of why competitors are encouraged to re-annotate the training set themselves with the help of more radiologists.  That's leaving the question of the bounding boxes aside entirely.</p>\n\n<p>Not only that, but no use of the NIH labels - or any other outside labels - is allowed for the test set.</p>\n\n<p>So a) the labels used here are not exposed, and b) if anyone is found to have used the NIH labels to label the test set (however the labels might disagree), they would be disqualified from earning prizes.</p>",
          "rawMarkdown": "This data set does not use the NIH labels.  The criteria for whether or not something is labeled with \"pneumonia\" in this particular set is multi-tiered and most certainly does *not* match the criteria used in the labels you're looking at.\n\nHow many radiologists provided the labels you're looking at?  How did they agree on a final diagnosis?  Doctor-doctor agreement for this problem is not exceedingly high, especially if low-grade cases are included.  Were they in your set?  Do you know what severity was assigned to the labels you're looking at?  The host here spent a great deal of effort with multiple passes, multiple doctors, and adjudication to ensure very high-quality labels.  If you're looking at one radiologist's labels (or worse, labels created without a radiologist) you're not looking at remotely the same set of labels used here.  Which is part of why competitors are encouraged to re-annotate the training set themselves with the help of more radiologists.  That's leaving the question of the bounding boxes aside entirely.\n\nNot only that, but no use of the NIH labels - or any other outside labels - is allowed for the test set.\n\nSo a) the labels used here are not exposed, and b) if anyone is found to have used the NIH labels to label the test set (however the labels might disagree), they would be disqualified from earning prizes.",
          "votes": 3
        },
        {
          "id": 381897,
          "postDate": "2018-09-05T11:32:56.817Z",
          "content": "<p>I would also encourage interested people to read about the annotation methods used for this competition <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723\">here</a>.</p>",
          "rawMarkdown": "I would also encourage interested people to read about the annotation methods used for this competition [here][1].\n\n\n  [1]: https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723"
        },
        {
          "id": 381903,
          "postDate": "2018-09-05T11:43:47.557Z",
          "content": "<p>I'm looking at the labels provided by NIH as per competition <a href=\"https://nihcc.app.box.com/v/ChestXray-NIHCC/folder/36938765345\">\"Original data source\"</a> link.</p>\n\n<p>One could train a classifier based on NIH original data and use the labels provide by NIHs to classify if pneumonia or not. Next, train a box/segmentation model only for \"pneumonia\" images - provide the classifier does a good job on \"test\" set, it would simplify the work of box/segmentation model.</p>\n\n<p>So what you are saying is that the original data source <strong>cannot</strong> be used for training a classifier because it leaks information from the \"Test\" set. </p>\n\n<p>If so, I want to see how that's going to be enforced at the end of the competition. I know this competition requires the \"model(s)\" and obviously the \"weights\" to generate the results. But the \"weights\" could have seen NIH data. How do you plan in evaluating that's not the case?</p>\n\n<p>Are you going to spend the time in re-training the models with the data provided in this competition? How are you going to ensure you get the same random initialization of the parameters, provided the models do not use any pre-trained models? Etc? Bottom line I don't see how you can ensure that winning model hasn't used the NIH data set and for that reason I strong believe that the classification part has a huge leak. As stated, unless most of the images selected to be in the \"Test\" had been mislabeled by NIH you have a leak.</p>",
          "rawMarkdown": "I'm looking at the labels provided by NIH as per competition [\"Original data source\"][1] link.\n\n One could train a classifier based on NIH original data and use the labels provide by NIHs to classify if pneumonia or not. Next, train a box/segmentation model only for \"pneumonia\" images - provide the classifier does a good job on \"test\" set, it would simplify the work of box/segmentation model.\n\nSo what you are saying is that the original data source **cannot** be used for training a classifier because it leaks information from the \"Test\" set. \n\nIf so, I want to see how that's going to be enforced at the end of the competition. I know this competition requires the \"model(s)\" and obviously the \"weights\" to generate the results. But the \"weights\" could have seen NIH data. How do you plan in evaluating that's not the case?\n\nAre you going to spend the time in re-training the models with the data provided in this competition? How are you going to ensure you get the same random initialization of the parameters, provided the models do not use any pre-trained models? Etc? Bottom line I don't see how you can ensure that winning model hasn't used the NIH data set and for that reason I strong believe that the classification part has a huge leak. As stated, unless most of the images selected to be in the \"Test\" had been mislabeled by NIH you have a leak.\n\n  [1]: https://nihcc.app.box.com/v/ChestXray-NIHCC/folder/36938765345",
          "votes": 3
        },
        {
          "id": 381924,
          "postDate": "2018-09-05T12:11:57.073Z",
          "content": "<p>Pipeline</p>\n\n<p>Test data -&gt; Classifier Model\n-&gt; Pneumonia? \n-&gt; YES -&gt; Box/Segmentation Model (trained only on pneumonia data) -&gt; ensemble ... etc \n-&gt; NO - No Box</p>\n\n<p>If the Classifier Model is overfitted (memorized) NIH data the problem is simpler. So ... leak ...</p>",
          "rawMarkdown": "Pipeline\n\nTest data -&gt; Classifier Model\n-&gt; Pneumonia? \n-&gt; YES -&gt; Box/Segmentation Model (trained only on pneumonia data) -&gt; ensemble ... etc \n-&gt; NO - No Box\n\nIf the Classifier Model is overfitted (memorized) NIH data the problem is simpler. So ... leak ...",
          "votes": 1
        },
        {
          "id": 381977,
          "postDate": "2018-09-05T13:23:56.730Z",
          "content": "<p>Mihai -\n   It is widely perceived by radiologists who practice machine learning as well as data scientists like Jeremy Howard that the NIH Chest X-ray 14 dataset was a \"first effort\" dataset, and therefore suffered from a older, and unfortunately substantially inaccurate, NLP classification system which substantially decreased the accuracy of labels, lower than I think many of us would like to have seen.  ChestXRay14 was released to a great deal of fanfare, along with a sensational statement about CheXNet, which is probably why the RSNA has chosen this challenge to make a more fair assessment.</p>\n\n<p>Here are <a href=\"https://n2value.com/blog/chexnet-a-brief-evaluation/\">my observations on the dataset</a>.  This <a href=\"https://n2value.com/blog/are-computers-better-than-doctors-will-the-computer-see-you-now-what-we-learnt-from-the-chexnet-paper-for-pneumonia-diagnosis/\">online meetup discussed important points of the dataset</a>.  And Luke Oakden Rayner <a href=\"https://lukeoakdenrayner.wordpress.com/2017/12/18/the-chestxray14-dataset-problems/\">has written about it extensively, as well as providing his own estimates of the accuracy of the dataset</a>.  A SIIM-centric working group spent time on this dataset to re-label pneumothoraces (data I believe is being presented in a few days at C-MIMI) in large because of these inaccuracies.  I suspect most professional ML teams who are using the CXR-14 dataset in production have needed to hire radiologists to re-label.  </p>\n\n<p>If you're looking to use CXR-14 as an opportunity for transfer learning, its not entirely unreasonable, try it!  But be very clear about the underlying strata you are working with.  </p>",
          "rawMarkdown": "Mihai -\n   It is widely perceived by radiologists who practice machine learning as well as data scientists like Jeremy Howard that the NIH Chest X-ray 14 dataset was a \"first effort\" dataset, and therefore suffered from a older, and unfortunately substantially inaccurate, NLP classification system which substantially decreased the accuracy of labels, lower than I think many of us would like to have seen.  ChestXRay14 was released to a great deal of fanfare, along with a sensational statement about CheXNet, which is probably why the RSNA has chosen this challenge to make a more fair assessment.\n\nHere are [my observations on the dataset][1].  This [online meetup discussed important points of the dataset][2].  And Luke Oakden Rayner [has written about it extensively, as well as providing his own estimates of the accuracy of the dataset][3].  A SIIM-centric working group spent time on this dataset to re-label pneumothoraces (data I believe is being presented in a few days at C-MIMI) in large because of these inaccuracies.  I suspect most professional ML teams who are using the CXR-14 dataset in production have needed to hire radiologists to re-label.  \n\nIf you're looking to use CXR-14 as an opportunity for transfer learning, its not entirely unreasonable, try it!  But be very clear about the underlying strata you are working with.  \n\n\n  [1]: https://n2value.com/blog/chexnet-a-brief-evaluation/\n  [2]: https://n2value.com/blog/are-computers-better-than-doctors-will-the-computer-see-you-now-what-we-learnt-from-the-chexnet-paper-for-pneumonia-diagnosis/\n  [3]: https://lukeoakdenrayner.wordpress.com/2017/12/18/the-chestxray14-dataset-problems/",
          "votes": 2
        },
        {
          "id": 382047,
          "postDate": "2018-09-05T15:53:09.320Z",
          "content": "<p>@drsxr - excellent post. </p>\n\n<p>Now here's an idea, let's check how many from \"Train\" set are mislabeled in NIH set :-) Some insights can be obtained from that.  </p>",
          "rawMarkdown": "@drsxr - excellent post. \n\nNow here's an idea, let's check how many from \"Train\" set are mislabeled in NIH set :-) Some insights can be obtained from that.  ",
          "votes": 2
        },
        {
          "id": 382510,
          "postDate": "2018-09-06T13:56:32.990Z",
          "content": "<p><a href=\"/philculliton\">@philculliton</a>. Please clarify on how does the host disqualify a team. At this stage, I don't know what exactly a \"model\" is, which is required to be uploaded after stage 1. If the host receive a code,</p>\n\n<p>1) will the host re-run the whole model (or, in other words, Machine Learning code) on the training set of this competition, then use that trained model predict on the test set of stage 1 to confirm that the result on the leaderboard is identical, then, use that model to predict the test set of stage 2? Or</p>\n\n<p>2) the host only require the trained model uploaded by the team (or with code also for verification that \"there is some model\", but the host will NOT re-run the code on the training set of this competition), then direct use that model to predict the test set of stage 1 to confirm?</p>\n\n<p>If case 2) is right, then I'm afraid the models may be trained with some sort of leak information obtained from NIH. There's no way the host know whether the team use any sort of external leak data or not.</p>\n\n<p>And there is other issue. Even if the host can manage to do the verification described in 1) for the the few top teams, the other teams will be exempted as the result of the burden of this check, and they will get Kaggle medals with some leak information. </p>\n\n<p>This competition therefore, has too many problems that concern, I believe, not only me.</p>",
          "rawMarkdown": "@philculliton. Please clarify on how does the host disqualify a team. At this stage, I don't know what exactly a \"model\" is, which is required to be uploaded after stage 1. If the host receive a code,\n\n1) will the host re-run the whole model (or, in other words, Machine Learning code) on the training set of this competition, then use that trained model predict on the test set of stage 1 to confirm that the result on the leaderboard is identical, then, use that model to predict the test set of stage 2? Or\n\n2) the host only require the trained model uploaded by the team (or with code also for verification that \"there is some model\", but the host will NOT re-run the code on the training set of this competition), then direct use that model to predict the test set of stage 1 to confirm?\n\nIf case 2) is right, then I'm afraid the models may be trained with some sort of leak information obtained from NIH. There's no way the host know whether the team use any sort of external leak data or not.\n\nAnd there is other issue. Even if the host can manage to do the verification described in 1) for the the few top teams, the other teams will be exempted as the result of the burden of this check, and they will get Kaggle medals with some leak information. \n\nThis competition therefore, has too many problems that concern, I believe, not only me."
        },
        {
          "id": 382569,
          "postDate": "2018-09-06T16:45:22.590Z",
          "content": "<p>@Kha Vo - Please review the rules and timeline for details on host verification. You are free to choose not to participate in the competition, if this is a design that does not appeal to you.</p>",
          "rawMarkdown": "@Kha Vo - Please review the rules and timeline for details on host verification. You are free to choose not to participate in the competition, if this is a design that does not appeal to you.",
          "votes": -3
        },
        {
          "id": 382577,
          "postDate": "2018-09-06T17:05:35.400Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 382585,
          "postDate": "2018-09-06T17:20:08.903Z",
          "content": "<p>I am asking for the sake of this competition, more especially, the fairness. I am a serious competitor and not f*<em>*</em>* around asking nonsense. There’s no way to stop a lot of cheatings around.  I feel your comment is a little bit subjectively sensitive.</p>\n\n<p>The rules did not clearly state how the host verify the integrity of the uploaded models.  That’s why I asked. How do you know if a model is trained by an external annotated dataset that partly overlaps with the test set?  How do you handle that and make the other honest teams feel relieved to participate?</p>\n\n<p>You cannot definitely say that we expect there’s no cheaters. There will be definitely not a few, but a lot. Only using some small leak may heavily affect the leaderboard. Can you imagine the consequence if some of the top teams all cheat? My question is a constructive question. </p>\n\n<p>There was a recent Kaggle competition with leak discovered during the competition, and exploited relentlessly by almost all teams.  </p>",
          "rawMarkdown": "I am asking for the sake of this competition, more especially, the fairness. I am a serious competitor and not f***** around asking nonsense. There’s no way to stop a lot of cheatings around.  I feel your comment is a little bit subjectively sensitive.\n\nThe rules did not clearly state how the host verify the integrity of the uploaded models.  That’s why I asked. How do you know if a model is trained by an external annotated dataset that partly overlaps with the test set?  How do you handle that and make the other honest teams feel relieved to participate?\n\nYou cannot definitely say that we expect there’s no cheaters. There will be definitely not a few, but a lot. Only using some small leak may heavily affect the leaderboard. Can you imagine the consequence if some of the top teams all cheat? My question is a constructive question. \n\nThere was a recent Kaggle competition with leak discovered during the competition, and exploited relentlessly by almost all teams.  ",
          "votes": -1
        },
        {
          "id": 382591,
          "postDate": "2018-09-06T17:26:36.263Z",
          "content": "<p>No offense or trivialization of your question was intended.</p>",
          "rawMarkdown": "No offense or trivialization of your question was intended.",
          "votes": 2
        },
        {
          "id": 382626,
          "postDate": "2018-09-06T18:23:25.060Z",
          "content": "<p>@Kha Vo - if you'll read the info from the links posted by @drsxr you'll find out that apparently many images from the original NIH data set aren't properly labeled. That's because they've used an NLP model to read the medical file associated with the image and auto-labeled the images. The set of images provided by RSNA have be correctly labeled using a team of radiologists. </p>\n\n<p>I gave it a bit more thought, using a clarifies over-fit on NIH set (even with correct labels) would fail on a Stage 2 - provided that there is a stage 2 and that the images are not part of the NIH set. That would filter out \"cheaters\" that exploit the leak. </p>\n\n<p>Ideally the Stage 2 images should be from NIH <strong>Test</strong> set and not Train set :-) Ideally should be images that weren't publicly available ...</p>",
          "rawMarkdown": "@Kha Vo - if you'll read the info from the links posted by @drsxr you'll find out that apparently many images from the original NIH data set aren't properly labeled. That's because they've used an NLP model to read the medical file associated with the image and auto-labeled the images. The set of images provided by RSNA have be correctly labeled using a team of radiologists. \n\nI gave it a bit more thought, using a clarifies over-fit on NIH set (even with correct labels) would fail on a Stage 2 - provided that there is a stage 2 and that the images are not part of the NIH set. That would filter out \"cheaters\" that exploit the leak. \n\nIdeally the Stage 2 images should be from NIH **Test** set and not Train set :-) Ideally should be images that weren't publicly available ...",
          "votes": 3
        },
        {
          "id": 382735,
          "postDate": "2018-09-07T02:17:53.150Z",
          "content": "<p>@Anouk. My last comment was not for you.</p>",
          "rawMarkdown": "@Anouk. My last comment was not for you."
        },
        {
          "id": 382743,
          "postDate": "2018-09-07T02:32:35.623Z",
          "content": "<p>@Kha Vo 😀</p>",
          "rawMarkdown": "@Kha Vo 😀",
          "votes": 1
        },
        {
          "id": 383166,
          "postDate": "2018-09-07T21:17:09.033Z",
          "content": "<p>I do not intend to be dismissive. Your questions are valid. You should expect that models will be evaluated both (A) to ensure the Stage 2 predictions were generated from the Stage 1 Uploaded model, and (B) to confirm compliance with restrictions around eligible data usage. This could mean going to the extent of running the whole model. The expectations for <a href=\"https://www.kaggle.com/WinningModelDocumentationGuidelines\">winners' model documentation</a> should help to clarify the extent of what could be investigated for those in prize-standing. Ultimately, verification is in the hands of the host, and the extent of enforcement will be at their discretion.</p>\n\n<p>We (Kaggle and its hosts) bring challenges to the platform because they represent real-world problems being solved by research or commercial organizations. We recognize that in order to make the problems posed realistic and applicable in advancing any field, we have to rely - in some part - on the integrity &amp; honesty of participants, and not just the threat or enforceability of verification &amp; legal stipulations. As participants, you should all be aware of the competition's <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/rules\">Code of Conduct</a> and recognize your contributions have the potential to make real impacts to healthcare. I hope we all share in that spirit of integrity.</p>",
          "rawMarkdown": "I do not intend to be dismissive. Your questions are valid. You should expect that models will be evaluated both (A) to ensure the Stage 2 predictions were generated from the Stage 1 Uploaded model, and (B) to confirm compliance with restrictions around eligible data usage. This could mean going to the extent of running the whole model. The expectations for [winners' model documentation][1] should help to clarify the extent of what could be investigated for those in prize-standing. Ultimately, verification is in the hands of the host, and the extent of enforcement will be at their discretion.\n\nWe (Kaggle and its hosts) bring challenges to the platform because they represent real-world problems being solved by research or commercial organizations. We recognize that in order to make the problems posed realistic and applicable in advancing any field, we have to rely - in some part - on the integrity &amp; honesty of participants, and not just the threat or enforceability of verification &amp; legal stipulations. As participants, you should all be aware of the competition's [Code of Conduct][2] and recognize your contributions have the potential to make real impacts to healthcare. I hope we all share in that spirit of integrity.\n\n\n[1]: https://www.kaggle.com/WinningModelDocumentationGuidelines\n[2]: https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/rules",
          "votes": 3
        },
        {
          "id": 383177,
          "postDate": "2018-09-07T22:33:08.810Z",
          "content": "<p>Thanks @Julia Elliott.</p>",
          "rawMarkdown": "Thanks @Julia Elliott."
        }
      ]
    },
    {
      "id": 382640,
      "postDate": "2018-09-06T19:18:02.797Z",
      "content": "<p>I share Kha Vo's (<a href=\"/khahuras\">@khahuras</a>) curiosity about how the model files will be verified.  The rules say the host team will \"verify the performance of the uploaded models matches the Stage 2 submission file.\"  But given the vagaries of GPUs, no matter how many random seeds we set, it's unlikely that most results can be reproduced exactly, and it's not reasonable to expect competitors to produce code that results in exactly reproducible results.  It would be helpful for us to have a better idea of what procedure will be followed and what criteria will be applied to determine if performance \"matches.\"</p>",
      "rawMarkdown": "I share Kha Vo's (@khahuras) curiosity about how the model files will be verified.  The rules say the host team will \"verify the performance of the uploaded models matches the Stage 2 submission file.\"  But given the vagaries of GPUs, no matter how many random seeds we set, it's unlikely that most results can be reproduced exactly, and it's not reasonable to expect competitors to produce code that results in exactly reproducible results.  It would be helpful for us to have a better idea of what procedure will be followed and what criteria will be applied to determine if performance \"matches.\"",
      "votes": 1
    },
    {
      "id": 381095,
      "postDate": "2018-09-04T03:40:26.910Z",
      "content": "<p>Just to make it a little more clearer what Anouk is trying to say. You get a PA view when the xrays hit your back and the film is placed on your chest. Therefore, PA view can only be done in patients who can stand. AP view is the opposite. The film is in the back and the rays hit your chest.</p>\n\n<p>AP view is performed in bedridden patients who can not stand. Therefore it is possible that pneumonia is correlated with the AP view because patients who can't stand may probably have pneumonia.</p>",
      "rawMarkdown": "Just to make it a little more clearer what Anouk is trying to say. You get a PA view when the xrays hit your back and the film is placed on your chest. Therefore, PA view can only be done in patients who can stand. AP view is the opposite. The film is in the back and the rays hit your chest.\n\nAP view is performed in bedridden patients who can not stand. Therefore it is possible that pneumonia is correlated with the AP view because patients who can't stand may probably have pneumonia.",
      "replies": [
        {
          "id": 381940,
          "postDate": "2018-09-05T12:26:52.553Z",
          "content": "<blockquote>\n  <p>may probably have pneumonia</p>\n</blockquote>\n\n<p>but not necessarily... Unfortunately in the case of the data we are provided, that's always the case (and a leak) - but in general it's not an indication that the patient has pneumonia ... it just states that for some reasons the x-ray is done differently. </p>",
          "rawMarkdown": "&gt; may probably have pneumonia\n\nbut not necessarily... Unfortunately in the case of the data we are provided, that's always the case (and a leak) - but in general it's not an indication that the patient has pneumonia ... it just states that for some reasons the x-ray is done differently. "
        },
        {
          "id": 382268,
          "postDate": "2018-09-06T01:51:16.693Z",
          "content": "<p>Mihai, that is why I have said 'probably'.\nHere are some reasons a bed-ridden patient has more probability of having pneumonia:\n 1. Poor immune function because of being bedridden. \n 2. Normally the there are microscopic hair-like structures in the trachea called cilia which clear bacteria from trachea and push them upwards. These are impaired in bedridden patients.\n 3. Having an tube in your trachea for breathing is another source of pneumonia., etc.</p>",
          "rawMarkdown": "Mihai, that is why I have said 'probably'.\nHere are some reasons a bed-ridden patient has more probability of having pneumonia:\n 1. Poor immune function because of being bedridden. \n 2. Normally the there are microscopic hair-like structures in the trachea called cilia which clear bacteria from trachea and push them upwards. These are impaired in bedridden patients.\n 3. Having an tube in your trachea for breathing is another source of pneumonia., etc."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 379551,
      "author_name": "Anouk Stein, MD",
      "author_url": "",
      "post_date": "2018-08-31T14:56:13.163000",
      "content": "<p>John Zech has a well-written blog post about this issue:  <a href=\"https://medium.com/@jrzech/what-are-radiological-deep-learning-models-actually-learning-f97a546c5b98\">https://medium.com/@jrzech/what-are-radiological-deep-learning-models-actually-learning-f97a546c5b98</a> It is correct to say that an algorithm that concludes that portable xrays are more likely to show pathology does not add anything to medical knowledge. Mihai is spot on when he emphasizes the bounding boxes over the classification.</p>\n\n<p>I am a radiologist employed by MD.ai</p>",
      "votes": 6,
      "replies": [
        {
          "id": 379761,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2018-08-31T22:29:49.793000",
          "content": "<p>as much as I wanted to be true (that boxes are more important than classification) - it is not entirely true. The model metric is very sensitive if you put at least one box on a \"good\" x-ray. so having a very good classifier on top of box detection is very important, as classifier would be a very good proxy for thresholding when to show a box and not. I would even start working at this problem from classification perspective (maybe even use whole NIH dataset for that) and only later focus on detection.</p>\n\n<p>So yeah. knowing if the image is PA/AP will be very handy, as @Giulia has stated!</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 379762,
          "author_name": "Anouk Stein, MD",
          "author_url": "",
          "post_date": "2018-08-31T22:33:18.790000",
          "content": "<p>I see your point. Even though there will be bias, classification will still be very useful. It would help to look at  the false positive rate to make sure it's  similar on PA and AP. Also, maybe redacting the 'Portable' label on the films themselves would make a difference. You'd probably need to make a similar redaction box for some non portable films. It might not make a difference because there are other signs that the film is AP/portable ie. larger heart silhouette, often less lung inflation because it's harder for the patient to take a deep breath.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 379863,
          "author_name": "Mihai Cvasnievschi",
          "author_url": "",
          "post_date": "2018-09-01T05:04:33.067000",
          "content": "<p>The classification problem is pretty \"simple\" ...  See <a href=\"https://github.com/Azure/AzureChestXRay\">here</a> some recent results ;-)</p>\n\n<p>The segmentation one is the more complicated one. But with tweaking it can be solve using something like <a href=\"https://github.com/matterport/Mask_RCNN\">this</a>.</p>\n\n<p><a href=\"/raddar\">@raddar</a>. you said it yourself, \"if you put at least one box\"</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 380890,
      "author_name": "Harm Buisman",
      "author_url": "",
      "post_date": "2018-09-03T16:49:08.043000",
      "content": "<p>How is this a data leak? This information is available to the doctor when he takes the image as well.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 380905,
          "author_name": "Anouk Stein, MD",
          "author_url": "",
          "post_date": "2018-09-03T17:30:21.337000",
          "content": "<p>Good point!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 380936,
          "author_name": "jtlowery",
          "author_url": "",
          "post_date": "2018-09-03T18:44:54.400000",
          "content": "<p>Agreed. I guess from the perspective that the task should only consider the image this could be considered a leak. But this is a feature which is included in the DICOM metadata which has clearly been altered as part of the dataset preparation (fields removed or set to the same value throughout dataset) so it was clearly intended that this feature would be accessible to us (not a leak). That's not to say that it isn't an important feature -- it clearly is.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 381014,
          "author_name": "Henrique Mendonça",
          "author_url": "",
          "post_date": "2018-09-03T22:36:55.807000",
          "content": "<p>It means the data is biased. In its essence there's nothing wrong with the view point, but as Giulia pointed out in her kernel, it reveals the ill/critical state of the patient (who can't even stand up). The latter correlates with harsher illnesses like pneumonia.\nRecommend reading the kernel ;)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 379252,
      "author_name": "Mihai Cvasnievschi",
      "author_url": "",
      "post_date": "2018-08-31T04:54:27.647000",
      "content": "<p>Don't worry, there's actually a huge \"classification\" leak because the data is based on NIH and the test data from this competition is present in the training data from NIH.</p>\n\n<p>Keep in mind what does this competition tries to solve! It's not the classification, it's the bounding boxes!! </p>",
      "votes": 4,
      "replies": [
        {
          "id": 381697,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2018-09-05T02:52:39.043000",
          "content": "<p>The images are based on NIH.  The data - the labels and bounding boxes - are newly created by the host team.  There isn't a leak.  It's just based on a common set of images.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 381850,
          "author_name": "Mihai Cvasnievschi",
          "author_url": "",
          "post_date": "2018-09-05T09:59:23.297000",
          "content": "<p>Do you mean you've found all the images form NIH  that are part of the \"test\" set as mislabeled? If not, it means that the presence of pneumonia in the \"test\" set of images is already labeled by NIH set. An that's a leak.  (ex. pID 2118026e-ed2f-4d1e-baa2-0e9e1004f72c in Test and    00006473_001 in NIH db)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 381895,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2018-09-05T11:29:38.067000",
          "content": "<p>This data set does not use the NIH labels.  The criteria for whether or not something is labeled with \"pneumonia\" in this particular set is multi-tiered and most certainly does <em>not</em> match the criteria used in the labels you're looking at.</p>\n\n<p>How many radiologists provided the labels you're looking at?  How did they agree on a final diagnosis?  Doctor-doctor agreement for this problem is not exceedingly high, especially if low-grade cases are included.  Were they in your set?  Do you know what severity was assigned to the labels you're looking at?  The host here spent a great deal of effort with multiple passes, multiple doctors, and adjudication to ensure very high-quality labels.  If you're looking at one radiologist's labels (or worse, labels created without a radiologist) you're not looking at remotely the same set of labels used here.  Which is part of why competitors are encouraged to re-annotate the training set themselves with the help of more radiologists.  That's leaving the question of the bounding boxes aside entirely.</p>\n\n<p>Not only that, but no use of the NIH labels - or any other outside labels - is allowed for the test set.</p>\n\n<p>So a) the labels used here are not exposed, and b) if anyone is found to have used the NIH labels to label the test set (however the labels might disagree), they would be disqualified from earning prizes.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 381897,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2018-09-05T11:32:56.817000",
          "content": "<p>I would also encourage interested people to read about the annotation methods used for this competition <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723\">here</a>.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 381903,
          "author_name": "Mihai Cvasnievschi",
          "author_url": "",
          "post_date": "2018-09-05T11:43:47.557000",
          "content": "<p>I'm looking at the labels provided by NIH as per competition <a href=\"https://nihcc.app.box.com/v/ChestXray-NIHCC/folder/36938765345\">\"Original data source\"</a> link.</p>\n\n<p>One could train a classifier based on NIH original data and use the labels provide by NIHs to classify if pneumonia or not. Next, train a box/segmentation model only for \"pneumonia\" images - provide the classifier does a good job on \"test\" set, it would simplify the work of box/segmentation model.</p>\n\n<p>So what you are saying is that the original data source <strong>cannot</strong> be used for training a classifier because it leaks information from the \"Test\" set. </p>\n\n<p>If so, I want to see how that's going to be enforced at the end of the competition. I know this competition requires the \"model(s)\" and obviously the \"weights\" to generate the results. But the \"weights\" could have seen NIH data. How do you plan in evaluating that's not the case?</p>\n\n<p>Are you going to spend the time in re-training the models with the data provided in this competition? How are you going to ensure you get the same random initialization of the parameters, provided the models do not use any pre-trained models? Etc? Bottom line I don't see how you can ensure that winning model hasn't used the NIH data set and for that reason I strong believe that the classification part has a huge leak. As stated, unless most of the images selected to be in the \"Test\" had been mislabeled by NIH you have a leak.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 381924,
          "author_name": "Mihai Cvasnievschi",
          "author_url": "",
          "post_date": "2018-09-05T12:11:57.073000",
          "content": "<p>Pipeline</p>\n\n<p>Test data -&gt; Classifier Model\n-&gt; Pneumonia? \n-&gt; YES -&gt; Box/Segmentation Model (trained only on pneumonia data) -&gt; ensemble ... etc \n-&gt; NO - No Box</p>\n\n<p>If the Classifier Model is overfitted (memorized) NIH data the problem is simpler. So ... leak ...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 381977,
          "author_name": "drsxr",
          "author_url": "",
          "post_date": "2018-09-05T13:23:56.730000",
          "content": "<p>Mihai -\n   It is widely perceived by radiologists who practice machine learning as well as data scientists like Jeremy Howard that the NIH Chest X-ray 14 dataset was a \"first effort\" dataset, and therefore suffered from a older, and unfortunately substantially inaccurate, NLP classification system which substantially decreased the accuracy of labels, lower than I think many of us would like to have seen.  ChestXRay14 was released to a great deal of fanfare, along with a sensational statement about CheXNet, which is probably why the RSNA has chosen this challenge to make a more fair assessment.</p>\n\n<p>Here are <a href=\"https://n2value.com/blog/chexnet-a-brief-evaluation/\">my observations on the dataset</a>.  This <a href=\"https://n2value.com/blog/are-computers-better-than-doctors-will-the-computer-see-you-now-what-we-learnt-from-the-chexnet-paper-for-pneumonia-diagnosis/\">online meetup discussed important points of the dataset</a>.  And Luke Oakden Rayner <a href=\"https://lukeoakdenrayner.wordpress.com/2017/12/18/the-chestxray14-dataset-problems/\">has written about it extensively, as well as providing his own estimates of the accuracy of the dataset</a>.  A SIIM-centric working group spent time on this dataset to re-label pneumothoraces (data I believe is being presented in a few days at C-MIMI) in large because of these inaccuracies.  I suspect most professional ML teams who are using the CXR-14 dataset in production have needed to hire radiologists to re-label.  </p>\n\n<p>If you're looking to use CXR-14 as an opportunity for transfer learning, its not entirely unreasonable, try it!  But be very clear about the underlying strata you are working with.  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 382047,
          "author_name": "Mihai Cvasnievschi",
          "author_url": "",
          "post_date": "2018-09-05T15:53:09.320000",
          "content": "<p>@drsxr - excellent post. </p>\n\n<p>Now here's an idea, let's check how many from \"Train\" set are mislabeled in NIH set :-) Some insights can be obtained from that.  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 382510,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2018-09-06T13:56:32.990000",
          "content": "<p><a href=\"/philculliton\">@philculliton</a>. Please clarify on how does the host disqualify a team. At this stage, I don't know what exactly a \"model\" is, which is required to be uploaded after stage 1. If the host receive a code,</p>\n\n<p>1) will the host re-run the whole model (or, in other words, Machine Learning code) on the training set of this competition, then use that trained model predict on the test set of stage 1 to confirm that the result on the leaderboard is identical, then, use that model to predict the test set of stage 2? Or</p>\n\n<p>2) the host only require the trained model uploaded by the team (or with code also for verification that \"there is some model\", but the host will NOT re-run the code on the training set of this competition), then direct use that model to predict the test set of stage 1 to confirm?</p>\n\n<p>If case 2) is right, then I'm afraid the models may be trained with some sort of leak information obtained from NIH. There's no way the host know whether the team use any sort of external leak data or not.</p>\n\n<p>And there is other issue. Even if the host can manage to do the verification described in 1) for the the few top teams, the other teams will be exempted as the result of the burden of this check, and they will get Kaggle medals with some leak information. </p>\n\n<p>This competition therefore, has too many problems that concern, I believe, not only me.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 382569,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2018-09-06T16:45:22.590000",
          "content": "<p>@Kha Vo - Please review the rules and timeline for details on host verification. You are free to choose not to participate in the competition, if this is a design that does not appeal to you.</p>",
          "votes": -3,
          "replies": []
        },
        {
          "id": 382577,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-09-06T17:05:35.400000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 382585,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2018-09-06T17:20:08.903000",
          "content": "<p>I am asking for the sake of this competition, more especially, the fairness. I am a serious competitor and not f*<em>*</em>* around asking nonsense. There’s no way to stop a lot of cheatings around.  I feel your comment is a little bit subjectively sensitive.</p>\n\n<p>The rules did not clearly state how the host verify the integrity of the uploaded models.  That’s why I asked. How do you know if a model is trained by an external annotated dataset that partly overlaps with the test set?  How do you handle that and make the other honest teams feel relieved to participate?</p>\n\n<p>You cannot definitely say that we expect there’s no cheaters. There will be definitely not a few, but a lot. Only using some small leak may heavily affect the leaderboard. Can you imagine the consequence if some of the top teams all cheat? My question is a constructive question. </p>\n\n<p>There was a recent Kaggle competition with leak discovered during the competition, and exploited relentlessly by almost all teams.  </p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 382591,
          "author_name": "Anouk Stein, MD",
          "author_url": "",
          "post_date": "2018-09-06T17:26:36.263000",
          "content": "<p>No offense or trivialization of your question was intended.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 382626,
          "author_name": "Mihai Cvasnievschi",
          "author_url": "",
          "post_date": "2018-09-06T18:23:25.060000",
          "content": "<p>@Kha Vo - if you'll read the info from the links posted by @drsxr you'll find out that apparently many images from the original NIH data set aren't properly labeled. That's because they've used an NLP model to read the medical file associated with the image and auto-labeled the images. The set of images provided by RSNA have be correctly labeled using a team of radiologists. </p>\n\n<p>I gave it a bit more thought, using a clarifies over-fit on NIH set (even with correct labels) would fail on a Stage 2 - provided that there is a stage 2 and that the images are not part of the NIH set. That would filter out \"cheaters\" that exploit the leak. </p>\n\n<p>Ideally the Stage 2 images should be from NIH <strong>Test</strong> set and not Train set :-) Ideally should be images that weren't publicly available ...</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 382735,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2018-09-07T02:17:53.150000",
          "content": "<p>@Anouk. My last comment was not for you.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 382743,
          "author_name": "Anouk Stein, MD",
          "author_url": "",
          "post_date": "2018-09-07T02:32:35.623000",
          "content": "<p>@Kha Vo 😀</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 383166,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2018-09-07T21:17:09.033000",
          "content": "<p>I do not intend to be dismissive. Your questions are valid. You should expect that models will be evaluated both (A) to ensure the Stage 2 predictions were generated from the Stage 1 Uploaded model, and (B) to confirm compliance with restrictions around eligible data usage. This could mean going to the extent of running the whole model. The expectations for <a href=\"https://www.kaggle.com/WinningModelDocumentationGuidelines\">winners' model documentation</a> should help to clarify the extent of what could be investigated for those in prize-standing. Ultimately, verification is in the hands of the host, and the extent of enforcement will be at their discretion.</p>\n\n<p>We (Kaggle and its hosts) bring challenges to the platform because they represent real-world problems being solved by research or commercial organizations. We recognize that in order to make the problems posed realistic and applicable in advancing any field, we have to rely - in some part - on the integrity &amp; honesty of participants, and not just the threat or enforceability of verification &amp; legal stipulations. As participants, you should all be aware of the competition's <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/rules\">Code of Conduct</a> and recognize your contributions have the potential to make real impacts to healthcare. I hope we all share in that spirit of integrity.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 383177,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-09-07T22:33:08.810000",
          "content": "<p>Thanks @Julia Elliott.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 382640,
      "author_name": "Andy Harless",
      "author_url": "",
      "post_date": "2018-09-06T19:18:02.797000",
      "content": "<p>I share Kha Vo's (<a href=\"/khahuras\">@khahuras</a>) curiosity about how the model files will be verified.  The rules say the host team will \"verify the performance of the uploaded models matches the Stage 2 submission file.\"  But given the vagaries of GPUs, no matter how many random seeds we set, it's unlikely that most results can be reproduced exactly, and it's not reasonable to expect competitors to produce code that results in exactly reproducible results.  It would be helpful for us to have a better idea of what procedure will be followed and what criteria will be applied to determine if performance \"matches.\"</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 381095,
      "author_name": "Pranay Aryal",
      "author_url": "",
      "post_date": "2018-09-04T03:40:26.910000",
      "content": "<p>Just to make it a little more clearer what Anouk is trying to say. You get a PA view when the xrays hit your back and the film is placed on your chest. Therefore, PA view can only be done in patients who can stand. AP view is the opposite. The film is in the back and the rays hit your chest.</p>\n\n<p>AP view is performed in bedridden patients who can not stand. Therefore it is possible that pneumonia is correlated with the AP view because patients who can't stand may probably have pneumonia.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 381940,
          "author_name": "Mihai Cvasnievschi",
          "author_url": "",
          "post_date": "2018-09-05T12:26:52.553000",
          "content": "<blockquote>\n  <p>may probably have pneumonia</p>\n</blockquote>\n\n<p>but not necessarily... Unfortunately in the case of the data we are provided, that's always the case (and a leak) - but in general it's not an indication that the patient has pneumonia ... it just states that for some reasons the x-ray is done differently. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 382268,
          "author_name": "Pranay Aryal",
          "author_url": "",
          "post_date": "2018-09-06T01:51:16.693000",
          "content": "<p>Mihai, that is why I have said 'probably'.\nHere are some reasons a bed-ridden patient has more probability of having pneumonia:\n 1. Poor immune function because of being bedridden. \n 2. Normally the there are microscopic hair-like structures in the trachea called cilia which clear bacteria from trachea and push them upwards. These are impaired in bedridden patients.\n 3. Having an tube in your trachea for breathing is another source of pneumonia., etc.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "379091": "The \"view position\" of the radiograph image appears to be really important in terms of target split. This information is extracted from the metadata header of the radiograph images. \n\nFor ViewPosition=='AP' the Target [0,1] split is 0.62 - 0.37\nFor ViewPosition=='PA' the Target [0,1] split is 0.91 - 0.09\n\nAP means Anterior-Posterior, whereas PA means Posterior-Anterior. This [webpage](https://www.med-ed.virginia.edu/courses/rad/cxr/technique3chest.html) explains that \"Whenever possible the patient should be imaged in an upright PA position.  AP views are less useful and should be reserved for very ill patients who cannot stand erect\". One way to interpret this target unbalance is that patients that are imaged in an AP position are those that are more ill, and therefore more likely to have contracted pneumonia. Note that the absolute split between AP and PA images is about 50-50, so the above consideration is extremely significant. \n\nYou can find more details at this kernel: https://www.kaggle.com/giuliasavorgnan/start-here-beginner-intro-to-lung-opacity-s1",
    "379551": "John Zech has a well-written blog post about this issue:  https://medium.com/@jrzech/what-are-radiological-deep-learning-models-actually-learning-f97a546c5b98 It is correct to say that an algorithm that concludes that portable xrays are more likely to show pathology does not add anything to medical knowledge. Mihai is spot on when he emphasizes the bounding boxes over the classification.\n\nI am a radiologist employed by MD.ai",
    "380890": "How is this a data leak? This information is available to the doctor when he takes the image as well.",
    "379252": "Don't worry, there's actually a huge \"classification\" leak because the data is based on NIH and the test data from this competition is present in the training data from NIH.\n\nKeep in mind what does this competition tries to solve! It's not the classification, it's the bounding boxes!! ",
    "382640": "I share Kha Vo's (@khahuras) curiosity about how the model files will be verified.  The rules say the host team will \"verify the performance of the uploaded models matches the Stage 2 submission file.\"  But given the vagaries of GPUs, no matter how many random seeds we set, it's unlikely that most results can be reproduced exactly, and it's not reasonable to expect competitors to produce code that results in exactly reproducible results.  It would be helpful for us to have a better idea of what procedure will be followed and what criteria will be applied to determine if performance \"matches.\"",
    "381095": "Just to make it a little more clearer what Anouk is trying to say. You get a PA view when the xrays hit your back and the film is placed on your chest. Therefore, PA view can only be done in patients who can stand. AP view is the opposite. The film is in the back and the rays hit your chest.\n\nAP view is performed in bedridden patients who can not stand. Therefore it is possible that pneumonia is correlated with the AP view because patients who can't stand may probably have pneumonia."
  }
}