{
  "id": 223025,
  "title": "Remapping the data back to NIH Chest X-ray dataset",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/223025",
  "author_name": "",
  "post_date": "2021-03-02T06:21:04.303084500Z",
  "votes": 14,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Like mentioned in the acknowledgment section, all of the X-rays here originally come from this <a href=\"https://www.kaggle.com/nih-chest-xrays/data\" target=\"_blank\">NIH Chest X-ray14 dataset</a> and have been reannotated with regard to the catheters. This original dataset comes with various other bits of information that are not provided to us like age, gender, viewpoint, etc. </p>\n<p>Like <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221808\" target=\"_blank\">some have pointed out</a> we can map the RANZCR images to the NIH dataset via some simple image hashing with fairly serviceable accuracy. If we are able to do this then we can do lots of interesting things with the associated data. Using the tabular data that can be extracted here alone I could train a simple NN model that achieves AUCs of </p>\n<p><code>[0.5504 0.6776 0.8025 0.6282 0.6336 0.7311 0.7172 0.5647 0.5686 0.546 0.6703]</code></p>\n<p>just with </p>\n<pre><code>['Nodule', 'Consolidation', 'Effusion', 'Cardiomegaly', 'Infiltration',\n       'Hernia', 'Pleural_Thickening', 'Emphysema', 'Pneumothorax',\n       'Pneumonia', 'Fibrosis', 'No Finding', 'Edema', 'Atelectasis', 'Mass', 'Follow-up #',\n       'Patient Age', 'Patient Gender', 'View Position', 'OriginalImage[Width',\n       'Height]', 'OriginalImagePixelSpacing[x', 'y]']\n</code></pre>\n<p>So there is not great signal but there is clearly some prior that can be added. It is somewhat questionable if this is a valid solution though given that on future X-rays we wont necessarily be guaranteed to have many of these annotations. The finding of pneumothorax and other conditions were doctor annotated and put into records so this technique would not extend out of this dataset. </p>\n<p>Even beyond this, being able to recover out the patient ID we can map the various images together and do some sort of comparison or apply some prior. On the train set we have the patient IDs but on the test set we do not unless we do this mapping technique. I am not sure exactly how that can be utilized but there is potentially some way you use multiple predictions for a single patient and compare their normal vs abnormal or know more confidently which catheter types are present. </p>",
  "messages": [
    {
      "id": "1222786",
      "postDate": "03/02/2021 06:21:04",
      "content": "<p>Like mentioned in the acknowledgment section, all of the X-rays here originally come from this <a href=\"https://www.kaggle.com/nih-chest-xrays/data\" target=\"_blank\">NIH Chest X-ray14 dataset</a> and have been reannotated with regard to the catheters. This original dataset comes with various other bits of information that are not provided to us like age, gender, viewpoint, etc. </p>\n<p>Like <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221808\" target=\"_blank\">some have pointed out</a> we can map the RANZCR images to the NIH dataset via some simple image hashing with fairly serviceable accuracy. If we are able to do this then we can do lots of interesting things with the associated data. Using the tabular data that can be extracted here alone I could train a simple NN model that achieves AUCs of </p>\n<p><code>[0.5504 0.6776 0.8025 0.6282 0.6336 0.7311 0.7172 0.5647 0.5686 0.546 0.6703]</code></p>\n<p>just with </p>\n<pre><code>['Nodule', 'Consolidation', 'Effusion', 'Cardiomegaly', 'Infiltration',\n       'Hernia', 'Pleural_Thickening', 'Emphysema', 'Pneumothorax',\n       'Pneumonia', 'Fibrosis', 'No Finding', 'Edema', 'Atelectasis', 'Mass', 'Follow-up #',\n       'Patient Age', 'Patient Gender', 'View Position', 'OriginalImage[Width',\n       'Height]', 'OriginalImagePixelSpacing[x', 'y]']\n</code></pre>\n<p>So there is not great signal but there is clearly some prior that can be added. It is somewhat questionable if this is a valid solution though given that on future X-rays we wont necessarily be guaranteed to have many of these annotations. The finding of pneumothorax and other conditions were doctor annotated and put into records so this technique would not extend out of this dataset. </p>\n<p>Even beyond this, being able to recover out the patient ID we can map the various images together and do some sort of comparison or apply some prior. On the train set we have the patient IDs but on the test set we do not unless we do this mapping technique. I am not sure exactly how that can be utilized but there is potentially some way you use multiple predictions for a single patient and compare their normal vs abnormal or know more confidently which catheter types are present. </p>",
      "rawMarkdown": "Like mentioned in the acknowledgment section, all of the X-rays here originally come from this [NIH Chest X-ray14 dataset](https://www.kaggle.com/nih-chest-xrays/data) and have been reannotated with regard to the catheters. This original dataset comes with various other bits of information that are not provided to us like age, gender, viewpoint, etc. \n\nLike [some have pointed out](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221808) we can map the RANZCR images to the NIH dataset via some simple image hashing with fairly serviceable accuracy. If we are able to do this then we can do lots of interesting things with the associated data. Using the tabular data that can be extracted here alone I could train a simple NN model that achieves AUCs of \n\n`[0.5504 0.6776 0.8025 0.6282 0.6336 0.7311 0.7172 0.5647 0.5686 0.546 0.6703]`\n\njust with \n\n```\n['Nodule', 'Consolidation', 'Effusion', 'Cardiomegaly', 'Infiltration',\n       'Hernia', 'Pleural_Thickening', 'Emphysema', 'Pneumothorax',\n       'Pneumonia', 'Fibrosis', 'No Finding', 'Edema', 'Atelectasis', 'Mass', 'Follow-up #',\n       'Patient Age', 'Patient Gender', 'View Position', 'OriginalImage[Width',\n       'Height]', 'OriginalImagePixelSpacing[x', 'y]']\n```\n\nSo there is not great signal but there is clearly some prior that can be added. It is somewhat questionable if this is a valid solution though given that on future X-rays we wont necessarily be guaranteed to have many of these annotations. The finding of pneumothorax and other conditions were doctor annotated and put into records so this technique would not extend out of this dataset. \n\nEven beyond this, being able to recover out the patient ID we can map the various images together and do some sort of comparison or apply some prior. On the train set we have the patient IDs but on the test set we do not unless we do this mapping technique. I am not sure exactly how that can be utilized but there is potentially some way you use multiple predictions for a single patient and compare their normal vs abnormal or know more confidently which catheter types are present.",
      "votes": null
    },
    {
      "id": "1222792",
      "postDate": "03/02/2021 06:25:32",
      "content": "<p>Also just to note it is possible to also apply this to the hidden test set fairly easily by just loading the hashes for the images of the full chest14 dataset and doing a simple comparison and then lookup in the csv. </p>",
      "rawMarkdown": "Also just to note it is possible to also apply this to the hidden test set fairly easily by just loading the hashes for the images of the full chest14 dataset and doing a simple comparison and then lookup in the csv.",
      "votes": null
    },
    {
      "id": "1222793",
      "postDate": "03/02/2021 06:26:37",
      "content": "<p>Initial results didnt show anything when I combined my pretrained image based model along with the last couple layers retrained with the additional features, but maybe there is some angle to eek more out. </p>",
      "rawMarkdown": "Initial results didnt show anything when I combined my pretrained image based model along with the last couple layers retrained with the additional features, but maybe there is some angle to eek more out.",
      "votes": null
    },
    {
      "id": "1226904",
      "postDate": "03/05/2021 02:14:07",
      "content": "<p>I like it. As you have likely suspected, patients with numerous CXRs will have mostly have the same lines, in very similar projections, but with one of these lines in an abnormal position (being rejigged before a repeat scan)</p>",
      "rawMarkdown": "I like it. As you have likely suspected, patients with numerous CXRs will have mostly have the same lines, in very similar projections, but with one of these lines in an abnormal position (being rejigged before a repeat scan)",
      "votes": null
    },
    {
      "id": "1226942",
      "postDate": "03/05/2021 03:36:53",
      "content": "<p>Yeah, I am not entirely sure what to do once the patient ID has been recovered though. Like we have 25 images for a single patient, they show various characteristics when predicting on the other 24 xrays. What do we do with that information for the current prediction?</p>",
      "rawMarkdown": "Yeah, I am not entirely sure what to do once the patient ID has been recovered though. Like we have 25 images for a single patient, they show various characteristics when predicting on the other 24 xrays. What do we do with that information for the current prediction?",
      "votes": null
    },
    {
      "id": "1228669",
      "postDate": "03/06/2021 16:29:06",
      "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a>   \"on the test set we do not unless we do this mapping technique.\"</p>\n<p>there is a loophole in the design of the challenge. it is a mistake to use Chest X-ray14 dataset as test set.<br>\nhere is how it work</p>\n<p>1) exploit meta label in generating pseudo label of Chest X-ray14 : label = net1(chest14 image, meta)<br>\n2) now train a classifier to reproduce  pseudo label based on image: label = net2(chest14 image)<br>\n3) for kaggle test, use label = net2(test image) . now test image =  chest14 image, meta information has been leaked via pseudo label.</p>\n<p>in fact, you can train a network to predict (or rather to memorized) the meta information. this works as long as there is any test image = chest14 image</p>\n<hr>\n<p>\"Also just to note it is possible to also apply this to the hidden test set fairly easily by just loading the hashes for the images of the full chest14 dataset and doing a simple comparison and then lookup in the csv.\"</p>\n<p>this is another shortcut</p>\n<hr>\n<p>the organizer should state explicitly such practices are prohibited.</p>",
      "rawMarkdown": "ryches   \"on the test set we do not unless we do this mapping technique.\"\n\nthere is a loophole in the design of the challenge. it is a mistake to use Chest X-ray14 dataset as test set.\nhere is how it work\n\n1) exploit meta label in generating pseudo label of Chest X-ray14 : label = net1(chest14 image, meta)\n2) now train a classifier to reproduce  pseudo label based on image: label = net2(chest14 image)\n3) for kaggle test, use label = net2(test image) . now test image =  chest14 image, meta information has been leaked via pseudo label.\n\nin fact, you can train a network to predict (or rather to memorized) the meta information. this works as long as there is any test image = chest14 image\n\n---\n\n\"Also just to note it is possible to also apply this to the hidden test set fairly easily by just loading the hashes for the images of the full chest14 dataset and doing a simple comparison and then lookup in the csv.\"\n\nthis is another shortcut\n\n\n---\n\nthe organizer should state explicitly such practices are prohibited.",
      "votes": null
    },
    {
      "id": "1228673",
      "postDate": "03/06/2021 16:31:28",
      "content": "<p>Yes, that is another possible approach, using our models prediction as a sort of hashing function on its own. but one issue is that the nih dataset has already been resized to be 1024x1024  so we need to match the resizing procedure that they used or the results will not exactly match up. </p>",
      "rawMarkdown": "Yes, that is another possible approach, using our models prediction as a sort of hashing function on its own. but one issue is that the nih dataset has already been resized to be 1024x1024  so we need to match the resizing procedure that they used or the results will not exactly match up.",
      "votes": null
    },
    {
      "id": "1228833",
      "postDate": "03/06/2021 20:18:07",
      "content": "<p>I also feel like this should be prohibited because it is giving us labels and information they dont really want us operating on but I dont see how it is much different from allowing external annotation like some seem to be exploring here. </p>",
      "rawMarkdown": "I also feel like this should be prohibited because it is giving us labels and information they dont really want us operating on but I dont see how it is much different from allowing external annotation like some seem to be exploring here.",
      "votes": null
    },
    {
      "id": "1228975",
      "postDate": "03/06/2021 23:53:40",
      "content": "<p>\"resizing procedure that they used or the results will not exactly match up.\"<br>\nfor this challenge, i think resizing all images to say 480x480 and applying python hashing may do the trick. </p>\n<p>i did not try for this competition, but this is what i did for others (cassava, tpu flower, etc) and i can find duplicate:</p>\n<ol>\n<li>resize to some small size, e.g. 24x24</li>\n<li>the values 24x24 gray values is already a hash. you can find duplicates by l2 difference (e.g. pixel difference within a certain threshold)</li>\n<li>if 24x24 is too small, i would try 48x48. the hash will be subset of 49x49 pixels for fast matching.</li>\n</ol>",
      "rawMarkdown": "\"resizing procedure that they used or the results will not exactly match up.\"\nfor this challenge, i think resizing all images to say 480x480 and applying python hashing may do the trick. \n\n\ni did not try for this competition, but this is what i did for others (cassava, tpu flower, etc) and i can find duplicate:\n\n1. resize to some small size, e.g. 24x24\n2. the values 24x24 gray values is already a hash. you can find duplicates by l2 difference (e.g. pixel difference within a certain threshold)\n3. if 24x24 is too small, i would try 48x48. the hash will be subset of 49x49 pixels for fast matching.",
      "votes": null
    },
    {
      "id": "1228982",
      "postDate": "03/07/2021 00:16:51",
      "content": "<p>I applied the imagehash library like mentioned in other threads. I believe it does that operation internally already</p>",
      "rawMarkdown": "I applied the imagehash library like mentioned in other threads. I believe it does that operation internally already",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1222792,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "03/02/2021 06:25:32",
      "content": "<p>Also just to note it is possible to also apply this to the hidden test set fairly easily by just loading the hashes for the images of the full chest14 dataset and doing a simple comparison and then lookup in the csv. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1222793,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "03/02/2021 06:26:37",
      "content": "<p>Initial results didnt show anything when I combined my pretrained image based model along with the last couple layers retrained with the additional features, but maybe there is some angle to eek more out. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1226904,
      "author_name": "reubenschmidt",
      "author_url": "",
      "post_date": "03/05/2021 02:14:07",
      "content": "<p>I like it. As you have likely suspected, patients with numerous CXRs will have mostly have the same lines, in very similar projections, but with one of these lines in an abnormal position (being rejigged before a repeat scan)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1226942,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "03/05/2021 03:36:53",
          "content": "<p>Yeah, I am not entirely sure what to do once the patient ID has been recovered though. Like we have 25 images for a single patient, they show various characteristics when predicting on the other 24 xrays. What do we do with that information for the current prediction?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1228669,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/06/2021 16:29:06",
      "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a>   \"on the test set we do not unless we do this mapping technique.\"</p>\n<p>there is a loophole in the design of the challenge. it is a mistake to use Chest X-ray14 dataset as test set.<br>\nhere is how it work</p>\n<p>1) exploit meta label in generating pseudo label of Chest X-ray14 : label = net1(chest14 image, meta)<br>\n2) now train a classifier to reproduce  pseudo label based on image: label = net2(chest14 image)<br>\n3) for kaggle test, use label = net2(test image) . now test image =  chest14 image, meta information has been leaked via pseudo label.</p>\n<p>in fact, you can train a network to predict (or rather to memorized) the meta information. this works as long as there is any test image = chest14 image</p>\n<hr>\n<p>\"Also just to note it is possible to also apply this to the hidden test set fairly easily by just loading the hashes for the images of the full chest14 dataset and doing a simple comparison and then lookup in the csv.\"</p>\n<p>this is another shortcut</p>\n<hr>\n<p>the organizer should state explicitly such practices are prohibited.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1228673,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "03/06/2021 16:31:28",
          "content": "<p>Yes, that is another possible approach, using our models prediction as a sort of hashing function on its own. but one issue is that the nih dataset has already been resized to be 1024x1024  so we need to match the resizing procedure that they used or the results will not exactly match up. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1228833,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "03/06/2021 20:18:07",
          "content": "<p>I also feel like this should be prohibited because it is giving us labels and information they dont really want us operating on but I dont see how it is much different from allowing external annotation like some seem to be exploring here. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1228975,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/06/2021 23:53:40",
          "content": "<p>\"resizing procedure that they used or the results will not exactly match up.\"<br>\nfor this challenge, i think resizing all images to say 480x480 and applying python hashing may do the trick. </p>\n<p>i did not try for this competition, but this is what i did for others (cassava, tpu flower, etc) and i can find duplicate:</p>\n<ol>\n<li>resize to some small size, e.g. 24x24</li>\n<li>the values 24x24 gray values is already a hash. you can find duplicates by l2 difference (e.g. pixel difference within a certain threshold)</li>\n<li>if 24x24 is too small, i would try 48x48. the hash will be subset of 49x49 pixels for fast matching.</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1228982,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "03/07/2021 00:16:51",
          "content": "<p>I applied the imagehash library like mentioned in other threads. I believe it does that operation internally already</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1222786": "Like mentioned in the acknowledgment section, all of the X-rays here originally come from this [NIH Chest X-ray14 dataset](https://www.kaggle.com/nih-chest-xrays/data) and have been reannotated with regard to the catheters. This original dataset comes with various other bits of information that are not provided to us like age, gender, viewpoint, etc. \n\nLike [some have pointed out](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221808) we can map the RANZCR images to the NIH dataset via some simple image hashing with fairly serviceable accuracy. If we are able to do this then we can do lots of interesting things with the associated data. Using the tabular data that can be extracted here alone I could train a simple NN model that achieves AUCs of \n\n`[0.5504 0.6776 0.8025 0.6282 0.6336 0.7311 0.7172 0.5647 0.5686 0.546 0.6703]`\n\njust with \n\n```\n['Nodule', 'Consolidation', 'Effusion', 'Cardiomegaly', 'Infiltration',\n       'Hernia', 'Pleural_Thickening', 'Emphysema', 'Pneumothorax',\n       'Pneumonia', 'Fibrosis', 'No Finding', 'Edema', 'Atelectasis', 'Mass', 'Follow-up #',\n       'Patient Age', 'Patient Gender', 'View Position', 'OriginalImage[Width',\n       'Height]', 'OriginalImagePixelSpacing[x', 'y]']\n```\n\nSo there is not great signal but there is clearly some prior that can be added. It is somewhat questionable if this is a valid solution though given that on future X-rays we wont necessarily be guaranteed to have many of these annotations. The finding of pneumothorax and other conditions were doctor annotated and put into records so this technique would not extend out of this dataset. \n\nEven beyond this, being able to recover out the patient ID we can map the various images together and do some sort of comparison or apply some prior. On the train set we have the patient IDs but on the test set we do not unless we do this mapping technique. I am not sure exactly how that can be utilized but there is potentially some way you use multiple predictions for a single patient and compare their normal vs abnormal or know more confidently which catheter types are present.",
    "1222792": "Also just to note it is possible to also apply this to the hidden test set fairly easily by just loading the hashes for the images of the full chest14 dataset and doing a simple comparison and then lookup in the csv.",
    "1222793": "Initial results didnt show anything when I combined my pretrained image based model along with the last couple layers retrained with the additional features, but maybe there is some angle to eek more out.",
    "1226904": "I like it. As you have likely suspected, patients with numerous CXRs will have mostly have the same lines, in very similar projections, but with one of these lines in an abnormal position (being rejigged before a repeat scan)",
    "1226942": "Yeah, I am not entirely sure what to do once the patient ID has been recovered though. Like we have 25 images for a single patient, they show various characteristics when predicting on the other 24 xrays. What do we do with that information for the current prediction?",
    "1228669": "ryches   \"on the test set we do not unless we do this mapping technique.\"\n\nthere is a loophole in the design of the challenge. it is a mistake to use Chest X-ray14 dataset as test set.\nhere is how it work\n\n1) exploit meta label in generating pseudo label of Chest X-ray14 : label = net1(chest14 image, meta)\n2) now train a classifier to reproduce  pseudo label based on image: label = net2(chest14 image)\n3) for kaggle test, use label = net2(test image) . now test image =  chest14 image, meta information has been leaked via pseudo label.\n\nin fact, you can train a network to predict (or rather to memorized) the meta information. this works as long as there is any test image = chest14 image\n\n---\n\n\"Also just to note it is possible to also apply this to the hidden test set fairly easily by just loading the hashes for the images of the full chest14 dataset and doing a simple comparison and then lookup in the csv.\"\n\nthis is another shortcut\n\n\n---\n\nthe organizer should state explicitly such practices are prohibited.",
    "1228673": "Yes, that is another possible approach, using our models prediction as a sort of hashing function on its own. but one issue is that the nih dataset has already been resized to be 1024x1024  so we need to match the resizing procedure that they used or the results will not exactly match up.",
    "1228833": "I also feel like this should be prohibited because it is giving us labels and information they dont really want us operating on but I dont see how it is much different from allowing external annotation like some seem to be exploring here.",
    "1228975": "\"resizing procedure that they used or the results will not exactly match up.\"\nfor this challenge, i think resizing all images to say 480x480 and applying python hashing may do the trick. \n\n\ni did not try for this competition, but this is what i did for others (cassava, tpu flower, etc) and i can find duplicate:\n\n1. resize to some small size, e.g. 24x24\n2. the values 24x24 gray values is already a hash. you can find duplicates by l2 difference (e.g. pixel difference within a certain threshold)\n3. if 24x24 is too small, i would try 48x48. the hash will be subset of 49x49 pixels for fast matching.",
    "1228982": "I applied the imagehash library like mentioned in other threads. I believe it does that operation internally already"
  },
  "source": "meta"
}