{
  "id": 84964,
  "title": "There are duplicates!",
  "url": "/competitions/histopathologic-cancer-detection/discussion/84964",
  "author_name": "",
  "post_date": "2019-03-20T20:42:07.791000600Z",
  "votes": 9,
  "comment_count": 22,
  "views": 0,
  "content": "<p><code>\nlabel                             1\nwsi      camelyon16_train_tumor_007\nName: c448ff7117b8382151d115fb54a63987e0fb5c0d, dtype: object to label                             1\nwsi      camelyon16_train_tumor_006\nName: d547e4b7204266e4412887ad82ce59bae5852ba5, dtype: object\n</code>\n<img src=\"https://i.postimg.cc/GpDTwnyS/download.png\" alt=\"\"></p>\n\n<p>This could explain why using WSI makes a difference!\nAlso, these duplicates images can inspire us about augmentations.\nWhat do you think?</p>",
  "messages": [
    {
      "id": "495220",
      "postDate": "03/20/2019 20:42:07",
      "content": "<p><code>\nlabel                             1\nwsi      camelyon16_train_tumor_007\nName: c448ff7117b8382151d115fb54a63987e0fb5c0d, dtype: object to label                             1\nwsi      camelyon16_train_tumor_006\nName: d547e4b7204266e4412887ad82ce59bae5852ba5, dtype: object\n</code>\n<img src=\"https://i.postimg.cc/GpDTwnyS/download.png\" alt=\"\"></p>\n\n<p>This could explain why using WSI makes a difference!\nAlso, these duplicates images can inspire us about augmentations.\nWhat do you think?</p>",
      "rawMarkdown": "```\nlabel                             1\nwsi      camelyon16_train_tumor_007\nName: c448ff7117b8382151d115fb54a63987e0fb5c0d, dtype: object to label                             1\nwsi      camelyon16_train_tumor_006\nName: d547e4b7204266e4412887ad82ce59bae5852ba5, dtype: object\n```\n![](https://i.postimg.cc/GpDTwnyS/download.png)\n\nThis could explain why using WSI makes a difference!\nAlso, these duplicates images can inspire us about augmentations.\nWhat do you think?",
      "votes": null
    },
    {
      "id": "495731",
      "postDate": "03/21/2019 13:37:39",
      "content": "<p>Yes you are right, if we don't split by WSI the model overfits to these duplicates</p>",
      "rawMarkdown": "Yes you are right, if we don't split by WSI the model overfits to these duplicates",
      "votes": null
    },
    {
      "id": "495767",
      "postDate": "03/21/2019 14:37:24",
      "content": "<p>The point is, even if we split according to wsi, we still cannot make a stable cv.</p>",
      "rawMarkdown": "The point is, even if we split according to wsi, we still cannot make a stable cv.",
      "votes": null
    },
    {
      "id": "495886",
      "postDate": "03/21/2019 16:33:03",
      "content": "<p>Good find! How did you discover this? I wonder how many more cross WSI duplicates there are.</p>",
      "rawMarkdown": "Good find! How did you discover this? I wonder how many more cross WSI duplicates there are.",
      "votes": null
    },
    {
      "id": "496260",
      "postDate": "03/22/2019 02:25:32",
      "content": "<p>Your trained model can tell you that! Looking at the probability I get from my model's prediction, I found some interesting patterns there. The playground competition is very boring compared to the <code>Featured</code> ones - so having some fun digging some golds is not a bad idea. (chat on twitter?)</p>\n\n<p>Also, do you see the LB 1st there? with a huge increase in LB score than the rest of us. I bet he understands the data a lot better than we do. </p>",
      "rawMarkdown": "Your trained model can tell you that! Looking at the probability I get from my model's prediction, I found some interesting patterns there. The playground competition is very boring compared to the `Featured` ones - so having some fun digging some golds is not a bad idea. (chat on twitter?)\n\nAlso, do you see the LB 1st there? with a huge increase in LB score than the rest of us. I bet he understands the data a lot better than we do.",
      "votes": null
    },
    {
      "id": "496429",
      "postDate": "03/22/2019 07:30:25",
      "content": "<p>Domain knowledge is definitely useful here. Also, understanding how the WSIs were extracted could be helpful.</p>\n\n<p>I extracted some feature vectors from the images and compared their distances. By just comparing the first 1000 training images to the rest, I found two duplicate sets. However, this is a very slow way of searching these and I'm probably not going to go through the whole set.</p>\n\n<p><code>\na5e032f36bcf57c281c0d6c482ccfa7ef9251cf5\nlabel:1\nwsi:camelyon16_train_tumor_021\nbcbaa35feec3d65df42799179719ededa3daa2e5\nlabel:1\nwsi:not found\n06b546ead8e2fa3c423da707da163d31f1d954df\nlabel:1\nwsi:camelyon16_train_tumor_027\n474059a719a0810382a7ed3cf0b9bc3097e44ae0\nlabel:1\nwsi:not found\n</code></p>\n\n<p><img src=\"https://i.imgur.com/cFC2gBt.png\" alt=\"Duplicates\"></p>",
      "rawMarkdown": "Domain knowledge is definitely useful here. Also, understanding how the WSIs were extracted could be helpful.\n\nI extracted some feature vectors from the images and compared their distances. By just comparing the first 1000 training images to the rest, I found two duplicate sets. However, this is a very slow way of searching these and I'm probably not going to go through the whole set.\n\n```\na5e032f36bcf57c281c0d6c482ccfa7ef9251cf5\nlabel:1\nwsi:camelyon16_train_tumor_021\nbcbaa35feec3d65df42799179719ededa3daa2e5\nlabel:1\nwsi:not found\n06b546ead8e2fa3c423da707da163d31f1d954df\nlabel:1\nwsi:camelyon16_train_tumor_027\n474059a719a0810382a7ed3cf0b9bc3097e44ae0\nlabel:1\nwsi:not found\n```\n\n![Duplicates](https://i.imgur.com/cFC2gBt.png)",
      "votes": null
    },
    {
      "id": "496477",
      "postDate": "03/22/2019 08:52:26",
      "content": "<p>Duplicates with the first 2000 training images (+ the ones above):</p>\n\n<p><code>\ne41069001765d923f2865de764f70be74afc1307\nlabel:1\nwsi: not found\n41f1ae2f73a5fe6a8f6642dca5c6e14e7175a3d7\nlabel:1\nwsi: not found\nd8f32c46bf7557821cfa5cf114bbb1a347ac6438\nlabel:1\nwsi: not found\n9e17709ee5f13555ae5506a325af007b4752e51f\nlabel:1\nwsi: camelyon16_train_tumor_050\n48f19a11abb3b96bb3166182dbf9291b241b81cc\nlabel:1\nwsi: camelyon16_train_tumor_029\neccd0bbb88dfe5fd3fbce6db705581b4a72c035b\nlabel:1\nwsi: camelyon16_train_tumor_029\nee18d2d7a5549aa0a8624870594127be1a74b777\nlabel:1\nwsi: not found\n7727fd6be2163869b61f99d7233d46e9bc4cada9\nlabel:1\nwsi: not found\nf238c8d42363a347e43599c95fa70bfbdcacb151\nlabel:1\nwsi: not found\n9be2f92b143860ae6b49e67b526502152058d996\nlabel:1\nwsi: camelyon16_train_tumor_092\nc256b7c830b0c3027779f23b77843b4c4f341603\nlabel:1\nwsi: camelyon16_train_tumor_061\nc45586ea9a193578191d0f3b22548f061e83b084\nlabel:1\nwsi: camelyon16_train_tumor_061\n</code></p>\n\n<p><img src=\"https://i.imgur.com/l8mEQIv.png\" alt=\"duplicates 1000-2000\"></p>",
      "rawMarkdown": "Duplicates with the first 2000 training images (+ the ones above):\n\n```\ne41069001765d923f2865de764f70be74afc1307\nlabel:1\nwsi: not found\n41f1ae2f73a5fe6a8f6642dca5c6e14e7175a3d7\nlabel:1\nwsi: not found\nd8f32c46bf7557821cfa5cf114bbb1a347ac6438\nlabel:1\nwsi: not found\n9e17709ee5f13555ae5506a325af007b4752e51f\nlabel:1\nwsi: camelyon16_train_tumor_050\n48f19a11abb3b96bb3166182dbf9291b241b81cc\nlabel:1\nwsi: camelyon16_train_tumor_029\neccd0bbb88dfe5fd3fbce6db705581b4a72c035b\nlabel:1\nwsi: camelyon16_train_tumor_029\nee18d2d7a5549aa0a8624870594127be1a74b777\nlabel:1\nwsi: not found\n7727fd6be2163869b61f99d7233d46e9bc4cada9\nlabel:1\nwsi: not found\nf238c8d42363a347e43599c95fa70bfbdcacb151\nlabel:1\nwsi: not found\n9be2f92b143860ae6b49e67b526502152058d996\nlabel:1\nwsi: camelyon16_train_tumor_092\nc256b7c830b0c3027779f23b77843b4c4f341603\nlabel:1\nwsi: camelyon16_train_tumor_061\nc45586ea9a193578191d0f3b22548f061e83b084\nlabel:1\nwsi: camelyon16_train_tumor_061\n```\n\n![duplicates 1000-2000](https://i.imgur.com/l8mEQIv.png)",
      "votes": null
    },
    {
      "id": "496482",
      "postDate": "03/22/2019 09:03:08",
      "content": "<p>The second id <code>d547e4b7204266e4412887ad82ce59bae5852ba5</code> is not the same as the <code>Image (2)</code>. This is reassuring because it means that there are yet no duplicates found with different WSIs.</p>",
      "rawMarkdown": "The second id `d547e4b7204266e4412887ad82ce59bae5852ba5` is not the same as the `Image (2)`. This is reassuring because it means that there are yet no duplicates found with different WSIs.",
      "votes": null
    },
    {
      "id": "496558",
      "postDate": "03/22/2019 10:34:21",
      "content": "<p>The first 10,000 test images don't seem to match with any of the train images. So far no test-train leakage found.</p>",
      "rawMarkdown": "The first 10,000 test images don't seem to match with any of the train images. So far no test-train leakage found.",
      "votes": null
    },
    {
      "id": "496591",
      "postDate": "03/22/2019 11:24:56",
      "content": "<p><img src=\"https://i.postimg.cc/3R2j4tCR/1.png\" alt=\"\">\nThis image will answer your question.</p>",
      "rawMarkdown": "![](https://i.postimg.cc/3R2j4tCR/1.png)\nThis image will answer your question.",
      "votes": null
    },
    {
      "id": "496596",
      "postDate": "03/22/2019 11:27:31",
      "content": "<p>Do you know how <a href=\"/sermakarevich\">@sermakarevich</a> (SM) obtained the WSI id? I looked into the original dataset and did not find an efficient way to do so. If we can replicate his process (he only used means, but it would be better to also use stdev), we can find more WSI ids and that will help us to find more duplicates (because you would expect WSI to duplicate too)</p>",
      "rawMarkdown": "Do you know how @sermakarevich (SM) obtained the WSI id? I looked into the original dataset and did not find an efficient way to do so. If we can replicate his process (he only used means, but it would be better to also use stdev), we can find more WSI ids and that will help us to find more duplicates (because you would expect WSI to duplicate too)",
      "votes": null
    },
    {
      "id": "496601",
      "postDate": "03/22/2019 11:35:38",
      "content": "<p>Oh! Thank you for the catch. I will look for the id of <code>Image (2)</code> later.</p>",
      "rawMarkdown": "Oh! Thank you for the catch. I will look for the id of `Image (2)` later.",
      "votes": null
    },
    {
      "id": "496620",
      "postDate": "03/22/2019 11:54:00",
      "content": "<p><code>Image (2)</code> is <code>7a0c08fce92d57350add2b13e4c6e825122e7b1a</code>\nInterestingly, they have the same WSI</p>",
      "rawMarkdown": "`Image (2)` is `7a0c08fce92d57350add2b13e4c6e825122e7b1a`\nInterestingly, they have the same WSI",
      "votes": null
    },
    {
      "id": "496634",
      "postDate": "03/22/2019 12:15:49",
      "content": "<p>Sorry, I don't know how to get the WSIs out of the original ds.</p>\n\n<p>If I have understood correctly, what's different in each WSI beside of the tissue sample, is the dye saturation. The coloring is not consistent and what affects the output are:\n- Staining mixture, the combination of hematoxylin &amp; eosin. Different batches have different binding properties.\n- Staining procedure, dipping time. The longer each tissue sample on a glass slide is dipped in the mixture, the more color it gets.\n- Instruments, different light microscope.\n- The dyes tend to faint over time so if the images are captured immediately vs. an hour later, they will look a bit different.</p>\n\n<p>Practically, all WSIs are different in the stain color intensities. If you could extract the amounts of stain in each WSI, perhaps it could be used to identify patch WSIs.\nThe color forming procedure in light microscopy is subtractive: </p>\n\n<p><code>light_that_gets_to_the_image = light_from_source - hematoxylin_amount*hematoxylin_absorption_factor - easin_amount*easin_absorption_factor</code></p>\n\n<p>One way of doing this is color deconvolution. I tried this but didn't have much luck. I used <a href=\"https://digitalslidearchive.github.io/HistomicsTK/examples/color-deconvolution.html\">HistomicsTK</a></p>",
      "rawMarkdown": "Sorry, I don't know how to get the WSIs out of the original ds.\n\nIf I have understood correctly, what's different in each WSI beside of the tissue sample, is the dye saturation. The coloring is not consistent and what affects the output are:\n- Staining mixture, the combination of hematoxylin &amp; eosin. Different batches have different binding properties.\n- Staining procedure, dipping time. The longer each tissue sample on a glass slide is dipped in the mixture, the more color it gets.\n- Instruments, different light microscope.\n- The dyes tend to faint over time so if the images are captured immediately vs. an hour later, they will look a bit different.\n\nPractically, all WSIs are different in the stain color intensities. If you could extract the amounts of stain in each WSI, perhaps it could be used to identify patch WSIs.\nThe color forming procedure in light microscopy is subtractive: \n\n```light_that_gets_to_the_image = light_from_source - hematoxylin_amount*hematoxylin_absorption_factor - easin_amount*easin_absorption_factor```\n\nOne way of doing this is color deconvolution. I tried this but didn't have much luck. I used [HistomicsTK](https://digitalslidearchive.github.io/HistomicsTK/examples/color-deconvolution.html)",
      "votes": null
    },
    {
      "id": "496715",
      "postDate": "03/22/2019 14:05:12",
      "content": "<p>The patches from the dataset do not add up to a WSI, so using the RGB color to identify WSI would have a lot of noise. As far as I know, the network does a pretty good job interpreting the difference in dye saturation, but finding the leakage is more important.</p>",
      "rawMarkdown": "The patches from the dataset do not add up to a WSI, so using the RGB color to identify WSI would have a lot of noise. As far as I know, the network does a pretty good job interpreting the difference in dye saturation, but finding the leakage is more important.",
      "votes": null
    },
    {
      "id": "496738",
      "postDate": "03/22/2019 14:27:12",
      "content": "<p>I believe this was stated in the overview for the competition.  The WSI labels are derived from PCam. </p>\n\n<p><code>The data for this competition is a slightly modified version of the PatchCamelyon (PCam) benchmark dataset (the original PCam dataset contains duplicate images due to its probabilistic sampling, however, the version presented on Kaggle does not contain duplicates).</code></p>",
      "rawMarkdown": "I believe this was stated in the overview for the competition.  The WSI labels are derived from PCam. \n\n` The data for this competition is a slightly modified version of the PatchCamelyon (PCam) benchmark dataset (the original PCam dataset contains duplicate images due to its probabilistic sampling, however, the version presented on Kaggle does not contain duplicates).`",
      "votes": null
    },
    {
      "id": "496880",
      "postDate": "03/22/2019 17:42:33",
      "content": "<p><a href=\"/qitvision\">@qitvision</a>  I have assigned all dataset images to their correct WSI so none remain as unknown. See my post <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/85283\">https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/85283</a> </p>",
      "rawMarkdown": "qitvision  I have assigned all dataset images to their correct WSI so none remain as unknown. See my post https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/85283",
      "votes": null
    },
    {
      "id": "496903",
      "postDate": "03/22/2019 18:17:15",
      "content": "<p>After <a href=\"/markdesimone\">@markdesimone</a> released <code>patch_id_wsi_full</code> (thanks), it seems like the duplicates are all from the same WSI id. 8/2000=0.4%(+-1%?) of the training images have duplicates. If the above is true, the expected increase in LB would be about 0.0005~0.0001... Not my priority to work on.</p>",
      "rawMarkdown": "After @markdesimone released `patch_id_wsi_full` (thanks), it seems like the duplicates are all from the same WSI id. 8/2000=0.4%(+-1%?) of the training images have duplicates. If the above is true, the expected increase in LB would be about 0.0005~0.0001... Not my priority to work on.",
      "votes": null
    },
    {
      "id": "496937",
      "postDate": "03/22/2019 18:52:15",
      "content": "<p>Nice, a clever solution!</p>",
      "rawMarkdown": "Nice, a clever solution!",
      "votes": null
    },
    {
      "id": "497013",
      "postDate": "03/22/2019 20:38:02",
      "content": "<p>Thanks, I just got started on this but your excellent kernel is a great starting point, thank you</p>",
      "rawMarkdown": "Thanks, I just got started on this but your excellent kernel is a great starting point, thank you",
      "votes": null
    },
    {
      "id": "497538",
      "postDate": "03/23/2019 16:41:05",
      "content": "<p><strong>Duplicate search</strong>:\n- found 396 duplicates in training set (this list is not exhaustive)\n- all duplicates have the same WSI\n- found 0 duplicates between train and test sets</p>\n\n<p>Training set duplicate id's are in <em>duplicates.csv</em></p>",
      "rawMarkdown": "**Duplicate search**:\n- found 396 duplicates in training set (this list is not exhaustive)\n- all duplicates have the same WSI\n- found 0 duplicates between train and test sets\n\nTraining set duplicate id's are in *duplicates.csv*",
      "votes": null
    },
    {
      "id": "497549",
      "postDate": "03/23/2019 16:49:27",
      "content": "<p>Thank you for sharing! How did you find them?</p>",
      "rawMarkdown": "Thank you for sharing! How did you find them?",
      "votes": null
    },
    {
      "id": "497571",
      "postDate": "03/23/2019 17:08:52",
      "content": "<p>I extracted color and texture features and then performed a nearest neighbor search between feature vectors. If the distance was less than a set threshold, it was a match. I set a quite low threshold to filter out false duplicates so this didn't find all the pairs. For example, this didn't find that duplicate pair you found.</p>",
      "rawMarkdown": "I extracted color and texture features and then performed a nearest neighbor search between feature vectors. If the distance was less than a set threshold, it was a match. I set a quite low threshold to filter out false duplicates so this didn't find all the pairs. For example, this didn't find that duplicate pair you found.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 495731,
      "author_name": "guntherthepenguin",
      "author_url": "",
      "post_date": "03/21/2019 13:37:39",
      "content": "<p>Yes you are right, if we don't split by WSI the model overfits to these duplicates</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 495767,
      "author_name": "kokecacao",
      "author_url": "",
      "post_date": "03/21/2019 14:37:24",
      "content": "<p>The point is, even if we split according to wsi, we still cannot make a stable cv.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 495886,
      "author_name": "qitvision",
      "author_url": "",
      "post_date": "03/21/2019 16:33:03",
      "content": "<p>Good find! How did you discover this? I wonder how many more cross WSI duplicates there are.</p>",
      "votes": null,
      "replies": [
        {
          "id": 496260,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "03/22/2019 02:25:32",
          "content": "<p>Your trained model can tell you that! Looking at the probability I get from my model's prediction, I found some interesting patterns there. The playground competition is very boring compared to the <code>Featured</code> ones - so having some fun digging some golds is not a bad idea. (chat on twitter?)</p>\n\n<p>Also, do you see the LB 1st there? with a huge increase in LB score than the rest of us. I bet he understands the data a lot better than we do. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496429,
          "author_name": "qitvision",
          "author_url": "",
          "post_date": "03/22/2019 07:30:25",
          "content": "<p>Domain knowledge is definitely useful here. Also, understanding how the WSIs were extracted could be helpful.</p>\n\n<p>I extracted some feature vectors from the images and compared their distances. By just comparing the first 1000 training images to the rest, I found two duplicate sets. However, this is a very slow way of searching these and I'm probably not going to go through the whole set.</p>\n\n<p><code>\na5e032f36bcf57c281c0d6c482ccfa7ef9251cf5\nlabel:1\nwsi:camelyon16_train_tumor_021\nbcbaa35feec3d65df42799179719ededa3daa2e5\nlabel:1\nwsi:not found\n06b546ead8e2fa3c423da707da163d31f1d954df\nlabel:1\nwsi:camelyon16_train_tumor_027\n474059a719a0810382a7ed3cf0b9bc3097e44ae0\nlabel:1\nwsi:not found\n</code></p>\n\n<p><img src=\"https://i.imgur.com/cFC2gBt.png\" alt=\"Duplicates\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496477,
          "author_name": "qitvision",
          "author_url": "",
          "post_date": "03/22/2019 08:52:26",
          "content": "<p>Duplicates with the first 2000 training images (+ the ones above):</p>\n\n<p><code>\ne41069001765d923f2865de764f70be74afc1307\nlabel:1\nwsi: not found\n41f1ae2f73a5fe6a8f6642dca5c6e14e7175a3d7\nlabel:1\nwsi: not found\nd8f32c46bf7557821cfa5cf114bbb1a347ac6438\nlabel:1\nwsi: not found\n9e17709ee5f13555ae5506a325af007b4752e51f\nlabel:1\nwsi: camelyon16_train_tumor_050\n48f19a11abb3b96bb3166182dbf9291b241b81cc\nlabel:1\nwsi: camelyon16_train_tumor_029\neccd0bbb88dfe5fd3fbce6db705581b4a72c035b\nlabel:1\nwsi: camelyon16_train_tumor_029\nee18d2d7a5549aa0a8624870594127be1a74b777\nlabel:1\nwsi: not found\n7727fd6be2163869b61f99d7233d46e9bc4cada9\nlabel:1\nwsi: not found\nf238c8d42363a347e43599c95fa70bfbdcacb151\nlabel:1\nwsi: not found\n9be2f92b143860ae6b49e67b526502152058d996\nlabel:1\nwsi: camelyon16_train_tumor_092\nc256b7c830b0c3027779f23b77843b4c4f341603\nlabel:1\nwsi: camelyon16_train_tumor_061\nc45586ea9a193578191d0f3b22548f061e83b084\nlabel:1\nwsi: camelyon16_train_tumor_061\n</code></p>\n\n<p><img src=\"https://i.imgur.com/l8mEQIv.png\" alt=\"duplicates 1000-2000\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496558,
          "author_name": "qitvision",
          "author_url": "",
          "post_date": "03/22/2019 10:34:21",
          "content": "<p>The first 10,000 test images don't seem to match with any of the train images. So far no test-train leakage found.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496591,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "03/22/2019 11:24:56",
          "content": "<p><img src=\"https://i.postimg.cc/3R2j4tCR/1.png\" alt=\"\">\nThis image will answer your question.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496596,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "03/22/2019 11:27:31",
          "content": "<p>Do you know how <a href=\"/sermakarevich\">@sermakarevich</a> (SM) obtained the WSI id? I looked into the original dataset and did not find an efficient way to do so. If we can replicate his process (he only used means, but it would be better to also use stdev), we can find more WSI ids and that will help us to find more duplicates (because you would expect WSI to duplicate too)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496634,
          "author_name": "qitvision",
          "author_url": "",
          "post_date": "03/22/2019 12:15:49",
          "content": "<p>Sorry, I don't know how to get the WSIs out of the original ds.</p>\n\n<p>If I have understood correctly, what's different in each WSI beside of the tissue sample, is the dye saturation. The coloring is not consistent and what affects the output are:\n- Staining mixture, the combination of hematoxylin &amp; eosin. Different batches have different binding properties.\n- Staining procedure, dipping time. The longer each tissue sample on a glass slide is dipped in the mixture, the more color it gets.\n- Instruments, different light microscope.\n- The dyes tend to faint over time so if the images are captured immediately vs. an hour later, they will look a bit different.</p>\n\n<p>Practically, all WSIs are different in the stain color intensities. If you could extract the amounts of stain in each WSI, perhaps it could be used to identify patch WSIs.\nThe color forming procedure in light microscopy is subtractive: </p>\n\n<p><code>light_that_gets_to_the_image = light_from_source - hematoxylin_amount*hematoxylin_absorption_factor - easin_amount*easin_absorption_factor</code></p>\n\n<p>One way of doing this is color deconvolution. I tried this but didn't have much luck. I used <a href=\"https://digitalslidearchive.github.io/HistomicsTK/examples/color-deconvolution.html\">HistomicsTK</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496715,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "03/22/2019 14:05:12",
          "content": "<p>The patches from the dataset do not add up to a WSI, so using the RGB color to identify WSI would have a lot of noise. As far as I know, the network does a pretty good job interpreting the difference in dye saturation, but finding the leakage is more important.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496903,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "03/22/2019 18:17:15",
          "content": "<p>After <a href=\"/markdesimone\">@markdesimone</a> released <code>patch_id_wsi_full</code> (thanks), it seems like the duplicates are all from the same WSI id. 8/2000=0.4%(+-1%?) of the training images have duplicates. If the above is true, the expected increase in LB would be about 0.0005~0.0001... Not my priority to work on.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 496482,
      "author_name": "qitvision",
      "author_url": "",
      "post_date": "03/22/2019 09:03:08",
      "content": "<p>The second id <code>d547e4b7204266e4412887ad82ce59bae5852ba5</code> is not the same as the <code>Image (2)</code>. This is reassuring because it means that there are yet no duplicates found with different WSIs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 496601,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "03/22/2019 11:35:38",
          "content": "<p>Oh! Thank you for the catch. I will look for the id of <code>Image (2)</code> later.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496620,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "03/22/2019 11:54:00",
          "content": "<p><code>Image (2)</code> is <code>7a0c08fce92d57350add2b13e4c6e825122e7b1a</code>\nInterestingly, they have the same WSI</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 496738,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "03/22/2019 14:27:12",
      "content": "<p>I believe this was stated in the overview for the competition.  The WSI labels are derived from PCam. </p>\n\n<p><code>The data for this competition is a slightly modified version of the PatchCamelyon (PCam) benchmark dataset (the original PCam dataset contains duplicate images due to its probabilistic sampling, however, the version presented on Kaggle does not contain duplicates).</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 496880,
      "author_name": "markdesimone",
      "author_url": "",
      "post_date": "03/22/2019 17:42:33",
      "content": "<p><a href=\"/qitvision\">@qitvision</a>  I have assigned all dataset images to their correct WSI so none remain as unknown. See my post <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/85283\">https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/85283</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 496937,
          "author_name": "qitvision",
          "author_url": "",
          "post_date": "03/22/2019 18:52:15",
          "content": "<p>Nice, a clever solution!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 497013,
          "author_name": "markdesimone",
          "author_url": "",
          "post_date": "03/22/2019 20:38:02",
          "content": "<p>Thanks, I just got started on this but your excellent kernel is a great starting point, thank you</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 497538,
      "author_name": "qitvision",
      "author_url": "",
      "post_date": "03/23/2019 16:41:05",
      "content": "<p><strong>Duplicate search</strong>:\n- found 396 duplicates in training set (this list is not exhaustive)\n- all duplicates have the same WSI\n- found 0 duplicates between train and test sets</p>\n\n<p>Training set duplicate id's are in <em>duplicates.csv</em></p>",
      "votes": null,
      "replies": [
        {
          "id": 497549,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "03/23/2019 16:49:27",
          "content": "<p>Thank you for sharing! How did you find them?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 497571,
          "author_name": "qitvision",
          "author_url": "",
          "post_date": "03/23/2019 17:08:52",
          "content": "<p>I extracted color and texture features and then performed a nearest neighbor search between feature vectors. If the distance was less than a set threshold, it was a match. I set a quite low threshold to filter out false duplicates so this didn't find all the pairs. For example, this didn't find that duplicate pair you found.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "495220": "```\nlabel                             1\nwsi      camelyon16_train_tumor_007\nName: c448ff7117b8382151d115fb54a63987e0fb5c0d, dtype: object to label                             1\nwsi      camelyon16_train_tumor_006\nName: d547e4b7204266e4412887ad82ce59bae5852ba5, dtype: object\n```\n![](https://i.postimg.cc/GpDTwnyS/download.png)\n\nThis could explain why using WSI makes a difference!\nAlso, these duplicates images can inspire us about augmentations.\nWhat do you think?",
    "495731": "Yes you are right, if we don't split by WSI the model overfits to these duplicates",
    "495767": "The point is, even if we split according to wsi, we still cannot make a stable cv.",
    "495886": "Good find! How did you discover this? I wonder how many more cross WSI duplicates there are.",
    "496260": "Your trained model can tell you that! Looking at the probability I get from my model's prediction, I found some interesting patterns there. The playground competition is very boring compared to the `Featured` ones - so having some fun digging some golds is not a bad idea. (chat on twitter?)\n\nAlso, do you see the LB 1st there? with a huge increase in LB score than the rest of us. I bet he understands the data a lot better than we do.",
    "496429": "Domain knowledge is definitely useful here. Also, understanding how the WSIs were extracted could be helpful.\n\nI extracted some feature vectors from the images and compared their distances. By just comparing the first 1000 training images to the rest, I found two duplicate sets. However, this is a very slow way of searching these and I'm probably not going to go through the whole set.\n\n```\na5e032f36bcf57c281c0d6c482ccfa7ef9251cf5\nlabel:1\nwsi:camelyon16_train_tumor_021\nbcbaa35feec3d65df42799179719ededa3daa2e5\nlabel:1\nwsi:not found\n06b546ead8e2fa3c423da707da163d31f1d954df\nlabel:1\nwsi:camelyon16_train_tumor_027\n474059a719a0810382a7ed3cf0b9bc3097e44ae0\nlabel:1\nwsi:not found\n```\n\n![Duplicates](https://i.imgur.com/cFC2gBt.png)",
    "496477": "Duplicates with the first 2000 training images (+ the ones above):\n\n```\ne41069001765d923f2865de764f70be74afc1307\nlabel:1\nwsi: not found\n41f1ae2f73a5fe6a8f6642dca5c6e14e7175a3d7\nlabel:1\nwsi: not found\nd8f32c46bf7557821cfa5cf114bbb1a347ac6438\nlabel:1\nwsi: not found\n9e17709ee5f13555ae5506a325af007b4752e51f\nlabel:1\nwsi: camelyon16_train_tumor_050\n48f19a11abb3b96bb3166182dbf9291b241b81cc\nlabel:1\nwsi: camelyon16_train_tumor_029\neccd0bbb88dfe5fd3fbce6db705581b4a72c035b\nlabel:1\nwsi: camelyon16_train_tumor_029\nee18d2d7a5549aa0a8624870594127be1a74b777\nlabel:1\nwsi: not found\n7727fd6be2163869b61f99d7233d46e9bc4cada9\nlabel:1\nwsi: not found\nf238c8d42363a347e43599c95fa70bfbdcacb151\nlabel:1\nwsi: not found\n9be2f92b143860ae6b49e67b526502152058d996\nlabel:1\nwsi: camelyon16_train_tumor_092\nc256b7c830b0c3027779f23b77843b4c4f341603\nlabel:1\nwsi: camelyon16_train_tumor_061\nc45586ea9a193578191d0f3b22548f061e83b084\nlabel:1\nwsi: camelyon16_train_tumor_061\n```\n\n![duplicates 1000-2000](https://i.imgur.com/l8mEQIv.png)",
    "496482": "The second id `d547e4b7204266e4412887ad82ce59bae5852ba5` is not the same as the `Image (2)`. This is reassuring because it means that there are yet no duplicates found with different WSIs.",
    "496558": "The first 10,000 test images don't seem to match with any of the train images. So far no test-train leakage found.",
    "496591": "![](https://i.postimg.cc/3R2j4tCR/1.png)\nThis image will answer your question.",
    "496596": "Do you know how @sermakarevich (SM) obtained the WSI id? I looked into the original dataset and did not find an efficient way to do so. If we can replicate his process (he only used means, but it would be better to also use stdev), we can find more WSI ids and that will help us to find more duplicates (because you would expect WSI to duplicate too)",
    "496601": "Oh! Thank you for the catch. I will look for the id of `Image (2)` later.",
    "496620": "`Image (2)` is `7a0c08fce92d57350add2b13e4c6e825122e7b1a`\nInterestingly, they have the same WSI",
    "496634": "Sorry, I don't know how to get the WSIs out of the original ds.\n\nIf I have understood correctly, what's different in each WSI beside of the tissue sample, is the dye saturation. The coloring is not consistent and what affects the output are:\n- Staining mixture, the combination of hematoxylin &amp; eosin. Different batches have different binding properties.\n- Staining procedure, dipping time. The longer each tissue sample on a glass slide is dipped in the mixture, the more color it gets.\n- Instruments, different light microscope.\n- The dyes tend to faint over time so if the images are captured immediately vs. an hour later, they will look a bit different.\n\nPractically, all WSIs are different in the stain color intensities. If you could extract the amounts of stain in each WSI, perhaps it could be used to identify patch WSIs.\nThe color forming procedure in light microscopy is subtractive: \n\n```light_that_gets_to_the_image = light_from_source - hematoxylin_amount*hematoxylin_absorption_factor - easin_amount*easin_absorption_factor```\n\nOne way of doing this is color deconvolution. I tried this but didn't have much luck. I used [HistomicsTK](https://digitalslidearchive.github.io/HistomicsTK/examples/color-deconvolution.html)",
    "496715": "The patches from the dataset do not add up to a WSI, so using the RGB color to identify WSI would have a lot of noise. As far as I know, the network does a pretty good job interpreting the difference in dye saturation, but finding the leakage is more important.",
    "496738": "I believe this was stated in the overview for the competition.  The WSI labels are derived from PCam. \n\n` The data for this competition is a slightly modified version of the PatchCamelyon (PCam) benchmark dataset (the original PCam dataset contains duplicate images due to its probabilistic sampling, however, the version presented on Kaggle does not contain duplicates).`",
    "496880": "qitvision  I have assigned all dataset images to their correct WSI so none remain as unknown. See my post https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/85283",
    "496903": "After @markdesimone released `patch_id_wsi_full` (thanks), it seems like the duplicates are all from the same WSI id. 8/2000=0.4%(+-1%?) of the training images have duplicates. If the above is true, the expected increase in LB would be about 0.0005~0.0001... Not my priority to work on.",
    "496937": "Nice, a clever solution!",
    "497013": "Thanks, I just got started on this but your excellent kernel is a great starting point, thank you",
    "497538": "**Duplicate search**:\n- found 396 duplicates in training set (this list is not exhaustive)\n- all duplicates have the same WSI\n- found 0 duplicates between train and test sets\n\nTraining set duplicate id's are in *duplicates.csv*",
    "497549": "Thank you for sharing! How did you find them?",
    "497571": "I extracted color and texture features and then performed a nearest neighbor search between feature vectors. If the distance was less than a set threshold, it was a match. I set a quite low threshold to filter out false duplicates so this didn't find all the pairs. For example, this didn't find that duplicate pair you found."
  },
  "source": "meta"
}