{
  "id": 72534,
  "title": "A list of identical and near-identical images",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/72534",
  "author_name": "",
  "post_date": "2018-11-24T07:13:06.478861800Z",
  "votes": 34,
  "comment_count": 18,
  "views": 0,
  "content": "<p>It seems that there is interest in this topic as it has been brought up <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69594\"><strong>here</strong></a>, but not fully described.</p>\n\n<p>The fastest way I know to find similar images is by image hashes - there is a nice Python implementation in <a href=\"https://github.com/JohannesBuchner/imagehash\"><strong>imagehash</strong></a> package. I have done it only on green images using <code>average hashing</code> and <code>perception hashing</code>. The latter is more reliable when it comes to finding true positive matches, while the former may be better at finding similar images at the expense of few bogus matches.</p>\n\n<p>First, here are two image pairs identified by Brian in <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69594\"><strong>this post</strong></a>.</p>\n\n<p>In train:\n<img src=\"https://i.postimg.cc/qMyxxzv8/example-train.png\" alt=\"enter image description here\"></p>\n\n<p>In test:\n<img src=\"https://i.postimg.cc/zXCTR6x7/example-test.png\" alt=\"enter image description here\"></p>\n\n<p>Train images seem identical but don't have the same brightness/amplitude. In microscopy parlance, that would mean the same field of cells but different exposure, or possibly two subsequent z-stacks during image acquisition. Test images do seem identical.</p>\n\n<p>I will attach my lists as text files. Their format is two columns of names like this:</p>\n\n<pre><code>565dbefe-bbc0-11e8-b2bb-ac1f6b6435d0   62c829c8-bbba-11e8-b2ba-ac1f6b6435d0\n0940c9ee-bbb7-11e8-b2ba-ac1f6b6435d0   40ba3872-bbc8-11e8-b2bc-ac1f6b6435d0\n3441eab0-bbb2-11e8-b2ba-ac1f6b6435d0   c4908034-bb9b-11e8-b2b9-ac1f6b6435d0\n90095b92-bba0-11e8-b2b9-ac1f6b6435d0   f46bceb2-bbb4-11e8-b2ba-ac1f6b6435d0\n</code></pre>\n\n<p>I also include rar-compressed side-by-side comparisons of all found pairs. Here are few examples that are identical according to <code>phash</code>:</p>\n\n<p><img src=\"https://i.postimg.cc/cLjDFVN3/train-identical-097.png\" alt=\"enter image description here\">\n<img src=\"https://i.postimg.cc/4yBPV05m/train-identical-094.png\" alt=\"enter image description here\"></p>\n\n<p>There is a slight difference in brightness, but it should be obvious that they are near-identical.</p>\n\n<p>Here are couple of duplicates that are truly identical:</p>\n\n<p><img src=\"https://i.postimg.cc/Bv48fvsf/train-identical-092.png\" alt=\"enter image description here\">\n<img src=\"https://i.postimg.cc/wTkMcpLW/test-identical-009.png\" alt=\"enter image description here\"></p>\n\n<p>And an example of similar pairs that <code>ahash</code> finds but <code>phash</code> doesn't, which would indicate that they are similar rather than identical. To my eye they seem different exposures of the same image.</p>\n\n<p><img src=\"https://i.postimg.cc/BnsHmchJ/train-similar-041.png\" alt=\"enter image description here\"></p>\n\n<p>I could not find any images between train and test that were identical - only within the two groups.</p>\n\n<p>If someone feels like playing with other hashes and finds other duplicates, please report them here.</p>",
  "messages": [
    {
      "id": "426935",
      "postDate": "11/24/2018 07:13:06",
      "content": "<p>It seems that there is interest in this topic as it has been brought up <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69594\"><strong>here</strong></a>, but not fully described.</p>\n\n<p>The fastest way I know to find similar images is by image hashes - there is a nice Python implementation in <a href=\"https://github.com/JohannesBuchner/imagehash\"><strong>imagehash</strong></a> package. I have done it only on green images using <code>average hashing</code> and <code>perception hashing</code>. The latter is more reliable when it comes to finding true positive matches, while the former may be better at finding similar images at the expense of few bogus matches.</p>\n\n<p>First, here are two image pairs identified by Brian in <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69594\"><strong>this post</strong></a>.</p>\n\n<p>In train:\n<img src=\"https://i.postimg.cc/qMyxxzv8/example-train.png\" alt=\"enter image description here\"></p>\n\n<p>In test:\n<img src=\"https://i.postimg.cc/zXCTR6x7/example-test.png\" alt=\"enter image description here\"></p>\n\n<p>Train images seem identical but don't have the same brightness/amplitude. In microscopy parlance, that would mean the same field of cells but different exposure, or possibly two subsequent z-stacks during image acquisition. Test images do seem identical.</p>\n\n<p>I will attach my lists as text files. Their format is two columns of names like this:</p>\n\n<pre><code>565dbefe-bbc0-11e8-b2bb-ac1f6b6435d0   62c829c8-bbba-11e8-b2ba-ac1f6b6435d0\n0940c9ee-bbb7-11e8-b2ba-ac1f6b6435d0   40ba3872-bbc8-11e8-b2bc-ac1f6b6435d0\n3441eab0-bbb2-11e8-b2ba-ac1f6b6435d0   c4908034-bb9b-11e8-b2b9-ac1f6b6435d0\n90095b92-bba0-11e8-b2b9-ac1f6b6435d0   f46bceb2-bbb4-11e8-b2ba-ac1f6b6435d0\n</code></pre>\n\n<p>I also include rar-compressed side-by-side comparisons of all found pairs. Here are few examples that are identical according to <code>phash</code>:</p>\n\n<p><img src=\"https://i.postimg.cc/cLjDFVN3/train-identical-097.png\" alt=\"enter image description here\">\n<img src=\"https://i.postimg.cc/4yBPV05m/train-identical-094.png\" alt=\"enter image description here\"></p>\n\n<p>There is a slight difference in brightness, but it should be obvious that they are near-identical.</p>\n\n<p>Here are couple of duplicates that are truly identical:</p>\n\n<p><img src=\"https://i.postimg.cc/Bv48fvsf/train-identical-092.png\" alt=\"enter image description here\">\n<img src=\"https://i.postimg.cc/wTkMcpLW/test-identical-009.png\" alt=\"enter image description here\"></p>\n\n<p>And an example of similar pairs that <code>ahash</code> finds but <code>phash</code> doesn't, which would indicate that they are similar rather than identical. To my eye they seem different exposures of the same image.</p>\n\n<p><img src=\"https://i.postimg.cc/BnsHmchJ/train-similar-041.png\" alt=\"enter image description here\"></p>\n\n<p>I could not find any images between train and test that were identical - only within the two groups.</p>\n\n<p>If someone feels like playing with other hashes and finds other duplicates, please report them here.</p>",
      "rawMarkdown": "It seems that there is interest in this topic as it has been brought up [__here__](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69594), but not fully described.\n\nThe fastest way I know to find similar images is by image hashes - there is a nice Python implementation in [__imagehash__](https://github.com/JohannesBuchner/imagehash) package. I have done it only on green images using `average hashing` and `perception hashing`. The latter is more reliable when it comes to finding true positive matches, while the former may be better at finding similar images at the expense of few bogus matches.\n\nFirst, here are two image pairs identified by Brian in [__this post__](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69594).\n\nIn train:\n![enter image description here][1]\n\nIn test:\n![enter image description here][2]\n\nTrain images seem identical but don't have the same brightness/amplitude. In microscopy parlance, that would mean the same field of cells but different exposure, or possibly two subsequent z-stacks during image acquisition. Test images do seem identical.\n\nI will attach my lists as text files. Their format is two columns of names like this:\n\n    565dbefe-bbc0-11e8-b2bb-ac1f6b6435d0   62c829c8-bbba-11e8-b2ba-ac1f6b6435d0\n    0940c9ee-bbb7-11e8-b2ba-ac1f6b6435d0   40ba3872-bbc8-11e8-b2bc-ac1f6b6435d0\n    3441eab0-bbb2-11e8-b2ba-ac1f6b6435d0   c4908034-bb9b-11e8-b2b9-ac1f6b6435d0\n    90095b92-bba0-11e8-b2b9-ac1f6b6435d0   f46bceb2-bbb4-11e8-b2ba-ac1f6b6435d0\n\nI also include rar-compressed side-by-side comparisons of all found pairs. Here are few examples that are identical according to `phash`:\n\n![enter image description here][3]\n![enter image description here][4]\n\nThere is a slight difference in brightness, but it should be obvious that they are near-identical.\n\nHere are couple of duplicates that are truly identical:\n\n![enter image description here][5]\n![enter image description here][6]\n\nAnd an example of similar pairs that `ahash` finds but `phash` doesn't, which would indicate that they are similar rather than identical. To my eye they seem different exposures of the same image.\n\n![enter image description here][7]\n\nI could not find any images between train and test that were identical - only within the two groups.\n\nIf someone feels like playing with other hashes and finds other duplicates, please report them here.\n\n\n  [1]: https://i.postimg.cc/qMyxxzv8/example-train.png\n  [2]: https://i.postimg.cc/zXCTR6x7/example-test.png\n  [3]: https://i.postimg.cc/cLjDFVN3/train-identical-097.png\n  [4]: https://i.postimg.cc/4yBPV05m/train-identical-094.png\n  [5]: https://i.postimg.cc/Bv48fvsf/train-identical-092.png\n  [6]: https://i.postimg.cc/wTkMcpLW/test-identical-009.png\n  [7]: https://i.postimg.cc/BnsHmchJ/train-similar-041.png",
      "votes": null
    },
    {
      "id": "426936",
      "postDate": "11/24/2018 07:15:48",
      "content": "<p>Attachments are here. Rar files were too big for Kaggle (~25 and 50 Mb).</p>",
      "rawMarkdown": "Attachments are here. Rar files were too big for Kaggle (~25 and 50 Mb).",
      "votes": null
    },
    {
      "id": "426941",
      "postDate": "11/24/2018 07:46:21",
      "content": "<p>nice post. I tried a couple you mention and I get the same images (I still don't get the test/train thing of brian - hope I didn't corrupt my data somehow).</p>\n\n<p>One thing that might be interesting: put these images aside before training and then, after training, use the model to classify them. Do they get the same labels?</p>",
      "rawMarkdown": "nice post. I tried a couple you mention and I get the same images (I still don't get the test/train thing of brian - hope I didn't corrupt my data somehow).\n\nOne thing that might be interesting: put these images aside before training and then, after training, use the model to classify them. Do they get the same labels?",
      "votes": null
    },
    {
      "id": "426943",
      "postDate": "11/24/2018 08:05:32",
      "content": "<p>Thank you so much, quite helpful post.</p>",
      "rawMarkdown": "Thank you so much, quite helpful post.",
      "votes": null
    },
    {
      "id": "427105",
      "postDate": "11/24/2018 15:36:04",
      "content": "<p>Great share! Thanks a lot!</p>",
      "rawMarkdown": "Great share! Thanks a lot!",
      "votes": null
    },
    {
      "id": "427191",
      "postDate": "11/24/2018 20:31:08",
      "content": "<p>Oh, thanks so much!</p>\n\n<p>I would keep in mind that if similar numbers of \"duplicates\" appear in the final evaluation set (and follow the the same distribution), then training with the \"duplicate\" images included may actually be better... (?)</p>",
      "rawMarkdown": "Oh, thanks so much!\n\nI would keep in mind that if similar numbers of \"duplicates\" appear in the final evaluation set (and follow the the same distribution), then training with the \"duplicate\" images included may actually be better... (?)",
      "votes": null
    },
    {
      "id": "428145",
      "postDate": "11/26/2018 20:41:48",
      "content": "<p>Here is my list of close matches in the training set of 31k images. Identified by image id and group id, if they are in the same group they are similar images.</p>",
      "rawMarkdown": "Here is my list of close matches in the training set of 31k images. Identified by image id and group id, if they are in the same group they are similar images.",
      "votes": null
    },
    {
      "id": "428342",
      "postDate": "11/27/2018 05:42:24",
      "content": "<p>More similar images in test data:</p>\n\n<pre><code>0774284e-bad7-11e8-b2b9-ac1f6b6435d0   34aa05de-bad4-11e8-b2b8-ac1f6b6435d0\n8f0666da-bad9-11e8-b2b9-ac1f6b6435d0   dad043dc-bad5-11e8-b2b9-ac1f6b6435d0\n</code></pre>\n\n<p>The latter is a good example of different z-slices.</p>",
      "rawMarkdown": "More similar images in test data:\n\n    0774284e-bad7-11e8-b2b9-ac1f6b6435d0   34aa05de-bad4-11e8-b2b8-ac1f6b6435d0\n    8f0666da-bad9-11e8-b2b9-ac1f6b6435d0   dad043dc-bad5-11e8-b2b9-ac1f6b6435d0\n\nThe latter is a good example of different z-slices.",
      "votes": null
    },
    {
      "id": "428383",
      "postDate": "11/27/2018 07:23:27",
      "content": "<p>More similar image pairs in train data:</p>\n\n<pre><code>7116d6f4-bba7-11e8-b2ba-ac1f6b6435d0   a166d11a-bbca-11e8-b2bc-ac1f6b6435d0\n0ebd689e-bbc8-11e8-b2bc-ac1f6b6435d0   826e48c8-bbbc-11e8-b2ba-ac1f6b6435d0\n1826c3b4-bba3-11e8-b2b9-ac1f6b6435d0   f3db54f0-bbae-11e8-b2ba-ac1f6b6435d0\n2f7acfa4-bbc3-11e8-b2bc-ac1f6b6435d0   8832c642-bba0-11e8-b2b9-ac1f6b6435d0\n49366c04-bbb0-11e8-b2ba-ac1f6b6435d0   b4af5e20-bbbd-11e8-b2ba-ac1f6b6435d0\na7fcfccc-bbb9-11e8-b2ba-ac1f6b6435d0   d9ad48e0-bb9f-11e8-b2b9-ac1f6b6435d0\nb1371ab2-bb9f-11e8-b2b9-ac1f6b6435d0   ebfcde90-bbac-11e8-b2ba-ac1f6b6435d0\n589a93cc-bbca-11e8-b2bc-ac1f6b6435d0   b367962a-bbc9-11e8-b2bc-ac1f6b6435d0\na656e5f2-bb9d-11e8-b2b9-ac1f6b6435d0   cf78b532-bbb1-11e8-b2ba-ac1f6b6435d0\n589a93cc-bbca-11e8-b2bc-ac1f6b6435d0   b367962a-bbc9-11e8-b2bc-ac1f6b6435d0\n4d3ad8a6-bbc1-11e8-b2bb-ac1f6b6435d0   7927dc50-bba5-11e8-b2ba-ac1f6b6435d0\n575666ec-bba9-11e8-b2ba-ac1f6b6435d0   63d23f16-bba6-11e8-b2ba-ac1f6b6435d0\n5dc2623a-bbb6-11e8-b2ba-ac1f6b6435d0   81cea69c-bb9e-11e8-b2b9-ac1f6b6435d0\n80ea0dc4-bbb3-11e8-b2ba-ac1f6b6435d0   ad2eff22-bba7-11e8-b2ba-ac1f6b6435d0\n3e608f2e-bba8-11e8-b2ba-ac1f6b6435d0   8779d884-bba6-11e8-b2ba-ac1f6b6435d0\nd57bc36c-bba6-11e8-b2ba-ac1f6b6435d0   d79a532e-bbca-11e8-b2bc-ac1f6b6435d0\n73f49160-bbc8-11e8-b2bc-ac1f6b6435d0   bcfacfac-bbb7-11e8-b2ba-ac1f6b6435d0\n301bb49c-bbae-11e8-b2ba-ac1f6b6435d0   f95980c2-bbb4-11e8-b2ba-ac1f6b6435d0\n0858d008-bb9e-11e8-b2b9-ac1f6b6435d0   344ba91e-bbbd-11e8-b2ba-ac1f6b6435d0\n4a2c88e4-bba8-11e8-b2ba-ac1f6b6435d0   f53d5ae6-bbc2-11e8-b2bc-ac1f6b6435d0\n36e73f52-bba5-11e8-b2ba-ac1f6b6435d0   3abf8b36-bba8-11e8-b2ba-ac1f6b6435d0\n2ab9afbe-bbc6-11e8-b2bc-ac1f6b6435d0   50026656-bbb9-11e8-b2ba-ac1f6b6435d0\n6e0b2662-bbad-11e8-b2ba-ac1f6b6435d0   e5b9a9c2-bbbc-11e8-b2ba-ac1f6b6435d0\na0a93e74-bbc6-11e8-b2bc-ac1f6b6435d0   d3877738-bbb3-11e8-b2ba-ac1f6b6435d0\n</code></pre>",
      "rawMarkdown": "More similar image pairs in train data:\n\n    7116d6f4-bba7-11e8-b2ba-ac1f6b6435d0   a166d11a-bbca-11e8-b2bc-ac1f6b6435d0\n    0ebd689e-bbc8-11e8-b2bc-ac1f6b6435d0   826e48c8-bbbc-11e8-b2ba-ac1f6b6435d0\n    1826c3b4-bba3-11e8-b2b9-ac1f6b6435d0   f3db54f0-bbae-11e8-b2ba-ac1f6b6435d0\n    2f7acfa4-bbc3-11e8-b2bc-ac1f6b6435d0   8832c642-bba0-11e8-b2b9-ac1f6b6435d0\n    49366c04-bbb0-11e8-b2ba-ac1f6b6435d0   b4af5e20-bbbd-11e8-b2ba-ac1f6b6435d0\n    a7fcfccc-bbb9-11e8-b2ba-ac1f6b6435d0   d9ad48e0-bb9f-11e8-b2b9-ac1f6b6435d0\n    b1371ab2-bb9f-11e8-b2b9-ac1f6b6435d0   ebfcde90-bbac-11e8-b2ba-ac1f6b6435d0\n    589a93cc-bbca-11e8-b2bc-ac1f6b6435d0   b367962a-bbc9-11e8-b2bc-ac1f6b6435d0\n    a656e5f2-bb9d-11e8-b2b9-ac1f6b6435d0   cf78b532-bbb1-11e8-b2ba-ac1f6b6435d0\n    589a93cc-bbca-11e8-b2bc-ac1f6b6435d0   b367962a-bbc9-11e8-b2bc-ac1f6b6435d0\n    4d3ad8a6-bbc1-11e8-b2bb-ac1f6b6435d0   7927dc50-bba5-11e8-b2ba-ac1f6b6435d0\n    575666ec-bba9-11e8-b2ba-ac1f6b6435d0   63d23f16-bba6-11e8-b2ba-ac1f6b6435d0\n    5dc2623a-bbb6-11e8-b2ba-ac1f6b6435d0   81cea69c-bb9e-11e8-b2b9-ac1f6b6435d0\n    80ea0dc4-bbb3-11e8-b2ba-ac1f6b6435d0   ad2eff22-bba7-11e8-b2ba-ac1f6b6435d0\n    3e608f2e-bba8-11e8-b2ba-ac1f6b6435d0   8779d884-bba6-11e8-b2ba-ac1f6b6435d0\n    d57bc36c-bba6-11e8-b2ba-ac1f6b6435d0   d79a532e-bbca-11e8-b2bc-ac1f6b6435d0\n    73f49160-bbc8-11e8-b2bc-ac1f6b6435d0   bcfacfac-bbb7-11e8-b2ba-ac1f6b6435d0\n    301bb49c-bbae-11e8-b2ba-ac1f6b6435d0   f95980c2-bbb4-11e8-b2ba-ac1f6b6435d0\n    0858d008-bb9e-11e8-b2b9-ac1f6b6435d0   344ba91e-bbbd-11e8-b2ba-ac1f6b6435d0\n    4a2c88e4-bba8-11e8-b2ba-ac1f6b6435d0   f53d5ae6-bbc2-11e8-b2bc-ac1f6b6435d0\n    36e73f52-bba5-11e8-b2ba-ac1f6b6435d0   3abf8b36-bba8-11e8-b2ba-ac1f6b6435d0\n    2ab9afbe-bbc6-11e8-b2bc-ac1f6b6435d0   50026656-bbb9-11e8-b2ba-ac1f6b6435d0\n    6e0b2662-bbad-11e8-b2ba-ac1f6b6435d0   e5b9a9c2-bbbc-11e8-b2ba-ac1f6b6435d0\n    a0a93e74-bbc6-11e8-b2bc-ac1f6b6435d0   d3877738-bbb3-11e8-b2ba-ac1f6b6435d0",
      "votes": null
    },
    {
      "id": "429687",
      "postDate": "11/29/2018 07:29:51",
      "content": "<p>thanks for sharing,anyone thried theses images?</p>",
      "rawMarkdown": "thanks for sharing,anyone thried theses images?",
      "votes": null
    },
    {
      "id": "430876",
      "postDate": "12/01/2018 05:03:10",
      "content": "<p>Matching with External Data:</p>\n\n<p>Extra(GeneID_ Dir_ImageURL) </p>\n\n<p>color = [red,green,blue]</p>\n\n<p>\"<a href=\"http://v18.proteinatlas.org/images/\">http://v18.proteinatlas.org/images/</a>\" + replace(Dir_ImageURL,Dir/ImageURL + _color.jpg)</p>\n\n<p>Test image Id\nSimR,SimG,SimB ... checked similarity channel wise</p>\n\n<p>! Not all images are checked. Selected by threshold.\nif Sim_Red &lt; 14 &amp;&amp; Sim_Gleen &lt; 12 &amp;&amp; Sim_Blue &lt; 10: </p>",
      "rawMarkdown": "Matching with External Data:\n\nExtra(GeneID_ Dir_ImageURL) \n\ncolor = [red,green,blue]\n\n\"http://v18.proteinatlas.org/images/\" + replace(Dir_ImageURL,Dir/ImageURL + _color.jpg)\n\nTest image Id\nSimR,SimG,SimB ... checked similarity channel wise\n\n! Not all images are checked. Selected by threshold.\nif Sim_Red &lt; 14 &amp;&amp; Sim_Gleen &lt; 12 &amp;&amp; Sim_Blue &lt; 10:",
      "votes": null
    },
    {
      "id": "431254",
      "postDate": "12/02/2018 00:13:39",
      "content": "<p>Shoud I delete this list...\nIt leads to cheat, on the other hand,\nI'd thought it's usuful to remove matching data from extra-data-training-set\nand you can better estimate your models generalization ability.\nI'd be appreciated if anyone give me some advice or suggestions.</p>",
      "rawMarkdown": "Shoud I delete this list...\nIt leads to cheat, on the other hand,\nI'd thought it's usuful to remove matching data from extra-data-training-set\nand you can better estimate your models generalization ability.\nI'd be appreciated if anyone give me some advice or suggestions.",
      "votes": null
    },
    {
      "id": "431279",
      "postDate": "12/02/2018 01:31:08",
      "content": "<p>Such \"leaks\" are not uncommon in Kaggle comps and the best course of action is usually to get them out the open.  Thanks very much for sharing Tomomi; I think you did the right thing.   A discussion point that might help us all at this point would be speculation on why these data were not included in the training set.    My best guess is the official data are likely more recent and originally higher resolution, and the organizers wanted to provide a vetted and consistent set.    They surely must have known someone would find the additional HPA data, though (since Kagglers are notoriously good at uncovering such things), so I'm kind of puzzled.</p>",
      "rawMarkdown": "Such \"leaks\" are not uncommon in Kaggle comps and the best course of action is usually to get them out the open.  Thanks very much for sharing Tomomi; I think you did the right thing.   A discussion point that might help us all at this point would be speculation on why these data were not included in the training set.    My best guess is the official data are likely more recent and originally higher resolution, and the organizers wanted to provide a vetted and consistent set.    They surely must have known someone would find the additional HPA data, though (since Kagglers are notoriously good at uncovering such things), so I'm kind of puzzled.",
      "votes": null
    },
    {
      "id": "431296",
      "postDate": "12/02/2018 02:34:52",
      "content": "<p>Thank you @Russ W\nI feel safe, I can sleep well. Thank you so mach for your advice! </p>",
      "rawMarkdown": "Thank you @Russ W\nI feel safe, I can sleep well. Thank you so mach for your advice!",
      "votes": null
    },
    {
      "id": "431304",
      "postDate": "12/02/2018 02:48:07",
      "content": "<p>how do you calculate the SimRed, SimGreen or SimBlue?</p>",
      "rawMarkdown": "how do you calculate the SimRed, SimGreen or SimBlue?",
      "votes": null
    },
    {
      "id": "431320",
      "postDate": "12/02/2018 03:45:36",
      "content": "<p>Hi!, @Paul Chen</p>\n\n<p><a href=\"https://github.com/JohannesBuchner/imagehash\">https://github.com/JohannesBuchner/imagehash</a></p>\n\n<p>test = imagehash.phash(Image.open(TEST+row['Id']+'_red.png'))</p>\n\n<p>hpav18 = imagehash.phash(Image.open(HPAv18+row['Id']+'_red.png'))</p>\n\n<p>SimRed = hpav18 - test</p>",
      "rawMarkdown": "Hi!, @Paul Chen\n\nhttps://github.com/JohannesBuchner/imagehash\n\ntest = imagehash.phash(Image.open(TEST+row['Id']+'_red.png'))\n\nhpav18 = imagehash.phash(Image.open(HPAv18+row['Id']+'_red.png'))\n\nSimRed = hpav18 - test",
      "votes": null
    },
    {
      "id": "443114",
      "postDate": "12/21/2018 03:15:49",
      "content": "<p>hi Tilii:\nThank you for sharing!\nWhat is the impact of \"duplicates\" images in the training set? I think that \"duplicates\" images pair will affect the final model selection.</p>",
      "rawMarkdown": "hi Tilii:\nThank you for sharing!\nWhat is the impact of \"duplicates\" images in the training set? I think that \"duplicates\" images pair will affect the final model selection.",
      "votes": null
    },
    {
      "id": "443146",
      "postDate": "12/21/2018 05:17:35",
      "content": "<blockquote>\n  <p>What is the impact of \"duplicates\" images in the training set?</p>\n</blockquote>\n\n<p><a href=\"/femichen\">@femichen</a> I don't know for sure, but it seems like it won't matter much during the training. I think it is more important for predictions, as there are some similar images in test data that are considerably different in terms of quality. In such cases I would predict the class on better images and transfer that prediction to their mis-shaped twins. See <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74068\"><strong>here</strong></a> for more details.</p>",
      "rawMarkdown": "&gt; What is the impact of \"duplicates\" images in the training set?\n\n@femichen I don't know for sure, but it seems like it won't matter much during the training. I think it is more important for predictions, as there are some similar images in test data that are considerably different in terms of quality. In such cases I would predict the class on better images and transfer that prediction to their mis-shaped twins. See [__here__](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74068) for more details.",
      "votes": null
    },
    {
      "id": "443197",
      "postDate": "12/21/2018 07:43:18",
      "content": "<p>thank you Tilii, I think the way to use similar images in test data is very good.</p>",
      "rawMarkdown": "thank you Tilii, I think the way to use similar images in test data is very good.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 426936,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "11/24/2018 07:15:48",
      "content": "<p>Attachments are here. Rar files were too big for Kaggle (~25 and 50 Mb).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 426941,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "11/24/2018 07:46:21",
      "content": "<p>nice post. I tried a couple you mention and I get the same images (I still don't get the test/train thing of brian - hope I didn't corrupt my data somehow).</p>\n\n<p>One thing that might be interesting: put these images aside before training and then, after training, use the model to classify them. Do they get the same labels?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 426943,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "11/24/2018 08:05:32",
      "content": "<p>Thank you so much, quite helpful post.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 427105,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "11/24/2018 15:36:04",
      "content": "<p>Great share! Thanks a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 427191,
      "author_name": "dstjhb",
      "author_url": "",
      "post_date": "11/24/2018 20:31:08",
      "content": "<p>Oh, thanks so much!</p>\n\n<p>I would keep in mind that if similar numbers of \"duplicates\" appear in the final evaluation set (and follow the the same distribution), then training with the \"duplicate\" images included may actually be better... (?)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 428145,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "11/26/2018 20:41:48",
      "content": "<p>Here is my list of close matches in the training set of 31k images. Identified by image id and group id, if they are in the same group they are similar images.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 428342,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "11/27/2018 05:42:24",
      "content": "<p>More similar images in test data:</p>\n\n<pre><code>0774284e-bad7-11e8-b2b9-ac1f6b6435d0   34aa05de-bad4-11e8-b2b8-ac1f6b6435d0\n8f0666da-bad9-11e8-b2b9-ac1f6b6435d0   dad043dc-bad5-11e8-b2b9-ac1f6b6435d0\n</code></pre>\n\n<p>The latter is a good example of different z-slices.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 428383,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "11/27/2018 07:23:27",
      "content": "<p>More similar image pairs in train data:</p>\n\n<pre><code>7116d6f4-bba7-11e8-b2ba-ac1f6b6435d0   a166d11a-bbca-11e8-b2bc-ac1f6b6435d0\n0ebd689e-bbc8-11e8-b2bc-ac1f6b6435d0   826e48c8-bbbc-11e8-b2ba-ac1f6b6435d0\n1826c3b4-bba3-11e8-b2b9-ac1f6b6435d0   f3db54f0-bbae-11e8-b2ba-ac1f6b6435d0\n2f7acfa4-bbc3-11e8-b2bc-ac1f6b6435d0   8832c642-bba0-11e8-b2b9-ac1f6b6435d0\n49366c04-bbb0-11e8-b2ba-ac1f6b6435d0   b4af5e20-bbbd-11e8-b2ba-ac1f6b6435d0\na7fcfccc-bbb9-11e8-b2ba-ac1f6b6435d0   d9ad48e0-bb9f-11e8-b2b9-ac1f6b6435d0\nb1371ab2-bb9f-11e8-b2b9-ac1f6b6435d0   ebfcde90-bbac-11e8-b2ba-ac1f6b6435d0\n589a93cc-bbca-11e8-b2bc-ac1f6b6435d0   b367962a-bbc9-11e8-b2bc-ac1f6b6435d0\na656e5f2-bb9d-11e8-b2b9-ac1f6b6435d0   cf78b532-bbb1-11e8-b2ba-ac1f6b6435d0\n589a93cc-bbca-11e8-b2bc-ac1f6b6435d0   b367962a-bbc9-11e8-b2bc-ac1f6b6435d0\n4d3ad8a6-bbc1-11e8-b2bb-ac1f6b6435d0   7927dc50-bba5-11e8-b2ba-ac1f6b6435d0\n575666ec-bba9-11e8-b2ba-ac1f6b6435d0   63d23f16-bba6-11e8-b2ba-ac1f6b6435d0\n5dc2623a-bbb6-11e8-b2ba-ac1f6b6435d0   81cea69c-bb9e-11e8-b2b9-ac1f6b6435d0\n80ea0dc4-bbb3-11e8-b2ba-ac1f6b6435d0   ad2eff22-bba7-11e8-b2ba-ac1f6b6435d0\n3e608f2e-bba8-11e8-b2ba-ac1f6b6435d0   8779d884-bba6-11e8-b2ba-ac1f6b6435d0\nd57bc36c-bba6-11e8-b2ba-ac1f6b6435d0   d79a532e-bbca-11e8-b2bc-ac1f6b6435d0\n73f49160-bbc8-11e8-b2bc-ac1f6b6435d0   bcfacfac-bbb7-11e8-b2ba-ac1f6b6435d0\n301bb49c-bbae-11e8-b2ba-ac1f6b6435d0   f95980c2-bbb4-11e8-b2ba-ac1f6b6435d0\n0858d008-bb9e-11e8-b2b9-ac1f6b6435d0   344ba91e-bbbd-11e8-b2ba-ac1f6b6435d0\n4a2c88e4-bba8-11e8-b2ba-ac1f6b6435d0   f53d5ae6-bbc2-11e8-b2bc-ac1f6b6435d0\n36e73f52-bba5-11e8-b2ba-ac1f6b6435d0   3abf8b36-bba8-11e8-b2ba-ac1f6b6435d0\n2ab9afbe-bbc6-11e8-b2bc-ac1f6b6435d0   50026656-bbb9-11e8-b2ba-ac1f6b6435d0\n6e0b2662-bbad-11e8-b2ba-ac1f6b6435d0   e5b9a9c2-bbbc-11e8-b2ba-ac1f6b6435d0\na0a93e74-bbc6-11e8-b2bc-ac1f6b6435d0   d3877738-bbb3-11e8-b2ba-ac1f6b6435d0\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 429687,
      "author_name": "spytensor",
      "author_url": "",
      "post_date": "11/29/2018 07:29:51",
      "content": "<p>thanks for sharing,anyone thried theses images?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 430876,
      "author_name": "tomomimoriyama",
      "author_url": "",
      "post_date": "12/01/2018 05:03:10",
      "content": "<p>Matching with External Data:</p>\n\n<p>Extra(GeneID_ Dir_ImageURL) </p>\n\n<p>color = [red,green,blue]</p>\n\n<p>\"<a href=\"http://v18.proteinatlas.org/images/\">http://v18.proteinatlas.org/images/</a>\" + replace(Dir_ImageURL,Dir/ImageURL + _color.jpg)</p>\n\n<p>Test image Id\nSimR,SimG,SimB ... checked similarity channel wise</p>\n\n<p>! Not all images are checked. Selected by threshold.\nif Sim_Red &lt; 14 &amp;&amp; Sim_Gleen &lt; 12 &amp;&amp; Sim_Blue &lt; 10: </p>",
      "votes": null,
      "replies": [
        {
          "id": 431254,
          "author_name": "tomomimoriyama",
          "author_url": "",
          "post_date": "12/02/2018 00:13:39",
          "content": "<p>Shoud I delete this list...\nIt leads to cheat, on the other hand,\nI'd thought it's usuful to remove matching data from extra-data-training-set\nand you can better estimate your models generalization ability.\nI'd be appreciated if anyone give me some advice or suggestions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 431279,
          "author_name": "sasrdw",
          "author_url": "",
          "post_date": "12/02/2018 01:31:08",
          "content": "<p>Such \"leaks\" are not uncommon in Kaggle comps and the best course of action is usually to get them out the open.  Thanks very much for sharing Tomomi; I think you did the right thing.   A discussion point that might help us all at this point would be speculation on why these data were not included in the training set.    My best guess is the official data are likely more recent and originally higher resolution, and the organizers wanted to provide a vetted and consistent set.    They surely must have known someone would find the additional HPA data, though (since Kagglers are notoriously good at uncovering such things), so I'm kind of puzzled.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 431296,
          "author_name": "tomomimoriyama",
          "author_url": "",
          "post_date": "12/02/2018 02:34:52",
          "content": "<p>Thank you @Russ W\nI feel safe, I can sleep well. Thank you so mach for your advice! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 431304,
          "author_name": "dxchen",
          "author_url": "",
          "post_date": "12/02/2018 02:48:07",
          "content": "<p>how do you calculate the SimRed, SimGreen or SimBlue?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 431320,
          "author_name": "tomomimoriyama",
          "author_url": "",
          "post_date": "12/02/2018 03:45:36",
          "content": "<p>Hi!, @Paul Chen</p>\n\n<p><a href=\"https://github.com/JohannesBuchner/imagehash\">https://github.com/JohannesBuchner/imagehash</a></p>\n\n<p>test = imagehash.phash(Image.open(TEST+row['Id']+'_red.png'))</p>\n\n<p>hpav18 = imagehash.phash(Image.open(HPAv18+row['Id']+'_red.png'))</p>\n\n<p>SimRed = hpav18 - test</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 443114,
      "author_name": "femichen",
      "author_url": "",
      "post_date": "12/21/2018 03:15:49",
      "content": "<p>hi Tilii:\nThank you for sharing!\nWhat is the impact of \"duplicates\" images in the training set? I think that \"duplicates\" images pair will affect the final model selection.</p>",
      "votes": null,
      "replies": [
        {
          "id": 443146,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "12/21/2018 05:17:35",
          "content": "<blockquote>\n  <p>What is the impact of \"duplicates\" images in the training set?</p>\n</blockquote>\n\n<p><a href=\"/femichen\">@femichen</a> I don't know for sure, but it seems like it won't matter much during the training. I think it is more important for predictions, as there are some similar images in test data that are considerably different in terms of quality. In such cases I would predict the class on better images and transfer that prediction to their mis-shaped twins. See <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74068\"><strong>here</strong></a> for more details.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443197,
          "author_name": "femichen",
          "author_url": "",
          "post_date": "12/21/2018 07:43:18",
          "content": "<p>thank you Tilii, I think the way to use similar images in test data is very good.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "426935": "It seems that there is interest in this topic as it has been brought up [__here__](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69594), but not fully described.\n\nThe fastest way I know to find similar images is by image hashes - there is a nice Python implementation in [__imagehash__](https://github.com/JohannesBuchner/imagehash) package. I have done it only on green images using `average hashing` and `perception hashing`. The latter is more reliable when it comes to finding true positive matches, while the former may be better at finding similar images at the expense of few bogus matches.\n\nFirst, here are two image pairs identified by Brian in [__this post__](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69594).\n\nIn train:\n![enter image description here][1]\n\nIn test:\n![enter image description here][2]\n\nTrain images seem identical but don't have the same brightness/amplitude. In microscopy parlance, that would mean the same field of cells but different exposure, or possibly two subsequent z-stacks during image acquisition. Test images do seem identical.\n\nI will attach my lists as text files. Their format is two columns of names like this:\n\n    565dbefe-bbc0-11e8-b2bb-ac1f6b6435d0   62c829c8-bbba-11e8-b2ba-ac1f6b6435d0\n    0940c9ee-bbb7-11e8-b2ba-ac1f6b6435d0   40ba3872-bbc8-11e8-b2bc-ac1f6b6435d0\n    3441eab0-bbb2-11e8-b2ba-ac1f6b6435d0   c4908034-bb9b-11e8-b2b9-ac1f6b6435d0\n    90095b92-bba0-11e8-b2b9-ac1f6b6435d0   f46bceb2-bbb4-11e8-b2ba-ac1f6b6435d0\n\nI also include rar-compressed side-by-side comparisons of all found pairs. Here are few examples that are identical according to `phash`:\n\n![enter image description here][3]\n![enter image description here][4]\n\nThere is a slight difference in brightness, but it should be obvious that they are near-identical.\n\nHere are couple of duplicates that are truly identical:\n\n![enter image description here][5]\n![enter image description here][6]\n\nAnd an example of similar pairs that `ahash` finds but `phash` doesn't, which would indicate that they are similar rather than identical. To my eye they seem different exposures of the same image.\n\n![enter image description here][7]\n\nI could not find any images between train and test that were identical - only within the two groups.\n\nIf someone feels like playing with other hashes and finds other duplicates, please report them here.\n\n\n  [1]: https://i.postimg.cc/qMyxxzv8/example-train.png\n  [2]: https://i.postimg.cc/zXCTR6x7/example-test.png\n  [3]: https://i.postimg.cc/cLjDFVN3/train-identical-097.png\n  [4]: https://i.postimg.cc/4yBPV05m/train-identical-094.png\n  [5]: https://i.postimg.cc/Bv48fvsf/train-identical-092.png\n  [6]: https://i.postimg.cc/wTkMcpLW/test-identical-009.png\n  [7]: https://i.postimg.cc/BnsHmchJ/train-similar-041.png",
    "426936": "Attachments are here. Rar files were too big for Kaggle (~25 and 50 Mb).",
    "426941": "nice post. I tried a couple you mention and I get the same images (I still don't get the test/train thing of brian - hope I didn't corrupt my data somehow).\n\nOne thing that might be interesting: put these images aside before training and then, after training, use the model to classify them. Do they get the same labels?",
    "426943": "Thank you so much, quite helpful post.",
    "427105": "Great share! Thanks a lot!",
    "427191": "Oh, thanks so much!\n\nI would keep in mind that if similar numbers of \"duplicates\" appear in the final evaluation set (and follow the the same distribution), then training with the \"duplicate\" images included may actually be better... (?)",
    "428145": "Here is my list of close matches in the training set of 31k images. Identified by image id and group id, if they are in the same group they are similar images.",
    "428342": "More similar images in test data:\n\n    0774284e-bad7-11e8-b2b9-ac1f6b6435d0   34aa05de-bad4-11e8-b2b8-ac1f6b6435d0\n    8f0666da-bad9-11e8-b2b9-ac1f6b6435d0   dad043dc-bad5-11e8-b2b9-ac1f6b6435d0\n\nThe latter is a good example of different z-slices.",
    "428383": "More similar image pairs in train data:\n\n    7116d6f4-bba7-11e8-b2ba-ac1f6b6435d0   a166d11a-bbca-11e8-b2bc-ac1f6b6435d0\n    0ebd689e-bbc8-11e8-b2bc-ac1f6b6435d0   826e48c8-bbbc-11e8-b2ba-ac1f6b6435d0\n    1826c3b4-bba3-11e8-b2b9-ac1f6b6435d0   f3db54f0-bbae-11e8-b2ba-ac1f6b6435d0\n    2f7acfa4-bbc3-11e8-b2bc-ac1f6b6435d0   8832c642-bba0-11e8-b2b9-ac1f6b6435d0\n    49366c04-bbb0-11e8-b2ba-ac1f6b6435d0   b4af5e20-bbbd-11e8-b2ba-ac1f6b6435d0\n    a7fcfccc-bbb9-11e8-b2ba-ac1f6b6435d0   d9ad48e0-bb9f-11e8-b2b9-ac1f6b6435d0\n    b1371ab2-bb9f-11e8-b2b9-ac1f6b6435d0   ebfcde90-bbac-11e8-b2ba-ac1f6b6435d0\n    589a93cc-bbca-11e8-b2bc-ac1f6b6435d0   b367962a-bbc9-11e8-b2bc-ac1f6b6435d0\n    a656e5f2-bb9d-11e8-b2b9-ac1f6b6435d0   cf78b532-bbb1-11e8-b2ba-ac1f6b6435d0\n    589a93cc-bbca-11e8-b2bc-ac1f6b6435d0   b367962a-bbc9-11e8-b2bc-ac1f6b6435d0\n    4d3ad8a6-bbc1-11e8-b2bb-ac1f6b6435d0   7927dc50-bba5-11e8-b2ba-ac1f6b6435d0\n    575666ec-bba9-11e8-b2ba-ac1f6b6435d0   63d23f16-bba6-11e8-b2ba-ac1f6b6435d0\n    5dc2623a-bbb6-11e8-b2ba-ac1f6b6435d0   81cea69c-bb9e-11e8-b2b9-ac1f6b6435d0\n    80ea0dc4-bbb3-11e8-b2ba-ac1f6b6435d0   ad2eff22-bba7-11e8-b2ba-ac1f6b6435d0\n    3e608f2e-bba8-11e8-b2ba-ac1f6b6435d0   8779d884-bba6-11e8-b2ba-ac1f6b6435d0\n    d57bc36c-bba6-11e8-b2ba-ac1f6b6435d0   d79a532e-bbca-11e8-b2bc-ac1f6b6435d0\n    73f49160-bbc8-11e8-b2bc-ac1f6b6435d0   bcfacfac-bbb7-11e8-b2ba-ac1f6b6435d0\n    301bb49c-bbae-11e8-b2ba-ac1f6b6435d0   f95980c2-bbb4-11e8-b2ba-ac1f6b6435d0\n    0858d008-bb9e-11e8-b2b9-ac1f6b6435d0   344ba91e-bbbd-11e8-b2ba-ac1f6b6435d0\n    4a2c88e4-bba8-11e8-b2ba-ac1f6b6435d0   f53d5ae6-bbc2-11e8-b2bc-ac1f6b6435d0\n    36e73f52-bba5-11e8-b2ba-ac1f6b6435d0   3abf8b36-bba8-11e8-b2ba-ac1f6b6435d0\n    2ab9afbe-bbc6-11e8-b2bc-ac1f6b6435d0   50026656-bbb9-11e8-b2ba-ac1f6b6435d0\n    6e0b2662-bbad-11e8-b2ba-ac1f6b6435d0   e5b9a9c2-bbbc-11e8-b2ba-ac1f6b6435d0\n    a0a93e74-bbc6-11e8-b2bc-ac1f6b6435d0   d3877738-bbb3-11e8-b2ba-ac1f6b6435d0",
    "429687": "thanks for sharing,anyone thried theses images?",
    "430876": "Matching with External Data:\n\nExtra(GeneID_ Dir_ImageURL) \n\ncolor = [red,green,blue]\n\n\"http://v18.proteinatlas.org/images/\" + replace(Dir_ImageURL,Dir/ImageURL + _color.jpg)\n\nTest image Id\nSimR,SimG,SimB ... checked similarity channel wise\n\n! Not all images are checked. Selected by threshold.\nif Sim_Red &lt; 14 &amp;&amp; Sim_Gleen &lt; 12 &amp;&amp; Sim_Blue &lt; 10:",
    "431254": "Shoud I delete this list...\nIt leads to cheat, on the other hand,\nI'd thought it's usuful to remove matching data from extra-data-training-set\nand you can better estimate your models generalization ability.\nI'd be appreciated if anyone give me some advice or suggestions.",
    "431279": "Such \"leaks\" are not uncommon in Kaggle comps and the best course of action is usually to get them out the open.  Thanks very much for sharing Tomomi; I think you did the right thing.   A discussion point that might help us all at this point would be speculation on why these data were not included in the training set.    My best guess is the official data are likely more recent and originally higher resolution, and the organizers wanted to provide a vetted and consistent set.    They surely must have known someone would find the additional HPA data, though (since Kagglers are notoriously good at uncovering such things), so I'm kind of puzzled.",
    "431296": "Thank you @Russ W\nI feel safe, I can sleep well. Thank you so mach for your advice!",
    "431304": "how do you calculate the SimRed, SimGreen or SimBlue?",
    "431320": "Hi!, @Paul Chen\n\nhttps://github.com/JohannesBuchner/imagehash\n\ntest = imagehash.phash(Image.open(TEST+row['Id']+'_red.png'))\n\nhpav18 = imagehash.phash(Image.open(HPAv18+row['Id']+'_red.png'))\n\nSimRed = hpav18 - test",
    "443114": "hi Tilii:\nThank you for sharing!\nWhat is the impact of \"duplicates\" images in the training set? I think that \"duplicates\" images pair will affect the final model selection.",
    "443146": "&gt; What is the impact of \"duplicates\" images in the training set?\n\n@femichen I don't know for sure, but it seems like it won't matter much during the training. I think it is more important for predictions, as there are some similar images in test data that are considerably different in terms of quality. In such cases I would predict the class on better images and transfer that prediction to their mis-shaped twins. See [__here__](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74068) for more details.",
    "443197": "thank you Tilii, I think the way to use similar images in test data is very good."
  },
  "source": "meta"
}