{
  "id": 323906,
  "title": "Data leakage (image duplicates)",
  "url": "/competitions/herbarium-2022-fgvc9/discussion/323906",
  "author_name": "amirus",
  "post_date": "2022-05-09T06:19:02.278000",
  "votes": 0,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi, while exploring the dataset we have discovered some duplicated images between train-test splits. It seems like there are thousands of such duplicates…</p>\n<p>See a couple of examples below:<br>\n<a href=\"https://www.dropbox.com/s/rdx717r7tsfwjne/duplicate_example1.jpeg?dl=0\" target=\"_blank\">https://www.dropbox.com/s/rdx717r7tsfwjne/duplicate_example1.jpeg?dl=0</a><br>\n<a href=\"https://www.dropbox.com/s/nlxe1eljm9nyucn/duplicate_example2.jpeg?dl=0\" target=\"_blank\">https://www.dropbox.com/s/nlxe1eljm9nyucn/duplicate_example2.jpeg?dl=0</a></p>\n<p>I could share code for finding these duplicates if anyone is interested to validate our findings. </p>",
  "messages": [
    {
      "id": 1781992,
      "postDate": "2022-05-09T06:19:02.280Z",
      "content": "<p>Hi, while exploring the dataset we have discovered some duplicated images between train-test splits. It seems like there are thousands of such duplicates…</p>\n<p>See a couple of examples below:<br>\n<a href=\"https://www.dropbox.com/s/rdx717r7tsfwjne/duplicate_example1.jpeg?dl=0\" target=\"_blank\">https://www.dropbox.com/s/rdx717r7tsfwjne/duplicate_example1.jpeg?dl=0</a><br>\n<a href=\"https://www.dropbox.com/s/nlxe1eljm9nyucn/duplicate_example2.jpeg?dl=0\" target=\"_blank\">https://www.dropbox.com/s/nlxe1eljm9nyucn/duplicate_example2.jpeg?dl=0</a></p>\n<p>I could share code for finding these duplicates if anyone is interested to validate our findings. </p>",
      "rawMarkdown": "Hi, while exploring the dataset we have discovered some duplicated images between train-test splits. It seems like there are thousands of such duplicates...\n\nSee a couple of examples below:\nhttps://www.dropbox.com/s/rdx717r7tsfwjne/duplicate_example1.jpeg?dl=0\nhttps://www.dropbox.com/s/nlxe1eljm9nyucn/duplicate_example2.jpeg?dl=0\n\nI could share code for finding these duplicates if anyone is interested to validate our findings. \n"
    },
    {
      "id": 1782704,
      "postDate": "2022-05-09T19:41:20.123Z",
      "content": "<p>p.s. Regarding runtime it should be less than half an hour on a 32 core ec2 machine, total computation cost was 0.69$.</p>",
      "rawMarkdown": "p.s. Regarding runtime it should be less than half an hour on a 32 core ec2 machine, total computation cost was 0.69$."
    },
    {
      "id": 1782701,
      "postDate": "2022-05-09T19:37:22.100Z",
      "content": "<p><a href=\"https://www.kaggle.com/parkjohnychae\" target=\"_blank\">@parkjohnychae</a> <a href=\"https://www.kaggle.com/satoshidatamoto\" target=\"_blank\">@satoshidatamoto</a> <br>\nWe would be happy to share our code. <br>\nWe haven't built any website so you're welcome to email us directly and we'll send you the code for free.<br>\nIt basically allows you to find duplicates/near-duplicates in any image dataset.</p>\n<p>The processing time on this dataset (1,050,179) images) is:</p>\n<ul>\n<li>5,121 seconds on a 8 cores,32GB RAM GCP instance (N2-standard-8).</li>\n<li>1,598 seconds on a 32 cores, 128GB RAM GCP instance (N2-standard-32).</li>\n</ul>\n<p>We are still exploring and have only preliminary results:<br>\nThere are around 9,892 duplicated pairs (identical and near identical images).<br>\n~10K duplicated images in the training set only.<br>\n~500 duplicated images in the test set only<br>\n~6K duplicated images that are both in the train and test.</p>\n<p>We are still validating these but it seems to be roughly the numbers.</p>\n<p>Our emails are:<br>\n<a>amiralush@gmail.com</a> or <a>danny.bickson@gmail.com</a></p>",
      "rawMarkdown": "@parkjohnychae @satoshidatamoto \nWe would be happy to share our code. \nWe haven't built any website so you're welcome to email us directly and we'll send you the code for free.\nIt basically allows you to find duplicates/near-duplicates in any image dataset.\n\nThe processing time on this dataset (1,050,179) images) is:\n- 5,121 seconds on a 8 cores,32GB RAM GCP instance (N2-standard-8).\n- 1,598 seconds on a 32 cores, 128GB RAM GCP instance (N2-standard-32).\n\nWe are still exploring and have only preliminary results:\nThere are around 9,892 duplicated pairs (identical and near identical images).\n~10K duplicated images in the training set only.\n~500 duplicated images in the test set only\n~6K duplicated images that are both in the train and test.\n\nWe are still validating these but it seems to be roughly the numbers.\n\nOur emails are:\namiralush@gmail.com or danny.bickson@gmail.com",
      "replies": [
        {
          "id": 1782949,
          "postDate": "2022-05-10T02:36:54.930Z",
          "content": "<p>It would be great if you could share the code through the Kaggle notebook. No pressure, though!</p>\n<p>The image overlap rate is quite low (&lt;1%), even if you take your current estimation as a face value. And I guess it would take extra effort to validate your method. Since that's a small portion, it may not be much of an issue in terms of the competition, because it will affect all participants equally by only a small margin. We think this competition is fair in terms of that. Likewise, validating all 10K images might be too much unnecessary work, unless you are trying to benchmark your code. But again, thanks for pointing it out!</p>",
          "rawMarkdown": "It would be great if you could share the code through the Kaggle notebook. No pressure, though!\n\nThe image overlap rate is quite low (<1%), even if you take your current estimation as a face value. And I guess it would take extra effort to validate your method. Since that's a small portion, it may not be much of an issue in terms of the competition, because it will affect all participants equally by only a small margin. We think this competition is fair in terms of that. Likewise, validating all 10K images might be too much unnecessary work, unless you are trying to benchmark your code. But again, thanks for pointing it out!",
          "votes": 1
        },
        {
          "id": 1783956,
          "postDate": "2022-05-10T19:50:36.590Z",
          "content": "<p>Hi,<br>\nWe did a deeper analysis with a visual validation on the top ranking 10K identical pairs and here are the results:<br>\nThere are at-least 9,982 identical pairs, probably the number is higher.<br>\ntrain-train    6,284<br>\ntest-train     3,363  --&gt; ~1.7% of the test set.<br>\ntest-test       335</p>\n<p>In my opinion you can't really know what will be the effect on the results. I would be surprised if it will affect all participants equally and we are talking about a very small margin in the leaderboard between the participants.</p>\n<p>Here are the top 10K identical/near-identical pairs. You can see that the \"less\" similar batch (.9001-10000) looks pretty strong, so I guess you could also decrease this threshold and find more similar images.</p>\n<p><a href=\"https://www.dropbox.com/s/w3f0msk3q9z16bq/similarity_herbarium-2022-fgvc9_1_1000.html?dl=0\" target=\"_blank\">https://www.dropbox.com/s/w3f0msk3q9z16bq/similarity_herbarium-2022-fgvc9_1_1000.html?dl=0</a><br>\n<a href=\"https://www.dropbox.com/s/y0ug0hejomxv47y/similarity_herbarium-2022-fgvc9_1001_2000.html?dl=0\" target=\"_blank\">https://www.dropbox.com/s/y0ug0hejomxv47y/similarity_herbarium-2022-fgvc9_1001_2000.html?dl=0</a><br>\n<a href=\"https://www.dropbox.com/s/4nd8dqs2r3vzjyt/similarity_herbarium-2022-fgvc9_2001_3000.html?dl=0\" target=\"_blank\">https://www.dropbox.com/s/4nd8dqs2r3vzjyt/similarity_herbarium-2022-fgvc9_2001_3000.html?dl=0</a><br>\n<a href=\"https://www.dropbox.com/s/f74pzrhdn0fkyiu/similarity_herbarium-2022-fgvc9_3001_4000.html?dl=0\" target=\"_blank\">https://www.dropbox.com/s/f74pzrhdn0fkyiu/similarity_herbarium-2022-fgvc9_3001_4000.html?dl=0</a><br>\n<a href=\"https://www.dropbox.com/s/922sz1ibizgpggu/similarity_herbarium-2022-fgvc9_4001_5000.html?dl=0\" target=\"_blank\">https://www.dropbox.com/s/922sz1ibizgpggu/similarity_herbarium-2022-fgvc9_4001_5000.html?dl=0</a><br>\n<a href=\"https://www.dropbox.com/s/0jn5381588yq0ef/similarity_herbarium-2022-fgvc9_5001_6000.html?dl=0\" target=\"_blank\">https://www.dropbox.com/s/0jn5381588yq0ef/similarity_herbarium-2022-fgvc9_5001_6000.html?dl=0</a><br>\n<a href=\"https://www.dropbox.com/s/k9u3g28jo7thtxz/similarity_herbarium-2022-fgvc9_6001_7000.html?dl=0\" target=\"_blank\">https://www.dropbox.com/s/k9u3g28jo7thtxz/similarity_herbarium-2022-fgvc9_6001_7000.html?dl=0</a><br>\n<a href=\"https://www.dropbox.com/s/g2cfby6g8oyj94p/similarity_herbarium-2022-fgvc9_7001_8000.html?dl=0\" target=\"_blank\">https://www.dropbox.com/s/g2cfby6g8oyj94p/similarity_herbarium-2022-fgvc9_7001_8000.html?dl=0</a><br>\n<a href=\"https://www.dropbox.com/s/yk5t81u3v9lr73e/similarity_herbarium-2022-fgvc9_8001_9000.html?dl=0\" target=\"_blank\">https://www.dropbox.com/s/yk5t81u3v9lr73e/similarity_herbarium-2022-fgvc9_8001_9000.html?dl=0</a><br>\n<a href=\"https://www.dropbox.com/s/56l8da19c3c1s88/similarity_herbarium-2022-fgvc9_9001_10000.html?dl=0\" target=\"_blank\">https://www.dropbox.com/s/56l8da19c3c1s88/similarity_herbarium-2022-fgvc9_9001_10000.html?dl=0</a></p>",
          "rawMarkdown": "Hi,\nWe did a deeper analysis with a visual validation on the top ranking 10K identical pairs and here are the results:\nThere are at-least 9,982 identical pairs, probably the number is higher.\ntrain-train    6,284\ntest-train     3,363  --> ~1.7% of the test set.\ntest-test       335\n\nIn my opinion you can't really know what will be the effect on the results. I would be surprised if it will affect all participants equally and we are talking about a very small margin in the leaderboard between the participants.\n\nHere are the top 10K identical/near-identical pairs. You can see that the \"less\" similar batch (.9001-10000) looks pretty strong, so I guess you could also decrease this threshold and find more similar images.\n\nhttps://www.dropbox.com/s/w3f0msk3q9z16bq/similarity_herbarium-2022-fgvc9_1_1000.html?dl=0\nhttps://www.dropbox.com/s/y0ug0hejomxv47y/similarity_herbarium-2022-fgvc9_1001_2000.html?dl=0\nhttps://www.dropbox.com/s/4nd8dqs2r3vzjyt/similarity_herbarium-2022-fgvc9_2001_3000.html?dl=0\nhttps://www.dropbox.com/s/f74pzrhdn0fkyiu/similarity_herbarium-2022-fgvc9_3001_4000.html?dl=0\nhttps://www.dropbox.com/s/922sz1ibizgpggu/similarity_herbarium-2022-fgvc9_4001_5000.html?dl=0\nhttps://www.dropbox.com/s/0jn5381588yq0ef/similarity_herbarium-2022-fgvc9_5001_6000.html?dl=0\nhttps://www.dropbox.com/s/k9u3g28jo7thtxz/similarity_herbarium-2022-fgvc9_6001_7000.html?dl=0\nhttps://www.dropbox.com/s/g2cfby6g8oyj94p/similarity_herbarium-2022-fgvc9_7001_8000.html?dl=0\nhttps://www.dropbox.com/s/yk5t81u3v9lr73e/similarity_herbarium-2022-fgvc9_8001_9000.html?dl=0\nhttps://www.dropbox.com/s/56l8da19c3c1s88/similarity_herbarium-2022-fgvc9_9001_10000.html?dl=0"
        }
      ]
    },
    {
      "id": 1782679,
      "postDate": "2022-05-09T19:07:56.147Z",
      "content": "<p>Hi there - thanks for sharing your findings! This is definitely something worth exploring further. I'm curious to know if you've found any duplicates in the training data as well? If so, it could be indicative of a larger data leakage issue. In any case, thanks again for bringing this to everyone's attention.</p>",
      "rawMarkdown": "Hi there - thanks for sharing your findings! This is definitely something worth exploring further. I'm curious to know if you've found any duplicates in the training data as well? If so, it could be indicative of a larger data leakage issue. In any case, thanks again for bringing this to everyone's attention."
    },
    {
      "id": 1782605,
      "postDate": "2022-05-09T17:52:28.520Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/amirus\" target=\"_blank\">@amirus</a>! Thank you very much for identifying this problem. </p>\n<p>I checked the examples you shared with us. Those images are actually different images of the same specimens; they are treated differently in the record with different unique identifiers. There could be several reasons why duplicate images are treated as such in the herbarium perspective… This particular case seems to be related to plant samples in the packet. But still, it would be good to know the extent of overlapping images. </p>\n<p>Would you please share the code for finding the duplicates? What kind of hashing algorithm did you use? I think it would be nice for us to know exactly how many images may be overlapping between the test and the training dataset. </p>",
      "rawMarkdown": "Hi @amirus! Thank you very much for identifying this problem. \n\nI checked the examples you shared with us. Those images are actually different images of the same specimens; they are treated differently in the record with different unique identifiers. There could be several reasons why duplicate images are treated as such in the herbarium perspective... This particular case seems to be related to plant samples in the packet. But still, it would be good to know the extent of overlapping images. \n\nWould you please share the code for finding the duplicates? What kind of hashing algorithm did you use? I think it would be nice for us to know exactly how many images may be overlapping between the test and the training dataset. \n\n "
    },
    {
      "id": 1783834,
      "postDate": "2022-05-10T18:15:04.810Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1782704,
      "author_name": "GraphLab",
      "author_url": "",
      "post_date": "2022-05-09T19:41:20.123000",
      "content": "<p>p.s. Regarding runtime it should be less than half an hour on a 32 core ec2 machine, total computation cost was 0.69$.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1782701,
      "author_name": "amirus",
      "author_url": "",
      "post_date": "2022-05-09T19:37:22.100000",
      "content": "<p><a href=\"https://www.kaggle.com/parkjohnychae\" target=\"_blank\">@parkjohnychae</a> <a href=\"https://www.kaggle.com/satoshidatamoto\" target=\"_blank\">@satoshidatamoto</a> <br>\nWe would be happy to share our code. <br>\nWe haven't built any website so you're welcome to email us directly and we'll send you the code for free.<br>\nIt basically allows you to find duplicates/near-duplicates in any image dataset.</p>\n<p>The processing time on this dataset (1,050,179) images) is:</p>\n<ul>\n<li>5,121 seconds on a 8 cores,32GB RAM GCP instance (N2-standard-8).</li>\n<li>1,598 seconds on a 32 cores, 128GB RAM GCP instance (N2-standard-32).</li>\n</ul>\n<p>We are still exploring and have only preliminary results:<br>\nThere are around 9,892 duplicated pairs (identical and near identical images).<br>\n~10K duplicated images in the training set only.<br>\n~500 duplicated images in the test set only<br>\n~6K duplicated images that are both in the train and test.</p>\n<p>We are still validating these but it seems to be roughly the numbers.</p>\n<p>Our emails are:<br>\n<a>amiralush@gmail.com</a> or <a>danny.bickson@gmail.com</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1782949,
          "author_name": "John Park",
          "author_url": "",
          "post_date": "2022-05-10T02:36:54.930000",
          "content": "<p>It would be great if you could share the code through the Kaggle notebook. No pressure, though!</p>\n<p>The image overlap rate is quite low (&lt;1%), even if you take your current estimation as a face value. And I guess it would take extra effort to validate your method. Since that's a small portion, it may not be much of an issue in terms of the competition, because it will affect all participants equally by only a small margin. We think this competition is fair in terms of that. Likewise, validating all 10K images might be too much unnecessary work, unless you are trying to benchmark your code. But again, thanks for pointing it out!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1783956,
          "author_name": "amirus",
          "author_url": "",
          "post_date": "2022-05-10T19:50:36.590000",
          "content": "<p>Hi,<br>\nWe did a deeper analysis with a visual validation on the top ranking 10K identical pairs and here are the results:<br>\nThere are at-least 9,982 identical pairs, probably the number is higher.<br>\ntrain-train    6,284<br>\ntest-train     3,363  --&gt; ~1.7% of the test set.<br>\ntest-test       335</p>\n<p>In my opinion you can't really know what will be the effect on the results. I would be surprised if it will affect all participants equally and we are talking about a very small margin in the leaderboard between the participants.</p>\n<p>Here are the top 10K identical/near-identical pairs. You can see that the \"less\" similar batch (.9001-10000) looks pretty strong, so I guess you could also decrease this threshold and find more similar images.</p>\n<p><a href=\"https://www.dropbox.com/s/w3f0msk3q9z16bq/similarity_herbarium-2022-fgvc9_1_1000.html?dl=0\" target=\"_blank\">https://www.dropbox.com/s/w3f0msk3q9z16bq/similarity_herbarium-2022-fgvc9_1_1000.html?dl=0</a><br>\n<a href=\"https://www.dropbox.com/s/y0ug0hejomxv47y/similarity_herbarium-2022-fgvc9_1001_2000.html?dl=0\" target=\"_blank\">https://www.dropbox.com/s/y0ug0hejomxv47y/similarity_herbarium-2022-fgvc9_1001_2000.html?dl=0</a><br>\n<a href=\"https://www.dropbox.com/s/4nd8dqs2r3vzjyt/similarity_herbarium-2022-fgvc9_2001_3000.html?dl=0\" target=\"_blank\">https://www.dropbox.com/s/4nd8dqs2r3vzjyt/similarity_herbarium-2022-fgvc9_2001_3000.html?dl=0</a><br>\n<a href=\"https://www.dropbox.com/s/f74pzrhdn0fkyiu/similarity_herbarium-2022-fgvc9_3001_4000.html?dl=0\" target=\"_blank\">https://www.dropbox.com/s/f74pzrhdn0fkyiu/similarity_herbarium-2022-fgvc9_3001_4000.html?dl=0</a><br>\n<a href=\"https://www.dropbox.com/s/922sz1ibizgpggu/similarity_herbarium-2022-fgvc9_4001_5000.html?dl=0\" target=\"_blank\">https://www.dropbox.com/s/922sz1ibizgpggu/similarity_herbarium-2022-fgvc9_4001_5000.html?dl=0</a><br>\n<a href=\"https://www.dropbox.com/s/0jn5381588yq0ef/similarity_herbarium-2022-fgvc9_5001_6000.html?dl=0\" target=\"_blank\">https://www.dropbox.com/s/0jn5381588yq0ef/similarity_herbarium-2022-fgvc9_5001_6000.html?dl=0</a><br>\n<a href=\"https://www.dropbox.com/s/k9u3g28jo7thtxz/similarity_herbarium-2022-fgvc9_6001_7000.html?dl=0\" target=\"_blank\">https://www.dropbox.com/s/k9u3g28jo7thtxz/similarity_herbarium-2022-fgvc9_6001_7000.html?dl=0</a><br>\n<a href=\"https://www.dropbox.com/s/g2cfby6g8oyj94p/similarity_herbarium-2022-fgvc9_7001_8000.html?dl=0\" target=\"_blank\">https://www.dropbox.com/s/g2cfby6g8oyj94p/similarity_herbarium-2022-fgvc9_7001_8000.html?dl=0</a><br>\n<a href=\"https://www.dropbox.com/s/yk5t81u3v9lr73e/similarity_herbarium-2022-fgvc9_8001_9000.html?dl=0\" target=\"_blank\">https://www.dropbox.com/s/yk5t81u3v9lr73e/similarity_herbarium-2022-fgvc9_8001_9000.html?dl=0</a><br>\n<a href=\"https://www.dropbox.com/s/56l8da19c3c1s88/similarity_herbarium-2022-fgvc9_9001_10000.html?dl=0\" target=\"_blank\">https://www.dropbox.com/s/56l8da19c3c1s88/similarity_herbarium-2022-fgvc9_9001_10000.html?dl=0</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1782679,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-09T19:07:56.147000",
      "content": "<p>Hi there - thanks for sharing your findings! This is definitely something worth exploring further. I'm curious to know if you've found any duplicates in the training data as well? If so, it could be indicative of a larger data leakage issue. In any case, thanks again for bringing this to everyone's attention.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1782605,
      "author_name": "John Park",
      "author_url": "",
      "post_date": "2022-05-09T17:52:28.520000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/amirus\" target=\"_blank\">@amirus</a>! Thank you very much for identifying this problem. </p>\n<p>I checked the examples you shared with us. Those images are actually different images of the same specimens; they are treated differently in the record with different unique identifiers. There could be several reasons why duplicate images are treated as such in the herbarium perspective… This particular case seems to be related to plant samples in the packet. But still, it would be good to know the extent of overlapping images. </p>\n<p>Would you please share the code for finding the duplicates? What kind of hashing algorithm did you use? I think it would be nice for us to know exactly how many images may be overlapping between the test and the training dataset. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1783834,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T18:15:04.810000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1781992": "Hi, while exploring the dataset we have discovered some duplicated images between train-test splits. It seems like there are thousands of such duplicates...\n\nSee a couple of examples below:\nhttps://www.dropbox.com/s/rdx717r7tsfwjne/duplicate_example1.jpeg?dl=0\nhttps://www.dropbox.com/s/nlxe1eljm9nyucn/duplicate_example2.jpeg?dl=0\n\nI could share code for finding these duplicates if anyone is interested to validate our findings. \n",
    "1782704": "p.s. Regarding runtime it should be less than half an hour on a 32 core ec2 machine, total computation cost was 0.69$.",
    "1782701": "@parkjohnychae @satoshidatamoto \nWe would be happy to share our code. \nWe haven't built any website so you're welcome to email us directly and we'll send you the code for free.\nIt basically allows you to find duplicates/near-duplicates in any image dataset.\n\nThe processing time on this dataset (1,050,179) images) is:\n- 5,121 seconds on a 8 cores,32GB RAM GCP instance (N2-standard-8).\n- 1,598 seconds on a 32 cores, 128GB RAM GCP instance (N2-standard-32).\n\nWe are still exploring and have only preliminary results:\nThere are around 9,892 duplicated pairs (identical and near identical images).\n~10K duplicated images in the training set only.\n~500 duplicated images in the test set only\n~6K duplicated images that are both in the train and test.\n\nWe are still validating these but it seems to be roughly the numbers.\n\nOur emails are:\namiralush@gmail.com or danny.bickson@gmail.com",
    "1782679": "Hi there - thanks for sharing your findings! This is definitely something worth exploring further. I'm curious to know if you've found any duplicates in the training data as well? If so, it could be indicative of a larger data leakage issue. In any case, thanks again for bringing this to everyone's attention.",
    "1782605": "Hi @amirus! Thank you very much for identifying this problem. \n\nI checked the examples you shared with us. Those images are actually different images of the same specimens; they are treated differently in the record with different unique identifiers. There could be several reasons why duplicate images are treated as such in the herbarium perspective... This particular case seems to be related to plant samples in the packet. But still, it would be good to know the extent of overlapping images. \n\nWould you please share the code for finding the duplicates? What kind of hashing algorithm did you use? I think it would be nice for us to know exactly how many images may be overlapping between the test and the training dataset. \n\n ",
    "1783834": ""
  }
}