{
  "id": 158414,
  "title": "Train & Test Intersections",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/158414",
  "author_name": "",
  "post_date": "2020-06-14T09:00:39.150005200Z",
  "votes": 46,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Hi everyone!</p>\n\n<p>I solve task with searching duplicated images and found some samples that included in train and test sets. I would like to share with you.</p>\n\n<p>If you find more duplicates <code>train ⋂ test</code>, so lets to publish in this topic!</p>\n\n<p>Thank you!</p>\n\n<p>My samples:</p>\n\n<p>```\ntrain --&gt; test [target]</p>\n\n<p>ISIC_0014506_downsampled --&gt; ISIC_9353360 [1]\nISIC_0014518_downsampled --&gt; ISIC_9207777 [1]\nISIC_0014541_downsampled --&gt; ISIC_5224960 [1]\nISIC_0014542_downsampled --&gt; ISIC_6457527 [1]\nISIC_0030762 --&gt; ISIC_3689290 [0]\nISIC_0025740 --&gt; ISIC_3584949 [0]\nISIC_0014644_downsampled --&gt; ISIC_8347588 [0]\n```</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2F668648728455a7f757389a30d95ff265%2F2020-06-14%2000-31-29.png?generation=1592125301840175&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2Fcb7d15e2e26dbcdb8b38153b0662d082%2F2020-06-14%2000-31-51.png?generation=1592125312324240&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2F8daf7227b519b1a07e2079d753ca1d7c%2F2020-06-14%2000-32-02.png?generation=1592125321734523&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "885487",
      "postDate": "06/14/2020 09:00:39",
      "content": "<p>Hi everyone!</p>\n\n<p>I solve task with searching duplicated images and found some samples that included in train and test sets. I would like to share with you.</p>\n\n<p>If you find more duplicates <code>train ⋂ test</code>, so lets to publish in this topic!</p>\n\n<p>Thank you!</p>\n\n<p>My samples:</p>\n\n<p>```\ntrain --&gt; test [target]</p>\n\n<p>ISIC_0014506_downsampled --&gt; ISIC_9353360 [1]\nISIC_0014518_downsampled --&gt; ISIC_9207777 [1]\nISIC_0014541_downsampled --&gt; ISIC_5224960 [1]\nISIC_0014542_downsampled --&gt; ISIC_6457527 [1]\nISIC_0030762 --&gt; ISIC_3689290 [0]\nISIC_0025740 --&gt; ISIC_3584949 [0]\nISIC_0014644_downsampled --&gt; ISIC_8347588 [0]\n```</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2F668648728455a7f757389a30d95ff265%2F2020-06-14%2000-31-29.png?generation=1592125301840175&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2Fcb7d15e2e26dbcdb8b38153b0662d082%2F2020-06-14%2000-31-51.png?generation=1592125312324240&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2F8daf7227b519b1a07e2079d753ca1d7c%2F2020-06-14%2000-32-02.png?generation=1592125321734523&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi everyone!\n\nI solve task with searching duplicated images and found some samples that included in train and test sets. I would like to share with you.\n\nIf you find more duplicates `train ⋂ test`, so lets to publish in this topic!\n\nThank you!\n\nMy samples:\n\n```\ntrain --&gt; test [target]\n\nISIC_0014506_downsampled --&gt; ISIC_9353360 [1]\nISIC_0014518_downsampled --&gt; ISIC_9207777 [1]\nISIC_0014541_downsampled --&gt; ISIC_5224960 [1]\nISIC_0014542_downsampled --&gt; ISIC_6457527 [1]\nISIC_0030762 --&gt; ISIC_3689290 [0]\nISIC_0025740 --&gt; ISIC_3584949 [0]\nISIC_0014644_downsampled --&gt; ISIC_8347588 [0]\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2F668648728455a7f757389a30d95ff265%2F2020-06-14%2000-31-29.png?generation=1592125301840175&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2Fcb7d15e2e26dbcdb8b38153b0662d082%2F2020-06-14%2000-31-51.png?generation=1592125312324240&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2F8daf7227b519b1a07e2079d753ca1d7c%2F2020-06-14%2000-32-02.png?generation=1592125321734523&amp;alt=media)",
      "votes": null
    },
    {
      "id": "885545",
      "postDate": "06/14/2020 09:58:49",
      "content": "<p>Thanks <a href=\"/shonenkov\">@shonenkov</a> for sharing this.\nDid you perform the same method to find duplicates as in this kernel <a href=\"https://www.kaggle.com/shonenkov/dbscan-clustering-check-marking\">https://www.kaggle.com/shonenkov/dbscan-clustering-check-marking</a> i.e DBSCAN?\nThose 7 are the only duplicates it gave?</p>",
      "rawMarkdown": "Thanks @shonenkov for sharing this.\nDid you perform the same method to find duplicates as in this kernel https://www.kaggle.com/shonenkov/dbscan-clustering-check-marking i.e DBSCAN?\nThose 7 are the only duplicates it gave?",
      "votes": null
    },
    {
      "id": "885562",
      "postDate": "06/14/2020 10:15:18",
      "content": "<p>that explains:</p>\n\n<ol>\n<li><p>why external data improve results</p></li>\n<li><p>contrastive learning may give good results (minimizing distance of sample and its augmentation)</p></li>\n</ol>",
      "rawMarkdown": "that explains:\n\n1. why external data improve results\n\n2. contrastive learning may give good results (minimizing distance of sample and its augmentation)",
      "votes": null
    },
    {
      "id": "885604",
      "postDate": "06/14/2020 10:41:55",
      "content": "<p><a href=\"/optimo\">@optimo</a> Yes, I used DBSCAN for searching train-test. For this case it gave me ~20-60 samples ( for different values of epsilon), but only 7 is really duplicates.</p>\n\n<p>In this kernel are intersections train-train (~1130), and in the next version I will add intersections test-test.</p>\n\n<p>P.S.\nFor finding these samples you can use DBSCAN (such as my kernel) with eps [3.0, 5.0, 7.0] for merged train+test. In kernel I wouldn’t like to add this code, because it needs huge time (~2h by eps) for clustering on Kaggle kernels.</p>",
      "rawMarkdown": "optimo Yes, I used DBSCAN for searching train-test. For this case it gave me ~20-60 samples ( for different values of epsilon), but only 7 is really duplicates.\n\nIn this kernel are intersections train-train (~1130), and in the next version I will add intersections test-test.\n\nP.S.\nFor finding these samples you can use DBSCAN (such as my kernel) with eps [3.0, 5.0, 7.0] for merged train+test. In kernel I wouldn’t like to add this code, because it needs huge time (~2h by eps) for clustering on Kaggle kernels.",
      "votes": null
    },
    {
      "id": "885975",
      "postDate": "06/14/2020 16:00:42",
      "content": "<p>Is it allowed to use this kind of \"leak\" to generate predictions for the test set or is this considered \"labeling the test set\" ?</p>",
      "rawMarkdown": "Is it allowed to use this kind of \"leak\" to generate predictions for the test set or is this considered \"labeling the test set\" ?",
      "votes": null
    },
    {
      "id": "886029",
      "postDate": "06/14/2020 16:54:19",
      "content": "<p>In the past it would have been perfectly fine to use that... but given current state of how Kaggle started to interpret its own rules - I am not too sure now</p>\n\n<p>Refering to <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983</a></p>",
      "rawMarkdown": "In the past it would have been perfectly fine to use that... but given current state of how Kaggle started to interpret its own rules - I am not too sure now\n\nRefering to https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983",
      "votes": null
    },
    {
      "id": "886038",
      "postDate": "06/14/2020 17:01:00",
      "content": "<p>That was my exact thought 😄 </p>",
      "rawMarkdown": "That was my exact thought 😄",
      "votes": null
    },
    {
      "id": "886586",
      "postDate": "06/15/2020 06:28:37",
      "content": "<p>Please see <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/rules\">Rules</a>, where <strong>EXTERNAL DATA</strong> section clearly mentions <strong>No hand-labeling of the test set.</strong></p>",
      "rawMarkdown": "Please see [Rules](https://www.kaggle.com/c/siim-isic-melanoma-classification/rules), where **EXTERNAL DATA** section clearly mentions **No hand-labeling of the test set.**",
      "votes": null
    },
    {
      "id": "886694",
      "postDate": "06/15/2020 07:53:53",
      "content": "<p><a href=\"/sirishks\">@sirishks</a> it is no hand-labeling of the test set. These samples are from train set. Thank you for attention.</p>",
      "rawMarkdown": "sirishks it is no hand-labeling of the test set. These samples are from train set. Thank you for attention.",
      "votes": null
    },
    {
      "id": "886711",
      "postDate": "06/15/2020 08:19:13",
      "content": "<p>well <a href=\"/shonenkov\">@shonenkov</a> I guess it's not allowed to simply set those images to the corresponding value \"by hand\" in the test set, you'll need at least to say that your dbscan algorithm has a final say in your final predictions.</p>",
      "rawMarkdown": "well @shonenkov I guess it's not allowed to simply set those images to the corresponding value \"by hand\" in the test set, you'll need at least to say that your dbscan algorithm has a final say in your final predictions.",
      "votes": null
    },
    {
      "id": "886724",
      "postDate": "06/15/2020 08:33:21",
      "content": "<p><a href=\"/optimo\">@optimo</a> not necessary to set target values “by hand” 😂 enough to know about this leak. you can tune model or make “right” splitting train data or smth else. It is not against the rules if it is not double standard!  :troll:</p>",
      "rawMarkdown": "optimo not necessary to set target values “by hand” 😂 enough to know about this leak. you can tune model or make “right” splitting train data or smth else. It is not against the rules if it is not double standard!  :troll:",
      "votes": null
    },
    {
      "id": "889056",
      "postDate": "06/16/2020 18:35:12",
      "content": "<p>Really good work! Thanks!</p>\n\n<p>One more for the list:\n<code>ISIC_0014525_downsampled.jpg --&gt; ISIC_8372206.jpg [1]</code></p>",
      "rawMarkdown": "Really good work! Thanks!\n\nOne more for the list:\n`ISIC_0014525_downsampled.jpg --&gt; ISIC_8372206.jpg [1]`",
      "votes": null
    },
    {
      "id": "889266",
      "postDate": "06/16/2020 21:33:42",
      "content": "<p>A version of this question was also raised <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/157701#882521\">on this thread</a>. Thanks for sharing. The host is looking into this and will respond as soon as possible.</p>",
      "rawMarkdown": "A version of this question was also raised [on this thread](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/157701#882521). Thanks for sharing. The host is looking into this and will respond as soon as possible.",
      "votes": null
    },
    {
      "id": "891070",
      "postDate": "06/17/2020 21:49:04",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> thanks for the update and I look forward to the official response soon.</p>",
      "rawMarkdown": "juliaelliott thanks for the update and I look forward to the official response soon.",
      "votes": null
    },
    {
      "id": "903473",
      "postDate": "06/26/2020 21:18:54",
      "content": "<p>Alright, the host has <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\">issued this explanation</a> that addresses dupes, both that you have identified and some additional ones uncovered. The short of it is that they acknowledge these dupes exist and have confirmed they are okay to use. Obviously, manually labeling the test set with this knowledge is still prohibited, but these images are still free to include in training.</p>",
      "rawMarkdown": "Alright, the host has [issued this explanation](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943) that addresses dupes, both that you have identified and some additional ones uncovered. The short of it is that they acknowledge these dupes exist and have confirmed they are okay to use. Obviously, manually labeling the test set with this knowledge is still prohibited, but these images are still free to include in training.",
      "votes": null
    },
    {
      "id": "932218",
      "postDate": "07/16/2020 20:30:51",
      "content": "<p>Below is another test image that is contained in last years 2019 train dataset. I posted a notebook <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-using-cnn-embeddings\">here</a> showing how to use RAPIDS cuML kNN plus CNN image embeddings to find duplicates.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fcd82c82fa3f27bfa531f404a135fb76d%2FScreen%20Shot%202020-07-16%20at%201.29.16%20PM.png?generation=1594931385115776&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Below is another test image that is contained in last years 2019 train dataset. I posted a notebook [here][1] showing how to use RAPIDS cuML kNN plus CNN image embeddings to find duplicates.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fcd82c82fa3f27bfa531f404a135fb76d%2FScreen%20Shot%202020-07-16%20at%201.29.16%20PM.png?generation=1594931385115776&amp;alt=media)\n\n[1]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-using-cnn-embeddings",
      "votes": null
    },
    {
      "id": "932422",
      "postDate": "07/17/2020 03:55:30",
      "content": "<p>I think this is the image I mentioned above</p>",
      "rawMarkdown": "I think this is the image I mentioned above",
      "votes": null
    },
    {
      "id": "932426",
      "postDate": "07/17/2020 04:03:12",
      "content": "<p>Yes it is. Nice catch, you found it a month ago.</p>",
      "rawMarkdown": "Yes it is. Nice catch, you found it a month ago.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 885545,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "06/14/2020 09:58:49",
      "content": "<p>Thanks <a href=\"/shonenkov\">@shonenkov</a> for sharing this.\nDid you perform the same method to find duplicates as in this kernel <a href=\"https://www.kaggle.com/shonenkov/dbscan-clustering-check-marking\">https://www.kaggle.com/shonenkov/dbscan-clustering-check-marking</a> i.e DBSCAN?\nThose 7 are the only duplicates it gave?</p>",
      "votes": null,
      "replies": [
        {
          "id": 885604,
          "author_name": "shonenkov",
          "author_url": "",
          "post_date": "06/14/2020 10:41:55",
          "content": "<p><a href=\"/optimo\">@optimo</a> Yes, I used DBSCAN for searching train-test. For this case it gave me ~20-60 samples ( for different values of epsilon), but only 7 is really duplicates.</p>\n\n<p>In this kernel are intersections train-train (~1130), and in the next version I will add intersections test-test.</p>\n\n<p>P.S.\nFor finding these samples you can use DBSCAN (such as my kernel) with eps [3.0, 5.0, 7.0] for merged train+test. In kernel I wouldn’t like to add this code, because it needs huge time (~2h by eps) for clustering on Kaggle kernels.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 885562,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/14/2020 10:15:18",
      "content": "<p>that explains:</p>\n\n<ol>\n<li><p>why external data improve results</p></li>\n<li><p>contrastive learning may give good results (minimizing distance of sample and its augmentation)</p></li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 885975,
      "author_name": "jpbremer",
      "author_url": "",
      "post_date": "06/14/2020 16:00:42",
      "content": "<p>Is it allowed to use this kind of \"leak\" to generate predictions for the test set or is this considered \"labeling the test set\" ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 886029,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "06/14/2020 16:54:19",
          "content": "<p>In the past it would have been perfectly fine to use that... but given current state of how Kaggle started to interpret its own rules - I am not too sure now</p>\n\n<p>Refering to <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 886038,
          "author_name": "jpbremer",
          "author_url": "",
          "post_date": "06/14/2020 17:01:00",
          "content": "<p>That was my exact thought 😄 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 886586,
      "author_name": "sirishks",
      "author_url": "",
      "post_date": "06/15/2020 06:28:37",
      "content": "<p>Please see <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/rules\">Rules</a>, where <strong>EXTERNAL DATA</strong> section clearly mentions <strong>No hand-labeling of the test set.</strong></p>",
      "votes": null,
      "replies": [
        {
          "id": 886694,
          "author_name": "shonenkov",
          "author_url": "",
          "post_date": "06/15/2020 07:53:53",
          "content": "<p><a href=\"/sirishks\">@sirishks</a> it is no hand-labeling of the test set. These samples are from train set. Thank you for attention.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 886711,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "06/15/2020 08:19:13",
          "content": "<p>well <a href=\"/shonenkov\">@shonenkov</a> I guess it's not allowed to simply set those images to the corresponding value \"by hand\" in the test set, you'll need at least to say that your dbscan algorithm has a final say in your final predictions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 886724,
          "author_name": "shonenkov",
          "author_url": "",
          "post_date": "06/15/2020 08:33:21",
          "content": "<p><a href=\"/optimo\">@optimo</a> not necessary to set target values “by hand” 😂 enough to know about this leak. you can tune model or make “right” splitting train data or smth else. It is not against the rules if it is not double standard!  :troll:</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 889056,
      "author_name": "anjum48",
      "author_url": "",
      "post_date": "06/16/2020 18:35:12",
      "content": "<p>Really good work! Thanks!</p>\n\n<p>One more for the list:\n<code>ISIC_0014525_downsampled.jpg --&gt; ISIC_8372206.jpg [1]</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 889266,
      "author_name": "juliaelliott",
      "author_url": "",
      "post_date": "06/16/2020 21:33:42",
      "content": "<p>A version of this question was also raised <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/157701#882521\">on this thread</a>. Thanks for sharing. The host is looking into this and will respond as soon as possible.</p>",
      "votes": null,
      "replies": [
        {
          "id": 891070,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "06/17/2020 21:49:04",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> thanks for the update and I look forward to the official response soon.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 903473,
      "author_name": "juliaelliott",
      "author_url": "",
      "post_date": "06/26/2020 21:18:54",
      "content": "<p>Alright, the host has <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\">issued this explanation</a> that addresses dupes, both that you have identified and some additional ones uncovered. The short of it is that they acknowledge these dupes exist and have confirmed they are okay to use. Obviously, manually labeling the test set with this knowledge is still prohibited, but these images are still free to include in training.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 932218,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/16/2020 20:30:51",
      "content": "<p>Below is another test image that is contained in last years 2019 train dataset. I posted a notebook <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-using-cnn-embeddings\">here</a> showing how to use RAPIDS cuML kNN plus CNN image embeddings to find duplicates.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fcd82c82fa3f27bfa531f404a135fb76d%2FScreen%20Shot%202020-07-16%20at%201.29.16%20PM.png?generation=1594931385115776&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 932422,
          "author_name": "anjum48",
          "author_url": "",
          "post_date": "07/17/2020 03:55:30",
          "content": "<p>I think this is the image I mentioned above</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932426,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/17/2020 04:03:12",
          "content": "<p>Yes it is. Nice catch, you found it a month ago.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "885487": "Hi everyone!\n\nI solve task with searching duplicated images and found some samples that included in train and test sets. I would like to share with you.\n\nIf you find more duplicates `train ⋂ test`, so lets to publish in this topic!\n\nThank you!\n\nMy samples:\n\n```\ntrain --&gt; test [target]\n\nISIC_0014506_downsampled --&gt; ISIC_9353360 [1]\nISIC_0014518_downsampled --&gt; ISIC_9207777 [1]\nISIC_0014541_downsampled --&gt; ISIC_5224960 [1]\nISIC_0014542_downsampled --&gt; ISIC_6457527 [1]\nISIC_0030762 --&gt; ISIC_3689290 [0]\nISIC_0025740 --&gt; ISIC_3584949 [0]\nISIC_0014644_downsampled --&gt; ISIC_8347588 [0]\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2F668648728455a7f757389a30d95ff265%2F2020-06-14%2000-31-29.png?generation=1592125301840175&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2Fcb7d15e2e26dbcdb8b38153b0662d082%2F2020-06-14%2000-31-51.png?generation=1592125312324240&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2F8daf7227b519b1a07e2079d753ca1d7c%2F2020-06-14%2000-32-02.png?generation=1592125321734523&amp;alt=media)",
    "885545": "Thanks @shonenkov for sharing this.\nDid you perform the same method to find duplicates as in this kernel https://www.kaggle.com/shonenkov/dbscan-clustering-check-marking i.e DBSCAN?\nThose 7 are the only duplicates it gave?",
    "885562": "that explains:\n\n1. why external data improve results\n\n2. contrastive learning may give good results (minimizing distance of sample and its augmentation)",
    "885604": "optimo Yes, I used DBSCAN for searching train-test. For this case it gave me ~20-60 samples ( for different values of epsilon), but only 7 is really duplicates.\n\nIn this kernel are intersections train-train (~1130), and in the next version I will add intersections test-test.\n\nP.S.\nFor finding these samples you can use DBSCAN (such as my kernel) with eps [3.0, 5.0, 7.0] for merged train+test. In kernel I wouldn’t like to add this code, because it needs huge time (~2h by eps) for clustering on Kaggle kernels.",
    "885975": "Is it allowed to use this kind of \"leak\" to generate predictions for the test set or is this considered \"labeling the test set\" ?",
    "886029": "In the past it would have been perfectly fine to use that... but given current state of how Kaggle started to interpret its own rules - I am not too sure now\n\nRefering to https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983",
    "886038": "That was my exact thought 😄",
    "886586": "Please see [Rules](https://www.kaggle.com/c/siim-isic-melanoma-classification/rules), where **EXTERNAL DATA** section clearly mentions **No hand-labeling of the test set.**",
    "886694": "sirishks it is no hand-labeling of the test set. These samples are from train set. Thank you for attention.",
    "886711": "well @shonenkov I guess it's not allowed to simply set those images to the corresponding value \"by hand\" in the test set, you'll need at least to say that your dbscan algorithm has a final say in your final predictions.",
    "886724": "optimo not necessary to set target values “by hand” 😂 enough to know about this leak. you can tune model or make “right” splitting train data or smth else. It is not against the rules if it is not double standard!  :troll:",
    "889056": "Really good work! Thanks!\n\nOne more for the list:\n`ISIC_0014525_downsampled.jpg --&gt; ISIC_8372206.jpg [1]`",
    "889266": "A version of this question was also raised [on this thread](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/157701#882521). Thanks for sharing. The host is looking into this and will respond as soon as possible.",
    "891070": "juliaelliott thanks for the update and I look forward to the official response soon.",
    "903473": "Alright, the host has [issued this explanation](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943) that addresses dupes, both that you have identified and some additional ones uncovered. The short of it is that they acknowledge these dupes exist and have confirmed they are okay to use. Obviously, manually labeling the test set with this knowledge is still prohibited, but these images are still free to include in training.",
    "932218": "Below is another test image that is contained in last years 2019 train dataset. I posted a notebook [here][1] showing how to use RAPIDS cuML kNN plus CNN image embeddings to find duplicates.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fcd82c82fa3f27bfa531f404a135fb76d%2FScreen%20Shot%202020-07-16%20at%201.29.16%20PM.png?generation=1594931385115776&amp;alt=media)\n\n[1]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-using-cnn-embeddings",
    "932422": "I think this is the image I mentioned above",
    "932426": "Yes it is. Nice catch, you found it a month ago."
  },
  "source": "meta"
}