{
  "id": 161943,
  "title": "True duplicates in this dataset",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/161943",
  "author_name": "",
  "post_date": "2020-06-26T19:12:01.136710200Z",
  "votes": 65,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Dear SIIM/ISIC 2020 challenge participants,</p>\n\n<p>We're very grateful for the awesome feedback we have received so far, and especially to the users who have prodded us with information about duplicate images.</p>\n\n<p>Indeed, we discovered that based on an issue that happened during data ingestion into our archive, several hundred images were indeed stored twice under two internal names (ISIC_ID).</p>\n\n<p>Attached is a list of images that are these exact duplicates, with both ISIC_IDs and the partition in which they occur in--in all cases both copies are in the same partition! In other words, no leakage occurred in this manner that we are aware of.</p>\n\n<p>Please note that several threads by users have also demonstrated duplicates and near duplicates between the 2020 challenge dataset and images prior uploaded to the ISIC Archive. This was, to some extent, inevitable, since the archive doesn't yet support complex image duplicate detection, which is something we will be working on before next year's challenge. We are aware of some of these additional (near-) duplicates, but believe them to be few in number.</p>\n\n<p>Since the split of exact duplicates across train and test roughly matches the overall dataset split, and the total number of possible leaks is small, we believe that no substantial bias in favor of specific outcomes occurs, and we have agreed with the Kaggle team that the scoring will not be altered--which we assume will also make participation much easier.</p>\n\n<p>We would like to apologize that this mistake happened, and we will put in place at least two mechanisms and steps such that future datasets will definitely not have the same problem.</p>\n\n<p>And, again, thank you very much for such fantastic submissions to date, and also to the users who have provided some initial code for the detection of duplicates!</p>",
  "messages": [
    {
      "id": "903372",
      "postDate": "06/26/2020 19:12:01",
      "content": "<p>Dear SIIM/ISIC 2020 challenge participants,</p>\n\n<p>We're very grateful for the awesome feedback we have received so far, and especially to the users who have prodded us with information about duplicate images.</p>\n\n<p>Indeed, we discovered that based on an issue that happened during data ingestion into our archive, several hundred images were indeed stored twice under two internal names (ISIC_ID).</p>\n\n<p>Attached is a list of images that are these exact duplicates, with both ISIC_IDs and the partition in which they occur in--in all cases both copies are in the same partition! In other words, no leakage occurred in this manner that we are aware of.</p>\n\n<p>Please note that several threads by users have also demonstrated duplicates and near duplicates between the 2020 challenge dataset and images prior uploaded to the ISIC Archive. This was, to some extent, inevitable, since the archive doesn't yet support complex image duplicate detection, which is something we will be working on before next year's challenge. We are aware of some of these additional (near-) duplicates, but believe them to be few in number.</p>\n\n<p>Since the split of exact duplicates across train and test roughly matches the overall dataset split, and the total number of possible leaks is small, we believe that no substantial bias in favor of specific outcomes occurs, and we have agreed with the Kaggle team that the scoring will not be altered--which we assume will also make participation much easier.</p>\n\n<p>We would like to apologize that this mistake happened, and we will put in place at least two mechanisms and steps such that future datasets will definitely not have the same problem.</p>\n\n<p>And, again, thank you very much for such fantastic submissions to date, and also to the users who have provided some initial code for the detection of duplicates!</p>",
      "rawMarkdown": "Dear SIIM/ISIC 2020 challenge participants,\n\nWe're very grateful for the awesome feedback we have received so far, and especially to the users who have prodded us with information about duplicate images.\n\nIndeed, we discovered that based on an issue that happened during data ingestion into our archive, several hundred images were indeed stored twice under two internal names (ISIC_ID).\n\nAttached is a list of images that are these exact duplicates, with both ISIC_IDs and the partition in which they occur in--in all cases both copies are in the same partition! In other words, no leakage occurred in this manner that we are aware of.\n\nPlease note that several threads by users have also demonstrated duplicates and near duplicates between the 2020 challenge dataset and images prior uploaded to the ISIC Archive. This was, to some extent, inevitable, since the archive doesn't yet support complex image duplicate detection, which is something we will be working on before next year's challenge. We are aware of some of these additional (near-) duplicates, but believe them to be few in number.\n\nSince the split of exact duplicates across train and test roughly matches the overall dataset split, and the total number of possible leaks is small, we believe that no substantial bias in favor of specific outcomes occurs, and we have agreed with the Kaggle team that the scoring will not be altered--which we assume will also make participation much easier.\n\nWe would like to apologize that this mistake happened, and we will put in place at least two mechanisms and steps such that future datasets will definitely not have the same problem.\n\nAnd, again, thank you very much for such fantastic submissions to date, and also to the users who have provided some initial code for the detection of duplicates!",
      "votes": null
    },
    {
      "id": "903379",
      "postDate": "06/26/2020 19:26:25",
      "content": "<p>Thank you for this very helpful information! I am very glad that you provided the list of the duplicates -- now we can focus on building our models instead of hunting down all duplicate images. And it gives me a piece of mind that the test set seems to be leakage free. </p>",
      "rawMarkdown": "Thank you for this very helpful information! I am very glad that you provided the list of the duplicates -- now we can focus on building our models instead of hunting down all duplicate images. And it gives me a piece of mind that the test set seems to be leakage free.",
      "votes": null
    },
    {
      "id": "903422",
      "postDate": "06/26/2020 20:22:08",
      "content": "<p>Thanks for the duplicates list.</p>",
      "rawMarkdown": "Thanks for the duplicates list.",
      "votes": null
    },
    {
      "id": "903432",
      "postDate": "06/26/2020 20:27:42",
      "content": "<p>I have put the list of duplicates in a public data set for  easy access. Here is the link: <a href=\"https://www.kaggle.com/graf10a/siim-list-of-duplicates\">SIIM List of Duplicates</a>.</p>\n\n<p>UPDATE: And here is a simple kernel checking if the tabular data are the same for the original and the duplicate images:\n <a href=\"https://www.kaggle.com/graf10a/siim-tabular-data-for-duplicates\">SIIM Tabular Data for Duplicates</a>. \nI did find some entries that differ between these two sets of images. But the number is fairly small, so it should not cause us any trouble.</p>",
      "rawMarkdown": "I have put the list of duplicates in a public data set for  easy access. Here is the link: [SIIM List of Duplicates](https://www.kaggle.com/graf10a/siim-list-of-duplicates).\n\nUPDATE: And here is a simple kernel checking if the tabular data are the same for the original and the duplicate images:\n [SIIM Tabular Data for Duplicates](https://www.kaggle.com/graf10a/siim-tabular-data-for-duplicates). \nI did find some entries that differ between these two sets of images. But the number is fairly small, so it should not cause us any trouble.",
      "votes": null
    },
    {
      "id": "903516",
      "postDate": "06/26/2020 22:45:46",
      "content": "<p><a href=\"/jwebermsk\">@jwebermsk</a> Thank you for checking these duplicates!</p>",
      "rawMarkdown": "jwebermsk Thank you for checking these duplicates!",
      "votes": null
    },
    {
      "id": "903740",
      "postDate": "06/27/2020 04:35:46",
      "content": "<p><a href=\"/jwebermsk\">@jwebermsk</a> Thank you for listing duplicates in Train, duplicates in Test. However, one quick question - did you also check for duplicates between Train set and Test set? (because I don't see these in your attachment)</p>\n\n<p>@All Please don't forget to:\n1. Remove duplicates from <a href=\"https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery\">External data</a> you are using as well \n    i.e. Duplicates between Train-Malignant and External-Malignant, Train-Benign and External-Benign, Test-All vs External-All\n2. At the time of making predictions, validate whether you are predicting the same label/score for duplicate items! (and this is obvious, but <strong>don't throw away</strong> the Test set duplicates - you still need to submit predictions on them 😃)\n3. Use External data only in training folds, but not your validation folds</p>",
      "rawMarkdown": "jwebermsk Thank you for listing duplicates in Train, duplicates in Test. However, one quick question - did you also check for duplicates between Train set and Test set? (because I don't see these in your attachment)\n\n@All Please don't forget to:\n1. Remove duplicates from [External data](https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery) you are using as well \n    i.e. Duplicates between Train-Malignant and External-Malignant, Train-Benign and External-Benign, Test-All vs External-All\n2. At the time of making predictions, validate whether you are predicting the same label/score for duplicate items! (and this is obvious, but **don't throw away** the Test set duplicates - you still need to submit predictions on them 😃)\n3. Use External data only in training folds, but not your validation folds",
      "votes": null
    },
    {
      "id": "903904",
      "postDate": "06/27/2020 07:46:58",
      "content": "<p>A few finding about the duplicates:\n1. In the training data every duplicated couple belong to the same patient, i.e. if if <code>image_name(x)=image_name(y)</code> it means <code>patient_id(x)=patient_id(y)</code>. which means, there is no new data about the patient.</p>\n\n<ol>\n<li><p>For the test set it is a bit different, there are 2 couples of patients with duplicate images, i.e. they are the same patient: <code>(IP_4181913, IP_2799140) and (IP_8179715, IP_5758416)</code>. </p></li>\n<li><p><a href=\"/jwebermsk\">@jwebermsk</a> <a href=\"/juliaelliott\">@juliaelliott</a>  As for the potential exploit of this leak, if every patient is ether on the private LB or the public LB , then except for these 2 patient, the is no real leak. <strong>BUT</strong> if this is not the case, one can do <strong>LB probing</strong> using the duplicate list to discover which images have one duplicate on the public LB and the other on the private, and what is the target value for these images. Given the relatively small number of positive targets this may \nhave an impact on the final ranking. If this is the case, I would suggest to move these images out of the private LB.</p></li>\n</ol>",
      "rawMarkdown": "A few finding about the duplicates:\n1. In the training data every duplicated couple belong to the same patient, i.e. if if `image_name(x)=image_name(y)` it means `patient_id(x)=patient_id(y)`. which means, there is no new data about the patient.\n\n2. For the test set it is a bit different, there are 2 couples of patients with duplicate images, i.e. they are the same patient: `(IP_4181913, IP_2799140) and (IP_8179715, IP_5758416)`. \n\n3. @jwebermsk @juliaelliott  As for the potential exploit of this leak, if every patient is ether on the private LB or the public LB , then except for these 2 patient, the is no real leak. **BUT** if this is not the case, one can do **LB probing** using the duplicate list to discover which images have one duplicate on the public LB and the other on the private, and what is the target value for these images. Given the relatively small number of positive targets this may \nhave an impact on the final ranking. If this is the case, I would suggest to move these images out of the private LB.",
      "votes": null
    },
    {
      "id": "904255",
      "postDate": "06/27/2020 13:29:02",
      "content": "<p>Hello Sirish,</p>\n\n<p>Thanks for asking this to clarify how the detection was performed! :)</p>\n\n<p>The test for (exact) duplicates was run over the entire (train + test) dataset, and since we assigned images to partition by patient ID, only the two cases in <a href=\"/yuval6967\">@yuval6967</a>'s post (see above) could have led to a split between partitions.</p>\n\n<p>Luckily this did not happen (whew).</p>",
      "rawMarkdown": "Hello Sirish,\n\nThanks for asking this to clarify how the detection was performed! :)\n\nThe test for (exact) duplicates was run over the entire (train + test) dataset, and since we assigned images to partition by patient ID, only the two cases in @yuval6967's post (see above) could have led to a split between partitions.\n\nLuckily this did not happen (whew).",
      "votes": null
    },
    {
      "id": "905153",
      "postDate": "06/28/2020 10:03:28",
      "content": "<p>Thanks a lot, I filtered my train dataset to avoid problems with CV.</p>\n\n<p>Question - are there also duplicates between 2020 train and 2019 train from dataset mentioned in External Data Thread? <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154296\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154296</a></p>",
      "rawMarkdown": "Thanks a lot, I filtered my train dataset to avoid problems with CV.\n\nQuestion - are there also duplicates between 2020 train and 2019 train from dataset mentioned in External Data Thread? https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154296",
      "votes": null
    },
    {
      "id": "907425",
      "postDate": "06/30/2020 00:41:19",
      "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> Are you referring to <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158414\">these</a>, also <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/157701\">discussed here</a>? If so, yes, Jochen has acknowledged these as known in his note above, and they are also eligible to be used, as long as the test set is not manually hand-labeled: </p>\n\n<blockquote>\n  <p>Please note that several threads by users have also demonstrated duplicates and near duplicates between the 2020 challenge dataset and images prior uploaded to the ISIC Archive.</p>\n</blockquote>",
      "rawMarkdown": "jacekpoplawski Are you referring to [these](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158414), also [discussed here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/157701)? If so, yes, Jochen has acknowledged these as known in his note above, and they are also eligible to be used, as long as the test set is not manually hand-labeled: \n&gt; Please note that several threads by users have also demonstrated duplicates and near duplicates between the 2020 challenge dataset and images prior uploaded to the ISIC Archive.",
      "votes": null
    },
    {
      "id": "907551",
      "postDate": "06/30/2020 03:05:10",
      "content": "<p>Thank you Julia, I found the list and removed them from my 2019 dataset.</p>",
      "rawMarkdown": "Thank you Julia, I found the list and removed them from my 2019 dataset.",
      "votes": null
    },
    {
      "id": "922021",
      "postDate": "07/09/2020 18:34:50",
      "content": "<p>I have also compared 2019 comp data with 2020 comp data. And found 59 duplicates. (A few in my list are just very similar but most are exact duplicates). My list is <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910#920904\">here</a>. </p>\n\n<p>Note that the 2018 comp data is entirely contained within the 2019 comp data. And the 2017 comp malignant images are entirely contained within the 2019 comp data. (I have not compared the 2017 benign images but they are most likely contained too but this is less important to know).</p>",
      "rawMarkdown": "I have also compared 2019 comp data with 2020 comp data. And found 59 duplicates. (A few in my list are just very similar but most are exact duplicates). My list is [here][1]. \n\nNote that the 2018 comp data is entirely contained within the 2019 comp data. And the 2017 comp malignant images are entirely contained within the 2019 comp data. (I have not compared the 2017 benign images but they are most likely contained too but this is less important to know).\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910#920904",
      "votes": null
    },
    {
      "id": "922187",
      "postDate": "07/09/2020 21:42:20",
      "content": "<p>Thanks Jochen. I doubled checked the list. The list is missing 9 more pairs in the train set</p>\n\n<pre><code>01 ,ISIC_1642492, T=0, =&amp;gt; ,ISIC_8776686,\n02 ,ISIC_2074396, T=0, =&amp;gt; ,ISIC_8689583,\n03 ,ISIC_2718135, T=0, =&amp;gt; ,ISIC_9409255,\n04 ,ISIC_2887033, T=0, =&amp;gt; ,ISIC_5194874, (nearly exact)\n05 ,ISIC_5492174, T=0, =&amp;gt; ,ISIC_6705662,\n06 ,ISIC_6063252, T=0, =&amp;gt; ,ISIC_7195645,\n07 ,ISIC_6284722, T=0, =&amp;gt; ,ISIC_8262759,\n08 ,ISIC_6450285, T=0, =&amp;gt; ,ISIC_6548307,\n09 ,ISIC_7607101, T=0, =&amp;gt; ,ISIC_7675261,\n</code></pre>\n\n<p>And 2 more pairs in test</p>\n\n<pre><code>01 ,ISIC_2391447, =&amp;gt; ,ISIC_4718575,\n02 ,ISIC_3091968, =&amp;gt; ,ISIC_5914041, (nearly exact)\n</code></pre>\n\n<p>I also confirm that no test images have an exact pair in this year's competition train data </p>",
      "rawMarkdown": "Thanks Jochen. I doubled checked the list. The list is missing 9 more pairs in the train set\n\n    01 ,ISIC_1642492, T=0, =&gt; ,ISIC_8776686,\n    02 ,ISIC_2074396, T=0, =&gt; ,ISIC_8689583,\n    03 ,ISIC_2718135, T=0, =&gt; ,ISIC_9409255,\n    04 ,ISIC_2887033, T=0, =&gt; ,ISIC_5194874, (nearly exact)\n    05 ,ISIC_5492174, T=0, =&gt; ,ISIC_6705662,\n    06 ,ISIC_6063252, T=0, =&gt; ,ISIC_7195645,\n    07 ,ISIC_6284722, T=0, =&gt; ,ISIC_8262759,\n    08 ,ISIC_6450285, T=0, =&gt; ,ISIC_6548307,\n    09 ,ISIC_7607101, T=0, =&gt; ,ISIC_7675261,\n\nAnd 2 more pairs in test\n\n    01 ,ISIC_2391447, =&gt; ,ISIC_4718575,\n    02 ,ISIC_3091968, =&gt; ,ISIC_5914041, (nearly exact)\n\nI also confirm that no test images have an exact pair in this year's competition train data",
      "votes": null
    },
    {
      "id": "922194",
      "postDate": "07/09/2020 22:02:18",
      "content": "<p>Hey Chris,</p>\n\n<p>Thanks so much for looking into this as well! :)</p>\n\n<p>I just ran these through the PHash algorithm, and can confirm the 9 additional matches (identical hashes); for some reason however, the files that I was running the original search over weren't detected as duplicates (I was using md5sum)... I will spend some time tomorrow trying to figure out what happened.</p>\n\n<p>Cheers,\n/jochen</p>",
      "rawMarkdown": "Hey Chris,\n\nThanks so much for looking into this as well! :)\n\nI just ran these through the PHash algorithm, and can confirm the 9 additional matches (identical hashes); for some reason however, the files that I was running the original search over weren't detected as duplicates (I was using md5sum)... I will spend some time tomorrow trying to figure out what happened.\n\nCheers,\n/jochen",
      "votes": null
    },
    {
      "id": "922224",
      "postDate": "07/09/2020 23:29:45",
      "content": "<p>thanks a lot Chris! I will use your list</p>",
      "rawMarkdown": "thanks a lot Chris! I will use your list",
      "votes": null
    },
    {
      "id": "922418",
      "postDate": "07/10/2020 05:12:58",
      "content": "<p>Ok, i have also updated my popular training TFRecords and removed the 434 duplicate training images (discussion <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\">here</a>). This will help participants create a more reliable CV. And it will get word out about the duplicate images.</p>",
      "rawMarkdown": "Ok, i have also updated my popular training TFRecords and removed the 434 duplicate training images (discussion [here][1]). This will help participants create a more reliable CV. And it will get word out about the duplicate images.\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526",
      "votes": null
    },
    {
      "id": "963294",
      "postDate": "08/08/2020 21:25:10",
      "content": "<p>Thanks a lot !!</p>",
      "rawMarkdown": "Thanks a lot !!",
      "votes": null
    },
    {
      "id": "964064",
      "postDate": "08/09/2020 15:03:34",
      "content": "<p><a href=\"https://www.kaggle.com/graf10a\" target=\"_blank\">@graf10a</a> it seems this might be interesting for you: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930</a></p>",
      "rawMarkdown": "graf10a it seems this might be interesting for you: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930",
      "votes": null
    },
    {
      "id": "964657",
      "postDate": "08/10/2020 04:42:35",
      "content": "<p>For those interested in this topic have a look at the <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> analysis in here <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930\" target=\"_blank\">here</a>. It seems even knowing all the 107 images that are potential leak from the test data exploiting it doesnt give much advantage. </p>\n<p>Big thanks to <a href=\"https://www.kaggle.com/ucanmakeit175\" target=\"_blank\">@ucanmakeit175</a>  <a href=\"https://www.kaggle.com/benboren\" target=\"_blank\">@benboren</a>  <a href=\"https://www.kaggle.com/coreacasa\" target=\"_blank\">@coreacasa</a> <a href=\"https://www.kaggle.com/ajtryt2\" target=\"_blank\">@ajtryt2</a> and <a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a></p>",
      "rawMarkdown": "For those interested in this topic have a look at the @cdeotte analysis in here [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930). It seems even knowing all the 107 images that are potential leak from the test data exploiting it doesnt give much advantage. \n\nBig thanks to @ucanmakeit175  @benboren  @coreacasa @ajtryt2 and @sirishks",
      "votes": null
    },
    {
      "id": "970341",
      "postDate": "08/14/2020 11:23:19",
      "content": "<p>Thanks for you sharing!</p>",
      "rawMarkdown": "Thanks for you sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 964657,
      "author_name": "janidziak",
      "author_url": "",
      "post_date": "08/10/2020 04:42:35",
      "content": "<p>For those interested in this topic have a look at the <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> analysis in here <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930\" target=\"_blank\">here</a>. It seems even knowing all the 107 images that are potential leak from the test data exploiting it doesnt give much advantage. </p>\n<p>Big thanks to <a href=\"https://www.kaggle.com/ucanmakeit175\" target=\"_blank\">@ucanmakeit175</a>  <a href=\"https://www.kaggle.com/benboren\" target=\"_blank\">@benboren</a>  <a href=\"https://www.kaggle.com/coreacasa\" target=\"_blank\">@coreacasa</a> <a href=\"https://www.kaggle.com/ajtryt2\" target=\"_blank\">@ajtryt2</a> and <a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 970341,
      "author_name": "junwangai",
      "author_url": "",
      "post_date": "08/14/2020 11:23:19",
      "content": "<p>Thanks for you sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 903379,
      "author_name": "graf10a",
      "author_url": "",
      "post_date": "06/26/2020 19:26:25",
      "content": "<p>Thank you for this very helpful information! I am very glad that you provided the list of the duplicates -- now we can focus on building our models instead of hunting down all duplicate images. And it gives me a piece of mind that the test set seems to be leakage free. </p>",
      "votes": null,
      "replies": [
        {
          "id": 964064,
          "author_name": "janidziak",
          "author_url": "",
          "post_date": "08/09/2020 15:03:34",
          "content": "<p><a href=\"https://www.kaggle.com/graf10a\" target=\"_blank\">@graf10a</a> it seems this might be interesting for you: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 903422,
      "author_name": "yash612",
      "author_url": "",
      "post_date": "06/26/2020 20:22:08",
      "content": "<p>Thanks for the duplicates list.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 903432,
      "author_name": "graf10a",
      "author_url": "",
      "post_date": "06/26/2020 20:27:42",
      "content": "<p>I have put the list of duplicates in a public data set for  easy access. Here is the link: <a href=\"https://www.kaggle.com/graf10a/siim-list-of-duplicates\">SIIM List of Duplicates</a>.</p>\n\n<p>UPDATE: And here is a simple kernel checking if the tabular data are the same for the original and the duplicate images:\n <a href=\"https://www.kaggle.com/graf10a/siim-tabular-data-for-duplicates\">SIIM Tabular Data for Duplicates</a>. \nI did find some entries that differ between these two sets of images. But the number is fairly small, so it should not cause us any trouble.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 903516,
      "author_name": "ebouteillon",
      "author_url": "",
      "post_date": "06/26/2020 22:45:46",
      "content": "<p><a href=\"/jwebermsk\">@jwebermsk</a> Thank you for checking these duplicates!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 903740,
      "author_name": "sirishks",
      "author_url": "",
      "post_date": "06/27/2020 04:35:46",
      "content": "<p><a href=\"/jwebermsk\">@jwebermsk</a> Thank you for listing duplicates in Train, duplicates in Test. However, one quick question - did you also check for duplicates between Train set and Test set? (because I don't see these in your attachment)</p>\n\n<p>@All Please don't forget to:\n1. Remove duplicates from <a href=\"https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery\">External data</a> you are using as well \n    i.e. Duplicates between Train-Malignant and External-Malignant, Train-Benign and External-Benign, Test-All vs External-All\n2. At the time of making predictions, validate whether you are predicting the same label/score for duplicate items! (and this is obvious, but <strong>don't throw away</strong> the Test set duplicates - you still need to submit predictions on them 😃)\n3. Use External data only in training folds, but not your validation folds</p>",
      "votes": null,
      "replies": [
        {
          "id": 904255,
          "author_name": "jwebermsk",
          "author_url": "",
          "post_date": "06/27/2020 13:29:02",
          "content": "<p>Hello Sirish,</p>\n\n<p>Thanks for asking this to clarify how the detection was performed! :)</p>\n\n<p>The test for (exact) duplicates was run over the entire (train + test) dataset, and since we assigned images to partition by patient ID, only the two cases in <a href=\"/yuval6967\">@yuval6967</a>'s post (see above) could have led to a split between partitions.</p>\n\n<p>Luckily this did not happen (whew).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 903904,
      "author_name": "yuval6967",
      "author_url": "",
      "post_date": "06/27/2020 07:46:58",
      "content": "<p>A few finding about the duplicates:\n1. In the training data every duplicated couple belong to the same patient, i.e. if if <code>image_name(x)=image_name(y)</code> it means <code>patient_id(x)=patient_id(y)</code>. which means, there is no new data about the patient.</p>\n\n<ol>\n<li><p>For the test set it is a bit different, there are 2 couples of patients with duplicate images, i.e. they are the same patient: <code>(IP_4181913, IP_2799140) and (IP_8179715, IP_5758416)</code>. </p></li>\n<li><p><a href=\"/jwebermsk\">@jwebermsk</a> <a href=\"/juliaelliott\">@juliaelliott</a>  As for the potential exploit of this leak, if every patient is ether on the private LB or the public LB , then except for these 2 patient, the is no real leak. <strong>BUT</strong> if this is not the case, one can do <strong>LB probing</strong> using the duplicate list to discover which images have one duplicate on the public LB and the other on the private, and what is the target value for these images. Given the relatively small number of positive targets this may \nhave an impact on the final ranking. If this is the case, I would suggest to move these images out of the private LB.</p></li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 905153,
      "author_name": "jacekpoplawski",
      "author_url": "",
      "post_date": "06/28/2020 10:03:28",
      "content": "<p>Thanks a lot, I filtered my train dataset to avoid problems with CV.</p>\n\n<p>Question - are there also duplicates between 2020 train and 2019 train from dataset mentioned in External Data Thread? <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154296\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154296</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 907425,
          "author_name": "juliaelliott",
          "author_url": "",
          "post_date": "06/30/2020 00:41:19",
          "content": "<p><a href=\"/jacekpoplawski\">@jacekpoplawski</a> Are you referring to <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158414\">these</a>, also <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/157701\">discussed here</a>? If so, yes, Jochen has acknowledged these as known in his note above, and they are also eligible to be used, as long as the test set is not manually hand-labeled: </p>\n\n<blockquote>\n  <p>Please note that several threads by users have also demonstrated duplicates and near duplicates between the 2020 challenge dataset and images prior uploaded to the ISIC Archive.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 907551,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "06/30/2020 03:05:10",
          "content": "<p>Thank you Julia, I found the list and removed them from my 2019 dataset.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 922021,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/09/2020 18:34:50",
          "content": "<p>I have also compared 2019 comp data with 2020 comp data. And found 59 duplicates. (A few in my list are just very similar but most are exact duplicates). My list is <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910#920904\">here</a>. </p>\n\n<p>Note that the 2018 comp data is entirely contained within the 2019 comp data. And the 2017 comp malignant images are entirely contained within the 2019 comp data. (I have not compared the 2017 benign images but they are most likely contained too but this is less important to know).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 922224,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/09/2020 23:29:45",
          "content": "<p>thanks a lot Chris! I will use your list</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 922187,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/09/2020 21:42:20",
      "content": "<p>Thanks Jochen. I doubled checked the list. The list is missing 9 more pairs in the train set</p>\n\n<pre><code>01 ,ISIC_1642492, T=0, =&amp;gt; ,ISIC_8776686,\n02 ,ISIC_2074396, T=0, =&amp;gt; ,ISIC_8689583,\n03 ,ISIC_2718135, T=0, =&amp;gt; ,ISIC_9409255,\n04 ,ISIC_2887033, T=0, =&amp;gt; ,ISIC_5194874, (nearly exact)\n05 ,ISIC_5492174, T=0, =&amp;gt; ,ISIC_6705662,\n06 ,ISIC_6063252, T=0, =&amp;gt; ,ISIC_7195645,\n07 ,ISIC_6284722, T=0, =&amp;gt; ,ISIC_8262759,\n08 ,ISIC_6450285, T=0, =&amp;gt; ,ISIC_6548307,\n09 ,ISIC_7607101, T=0, =&amp;gt; ,ISIC_7675261,\n</code></pre>\n\n<p>And 2 more pairs in test</p>\n\n<pre><code>01 ,ISIC_2391447, =&amp;gt; ,ISIC_4718575,\n02 ,ISIC_3091968, =&amp;gt; ,ISIC_5914041, (nearly exact)\n</code></pre>\n\n<p>I also confirm that no test images have an exact pair in this year's competition train data </p>",
      "votes": null,
      "replies": [
        {
          "id": 922194,
          "author_name": "jwebermsk",
          "author_url": "",
          "post_date": "07/09/2020 22:02:18",
          "content": "<p>Hey Chris,</p>\n\n<p>Thanks so much for looking into this as well! :)</p>\n\n<p>I just ran these through the PHash algorithm, and can confirm the 9 additional matches (identical hashes); for some reason however, the files that I was running the original search over weren't detected as duplicates (I was using md5sum)... I will spend some time tomorrow trying to figure out what happened.</p>\n\n<p>Cheers,\n/jochen</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 922418,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/10/2020 05:12:58",
          "content": "<p>Ok, i have also updated my popular training TFRecords and removed the 434 duplicate training images (discussion <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\">here</a>). This will help participants create a more reliable CV. And it will get word out about the duplicate images.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 963294,
      "author_name": "ismaelkaissy",
      "author_url": "",
      "post_date": "08/08/2020 21:25:10",
      "content": "<p>Thanks a lot !!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "903372": "Dear SIIM/ISIC 2020 challenge participants,\n\nWe're very grateful for the awesome feedback we have received so far, and especially to the users who have prodded us with information about duplicate images.\n\nIndeed, we discovered that based on an issue that happened during data ingestion into our archive, several hundred images were indeed stored twice under two internal names (ISIC_ID).\n\nAttached is a list of images that are these exact duplicates, with both ISIC_IDs and the partition in which they occur in--in all cases both copies are in the same partition! In other words, no leakage occurred in this manner that we are aware of.\n\nPlease note that several threads by users have also demonstrated duplicates and near duplicates between the 2020 challenge dataset and images prior uploaded to the ISIC Archive. This was, to some extent, inevitable, since the archive doesn't yet support complex image duplicate detection, which is something we will be working on before next year's challenge. We are aware of some of these additional (near-) duplicates, but believe them to be few in number.\n\nSince the split of exact duplicates across train and test roughly matches the overall dataset split, and the total number of possible leaks is small, we believe that no substantial bias in favor of specific outcomes occurs, and we have agreed with the Kaggle team that the scoring will not be altered--which we assume will also make participation much easier.\n\nWe would like to apologize that this mistake happened, and we will put in place at least two mechanisms and steps such that future datasets will definitely not have the same problem.\n\nAnd, again, thank you very much for such fantastic submissions to date, and also to the users who have provided some initial code for the detection of duplicates!",
    "903379": "Thank you for this very helpful information! I am very glad that you provided the list of the duplicates -- now we can focus on building our models instead of hunting down all duplicate images. And it gives me a piece of mind that the test set seems to be leakage free.",
    "903422": "Thanks for the duplicates list.",
    "903432": "I have put the list of duplicates in a public data set for  easy access. Here is the link: [SIIM List of Duplicates](https://www.kaggle.com/graf10a/siim-list-of-duplicates).\n\nUPDATE: And here is a simple kernel checking if the tabular data are the same for the original and the duplicate images:\n [SIIM Tabular Data for Duplicates](https://www.kaggle.com/graf10a/siim-tabular-data-for-duplicates). \nI did find some entries that differ between these two sets of images. But the number is fairly small, so it should not cause us any trouble.",
    "903516": "jwebermsk Thank you for checking these duplicates!",
    "903740": "jwebermsk Thank you for listing duplicates in Train, duplicates in Test. However, one quick question - did you also check for duplicates between Train set and Test set? (because I don't see these in your attachment)\n\n@All Please don't forget to:\n1. Remove duplicates from [External data](https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery) you are using as well \n    i.e. Duplicates between Train-Malignant and External-Malignant, Train-Benign and External-Benign, Test-All vs External-All\n2. At the time of making predictions, validate whether you are predicting the same label/score for duplicate items! (and this is obvious, but **don't throw away** the Test set duplicates - you still need to submit predictions on them 😃)\n3. Use External data only in training folds, but not your validation folds",
    "903904": "A few finding about the duplicates:\n1. In the training data every duplicated couple belong to the same patient, i.e. if if `image_name(x)=image_name(y)` it means `patient_id(x)=patient_id(y)`. which means, there is no new data about the patient.\n\n2. For the test set it is a bit different, there are 2 couples of patients with duplicate images, i.e. they are the same patient: `(IP_4181913, IP_2799140) and (IP_8179715, IP_5758416)`. \n\n3. @jwebermsk @juliaelliott  As for the potential exploit of this leak, if every patient is ether on the private LB or the public LB , then except for these 2 patient, the is no real leak. **BUT** if this is not the case, one can do **LB probing** using the duplicate list to discover which images have one duplicate on the public LB and the other on the private, and what is the target value for these images. Given the relatively small number of positive targets this may \nhave an impact on the final ranking. If this is the case, I would suggest to move these images out of the private LB.",
    "904255": "Hello Sirish,\n\nThanks for asking this to clarify how the detection was performed! :)\n\nThe test for (exact) duplicates was run over the entire (train + test) dataset, and since we assigned images to partition by patient ID, only the two cases in @yuval6967's post (see above) could have led to a split between partitions.\n\nLuckily this did not happen (whew).",
    "905153": "Thanks a lot, I filtered my train dataset to avoid problems with CV.\n\nQuestion - are there also duplicates between 2020 train and 2019 train from dataset mentioned in External Data Thread? https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154296",
    "907425": "jacekpoplawski Are you referring to [these](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158414), also [discussed here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/157701)? If so, yes, Jochen has acknowledged these as known in his note above, and they are also eligible to be used, as long as the test set is not manually hand-labeled: \n&gt; Please note that several threads by users have also demonstrated duplicates and near duplicates between the 2020 challenge dataset and images prior uploaded to the ISIC Archive.",
    "907551": "Thank you Julia, I found the list and removed them from my 2019 dataset.",
    "922021": "I have also compared 2019 comp data with 2020 comp data. And found 59 duplicates. (A few in my list are just very similar but most are exact duplicates). My list is [here][1]. \n\nNote that the 2018 comp data is entirely contained within the 2019 comp data. And the 2017 comp malignant images are entirely contained within the 2019 comp data. (I have not compared the 2017 benign images but they are most likely contained too but this is less important to know).\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910#920904",
    "922187": "Thanks Jochen. I doubled checked the list. The list is missing 9 more pairs in the train set\n\n    01 ,ISIC_1642492, T=0, =&gt; ,ISIC_8776686,\n    02 ,ISIC_2074396, T=0, =&gt; ,ISIC_8689583,\n    03 ,ISIC_2718135, T=0, =&gt; ,ISIC_9409255,\n    04 ,ISIC_2887033, T=0, =&gt; ,ISIC_5194874, (nearly exact)\n    05 ,ISIC_5492174, T=0, =&gt; ,ISIC_6705662,\n    06 ,ISIC_6063252, T=0, =&gt; ,ISIC_7195645,\n    07 ,ISIC_6284722, T=0, =&gt; ,ISIC_8262759,\n    08 ,ISIC_6450285, T=0, =&gt; ,ISIC_6548307,\n    09 ,ISIC_7607101, T=0, =&gt; ,ISIC_7675261,\n\nAnd 2 more pairs in test\n\n    01 ,ISIC_2391447, =&gt; ,ISIC_4718575,\n    02 ,ISIC_3091968, =&gt; ,ISIC_5914041, (nearly exact)\n\nI also confirm that no test images have an exact pair in this year's competition train data",
    "922194": "Hey Chris,\n\nThanks so much for looking into this as well! :)\n\nI just ran these through the PHash algorithm, and can confirm the 9 additional matches (identical hashes); for some reason however, the files that I was running the original search over weren't detected as duplicates (I was using md5sum)... I will spend some time tomorrow trying to figure out what happened.\n\nCheers,\n/jochen",
    "922224": "thanks a lot Chris! I will use your list",
    "922418": "Ok, i have also updated my popular training TFRecords and removed the 434 duplicate training images (discussion [here][1]). This will help participants create a more reliable CV. And it will get word out about the duplicate images.\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526",
    "963294": "Thanks a lot !!",
    "964064": "graf10a it seems this might be interesting for you: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930",
    "964657": "For those interested in this topic have a look at the @cdeotte analysis in here [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930). It seems even knowing all the 107 images that are potential leak from the test data exploiting it doesnt give much advantage. \n\nBig thanks to @ucanmakeit175  @benboren  @coreacasa @ajtryt2 and @sirishks",
    "970341": "Thanks for you sharing!"
  },
  "source": "meta"
}