{
  "id": 171801,
  "title": "Exploiting duplicate images",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/171801",
  "author_name": "",
  "post_date": "2020-08-02T14:27:35.494870900Z",
  "votes": 18,
  "comment_count": 29,
  "views": 0,
  "content": "<p>A question to the organizers, @juliaelliott .</p>\n\n<p>In the test set we have duplicate images of 2 types:</p>\n\n<ol>\n<li><p><strong>Duplicates vs 2019 ISIC images</strong>, <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158414\">see the list here</a>. It includes 8 images, with 5 of them target 1. This can be directly copied into one's submission. </p></li>\n<li><p><strong>Duplicates vs the test set itself</strong>, as published by the organizers. It includes pairs which potentially have one of the images in private and one in public. The public one can be resolved unequivocally with a probing submission, and then the private gets its value.</p></li>\n</ol>\n\n<p>Is it allowed to use specifically these two types of leaks? The rule that can potentially prohibit it is, of course, the following: </p>\n\n<blockquote>\n  <p>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</p>\n</blockquote>\n\n<p>If only one of the two is allowed, what is the underlying difference? In case one of them or both are prohibited, can we have peace of mind that the verification team will pay special attention to those specific images, that their targets come purely from one's generic models and not the leaks?</p>",
  "messages": [
    {
      "id": "955293",
      "postDate": "08/02/2020 14:27:35",
      "content": "<p>A question to the organizers, @juliaelliott .</p>\n\n<p>In the test set we have duplicate images of 2 types:</p>\n\n<ol>\n<li><p><strong>Duplicates vs 2019 ISIC images</strong>, <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158414\">see the list here</a>. It includes 8 images, with 5 of them target 1. This can be directly copied into one's submission. </p></li>\n<li><p><strong>Duplicates vs the test set itself</strong>, as published by the organizers. It includes pairs which potentially have one of the images in private and one in public. The public one can be resolved unequivocally with a probing submission, and then the private gets its value.</p></li>\n</ol>\n\n<p>Is it allowed to use specifically these two types of leaks? The rule that can potentially prohibit it is, of course, the following: </p>\n\n<blockquote>\n  <p>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</p>\n</blockquote>\n\n<p>If only one of the two is allowed, what is the underlying difference? In case one of them or both are prohibited, can we have peace of mind that the verification team will pay special attention to those specific images, that their targets come purely from one's generic models and not the leaks?</p>",
      "rawMarkdown": "A question to the organizers, @juliaelliott .\n\nIn the test set we have duplicate images of 2 types:\n\n1. **Duplicates vs 2019 ISIC images**, [see the list here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158414). It includes 8 images, with 5 of them target 1. This can be directly copied into one's submission. \n\n2. **Duplicates vs the test set itself**, as published by the organizers. It includes pairs which potentially have one of the images in private and one in public. The public one can be resolved unequivocally with a probing submission, and then the private gets its value.\n\nIs it allowed to use specifically these two types of leaks? The rule that can potentially prohibit it is, of course, the following: \n\n&gt; Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\n\nIf only one of the two is allowed, what is the underlying difference? In case one of them or both are prohibited, can we have peace of mind that the verification team will pay special attention to those specific images, that their targets come purely from one's generic models and not the leaks?",
      "votes": null
    },
    {
      "id": "955465",
      "postDate": "08/02/2020 16:40:14",
      "content": "<p><a href=\"/zaharch\">@zaharch</a>  I believe that <strong>both options are not allowed</strong> because our code/input-data may not contain specific information derived from test data.</p>\n\n<p>However, you can use this <strong>knowledge</strong> informally, to verify your model's predictions 😏 </p>",
      "rawMarkdown": "zaharch  I believe that **both options are not allowed** because our code/input-data may not contain specific information derived from test data.\n\nHowever, you can use this **knowledge** informally, to verify your model's predictions 😏",
      "votes": null
    },
    {
      "id": "955484",
      "postDate": "08/02/2020 16:51:53",
      "content": "<p>The bigger question is whether the patients in public test are disjoint from the patients in private test. For example, the host says <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943#904255\">here</a>:\n&gt; The test for (exact) duplicates was run over the entire (train + test) dataset, and since we assigned images to partition by patient ID, only the two cases in <a href=\"/yuval6967\">@yuval6967</a>'s post (see above) could have led to a split between partitions.</p>\n\n<p><a href=\"/jwebermsk\">@jwebermsk</a> are you implying that the public and private test were assigned by patient ID too? If so, then duplicates will stay either all in public or all in private and no leakage will occur. If patients have images in both public and private, then there will be more leakage than just duplicate images.</p>",
      "rawMarkdown": "The bigger question is whether the patients in public test are disjoint from the patients in private test. For example, the host says [here][1]:\n&gt; The test for (exact) duplicates was run over the entire (train + test) dataset, and since we assigned images to partition by patient ID, only the two cases in @yuval6967's post (see above) could have led to a split between partitions.\n\n@jwebermsk are you implying that the public and private test were assigned by patient ID too? If so, then duplicates will stay either all in public or all in private and no leakage will occur. If patients have images in both public and private, then there will be more leakage than just duplicate images.\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943#904255",
      "votes": null
    },
    {
      "id": "955502",
      "postDate": "08/02/2020 17:04:56",
      "content": "<p>Why do you say that the code and the input should not contain the test data? It is allowed to pseudo-label, to match distributions between the train and test images, to make post-processing like automatically finding similar images in the test and averaging their predictions etc. </p>\n\n<p>The two types of the leaks are also don't need to be thought of as <em>specific</em> information. Because there are perfectly generic and legitimate algorithms like finding similar images between the train and the test and setting the labels accordingly (it is called k-NN). </p>\n\n<p>Finally, why do you say that it is OK informally, but not OK formally?</p>\n\n<p>Despite disagreement, thank you for sharing your opinion, appreciate it. If you have counter arguments please share.</p>",
      "rawMarkdown": "Why do you say that the code and the input should not contain the test data? It is allowed to pseudo-label, to match distributions between the train and test images, to make post-processing like automatically finding similar images in the test and averaging their predictions etc. \n\nThe two types of the leaks are also don't need to be thought of as *specific* information. Because there are perfectly generic and legitimate algorithms like finding similar images between the train and the test and setting the labels accordingly (it is called k-NN). \n\nFinally, why do you say that it is OK informally, but not OK formally?\n\nDespite disagreement, thank you for sharing your opinion, appreciate it. If you have counter arguments please share.",
      "votes": null
    },
    {
      "id": "955514",
      "postDate": "08/02/2020 17:14:20",
      "content": "<p><a href=\"/zaharch\">@zaharch</a> for the above questions 1&amp;2, I believe that one <strong>may not write code such as</strong>:\n<code>if image in list, then target=1</code>\nFor example, you may not use <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215#931741\">target=1 for ISIC_9207777.jpg</a> even though you know that it is an M! (you must use the output of your model prediction + general post-processing steps only)</p>\n\n<p>However, your point is correct that psuedo-labeling, <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/170846#950967\">psuedo-GAN-generation</a>, kNN are all approved methods, as long as a participant is not cleverly disguising hand-labeling/human-predictions, even if such information is public knowledge and the list is obtained from the Discussion forums.</p>",
      "rawMarkdown": "zaharch for the above questions 1&amp;2, I believe that one **may not write code such as**:\n`if image in list, then target=1`\nFor example, you may not use [target=1 for ISIC_9207777.jpg](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215#931741) even though you know that it is an M! (you must use the output of your model prediction + general post-processing steps only)\n\nHowever, your point is correct that psuedo-labeling, [psuedo-GAN-generation](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/170846#950967), kNN are all approved methods, as long as a participant is not cleverly disguising hand-labeling/human-predictions, even if such information is public knowledge and the list is obtained from the Discussion forums.",
      "votes": null
    },
    {
      "id": "955592",
      "postDate": "08/02/2020 18:17:35",
      "content": "<p>This is not manual labeling, so it should be allowed. You can write a few lines of code that compare images and detect duplicates and label them accordingly. In basic terms, this is what a ML model does.</p>",
      "rawMarkdown": "This is not manual labeling, so it should be allowed. You can write a few lines of code that compare images and detect duplicates and label them accordingly. In basic terms, this is what a ML model does.",
      "votes": null
    },
    {
      "id": "955628",
      "postDate": "08/02/2020 18:48:51",
      "content": "<p>I feel like this should not be allowed but there's a real struggle because how can Kaggle find out about this if you are not one of the winning teams?\n5-10 targets set to one makes quite a difference of private score I think, so now that the idea is publicly out how can you prevent everyone from doing this in order to win medals?</p>\n\n<p>I would say that the best thing to do is remove all known duplicates from final scoring, but I guess that's not going to happen...</p>",
      "rawMarkdown": "I feel like this should not be allowed but there's a real struggle because how can Kaggle find out about this if you are not one of the winning teams?\n5-10 targets set to one makes quite a difference of private score I think, so now that the idea is publicly out how can you prevent everyone from doing this in order to win medals?\n\nI would say that the best thing to do is remove all known duplicates from final scoring, but I guess that's not going to happen...",
      "votes": null
    },
    {
      "id": "955635",
      "postDate": "08/02/2020 19:01:11",
      "content": "<p>I agree that removing all the duplicates from the test seems to be the correct way. But they decided not to do it, the duplicates have being a known issue for 2 months.</p>",
      "rawMarkdown": "I agree that removing all the duplicates from the test seems to be the correct way. But they decided not to do it, the duplicates have being a known issue for 2 months.",
      "votes": null
    },
    {
      "id": "955653",
      "postDate": "08/02/2020 19:26:37",
      "content": "<p>Of course you are allowed to use information from test. </p>",
      "rawMarkdown": "Of course you are allowed to use information from test.",
      "votes": null
    },
    {
      "id": "955669",
      "postDate": "08/02/2020 20:10:34",
      "content": "<p>Looks like I'll have to burn submission to test the truth value of test duplicates.  I wish someone share these.</p>",
      "rawMarkdown": "Looks like I'll have to burn submission to test the truth value of test duplicates.  I wish someone share these.",
      "votes": null
    },
    {
      "id": "955683",
      "postDate": "08/02/2020 20:18:59",
      "content": "<p>Not sure we will have time for that so I can just support your wish.</p>",
      "rawMarkdown": "Not sure we will have time for that so I can just support your wish.",
      "votes": null
    },
    {
      "id": "955693",
      "postDate": "08/02/2020 20:31:24",
      "content": "<p>There are only 6 known test answers (five target=1, one target=0). I posted them <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\">here</a> (i.e. 6 test images exist in last year's 2019 data). There are no other duplicates between test and anything outside of test. (I've compared test with 2020 train, 2019, 2018, 2017 and ISIC-archive) </p>\n\n<p>The other duplicates being discussed here are duplicates within test. There may be 100 images (i forget) in test that have a duplicate in test. The full list has been provided by host <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\">here</a>. So the concern is that one copy is in public and one copy is in private. If that were the case, then we would want to overfit public LB.</p>",
      "rawMarkdown": "There are only 6 known test answers (five target=1, one target=0). I posted them [here][1] (i.e. 6 test images exist in last year's 2019 data). There are no other duplicates between test and anything outside of test. (I've compared test with 2020 train, 2019, 2018, 2017 and ISIC-archive) \n\nThe other duplicates being discussed here are duplicates within test. There may be 100 images (i forget) in test that have a duplicate in test. The full list has been provided by host [here][2]. So the concern is that one copy is in public and one copy is in private. If that were the case, then we would want to overfit public LB.\n\n[1]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943",
      "votes": null
    },
    {
      "id": "955707",
      "postDate": "08/02/2020 20:44:49",
      "content": "<p>And that's what <a href=\"/cpmpml\">@cpmpml</a> refers to - namely probing these 100 images. </p>",
      "rawMarkdown": "And that's what @cpmpml refers to - namely probing these 100 images.",
      "votes": null
    },
    {
      "id": "955714",
      "postDate": "08/02/2020 20:57:51",
      "content": "<p>Chris, it is not 6, it is 8.</p>\n\n<p><code>\n['ISIC_0014506', 'ISIC_9353360'],\n['ISIC_0014518', 'ISIC_9207777'],\n['ISIC_0014541', 'ISIC_5224960'],\n['ISIC_0014542', 'ISIC_6457527'],\n['ISIC_0030762', 'ISIC_3689290'],\n['ISIC_0025740', 'ISIC_3584949'],\n['ISIC_0014644', 'ISIC_8347588'],\n['ISIC_0014525', 'ISIC_8372206']\n</code></p>",
      "rawMarkdown": "Chris, it is not 6, it is 8.\n\n```\n['ISIC_0014506', 'ISIC_9353360'],\n['ISIC_0014518', 'ISIC_9207777'],\n['ISIC_0014541', 'ISIC_5224960'],\n['ISIC_0014542', 'ISIC_6457527'],\n['ISIC_0030762', 'ISIC_3689290'],\n['ISIC_0025740', 'ISIC_3584949'],\n['ISIC_0014644', 'ISIC_8347588'],\n['ISIC_0014525', 'ISIC_8372206']\n```",
      "votes": null
    },
    {
      "id": "955735",
      "postDate": "08/02/2020 21:40:53",
      "content": "<p><a href=\"/sirishks\">@sirishks</a> it seems that you might help us here?\nYou know if your current score is from a real model you are going to save a looot of lives! (and I’m going to love reading your final write up)\nOtherwise I’m sure you can help.</p>",
      "rawMarkdown": "sirishks it seems that you might help us here?\nYou know if your current score is from a real model you are going to save a looot of lives! (and I’m going to love reading your final write up)\nOtherwise I’m sure you can help.",
      "votes": null
    },
    {
      "id": "955802",
      "postDate": "08/03/2020 01:01:34",
      "content": "<p>&gt; Chris, it is not 6, it is 8.</p>\n\n<p>Be careful. 6 are exact duplicates. 2 are very similar. You take a risk using the similar ones.</p>",
      "rawMarkdown": "&gt; Chris, it is not 6, it is 8.\n\nBe careful. 6 are exact duplicates. 2 are very similar. You take a risk using the similar ones.",
      "votes": null
    },
    {
      "id": "955805",
      "postDate": "08/03/2020 01:03:40",
      "content": "<blockquote>\n  <p>And that's what <a href=\"/cpmpml\">@cpmpml</a> refers to - namely probing these 100 images.</p>\n</blockquote>\n\n<p>Wow! I always assumed that the duplicates will both be in public or both be in private because they are the same patient. If they are split between public and private, that means that patients are being split between public and private.</p>",
      "rawMarkdown": "&gt; And that's what @cpmpml refers to - namely probing these 100 images.\n\nWow! I always assumed that the duplicates will both be in public or both be in private because they are the same patient. If they are split between public and private, that means that patients are being split between public and private.",
      "votes": null
    },
    {
      "id": "955839",
      "postDate": "08/03/2020 02:10:51",
      "content": "<p>If the same <code>patient_id</code> is never in both public and private, then we only need to LB probe 20 duplicate test images posted <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930\">here</a></p>",
      "rawMarkdown": "If the same `patient_id` is never in both public and private, then we only need to LB probe 20 duplicate test images posted [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930",
      "votes": null
    },
    {
      "id": "955982",
      "postDate": "08/03/2020 06:07:45",
      "content": "<p>If the hosts don't provide a timely response, then I see a potential gold notebook or discussion, where the editor shares the results of LB probing.</p>",
      "rawMarkdown": "If the hosts don't provide a timely response, then I see a potential gold notebook or discussion, where the editor shares the results of LB probing.",
      "votes": null
    },
    {
      "id": "956443",
      "postDate": "08/03/2020 14:01:57",
      "content": "<p><a href=\"/optimo\">@optimo</a>  <a href=\"/rohitagarwal\">@rohitagarwal</a>  See <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930#956437\">my reply</a> along with calculation details in the new topic created by Chris.</p>\n\n<p><a href=\"/rohitagarwal\">@rohitagarwal</a> should I expect a \"Gold\" for this reply 😂 ?</p>",
      "rawMarkdown": "optimo  @rohitagarwal  See [my reply](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930#956437) along with calculation details in the new topic created by Chris.\n\n@rohitagarwal should I expect a \"Gold\" for this reply 😂 ?",
      "votes": null
    },
    {
      "id": "956587",
      "postDate": "08/03/2020 16:03:17",
      "content": "<p>Where is it written that duplicates can't occur in both public and private? I wish it was the case.</p>",
      "rawMarkdown": "Where is it written that duplicates can't occur in both public and private? I wish it was the case.",
      "votes": null
    },
    {
      "id": "956592",
      "postDate": "08/03/2020 16:04:54",
      "content": "<p>I designed an experiment based on Sirish's experiment to test all 214 duplicates in test <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930#956548\">here</a>. With a little help from the community, we can do it with 4 submissions.</p>",
      "rawMarkdown": "I designed an experiment based on Sirish's experiment to test all 214 duplicates in test [here][1]. With a little help from the community, we can do it with 4 submissions.\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930#956548",
      "votes": null
    },
    {
      "id": "956698",
      "postDate": "08/03/2020 17:39:29",
      "content": "<p>Man, I was your first upvote on that reply.</p>",
      "rawMarkdown": "Man, I was your first upvote on that reply.",
      "votes": null
    },
    {
      "id": "957726",
      "postDate": "08/04/2020 14:11:34",
      "content": "<p>UPDATE: All 214 duplicate test images have been probed. None of them are both in public and malignant. Therefore no duplicate will give anyone an advantage on private LB. Probes posted <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930\">here</a></p>",
      "rawMarkdown": "UPDATE: All 214 duplicate test images have been probed. None of them are both in public and malignant. Therefore no duplicate will give anyone an advantage on private LB. Probes posted [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930",
      "votes": null
    },
    {
      "id": "958949",
      "postDate": "08/05/2020 08:31:10",
      "content": "<p>Checking my own model, it strongly predicts correct class for all of those, and I assume it does for most.......so they aren't really borderline cases, they are likely not to give much trouble.</p>",
      "rawMarkdown": "Checking my own model, it strongly predicts correct class for all of those, and I assume it does for most.......so they aren't really borderline cases, they are likely not to give much trouble.",
      "votes": null
    },
    {
      "id": "959074",
      "postDate": "08/05/2020 10:14:44",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> I wanted to ask you a question about DBScan for finding duplicates.</p>\n\n<p>It seems that this approach search for near exact duplicates, but it does not include \"similar duplicates\" (image rotation, shifts, zoom etc...) But it has been shown a long time ago that similar duplicates do exist (see here <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155306#869849\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155306#869849</a>).</p>\n\n<p>Have you tried to perform a \"similar duplicates\" search? Independently of potential leakage, I'm interested in what would be a good approach to find such transformation in a database? I think it would also be very useful for future competitions. Any idea?</p>",
      "rawMarkdown": "cdeotte I wanted to ask you a question about DBScan for finding duplicates.\n\nIt seems that this approach search for near exact duplicates, but it does not include \"similar duplicates\" (image rotation, shifts, zoom etc...) But it has been shown a long time ago that similar duplicates do exist (see here https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155306#869849).\n\nHave you tried to perform a \"similar duplicates\" search? Independently of potential leakage, I'm interested in what would be a good approach to find such transformation in a database? I think it would also be very useful for future competitions. Any idea?",
      "votes": null
    },
    {
      "id": "959087",
      "postDate": "08/05/2020 10:26:51",
      "content": "<p><a href=\"/optimo\">@optimo</a>  You can use <strong>CNN</strong> for finding duplicates.</p>\n\n<p>For example, <a href=\"https://github.com/idealo/imagededup\">Imagededup</a> has CNN option also.</p>",
      "rawMarkdown": "optimo  You can use **CNN** for finding duplicates.\n\nFor example, [Imagededup](https://github.com/idealo/imagededup) has CNN option also.",
      "votes": null
    },
    {
      "id": "959361",
      "postDate": "08/05/2020 14:28:01",
      "content": "<p>Using RAPIDS cuML kNN, you can find \"similar duplicates\" (image rotation, shifts, zooms, etc). Just increase the distance threshold for similarity in my notebook <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\">here</a>. (In Flower comp, i found similar rotated flowers using this method).</p>\n\n<p>If you want to encourage that notebook to find more \"similar duplicates\", then perhaps you can create a for-loop that extracts embeddings and performs kNN. Each time you extract embeddings, apply different random data augmentation of rotation, zoom, etc. beforehand (to \"one set\").</p>\n\n<p>(By \"one set\", i mean if you're trying to find duplicates of test in 2019. Then leave 2019 alone, and apply random augmentation to test only).</p>",
      "rawMarkdown": "Using RAPIDS cuML kNN, you can find \"similar duplicates\" (image rotation, shifts, zooms, etc). Just increase the distance threshold for similarity in my notebook [here][1]. (In Flower comp, i found similar rotated flowers using this method).\n\nIf you want to encourage that notebook to find more \"similar duplicates\", then perhaps you can create a for-loop that extracts embeddings and performs kNN. Each time you extract embeddings, apply different random data augmentation of rotation, zoom, etc. beforehand (to \"one set\").\n\n(By \"one set\", i mean if you're trying to find duplicates of test in 2019. Then leave 2019 alone, and apply random augmentation to test only).\n\n[1]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates",
      "votes": null
    },
    {
      "id": "959480",
      "postDate": "08/05/2020 16:04:14",
      "content": "<p>Thanks I was thinking about something like this, will share if I do try it!</p>",
      "rawMarkdown": "Thanks I was thinking about something like this, will share if I do try it!",
      "votes": null
    },
    {
      "id": "959706",
      "postDate": "08/05/2020 20:48:23",
      "content": "<p><a href=\"/zaharch\">@zaharch</a> - If you are not hand-labeling the test set, use of duplicates knowledge is fine. For example, the method <a href=\"/philippsinger\">@philippsinger</a> describes could be a valid, non-prohibited method, as long as it is fully automated as he describes and you are not anchoring on the manual labeling of any part of this competition's test set (public &amp; private included) -- <strong>unless</strong> it is already provided as labeled outside of this competition's test set. For example: The known label of a duplicate from 2019 or from this competition's training set that also appears in this year's test set would be okay to use, because you haven't manually labeled the test set. </p>",
      "rawMarkdown": "zaharch - If you are not hand-labeling the test set, use of duplicates knowledge is fine. For example, the method @philippsinger describes could be a valid, non-prohibited method, as long as it is fully automated as he describes and you are not anchoring on the manual labeling of any part of this competition's test set (public &amp; private included) -- **unless** it is already provided as labeled outside of this competition's test set. For example: The known label of a duplicate from 2019 or from this competition's training set that also appears in this year's test set would be okay to use, because you haven't manually labeled the test set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 955465,
      "author_name": "sirishks",
      "author_url": "",
      "post_date": "08/02/2020 16:40:14",
      "content": "<p><a href=\"/zaharch\">@zaharch</a>  I believe that <strong>both options are not allowed</strong> because our code/input-data may not contain specific information derived from test data.</p>\n\n<p>However, you can use this <strong>knowledge</strong> informally, to verify your model's predictions 😏 </p>",
      "votes": null,
      "replies": [
        {
          "id": 955502,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "08/02/2020 17:04:56",
          "content": "<p>Why do you say that the code and the input should not contain the test data? It is allowed to pseudo-label, to match distributions between the train and test images, to make post-processing like automatically finding similar images in the test and averaging their predictions etc. </p>\n\n<p>The two types of the leaks are also don't need to be thought of as <em>specific</em> information. Because there are perfectly generic and legitimate algorithms like finding similar images between the train and the test and setting the labels accordingly (it is called k-NN). </p>\n\n<p>Finally, why do you say that it is OK informally, but not OK formally?</p>\n\n<p>Despite disagreement, thank you for sharing your opinion, appreciate it. If you have counter arguments please share.</p>",
          "votes": null,
          "replies": [
            {
              "id": 955514,
              "author_name": "sirishks",
              "author_url": "",
              "post_date": "08/02/2020 17:14:20",
              "content": "<p><a href=\"/zaharch\">@zaharch</a> for the above questions 1&amp;2, I believe that one <strong>may not write code such as</strong>:\n<code>if image in list, then target=1</code>\nFor example, you may not use <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215#931741\">target=1 for ISIC_9207777.jpg</a> even though you know that it is an M! (you must use the output of your model prediction + general post-processing steps only)</p>\n\n<p>However, your point is correct that psuedo-labeling, <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/170846#950967\">psuedo-GAN-generation</a>, kNN are all approved methods, as long as a participant is not cleverly disguising hand-labeling/human-predictions, even if such information is public knowledge and the list is obtained from the Discussion forums.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 955653,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "08/02/2020 19:26:37",
          "content": "<p>Of course you are allowed to use information from test. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 955484,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/02/2020 16:51:53",
      "content": "<p>The bigger question is whether the patients in public test are disjoint from the patients in private test. For example, the host says <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943#904255\">here</a>:\n&gt; The test for (exact) duplicates was run over the entire (train + test) dataset, and since we assigned images to partition by patient ID, only the two cases in <a href=\"/yuval6967\">@yuval6967</a>'s post (see above) could have led to a split between partitions.</p>\n\n<p><a href=\"/jwebermsk\">@jwebermsk</a> are you implying that the public and private test were assigned by patient ID too? If so, then duplicates will stay either all in public or all in private and no leakage will occur. If patients have images in both public and private, then there will be more leakage than just duplicate images.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 955592,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "08/02/2020 18:17:35",
      "content": "<p>This is not manual labeling, so it should be allowed. You can write a few lines of code that compare images and detect duplicates and label them accordingly. In basic terms, this is what a ML model does.</p>",
      "votes": null,
      "replies": [
        {
          "id": 958949,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/05/2020 08:31:10",
          "content": "<p>Checking my own model, it strongly predicts correct class for all of those, and I assume it does for most.......so they aren't really borderline cases, they are likely not to give much trouble.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 955628,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "08/02/2020 18:48:51",
      "content": "<p>I feel like this should not be allowed but there's a real struggle because how can Kaggle find out about this if you are not one of the winning teams?\n5-10 targets set to one makes quite a difference of private score I think, so now that the idea is publicly out how can you prevent everyone from doing this in order to win medals?</p>\n\n<p>I would say that the best thing to do is remove all known duplicates from final scoring, but I guess that's not going to happen...</p>",
      "votes": null,
      "replies": [
        {
          "id": 955635,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "08/02/2020 19:01:11",
          "content": "<p>I agree that removing all the duplicates from the test seems to be the correct way. But they decided not to do it, the duplicates have being a known issue for 2 months.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 956443,
          "author_name": "sirishks",
          "author_url": "",
          "post_date": "08/03/2020 14:01:57",
          "content": "<p><a href=\"/optimo\">@optimo</a>  <a href=\"/rohitagarwal\">@rohitagarwal</a>  See <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930#956437\">my reply</a> along with calculation details in the new topic created by Chris.</p>\n\n<p><a href=\"/rohitagarwal\">@rohitagarwal</a> should I expect a \"Gold\" for this reply 😂 ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 956698,
          "author_name": "rohitagarwal",
          "author_url": "",
          "post_date": "08/03/2020 17:39:29",
          "content": "<p>Man, I was your first upvote on that reply.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 955669,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/02/2020 20:10:34",
      "content": "<p>Looks like I'll have to burn submission to test the truth value of test duplicates.  I wish someone share these.</p>",
      "votes": null,
      "replies": [
        {
          "id": 955683,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "08/02/2020 20:18:59",
          "content": "<p>Not sure we will have time for that so I can just support your wish.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 955693,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/02/2020 20:31:24",
          "content": "<p>There are only 6 known test answers (five target=1, one target=0). I posted them <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\">here</a> (i.e. 6 test images exist in last year's 2019 data). There are no other duplicates between test and anything outside of test. (I've compared test with 2020 train, 2019, 2018, 2017 and ISIC-archive) </p>\n\n<p>The other duplicates being discussed here are duplicates within test. There may be 100 images (i forget) in test that have a duplicate in test. The full list has been provided by host <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\">here</a>. So the concern is that one copy is in public and one copy is in private. If that were the case, then we would want to overfit public LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 955707,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "08/02/2020 20:44:49",
          "content": "<p>And that's what <a href=\"/cpmpml\">@cpmpml</a> refers to - namely probing these 100 images. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 955714,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "08/02/2020 20:57:51",
          "content": "<p>Chris, it is not 6, it is 8.</p>\n\n<p><code>\n['ISIC_0014506', 'ISIC_9353360'],\n['ISIC_0014518', 'ISIC_9207777'],\n['ISIC_0014541', 'ISIC_5224960'],\n['ISIC_0014542', 'ISIC_6457527'],\n['ISIC_0030762', 'ISIC_3689290'],\n['ISIC_0025740', 'ISIC_3584949'],\n['ISIC_0014644', 'ISIC_8347588'],\n['ISIC_0014525', 'ISIC_8372206']\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 955735,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "08/02/2020 21:40:53",
          "content": "<p><a href=\"/sirishks\">@sirishks</a> it seems that you might help us here?\nYou know if your current score is from a real model you are going to save a looot of lives! (and I’m going to love reading your final write up)\nOtherwise I’m sure you can help.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 955802,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/03/2020 01:01:34",
          "content": "<p>&gt; Chris, it is not 6, it is 8.</p>\n\n<p>Be careful. 6 are exact duplicates. 2 are very similar. You take a risk using the similar ones.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 955805,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/03/2020 01:03:40",
          "content": "<blockquote>\n  <p>And that's what <a href=\"/cpmpml\">@cpmpml</a> refers to - namely probing these 100 images.</p>\n</blockquote>\n\n<p>Wow! I always assumed that the duplicates will both be in public or both be in private because they are the same patient. If they are split between public and private, that means that patients are being split between public and private.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 955839,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/03/2020 02:10:51",
          "content": "<p>If the same <code>patient_id</code> is never in both public and private, then we only need to LB probe 20 duplicate test images posted <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930\">here</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 956587,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/03/2020 16:03:17",
          "content": "<p>Where is it written that duplicates can't occur in both public and private? I wish it was the case.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 955982,
      "author_name": "rohitagarwal",
      "author_url": "",
      "post_date": "08/03/2020 06:07:45",
      "content": "<p>If the hosts don't provide a timely response, then I see a potential gold notebook or discussion, where the editor shares the results of LB probing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 956592,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/03/2020 16:04:54",
      "content": "<p>I designed an experiment based on Sirish's experiment to test all 214 duplicates in test <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930#956548\">here</a>. With a little help from the community, we can do it with 4 submissions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 957726,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/04/2020 14:11:34",
      "content": "<p>UPDATE: All 214 duplicate test images have been probed. None of them are both in public and malignant. Therefore no duplicate will give anyone an advantage on private LB. Probes posted <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 959074,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "08/05/2020 10:14:44",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> I wanted to ask you a question about DBScan for finding duplicates.</p>\n\n<p>It seems that this approach search for near exact duplicates, but it does not include \"similar duplicates\" (image rotation, shifts, zoom etc...) But it has been shown a long time ago that similar duplicates do exist (see here <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155306#869849\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155306#869849</a>).</p>\n\n<p>Have you tried to perform a \"similar duplicates\" search? Independently of potential leakage, I'm interested in what would be a good approach to find such transformation in a database? I think it would also be very useful for future competitions. Any idea?</p>",
          "votes": null,
          "replies": [
            {
              "id": 959087,
              "author_name": "sirishks",
              "author_url": "",
              "post_date": "08/05/2020 10:26:51",
              "content": "<p><a href=\"/optimo\">@optimo</a>  You can use <strong>CNN</strong> for finding duplicates.</p>\n\n<p>For example, <a href=\"https://github.com/idealo/imagededup\">Imagededup</a> has CNN option also.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 959361,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/05/2020 14:28:01",
          "content": "<p>Using RAPIDS cuML kNN, you can find \"similar duplicates\" (image rotation, shifts, zooms, etc). Just increase the distance threshold for similarity in my notebook <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\">here</a>. (In Flower comp, i found similar rotated flowers using this method).</p>\n\n<p>If you want to encourage that notebook to find more \"similar duplicates\", then perhaps you can create a for-loop that extracts embeddings and performs kNN. Each time you extract embeddings, apply different random data augmentation of rotation, zoom, etc. beforehand (to \"one set\").</p>\n\n<p>(By \"one set\", i mean if you're trying to find duplicates of test in 2019. Then leave 2019 alone, and apply random augmentation to test only).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 959480,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "08/05/2020 16:04:14",
          "content": "<p>Thanks I was thinking about something like this, will share if I do try it!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 959706,
      "author_name": "juliaelliott",
      "author_url": "",
      "post_date": "08/05/2020 20:48:23",
      "content": "<p><a href=\"/zaharch\">@zaharch</a> - If you are not hand-labeling the test set, use of duplicates knowledge is fine. For example, the method <a href=\"/philippsinger\">@philippsinger</a> describes could be a valid, non-prohibited method, as long as it is fully automated as he describes and you are not anchoring on the manual labeling of any part of this competition's test set (public &amp; private included) -- <strong>unless</strong> it is already provided as labeled outside of this competition's test set. For example: The known label of a duplicate from 2019 or from this competition's training set that also appears in this year's test set would be okay to use, because you haven't manually labeled the test set. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "955293": "A question to the organizers, @juliaelliott .\n\nIn the test set we have duplicate images of 2 types:\n\n1. **Duplicates vs 2019 ISIC images**, [see the list here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158414). It includes 8 images, with 5 of them target 1. This can be directly copied into one's submission. \n\n2. **Duplicates vs the test set itself**, as published by the organizers. It includes pairs which potentially have one of the images in private and one in public. The public one can be resolved unequivocally with a probing submission, and then the private gets its value.\n\nIs it allowed to use specifically these two types of leaks? The rule that can potentially prohibit it is, of course, the following: \n\n&gt; Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\n\nIf only one of the two is allowed, what is the underlying difference? In case one of them or both are prohibited, can we have peace of mind that the verification team will pay special attention to those specific images, that their targets come purely from one's generic models and not the leaks?",
    "955465": "zaharch  I believe that **both options are not allowed** because our code/input-data may not contain specific information derived from test data.\n\nHowever, you can use this **knowledge** informally, to verify your model's predictions 😏",
    "955484": "The bigger question is whether the patients in public test are disjoint from the patients in private test. For example, the host says [here][1]:\n&gt; The test for (exact) duplicates was run over the entire (train + test) dataset, and since we assigned images to partition by patient ID, only the two cases in @yuval6967's post (see above) could have led to a split between partitions.\n\n@jwebermsk are you implying that the public and private test were assigned by patient ID too? If so, then duplicates will stay either all in public or all in private and no leakage will occur. If patients have images in both public and private, then there will be more leakage than just duplicate images.\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943#904255",
    "955502": "Why do you say that the code and the input should not contain the test data? It is allowed to pseudo-label, to match distributions between the train and test images, to make post-processing like automatically finding similar images in the test and averaging their predictions etc. \n\nThe two types of the leaks are also don't need to be thought of as *specific* information. Because there are perfectly generic and legitimate algorithms like finding similar images between the train and the test and setting the labels accordingly (it is called k-NN). \n\nFinally, why do you say that it is OK informally, but not OK formally?\n\nDespite disagreement, thank you for sharing your opinion, appreciate it. If you have counter arguments please share.",
    "955514": "zaharch for the above questions 1&amp;2, I believe that one **may not write code such as**:\n`if image in list, then target=1`\nFor example, you may not use [target=1 for ISIC_9207777.jpg](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215#931741) even though you know that it is an M! (you must use the output of your model prediction + general post-processing steps only)\n\nHowever, your point is correct that psuedo-labeling, [psuedo-GAN-generation](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/170846#950967), kNN are all approved methods, as long as a participant is not cleverly disguising hand-labeling/human-predictions, even if such information is public knowledge and the list is obtained from the Discussion forums.",
    "955592": "This is not manual labeling, so it should be allowed. You can write a few lines of code that compare images and detect duplicates and label them accordingly. In basic terms, this is what a ML model does.",
    "955628": "I feel like this should not be allowed but there's a real struggle because how can Kaggle find out about this if you are not one of the winning teams?\n5-10 targets set to one makes quite a difference of private score I think, so now that the idea is publicly out how can you prevent everyone from doing this in order to win medals?\n\nI would say that the best thing to do is remove all known duplicates from final scoring, but I guess that's not going to happen...",
    "955635": "I agree that removing all the duplicates from the test seems to be the correct way. But they decided not to do it, the duplicates have being a known issue for 2 months.",
    "955653": "Of course you are allowed to use information from test.",
    "955669": "Looks like I'll have to burn submission to test the truth value of test duplicates.  I wish someone share these.",
    "955683": "Not sure we will have time for that so I can just support your wish.",
    "955693": "There are only 6 known test answers (five target=1, one target=0). I posted them [here][1] (i.e. 6 test images exist in last year's 2019 data). There are no other duplicates between test and anything outside of test. (I've compared test with 2020 train, 2019, 2018, 2017 and ISIC-archive) \n\nThe other duplicates being discussed here are duplicates within test. There may be 100 images (i forget) in test that have a duplicate in test. The full list has been provided by host [here][2]. So the concern is that one copy is in public and one copy is in private. If that were the case, then we would want to overfit public LB.\n\n[1]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943",
    "955707": "And that's what @cpmpml refers to - namely probing these 100 images.",
    "955714": "Chris, it is not 6, it is 8.\n\n```\n['ISIC_0014506', 'ISIC_9353360'],\n['ISIC_0014518', 'ISIC_9207777'],\n['ISIC_0014541', 'ISIC_5224960'],\n['ISIC_0014542', 'ISIC_6457527'],\n['ISIC_0030762', 'ISIC_3689290'],\n['ISIC_0025740', 'ISIC_3584949'],\n['ISIC_0014644', 'ISIC_8347588'],\n['ISIC_0014525', 'ISIC_8372206']\n```",
    "955735": "sirishks it seems that you might help us here?\nYou know if your current score is from a real model you are going to save a looot of lives! (and I’m going to love reading your final write up)\nOtherwise I’m sure you can help.",
    "955802": "&gt; Chris, it is not 6, it is 8.\n\nBe careful. 6 are exact duplicates. 2 are very similar. You take a risk using the similar ones.",
    "955805": "&gt; And that's what @cpmpml refers to - namely probing these 100 images.\n\nWow! I always assumed that the duplicates will both be in public or both be in private because they are the same patient. If they are split between public and private, that means that patients are being split between public and private.",
    "955839": "If the same `patient_id` is never in both public and private, then we only need to LB probe 20 duplicate test images posted [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930",
    "955982": "If the hosts don't provide a timely response, then I see a potential gold notebook or discussion, where the editor shares the results of LB probing.",
    "956443": "optimo  @rohitagarwal  See [my reply](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930#956437) along with calculation details in the new topic created by Chris.\n\n@rohitagarwal should I expect a \"Gold\" for this reply 😂 ?",
    "956587": "Where is it written that duplicates can't occur in both public and private? I wish it was the case.",
    "956592": "I designed an experiment based on Sirish's experiment to test all 214 duplicates in test [here][1]. With a little help from the community, we can do it with 4 submissions.\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930#956548",
    "956698": "Man, I was your first upvote on that reply.",
    "957726": "UPDATE: All 214 duplicate test images have been probed. None of them are both in public and malignant. Therefore no duplicate will give anyone an advantage on private LB. Probes posted [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930",
    "958949": "Checking my own model, it strongly predicts correct class for all of those, and I assume it does for most.......so they aren't really borderline cases, they are likely not to give much trouble.",
    "959074": "cdeotte I wanted to ask you a question about DBScan for finding duplicates.\n\nIt seems that this approach search for near exact duplicates, but it does not include \"similar duplicates\" (image rotation, shifts, zoom etc...) But it has been shown a long time ago that similar duplicates do exist (see here https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155306#869849).\n\nHave you tried to perform a \"similar duplicates\" search? Independently of potential leakage, I'm interested in what would be a good approach to find such transformation in a database? I think it would also be very useful for future competitions. Any idea?",
    "959087": "optimo  You can use **CNN** for finding duplicates.\n\nFor example, [Imagededup](https://github.com/idealo/imagededup) has CNN option also.",
    "959361": "Using RAPIDS cuML kNN, you can find \"similar duplicates\" (image rotation, shifts, zooms, etc). Just increase the distance threshold for similarity in my notebook [here][1]. (In Flower comp, i found similar rotated flowers using this method).\n\nIf you want to encourage that notebook to find more \"similar duplicates\", then perhaps you can create a for-loop that extracts embeddings and performs kNN. Each time you extract embeddings, apply different random data augmentation of rotation, zoom, etc. beforehand (to \"one set\").\n\n(By \"one set\", i mean if you're trying to find duplicates of test in 2019. Then leave 2019 alone, and apply random augmentation to test only).\n\n[1]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates",
    "959480": "Thanks I was thinking about something like this, will share if I do try it!",
    "959706": "zaharch - If you are not hand-labeling the test set, use of duplicates knowledge is fine. For example, the method @philippsinger describes could be a valid, non-prohibited method, as long as it is fully automated as he describes and you are not anchoring on the manual labeling of any part of this competition's test set (public &amp; private included) -- **unless** it is already provided as labeled outside of this competition's test set. For example: The known label of a duplicate from 2019 or from this competition's training set that also appears in this year's test set would be okay to use, because you haven't manually labeled the test set."
  },
  "source": "meta"
}