{
  "id": 34788,
  "title": "Stage 2 data is not purely unseen data ?",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/34788",
  "author_name": "",
  "post_date": "2017-06-15T16:49:04.779063Z",
  "votes": 31,
  "comment_count": 28,
  "views": 0,
  "content": "<p>It seems quite many Stage 2 \"unseen data\" already appeared in Stage 1 train/test set.... i.e. leaks\nFor examples :</p>\n\n<ul>\n<li>test_stg2/10698.jpg  --&gt; train/Type_2/485.jpg </li>\n<li>test_stg2/11249.jpg   --&gt; train/Type_2/485.jpg </li>\n<li>test_stg2/13364.jpg --&gt; train/Type_1/27.jpg</li>\n<li>test_stg2/10898.jpg --&gt; train/Type_1/265.jpg  </li>\n<li>test_stg2/13456.jpg --&gt; test/0.jpg</li>\n<li>test_stg2/10398.jpg --&gt; test/98.jpg</li>\n<li>test_stg2/10844.jpg --&gt; test/327.jpg</li>\n<li>test_stg2/11334.jpg --&gt; test/1068.jpg</li>\n<li>etc... I am pretty sure there are more we can find out</li>\n</ul>\n\n<p>Its a bit disappointing , because that means models that overfit to <strong>REMEMBER</strong> stage 1 data could out-perform those models that are actually <strong>PREDICTING</strong>..... which defect the purpose of this challenge.</p>\n\n<p>Would the organizer/admin consider excluding those leaks in Stage 2 scoring  ?</p>",
  "messages": [
    {
      "id": "193117",
      "postDate": "06/15/2017 16:49:04",
      "content": "<p>It seems quite many Stage 2 \"unseen data\" already appeared in Stage 1 train/test set.... i.e. leaks\nFor examples :</p>\n\n<ul>\n<li>test_stg2/10698.jpg  --&gt; train/Type_2/485.jpg </li>\n<li>test_stg2/11249.jpg   --&gt; train/Type_2/485.jpg </li>\n<li>test_stg2/13364.jpg --&gt; train/Type_1/27.jpg</li>\n<li>test_stg2/10898.jpg --&gt; train/Type_1/265.jpg  </li>\n<li>test_stg2/13456.jpg --&gt; test/0.jpg</li>\n<li>test_stg2/10398.jpg --&gt; test/98.jpg</li>\n<li>test_stg2/10844.jpg --&gt; test/327.jpg</li>\n<li>test_stg2/11334.jpg --&gt; test/1068.jpg</li>\n<li>etc... I am pretty sure there are more we can find out</li>\n</ul>\n\n<p>Its a bit disappointing , because that means models that overfit to <strong>REMEMBER</strong> stage 1 data could out-perform those models that are actually <strong>PREDICTING</strong>..... which defect the purpose of this challenge.</p>\n\n<p>Would the organizer/admin consider excluding those leaks in Stage 2 scoring  ?</p>",
      "rawMarkdown": "It seems quite many Stage 2 \"unseen data\" already appeared in Stage 1 train/test set.... i.e. leaks\nFor examples :\n\n - test_stg2/10698.jpg  --&gt; train/Type_2/485.jpg \n - test_stg2/11249.jpg   --&gt; train/Type_2/485.jpg \n - test_stg2/13364.jpg --&gt; train/Type_1/27.jpg\n - test_stg2/10898.jpg --&gt; train/Type_1/265.jpg\t\n - test_stg2/13456.jpg --&gt; test/0.jpg\n - test_stg2/10398.jpg --&gt; test/98.jpg\n - test_stg2/10844.jpg --&gt; test/327.jpg\n - test_stg2/11334.jpg --&gt; test/1068.jpg\n - etc... I am pretty sure there are more we can find out\n\nIts a bit disappointing , because that means models that overfit to **REMEMBER** stage 1 data could out-perform those models that are actually **PREDICTING**..... which defect the purpose of this challenge.\n\nWould the organizer/admin consider excluding those leaks in Stage 2 scoring  ?",
      "votes": null
    },
    {
      "id": "193128",
      "postDate": "06/15/2017 17:14:18",
      "content": "<p>I believe someone had written a method to determine images that are similar. I'm assuming that these images are exact matches but I wonder if there are any images like in the additional dataset that are of the same patient but a slightly different image(camera moved slightly between taking images).</p>",
      "rawMarkdown": "I believe someone had written a method to determine images that are similar. I'm assuming that these images are exact matches but I wonder if there are any images like in the additional dataset that are of the same patient but a slightly different image(camera moved slightly between taking images).",
      "votes": null
    },
    {
      "id": "193140",
      "postDate": "06/15/2017 17:28:21",
      "content": "<p>Oh! This is very serious! Both Yau Ben-Or and Wendy Kan promised that all stage2 images will be unique in this thread <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/32363\">https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/32363</a> \nI hope organizers will find all duplicates and will exclude them from final scoring..</p>",
      "rawMarkdown": "Oh! This is very serious! Both Yau Ben-Or and Wendy Kan promised that all stage2 images will be unique in this thread https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/32363 \nI hope organizers will find all duplicates and will exclude them from final scoring..",
      "votes": null
    },
    {
      "id": "193146",
      "postDate": "06/15/2017 17:38:24",
      "content": "<p>This is quite bad to hear. One of the things making the leaderboard unreliable during the first stage of the competition was the existence of information \"leaks\". It was unclear to me whether the second stage would be \"leak free\" or not, but it seemed more in the vein of creating generalizable learners if the second stage did not contain any images of women seen in any of the previously release data. Personally, I find this against the spirit of the competition and hope that the organizers comment on this discovery.</p>",
      "rawMarkdown": "This is quite bad to hear. One of the things making the leaderboard unreliable during the first stage of the competition was the existence of information \"leaks\". It was unclear to me whether the second stage would be \"leak free\" or not, but it seemed more in the vein of creating generalizable learners if the second stage did not contain any images of women seen in any of the previously release data. Personally, I find this against the spirit of the competition and hope that the organizers comment on this discovery.",
      "votes": null
    },
    {
      "id": "193149",
      "postDate": "06/15/2017 17:46:09",
      "content": "<p>I really hope these photos will be excluded from final scoring... who knows what else we got there...</p>",
      "rawMarkdown": "I really hope these photos will be excluded from final scoring... who knows what else we got there...",
      "votes": null
    },
    {
      "id": "193154",
      "postDate": "06/15/2017 18:02:42",
      "content": "<p>Comparing the predictions in my submission file, I notice that some of the images that you've provided(specifically test_stg2/13456.jpg --&gt; test/0.jpg, test_stg2/10398.jpg --&gt; test/98.jpg, test_stg2/10844.jpg --&gt; test/327.jpg), the predictions are incredibly close but not the exact same, which would imply that the images are very similar but not the exact same, perhaps how they went unnoticed. </p>",
      "rawMarkdown": "Comparing the predictions in my submission file, I notice that some of the images that you've provided(specifically test_stg2/13456.jpg --&gt; test/0.jpg, test_stg2/10398.jpg --&gt; test/98.jpg, test_stg2/10844.jpg --&gt; test/327.jpg), the predictions are incredibly close but not the exact same, which would imply that the images are very similar but not the exact same, perhaps how they went unnoticed.",
      "votes": null
    },
    {
      "id": "193159",
      "postDate": "06/15/2017 18:18:10",
      "content": "<p>I think that a leak doesn't have to be from exact picture matches - multiple pictures from the same individual should be viewed as a leak. This would be like if AlphaGo evaluated their state evaluation model by training and testing from board states from the same game.</p>",
      "rawMarkdown": "I think that a leak doesn't have to be from exact picture matches - multiple pictures from the same individual should be viewed as a leak. This would be like if AlphaGo evaluated their state evaluation model by training and testing from board states from the same game.",
      "votes": null
    },
    {
      "id": "193161",
      "postDate": "06/15/2017 18:21:35",
      "content": "<p>I completely agree, I'm just curious as to how it went unnoticed since there was a similar issue when the competition first launched. </p>",
      "rawMarkdown": "I completely agree, I'm just curious as to how it went unnoticed since there was a similar issue when the competition first launched.",
      "votes": null
    },
    {
      "id": "193164",
      "postDate": "06/15/2017 18:28:28",
      "content": "<p>@Admins: can you please put a hold on the competition until you have sorted out this leak? The implication is that the best performances will have an image matching component that is entirely irrelevant to anything you want to achieve. I asked about this here <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/32363\">https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/32363</a> specifically because it was critical to the model we needed to upload. I would second all previous suggestions that you go through the stage 2 test set and screen everything to ensure that they are unique patients and images, as was promised.</p>",
      "rawMarkdown": "Admins: can you please put a hold on the competition until you have sorted out this leak? The implication is that the best performances will have an image matching component that is entirely irrelevant to anything you want to achieve. I asked about this here https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/32363 specifically because it was critical to the model we needed to upload. I would second all previous suggestions that you go through the stage 2 test set and screen everything to ensure that they are unique patients and images, as was promised.",
      "votes": null
    },
    {
      "id": "193169",
      "postDate": "06/15/2017 18:49:54",
      "content": "<p>Ah I see what you are saying. Right, that is surprising.</p>",
      "rawMarkdown": "Ah I see what you are saying. Right, that is surprising.",
      "votes": null
    },
    {
      "id": "193176",
      "postDate": "06/15/2017 19:04:37",
      "content": "<p>I ran a simple histogram comparison script to compare test_stg2 images with the ones from train and test_stg1. It seems to me that around 50% of the images have a leaked label. \nI'm adding a text file with lines such as follows:</p>\n\n<p>test_stg2/13295.jpg test/96.jpg 0.004</p>\n\n<p>This means that 13295.jpg from the test_stg2 has a very similar histogram (distance 0.004) to the image 96.jpg from test_stg1. There might be some false matches (especially if the distance is very close to 0.1), however, I think for the most cases it's just a slight variation of the same image.</p>",
      "rawMarkdown": "I ran a simple histogram comparison script to compare test_stg2 images with the ones from train and test_stg1. It seems to me that around 50% of the images have a leaked label. \nI'm adding a text file with lines such as follows:\n\ntest_stg2/13295.jpg test/96.jpg 0.004\n\nThis means that 13295.jpg from the test_stg2 has a very similar histogram (distance 0.004) to the image 96.jpg from test_stg1. There might be some false matches (especially if the distance is very close to 0.1), however, I think for the most cases it's just a slight variation of the same image.",
      "votes": null
    },
    {
      "id": "193186",
      "postDate": "06/15/2017 19:31:28",
      "content": "<p>Thanks for this. <br>\nI've ran a similar analysis and the results are similar - what a shame! <br>\nUpdate: <br>\nkernel for analysis: <a href=\"https://www.kaggle.com/kylehounslow/a-method-for-finding-leaked-images-in-test-set\">https://www.kaggle.com/kylehounslow/a-method-for-finding-leaked-images-in-test-set</a></p>",
      "rawMarkdown": "Thanks for this.  \nI've ran a similar analysis and the results are similar - what a shame!  \nUpdate:    \nkernel for analysis: https://www.kaggle.com/kylehounslow/a-method-for-finding-leaked-images-in-test-set",
      "votes": null
    },
    {
      "id": "193207",
      "postDate": "06/15/2017 19:39:45",
      "content": "<p>I agree that these leaks are really sad for the competition, and also probably for the quality of the final model.</p>\n\n<p>But, Scotty... How do you find such similar looking images?</p>",
      "rawMarkdown": "I agree that these leaks are really sad for the competition, and also probably for the quality of the final model.\n\nBut, Scotty... How do you find such similar looking images?",
      "votes": null
    },
    {
      "id": "193214",
      "postDate": "06/15/2017 20:00:02",
      "content": "<p>Great analysis! Disturbing results..</p>",
      "rawMarkdown": "Great analysis! Disturbing results..",
      "votes": null
    },
    {
      "id": "193233",
      "postDate": "06/15/2017 20:59:38",
      "content": "<p>Thanks for reporting. We are looking into this issue. In particular:</p>\n\n<ol>\n<li><p>We checked for duplicates, but these images slipped through since they were similar but not identical. </p></li>\n<li><p>We will set those duplicated images to be ignored in the final leaderboard scoring. </p></li>\n<li><p>We are checking with the data provider to see where the issue occurred. </p></li>\n<li><p>Please continue with the competition according to the current timeline. </p></li>\n</ol>",
      "rawMarkdown": "Thanks for reporting. We are looking into this issue. In particular:\n\n1. We checked for duplicates, but these images slipped through since they were similar but not identical. \n\n2. We will set those duplicated images to be ignored in the final leaderboard scoring. \n\n3. We are checking with the data provider to see where the issue occurred. \n\n4. Please continue with the competition according to the current timeline.",
      "votes": null
    },
    {
      "id": "193236",
      "postDate": "06/15/2017 21:03:36",
      "content": "<p>I recommend using histogram distance,  for example compute 12x12x12 RGB histogram so that each image is mapped to a  1728 dimensional vector and compare with Bhattacharyya distance. </p>",
      "rawMarkdown": "I recommend using histogram distance,  for example compute 12x12x12 RGB histogram so that each image is mapped to a  1728 dimensional vector and compare with Bhattacharyya distance.",
      "votes": null
    },
    {
      "id": "193240",
      "postDate": "06/15/2017 21:10:09",
      "content": "<p>Ideally duplicates are detected by patient ID, which the data provider should have.</p>",
      "rawMarkdown": "Ideally duplicates are detected by patient ID, which the data provider should have.",
      "votes": null
    },
    {
      "id": "193242",
      "postDate": "06/15/2017 21:21:57",
      "content": "<p>Hey a question on that. \nWhy is the prediction file expecting 4018 predictions? Where you're giving 3506 photos to predict?\nShould we do the look ups for other photos?</p>",
      "rawMarkdown": "Hey a question on that. \nWhy is the prediction file expecting 4018 predictions? Where you're giving 3506 photos to predict?\nShould we do the look ups for other photos?",
      "votes": null
    },
    {
      "id": "193246",
      "postDate": "06/15/2017 21:26:59",
      "content": "<p>The other images are simply the images from the stage 1 test set. You need to combine the two test sets, and the public leaderboard will be scored on the stage 1 test set whereas the private leaderboard will be the other 3506 images.</p>",
      "rawMarkdown": "The other images are simply the images from the stage 1 test set. You need to combine the two test sets, and the public leaderboard will be scored on the stage 1 test set whereas the private leaderboard will be the other 3506 images.",
      "votes": null
    },
    {
      "id": "193255",
      "postDate": "06/15/2017 21:57:32",
      "content": "<p>Thanks Mate.\nSo we won't know results for sure until the end of competition?</p>",
      "rawMarkdown": "Thanks Mate.\nSo we won't know results for sure until the end of competition?",
      "votes": null
    },
    {
      "id": "193265",
      "postDate": "06/15/2017 22:22:16",
      "content": "<p>@Wendy Kan I have made a Kernel for finding duplicates using histogram: <a href=\"https://www.kaggle.com/kylehounslow/a-method-for-finding-leaked-images-in-test-set/\">https://www.kaggle.com/kylehounslow/a-method-for-finding-leaked-images-in-test-set/</a></p>",
      "rawMarkdown": "Wendy Kan I have made a Kernel for finding duplicates using histogram: https://www.kaggle.com/kylehounslow/a-method-for-finding-leaked-images-in-test-set/",
      "votes": null
    },
    {
      "id": "193361",
      "postDate": "06/16/2017 06:17:03",
      "content": "<p>We need to incentivize creating useful tools to battle cancer, curating this dataset should be a top priority.</p>",
      "rawMarkdown": "We need to incentivize creating useful tools to battle cancer, curating this dataset should be a top priority.",
      "votes": null
    },
    {
      "id": "193537",
      "postDate": "06/16/2017 20:05:43",
      "content": "<p>@Wendy: Are we allowed to augment our training dataset with these leaked images, and retrain?</p>",
      "rawMarkdown": "Wendy: Are we allowed to augment our training dataset with these leaked images, and retrain?",
      "votes": null
    },
    {
      "id": "193645",
      "postDate": "06/17/2017 10:36:41",
      "content": "<p>I've found 1781 duplicates (100% manually checked) and attached here.\nUsed my algorithm for cleaning additional data:\n- used 12288 features from resnet50, vgg16 and vgg19 for 4 sectors of image (2*2), normalized with std\n- green only channel (because there are duplicates for green images)\nThen I've spend few hours to check all found pairs and removed false duplicates.\nHope this will be helpful for organizers. May be more images, but these definitly should be excluded.</p>\n\n<p>My aunt have died from Cervical Cancer during this competition, so I've put a lot of effort to this competition. And I really want True best solutions to win, not overfited ones.</p>\n\n<p>Edit: uploaded wrong file - fixed</p>",
      "rawMarkdown": "I've found 1781 duplicates (100% manually checked) and attached here.\nUsed my algorithm for cleaning additional data:\n- used 12288 features from resnet50, vgg16 and vgg19 for 4 sectors of image (2*2), normalized with std\n- green only channel (because there are duplicates for green images)\nThen I've spend few hours to check all found pairs and removed false duplicates.\nHope this will be helpful for organizers. May be more images, but these definitly should be excluded.\n\nMy aunt have died from Cervical Cancer during this competition, so I've put a lot of effort to this competition. And I really want True best solutions to win, not overfited ones.\n\nEdit: uploaded wrong file - fixed",
      "votes": null
    },
    {
      "id": "193678",
      "postDate": "06/17/2017 14:20:31",
      "content": "<p>Very sorry to hear about your aunt. Really hope this competition will be of use. @Wendy </p>",
      "rawMarkdown": "Very sorry to hear about your aunt. Really hope this competition will be of use. @Wendy",
      "votes": null
    },
    {
      "id": "193966",
      "postDate": "06/19/2017 01:01:47",
      "content": "<p>My condolences.  Thank you for your dedication. </p>",
      "rawMarkdown": "My condolences.  Thank you for your dedication.",
      "votes": null
    },
    {
      "id": "194919",
      "postDate": "06/22/2017 07:48:17",
      "content": "<p>Now the stage 2 result is out, congrats to the winning teams! </p>\n\n<p>But what is the final conclusion on which images were included / exclude in the final score of stage 2 ?  </p>\n\n<p>Is the result already excluding the leak?\nThanks!</p>",
      "rawMarkdown": "Now the stage 2 result is out, congrats to the winning teams! \n\nBut what is the final conclusion on which images were included / exclude in the final score of stage 2 ?  \n\nIs the result already excluding the leak?\nThanks!",
      "votes": null
    },
    {
      "id": "195007",
      "postDate": "06/22/2017 15:22:53",
      "content": "<p>I think they cut about half of the images from the final scoring. I am making this guess based on that the message above the leaderboard reads:</p>\n\n<blockquote>\n  <p>This leaderboard is calculated with approximately 25% of the test data.\n  The final results will be based on the other 75%, so the final standings may be different.</p>\n</blockquote>\n\n<p>But at the start of stage 2 it read 13% and 87% respectively. </p>",
      "rawMarkdown": "I think they cut about half of the images from the final scoring. I am making this guess based on that the message above the leaderboard reads:\n\n&gt; This leaderboard is calculated with approximately 25% of the test data.\n&gt; The final results will be based on the other 75%, so the final standings may be different.\n\nBut at the start of stage 2 it read 13% and 87% respectively.",
      "votes": null
    },
    {
      "id": "195298",
      "postDate": "06/23/2017 07:34:56",
      "content": "<p>Oh thanks gkericks !</p>",
      "rawMarkdown": "Oh thanks gkericks !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 193128,
      "author_name": "timothyrimnac",
      "author_url": "",
      "post_date": "06/15/2017 17:14:18",
      "content": "<p>I believe someone had written a method to determine images that are similar. I'm assuming that these images are exact matches but I wonder if there are any images like in the additional dataset that are of the same patient but a slightly different image(camera moved slightly between taking images).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 193140,
      "author_name": "victorsd",
      "author_url": "",
      "post_date": "06/15/2017 17:28:21",
      "content": "<p>Oh! This is very serious! Both Yau Ben-Or and Wendy Kan promised that all stage2 images will be unique in this thread <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/32363\">https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/32363</a> \nI hope organizers will find all duplicates and will exclude them from final scoring..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 193146,
      "author_name": "gkericks",
      "author_url": "",
      "post_date": "06/15/2017 17:38:24",
      "content": "<p>This is quite bad to hear. One of the things making the leaderboard unreliable during the first stage of the competition was the existence of information \"leaks\". It was unclear to me whether the second stage would be \"leak free\" or not, but it seemed more in the vein of creating generalizable learners if the second stage did not contain any images of women seen in any of the previously release data. Personally, I find this against the spirit of the competition and hope that the organizers comment on this discovery.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 193149,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "06/15/2017 17:46:09",
      "content": "<p>I really hope these photos will be excluded from final scoring... who knows what else we got there...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 193154,
      "author_name": "timothyrimnac",
      "author_url": "",
      "post_date": "06/15/2017 18:02:42",
      "content": "<p>Comparing the predictions in my submission file, I notice that some of the images that you've provided(specifically test_stg2/13456.jpg --&gt; test/0.jpg, test_stg2/10398.jpg --&gt; test/98.jpg, test_stg2/10844.jpg --&gt; test/327.jpg), the predictions are incredibly close but not the exact same, which would imply that the images are very similar but not the exact same, perhaps how they went unnoticed. </p>",
      "votes": null,
      "replies": [
        {
          "id": 193159,
          "author_name": "gkericks",
          "author_url": "",
          "post_date": "06/15/2017 18:18:10",
          "content": "<p>I think that a leak doesn't have to be from exact picture matches - multiple pictures from the same individual should be viewed as a leak. This would be like if AlphaGo evaluated their state evaluation model by training and testing from board states from the same game.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 193161,
          "author_name": "timothyrimnac",
          "author_url": "",
          "post_date": "06/15/2017 18:21:35",
          "content": "<p>I completely agree, I'm just curious as to how it went unnoticed since there was a similar issue when the competition first launched. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 193169,
          "author_name": "gkericks",
          "author_url": "",
          "post_date": "06/15/2017 18:49:54",
          "content": "<p>Ah I see what you are saying. Right, that is surprising.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 193164,
      "author_name": "fergusoci",
      "author_url": "",
      "post_date": "06/15/2017 18:28:28",
      "content": "<p>@Admins: can you please put a hold on the competition until you have sorted out this leak? The implication is that the best performances will have an image matching component that is entirely irrelevant to anything you want to achieve. I asked about this here <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/32363\">https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/32363</a> specifically because it was critical to the model we needed to upload. I would second all previous suggestions that you go through the stage 2 test set and screen everything to ensure that they are unique patients and images, as was promised.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 193176,
      "author_name": "bobutis",
      "author_url": "",
      "post_date": "06/15/2017 19:04:37",
      "content": "<p>I ran a simple histogram comparison script to compare test_stg2 images with the ones from train and test_stg1. It seems to me that around 50% of the images have a leaked label. \nI'm adding a text file with lines such as follows:</p>\n\n<p>test_stg2/13295.jpg test/96.jpg 0.004</p>\n\n<p>This means that 13295.jpg from the test_stg2 has a very similar histogram (distance 0.004) to the image 96.jpg from test_stg1. There might be some false matches (especially if the distance is very close to 0.1), however, I think for the most cases it's just a slight variation of the same image.</p>",
      "votes": null,
      "replies": [
        {
          "id": 193186,
          "author_name": "kylehounslow",
          "author_url": "",
          "post_date": "06/15/2017 19:31:28",
          "content": "<p>Thanks for this. <br>\nI've ran a similar analysis and the results are similar - what a shame! <br>\nUpdate: <br>\nkernel for analysis: <a href=\"https://www.kaggle.com/kylehounslow/a-method-for-finding-leaked-images-in-test-set\">https://www.kaggle.com/kylehounslow/a-method-for-finding-leaked-images-in-test-set</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 193214,
          "author_name": "timothyrimnac",
          "author_url": "",
          "post_date": "06/15/2017 20:00:02",
          "content": "<p>Great analysis! Disturbing results..</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 193207,
      "author_name": "oysteijo",
      "author_url": "",
      "post_date": "06/15/2017 19:39:45",
      "content": "<p>I agree that these leaks are really sad for the competition, and also probably for the quality of the final model.</p>\n\n<p>But, Scotty... How do you find such similar looking images?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 193233,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "06/15/2017 20:59:38",
      "content": "<p>Thanks for reporting. We are looking into this issue. In particular:</p>\n\n<ol>\n<li><p>We checked for duplicates, but these images slipped through since they were similar but not identical. </p></li>\n<li><p>We will set those duplicated images to be ignored in the final leaderboard scoring. </p></li>\n<li><p>We are checking with the data provider to see where the issue occurred. </p></li>\n<li><p>Please continue with the competition according to the current timeline. </p></li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 193236,
          "author_name": "bobutis",
          "author_url": "",
          "post_date": "06/15/2017 21:03:36",
          "content": "<p>I recommend using histogram distance,  for example compute 12x12x12 RGB histogram so that each image is mapped to a  1728 dimensional vector and compare with Bhattacharyya distance. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 193240,
          "author_name": "kubilai",
          "author_url": "",
          "post_date": "06/15/2017 21:10:09",
          "content": "<p>Ideally duplicates are detected by patient ID, which the data provider should have.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 193265,
          "author_name": "kylehounslow",
          "author_url": "",
          "post_date": "06/15/2017 22:22:16",
          "content": "<p>@Wendy Kan I have made a Kernel for finding duplicates using histogram: <a href=\"https://www.kaggle.com/kylehounslow/a-method-for-finding-leaked-images-in-test-set/\">https://www.kaggle.com/kylehounslow/a-method-for-finding-leaked-images-in-test-set/</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 193537,
          "author_name": "oysteijo",
          "author_url": "",
          "post_date": "06/16/2017 20:05:43",
          "content": "<p>@Wendy: Are we allowed to augment our training dataset with these leaked images, and retrain?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 193242,
      "author_name": "dziduchm",
      "author_url": "",
      "post_date": "06/15/2017 21:21:57",
      "content": "<p>Hey a question on that. \nWhy is the prediction file expecting 4018 predictions? Where you're giving 3506 photos to predict?\nShould we do the look ups for other photos?</p>",
      "votes": null,
      "replies": [
        {
          "id": 193246,
          "author_name": "timothyrimnac",
          "author_url": "",
          "post_date": "06/15/2017 21:26:59",
          "content": "<p>The other images are simply the images from the stage 1 test set. You need to combine the two test sets, and the public leaderboard will be scored on the stage 1 test set whereas the private leaderboard will be the other 3506 images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 193255,
          "author_name": "dziduchm",
          "author_url": "",
          "post_date": "06/15/2017 21:57:32",
          "content": "<p>Thanks Mate.\nSo we won't know results for sure until the end of competition?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 193361,
      "author_name": "tanstaafl",
      "author_url": "",
      "post_date": "06/16/2017 06:17:03",
      "content": "<p>We need to incentivize creating useful tools to battle cancer, curating this dataset should be a top priority.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 193645,
      "author_name": "victorsd",
      "author_url": "",
      "post_date": "06/17/2017 10:36:41",
      "content": "<p>I've found 1781 duplicates (100% manually checked) and attached here.\nUsed my algorithm for cleaning additional data:\n- used 12288 features from resnet50, vgg16 and vgg19 for 4 sectors of image (2*2), normalized with std\n- green only channel (because there are duplicates for green images)\nThen I've spend few hours to check all found pairs and removed false duplicates.\nHope this will be helpful for organizers. May be more images, but these definitly should be excluded.</p>\n\n<p>My aunt have died from Cervical Cancer during this competition, so I've put a lot of effort to this competition. And I really want True best solutions to win, not overfited ones.</p>\n\n<p>Edit: uploaded wrong file - fixed</p>",
      "votes": null,
      "replies": [
        {
          "id": 193678,
          "author_name": "bvineeth007",
          "author_url": "",
          "post_date": "06/17/2017 14:20:31",
          "content": "<p>Very sorry to hear about your aunt. Really hope this competition will be of use. @Wendy </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 193966,
          "author_name": "kubilai",
          "author_url": "",
          "post_date": "06/19/2017 01:01:47",
          "content": "<p>My condolences.  Thank you for your dedication. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 194919,
      "author_name": "scottykwok",
      "author_url": "",
      "post_date": "06/22/2017 07:48:17",
      "content": "<p>Now the stage 2 result is out, congrats to the winning teams! </p>\n\n<p>But what is the final conclusion on which images were included / exclude in the final score of stage 2 ?  </p>\n\n<p>Is the result already excluding the leak?\nThanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 195007,
          "author_name": "gkericks",
          "author_url": "",
          "post_date": "06/22/2017 15:22:53",
          "content": "<p>I think they cut about half of the images from the final scoring. I am making this guess based on that the message above the leaderboard reads:</p>\n\n<blockquote>\n  <p>This leaderboard is calculated with approximately 25% of the test data.\n  The final results will be based on the other 75%, so the final standings may be different.</p>\n</blockquote>\n\n<p>But at the start of stage 2 it read 13% and 87% respectively. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 195298,
          "author_name": "scottykwok",
          "author_url": "",
          "post_date": "06/23/2017 07:34:56",
          "content": "<p>Oh thanks gkericks !</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "193117": "It seems quite many Stage 2 \"unseen data\" already appeared in Stage 1 train/test set.... i.e. leaks\nFor examples :\n\n - test_stg2/10698.jpg  --&gt; train/Type_2/485.jpg \n - test_stg2/11249.jpg   --&gt; train/Type_2/485.jpg \n - test_stg2/13364.jpg --&gt; train/Type_1/27.jpg\n - test_stg2/10898.jpg --&gt; train/Type_1/265.jpg\t\n - test_stg2/13456.jpg --&gt; test/0.jpg\n - test_stg2/10398.jpg --&gt; test/98.jpg\n - test_stg2/10844.jpg --&gt; test/327.jpg\n - test_stg2/11334.jpg --&gt; test/1068.jpg\n - etc... I am pretty sure there are more we can find out\n\nIts a bit disappointing , because that means models that overfit to **REMEMBER** stage 1 data could out-perform those models that are actually **PREDICTING**..... which defect the purpose of this challenge.\n\nWould the organizer/admin consider excluding those leaks in Stage 2 scoring  ?",
    "193128": "I believe someone had written a method to determine images that are similar. I'm assuming that these images are exact matches but I wonder if there are any images like in the additional dataset that are of the same patient but a slightly different image(camera moved slightly between taking images).",
    "193140": "Oh! This is very serious! Both Yau Ben-Or and Wendy Kan promised that all stage2 images will be unique in this thread https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/32363 \nI hope organizers will find all duplicates and will exclude them from final scoring..",
    "193146": "This is quite bad to hear. One of the things making the leaderboard unreliable during the first stage of the competition was the existence of information \"leaks\". It was unclear to me whether the second stage would be \"leak free\" or not, but it seemed more in the vein of creating generalizable learners if the second stage did not contain any images of women seen in any of the previously release data. Personally, I find this against the spirit of the competition and hope that the organizers comment on this discovery.",
    "193149": "I really hope these photos will be excluded from final scoring... who knows what else we got there...",
    "193154": "Comparing the predictions in my submission file, I notice that some of the images that you've provided(specifically test_stg2/13456.jpg --&gt; test/0.jpg, test_stg2/10398.jpg --&gt; test/98.jpg, test_stg2/10844.jpg --&gt; test/327.jpg), the predictions are incredibly close but not the exact same, which would imply that the images are very similar but not the exact same, perhaps how they went unnoticed.",
    "193159": "I think that a leak doesn't have to be from exact picture matches - multiple pictures from the same individual should be viewed as a leak. This would be like if AlphaGo evaluated their state evaluation model by training and testing from board states from the same game.",
    "193161": "I completely agree, I'm just curious as to how it went unnoticed since there was a similar issue when the competition first launched.",
    "193164": "Admins: can you please put a hold on the competition until you have sorted out this leak? The implication is that the best performances will have an image matching component that is entirely irrelevant to anything you want to achieve. I asked about this here https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/32363 specifically because it was critical to the model we needed to upload. I would second all previous suggestions that you go through the stage 2 test set and screen everything to ensure that they are unique patients and images, as was promised.",
    "193169": "Ah I see what you are saying. Right, that is surprising.",
    "193176": "I ran a simple histogram comparison script to compare test_stg2 images with the ones from train and test_stg1. It seems to me that around 50% of the images have a leaked label. \nI'm adding a text file with lines such as follows:\n\ntest_stg2/13295.jpg test/96.jpg 0.004\n\nThis means that 13295.jpg from the test_stg2 has a very similar histogram (distance 0.004) to the image 96.jpg from test_stg1. There might be some false matches (especially if the distance is very close to 0.1), however, I think for the most cases it's just a slight variation of the same image.",
    "193186": "Thanks for this.  \nI've ran a similar analysis and the results are similar - what a shame!  \nUpdate:    \nkernel for analysis: https://www.kaggle.com/kylehounslow/a-method-for-finding-leaked-images-in-test-set",
    "193207": "I agree that these leaks are really sad for the competition, and also probably for the quality of the final model.\n\nBut, Scotty... How do you find such similar looking images?",
    "193214": "Great analysis! Disturbing results..",
    "193233": "Thanks for reporting. We are looking into this issue. In particular:\n\n1. We checked for duplicates, but these images slipped through since they were similar but not identical. \n\n2. We will set those duplicated images to be ignored in the final leaderboard scoring. \n\n3. We are checking with the data provider to see where the issue occurred. \n\n4. Please continue with the competition according to the current timeline.",
    "193236": "I recommend using histogram distance,  for example compute 12x12x12 RGB histogram so that each image is mapped to a  1728 dimensional vector and compare with Bhattacharyya distance.",
    "193240": "Ideally duplicates are detected by patient ID, which the data provider should have.",
    "193242": "Hey a question on that. \nWhy is the prediction file expecting 4018 predictions? Where you're giving 3506 photos to predict?\nShould we do the look ups for other photos?",
    "193246": "The other images are simply the images from the stage 1 test set. You need to combine the two test sets, and the public leaderboard will be scored on the stage 1 test set whereas the private leaderboard will be the other 3506 images.",
    "193255": "Thanks Mate.\nSo we won't know results for sure until the end of competition?",
    "193265": "Wendy Kan I have made a Kernel for finding duplicates using histogram: https://www.kaggle.com/kylehounslow/a-method-for-finding-leaked-images-in-test-set/",
    "193361": "We need to incentivize creating useful tools to battle cancer, curating this dataset should be a top priority.",
    "193537": "Wendy: Are we allowed to augment our training dataset with these leaked images, and retrain?",
    "193645": "I've found 1781 duplicates (100% manually checked) and attached here.\nUsed my algorithm for cleaning additional data:\n- used 12288 features from resnet50, vgg16 and vgg19 for 4 sectors of image (2*2), normalized with std\n- green only channel (because there are duplicates for green images)\nThen I've spend few hours to check all found pairs and removed false duplicates.\nHope this will be helpful for organizers. May be more images, but these definitly should be excluded.\n\nMy aunt have died from Cervical Cancer during this competition, so I've put a lot of effort to this competition. And I really want True best solutions to win, not overfited ones.\n\nEdit: uploaded wrong file - fixed",
    "193678": "Very sorry to hear about your aunt. Really hope this competition will be of use. @Wendy",
    "193966": "My condolences.  Thank you for your dedication.",
    "194919": "Now the stage 2 result is out, congrats to the winning teams! \n\nBut what is the final conclusion on which images were included / exclude in the final score of stage 2 ?  \n\nIs the result already excluding the leak?\nThanks!",
    "195007": "I think they cut about half of the images from the final scoring. I am making this guess based on that the message above the leaderboard reads:\n\n&gt; This leaderboard is calculated with approximately 25% of the test data.\n&gt; The final results will be based on the other 75%, so the final standings may be different.\n\nBut at the start of stage 2 it read 13% and 87% respectively.",
    "195298": "Oh thanks gkericks !"
  },
  "source": "meta"
}