{
  "id": 32427,
  "title": "Test images in additional data",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/32427",
  "author_name": "",
  "post_date": "2017-05-02T22:01:38.966233500Z",
  "votes": 24,
  "comment_count": 10,
  "views": 0,
  "content": "<p>For those who use additional images for training, be aware of the fact that there are indeed a lot of almost duplicates from the test set (cf the thread - MobileODT Webinar for Kaggle). \nHence, I guess we (those using those images) are overfitting badly.</p>\n\n<p>I've made a simple detection of almost duplicates, with the following process:</p>\n\n<ul>\n<li><p>compare the average absolute differences of features extracted by VGG16, on 448x448 cropped images</p></li>\n<li><p>for those under a threhsold from the first step, compare raw pixel differences - average of absolute differences</p></li>\n</ul>\n\n<p>Here is the list that I get - you may want to remove those images from the training set to get a more realistic LB score. \nIt's probably not exhaustive since my method is highly heuristic now. </p>",
  "messages": [
    {
      "id": "179818",
      "postDate": "05/02/2017 22:01:38",
      "content": "<p>For those who use additional images for training, be aware of the fact that there are indeed a lot of almost duplicates from the test set (cf the thread - MobileODT Webinar for Kaggle). \nHence, I guess we (those using those images) are overfitting badly.</p>\n\n<p>I've made a simple detection of almost duplicates, with the following process:</p>\n\n<ul>\n<li><p>compare the average absolute differences of features extracted by VGG16, on 448x448 cropped images</p></li>\n<li><p>for those under a threhsold from the first step, compare raw pixel differences - average of absolute differences</p></li>\n</ul>\n\n<p>Here is the list that I get - you may want to remove those images from the training set to get a more realistic LB score. \nIt's probably not exhaustive since my method is highly heuristic now. </p>",
      "rawMarkdown": "For those who use additional images for training, be aware of the fact that there are indeed a lot of almost duplicates from the test set (cf the thread - MobileODT Webinar for Kaggle). \nHence, I guess we (those using those images) are overfitting badly.\n\nI've made a simple detection of almost duplicates, with the following process:\n\n- compare the average absolute differences of features extracted by VGG16, on 448x448 cropped images\n\n- for those under a threhsold from the first step, compare raw pixel differences - average of absolute differences\n\n\nHere is the list that I get - you may want to remove those images from the training set to get a more realistic LB score. \nIt's probably not exhaustive since my method is highly heuristic now.",
      "votes": null
    },
    {
      "id": "180503",
      "postDate": "05/05/2017 17:48:26",
      "content": "<p>Thanks @Alchemist for the update!!!</p>",
      "rawMarkdown": "Thanks @Alchemist for the update!!!",
      "votes": null
    },
    {
      "id": "180518",
      "postDate": "05/05/2017 19:06:15",
      "content": "<p>Thanks @Alchemist. I found some more using the structural similarity index and appended them to your list. Unfortunately, I didn't record which test items they matched (so the first column in empty) but they definitely do!</p>",
      "rawMarkdown": "Thanks @Alchemist. I found some more using the structural similarity index and appended them to your list. Unfortunately, I didn't record which test items they matched (so the first column in empty) but they definitely do!",
      "votes": null
    },
    {
      "id": "180561",
      "postDate": "05/05/2017 22:10:29",
      "content": "<p>I've had a hard time keeping track of all the dataset updates in this competition, but are you sure that additional/Type_3/788.jpg is a valid image? I thought I had the most recent version.</p>\n\n<p>I created a quick script to eliminate this duplicate and it failed.</p>\n\n<p>fixedlabelsv2 says that 788 was originally classified as type 2 but need to be moved to type 1, which is what my dataset reflects. When I compare type1/788.jpg to 209.jpg, as suggested by your output, they are not the same.</p>",
      "rawMarkdown": "I've had a hard time keeping track of all the dataset updates in this competition, but are you sure that additional/Type_3/788.jpg is a valid image? I thought I had the most recent version.\n\n I created a quick script to eliminate this duplicate and it failed.\n\nfixedlabelsv2 says that 788 was originally classified as type 2 but need to be moved to type 1, which is what my dataset reflects. When I compare type1/788.jpg to 209.jpg, as suggested by your output, they are not the same.",
      "votes": null
    },
    {
      "id": "180562",
      "postDate": "05/05/2017 22:24:03",
      "content": "<p>I've downloaded the data on 1/4/2017 - I'm not aware of any changes after that date. </p>",
      "rawMarkdown": "I've downloaded the data on 1/4/2017 - I'm not aware of any changes after that date.",
      "votes": null
    },
    {
      "id": "180564",
      "postDate": "05/05/2017 22:37:04",
      "content": "<p>Thanks for pointing this out @MitchellWagner. I don't know about @Alchemist but I was actually using the old data!</p>",
      "rawMarkdown": "Thanks for pointing this out @MitchellWagner. I don't know about @Alchemist but I was actually using the old data!",
      "votes": null
    },
    {
      "id": "180574",
      "postDate": "05/06/2017 00:15:25",
      "content": "<p>I downloaded April 10, and after extracting the files from the Type_3 additional data again, I can confirm that additional/Type_3/788.jpg is not among the files.</p>\n\n<p>Either way, thank you for calling attention to this! It's helpful to realize that this is indeed a problem, and the rest of the images I spot-checked were indeed duplicates.</p>",
      "rawMarkdown": "I downloaded April 10, and after extracting the files from the Type_3 additional data again, I can confirm that additional/Type_3/788.jpg is not among the files.\n\nEither way, thank you for calling attention to this! It's helpful to realize that this is indeed a problem, and the rest of the images I spot-checked were indeed duplicates.",
      "votes": null
    },
    {
      "id": "180649",
      "postDate": "05/06/2017 11:15:56",
      "content": "<p>I've double checked and found the error - I've replaced manually the paths before uploading the file, since in my code I merge train / additional (additional being with a prefix \"a_\"). </p>\n\n<p>It turns out I use the most recent data, however there were 2 occurencies where the algo spotted duplicates between test and train (which are not real duplicates btw). Those lines:\ntest/209.jpg,additional/Type_3/788.jpg\ntest/209.jpg,additional/Type_3/118.jpg</p>\n\n<p>were actually:\ntest/209.jpg,train/Type_3/788.jpg\ntest/209.jpg,train/Type_3/118.jpg</p>\n\n<p>I've uploaded the corrected file - apologies for the confusion and thanks for spotting this !</p>",
      "rawMarkdown": "I've double checked and found the error - I've replaced manually the paths before uploading the file, since in my code I merge train / additional (additional being with a prefix \"a_\"). \n\nIt turns out I use the most recent data, however there were 2 occurencies where the algo spotted duplicates between test and train (which are not real duplicates btw). Those lines:\ntest/209.jpg,additional/Type_3/788.jpg\ntest/209.jpg,additional/Type_3/118.jpg\n\nwere actually:\ntest/209.jpg,train/Type_3/788.jpg\ntest/209.jpg,train/Type_3/118.jpg\n\nI've uploaded the corrected file - apologies for the confusion and thanks for spotting this !",
      "votes": null
    },
    {
      "id": "182516",
      "postDate": "05/14/2017 09:59:56",
      "content": "<p>Hi,</p>\n\n<p>indeed some of your files looks very good ! </p>\n\n<p>I've uploaded a v3 with some improvements:</p>\n\n<ul>\n<li><p>while calculating the image similarity I try to match smaller images with sliding them across the other one. Many images are actually duplicates, but with a small shift - this helps identify them (although it doesn't identify well the ones that are slightly rotated)</p></li>\n<li><p>the VGG16 feature difference actually is much poorer predictor of similarity than just the raw pixel diffs (once the sliding procedure applied) - I only take the images based on the raw pixel diffs now</p></li>\n</ul>\n\n<p>There are some false positives now, but I believe it's much more complete now.</p>",
      "rawMarkdown": "Hi,\n\nindeed some of your files looks very good ! \n\nI've uploaded a v3 with some improvements:\n\n- while calculating the image similarity I try to match smaller images with sliding them across the other one. Many images are actually duplicates, but with a small shift - this helps identify them (although it doesn't identify well the ones that are slightly rotated)\n\n- the VGG16 feature difference actually is much poorer predictor of similarity than just the raw pixel diffs (once the sliding procedure applied) - I only take the images based on the raw pixel diffs now\n\n\nThere are some false positives now, but I believe it's much more complete now.",
      "votes": null
    },
    {
      "id": "182532",
      "postDate": "05/14/2017 12:12:56",
      "content": "<p>Awesome.</p>",
      "rawMarkdown": "Awesome.",
      "votes": null
    },
    {
      "id": "183646",
      "postDate": "05/18/2017 20:25:02",
      "content": "<p>Some of those additional images are really bad.  Check out 4065.jpg from the additional Type_1 data. It's part of the Motorola label.</p>",
      "rawMarkdown": "Some of those additional images are really bad.  Check out 4065.jpg from the additional Type_1 data. It's part of the Motorola label.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 180503,
      "author_name": "yadavsarthak",
      "author_url": "",
      "post_date": "05/05/2017 17:48:26",
      "content": "<p>Thanks @Alchemist for the update!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 180518,
      "author_name": "bk0000",
      "author_url": "",
      "post_date": "05/05/2017 19:06:15",
      "content": "<p>Thanks @Alchemist. I found some more using the structural similarity index and appended them to your list. Unfortunately, I didn't record which test items they matched (so the first column in empty) but they definitely do!</p>",
      "votes": null,
      "replies": [
        {
          "id": 182516,
          "author_name": "alchemist",
          "author_url": "",
          "post_date": "05/14/2017 09:59:56",
          "content": "<p>Hi,</p>\n\n<p>indeed some of your files looks very good ! </p>\n\n<p>I've uploaded a v3 with some improvements:</p>\n\n<ul>\n<li><p>while calculating the image similarity I try to match smaller images with sliding them across the other one. Many images are actually duplicates, but with a small shift - this helps identify them (although it doesn't identify well the ones that are slightly rotated)</p></li>\n<li><p>the VGG16 feature difference actually is much poorer predictor of similarity than just the raw pixel diffs (once the sliding procedure applied) - I only take the images based on the raw pixel diffs now</p></li>\n</ul>\n\n<p>There are some false positives now, but I believe it's much more complete now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 182532,
          "author_name": "bk0000",
          "author_url": "",
          "post_date": "05/14/2017 12:12:56",
          "content": "<p>Awesome.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 180561,
      "author_name": "mitchwagner",
      "author_url": "",
      "post_date": "05/05/2017 22:10:29",
      "content": "<p>I've had a hard time keeping track of all the dataset updates in this competition, but are you sure that additional/Type_3/788.jpg is a valid image? I thought I had the most recent version.</p>\n\n<p>I created a quick script to eliminate this duplicate and it failed.</p>\n\n<p>fixedlabelsv2 says that 788 was originally classified as type 2 but need to be moved to type 1, which is what my dataset reflects. When I compare type1/788.jpg to 209.jpg, as suggested by your output, they are not the same.</p>",
      "votes": null,
      "replies": [
        {
          "id": 180562,
          "author_name": "alchemist",
          "author_url": "",
          "post_date": "05/05/2017 22:24:03",
          "content": "<p>I've downloaded the data on 1/4/2017 - I'm not aware of any changes after that date. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 180564,
          "author_name": "bk0000",
          "author_url": "",
          "post_date": "05/05/2017 22:37:04",
          "content": "<p>Thanks for pointing this out @MitchellWagner. I don't know about @Alchemist but I was actually using the old data!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 180574,
          "author_name": "mitchwagner",
          "author_url": "",
          "post_date": "05/06/2017 00:15:25",
          "content": "<p>I downloaded April 10, and after extracting the files from the Type_3 additional data again, I can confirm that additional/Type_3/788.jpg is not among the files.</p>\n\n<p>Either way, thank you for calling attention to this! It's helpful to realize that this is indeed a problem, and the rest of the images I spot-checked were indeed duplicates.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 180649,
          "author_name": "alchemist",
          "author_url": "",
          "post_date": "05/06/2017 11:15:56",
          "content": "<p>I've double checked and found the error - I've replaced manually the paths before uploading the file, since in my code I merge train / additional (additional being with a prefix \"a_\"). </p>\n\n<p>It turns out I use the most recent data, however there were 2 occurencies where the algo spotted duplicates between test and train (which are not real duplicates btw). Those lines:\ntest/209.jpg,additional/Type_3/788.jpg\ntest/209.jpg,additional/Type_3/118.jpg</p>\n\n<p>were actually:\ntest/209.jpg,train/Type_3/788.jpg\ntest/209.jpg,train/Type_3/118.jpg</p>\n\n<p>I've uploaded the corrected file - apologies for the confusion and thanks for spotting this !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 183646,
      "author_name": "ekkus93",
      "author_url": "",
      "post_date": "05/18/2017 20:25:02",
      "content": "<p>Some of those additional images are really bad.  Check out 4065.jpg from the additional Type_1 data. It's part of the Motorola label.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "179818": "For those who use additional images for training, be aware of the fact that there are indeed a lot of almost duplicates from the test set (cf the thread - MobileODT Webinar for Kaggle). \nHence, I guess we (those using those images) are overfitting badly.\n\nI've made a simple detection of almost duplicates, with the following process:\n\n- compare the average absolute differences of features extracted by VGG16, on 448x448 cropped images\n\n- for those under a threhsold from the first step, compare raw pixel differences - average of absolute differences\n\n\nHere is the list that I get - you may want to remove those images from the training set to get a more realistic LB score. \nIt's probably not exhaustive since my method is highly heuristic now.",
    "180503": "Thanks @Alchemist for the update!!!",
    "180518": "Thanks @Alchemist. I found some more using the structural similarity index and appended them to your list. Unfortunately, I didn't record which test items they matched (so the first column in empty) but they definitely do!",
    "180561": "I've had a hard time keeping track of all the dataset updates in this competition, but are you sure that additional/Type_3/788.jpg is a valid image? I thought I had the most recent version.\n\n I created a quick script to eliminate this duplicate and it failed.\n\nfixedlabelsv2 says that 788 was originally classified as type 2 but need to be moved to type 1, which is what my dataset reflects. When I compare type1/788.jpg to 209.jpg, as suggested by your output, they are not the same.",
    "180562": "I've downloaded the data on 1/4/2017 - I'm not aware of any changes after that date.",
    "180564": "Thanks for pointing this out @MitchellWagner. I don't know about @Alchemist but I was actually using the old data!",
    "180574": "I downloaded April 10, and after extracting the files from the Type_3 additional data again, I can confirm that additional/Type_3/788.jpg is not among the files.\n\nEither way, thank you for calling attention to this! It's helpful to realize that this is indeed a problem, and the rest of the images I spot-checked were indeed duplicates.",
    "180649": "I've double checked and found the error - I've replaced manually the paths before uploading the file, since in my code I merge train / additional (additional being with a prefix \"a_\"). \n\nIt turns out I use the most recent data, however there were 2 occurencies where the algo spotted duplicates between test and train (which are not real duplicates btw). Those lines:\ntest/209.jpg,additional/Type_3/788.jpg\ntest/209.jpg,additional/Type_3/118.jpg\n\nwere actually:\ntest/209.jpg,train/Type_3/788.jpg\ntest/209.jpg,train/Type_3/118.jpg\n\nI've uploaded the corrected file - apologies for the confusion and thanks for spotting this !",
    "182516": "Hi,\n\nindeed some of your files looks very good ! \n\nI've uploaded a v3 with some improvements:\n\n- while calculating the image similarity I try to match smaller images with sliding them across the other one. Many images are actually duplicates, but with a small shift - this helps identify them (although it doesn't identify well the ones that are slightly rotated)\n\n- the VGG16 feature difference actually is much poorer predictor of similarity than just the raw pixel diffs (once the sliding procedure applied) - I only take the images based on the raw pixel diffs now\n\n\nThere are some false positives now, but I believe it's much more complete now.",
    "182532": "Awesome.",
    "183646": "Some of those additional images are really bad.  Check out 4065.jpg from the additional Type_1 data. It's part of the Motorola label."
  },
  "source": "meta"
}