{
  "id": 100489,
  "title": "Removing duplicates",
  "url": "/competitions/aptos2019-blindness-detection/discussion/100489",
  "author_name": "",
  "post_date": "2019-07-18T20:55:10.939197400Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Does anybody have a snippet for removing duplicate images (other than one in each set of duplicates) from their data?</p>\n\n<p>Should I remove duplicates in training but leave them for testing? What are others doing to clean up this dataset?</p>",
  "messages": [
    {
      "id": "579452",
      "postDate": "07/18/2019 20:55:10",
      "content": "<p>Does anybody have a snippet for removing duplicate images (other than one in each set of duplicates) from their data?</p>\n\n<p>Should I remove duplicates in training but leave them for testing? What are others doing to clean up this dataset?</p>",
      "rawMarkdown": "Does anybody have a snippet for removing duplicate images (other than one in each set of duplicates) from their data?\n\nShould I remove duplicates in training but leave them for testing? What are others doing to clean up this dataset?",
      "votes": null
    },
    {
      "id": "580371",
      "postDate": "07/20/2019 04:03:06",
      "content": "<p>To me it seems that you would want to remove the duplicates and so I have tried it but unfortunately my results have been inconclusive - at least where I am (I am a beginner) the results with and without duplicates are not that much different. </p>",
      "rawMarkdown": "To me it seems that you would want to remove the duplicates and so I have tried it but unfortunately my results have been inconclusive - at least where I am (I am a beginner) the results with and without duplicates are not that much different.",
      "votes": null
    },
    {
      "id": "580456",
      "postDate": "07/20/2019 07:20:10",
      "content": "<p>On the one hand, there are not that many duplicates (262), especially if we include the big dataset from the previous competition.\nOn the other hand, the duplicates definitely contain information about uncertainty (that is part of the domain here, as often enough the medical specialists don't agree on the same image) and certainty, so in my imagination they should help a deep learning model (we're not working with SVMs any longer).</p>\n\n<p>If we look to the previous competition, we even have intentional kind of duplicates with right/left eye. Of course these are different images, but the diagnosis for left/right eye and the structural failures we can see in the eyes should be highly correlated. I think (without having already a working for it), there's more potential in analyzing and using duplicates we find instead of just removing them.</p>",
      "rawMarkdown": "On the one hand, there are not that many duplicates (262), especially if we include the big dataset from the previous competition.\nOn the other hand, the duplicates definitely contain information about uncertainty (that is part of the domain here, as often enough the medical specialists don't agree on the same image) and certainty, so in my imagination they should help a deep learning model (we're not working with SVMs any longer).\n\nIf we look to the previous competition, we even have intentional kind of duplicates with right/left eye. Of course these are different images, but the diagnosis for left/right eye and the structural failures we can see in the eyes should be highly correlated. I think (without having already a working for it), there's more potential in analyzing and using duplicates we find instead of just removing them.",
      "votes": null
    },
    {
      "id": "580523",
      "postDate": "07/20/2019 09:35:30",
      "content": "<p>You make a good point about uncertainty that I have considered as well. Especially after data augmentation, I am thinking now that there is probably no benefit to removing the duplicate images.</p>",
      "rawMarkdown": "You make a good point about uncertainty that I have considered as well. Especially after data augmentation, I am thinking now that there is probably no benefit to removing the duplicate images.",
      "votes": null
    },
    {
      "id": "611882",
      "postDate": "08/29/2019 13:38:46",
      "content": "<p>this is a great kernel to remove duplicates: \n<a href=\"https://www.kaggle.com/manojprabhaakr/similar-duplicate-images-in-aptos-data/notebook\">https://www.kaggle.com/manojprabhaakr/similar-duplicate-images-in-aptos-data/notebook</a>\nthanks <a href=\"/manojprabhaakr\">@manojprabhaakr</a> </p>",
      "rawMarkdown": "this is a great kernel to remove duplicates: \nhttps://www.kaggle.com/manojprabhaakr/similar-duplicate-images-in-aptos-data/notebook\nthanks @manojprabhaakr",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 580371,
      "author_name": "edchepen",
      "author_url": "",
      "post_date": "07/20/2019 04:03:06",
      "content": "<p>To me it seems that you would want to remove the duplicates and so I have tried it but unfortunately my results have been inconclusive - at least where I am (I am a beginner) the results with and without duplicates are not that much different. </p>",
      "votes": null,
      "replies": [
        {
          "id": 580456,
          "author_name": "hanfried",
          "author_url": "",
          "post_date": "07/20/2019 07:20:10",
          "content": "<p>On the one hand, there are not that many duplicates (262), especially if we include the big dataset from the previous competition.\nOn the other hand, the duplicates definitely contain information about uncertainty (that is part of the domain here, as often enough the medical specialists don't agree on the same image) and certainty, so in my imagination they should help a deep learning model (we're not working with SVMs any longer).</p>\n\n<p>If we look to the previous competition, we even have intentional kind of duplicates with right/left eye. Of course these are different images, but the diagnosis for left/right eye and the structural failures we can see in the eyes should be highly correlated. I think (without having already a working for it), there's more potential in analyzing and using duplicates we find instead of just removing them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 580523,
          "author_name": "nwickman",
          "author_url": "",
          "post_date": "07/20/2019 09:35:30",
          "content": "<p>You make a good point about uncertainty that I have considered as well. Especially after data augmentation, I am thinking now that there is probably no benefit to removing the duplicate images.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 611882,
      "author_name": "anuragtr",
      "author_url": "",
      "post_date": "08/29/2019 13:38:46",
      "content": "<p>this is a great kernel to remove duplicates: \n<a href=\"https://www.kaggle.com/manojprabhaakr/similar-duplicate-images-in-aptos-data/notebook\">https://www.kaggle.com/manojprabhaakr/similar-duplicate-images-in-aptos-data/notebook</a>\nthanks <a href=\"/manojprabhaakr\">@manojprabhaakr</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "579452": "Does anybody have a snippet for removing duplicate images (other than one in each set of duplicates) from their data?\n\nShould I remove duplicates in training but leave them for testing? What are others doing to clean up this dataset?",
    "580371": "To me it seems that you would want to remove the duplicates and so I have tried it but unfortunately my results have been inconclusive - at least where I am (I am a beginner) the results with and without duplicates are not that much different.",
    "580456": "On the one hand, there are not that many duplicates (262), especially if we include the big dataset from the previous competition.\nOn the other hand, the duplicates definitely contain information about uncertainty (that is part of the domain here, as often enough the medical specialists don't agree on the same image) and certainty, so in my imagination they should help a deep learning model (we're not working with SVMs any longer).\n\nIf we look to the previous competition, we even have intentional kind of duplicates with right/left eye. Of course these are different images, but the diagnosis for left/right eye and the structural failures we can see in the eyes should be highly correlated. I think (without having already a working for it), there's more potential in analyzing and using duplicates we find instead of just removing them.",
    "580523": "You make a good point about uncertainty that I have considered as well. Especially after data augmentation, I am thinking now that there is probably no benefit to removing the duplicate images.",
    "611882": "this is a great kernel to remove duplicates: \nhttps://www.kaggle.com/manojprabhaakr/similar-duplicate-images-in-aptos-data/notebook\nthanks @manojprabhaakr"
  },
  "source": "meta"
}