{
  "id": 230853,
  "title": "Discovered 270 more pairs of duplicates and some BAD images.",
  "url": "/competitions/plant-pathology-2021-fgvc8/discussion/230853",
  "author_name": "Anatoly Pavlov",
  "post_date": "2021-04-05T21:45:28.851000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi There,</p>\n<p>There were 50 duplicate pairs found before in this <a href=\"https://www.kaggle.com/nickuzmenkov/pp2021-duplicates-revealing\" target=\"_blank\">notebook</a> by <a href=\"https://www.kaggle.com/nickuzmenkov\" target=\"_blank\">@nickuzmenkov</a> . Here, I am sharing my find of at least 270 more duplicate pairs via image embedding and similarity method which I used and described in all details in this <a href=\"https://www.kaggle.com/datasciencegeek/detect-duplicates-via-image-embeddings\" target=\"_blank\">notebook</a>.</p>\n<p>The duplicates I found can be split into two categories:</p>\n<p>1) Exact duplicates.</p>\n<p>2) Images of the same leaf but taken differently, either at different angles, leaf location within image, or other conditions that vary.</p>\n<p>The second class has rather various duplicate images with respect to how much they close or differ from one another. I have not decided yet how exactly I am going to deal with them when it comes to model development.</p>\n<p>I also discovered some bad images, just 4 of them, which better to be removed from model training. In the output of my notebook mentioned above you can find two CSV files one with all duplicates found so far and in another file id's of bad images.</p>",
  "messages": [
    {
      "id": 1264094,
      "postDate": "2021-04-05T21:45:28.850Z",
      "content": "<p>Hi There,</p>\n<p>There were 50 duplicate pairs found before in this <a href=\"https://www.kaggle.com/nickuzmenkov/pp2021-duplicates-revealing\" target=\"_blank\">notebook</a> by <a href=\"https://www.kaggle.com/nickuzmenkov\" target=\"_blank\">@nickuzmenkov</a> . Here, I am sharing my find of at least 270 more duplicate pairs via image embedding and similarity method which I used and described in all details in this <a href=\"https://www.kaggle.com/datasciencegeek/detect-duplicates-via-image-embeddings\" target=\"_blank\">notebook</a>.</p>\n<p>The duplicates I found can be split into two categories:</p>\n<p>1) Exact duplicates.</p>\n<p>2) Images of the same leaf but taken differently, either at different angles, leaf location within image, or other conditions that vary.</p>\n<p>The second class has rather various duplicate images with respect to how much they close or differ from one another. I have not decided yet how exactly I am going to deal with them when it comes to model development.</p>\n<p>I also discovered some bad images, just 4 of them, which better to be removed from model training. In the output of my notebook mentioned above you can find two CSV files one with all duplicates found so far and in another file id's of bad images.</p>",
      "rawMarkdown": "Hi There,\n\nThere were 50 duplicate pairs found before in this [notebook](https://www.kaggle.com/nickuzmenkov/pp2021-duplicates-revealing) by @nickuzmenkov . Here, I am sharing my find of at least 270 more duplicate pairs via image embedding and similarity method which I used and described in all details in this [notebook](https://www.kaggle.com/datasciencegeek/detect-duplicates-via-image-embeddings).\n\nThe duplicates I found can be split into two categories:\n\n1) Exact duplicates.\n\n2) Images of the same leaf but taken differently, either at different angles, leaf location within image, or other conditions that vary.\n\nThe second class has rather various duplicate images with respect to how much they close or differ from one another. I have not decided yet how exactly I am going to deal with them when it comes to model development.\n\nI also discovered some bad images, just 4 of them, which better to be removed from model training. In the output of my notebook mentioned above you can find two CSV files one with all duplicates found so far and in another file id's of bad images.",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1264094": "Hi There,\n\nThere were 50 duplicate pairs found before in this [notebook](https://www.kaggle.com/nickuzmenkov/pp2021-duplicates-revealing) by @nickuzmenkov . Here, I am sharing my find of at least 270 more duplicate pairs via image embedding and similarity method which I used and described in all details in this [notebook](https://www.kaggle.com/datasciencegeek/detect-duplicates-via-image-embeddings).\n\nThe duplicates I found can be split into two categories:\n\n1) Exact duplicates.\n\n2) Images of the same leaf but taken differently, either at different angles, leaf location within image, or other conditions that vary.\n\nThe second class has rather various duplicate images with respect to how much they close or differ from one another. I have not decided yet how exactly I am going to deal with them when it comes to model development.\n\nI also discovered some bad images, just 4 of them, which better to be removed from model training. In the output of my notebook mentioned above you can find two CSV files one with all duplicates found so far and in another file id's of bad images."
  }
}