{
  "id": 210723,
  "title": "Searching for duplicated images with DBSCAN",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/210723",
  "author_name": "Alexey Pronin",
  "post_date": "2021-01-12T03:32:52.279000",
  "votes": 12,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I released a new TensorFlow notebook demonstrating how to search for duplicated images using image embeddings and the DBSCAN clustering algorithm. Here is the link:</p>\n<p><a href=\"https://www.kaggle.com/graf10a/cldc-image-duplicates-with-dbscan\" target=\"_blank\">\nCLDC_Image_Duplicates_with_DBSCAN\n</a></p>\n<p>Some competitors already reported about the presence of duplicated images in the current competition data. See, for example, <a href=\"https://www.kaggle.com/nakajima/duplicate-train-images?scriptVersionId=47295222\" target=\"_blank\">this public kernel</a>. To the best of my knowledge (please correct me if I am missing something!) 2 pairs of duplicated images were reported so far. With my method I found 3 more pairs (see below), which is nice! Also, the method did not yield any false positive, so its precision is 100% (with unknown recall). Anyhow, I think it looks very promising and decided to share it. Also, I recall that some of the high LB rank participants mentioned that they were training their models on the combined 2020/2019 data with duplicates removed with DBSCAN. Maybe this is going to be our key to success?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1461289%2Ff13fc114c39d0fe77ad36f7a8265f24d%2FDuplicates.png?generation=1610643872443693&amp;alt=media\" alt=\"\"></p>\n<p>Enjoy!</p>",
  "messages": [
    {
      "id": 1149653,
      "postDate": "2021-01-12T03:32:52.280Z",
      "content": "<p>I released a new TensorFlow notebook demonstrating how to search for duplicated images using image embeddings and the DBSCAN clustering algorithm. Here is the link:</p>\n<p><a href=\"https://www.kaggle.com/graf10a/cldc-image-duplicates-with-dbscan\" target=\"_blank\">\nCLDC_Image_Duplicates_with_DBSCAN\n</a></p>\n<p>Some competitors already reported about the presence of duplicated images in the current competition data. See, for example, <a href=\"https://www.kaggle.com/nakajima/duplicate-train-images?scriptVersionId=47295222\" target=\"_blank\">this public kernel</a>. To the best of my knowledge (please correct me if I am missing something!) 2 pairs of duplicated images were reported so far. With my method I found 3 more pairs (see below), which is nice! Also, the method did not yield any false positive, so its precision is 100% (with unknown recall). Anyhow, I think it looks very promising and decided to share it. Also, I recall that some of the high LB rank participants mentioned that they were training their models on the combined 2020/2019 data with duplicates removed with DBSCAN. Maybe this is going to be our key to success?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1461289%2Ff13fc114c39d0fe77ad36f7a8265f24d%2FDuplicates.png?generation=1610643872443693&amp;alt=media\" alt=\"\"></p>\n<p>Enjoy!</p>",
      "rawMarkdown": "I released a new TensorFlow notebook demonstrating how to search for duplicated images using image embeddings and the DBSCAN clustering algorithm. Here is the link:\n\n[\nCLDC_Image_Duplicates_with_DBSCAN\n](https://www.kaggle.com/graf10a/cldc-image-duplicates-with-dbscan)\n\nSome competitors already reported about the presence of duplicated images in the current competition data. See, for example, [this public kernel](https://www.kaggle.com/nakajima/duplicate-train-images?scriptVersionId=47295222). To the best of my knowledge (please correct me if I am missing something!) 2 pairs of duplicated images were reported so far. With my method I found 3 more pairs (see below), which is nice! Also, the method did not yield any false positive, so its precision is 100% (with unknown recall). Anyhow, I think it looks very promising and decided to share it. Also, I recall that some of the high LB rank participants mentioned that they were training their models on the combined 2020/2019 data with duplicates removed with DBSCAN. Maybe this is going to be our key to success?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1461289%2Ff13fc114c39d0fe77ad36f7a8265f24d%2FDuplicates.png?generation=1610643872443693&alt=media)\n\nEnjoy!",
      "votes": 12
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1149653": "I released a new TensorFlow notebook demonstrating how to search for duplicated images using image embeddings and the DBSCAN clustering algorithm. Here is the link:\n\n[\nCLDC_Image_Duplicates_with_DBSCAN\n](https://www.kaggle.com/graf10a/cldc-image-duplicates-with-dbscan)\n\nSome competitors already reported about the presence of duplicated images in the current competition data. See, for example, [this public kernel](https://www.kaggle.com/nakajima/duplicate-train-images?scriptVersionId=47295222). To the best of my knowledge (please correct me if I am missing something!) 2 pairs of duplicated images were reported so far. With my method I found 3 more pairs (see below), which is nice! Also, the method did not yield any false positive, so its precision is 100% (with unknown recall). Anyhow, I think it looks very promising and decided to share it. Also, I recall that some of the high LB rank participants mentioned that they were training their models on the combined 2020/2019 data with duplicates removed with DBSCAN. Maybe this is going to be our key to success?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1461289%2Ff13fc114c39d0fe77ad36f7a8265f24d%2FDuplicates.png?generation=1610643872443693&alt=media)\n\nEnjoy!"
  }
}