{
  "id": 227822,
  "title": "Duplicate Images in Training Data",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/227822",
  "author_name": "Rai",
  "post_date": "2021-03-22T12:21:40.948000",
  "votes": 9,
  "comment_count": 2,
  "views": 0,
  "content": "<h2>Duplicate Images in Training Data</h2>\n<p>Most of the top solutions from the previous HPA competition mention identifying and handling duplicates. I decided to check this year's training data, and sure enough, there are duplicate images in the training data.&nbsp; Here are a few examples:</p>\n<p><img src=\"https://i.imgur.com/C6ja0IU.png\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/rai555/hpa-duplicate-images-in-train\" target=\"_blank\">This notebook</a> details how I identified the duplicates, list the images etc. I'm hoping that effectively handling the duplicates&nbsp;will result in a more stable CV and faster training - every&nbsp;little bit helps with such a big dataset!&nbsp;</p>",
  "messages": [
    {
      "id": 1248187,
      "postDate": "2021-03-22T12:21:40.950Z",
      "content": "<h2>Duplicate Images in Training Data</h2>\n<p>Most of the top solutions from the previous HPA competition mention identifying and handling duplicates. I decided to check this year's training data, and sure enough, there are duplicate images in the training data.&nbsp; Here are a few examples:</p>\n<p><img src=\"https://i.imgur.com/C6ja0IU.png\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/rai555/hpa-duplicate-images-in-train\" target=\"_blank\">This notebook</a> details how I identified the duplicates, list the images etc. I'm hoping that effectively handling the duplicates&nbsp;will result in a more stable CV and faster training - every&nbsp;little bit helps with such a big dataset!&nbsp;</p>",
      "rawMarkdown": "## Duplicate Images in Training Data\n\nMost of the top solutions from the previous HPA competition mention identifying and handling duplicates. I decided to check this year's training data, and sure enough, there are duplicate images in the training data.  Here are a few examples:\n\n![](https://i.imgur.com/C6ja0IU.png)\n\n[This notebook](https://www.kaggle.com/rai555/hpa-duplicate-images-in-train) details how I identified the duplicates, list the images etc. I'm hoping that effectively handling the duplicates will result in a more stable CV and faster training - every little bit helps with such a big dataset! ",
      "votes": 9
    },
    {
      "id": 1248368,
      "postDate": "2021-03-22T14:44:26.213Z",
      "content": "<p>It seems to me that the impact of duplicates does depend on the level of augmentation being used to train the model.   If your using heavy augmentation than I would not expect much change in CV.</p>\n<p>Of course if the duplicates have different labels than that's a whole new kettle of fish.  I believe other discussion posts on duplicates have indicated that some do have different labels.  You might want to include label information - for sure I think removal of duplicates with different labels is an important consideration.</p>",
      "rawMarkdown": "It seems to me that the impact of duplicates does depend on the level of augmentation being used to train the model.   If your using heavy augmentation than I would not expect much change in CV.\n\nOf course if the duplicates have different labels than that's a whole new kettle of fish.  I believe other discussion posts on duplicates have indicated that some do have different labels.  You might want to include label information - for sure I think removal of duplicates with different labels is an important consideration.",
      "votes": 1,
      "replies": [
        {
          "id": 1248656,
          "postDate": "2021-03-22T18:26:56.230Z",
          "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> I've just started my EDA for this competition&nbsp;and I seem to have missed the other discussions about&nbsp;duplicates. Thanks for letting me know - I'll go find them!</p>\n<p>It seems like the duplicates have the same labels but I haven't&nbsp;had the time to look too closely. I'll update the post and the notebooks with labels in a bit.</p>",
          "rawMarkdown": "@pcjimmmy I've just started my EDA for this competition and I seem to have missed the other discussions about duplicates. Thanks for letting me know - I'll go find them!\n\nIt seems like the duplicates have the same labels but I haven't had the time to look too closely. I'll update the post and the notebooks with labels in a bit.\n\n",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1248368,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2021-03-22T14:44:26.213000",
      "content": "<p>It seems to me that the impact of duplicates does depend on the level of augmentation being used to train the model.   If your using heavy augmentation than I would not expect much change in CV.</p>\n<p>Of course if the duplicates have different labels than that's a whole new kettle of fish.  I believe other discussion posts on duplicates have indicated that some do have different labels.  You might want to include label information - for sure I think removal of duplicates with different labels is an important consideration.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1248656,
          "author_name": "Rai",
          "author_url": "",
          "post_date": "2021-03-22T18:26:56.230000",
          "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> I've just started my EDA for this competition&nbsp;and I seem to have missed the other discussions about&nbsp;duplicates. Thanks for letting me know - I'll go find them!</p>\n<p>It seems like the duplicates have the same labels but I haven't&nbsp;had the time to look too closely. I'll update the post and the notebooks with labels in a bit.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1248187": "## Duplicate Images in Training Data\n\nMost of the top solutions from the previous HPA competition mention identifying and handling duplicates. I decided to check this year's training data, and sure enough, there are duplicate images in the training data.  Here are a few examples:\n\n![](https://i.imgur.com/C6ja0IU.png)\n\n[This notebook](https://www.kaggle.com/rai555/hpa-duplicate-images-in-train) details how I identified the duplicates, list the images etc. I'm hoping that effectively handling the duplicates will result in a more stable CV and faster training - every little bit helps with such a big dataset! ",
    "1248368": "It seems to me that the impact of duplicates does depend on the level of augmentation being used to train the model.   If your using heavy augmentation than I would not expect much change in CV.\n\nOf course if the duplicates have different labels than that's a whole new kettle of fish.  I believe other discussion posts on duplicates have indicated that some do have different labels.  You might want to include label information - for sure I think removal of duplicates with different labels is an important consideration."
  }
}