{
  "id": 102648,
  "title": "What do you do on \"duplicate\" train or test image",
  "url": "/competitions/aptos2019-blindness-detection/discussion/102648",
  "author_name": "",
  "post_date": "2019-08-03T14:49:49.750186Z",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>By taking image hash, it's easy to identify visually similar images.</p>\n\n<p><a href=\"https://www.kaggle.com/h4211819/more-information-about-duplicate\">https://www.kaggle.com/h4211819/more-information-about-duplicate</a></p>\n\n<p>I believe one of the reason why it's so easy get 0.7x in LB or 0.9x for CV without doing much data augmentation is because <strong><em>image leakage</em></strong> into test set.</p>\n\n<p>A simple way would be applying augmentation to train data to make it looks different with test set, what other technique can be applied? </p>\n\n<p>Share here. </p>",
  "messages": [
    {
      "id": "591360",
      "postDate": "08/03/2019 14:49:49",
      "content": "<p>By taking image hash, it's easy to identify visually similar images.</p>\n\n<p><a href=\"https://www.kaggle.com/h4211819/more-information-about-duplicate\">https://www.kaggle.com/h4211819/more-information-about-duplicate</a></p>\n\n<p>I believe one of the reason why it's so easy get 0.7x in LB or 0.9x for CV without doing much data augmentation is because <strong><em>image leakage</em></strong> into test set.</p>\n\n<p>A simple way would be applying augmentation to train data to make it looks different with test set, what other technique can be applied? </p>\n\n<p>Share here. </p>",
      "rawMarkdown": "By taking image hash, it's easy to identify visually similar images.\n\nhttps://www.kaggle.com/h4211819/more-information-about-duplicate\n\nI believe one of the reason why it's so easy get 0.7x in LB or 0.9x for CV without doing much data augmentation is because ***image leakage*** into test set.\n\nA simple way would be applying augmentation to train data to make it looks different with test set, what other technique can be applied? \n\nShare here.",
      "votes": null
    },
    {
      "id": "592519",
      "postDate": "08/05/2019 12:16:46",
      "content": "<p>If i remove those \"duplicate\" from train set, can see the LB 0.7xx drop to 0.6xx </p>",
      "rawMarkdown": "If i remove those \"duplicate\" from train set, can see the LB 0.7xx drop to 0.6xx",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 592519,
      "author_name": "ccjoshua",
      "author_url": "",
      "post_date": "08/05/2019 12:16:46",
      "content": "<p>If i remove those \"duplicate\" from train set, can see the LB 0.7xx drop to 0.6xx </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "591360": "By taking image hash, it's easy to identify visually similar images.\n\nhttps://www.kaggle.com/h4211819/more-information-about-duplicate\n\nI believe one of the reason why it's so easy get 0.7x in LB or 0.9x for CV without doing much data augmentation is because ***image leakage*** into test set.\n\nA simple way would be applying augmentation to train data to make it looks different with test set, what other technique can be applied? \n\nShare here.",
    "592519": "If i remove those \"duplicate\" from train set, can see the LB 0.7xx drop to 0.6xx"
  },
  "source": "meta"
}