{
  "id": 672414,
  "title": "Duplicated images and resizing images",
  "url": "/competitions/jaguar-re-id/discussion/672414",
  "author_name": "",
  "post_date": "2026-02-08T05:25:38.603756600Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I noticed in train data there are many images duplicated. this means make a model be over fitted in learning. Delete the duplicated ones cause an issue because our data is imbalance so augmentation is best solution.\nAnother problem about size of images is not fixed must be solve it . \nThis is my initial analysis:\n<a href=\"https://www.kaggle.com/code/aliwannous2021/eda-of-jrc\" target=\"_blank\">https://www.kaggle.com/code/aliwannous2021/eda-of-jrc</a></p>",
  "messages": [
    {
      "id": "3403291",
      "postDate": "02/08/2026 05:25:38",
      "content": "<p>I noticed in train data there are many images duplicated. this means make a model be over fitted in learning. Delete the duplicated ones cause an issue because our data is imbalance so augmentation is best solution.\nAnother problem about size of images is not fixed must be solve it . \nThis is my initial analysis:\n<a href=\"https://www.kaggle.com/code/aliwannous2021/eda-of-jrc\" target=\"_blank\">https://www.kaggle.com/code/aliwannous2021/eda-of-jrc</a></p>",
      "rawMarkdown": "I noticed in train data there are many images duplicated. this means make a model be over fitted in learning. Delete the duplicated ones cause an issue because our data is imbalance so augmentation is best solution.\nAnother problem about size of images is not fixed must be solve it . \nThis is my initial analysis:\nhttps://www.kaggle.com/code/aliwannous2021/eda-of-jrc",
      "votes": null
    },
    {
      "id": "3403418",
      "postDate": "02/08/2026 12:38:27",
      "content": "<p>In the overview, the reason for many near-duplicates is stated: \"Training set considerations. The training set contains near-duplicate images where the same jaguar was captured multiple times in quick succession by the same photographer. These near-duplicates can help models learn robust features but may also lead to overfitting if not handled carefully. Note that near-duplicates between the training and test sets have been removed to ensure fair evaluation.\"</p>",
      "rawMarkdown": "In the overview, the reason for many near-duplicates is stated: \"Training set considerations. The training set contains near-duplicate images where the same jaguar was captured multiple times in quick succession by the same photographer. These near-duplicates can help models learn robust features but may also lead to overfitting if not handled carefully. Note that near-duplicates between the training and test sets have been removed to ensure fair evaluation.\"",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3403418,
      "author_name": "tonyhauptmann",
      "author_url": "",
      "post_date": "02/08/2026 12:38:27",
      "content": "<p>In the overview, the reason for many near-duplicates is stated: \"Training set considerations. The training set contains near-duplicate images where the same jaguar was captured multiple times in quick succession by the same photographer. These near-duplicates can help models learn robust features but may also lead to overfitting if not handled carefully. Note that near-duplicates between the training and test sets have been removed to ensure fair evaluation.\"</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3403291": "I noticed in train data there are many images duplicated. this means make a model be over fitted in learning. Delete the duplicated ones cause an issue because our data is imbalance so augmentation is best solution.\nAnother problem about size of images is not fixed must be solve it . \nThis is my initial analysis:\nhttps://www.kaggle.com/code/aliwannous2021/eda-of-jrc",
    "3403418": "In the overview, the reason for many near-duplicates is stated: \"Training set considerations. The training set contains near-duplicate images where the same jaguar was captured multiple times in quick succession by the same photographer. These near-duplicates can help models learn robust features but may also lead to overfitting if not handled carefully. Note that near-duplicates between the training and test sets have been removed to ensure fair evaluation.\""
  },
  "source": "meta"
}