{
  "id": 573943,
  "title": "Using noisy data in model training",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/573943",
  "author_name": "",
  "post_date": "2025-04-18T22:08:05.484384800Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I used noisy data during the training phase with YOLO, but it resulted in worse performance during the evaluation phase. Was this due to too much noise or too little noise? Do similar/duplicate data reduce the model's performance?</p>",
  "messages": [
    {
      "id": "3182186",
      "postDate": "04/18/2025 22:08:05",
      "content": "<p>I used noisy data during the training phase with YOLO, but it resulted in worse performance during the evaluation phase. Was this due to too much noise or too little noise? Do similar/duplicate data reduce the model's performance?</p>",
      "rawMarkdown": "I used noisy data during the training phase with YOLO, but it resulted in worse performance during the evaluation phase. Was this due to too much noise or too little noise? Do similar/duplicate data reduce the model's performance?",
      "votes": null
    },
    {
      "id": "3182725",
      "postDate": "04/19/2025 18:25:16",
      "content": "<p>Some fundamental info on Noisy data:</p>\n<p>When we train a model, it learns patterns from the data and tries to form a mathematical equation or function that best represents those patterns.</p>\n<p>However, real-world data often contains noise—these are data points that do not follow the general trend or are the result of errors, or inconsistencies (e.g., data entry mistakes or unusual values). Noise does not represent the true underlying pattern and can mislead the model.</p>\n<p>Hyperparameter tuning can help manage this by adjusting the model’s complexity so that it focuses on learning the true patterns in the data, rather than fitting the noise.</p>",
      "rawMarkdown": "Some fundamental info on Noisy data:\n\nWhen we train a model, it learns patterns from the data and tries to form a mathematical equation or function that best represents those patterns.\n\nHowever, real-world data often contains noise—these are data points that do not follow the general trend or are the result of errors, or inconsistencies (e.g., data entry mistakes or unusual values). Noise does not represent the true underlying pattern and can mislead the model.\n\nHyperparameter tuning can help manage this by adjusting the model’s complexity so that it focuses on learning the true patterns in the data, rather than fitting the noise.",
      "votes": null
    },
    {
      "id": "3182767",
      "postDate": "04/19/2025 19:24:45",
      "content": "<p>So, does using a limited amount of noisy data help the model?</p>",
      "rawMarkdown": "So, does using a limited amount of noisy data help the model?",
      "votes": null
    },
    {
      "id": "3182973",
      "postDate": "04/20/2025 07:45:30",
      "content": "<p>I think not</p>",
      "rawMarkdown": "I think not",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3182725,
      "author_name": "atulkgoyl",
      "author_url": "",
      "post_date": "04/19/2025 18:25:16",
      "content": "<p>Some fundamental info on Noisy data:</p>\n<p>When we train a model, it learns patterns from the data and tries to form a mathematical equation or function that best represents those patterns.</p>\n<p>However, real-world data often contains noise—these are data points that do not follow the general trend or are the result of errors, or inconsistencies (e.g., data entry mistakes or unusual values). Noise does not represent the true underlying pattern and can mislead the model.</p>\n<p>Hyperparameter tuning can help manage this by adjusting the model’s complexity so that it focuses on learning the true patterns in the data, rather than fitting the noise.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3182767,
          "author_name": "ali78kabirzadeh",
          "author_url": "",
          "post_date": "04/19/2025 19:24:45",
          "content": "<p>So, does using a limited amount of noisy data help the model?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3182973,
              "author_name": "atulkgoyl",
              "author_url": "",
              "post_date": "04/20/2025 07:45:30",
              "content": "<p>I think not</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3182186": "I used noisy data during the training phase with YOLO, but it resulted in worse performance during the evaluation phase. Was this due to too much noise or too little noise? Do similar/duplicate data reduce the model's performance?",
    "3182725": "Some fundamental info on Noisy data:\n\nWhen we train a model, it learns patterns from the data and tries to form a mathematical equation or function that best represents those patterns.\n\nHowever, real-world data often contains noise—these are data points that do not follow the general trend or are the result of errors, or inconsistencies (e.g., data entry mistakes or unusual values). Noise does not represent the true underlying pattern and can mislead the model.\n\nHyperparameter tuning can help manage this by adjusting the model’s complexity so that it focuses on learning the true patterns in the data, rather than fitting the noise.",
    "3182767": "So, does using a limited amount of noisy data help the model?",
    "3182973": "I think not"
  },
  "source": "meta"
}