{
  "id": 398725,
  "title": "More clarification on the noise in the data",
  "url": "/competitions/early-detection-of-3d-printing-issues/discussion/398725",
  "author_name": "Kenneth",
  "post_date": "2023-03-31T12:42:11.380000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>As pointed out in my previous post, the data has a lot of noise. One more thing to clarify about the noise in the data: it's perfectly fine to work on the data in order to get a higher score. It's not against the rules of this competition.</p>\n<p>If you believe cleaning up the training data will help, do it. If you believe adding even more noise to the data, such as augmenting the training data, will result in a high score, go ahead. You can selectively train on only part of the data if this help.</p>\n<p>However, just as stated in the rules, don't do the things that feel like \"cheating\" to you. For instance, since the test dataset is relatively small, you may be able to find a subset in the training data that will train a model that overfits the test data (I don't know if it's true as I never tried it). This feels like cheating to me. And your source code will show it. So don't do it.</p>\n<p>Please comment below if you need further clarifications, especially if you don't think the rules around \"cheating\" are not clear.</p>",
  "messages": [
    {
      "id": 2204245,
      "postDate": "2023-03-31T12:42:11.380Z",
      "content": "<p>As pointed out in my previous post, the data has a lot of noise. One more thing to clarify about the noise in the data: it's perfectly fine to work on the data in order to get a higher score. It's not against the rules of this competition.</p>\n<p>If you believe cleaning up the training data will help, do it. If you believe adding even more noise to the data, such as augmenting the training data, will result in a high score, go ahead. You can selectively train on only part of the data if this help.</p>\n<p>However, just as stated in the rules, don't do the things that feel like \"cheating\" to you. For instance, since the test dataset is relatively small, you may be able to find a subset in the training data that will train a model that overfits the test data (I don't know if it's true as I never tried it). This feels like cheating to me. And your source code will show it. So don't do it.</p>\n<p>Please comment below if you need further clarifications, especially if you don't think the rules around \"cheating\" are not clear.</p>",
      "rawMarkdown": "As pointed out in my previous post, the data has a lot of noise. One more thing to clarify about the noise in the data: it's perfectly fine to work on the data in order to get a higher score. It's not against the rules of this competition.\n\nIf you believe cleaning up the training data will help, do it. If you believe adding even more noise to the data, such as augmenting the training data, will result in a high score, go ahead. You can selectively train on only part of the data if this help.\n\nHowever, just as stated in the rules, don't do the things that feel like \"cheating\" to you. For instance, since the test dataset is relatively small, you may be able to find a subset in the training data that will train a model that overfits the test data (I don't know if it's true as I never tried it). This feels like cheating to me. And your source code will show it. So don't do it.\n\nPlease comment below if you need further clarifications, especially if you don't think the rules around \"cheating\" are not clear.",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2204245": "As pointed out in my previous post, the data has a lot of noise. One more thing to clarify about the noise in the data: it's perfectly fine to work on the data in order to get a higher score. It's not against the rules of this competition.\n\nIf you believe cleaning up the training data will help, do it. If you believe adding even more noise to the data, such as augmenting the training data, will result in a high score, go ahead. You can selectively train on only part of the data if this help.\n\nHowever, just as stated in the rules, don't do the things that feel like \"cheating\" to you. For instance, since the test dataset is relatively small, you may be able to find a subset in the training data that will train a model that overfits the test data (I don't know if it's true as I never tried it). This feels like cheating to me. And your source code will show it. So don't do it.\n\nPlease comment below if you need further clarifications, especially if you don't think the rules around \"cheating\" are not clear."
  }
}