{
  "id": 403375,
  "title": "Are these training datasets actually representative of the full data?",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/403375",
  "author_name": "",
  "post_date": "2023-04-22T18:04:11.081143700Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I believe other people may be having these issues as well but I need a sanity check. Does anyone else have the experience where when running their models on a train-test-split they do significantly better than on the private testing data? Any theories as to why this could be? My thought is maybe the section of data we are given to train on isn't sufficiently representative of the full data. Makes me wish there were more data points we could work with, especially for certain events.</p>",
  "messages": [
    {
      "id": "2230759",
      "postDate": "04/22/2023 18:04:11",
      "content": "<p>I believe other people may be having these issues as well but I need a sanity check. Does anyone else have the experience where when running their models on a train-test-split they do significantly better than on the private testing data? Any theories as to why this could be? My thought is maybe the section of data we are given to train on isn't sufficiently representative of the full data. Makes me wish there were more data points we could work with, especially for certain events.</p>",
      "rawMarkdown": "I believe other people may be having these issues as well but I need a sanity check. Does anyone else have the experience where when running their models on a train-test-split they do significantly better than on the private testing data? Any theories as to why this could be? My thought is maybe the section of data we are given to train on isn't sufficiently representative of the full data. Makes me wish there were more data points we could work with, especially for certain events.",
      "votes": null
    },
    {
      "id": "2237337",
      "postDate": "04/27/2023 14:45:53",
      "content": "<p>How much is \"significantly better\"?<br>\nMy CV does +0.03 more than my public score.</p>\n<p>I think it's no problem.</p>",
      "rawMarkdown": "How much is \"significantly better\"?\nMy CV does +0.03 more than my public score.\n\nI think it's no problem.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2237337,
      "author_name": "avivlevi815",
      "author_url": "",
      "post_date": "04/27/2023 14:45:53",
      "content": "<p>How much is \"significantly better\"?<br>\nMy CV does +0.03 more than my public score.</p>\n<p>I think it's no problem.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2230759": "I believe other people may be having these issues as well but I need a sanity check. Does anyone else have the experience where when running their models on a train-test-split they do significantly better than on the private testing data? Any theories as to why this could be? My thought is maybe the section of data we are given to train on isn't sufficiently representative of the full data. Makes me wish there were more data points we could work with, especially for certain events.",
    "2237337": "How much is \"significantly better\"?\nMy CV does +0.03 more than my public score.\n\nI think it's no problem."
  },
  "source": "meta"
}