{
  "id": 92740,
  "title": "Overfitting Insights",
  "url": "/competitions/avito-demand-prediction/discussion/92740",
  "author_name": "Miguel Niblock",
  "post_date": "2019-05-19T21:50:19.146000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>After working on this for a couple of weeks for a final project, Im currently getting very consistent cross-validated RMSE scores of 0.21 on the train data, but when I submit my test predictions, they land on the range of 0.23-0.24. </p>\n\n<p>This puzzles me because I've learned to identify ovrefitting through cross-validation. How could I combat overfitting if the cross-validation shows no signs of it whatsoever?</p>\n\n<p>Another explanation is that the test data must be very different from the train set. If they were the same, then I'd be scoring 0.21 and would be in first place.</p>\n\n<p>I also wonder if people in the first places are also getting better scores in their own train validation, and their submissions just happen to land around the 0.22 range.</p>",
  "messages": [
    {
      "id": 533740,
      "postDate": "2019-05-19T21:50:19.147Z",
      "content": "<p>After working on this for a couple of weeks for a final project, Im currently getting very consistent cross-validated RMSE scores of 0.21 on the train data, but when I submit my test predictions, they land on the range of 0.23-0.24. </p>\n\n<p>This puzzles me because I've learned to identify ovrefitting through cross-validation. How could I combat overfitting if the cross-validation shows no signs of it whatsoever?</p>\n\n<p>Another explanation is that the test data must be very different from the train set. If they were the same, then I'd be scoring 0.21 and would be in first place.</p>\n\n<p>I also wonder if people in the first places are also getting better scores in their own train validation, and their submissions just happen to land around the 0.22 range.</p>",
      "rawMarkdown": "After working on this for a couple of weeks for a final project, Im currently getting very consistent cross-validated RMSE scores of 0.21 on the train data, but when I submit my test predictions, they land on the range of 0.23-0.24. \n\nThis puzzles me because I've learned to identify ovrefitting through cross-validation. How could I combat overfitting if the cross-validation shows no signs of it whatsoever?\n\nAnother explanation is that the test data must be very different from the train set. If they were the same, then I'd be scoring 0.21 and would be in first place.\n\nI also wonder if people in the first places are also getting better scores in their own train validation, and their submissions just happen to land around the 0.22 range."
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "533740": "After working on this for a couple of weeks for a final project, Im currently getting very consistent cross-validated RMSE scores of 0.21 on the train data, but when I submit my test predictions, they land on the range of 0.23-0.24. \n\nThis puzzles me because I've learned to identify ovrefitting through cross-validation. How could I combat overfitting if the cross-validation shows no signs of it whatsoever?\n\nAnother explanation is that the test data must be very different from the train set. If they were the same, then I'd be scoring 0.21 and would be in first place.\n\nI also wonder if people in the first places are also getting better scores in their own train validation, and their submissions just happen to land around the 0.22 range."
  }
}