{
  "id": 251967,
  "title": "How to not overfit",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/251967",
  "author_name": "",
  "post_date": "2021-07-09T17:58:56.348060200Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I'm a noob to machine learning, I've done like half of fast ai's course and some Kaggle competitions. I see that there's lots of data provided so why not use all of it if it seems to help me climb the leaderboard? How does one know how much data is good to use without overfitting?</p>",
  "messages": [
    {
      "id": "1382346",
      "postDate": "07/09/2021 17:58:56",
      "content": "<p>I'm a noob to machine learning, I've done like half of fast ai's course and some Kaggle competitions. I see that there's lots of data provided so why not use all of it if it seems to help me climb the leaderboard? How does one know how much data is good to use without overfitting?</p>",
      "rawMarkdown": "I'm a noob to machine learning, I've done like half of fast ai's course and some Kaggle competitions. I see that there's lots of data provided so why not use all of it if it seems to help me climb the leaderboard? How does one know how much data is good to use without overfitting?",
      "votes": null
    },
    {
      "id": "1382384",
      "postDate": "07/09/2021 18:53:42",
      "content": "<p>Have a Cross-Validation approach, i.e choose your validation set or validation sets in such a way that, if the score on the validation set(s) improves then your leaderboard score also improves, or at least not declines. Test all your features locally for improvement on your validation set score. Overfitting happens if you choose features or blending weights that are only good for the public leaderboard and may not be for the private leaderboard. If you suspect that a certain feature has a different distribution in future data (private leaderboard) compared to training data and public leaderboard data if you use such a feature then the gains from that feature will not occur on the private leader board (Data).</p>",
      "rawMarkdown": "Have a Cross-Validation approach, i.e choose your validation set or validation sets in such a way that, if the score on the validation set(s) improves then your leaderboard score also improves, or at least not declines. Test all your features locally for improvement on your validation set score. Overfitting happens if you choose features or blending weights that are only good for the public leaderboard and may not be for the private leaderboard. If you suspect that a certain feature has a different distribution in future data (private leaderboard) compared to training data and public leaderboard data if you use such a feature then the gains from that feature will not occur on the private leader board (Data).",
      "votes": null
    },
    {
      "id": "1382425",
      "postDate": "07/09/2021 20:42:48",
      "content": "<p>The larger your feature space the greater the chance your model picks up noise rather than signal. <br>\nI recently tried the approach of chucking every feature at my model. It resulted in both poorer LB and CV than my more simple, but did improve the April holdout.<br>\nIn terms of number of observations, I've found including as many of the observations as possible to be better in my experiments. Some people have found only including the players in test set (there is an indicator for this in players csv) have improved their models but I havent found this to be the case.</p>",
      "rawMarkdown": "The larger your feature space the greater the chance your model picks up noise rather than signal. \nI recently tried the approach of chucking every feature at my model. It resulted in both poorer LB and CV than my more simple, but did improve the April holdout.\nIn terms of number of observations, I've found including as many of the observations as possible to be better in my experiments. Some people have found only including the players in test set (there is an indicator for this in players csv) have improved their models but I havent found this to be the case.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1382384,
      "author_name": "chinta",
      "author_url": "",
      "post_date": "07/09/2021 18:53:42",
      "content": "<p>Have a Cross-Validation approach, i.e choose your validation set or validation sets in such a way that, if the score on the validation set(s) improves then your leaderboard score also improves, or at least not declines. Test all your features locally for improvement on your validation set score. Overfitting happens if you choose features or blending weights that are only good for the public leaderboard and may not be for the private leaderboard. If you suspect that a certain feature has a different distribution in future data (private leaderboard) compared to training data and public leaderboard data if you use such a feature then the gains from that feature will not occur on the private leader board (Data).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1382425,
      "author_name": "jacobhowardparker",
      "author_url": "",
      "post_date": "07/09/2021 20:42:48",
      "content": "<p>The larger your feature space the greater the chance your model picks up noise rather than signal. <br>\nI recently tried the approach of chucking every feature at my model. It resulted in both poorer LB and CV than my more simple, but did improve the April holdout.<br>\nIn terms of number of observations, I've found including as many of the observations as possible to be better in my experiments. Some people have found only including the players in test set (there is an indicator for this in players csv) have improved their models but I havent found this to be the case.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1382346": "I'm a noob to machine learning, I've done like half of fast ai's course and some Kaggle competitions. I see that there's lots of data provided so why not use all of it if it seems to help me climb the leaderboard? How does one know how much data is good to use without overfitting?",
    "1382384": "Have a Cross-Validation approach, i.e choose your validation set or validation sets in such a way that, if the score on the validation set(s) improves then your leaderboard score also improves, or at least not declines. Test all your features locally for improvement on your validation set score. Overfitting happens if you choose features or blending weights that are only good for the public leaderboard and may not be for the private leaderboard. If you suspect that a certain feature has a different distribution in future data (private leaderboard) compared to training data and public leaderboard data if you use such a feature then the gains from that feature will not occur on the private leader board (Data).",
    "1382425": "The larger your feature space the greater the chance your model picks up noise rather than signal. \nI recently tried the approach of chucking every feature at my model. It resulted in both poorer LB and CV than my more simple, but did improve the April holdout.\nIn terms of number of observations, I've found including as many of the observations as possible to be better in my experiments. Some people have found only including the players in test set (there is an indicator for this in players csv) have improved their models but I havent found this to be the case."
  },
  "source": "meta"
}