{
  "id": 339573,
  "title": "Bagging or full data training",
  "url": "/competitions/amex-default-prediction/discussion/339573",
  "author_name": "delai50",
  "post_date": "2022-07-25T13:52:53.119000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I have always had this question about which strategy is better for inference:</p>\n<ul>\n<li>Option 1: average of models trained on the full dataset with different model random seeds.</li>\n<li>Option 2: average of models trained on each fold with different CV split random seeds.</li>\n</ul>\n<p>Let’s say that the available computation time is fixed, I see the following features:</p>\n<ul>\n<li>Option 1: more accurate models, less models trained in that fixed time, less variance between models.</li>\n<li>Option 2: less accurate models, more models trained in that fixed time, more variance between models.</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> shared in this <a href=\"https://www.youtube.com/watch?v=d7YvGX3oi3A&amp;t=2776s\" target=\"_blank\">video</a> the following trick:</p>\n<blockquote>\n  <p>A trick that works nearly every time is the fact that you can have a local validation at 5 folds, that means that you only training models on 80% of the data and you are inferring on this out of fold. But then when it is time to make your submission to the competition, don’t take the ensemble of those 5 models, what you do is you just remember all the hyperparameters that work, and now what you do is build a new neural network and you train it in a 100% of the data with those parameters. You just keep randomizing seeds, say you initialize it 5 times, so you make 5 models that are trained on 100% of the data and you submit those 5 models instead of 5 models that were trained on 80% of the data and that nearly always increases your score like 10%.</p>\n</blockquote>\n<p>However, I think it is not a fair comparison because the time you spent training models on 100% data is higher, and during the same time you could have trained a more 80% data models.</p>\n<p>Something that I never seen before until this <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <a href=\"https://www.kaggle.com/competitions/nbme-score-clinical-patient-notes/discussion/323085\" target=\"_blank\">solution</a> was a mix of the 2 strategies. If I understood correctly, he used both the fold models and the full data trained models.</p>\n<p>I have never conducted such a study yet, so how do you usually make your final submission?</p>\n<p>Cheers!</p>",
  "messages": [
    {
      "id": 1870362,
      "postDate": "2022-07-25T13:52:53.120Z",
      "content": "<p>I have always had this question about which strategy is better for inference:</p>\n<ul>\n<li>Option 1: average of models trained on the full dataset with different model random seeds.</li>\n<li>Option 2: average of models trained on each fold with different CV split random seeds.</li>\n</ul>\n<p>Let’s say that the available computation time is fixed, I see the following features:</p>\n<ul>\n<li>Option 1: more accurate models, less models trained in that fixed time, less variance between models.</li>\n<li>Option 2: less accurate models, more models trained in that fixed time, more variance between models.</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> shared in this <a href=\"https://www.youtube.com/watch?v=d7YvGX3oi3A&amp;t=2776s\" target=\"_blank\">video</a> the following trick:</p>\n<blockquote>\n  <p>A trick that works nearly every time is the fact that you can have a local validation at 5 folds, that means that you only training models on 80% of the data and you are inferring on this out of fold. But then when it is time to make your submission to the competition, don’t take the ensemble of those 5 models, what you do is you just remember all the hyperparameters that work, and now what you do is build a new neural network and you train it in a 100% of the data with those parameters. You just keep randomizing seeds, say you initialize it 5 times, so you make 5 models that are trained on 100% of the data and you submit those 5 models instead of 5 models that were trained on 80% of the data and that nearly always increases your score like 10%.</p>\n</blockquote>\n<p>However, I think it is not a fair comparison because the time you spent training models on 100% data is higher, and during the same time you could have trained a more 80% data models.</p>\n<p>Something that I never seen before until this <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <a href=\"https://www.kaggle.com/competitions/nbme-score-clinical-patient-notes/discussion/323085\" target=\"_blank\">solution</a> was a mix of the 2 strategies. If I understood correctly, he used both the fold models and the full data trained models.</p>\n<p>I have never conducted such a study yet, so how do you usually make your final submission?</p>\n<p>Cheers!</p>",
      "rawMarkdown": "I have always had this question about which strategy is better for inference:\n\n- Option 1: average of models trained on the full dataset with different model random seeds.\n- Option 2: average of models trained on each fold with different CV split random seeds.\n\nLet’s say that the available computation time is fixed, I see the following features:\n\n- Option 1: more accurate models, less models trained in that fixed time, less variance between models.\n- Option 2: less accurate models, more models trained in that fixed time, more variance between models.\n\n@cdeotte shared in this [video](https://www.youtube.com/watch?v=d7YvGX3oi3A&t=2776s) the following trick:\n> A trick that works nearly every time is the fact that you can have a local validation at 5 folds, that means that you only training models on 80% of the data and you are inferring on this out of fold. But then when it is time to make your submission to the competition, don’t take the ensemble of those 5 models, what you do is you just remember all the hyperparameters that work, and now what you do is build a new neural network and you train it in a 100% of the data with those parameters. You just keep randomizing seeds, say you initialize it 5 times, so you make 5 models that are trained on 100% of the data and you submit those 5 models instead of 5 models that were trained on 80% of the data and that nearly always increases your score like 10%.\n\nHowever, I think it is not a fair comparison because the time you spent training models on 100% data is higher, and during the same time you could have trained a more 80% data models.\n\nSomething that I never seen before until this @cpmpml [solution](https://www.kaggle.com/competitions/nbme-score-clinical-patient-notes/discussion/323085) was a mix of the 2 strategies. If I understood correctly, he used both the fold models and the full data trained models.\n\nI have never conducted such a study yet, so how do you usually make your final submission?\n\nCheers!",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1870362": "I have always had this question about which strategy is better for inference:\n\n- Option 1: average of models trained on the full dataset with different model random seeds.\n- Option 2: average of models trained on each fold with different CV split random seeds.\n\nLet’s say that the available computation time is fixed, I see the following features:\n\n- Option 1: more accurate models, less models trained in that fixed time, less variance between models.\n- Option 2: less accurate models, more models trained in that fixed time, more variance between models.\n\n@cdeotte shared in this [video](https://www.youtube.com/watch?v=d7YvGX3oi3A&t=2776s) the following trick:\n> A trick that works nearly every time is the fact that you can have a local validation at 5 folds, that means that you only training models on 80% of the data and you are inferring on this out of fold. But then when it is time to make your submission to the competition, don’t take the ensemble of those 5 models, what you do is you just remember all the hyperparameters that work, and now what you do is build a new neural network and you train it in a 100% of the data with those parameters. You just keep randomizing seeds, say you initialize it 5 times, so you make 5 models that are trained on 100% of the data and you submit those 5 models instead of 5 models that were trained on 80% of the data and that nearly always increases your score like 10%.\n\nHowever, I think it is not a fair comparison because the time you spent training models on 100% data is higher, and during the same time you could have trained a more 80% data models.\n\nSomething that I never seen before until this @cpmpml [solution](https://www.kaggle.com/competitions/nbme-score-clinical-patient-notes/discussion/323085) was a mix of the 2 strategies. If I understood correctly, he used both the fold models and the full data trained models.\n\nI have never conducted such a study yet, so how do you usually make your final submission?\n\nCheers!"
  }
}