{
  "id": 226499,
  "title": "Should I always prepare a holdout subset for validation when beginning a new competition?",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/226499",
  "author_name": "",
  "post_date": "2021-03-16T16:34:35.440933600Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi! I'm asking this question for one of the final steps where you ensemble all your best models and make the final predictions. </p>\n<p>Say I use the same 5 folds for all my models and save all the OOF predictions for these models, which are a good indication for model performance. Now I want to ensemble these models. A simple average of all the OOF predictions will usually give better result than any individual model, assuming they have good correlation. However, if I want to try something fancier (even just weighted average), then I am potentially risking overfitting. Is that correct? (Because I think your ensemble algorithm will always look for the best parameters to give the highest ensemble score). Simple models like logistic regression are less likely to overfit, I guess, but it's still possible.</p>\n<p>So should I save a subset of the original training data which I never use when training all my models and use this as an indicator whether my final ensemble overfits? I believe technically the public test dataset serves this purpose, but one can only make 5 submission which seems a bit limited.</p>\n<p>That said, if I did save a subset for the aforementioned purpose, I would be using fewer training samples compared with other people during normal training, which might lead to underperformance?</p>\n<p>Really curious and would appreciate your thoughts on this. Thanks so much in advance!</p>",
  "messages": [
    {
      "id": "1240818",
      "postDate": "03/16/2021 16:34:35",
      "content": "<p>Hi! I'm asking this question for one of the final steps where you ensemble all your best models and make the final predictions. </p>\n<p>Say I use the same 5 folds for all my models and save all the OOF predictions for these models, which are a good indication for model performance. Now I want to ensemble these models. A simple average of all the OOF predictions will usually give better result than any individual model, assuming they have good correlation. However, if I want to try something fancier (even just weighted average), then I am potentially risking overfitting. Is that correct? (Because I think your ensemble algorithm will always look for the best parameters to give the highest ensemble score). Simple models like logistic regression are less likely to overfit, I guess, but it's still possible.</p>\n<p>So should I save a subset of the original training data which I never use when training all my models and use this as an indicator whether my final ensemble overfits? I believe technically the public test dataset serves this purpose, but one can only make 5 submission which seems a bit limited.</p>\n<p>That said, if I did save a subset for the aforementioned purpose, I would be using fewer training samples compared with other people during normal training, which might lead to underperformance?</p>\n<p>Really curious and would appreciate your thoughts on this. Thanks so much in advance!</p>",
      "rawMarkdown": "Hi! I'm asking this question for one of the final steps where you ensemble all your best models and make the final predictions. \n\nSay I use the same 5 folds for all my models and save all the OOF predictions for these models, which are a good indication for model performance. Now I want to ensemble these models. A simple average of all the OOF predictions will usually give better result than any individual model, assuming they have good correlation. However, if I want to try something fancier (even just weighted average), then I am potentially risking overfitting. Is that correct? (Because I think your ensemble algorithm will always look for the best parameters to give the highest ensemble score). Simple models like logistic regression are less likely to overfit, I guess, but it's still possible.\n\nSo should I save a subset of the original training data which I never use when training all my models and use this as an indicator whether my final ensemble overfits? I believe technically the public test dataset serves this purpose, but one can only make 5 submission which seems a bit limited.\n\nThat said, if I did save a subset for the aforementioned purpose, I would be using fewer training samples compared with other people during normal training, which might lead to underperformance?\n\nReally curious and would appreciate your thoughts on this. Thanks so much in advance!",
      "votes": null
    },
    {
      "id": "1240846",
      "postDate": "03/16/2021 16:56:53",
      "content": "<p>The standard approach to avoid overfitting when doing the next level of combining models is to, again, use cross-validation to set hyperparameters. E.g. you would not just take all the out-of-fold predictions of your models and find the averaging weights that optimize the CV score across all out of folds. That can work decently, if there's a huge amount of data, but the smaller the data are, the more likely this is overfitting to the out-of-fold predictions. </p>\n<p>You would use CV to instead find the level of penalization (e.g. penalizing how much you go away from a simple average) that optimizes out-of-fold predictions: i.e. you optimize on the original training fold with some penalization, then apply that averaging approach to the original validation fold and see what's the best hyperparameters for that.</p>\n<p>It's my impression is that this approach is pretty successful on Kaggle and that this way one largely gets away with not having an extra independent validation set (and to some extent the public LB serves as that). \"Gets away with\" meaning \"it tends to beat the alternatives on the private LB\", not we have a completely unbiased assessment of model performance via the CV, of course.</p>",
      "rawMarkdown": "The standard approach to avoid overfitting when doing the next level of combining models is to, again, use cross-validation to set hyperparameters. E.g. you would not just take all the out-of-fold predictions of your models and find the averaging weights that optimize the CV score across all out of folds. That can work decently, if there's a huge amount of data, but the smaller the data are, the more likely this is overfitting to the out-of-fold predictions. \n\nYou would use CV to instead find the level of penalization (e.g. penalizing how much you go away from a simple average) that optimizes out-of-fold predictions: i.e. you optimize on the original training fold with some penalization, then apply that averaging approach to the original validation fold and see what's the best hyperparameters for that.\n\nIt's my impression is that this approach is pretty successful on Kaggle and that this way one largely gets away with not having an extra independent validation set (and to some extent the public LB serves as that). \"Gets away with\" meaning \"it tends to beat the alternatives on the private LB\", not we have a completely unbiased assessment of model performance via the CV, of course.",
      "votes": null
    },
    {
      "id": "1240879",
      "postDate": "03/16/2021 17:25:56",
      "content": "<p>Thanks a lot for your reply! That clarifies things for me. I just have two follow up qustions:</p>\n<ol>\n<li><p>When I make CV again for my OOF predictions, should I use stratified CV? i.e. if my OOF predictions have 5000 entries (~1000 from each training fold - 5 folds in total), should my next-level CV have ~200 entries from each of these 5 folds for a total of 1000 entries per fold?</p></li>\n<li><p>Can you specify some penalization methods you mentioned? Does l2 regularization in a logistic regression model count?</p></li>\n</ol>\n<p>Thanks again!</p>",
      "rawMarkdown": "Thanks a lot for your reply! That clarifies things for me. I just have two follow up qustions:\n\n1. When I make CV again for my OOF predictions, should I use stratified CV? i.e. if my OOF predictions have 5000 entries (~1000 from each training fold - 5 folds in total), should my next-level CV have ~200 entries from each of these 5 folds for a total of 1000 entries per fold?\n\n2. Can you specify some penalization methods you mentioned? Does l2 regularization in a logistic regression model count?\n\nThanks again!",
      "votes": null
    },
    {
      "id": "1240896",
      "postDate": "03/16/2021 17:58:37",
      "content": "<p>With CV, I actually mean re-using the original CV at the next level of ensembling/stacking. This is a <a href=\"https://datasciblog.github.io/2016/12/27/a-kagglers-guide-to-model-stacking-in-practice/\" target=\"_blank\">nice blog post</a> on the topic that's pretty readable (what I mention is not the only option discussed, but what I describe is discussed in <a href=\"https://datasciblog.github.io/2016/12/27/a-kagglers-guide-to-model-stacking-in-practice/\" target=\"_blank\">this section</a>). There's of course also a good bit of discussion around on the Kaggle discussion forums (e.g. <a href=\"https://www.kaggle.com/general/18793\" target=\"_blank\">this</a>).</p>\n<p>Certainly, the parameter controlling the strength of L2 (and/or L1) regularization could be one sensible hyperparameter for regularizing/penalizing a model.</p>\n<p>It's of course an interesting question how much of an overfitting risk there is in this competition, if one does not take this approach. Given that I got quite different answers on how to average depending on how I approached it, I think it may matter to some extent. If the shake-up favors me, I'll take that as an indication that I got it right, if not…</p>",
      "rawMarkdown": "With CV, I actually mean re-using the original CV at the next level of ensembling/stacking. This is a [nice blog post](https://datasciblog.github.io/2016/12/27/a-kagglers-guide-to-model-stacking-in-practice/) on the topic that's pretty readable (what I mention is not the only option discussed, but what I describe is discussed in [this section](https://datasciblog.github.io/2016/12/27/a-kagglers-guide-to-model-stacking-in-practice/)). There's of course also a good bit of discussion around on the Kaggle discussion forums (e.g. [this](https://www.kaggle.com/general/18793)).\n\nCertainly, the parameter controlling the strength of L2 (and/or L1) regularization could be one sensible hyperparameter for regularizing/penalizing a model.\n\nIt's of course an interesting question how much of an overfitting risk there is in this competition, if one does not take this approach. Given that I got quite different answers on how to average depending on how I approached it, I think it may matter to some extent. If the shake-up favors me, I'll take that as an indication that I got it right, if not...",
      "votes": null
    },
    {
      "id": "1241435",
      "postDate": "03/17/2021 03:28:46",
      "content": "<p>most of the time, you should.</p>",
      "rawMarkdown": "most of the time, you should.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1240846,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "03/16/2021 16:56:53",
      "content": "<p>The standard approach to avoid overfitting when doing the next level of combining models is to, again, use cross-validation to set hyperparameters. E.g. you would not just take all the out-of-fold predictions of your models and find the averaging weights that optimize the CV score across all out of folds. That can work decently, if there's a huge amount of data, but the smaller the data are, the more likely this is overfitting to the out-of-fold predictions. </p>\n<p>You would use CV to instead find the level of penalization (e.g. penalizing how much you go away from a simple average) that optimizes out-of-fold predictions: i.e. you optimize on the original training fold with some penalization, then apply that averaging approach to the original validation fold and see what's the best hyperparameters for that.</p>\n<p>It's my impression is that this approach is pretty successful on Kaggle and that this way one largely gets away with not having an extra independent validation set (and to some extent the public LB serves as that). \"Gets away with\" meaning \"it tends to beat the alternatives on the private LB\", not we have a completely unbiased assessment of model performance via the CV, of course.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1240879,
          "author_name": "yinanli516",
          "author_url": "",
          "post_date": "03/16/2021 17:25:56",
          "content": "<p>Thanks a lot for your reply! That clarifies things for me. I just have two follow up qustions:</p>\n<ol>\n<li><p>When I make CV again for my OOF predictions, should I use stratified CV? i.e. if my OOF predictions have 5000 entries (~1000 from each training fold - 5 folds in total), should my next-level CV have ~200 entries from each of these 5 folds for a total of 1000 entries per fold?</p></li>\n<li><p>Can you specify some penalization methods you mentioned? Does l2 regularization in a logistic regression model count?</p></li>\n</ol>\n<p>Thanks again!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1240896,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "03/16/2021 17:58:37",
          "content": "<p>With CV, I actually mean re-using the original CV at the next level of ensembling/stacking. This is a <a href=\"https://datasciblog.github.io/2016/12/27/a-kagglers-guide-to-model-stacking-in-practice/\" target=\"_blank\">nice blog post</a> on the topic that's pretty readable (what I mention is not the only option discussed, but what I describe is discussed in <a href=\"https://datasciblog.github.io/2016/12/27/a-kagglers-guide-to-model-stacking-in-practice/\" target=\"_blank\">this section</a>). There's of course also a good bit of discussion around on the Kaggle discussion forums (e.g. <a href=\"https://www.kaggle.com/general/18793\" target=\"_blank\">this</a>).</p>\n<p>Certainly, the parameter controlling the strength of L2 (and/or L1) regularization could be one sensible hyperparameter for regularizing/penalizing a model.</p>\n<p>It's of course an interesting question how much of an overfitting risk there is in this competition, if one does not take this approach. Given that I got quite different answers on how to average depending on how I approached it, I think it may matter to some extent. If the shake-up favors me, I'll take that as an indication that I got it right, if not…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1241435,
      "author_name": "moewie94",
      "author_url": "",
      "post_date": "03/17/2021 03:28:46",
      "content": "<p>most of the time, you should.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1240818": "Hi! I'm asking this question for one of the final steps where you ensemble all your best models and make the final predictions. \n\nSay I use the same 5 folds for all my models and save all the OOF predictions for these models, which are a good indication for model performance. Now I want to ensemble these models. A simple average of all the OOF predictions will usually give better result than any individual model, assuming they have good correlation. However, if I want to try something fancier (even just weighted average), then I am potentially risking overfitting. Is that correct? (Because I think your ensemble algorithm will always look for the best parameters to give the highest ensemble score). Simple models like logistic regression are less likely to overfit, I guess, but it's still possible.\n\nSo should I save a subset of the original training data which I never use when training all my models and use this as an indicator whether my final ensemble overfits? I believe technically the public test dataset serves this purpose, but one can only make 5 submission which seems a bit limited.\n\nThat said, if I did save a subset for the aforementioned purpose, I would be using fewer training samples compared with other people during normal training, which might lead to underperformance?\n\nReally curious and would appreciate your thoughts on this. Thanks so much in advance!",
    "1240846": "The standard approach to avoid overfitting when doing the next level of combining models is to, again, use cross-validation to set hyperparameters. E.g. you would not just take all the out-of-fold predictions of your models and find the averaging weights that optimize the CV score across all out of folds. That can work decently, if there's a huge amount of data, but the smaller the data are, the more likely this is overfitting to the out-of-fold predictions. \n\nYou would use CV to instead find the level of penalization (e.g. penalizing how much you go away from a simple average) that optimizes out-of-fold predictions: i.e. you optimize on the original training fold with some penalization, then apply that averaging approach to the original validation fold and see what's the best hyperparameters for that.\n\nIt's my impression is that this approach is pretty successful on Kaggle and that this way one largely gets away with not having an extra independent validation set (and to some extent the public LB serves as that). \"Gets away with\" meaning \"it tends to beat the alternatives on the private LB\", not we have a completely unbiased assessment of model performance via the CV, of course.",
    "1240879": "Thanks a lot for your reply! That clarifies things for me. I just have two follow up qustions:\n\n1. When I make CV again for my OOF predictions, should I use stratified CV? i.e. if my OOF predictions have 5000 entries (~1000 from each training fold - 5 folds in total), should my next-level CV have ~200 entries from each of these 5 folds for a total of 1000 entries per fold?\n\n2. Can you specify some penalization methods you mentioned? Does l2 regularization in a logistic regression model count?\n\nThanks again!",
    "1240896": "With CV, I actually mean re-using the original CV at the next level of ensembling/stacking. This is a [nice blog post](https://datasciblog.github.io/2016/12/27/a-kagglers-guide-to-model-stacking-in-practice/) on the topic that's pretty readable (what I mention is not the only option discussed, but what I describe is discussed in [this section](https://datasciblog.github.io/2016/12/27/a-kagglers-guide-to-model-stacking-in-practice/)). There's of course also a good bit of discussion around on the Kaggle discussion forums (e.g. [this](https://www.kaggle.com/general/18793)).\n\nCertainly, the parameter controlling the strength of L2 (and/or L1) regularization could be one sensible hyperparameter for regularizing/penalizing a model.\n\nIt's of course an interesting question how much of an overfitting risk there is in this competition, if one does not take this approach. Given that I got quite different answers on how to average depending on how I approached it, I think it may matter to some extent. If the shake-up favors me, I'll take that as an indication that I got it right, if not...",
    "1241435": "most of the time, you should."
  },
  "source": "meta"
}