{
  "id": 175614,
  "title": "How To CV and How To Ensemble OOF Files",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175614",
  "author_name": "Chris Deotte",
  "post_date": "2020-08-18T19:17:58.140000",
  "votes": 360,
  "comment_count": 89,
  "views": 0,
  "content": "<p>One thing I learned at Kaggle is how to ensemble models using OOF files. This was very helpful in Melanoma Competition. I would like to share a simple procedure below called <code>hill climbing</code> and I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private\" target=\"_blank\">here</a>.  </p>\n<h1>Cross Validation</h1>\n<p>The acronym CV refers to cross validation. We start with the full training dataset picture below to the far left. Next, we divide it into 5 subsets, called  <code>Fold 1, Fold 2, Fold 3, Fold 4, Fold 5</code>. Then we train 5 models. We train our first model using data from Folds 2-5 and predict Fold 1. Next we train model 2 using Folds 1, 3, 4, 5 and predict 2. Next 1, 2, 4, 5 and predict 3, etc etc.</p>\n<p>Afterward we have predictions for every training image. This compete set of predictions is called OOF, \"out of fold\" predictions. It is a good practice to save these predictions for every model you build during a competition as <code>oof.csv</code>.</p>\n<p>The CV score (or OOF AUC) is then calculated with <code>OOF_AUC = roc_auc_score( train.target, oof.prediction)</code>. And this is the best indicator of how your model performs. It is a better indicator than public LB.</p>\n<p>During the competition, every model should use the same folds. This is accomplished by using the same seed with <code>sklearn.model_selection.KFold(n_splits = 5, shuffle = True, random_seed = 42)</code></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F79668c5b80f69cbe78b5ee91deb6798b%2Fkfold.png?generation=1597776677559528&amp;alt=media\" alt=\"\"></p>\n<h1>Submission Files</h1>\n<p>For each of the 5 fold models above, we predict the test images. Therefore we have 5 predictions for each test image. We take the average of these 5 sets of predictions and this is our <code>submission.csv</code> file that we submit to Kaggle. When you submit this to Kaggle, your LB score should be similar to your CV score.</p>\n<h1>Model Ensemble</h1>\n<p>Now say that you build 2 models (that means that you did 5 KFold twice). You now have <code>oof_1.csv</code>, <code>oof_2.csv</code>, <code>sub_1.csv</code>, and <code>sub_2.csv</code>. How do we blend the two models?</p>\n<p>We find the weight <code>w</code> such that <code>w * oof_1.predictions + (1-w) * oof_2.predictions</code> has the largest AUC.</p>\n<pre><code> all = []\n for w in [0.00, 0.01, 0.02, ..., 0.98, 0.99, 1.00]:\n     ensemble_pred = w * oof_1.predictions + (1-w) * oof_2.predictions\n     ensemble_auc = roc_auc_score( oof.target , ensemble_pred )\n     all.append( ensemble_auc )\n best_weight = np.argmax( all ) / 100.\n</code></pre>\n<p>Then our submission to kaggle will be</p>\n<pre><code> kaggle_sub = best_weight * sub_1.target + (1-best_weight) * sub_2.target\n</code></pre>\n<h1>Starter Notebook</h1>\n<p>After weeks working on a competition, we will have more than 2 models. So we need a more sophisticated approach than a single for-loop. </p>\n<p>The simplest approach is <code>hill climbing</code> (or <code>forward selection</code>). Start with the one model that has highest CV score. Next iterate through all your additional models and find the one model that combines with the first model to generate the highest two model ensemble CV score. Then search for the best third model. Repeat until ensemble CV does not increase anymore.</p>\n<p>I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private\" target=\"_blank\">here</a> which includes all my <code>oof.csv</code> and <code>sub.csv</code> for this Melanoma comp. There are 39 models. Forward selection chooses 8 of them and the resultant ensemble has OOF CV 0.950, Public LB 0.958, and Private LB 0.942</p>",
  "messages": [
    {
      "id": 976295,
      "postDate": "2020-08-18T19:17:58.140Z",
      "content": "<p>One thing I learned at Kaggle is how to ensemble models using OOF files. This was very helpful in Melanoma Competition. I would like to share a simple procedure below called <code>hill climbing</code> and I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private\" target=\"_blank\">here</a>.  </p>\n<h1>Cross Validation</h1>\n<p>The acronym CV refers to cross validation. We start with the full training dataset picture below to the far left. Next, we divide it into 5 subsets, called  <code>Fold 1, Fold 2, Fold 3, Fold 4, Fold 5</code>. Then we train 5 models. We train our first model using data from Folds 2-5 and predict Fold 1. Next we train model 2 using Folds 1, 3, 4, 5 and predict 2. Next 1, 2, 4, 5 and predict 3, etc etc.</p>\n<p>Afterward we have predictions for every training image. This compete set of predictions is called OOF, \"out of fold\" predictions. It is a good practice to save these predictions for every model you build during a competition as <code>oof.csv</code>.</p>\n<p>The CV score (or OOF AUC) is then calculated with <code>OOF_AUC = roc_auc_score( train.target, oof.prediction)</code>. And this is the best indicator of how your model performs. It is a better indicator than public LB.</p>\n<p>During the competition, every model should use the same folds. This is accomplished by using the same seed with <code>sklearn.model_selection.KFold(n_splits = 5, shuffle = True, random_seed = 42)</code></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F79668c5b80f69cbe78b5ee91deb6798b%2Fkfold.png?generation=1597776677559528&amp;alt=media\" alt=\"\"></p>\n<h1>Submission Files</h1>\n<p>For each of the 5 fold models above, we predict the test images. Therefore we have 5 predictions for each test image. We take the average of these 5 sets of predictions and this is our <code>submission.csv</code> file that we submit to Kaggle. When you submit this to Kaggle, your LB score should be similar to your CV score.</p>\n<h1>Model Ensemble</h1>\n<p>Now say that you build 2 models (that means that you did 5 KFold twice). You now have <code>oof_1.csv</code>, <code>oof_2.csv</code>, <code>sub_1.csv</code>, and <code>sub_2.csv</code>. How do we blend the two models?</p>\n<p>We find the weight <code>w</code> such that <code>w * oof_1.predictions + (1-w) * oof_2.predictions</code> has the largest AUC.</p>\n<pre><code> all = []\n for w in [0.00, 0.01, 0.02, ..., 0.98, 0.99, 1.00]:\n     ensemble_pred = w * oof_1.predictions + (1-w) * oof_2.predictions\n     ensemble_auc = roc_auc_score( oof.target , ensemble_pred )\n     all.append( ensemble_auc )\n best_weight = np.argmax( all ) / 100.\n</code></pre>\n<p>Then our submission to kaggle will be</p>\n<pre><code> kaggle_sub = best_weight * sub_1.target + (1-best_weight) * sub_2.target\n</code></pre>\n<h1>Starter Notebook</h1>\n<p>After weeks working on a competition, we will have more than 2 models. So we need a more sophisticated approach than a single for-loop. </p>\n<p>The simplest approach is <code>hill climbing</code> (or <code>forward selection</code>). Start with the one model that has highest CV score. Next iterate through all your additional models and find the one model that combines with the first model to generate the highest two model ensemble CV score. Then search for the best third model. Repeat until ensemble CV does not increase anymore.</p>\n<p>I posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private\" target=\"_blank\">here</a> which includes all my <code>oof.csv</code> and <code>sub.csv</code> for this Melanoma comp. There are 39 models. Forward selection chooses 8 of them and the resultant ensemble has OOF CV 0.950, Public LB 0.958, and Private LB 0.942</p>",
      "rawMarkdown": "One thing I learned at Kaggle is how to ensemble models using OOF files. This was very helpful in Melanoma Competition. I would like to share a simple procedure below called `hill climbing` and I posted a starter notebook [here][1].  \n\n# Cross Validation\nThe acronym CV refers to cross validation. We start with the full training dataset picture below to the far left. Next, we divide it into 5 subsets, called  `Fold 1, Fold 2, Fold 3, Fold 4, Fold 5`. Then we train 5 models. We train our first model using data from Folds 2-5 and predict Fold 1. Next we train model 2 using Folds 1, 3, 4, 5 and predict 2. Next 1, 2, 4, 5 and predict 3, etc etc.\n\nAfterward we have predictions for every training image. This compete set of predictions is called OOF, \"out of fold\" predictions. It is a good practice to save these predictions for every model you build during a competition as `oof.csv`.\n\nThe CV score (or OOF AUC) is then calculated with `OOF_AUC = roc_auc_score( train.target, oof.prediction)`. And this is the best indicator of how your model performs. It is a better indicator than public LB.\n\nDuring the competition, every model should use the same folds. This is accomplished by using the same seed with `sklearn.model_selection.KFold(n_splits = 5, shuffle = True, random_seed = 42)`\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F79668c5b80f69cbe78b5ee91deb6798b%2Fkfold.png?generation=1597776677559528&alt=media)\n\n# Submission Files\nFor each of the 5 fold models above, we predict the test images. Therefore we have 5 predictions for each test image. We take the average of these 5 sets of predictions and this is our `submission.csv` file that we submit to Kaggle. When you submit this to Kaggle, your LB score should be similar to your CV score.\n\n# Model Ensemble\nNow say that you build 2 models (that means that you did 5 KFold twice). You now have `oof_1.csv`, `oof_2.csv`, `sub_1.csv`, and `sub_2.csv`. How do we blend the two models?\n\nWe find the weight `w` such that `w * oof_1.predictions + (1-w) * oof_2.predictions` has the largest AUC.\n\n     all = []\n     for w in [0.00, 0.01, 0.02, ..., 0.98, 0.99, 1.00]:\n         ensemble_pred = w * oof_1.predictions + (1-w) * oof_2.predictions\n         ensemble_auc = roc_auc_score( oof.target , ensemble_pred )\n         all.append( ensemble_auc )\n     best_weight = np.argmax( all ) / 100.\n\nThen our submission to kaggle will be\n\n     kaggle_sub = best_weight * sub_1.target + (1-best_weight) * sub_2.target\n\n# Starter Notebook\nAfter weeks working on a competition, we will have more than 2 models. So we need a more sophisticated approach than a single for-loop. \n\nThe simplest approach is `hill climbing` (or `forward selection`). Start with the one model that has highest CV score. Next iterate through all your additional models and find the one model that combines with the first model to generate the highest two model ensemble CV score. Then search for the best third model. Repeat until ensemble CV does not increase anymore.\n\nI posted a starter notebook [here][1] which includes all my `oof.csv` and `sub.csv` for this Melanoma comp. There are 39 models. Forward selection chooses 8 of them and the resultant ensemble has OOF CV 0.950, Public LB 0.958, and Private LB 0.942\n\n[1]: https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private",
      "votes": 359
    },
    {
      "id": 976302,
      "postDate": "2020-08-18T19:21:45.140Z",
      "content": "<p>Kagglers refer to this as <code>hill climbing</code> just as a side note if others see this phrase somewhere else.</p>",
      "rawMarkdown": "Kagglers refer to this as `hill climbing` just as a side note if others see this phrase somewhere else.",
      "votes": 13,
      "replies": [
        {
          "id": 976313,
          "postDate": "2020-08-18T19:29:00.980Z",
          "content": "<p>thanks. I updated my post.</p>",
          "rawMarkdown": "thanks. I updated my post.",
          "votes": 3
        },
        {
          "id": 984401,
          "postDate": "2020-08-25T04:46:25.477Z",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> When you and others use <code>hill climbing</code>, do you continue adding models as long as CV increases, or do you only add a new model if it increases CV by some tolerance?</p>\n<p>For my submission, i used <code>TOL = 0.0003</code> thinking I would prevent noise. But if you compare <code>hill climbing</code> to XGB stacking, then XGB uses <code>TOL = 0</code>, right? XGB just keeps adding as long as CV increases.</p>",
          "rawMarkdown": "@philippsinger When you and others use `hill climbing`, do you continue adding models as long as CV increases, or do you only add a new model if it increases CV by some tolerance?\n\nFor my submission, i used `TOL = 0.0003` thinking I would prevent noise. But if you compare `hill climbing` to XGB stacking, then XGB uses `TOL = 0`, right? XGB just keeps adding as long as CV increases.",
          "votes": 2
        },
        {
          "id": 984824,
          "postDate": "2020-08-25T10:07:07.807Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I usually apply it for a certain number of iterations. To avoid overfitting you can also add certain restrictions, like a single model can only be added once or twice.</p>",
          "rawMarkdown": "@cdeotte I usually apply it for a certain number of iterations. To avoid overfitting you can also add certain restrictions, like a single model can only be added once or twice.",
          "votes": 1
        }
      ]
    },
    {
      "id": 979176,
      "postDate": "2020-08-20T17:12:15.270Z",
      "content": "<p>Thanks for sharing this! Good explanation for CV!</p>\n<p>After reading some discussions in the past few days after the competition end, I gradually realized that what more people need is not the top solutions, but the basic ML knowledge, just like this posts.</p>\n<p>I see many people still don't understand why they should not trust public LB score (especially in this comp). But I don’t know how to explain to them in a simple and understandable way.</p>\n<p>I think your post can be a good explanation for it as well. Thanks again!</p>",
      "rawMarkdown": "Thanks for sharing this! Good explanation for CV!\n\nAfter reading some discussions in the past few days after the competition end, I gradually realized that what more people need is not the top solutions, but the basic ML knowledge, just like this posts.\n\nI see many people still don't understand why they should not trust public LB score (especially in this comp). But I don’t know how to explain to them in a simple and understandable way.\n\nI think your post can be a good explanation for it as well. Thanks again!",
      "votes": 14,
      "replies": [
        {
          "id": 979203,
          "postDate": "2020-08-20T17:31:26.207Z",
          "content": "<p>Thanks Qishen. I updated the discussion title to `\"How To CV and How To Ensemble OOF\". I thought i was just explaining ensemble, but you are right, this is a nice explanation of CV too.</p>\n<p>Congrats again on you and your team's awesome model with amazing 0.960 CV!</p>",
          "rawMarkdown": "Thanks Qishen. I updated the discussion title to `\"How To CV and How To Ensemble OOF\". I thought i was just explaining ensemble, but you are right, this is a nice explanation of CV too.\n\nCongrats again on you and your team's awesome model with amazing 0.960 CV!",
          "votes": 7
        }
      ]
    },
    {
      "id": 1654315,
      "postDate": "2022-01-18T11:35:17.327Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, your Hill Climbing approach is a great trick. I'll for sure try it out in my next competition.<br>\nI have small doubt, in the post you've linked you say </p>\n<blockquote>\n  <p>During the competition, every model should use the same folds. </p>\n</blockquote>\n<p>Does that really matter when your oof predictions have all the train samples anyway? I specifically ask in a context of ensembling models as a team, where different members worked with different splits.</p>",
      "rawMarkdown": "Thanks @cdeotte, your Hill Climbing approach is a great trick. I'll for sure try it out in my next competition.\nI have small doubt, in the post you've linked you say \n>During the competition, every model should use the same folds. \n\nDoes that really matter when your oof predictions have all the train samples anyway? I specifically ask in a context of ensembling models as a team, where different members worked with different splits.",
      "votes": 2
    },
    {
      "id": 988805,
      "postDate": "2020-08-28T09:42:28.270Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, have you seen this paper \"On Cross-Validation and Stacking: Building seemingly predictive models on random data\"?<br>\n<a href=\"https://www.kdd.org/exploration_files/v12-02-4-UR-Perlich.pdf\" target=\"_blank\">https://www.kdd.org/exploration_files/v12-02-4-UR-Perlich.pdf</a> </p>\n<p>I'd be interested to know what you and other Kagglers think about it</p>",
      "rawMarkdown": "Hi @cdeotte, have you seen this paper \"On Cross-Validation and Stacking: Building seemingly predictive models on random data\"?\nhttps://www.kdd.org/exploration_files/v12-02-4-UR-Perlich.pdf \n\nI'd be interested to know what you and other Kagglers think about it",
      "votes": 3,
      "replies": [
        {
          "id": 989300,
          "postDate": "2020-08-28T17:42:23.213Z",
          "content": "<p>Thanks for the paper, i'll check it out and post my reaction here.</p>",
          "rawMarkdown": "Thanks for the paper, i'll check it out and post my reaction here.",
          "votes": 4
        }
      ]
    },
    {
      "id": 979325,
      "postDate": "2020-08-20T19:03:15.597Z",
      "content": "<p>Hi Chris, the explanation is very clear but I have a question and after reading the post a couple of times I can't figure out the answer on my own. Talking about the submission file, you say</p>\n<blockquote>\n  <p>For each of the 5 fold models above, we predict the test images. Therefore we have 5 predictions for each test image. We take the average of these 5 sets of predictions and this is our submission.csv file that we submit to Kaggle. When you submit this to Kaggle, your LB score should be similar to your CV score.</p>\n</blockquote>\n<p>and this makes perfect sense, but why do you never train a model using all train data once you have a high CV score for that model?  <br>\nIs it because you would not know how it performs when trained with an additional fold, making all the work useless and destroying the ensembling process?  <br>\nI am asking because so far, when I perform cross-validation, I always fit the model on all train data with the best parameters to make the test predictions and so your way of doing it by averaging is completely new to me.<br>\nThank you!</p>",
      "rawMarkdown": "Hi Chris, the explanation is very clear but I have a question and after reading the post a couple of times I can't figure out the answer on my own. Talking about the submission file, you say\n> For each of the 5 fold models above, we predict the test images. Therefore we have 5 predictions for each test image. We take the average of these 5 sets of predictions and this is our submission.csv file that we submit to Kaggle. When you submit this to Kaggle, your LB score should be similar to your CV score.\n\nand this makes perfect sense, but why do you never train a model using all train data once you have a high CV score for that model?  \nIs it because you would not know how it performs when trained with an additional fold, making all the work useless and destroying the ensembling process?  \nI am asking because so far, when I perform cross-validation, I always fit the model on all train data with the best parameters to make the test predictions and so your way of doing it by averaging is completely new to me.\nThank you!",
      "votes": 4,
      "replies": [
        {
          "id": 979396,
          "postDate": "2020-08-20T19:49:16.223Z",
          "content": "<p>Doing it as I describe associates the CV to LB score. Because the identical models that produced the CV score are the same models producing the LB score. (When you run twice as your describe, you don't have this guarantee). That being said, I sometimes do what you suggest. Both ways have pros and cons.</p>\n<p>Note that the way I describe in this post is a form of bagging similar to a random forest. Each of the 5 fold models are training on a random 80% of the train data. Therefore even if each of the five fold models overfits their 80% train data, the ensemble of the 5 folds will not overfit the entire 100% train data. So predicting test using CV is also a way to prevent overfitting the train data (i.e. generalizing better to test data) in the same way that a random forest uses bagging to prevent overfitting train data.</p>",
          "rawMarkdown": "Doing it as I describe associates the CV to LB score. Because the identical models that produced the CV score are the same models producing the LB score. (When you run twice as your describe, you don't have this guarantee). That being said, I sometimes do what you suggest. Both ways have pros and cons.\n\nNote that the way I describe in this post is a form of bagging similar to a random forest. Each of the 5 fold models are training on a random 80% of the train data. Therefore even if each of the five fold models overfits their 80% train data, the ensemble of the 5 folds will not overfit the entire 100% train data. So predicting test using CV is also a way to prevent overfitting the train data (i.e. generalizing better to test data) in the same way that a random forest uses bagging to prevent overfitting train data.",
          "votes": 5
        },
        {
          "id": 979418,
          "postDate": "2020-08-20T20:31:06.320Z",
          "content": "<p>Ok, got it, thank you so much.  <br>\nBy doing this, you are training on less data but as a tradeoff you have a more reliable estimate of the performance of your model and also a free ensemble of 5 classifiers which is never a bad thing.  <br>\nI guess I'll try to apply this method in House Prices competition when I have my final model ready!</p>",
          "rawMarkdown": "Ok, got it, thank you so much.  \nBy doing this, you are training on less data but as a tradeoff you have a more reliable estimate of the performance of your model and also a free ensemble of 5 classifiers which is never a bad thing.  \nI guess I'll try to apply this method in House Prices competition when I have my final model ready!",
          "votes": 1
        }
      ]
    },
    {
      "id": 976417,
      "postDate": "2020-08-18T20:43:15.087Z",
      "content": "<p>The strategy I usually go for is finding a set of models that give the best CV when averaged together.</p>\n<p>Do you have any idea how your ensemble performs (especially on the leaderboard) when giving equal weights to every model ? <br>\nWhat about using the 8  (or any number of) models that have the best CV when averging the predictions together ? </p>\n<p>This would be nice to see how much hill climbing actually helps.</p>",
      "rawMarkdown": "The strategy I usually go for is finding a set of models that give the best CV when averaged together.\n\nDo you have any idea how your ensemble performs (especially on the leaderboard) when giving equal weights to every model ? \nWhat about using the 8  (or any number of) models that have the best CV when averging the predictions together ? \n\nThis would be nice to see how much hill climbing actually helps.",
      "votes": 4,
      "replies": [
        {
          "id": 976425,
          "postDate": "2020-08-18T20:53:24.477Z",
          "content": "<p>Great suggestions Theo. I will investigate this.</p>\n<p>I should probably change the language in my post. I should be more clear that high climbing is only one way and there are many others like the ones you suggest.</p>",
          "rawMarkdown": "Great suggestions Theo. I will investigate this.\n\nI should probably change the language in my post. I should be more clear that high climbing is only one way and there are many others like the ones you suggest.",
          "votes": 1
        },
        {
          "id": 976447,
          "postDate": "2020-08-18T21:22:20.207Z",
          "content": "<blockquote>\n  <p>I should probably change the language in my post. I should be more clear that high climbing is only one way and there are many others like the ones you suggest.</p>\n</blockquote>\n<p>No worries, this was clear to me :)</p>",
          "rawMarkdown": "> I should probably change the language in my post. I should be more clear that high climbing is only one way and there are many others like the ones you suggest.\n\nNo worries, this was clear to me :)",
          "votes": 1
        },
        {
          "id": 976452,
          "postDate": "2020-08-18T21:29:34.490Z",
          "content": "<p>Interesting! thanks for sharing. </p>\n<ol>\n<li>In my case, I used a kernel compute combined OOF for 6 models. </li>\n<li>In the Kernel 4 types of combined OOF are computed. </li>\n<li>OOF_avg_auc, OOF_rank_auc, OOF_pow_auc and OOF_bo_auc</li>\n<li>My heuristics are: I should pick the combination which has :</li>\n<li>highest OOF_bo_auc</li>\n<li>Or pick model which has highest average of all 4 OOF. (since if it is a good model, all OOF should be high ? )</li>\n<li>My top combinations were: </li>\n</ol>\n<table>\n<thead>\n<tr>\n<th>no_of_models</th>\n<th>OOF_avg_auc</th>\n<th>OOF_rank_auc</th>\n<th>OOF_pow_auc</th>\n<th>OOF_bo_auc</th>\n<th>Overall_mean</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>8</td>\n<td>0.93898</td>\n<td>0.93993</td>\n<td>0.937304</td>\n<td>0.93963</td>\n<td>0.93963</td>\n<td>0.9522</td>\n<td>0.9409</td>\n</tr>\n<tr>\n<td>8</td>\n<td>0.93787</td>\n<td>0.93806</td>\n<td>0.93652</td>\n<td>0.93918</td>\n<td>0.93791</td>\n<td>0.9518</td>\n<td>0.9411</td>\n</tr>\n<tr>\n<td>6</td>\n<td>0.938307</td>\n<td>0.938737</td>\n<td>0.93630</td>\n<td>0.93900</td>\n<td>0.93808</td>\n<td>0.9512</td>\n<td>0.9413</td>\n</tr>\n</tbody>\n</table>\n<ol>\n<li>I have a total of 12 models. </li>\n<li>So for ensemble of 6 models i thought i have to check 12<em>11</em>10<em>8</em>9*7 (== 12! - 6! ) combinations. </li>\n<li>As I understand, hill climbing will help to avoid checking 12<em>11</em>10<em>8</em>9*7 combinations. I will try out hill climbing for my models. </li>\n</ol>",
          "rawMarkdown": "Interesting! thanks for sharing. \n1. In my case, I used a kernel compute combined OOF for 6 models. \n2. In the Kernel 4 types of combined OOF are computed. \n3. OOF_avg_auc, OOF_rank_auc, OOF_pow_auc and OOF_bo_auc\n4. My heuristics are: I should pick the combination which has :\n5. highest OOF_bo_auc\n6. Or pick model which has highest average of all 4 OOF. (since if it is a good model, all OOF should be high ? )\n7. My top combinations were: \n\n| no_of_models | OOF_avg_auc |OOF_rank_auc|OOF_pow_auc|OOF_bo_auc|Overall_mean|Public LB| Private LB|\n| ---------------|---------------|---------------|--------------|--------------|-------------|----------|-----------|\n| 8                      | 0.93898|0.93993|0.937304|0.93963|0.93963|0.9522|0.9409|\n|8\t                   |0.93787|0.93806|0.93652|0.93918|0.93791|0.9518|0.9411|\n| 6                      |0.938307| 0.938737|\t0.93630|\t0.93900| 0.93808|0.9512|0.9413|\n\n1. I have a total of 12 models. \n2. So for ensemble of 6 models i thought i have to check 12*11*10*8*9*7 (== 12! - 6! ) combinations. \n3. As I understand, hill climbing will help to avoid checking 12*11*10*8*9*7 combinations. I will try out hill climbing for my models. ",
          "votes": 1
        },
        {
          "id": 976459,
          "postDate": "2020-08-18T21:50:44.220Z",
          "content": "<blockquote>\n  <p>What about using the 8 (or any number of) models that have the best CV when averging the predictions together ?</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> I just averaged all 39 models and it performed identical to <code>hill climbing</code>. Also I randomly chose 8 repeatedly 10,000 times and just averaged them. Then used the 8 with best CV. That also performed the same. So maybe everything just does the same haha.</p>\n<p>I know if you only have 2 or 3 models, then finding weights is better than just averaging them but perhaps when you have dozens of models just averaging works as well as finding weights.</p>\n<p>(I've never had so many models before so I don't have much experience what is best when you have dozens of models)</p>",
          "rawMarkdown": "> What about using the 8 (or any number of) models that have the best CV when averging the predictions together ?\n\n@theoviel I just averaged all 39 models and it performed identical to `hill climbing`. Also I randomly chose 8 repeatedly 10,000 times and just averaged them. Then used the 8 with best CV. That also performed the same. So maybe everything just does the same haha.\n\nI know if you only have 2 or 3 models, then finding weights is better than just averaging them but perhaps when you have dozens of models just averaging works as well as finding weights.\n\n(I've never had so many models before so I don't have much experience what is best when you have dozens of models)",
          "votes": 6
        },
        {
          "id": 976462,
          "postDate": "2020-08-18T21:54:49.323Z",
          "content": "<p>But i guess there is an advantage to hill climbing or your select 8 method. In the end, you will have fewer models in your ensemble which is usually a good thing.</p>",
          "rawMarkdown": "But i guess there is an advantage to hill climbing or your select 8 method. In the end, you will have fewer models in your ensemble which is usually a good thing.",
          "votes": 2
        },
        {
          "id": 977039,
          "postDate": "2020-08-19T08:52:27.083Z",
          "content": "<p>Interesting, thanks a lot for checking </p>",
          "rawMarkdown": "Interesting, thanks a lot for checking ",
          "votes": 1
        },
        {
          "id": 983075,
          "postDate": "2020-08-24T03:09:30.487Z",
          "content": "<p>I discovered that if i set <code>TOL = 0</code> to allow hill climbing to add any new model that increases CV AUC greater than 0, then the CV and private LB will become higher than averaging all or using a random 8. So perhaps to get the most out of high climbing, you need to set <code>TOL = 0</code>. I'm still unsure what is the optimal value because if an additional model only increase CV AUC 0.0001 it may just be random noise. (That's why I initially choose <code>TOL = 0.0003</code> in my hill climbing notebook. Perhaps I should analyze this mathematically).</p>",
          "rawMarkdown": "I discovered that if i set `TOL = 0` to allow hill climbing to add any new model that increases CV AUC greater than 0, then the CV and private LB will become higher than averaging all or using a random 8. So perhaps to get the most out of high climbing, you need to set `TOL = 0`. I'm still unsure what is the optimal value because if an additional model only increase CV AUC 0.0001 it may just be random noise. (That's why I initially choose `TOL = 0.0003` in my hill climbing notebook. Perhaps I should analyze this mathematically).",
          "votes": 1
        }
      ]
    },
    {
      "id": 998817,
      "postDate": "2020-09-05T05:40:33.813Z",
      "content": "<p>Suppose we are using a model and I have to choose the best hyperparameters for that model. Now, will it be right that for each hyperparameter I do a new cv and note down the best parameters? By doing this am I introducing any leakage?</p>",
      "rawMarkdown": "Suppose we are using a model and I have to choose the best hyperparameters for that model. Now, will it be right that for each hyperparameter I do a new cv and note down the best parameters? By doing this am I introducing any leakage?",
      "votes": 1,
      "replies": [
        {
          "id": 998899,
          "postDate": "2020-09-05T07:15:17.133Z",
          "content": "<p>That is not right in my opinion since the hyperparameters like this themselves might not give the best result in tandem. Choose the cv set with the least loss value and work on that but remember changing the hyperparameters too much would make it biased to your cv set and then might fail at generalizing. There is always another way to make sure, try it yourself once. I hope this helps : ).</p>",
          "rawMarkdown": "That is not right in my opinion since the hyperparameters like this themselves might not give the best result in tandem. Choose the cv set with the least loss value and work on that but remember changing the hyperparameters too much would make it biased to your cv set and then might fail at generalizing. There is always another way to make sure, try it yourself once. I hope this helps : ).",
          "votes": 1
        }
      ]
    },
    {
      "id": 997407,
      "postDate": "2020-09-04T03:09:53.400Z",
      "content": "<p>Thanks for sharing. I do follow these approaches on Stacking while working on the projects.<br>\nHave you tried Voting Classifiers as well?</p>",
      "rawMarkdown": "Thanks for sharing. I do follow these approaches on Stacking while working on the projects.\nHave you tried Voting Classifiers as well?",
      "votes": 1
    },
    {
      "id": 988133,
      "postDate": "2020-08-27T19:59:22.867Z",
      "content": "<p>A good explaination, was very helpful.</p>",
      "rawMarkdown": "A good explaination, was very helpful.",
      "votes": 1
    },
    {
      "id": 987504,
      "postDate": "2020-08-27T10:04:31.347Z",
      "content": "<p>Good one</p>",
      "rawMarkdown": "Good one",
      "votes": 1
    },
    {
      "id": 983330,
      "postDate": "2020-08-24T07:37:11.813Z",
      "content": "<p>Very crisp, clear &amp; to the point explanation. I am a novice at kaggle and I am learning amazing stuff reading your notebook &amp; discussion. Thanks!</p>",
      "rawMarkdown": "Very crisp, clear & to the point explanation. I am a novice at kaggle and I am learning amazing stuff reading your notebook & discussion. Thanks!",
      "votes": 1
    },
    {
      "id": 983068,
      "postDate": "2020-08-24T03:03:32.770Z",
      "content": "<p>Thanks for sharing great knowledge about ensembling! </p>\n<p>I didn't know hill climbing approach to ensemble the model and helped me a lot! </p>",
      "rawMarkdown": "Thanks for sharing great knowledge about ensembling! \n\nI didn't know hill climbing approach to ensemble the model and helped me a lot! ",
      "votes": 1
    },
    {
      "id": 979255,
      "postDate": "2020-08-20T18:06:47.880Z",
      "content": "<p>Because of hardware limitations I have been trying this….</p>\n<p>for each epoch in 1:30<br>\ntrain for 100 steps; in each  step use a batch of n images that is a random sample of the train set with augmentations (resampled at each step)<br>\nvalidate for 30 steps; in each step use a batch of n images that is a random sample of the train set without augmentations (resampled at each step)</p>\n<p>n is 16 32 64 depending on image size</p>\n<p>Different runs for the same model got consistent scores between \"this cv score\" and private set… not with public set thas was a bit lower…</p>\n<p>changing the set of images at each step may result in underfitting? it certainly did not overfit</p>",
      "rawMarkdown": "Because of hardware limitations I have been trying this....\n\nfor each epoch in 1:30\ntrain for 100 steps; in each  step use a batch of n images that is a random sample of the train set with augmentations (resampled at each step)\nvalidate for 30 steps; in each step use a batch of n images that is a random sample of the train set without augmentations (resampled at each step)\n\nn is 16 32 64 depending on image size\n\nDifferent runs for the same model got consistent scores between \"this cv score\" and private set... not with public set thas was a bit lower...\n\nchanging the set of images at each step may result in underfitting? it certainly did not overfit",
      "votes": 1,
      "replies": [
        {
          "id": 979275,
          "postDate": "2020-08-20T18:23:29.133Z",
          "content": "<p>Nice trick Marcelo. </p>\n<p>Note that if a full epoch has 1000 steps and you are only training on 100 steps, then after 10 of your epochs, you will do the same thing as 1 of the original epochs. The difference would be that you are changing the learning rate more often than it would occur in the original epoch.</p>\n<p>Another trick to speed up training is to train on random crops. So if you wish to train a model on 384x384. Then you randomly crop each train image as 192x192 (as data augmentation). Then after training for a while your model will see the entire 384x384 image, but you train fast because your are only processing 192x192 </p>",
          "rawMarkdown": "Nice trick Marcelo. \n\nNote that if a full epoch has 1000 steps and you are only training on 100 steps, then after 10 of your epochs, you will do the same thing as 1 of the original epochs. The difference would be that you are changing the learning rate more often than it would occur in the original epoch.\n\nAnother trick to speed up training is to train on random crops. So if you wish to train a model on 384x384. Then you randomly crop each train image as 192x192 (as data augmentation). Then after training for a while your model will see the entire 384x384 image, but you train fast because your are only processing 192x192 ",
          "votes": 1
        },
        {
          "id": 979677,
          "postDate": "2020-08-21T03:55:37.117Z",
          "content": "<p>That seems very interesting! Can an EffNet model trained with one image size tranfer weights to an EffNet model to predict a larger image size? or it only applies to simpler models?</p>",
          "rawMarkdown": "That seems very interesting! Can an EffNet model trained with one image size tranfer weights to an EffNet model to predict a larger image size? or it only applies to simpler models?",
          "votes": 1
        },
        {
          "id": 979693,
          "postDate": "2020-08-21T04:14:32.493Z",
          "content": "<p>Yes. You can train an EffNet with 192x192 and then predict with 384x384. (This is what I did in Steel Comp image segmetation, explained <a href=\"https://www.kaggle.com/c/severstal-steel-defect-detection/discussion/114321\" target=\"_blank\">here</a>). You can even train where each batch has a different size. One batch can be 256x256 and another batch can be 512x512. An EffNet model doesn't care about input size. An EffNet is just a bunch of convolution filters that get trained. Convolution filters are little 3x3 \"windows\" that move over an image.</p>\n<p>However, I'm not suggesting mixing training sizes and inference sizes in Melanoma Comp (even though it is possible but maybe it would be good to investigate). I'm suggesting using 384x384 but during training always show it a different random 192x192 crop. It will then learn to classify images based on just using a 192x192 piece of a full image. Then during inference, you use TTA=21, so you will have your CNN predict on 21 different 192x192 random pieces (of the same image) and then you average those 21 predictions as your final prediction for that single image. (So the model is always training and inferring on 192x192 pieces).</p>",
          "rawMarkdown": "Yes. You can train an EffNet with 192x192 and then predict with 384x384. (This is what I did in Steel Comp image segmetation, explained [here][1]). You can even train where each batch has a different size. One batch can be 256x256 and another batch can be 512x512. An EffNet model doesn't care about input size. An EffNet is just a bunch of convolution filters that get trained. Convolution filters are little 3x3 \"windows\" that move over an image.\n\nHowever, I'm not suggesting mixing training sizes and inference sizes in Melanoma Comp (even though it is possible but maybe it would be good to investigate). I'm suggesting using 384x384 but during training always show it a different random 192x192 crop. It will then learn to classify images based on just using a 192x192 piece of a full image. Then during inference, you use TTA=21, so you will have your CNN predict on 21 different 192x192 random pieces (of the same image) and then you average those 21 predictions as your final prediction for that single image. (So the model is always training and inferring on 192x192 pieces).\n\n[1]: https://www.kaggle.com/c/severstal-steel-defect-detection/discussion/114321",
          "votes": 2
        },
        {
          "id": 980236,
          "postDate": "2020-08-21T12:33:33.543Z",
          "content": "<p>Is it possible to train EffNet on 192x192 images just to learn the weights of the kernels to speed up learning at first and then only retrain the FC layers at the tip of the architecture to predict for 384x384 images at the end of training?  Are there any drawbacks to this approach?</p>",
          "rawMarkdown": "Is it possible to train EffNet on 192x192 images just to learn the weights of the kernels to speed up learning at first and then only retrain the FC layers at the tip of the architecture to predict for 384x384 images at the end of training?  Are there any drawbacks to this approach?"
        },
        {
          "id": 980295,
          "postDate": "2020-08-21T13:24:39.770Z",
          "content": "<p>Thank you very much Sir…<br>\nI found this possibility is very suitable for fine-grained images that are more or less homogeneous. With this approach I could train models with small  cropos much more quickly. I will give it a try with data from <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification\" target=\"_blank\">human protein classification</a> to see how it works …</p>",
          "rawMarkdown": "Thank you very much Sir...\nI found this possibility is very suitable for fine-grained images that are more or less homogeneous. With this approach I could train models with small  cropos much more quickly. I will give it a try with data from [human protein classification](https://www.kaggle.com/c/human-protein-atlas-image-classification) to see how it works ...",
          "votes": 1
        },
        {
          "id": 980521,
          "postDate": "2020-08-21T16:56:44.763Z",
          "content": "<blockquote>\n  <p>Is it possible to train EffNet on 192x192 images just to learn the weights of the kernels to speed up learning at first and then only retrain the FC layers at the tip of the architecture to predict for 384x384 images at the end of training? Are there any drawbacks to this approach?</p>\n</blockquote>\n<p>Yes, you can do this. The best way to see how well it works is try it and evaluate CV LB.</p>",
          "rawMarkdown": "> Is it possible to train EffNet on 192x192 images just to learn the weights of the kernels to speed up learning at first and then only retrain the FC layers at the tip of the architecture to predict for 384x384 images at the end of training? Are there any drawbacks to this approach?\n\nYes, you can do this. The best way to see how well it works is try it and evaluate CV LB."
        },
        {
          "id": 980720,
          "postDate": "2020-08-21T19:52:49.990Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 980721,
          "postDate": "2020-08-21T19:52:49.997Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 980722,
          "postDate": "2020-08-21T19:52:50.007Z",
          "content": "<p>And one more thing: The LB scores are very close to one another (.0005 isn't much especially when you take into consideration that test accuracy isn't true real world accuracy) so there must be some randomness  involved to who might be 1st beyond just skill. So how do hosts deal with this kind of randomness? Do they deem it insignificant? Is it insignificant?</p>",
          "rawMarkdown": "And one more thing: The LB scores are very close to one another (.0005 isn't much especially when you take into consideration that test accuracy isn't true real world accuracy) so there must be some randomness  involved to who might be 1st beyond just skill. So how do hosts deal with this kind of randomness? Do they deem it insignificant? Is it insignificant?",
          "votes": 1
        },
        {
          "id": 980885,
          "postDate": "2020-08-22T01:02:56.293Z",
          "content": "<p>Data science is both skill and luck. For example, I built a model from meta data: age, sex, site, image size. (By itself it had CV 0.77 LB 0.77). When I added it to my ensemble, It increased my CV by 0.0005. The question is do I include it in my final submission with such a small CV increase? How much can I gain, and how much do I risk?</p>\n<p>I simulated 1000 private leaderboards. You can see below that in some simulated private leaderboards, including the meta model (as 90% image 10% meta) will decrease my private LB score as much as AUC 0.002! There is also the chance it could increase my private LB 0.003! </p>\n<p>The expected increase in private LB score is 0.0005 with standard deviation 0.0008. That means that it will increase my score in 75% of private leaderboards. Should I take the risk? The best we can do is calculate expected values, variances, and risk and decide whether to include things in our models.</p>\n<p>I did include it and it lowered my private LB by 0.0002 but there was the chance that it could have increased me to Gold !!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Feb01f9d371f4a49be62093d5d2b638bd%2Fmeta.png?generation=1598057575239798&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Data science is both skill and luck. For example, I built a model from meta data: age, sex, site, image size. (By itself it had CV 0.77 LB 0.77). When I added it to my ensemble, It increased my CV by 0.0005. The question is do I include it in my final submission with such a small CV increase? How much can I gain, and how much do I risk?\n\nI simulated 1000 private leaderboards. You can see below that in some simulated private leaderboards, including the meta model (as 90% image 10% meta) will decrease my private LB score as much as AUC 0.002! There is also the chance it could increase my private LB 0.003! \n\nThe expected increase in private LB score is 0.0005 with standard deviation 0.0008. That means that it will increase my score in 75% of private leaderboards. Should I take the risk? The best we can do is calculate expected values, variances, and risk and decide whether to include things in our models.\n\nI did include it and it lowered my private LB by 0.0002 but there was the chance that it could have increased me to Gold !!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Feb01f9d371f4a49be62093d5d2b638bd%2Fmeta.png?generation=1598057575239798&alt=media)",
          "votes": 2
        },
        {
          "id": 981245,
          "postDate": "2020-08-22T10:11:42.923Z",
          "content": "<p>Thank you so much for investing your time into replying to such a newbie I really appreciate it. Due to the MASSIVE information spread about this field, it's really hard to get yourself around. If it wasn't for people like you, new comers won't have any chance keeping up. Especially if they were self-taught like me. Thank you so much</p>",
          "rawMarkdown": "Thank you so much for investing your time into replying to such a newbie I really appreciate it. Due to the MASSIVE information spread about this field, it's really hard to get yourself around. If it wasn't for people like you, new comers won't have any chance keeping up. Especially if they were self-taught like me. Thank you so much",
          "votes": 2
        },
        {
          "id": 1253672,
          "postDate": "2021-03-26T23:35:26.560Z",
          "content": "<blockquote>\n  <p>I simulated 1000 private leaderboards.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> may I ask you how one could simulate private leaderboards?</p>",
          "rawMarkdown": "> I simulated 1000 private leaderboards.\n\n@cdeotte may I ask you how one could simulate private leaderboards?"
        }
      ]
    },
    {
      "id": 977172,
      "postDate": "2020-08-19T10:35:08.297Z",
      "content": "<p>Thanks a lot for sharing and it is really good explanation for oof ensemble!</p>\n<p>May I ask something:<br>\nSome Kagglers save their folds to <code>folds.csv</code> and just load <code>folds.csv</code> for more experiments.</p>\n<p>If I use same seed(<code>e.g. random_seed = 42</code>), is it same folds result?</p>\n<p><code>new_folds = sklearn.model_selection.KFold(n_splits = 5, shuffle = True, random_seed = 42)</code><br>\n<code>fold.csv == new_folds</code></p>\n<p>Thanks again <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>.</p>",
      "rawMarkdown": "Thanks a lot for sharing and it is really good explanation for oof ensemble!\n\nMay I ask something:\nSome Kagglers save their folds to `folds.csv` and just load `folds.csv` for more experiments.\n\nIf I use same seed(`e.g. random_seed = 42`), is it same folds result?\n\n`new_folds = sklearn.model_selection.KFold(n_splits = 5, shuffle = True, random_seed = 42)`\n`fold.csv == new_folds`\n\nThanks again @cdeotte.",
      "votes": 1,
      "replies": [
        {
          "id": 977791,
          "postDate": "2020-08-19T17:53:07.827Z",
          "content": "<p>Yes. If two different computers use the same <code>random_seed = 42</code> then they will use the same folds. This is how working on a Kaggle team does it. Every team member uses the same seed on their own computer. Then you can compare everyone's CV score and use everyone's OOF to ensemble your final submission.</p>",
          "rawMarkdown": "Yes. If two different computers use the same `random_seed = 42` then they will use the same folds. This is how working on a Kaggle team does it. Every team member uses the same seed on their own computer. Then you can compare everyone's CV score and use everyone's OOF to ensemble your final submission.",
          "votes": 5
        },
        {
          "id": 979216,
          "postDate": "2020-08-20T17:40:45.370Z",
          "content": "<p>Thanks a lot for reply <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> ! <br>\nI want to use what I have learned from you in the next competition. :) </p>",
          "rawMarkdown": "Thanks a lot for reply @cdeotte ! \nI want to use what I have learned from you in the next competition. :) \n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 976954,
      "postDate": "2020-08-19T07:37:31.457Z",
      "content": "<p>What clarity!<br>\nI knew that intuitively, but I wasn't confident enough in my methods to implement it. I want to take advantage of this in the future.</p>",
      "rawMarkdown": "What clarity!\nI knew that intuitively, but I wasn't confident enough in my methods to implement it. I want to take advantage of this in the future.",
      "votes": 1
    },
    {
      "id": 976589,
      "postDate": "2020-08-19T01:07:29.990Z",
      "content": "<p>Simple yet very efficient, thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , for me it seems obvious now, but before this competition I never saved my model's OOF predictions, after seeing your public notebook this is a practice that I will keep from now on!</p>",
      "rawMarkdown": "Simple yet very efficient, thanks @cdeotte , for me it seems obvious now, but before this competition I never saved my model's OOF predictions, after seeing your public notebook this is a practice that I will keep from now on!",
      "votes": 1,
      "replies": [
        {
          "id": 976604,
          "postDate": "2020-08-19T01:19:06.400Z",
          "content": "<p>Yes. For the past 2 months, I save every model's (1) OOF (2) Submission CSV (3) Model weights. These things always come in handy later.</p>\n<p>For example, maybe we want to build a model that reuses 3 models' weights then build a new model using those three backbones, concatenate the outputs and train a new MLP on top of that.</p>\n<p>Or maybe we want to extract embeddings from old model weight backbones, then use XGB to build a head, etc etc.</p>",
          "rawMarkdown": "Yes. For the past 2 months, I save every model's (1) OOF (2) Submission CSV (3) Model weights. These things always come in handy later.\n\nFor example, maybe we want to build a model that reuses 3 models' weights then build a new model using those three backbones, concatenate the outputs and train a new MLP on top of that.\n\nOr maybe we want to extract embeddings from old model weight backbones, then use XGB to build a head, etc etc.",
          "votes": 5
        }
      ]
    },
    {
      "id": 976583,
      "postDate": "2020-08-19T00:52:34.967Z",
      "content": "<p>Could you clarify why you do a for loop instead of fitting a logistic regression or other optimization approaches?</p>",
      "rawMarkdown": "Could you clarify why you do a for loop instead of fitting a logistic regression or other optimization approaches?",
      "votes": 1,
      "replies": [
        {
          "id": 976598,
          "postDate": "2020-08-19T01:15:16.573Z",
          "content": "<p>I've tried logistic regression, linear regression and XGB. For some reason <code>hill climbing</code> achieves a higher ensemble CV. </p>\n<p>(Not sure why. Maybe I'm doing the other methods wrong).</p>",
          "rawMarkdown": "I've tried logistic regression, linear regression and XGB. For some reason `hill climbing` achieves a higher ensemble CV. \n\n(Not sure why. Maybe I'm doing the other methods wrong).",
          "votes": 1
        },
        {
          "id": 976688,
          "postDate": "2020-08-19T03:30:10.147Z",
          "content": "<p>Very interesting. I see many top kagglers use the same approach as yours, so there might be some underlying reasons.</p>\n<p>By the way thank you for all of your hard work! </p>",
          "rawMarkdown": "Very interesting. I see many top kagglers use the same approach as yours, so there might be some underlying reasons.\n\nBy the way thank you for all of your hard work! ",
          "votes": 1
        },
        {
          "id": 983070,
          "postDate": "2020-08-24T03:04:21.437Z",
          "content": "<p>Note that my forward selection notebook <a href=\"https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private\" target=\"_blank\">here</a> uses <code>TOL = 0.0003</code>. I'm still unsure what is the best selection for this variable. (This is CV AUC threshold increase for adding new models and I didn't want to clutter the ensemble with unnecessary additional models). Also I use <code>DUPLICATES = False</code>. That means that after it adds a model, it won't add that model again.</p>\n<p>But now that I think about it, it would be better to allow adding the same model again because that is equivalent to adding more weight to that model's coefficient. (Kind of how logistic regression or linear regression adjusts weights to find optimal).</p>\n<p>So, FYI, if you use <code>TOL = 0</code> and <code>DUPLICATES = True</code> in my public notebook, the CV gets up to 0.9507 and the private LB gets up to 0.9433! Therefore if you try to compare high climbing to logistic regression or linear regression, those are the CV and LB, you need to beat.</p>",
          "rawMarkdown": "Note that my forward selection notebook [here][1] uses `TOL = 0.0003`. I'm still unsure what is the best selection for this variable. (This is CV AUC threshold increase for adding new models and I didn't want to clutter the ensemble with unnecessary additional models). Also I use `DUPLICATES = False`. That means that after it adds a model, it won't add that model again.\n\nBut now that I think about it, it would be better to allow adding the same model again because that is equivalent to adding more weight to that model's coefficient. (Kind of how logistic regression or linear regression adjusts weights to find optimal).\n\nSo, FYI, if you use `TOL = 0` and `DUPLICATES = True` in my public notebook, the CV gets up to 0.9507 and the private LB gets up to 0.9433! Therefore if you try to compare high climbing to logistic regression or linear regression, those are the CV and LB, you need to beat.\n\n[1]: https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private",
          "votes": 1
        }
      ]
    },
    {
      "id": 976448,
      "postDate": "2020-08-18T21:25:41.197Z",
      "content": "<blockquote>\n  <p>Also every model should use the same folds. </p>\n</blockquote>\n<p>Is it necessary to have the same folds? I guess we can have different folds OOF csv files. Will it cause leak?</p>",
      "rawMarkdown": "> Also every model should use the same folds. \n\nIs it necessary to have the same folds? I guess we can have different folds OOF csv files. Will it cause leak?",
      "votes": 1,
      "replies": [
        {
          "id": 976454,
          "postDate": "2020-08-18T21:38:55.690Z",
          "content": "<p>It won't cause a leak but you will get artificial CV gains. If one OOF uses folds A, B, C, D, E. Then every train image in fold A is predicted using 80% of train (B, C, D, E). Let's call one image in fold A, <code>img_123</code>. If a second OOF uses folds F, G, H, I, J. Then <code>img_123</code> is in fold G and it gets predicted with the 80% of train in F, H, I, J. </p>\n<p>The 80% that predicted <code>img_123</code> in OOF_1 is different than the 80% that predicted <code>img_123</code> in OOF_2. Therefore you will most certainly get an increase when you average OOF_1 and OOF_2. So using different folds gives increases even when the two models are similar and not diverse.</p>\n<p>However when you blend SUB_1 and SUB_2 you will not get an increase because both are already similar. In conclusion, using different folds gives increases in CV which do not give increases in SUB blending.</p>",
          "rawMarkdown": "It won't cause a leak but you will get artificial CV gains. If one OOF uses folds A, B, C, D, E. Then every train image in fold A is predicted using 80% of train (B, C, D, E). Let's call one image in fold A, `img_123`. If a second OOF uses folds F, G, H, I, J. Then `img_123` is in fold G and it gets predicted with the 80% of train in F, H, I, J. \n\nThe 80% that predicted `img_123` in OOF_1 is different than the 80% that predicted `img_123` in OOF_2. Therefore you will most certainly get an increase when you average OOF_1 and OOF_2. So using different folds gives increases even when the two models are similar and not diverse.\n\nHowever when you blend SUB_1 and SUB_2 you will not get an increase because both are already similar. In conclusion, using different folds gives increases in CV which do not give increases in SUB blending.",
          "votes": 7
        },
        {
          "id": 976461,
          "postDate": "2020-08-18T21:53:53.423Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I can see the reason why CV score increases. But I think SUB_1 and SUB_2 will be slightly different because models are trained on different folds of data. So my feeling is that if we run different folds OOF ensemble, CV will increase and LB will increase small or no increase.</p>",
          "rawMarkdown": "Thanks @cdeotte I can see the reason why CV score increases. But I think SUB_1 and SUB_2 will be slightly different because models are trained on different folds of data. So my feeling is that if we run different folds OOF ensemble, CV will increase and LB will increase small or no increase."
        },
        {
          "id": 976638,
          "postDate": "2020-08-19T02:12:07.977Z",
          "content": "<p>SUB_1 and SUB2_1 may be slightly different. But if their combination increases public LB, it is a different reason than combining OOF_1 and OOF_2. Therefore the weights that optimize the increase of OOF_1 plus OOF_2 isn't necessarily the same weights that would optimize SUB_1 plus SUB_2.</p>",
          "rawMarkdown": "SUB_1 and SUB2_1 may be slightly different. But if their combination increases public LB, it is a different reason than combining OOF_1 and OOF_2. Therefore the weights that optimize the increase of OOF_1 plus OOF_2 isn't necessarily the same weights that would optimize SUB_1 plus SUB_2.",
          "votes": 1
        }
      ]
    },
    {
      "id": 976391,
      "postDate": "2020-08-18T20:18:19.303Z",
      "content": "<p>This is great. As a beginner, ensembling has always puzzled me, What are more creative ways to do ensembling?</p>",
      "rawMarkdown": "This is great. As a beginner, ensembling has always puzzled me, What are more creative ways to do ensembling?",
      "votes": 1,
      "replies": [
        {
          "id": 976412,
          "postDate": "2020-08-18T20:34:40.413Z",
          "content": "<p>What we're doing here is similar to linear regression or logistic regression. We are using the OOF as \"train data\" where each OOF is one column feature and we are predicting target. Since our model is linear, we don't worry above overfitting (which seldom occurs for linear models).</p>\n<p>Instead of linear regression, we can use any model as our \"ensemble model\". For example we could train an XGB using the OOF to predict target. Also we can add additional features. For example, some columns can be the OOF and other columns can be meta data like age, sex, site, etc.</p>\n<p>Essentially this is stacking.</p>\n<pre><code>X_train[:,0] = oof_1.predictions\nX_train[:,1] = oof_2.predictions\nX_train[:,2] = oof_3.predictions\nX_train[:,3] = oof.age\nX_train[:,4] = oof.sex\ny_train = oof.target\n\nmodel.fit(X_train,y_train)\n\nX_test[:,0] = sub_1.target\nX_test[:,1] = sub_2.target\nX_test[:,2] = sub_3.target\nX_test[:,3] = sub.age\nX_test[:,4] = sub.sex\n\nkaggle_sub = model.predict(X_test)\n</code></pre>",
          "rawMarkdown": "What we're doing here is similar to linear regression or logistic regression. We are using the OOF as \"train data\" where each OOF is one column feature and we are predicting target. Since our model is linear, we don't worry above overfitting (which seldom occurs for linear models).\n\nInstead of linear regression, we can use any model as our \"ensemble model\". For example we could train an XGB using the OOF to predict target. Also we can add additional features. For example, some columns can be the OOF and other columns can be meta data like age, sex, site, etc.\n\nEssentially this is stacking.\n\n    X_train[:,0] = oof_1.predictions\n    X_train[:,1] = oof_2.predictions\n    X_train[:,2] = oof_3.predictions\n    X_train[:,3] = oof.age\n    X_train[:,4] = oof.sex\n    y_train = oof.target\n\n    model.fit(X_train,y_train)\n\n    X_test[:,0] = sub_1.target\n    X_test[:,1] = sub_2.target\n    X_test[:,2] = sub_3.target\n    X_test[:,3] = sub.age\n    X_test[:,4] = sub.sex\n\n    kaggle_sub = model.predict(X_test)",
          "votes": 5
        },
        {
          "id": 976413,
          "postDate": "2020-08-18T20:35:03.787Z",
          "content": "<p><code>Afterward we have predictions for every training image. This complete set of predictions is called OOF,</code> i think it should be test images</p>",
          "rawMarkdown": "`Afterward we have predictions for every training image. This complete set of predictions is called OOF,` i think it should be test images"
        },
        {
          "id": 976416,
          "postDate": "2020-08-18T20:38:07.037Z",
          "content": "<p>No. Fold 1, Fold 2, Fold 3, Fold 4, Fold 5 are each 20% of the train images. So when we build a model with Fold 2, 3, 4, 5 and predict Fold 1. We now have predictions for the 20% of train images in Fold 1.</p>\n<p>After we do this five times, we have predictions for 100% of train images. These are called OOF</p>\n<p>Additionally, each fold model also makes predictions on the test images. After averaging these 5 sets of predictions, these are called SUB</p>",
          "rawMarkdown": "No. Fold 1, Fold 2, Fold 3, Fold 4, Fold 5 are each 20% of the train images. So when we build a model with Fold 2, 3, 4, 5 and predict Fold 1. We now have predictions for the 20% of train images in Fold 1.\n\nAfter we do this five times, we have predictions for 100% of train images. These are called OOF\n\nAdditionally, each fold model also makes predictions on the test images. After averaging these 5 sets of predictions, these are called SUB",
          "votes": 3
        }
      ]
    },
    {
      "id": 976370,
      "postDate": "2020-08-18T20:13:25.350Z",
      "content": "<p>Gonna bookmark this! thanks for the detailed explanation.</p>",
      "rawMarkdown": "Gonna bookmark this! thanks for the detailed explanation.",
      "votes": 1
    },
    {
      "id": 976311,
      "postDate": "2020-08-18T19:25:47.447Z",
      "content": "<p>thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> ! <br>\nI used your starter notebook and had 12 models. <br>\ni did ensembling using another starter kernel. <br>\nthis is my final ensemble with 6 models (Private LB = 0.9413). <br>\n<a href=\"https://www.kaggle.com/mpsampat/final-simple-oof-ensembling-methods-6-models\" target=\"_blank\">https://www.kaggle.com/mpsampat/final-simple-oof-ensembling-methods-6-models</a></p>\n<p>I had a hard time to find the optimal subset of models. Your starter notebook will be very handy in learning this task! <br>\nthanks again for sharing this and other notebooks! <br>\nCheers! </p>",
      "rawMarkdown": "thanks @cdeotte ! \nI used your starter notebook and had 12 models. \ni did ensembling using another starter kernel. \nthis is my final ensemble with 6 models (Private LB = 0.9413). \nhttps://www.kaggle.com/mpsampat/final-simple-oof-ensembling-methods-6-models\n\nI had a hard time to find the optimal subset of models. Your starter notebook will be very handy in learning this task! \nthanks again for sharing this and other notebooks! \nCheers! ",
      "votes": 1
    },
    {
      "id": 985887,
      "postDate": "2020-08-26T04:53:38.763Z",
      "content": "<blockquote>\n  <p>For each of the 5 fold models above, we predict the test images. Therefore we have 5 predictions for each test image.</p>\n</blockquote>\n<p>Amazing article. I keep revisiting this again &amp; again. I was wondering how there are 5 predictions for the test image since the 5 different folds will be a part of single model? So essentially, one model will predict one output for a single test image? I understand for different models the prediction of test image will vary but how is it varying for single model (and different folds)? Do we take the average for test image output for different models or for different folds of the same model?</p>",
      "rawMarkdown": "> For each of the 5 fold models above, we predict the test images. Therefore we have 5 predictions for each test image.\n\n\nAmazing article. I keep revisiting this again & again. I was wondering how there are 5 predictions for the test image since the 5 different folds will be a part of single model? So essentially, one model will predict one output for a single test image? I understand for different models the prediction of test image will vary but how is it varying for single model (and different folds)? Do we take the average for test image output for different models or for different folds of the same model?",
      "votes": 2,
      "replies": [
        {
          "id": 986903,
          "postDate": "2020-08-26T21:34:31.537Z",
          "content": "<p>In 5 KFold, you actually build 5 models. All models use the same architecture and hyperparameters so they will perform similarily, but each is trained on a different 80% of data. (Their 4 folds).</p>\n<p>So after performing 5 KFold, we have 5 models and each model predicts all the test images. Therefore each test image has 5 predictions (where each pred is between 0 and 1). We take the simple average of these 5 numbers as the prediction for test.</p>",
          "rawMarkdown": "In 5 KFold, you actually build 5 models. All models use the same architecture and hyperparameters so they will perform similarily, but each is trained on a different 80% of data. (Their 4 folds).\n\nSo after performing 5 KFold, we have 5 models and each model predicts all the test images. Therefore each test image has 5 predictions (where each pred is between 0 and 1). We take the simple average of these 5 numbers as the prediction for test.",
          "votes": 1
        },
        {
          "id": 987068,
          "postDate": "2020-08-27T00:48:28.893Z",
          "content": "<p><a href=\"https://www.kaggle.com/ujjwal29jain\" target=\"_blank\">@ujjwal29jain</a> You can think about CV KFold like a random forest. After building a random forest, you call it \"one model\" but really, the forest is many models each trained with a different 80% of the data and then ensembled together.</p>\n<p>Similarily, KFold is K models each trained with a different 80% of the data and then ensembled together. (And both help prevent overfitting and generalize better to unseen data).</p>",
          "rawMarkdown": "@ujjwal29jain You can think about CV KFold like a random forest. After building a random forest, you call it \"one model\" but really, the forest is many models each trained with a different 80% of the data and then ensembled together.\n\nSimilarily, KFold is K models each trained with a different 80% of the data and then ensembled together. (And both help prevent overfitting and generalize better to unseen data).",
          "votes": 3
        },
        {
          "id": 1110842,
          "postDate": "2020-12-13T06:07:41.173Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Dear Chris, was revisiting your old posts, and as usual, I have noticed that for such classification problems (and even the recent Cassava Comp), I generally build 5 folds and average out the predictions towards the end when I inference. Usually, this produces better results - but is there any statistical justification for this. Appreciate it if you can link me to one article that explains so…</p>",
          "rawMarkdown": "@cdeotte Dear Chris, was revisiting your old posts, and as usual, I have noticed that for such classification problems (and even the recent Cassava Comp), I generally build 5 folds and average out the predictions towards the end when I inference. Usually, this produces better results - but is there any statistical justification for this. Appreciate it if you can link me to one article that explains so..."
        }
      ]
    },
    {
      "id": 976706,
      "postDate": "2020-08-19T03:49:00.227Z",
      "content": "<p>Thanks for sharing Chris! Probably the clearest explanation of this I’ve seen.</p>\n<p>Is it necessary to use the same folds you used to train your models or can this CV scheme be independent of the initial scheme?</p>",
      "rawMarkdown": "Thanks for sharing Chris! Probably the clearest explanation of this I’ve seen.\n\nIs it necessary to use the same folds you used to train your models or can this CV scheme be independent of the initial scheme?",
      "votes": 2,
      "replies": [
        {
          "id": 976758,
          "postDate": "2020-08-19T04:34:36.367Z",
          "content": "<p>You should always use the same folds. Otherwise if you ensemble models using different folds, you will always see a CV increase but the LB may not increase.</p>",
          "rawMarkdown": "You should always use the same folds. Otherwise if you ensemble models using different folds, you will always see a CV increase but the LB may not increase.",
          "votes": 3
        },
        {
          "id": 976947,
          "postDate": "2020-08-19T07:30:46.203Z",
          "content": "<p>Chris, can you please elaborate this a bit further? </p>\n<p>Because, in this competition I blended different models using oof (each having a different seed value, merged oof files using image id) and it seems to have increased both CV and LB(public and private).</p>",
          "rawMarkdown": "Chris, can you please elaborate this a bit further? \n\nBecause, in this competition I blended different models using oof (each having a different seed value, merged oof files using image id) and it seems to have increased both CV and LB(public and private).",
          "votes": 2
        },
        {
          "id": 981999,
          "postDate": "2020-08-23T00:02:14.210Z",
          "content": "<p><a href=\"https://www.kaggle.com/manjeshg03\" target=\"_blank\">@manjeshg03</a> Using different seeds does add variety to your models. However it is hard to compare the models. For example if you have one model with CV 0.920 using seed=42 and another with CV 0.922 with seed=123, which is better? We don't know.</p>\n<p>Also if you ensemble two identical models where one uses seed=42 and the another uses seed=123, the ensemble CV score will always increase. This is because a 5 fold model only uses 80% of train data to predict the OOF. When you ensemble two different seeds, the result are predictions that uses more than 80% (most likely around 95%) of train data for each OOF prediction.</p>\n<p>However when you ensemble the submission.csv files of two identical models where one uses seed=42 and the another uses seed=123, you won't automatically have an increase. Because each model already uses the full 100% train to make the submission file (as result of blending the 5 folds together). So, the CV will always increase but the LB will <strong>not</strong> always increase. This makes using different seeds a confusing situation. (Not necessarily a bad situation, but certainly a confusing situation to evaluate CV LB).</p>",
          "rawMarkdown": "@manjeshg03 Using different seeds does add variety to your models. However it is hard to compare the models. For example if you have one model with CV 0.920 using seed=42 and another with CV 0.922 with seed=123, which is better? We don't know.\n\nAlso if you ensemble two identical models where one uses seed=42 and the another uses seed=123, the ensemble CV score will always increase. This is because a 5 fold model only uses 80% of train data to predict the OOF. When you ensemble two different seeds, the result are predictions that uses more than 80% (most likely around 95%) of train data for each OOF prediction.\n\nHowever when you ensemble the submission.csv files of two identical models where one uses seed=42 and the another uses seed=123, you won't automatically have an increase. Because each model already uses the full 100% train to make the submission file (as result of blending the 5 folds together). So, the CV will always increase but the LB will **not** always increase. This makes using different seeds a confusing situation. (Not necessarily a bad situation, but certainly a confusing situation to evaluate CV LB).",
          "votes": 11
        },
        {
          "id": 984837,
          "postDate": "2020-08-25T10:20:51.067Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Got the point. Thanks a lot!</p>",
          "rawMarkdown": "@cdeotte Got the point. Thanks a lot!",
          "votes": 1
        }
      ]
    },
    {
      "id": 976329,
      "postDate": "2020-08-18T19:42:00.883Z",
      "content": "<p>Nice brute force algorithm. I will definitely give it a try to improve my optimization techniques. How would you do this for OOF predictions for models of different seeds?</p>",
      "rawMarkdown": "Nice brute force algorithm. I will definitely give it a try to improve my optimization techniques. How would you do this for OOF predictions for models of different seeds?",
      "votes": 2,
      "replies": [
        {
          "id": 976639,
          "postDate": "2020-08-19T02:12:34.507Z",
          "content": "<p>I would suggest training a NN on the OOF predictions, adding a small amount of noise as data augmentation. </p>\n<p>As Chris has pointed out below - if you train exactly the same architecture twice, with different seeds each time, then the OOF predictions of the two trained models will be different, but their submissions will be similar. </p>\n<p>I think this can be overcome using noise because I found, with my own submissions, that I could almost perfectly distinguish the combined OOF predictions ([oof_1,oof_2,…,oof_n]) from the combined submissions ([sub_1,sub_2,…,sub_n]). Meaning that the distributions were slightly different. This isn't surprising because each OOF prediction is made by a single model, while each submission prediction comes from a combination of models. However, I found that I could partially avoid this issue by using a small amount of gaussian noise as augmentation while training a NN. With a small amount of noise the two sets could not be distinguished very well. Performing TTA with this noise and using stratified cross validation I trained a model that combined my OOF predictions slightly better than a straightforward linear blend in terms of CV score.</p>\n<p>What I am saying is that if the right amount of noise is added (something with less variance than that of the model that is having noise added to it) then the combined OOF predictions are difficult to distinguish from the combined submissions. As this was the case when using the same seed every time, I imagine it will also be the case when you do not have the same seed. And, in your case, it would also ensure that any of the OOF predictions that are similar become difficult to distinguish, removing their ability to provide artificial gains in the combined OOF predictions.</p>\n<p>I hope that makes sense, and it would be interesting to know if it works for you!</p>",
          "rawMarkdown": "I would suggest training a NN on the OOF predictions, adding a small amount of noise as data augmentation. \n\nAs Chris has pointed out below - if you train exactly the same architecture twice, with different seeds each time, then the OOF predictions of the two trained models will be different, but their submissions will be similar. \n\nI think this can be overcome using noise because I found, with my own submissions, that I could almost perfectly distinguish the combined OOF predictions ([oof_1,oof_2,...,oof_n]) from the combined submissions ([sub_1,sub_2,...,sub_n]). Meaning that the distributions were slightly different. This isn't surprising because each OOF prediction is made by a single model, while each submission prediction comes from a combination of models. However, I found that I could partially avoid this issue by using a small amount of gaussian noise as augmentation while training a NN. With a small amount of noise the two sets could not be distinguished very well. Performing TTA with this noise and using stratified cross validation I trained a model that combined my OOF predictions slightly better than a straightforward linear blend in terms of CV score.\n\nWhat I am saying is that if the right amount of noise is added (something with less variance than that of the model that is having noise added to it) then the combined OOF predictions are difficult to distinguish from the combined submissions. As this was the case when using the same seed every time, I imagine it will also be the case when you do not have the same seed. And, in your case, it would also ensure that any of the OOF predictions that are similar become difficult to distinguish, removing their ability to provide artificial gains in the combined OOF predictions.\n\nI hope that makes sense, and it would be interesting to know if it works for you!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2919597,
      "postDate": "2024-07-13T03:10:49.583Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, this was a nice discussion. But I was wondering maybe if you could give some advice in context of code competitions, like the <a href=\"https://www.kaggle.com/competitions/isic-2024-challenge\" target=\"_blank\">current ISIC competition</a>. Here, we do not have the entire test set, and hence need to do the full inference within the limit of 9 hours. </p>\n<p>Say, I have trained using 5-fold cross validation. So essentially I have 5 models to use for inference. Which one should I choose? Infering using all of them and then averaging just takes too much time.</p>",
      "rawMarkdown": "@cdeotte, this was a nice discussion. But I was wondering maybe if you could give some advice in context of code competitions, like the [current ISIC competition](https://www.kaggle.com/competitions/isic-2024-challenge). Here, we do not have the entire test set, and hence need to do the full inference within the limit of 9 hours. \n\nSay, I have trained using 5-fold cross validation. So essentially I have 5 models to use for inference. Which one should I choose? Infering using all of them and then averaging just takes too much time."
    },
    {
      "id": 2017546,
      "postDate": "2022-11-04T23:54:03.187Z",
      "content": "<p>The best concept I've ever seen. And I always learn new things from you.</p>",
      "rawMarkdown": "The best concept I've ever seen. And I always learn new things from you."
    },
    {
      "id": 1871067,
      "postDate": "2022-07-26T04:03:15.060Z",
      "content": "<p>Thank you for sharing your Hill Climbing approach.</p>",
      "rawMarkdown": "Thank you for sharing your Hill Climbing approach."
    },
    {
      "id": 1822540,
      "postDate": "2022-06-16T13:40:08.877Z",
      "content": "<p>thanks for sharing. hill climbing,good job.</p>",
      "rawMarkdown": "thanks for sharing. hill climbing,good job."
    },
    {
      "id": 1729854,
      "postDate": "2022-03-20T15:53:22.940Z",
      "content": "<p>Thankyou helped a lot</p>",
      "rawMarkdown": "Thankyou helped a lot"
    },
    {
      "id": 1654624,
      "postDate": "2022-01-18T17:11:02.800Z",
      "content": "<p>Thanks for sharing this!<br>\nIts been really helpful.</p>",
      "rawMarkdown": "Thanks for sharing this!\nIts been really helpful."
    },
    {
      "id": 1118700,
      "postDate": "2020-12-19T10:11:24.077Z",
      "content": "<p>Great post! Crisp and Clear! Thanks for sharing</p>",
      "rawMarkdown": "Great post! Crisp and Clear! Thanks for sharing"
    },
    {
      "id": 976390,
      "postDate": "2020-08-18T20:18:06.273Z",
      "content": "<p>Are you by finding these optimal weights not overfitting the validation set?</p>",
      "rawMarkdown": "Are you by finding these optimal weights not overfitting the validation set?",
      "replies": [
        {
          "id": 976409,
          "postDate": "2020-08-18T20:26:26.717Z",
          "content": "<p>These weights create a linear model. A line rarely overfits anything.</p>\n<p>If you trained XGB using the OOF files, then you could overfit validation and must be careful. </p>",
          "rawMarkdown": "These weights create a linear model. A line rarely overfits anything.\n\nIf you trained XGB using the OOF files, then you could overfit validation and must be careful. \n\n",
          "votes": 7
        },
        {
          "id": 977235,
          "postDate": "2020-08-19T11:13:17.920Z",
          "content": "<p>Buts let's assume we are ensembling multiple models:</p>\n<p>submission = 0.45 * model1 + 0.03 * model2 + … + 0.11 modelx</p>\n<p>This way you are just selecting the highest score on the validation set, which may be overfitting it and not generalizing well at all.</p>\n<p>I guess what you do is not add all the models at once?</p>",
          "rawMarkdown": "Buts let's assume we are ensembling multiple models:\n\nsubmission = 0.45 * model1 + 0.03 * model2 + ... + 0.11 modelx\n\nThis way you are just selecting the highest score on the validation set, which may be overfitting it and not generalizing well at all.\n\nI guess what you do is not add all the models at once?"
        }
      ]
    },
    {
      "id": 1946594,
      "postDate": "2022-09-19T23:39:58.900Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1947541,
          "postDate": "2022-09-20T14:14:43.423Z",
          "content": "<p>No, because the model will already have been trained on the validation data. Doing this will introduce leakage</p>",
          "rawMarkdown": "No, because the model will already have been trained on the validation data. Doing this will introduce leakage",
          "votes": 1
        }
      ]
    },
    {
      "id": 987263,
      "postDate": "2020-08-27T05:58:34.753Z",
      "content": "<p>Great stuff. Thanks for explaining </p>",
      "rawMarkdown": "Great stuff. Thanks for explaining ",
      "votes": 1
    },
    {
      "id": 986044,
      "postDate": "2020-08-26T06:57:01.327Z",
      "content": "<p>nice explanation <br>\nThanks for sharing</p>",
      "rawMarkdown": "nice explanation \nThanks for sharing",
      "votes": 1
    },
    {
      "id": 983744,
      "postDate": "2020-08-24T14:54:53.317Z",
      "content": "<p>thanks for sharing this!</p>",
      "rawMarkdown": "thanks for sharing this!",
      "votes": 1
    },
    {
      "id": 976322,
      "postDate": "2020-08-18T19:35:36.740Z",
      "content": "<p>Thanks sir</p>",
      "rawMarkdown": "Thanks sir",
      "votes": 1
    },
    {
      "id": 976303,
      "postDate": "2020-08-18T19:22:13.703Z",
      "content": "<p>This the light!! Thanks!</p>",
      "rawMarkdown": "This the light!! Thanks!",
      "votes": 1
    },
    {
      "id": 1797454,
      "postDate": "2022-05-22T01:26:52.757Z",
      "content": "<p>Thanks for sharing, extremely helpful!!</p>",
      "rawMarkdown": "Thanks for sharing, extremely helpful!!"
    }
  ],
  "comments": [
    {
      "id": 976302,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-08-18T19:21:45.140000",
      "content": "<p>Kagglers refer to this as <code>hill climbing</code> just as a side note if others see this phrase somewhere else.</p>",
      "votes": 13,
      "replies": [
        {
          "id": 976313,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-18T19:29:00.980000",
          "content": "<p>thanks. I updated my post.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 984401,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-25T04:46:25.477000",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> When you and others use <code>hill climbing</code>, do you continue adding models as long as CV increases, or do you only add a new model if it increases CV by some tolerance?</p>\n<p>For my submission, i used <code>TOL = 0.0003</code> thinking I would prevent noise. But if you compare <code>hill climbing</code> to XGB stacking, then XGB uses <code>TOL = 0</code>, right? XGB just keeps adding as long as CV increases.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 984824,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-08-25T10:07:07.807000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I usually apply it for a certain number of iterations. To avoid overfitting you can also add certain restrictions, like a single model can only be added once or twice.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 979176,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2020-08-20T17:12:15.270000",
      "content": "<p>Thanks for sharing this! Good explanation for CV!</p>\n<p>After reading some discussions in the past few days after the competition end, I gradually realized that what more people need is not the top solutions, but the basic ML knowledge, just like this posts.</p>\n<p>I see many people still don't understand why they should not trust public LB score (especially in this comp). But I don’t know how to explain to them in a simple and understandable way.</p>\n<p>I think your post can be a good explanation for it as well. Thanks again!</p>",
      "votes": 14,
      "replies": [
        {
          "id": 979203,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-20T17:31:26.207000",
          "content": "<p>Thanks Qishen. I updated the discussion title to `\"How To CV and How To Ensemble OOF\". I thought i was just explaining ensemble, but you are right, this is a nice explanation of CV too.</p>\n<p>Congrats again on you and your team's awesome model with amazing 0.960 CV!</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 1654315,
      "author_name": "Slawek Biel",
      "author_url": "",
      "post_date": "2022-01-18T11:35:17.327000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, your Hill Climbing approach is a great trick. I'll for sure try it out in my next competition.<br>\nI have small doubt, in the post you've linked you say </p>\n<blockquote>\n  <p>During the competition, every model should use the same folds. </p>\n</blockquote>\n<p>Does that really matter when your oof predictions have all the train samples anyway? I specifically ask in a context of ensembling models as a team, where different members worked with different splits.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 988805,
      "author_name": "datasaurus",
      "author_url": "",
      "post_date": "2020-08-28T09:42:28.270000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, have you seen this paper \"On Cross-Validation and Stacking: Building seemingly predictive models on random data\"?<br>\n<a href=\"https://www.kdd.org/exploration_files/v12-02-4-UR-Perlich.pdf\" target=\"_blank\">https://www.kdd.org/exploration_files/v12-02-4-UR-Perlich.pdf</a> </p>\n<p>I'd be interested to know what you and other Kagglers think about it</p>",
      "votes": 3,
      "replies": [
        {
          "id": 989300,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-28T17:42:23.213000",
          "content": "<p>Thanks for the paper, i'll check it out and post my reaction here.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 979325,
      "author_name": "Massimiliano Viola",
      "author_url": "",
      "post_date": "2020-08-20T19:03:15.597000",
      "content": "<p>Hi Chris, the explanation is very clear but I have a question and after reading the post a couple of times I can't figure out the answer on my own. Talking about the submission file, you say</p>\n<blockquote>\n  <p>For each of the 5 fold models above, we predict the test images. Therefore we have 5 predictions for each test image. We take the average of these 5 sets of predictions and this is our submission.csv file that we submit to Kaggle. When you submit this to Kaggle, your LB score should be similar to your CV score.</p>\n</blockquote>\n<p>and this makes perfect sense, but why do you never train a model using all train data once you have a high CV score for that model?  <br>\nIs it because you would not know how it performs when trained with an additional fold, making all the work useless and destroying the ensembling process?  <br>\nI am asking because so far, when I perform cross-validation, I always fit the model on all train data with the best parameters to make the test predictions and so your way of doing it by averaging is completely new to me.<br>\nThank you!</p>",
      "votes": 4,
      "replies": [
        {
          "id": 979396,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-20T19:49:16.223000",
          "content": "<p>Doing it as I describe associates the CV to LB score. Because the identical models that produced the CV score are the same models producing the LB score. (When you run twice as your describe, you don't have this guarantee). That being said, I sometimes do what you suggest. Both ways have pros and cons.</p>\n<p>Note that the way I describe in this post is a form of bagging similar to a random forest. Each of the 5 fold models are training on a random 80% of the train data. Therefore even if each of the five fold models overfits their 80% train data, the ensemble of the 5 folds will not overfit the entire 100% train data. So predicting test using CV is also a way to prevent overfitting the train data (i.e. generalizing better to test data) in the same way that a random forest uses bagging to prevent overfitting train data.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 979418,
          "author_name": "Massimiliano Viola",
          "author_url": "",
          "post_date": "2020-08-20T20:31:06.320000",
          "content": "<p>Ok, got it, thank you so much.  <br>\nBy doing this, you are training on less data but as a tradeoff you have a more reliable estimate of the performance of your model and also a free ensemble of 5 classifiers which is never a bad thing.  <br>\nI guess I'll try to apply this method in House Prices competition when I have my final model ready!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 976417,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2020-08-18T20:43:15.087000",
      "content": "<p>The strategy I usually go for is finding a set of models that give the best CV when averaged together.</p>\n<p>Do you have any idea how your ensemble performs (especially on the leaderboard) when giving equal weights to every model ? <br>\nWhat about using the 8  (or any number of) models that have the best CV when averging the predictions together ? </p>\n<p>This would be nice to see how much hill climbing actually helps.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 976425,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-18T20:53:24.477000",
          "content": "<p>Great suggestions Theo. I will investigate this.</p>\n<p>I should probably change the language in my post. I should be more clear that high climbing is only one way and there are many others like the ones you suggest.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 976447,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2020-08-18T21:22:20.207000",
          "content": "<blockquote>\n  <p>I should probably change the language in my post. I should be more clear that high climbing is only one way and there are many others like the ones you suggest.</p>\n</blockquote>\n<p>No worries, this was clear to me :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 976452,
          "author_name": "Mehul Sampat",
          "author_url": "",
          "post_date": "2020-08-18T21:29:34.490000",
          "content": "<p>Interesting! thanks for sharing. </p>\n<ol>\n<li>In my case, I used a kernel compute combined OOF for 6 models. </li>\n<li>In the Kernel 4 types of combined OOF are computed. </li>\n<li>OOF_avg_auc, OOF_rank_auc, OOF_pow_auc and OOF_bo_auc</li>\n<li>My heuristics are: I should pick the combination which has :</li>\n<li>highest OOF_bo_auc</li>\n<li>Or pick model which has highest average of all 4 OOF. (since if it is a good model, all OOF should be high ? )</li>\n<li>My top combinations were: </li>\n</ol>\n<table>\n<thead>\n<tr>\n<th>no_of_models</th>\n<th>OOF_avg_auc</th>\n<th>OOF_rank_auc</th>\n<th>OOF_pow_auc</th>\n<th>OOF_bo_auc</th>\n<th>Overall_mean</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>8</td>\n<td>0.93898</td>\n<td>0.93993</td>\n<td>0.937304</td>\n<td>0.93963</td>\n<td>0.93963</td>\n<td>0.9522</td>\n<td>0.9409</td>\n</tr>\n<tr>\n<td>8</td>\n<td>0.93787</td>\n<td>0.93806</td>\n<td>0.93652</td>\n<td>0.93918</td>\n<td>0.93791</td>\n<td>0.9518</td>\n<td>0.9411</td>\n</tr>\n<tr>\n<td>6</td>\n<td>0.938307</td>\n<td>0.938737</td>\n<td>0.93630</td>\n<td>0.93900</td>\n<td>0.93808</td>\n<td>0.9512</td>\n<td>0.9413</td>\n</tr>\n</tbody>\n</table>\n<ol>\n<li>I have a total of 12 models. </li>\n<li>So for ensemble of 6 models i thought i have to check 12<em>11</em>10<em>8</em>9*7 (== 12! - 6! ) combinations. </li>\n<li>As I understand, hill climbing will help to avoid checking 12<em>11</em>10<em>8</em>9*7 combinations. I will try out hill climbing for my models. </li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 976459,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-18T21:50:44.220000",
          "content": "<blockquote>\n  <p>What about using the 8 (or any number of) models that have the best CV when averging the predictions together ?</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> I just averaged all 39 models and it performed identical to <code>hill climbing</code>. Also I randomly chose 8 repeatedly 10,000 times and just averaged them. Then used the 8 with best CV. That also performed the same. So maybe everything just does the same haha.</p>\n<p>I know if you only have 2 or 3 models, then finding weights is better than just averaging them but perhaps when you have dozens of models just averaging works as well as finding weights.</p>\n<p>(I've never had so many models before so I don't have much experience what is best when you have dozens of models)</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 976462,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-18T21:54:49.323000",
          "content": "<p>But i guess there is an advantage to hill climbing or your select 8 method. In the end, you will have fewer models in your ensemble which is usually a good thing.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 977039,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2020-08-19T08:52:27.083000",
          "content": "<p>Interesting, thanks a lot for checking </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 983075,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-24T03:09:30.487000",
          "content": "<p>I discovered that if i set <code>TOL = 0</code> to allow hill climbing to add any new model that increases CV AUC greater than 0, then the CV and private LB will become higher than averaging all or using a random 8. So perhaps to get the most out of high climbing, you need to set <code>TOL = 0</code>. I'm still unsure what is the optimal value because if an additional model only increase CV AUC 0.0001 it may just be random noise. (That's why I initially choose <code>TOL = 0.0003</code> in my hill climbing notebook. Perhaps I should analyze this mathematically).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 998817,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2020-09-05T05:40:33.813000",
      "content": "<p>Suppose we are using a model and I have to choose the best hyperparameters for that model. Now, will it be right that for each hyperparameter I do a new cv and note down the best parameters? By doing this am I introducing any leakage?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 998899,
          "author_name": "NoMalady",
          "author_url": "",
          "post_date": "2020-09-05T07:15:17.133000",
          "content": "<p>That is not right in my opinion since the hyperparameters like this themselves might not give the best result in tandem. Choose the cv set with the least loss value and work on that but remember changing the hyperparameters too much would make it biased to your cv set and then might fail at generalizing. There is always another way to make sure, try it yourself once. I hope this helps : ).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 997407,
      "author_name": "Blesson Densil",
      "author_url": "",
      "post_date": "2020-09-04T03:09:53.400000",
      "content": "<p>Thanks for sharing. I do follow these approaches on Stacking while working on the projects.<br>\nHave you tried Voting Classifiers as well?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 988133,
      "author_name": "Shri Adke",
      "author_url": "",
      "post_date": "2020-08-27T19:59:22.867000",
      "content": "<p>A good explaination, was very helpful.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 987504,
      "author_name": "Ravi Singh",
      "author_url": "",
      "post_date": "2020-08-27T10:04:31.347000",
      "content": "<p>Good one</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 983330,
      "author_name": "UjjwalJain",
      "author_url": "",
      "post_date": "2020-08-24T07:37:11.813000",
      "content": "<p>Very crisp, clear &amp; to the point explanation. I am a novice at kaggle and I am learning amazing stuff reading your notebook &amp; discussion. Thanks!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 983068,
      "author_name": "lr",
      "author_url": "",
      "post_date": "2020-08-24T03:03:32.770000",
      "content": "<p>Thanks for sharing great knowledge about ensembling! </p>\n<p>I didn't know hill climbing approach to ensemble the model and helped me a lot! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 979255,
      "author_name": "Marcelo Kittlein",
      "author_url": "",
      "post_date": "2020-08-20T18:06:47.880000",
      "content": "<p>Because of hardware limitations I have been trying this….</p>\n<p>for each epoch in 1:30<br>\ntrain for 100 steps; in each  step use a batch of n images that is a random sample of the train set with augmentations (resampled at each step)<br>\nvalidate for 30 steps; in each step use a batch of n images that is a random sample of the train set without augmentations (resampled at each step)</p>\n<p>n is 16 32 64 depending on image size</p>\n<p>Different runs for the same model got consistent scores between \"this cv score\" and private set… not with public set thas was a bit lower…</p>\n<p>changing the set of images at each step may result in underfitting? it certainly did not overfit</p>",
      "votes": 1,
      "replies": [
        {
          "id": 979275,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-20T18:23:29.133000",
          "content": "<p>Nice trick Marcelo. </p>\n<p>Note that if a full epoch has 1000 steps and you are only training on 100 steps, then after 10 of your epochs, you will do the same thing as 1 of the original epochs. The difference would be that you are changing the learning rate more often than it would occur in the original epoch.</p>\n<p>Another trick to speed up training is to train on random crops. So if you wish to train a model on 384x384. Then you randomly crop each train image as 192x192 (as data augmentation). Then after training for a while your model will see the entire 384x384 image, but you train fast because your are only processing 192x192 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 979677,
          "author_name": "Marcelo Kittlein",
          "author_url": "",
          "post_date": "2020-08-21T03:55:37.117000",
          "content": "<p>That seems very interesting! Can an EffNet model trained with one image size tranfer weights to an EffNet model to predict a larger image size? or it only applies to simpler models?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 979693,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-21T04:14:32.493000",
          "content": "<p>Yes. You can train an EffNet with 192x192 and then predict with 384x384. (This is what I did in Steel Comp image segmetation, explained <a href=\"https://www.kaggle.com/c/severstal-steel-defect-detection/discussion/114321\" target=\"_blank\">here</a>). You can even train where each batch has a different size. One batch can be 256x256 and another batch can be 512x512. An EffNet model doesn't care about input size. An EffNet is just a bunch of convolution filters that get trained. Convolution filters are little 3x3 \"windows\" that move over an image.</p>\n<p>However, I'm not suggesting mixing training sizes and inference sizes in Melanoma Comp (even though it is possible but maybe it would be good to investigate). I'm suggesting using 384x384 but during training always show it a different random 192x192 crop. It will then learn to classify images based on just using a 192x192 piece of a full image. Then during inference, you use TTA=21, so you will have your CNN predict on 21 different 192x192 random pieces (of the same image) and then you average those 21 predictions as your final prediction for that single image. (So the model is always training and inferring on 192x192 pieces).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 980236,
          "author_name": "DarkCube",
          "author_url": "",
          "post_date": "2020-08-21T12:33:33.543000",
          "content": "<p>Is it possible to train EffNet on 192x192 images just to learn the weights of the kernels to speed up learning at first and then only retrain the FC layers at the tip of the architecture to predict for 384x384 images at the end of training?  Are there any drawbacks to this approach?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 980295,
          "author_name": "Marcelo Kittlein",
          "author_url": "",
          "post_date": "2020-08-21T13:24:39.770000",
          "content": "<p>Thank you very much Sir…<br>\nI found this possibility is very suitable for fine-grained images that are more or less homogeneous. With this approach I could train models with small  cropos much more quickly. I will give it a try with data from <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification\" target=\"_blank\">human protein classification</a> to see how it works …</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 980521,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-21T16:56:44.763000",
          "content": "<blockquote>\n  <p>Is it possible to train EffNet on 192x192 images just to learn the weights of the kernels to speed up learning at first and then only retrain the FC layers at the tip of the architecture to predict for 384x384 images at the end of training? Are there any drawbacks to this approach?</p>\n</blockquote>\n<p>Yes, you can do this. The best way to see how well it works is try it and evaluate CV LB.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 980720,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-21T19:52:49.990000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 980721,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-21T19:52:49.997000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 980722,
          "author_name": "DarkCube",
          "author_url": "",
          "post_date": "2020-08-21T19:52:50.007000",
          "content": "<p>And one more thing: The LB scores are very close to one another (.0005 isn't much especially when you take into consideration that test accuracy isn't true real world accuracy) so there must be some randomness  involved to who might be 1st beyond just skill. So how do hosts deal with this kind of randomness? Do they deem it insignificant? Is it insignificant?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 980885,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-22T01:02:56.293000",
          "content": "<p>Data science is both skill and luck. For example, I built a model from meta data: age, sex, site, image size. (By itself it had CV 0.77 LB 0.77). When I added it to my ensemble, It increased my CV by 0.0005. The question is do I include it in my final submission with such a small CV increase? How much can I gain, and how much do I risk?</p>\n<p>I simulated 1000 private leaderboards. You can see below that in some simulated private leaderboards, including the meta model (as 90% image 10% meta) will decrease my private LB score as much as AUC 0.002! There is also the chance it could increase my private LB 0.003! </p>\n<p>The expected increase in private LB score is 0.0005 with standard deviation 0.0008. That means that it will increase my score in 75% of private leaderboards. Should I take the risk? The best we can do is calculate expected values, variances, and risk and decide whether to include things in our models.</p>\n<p>I did include it and it lowered my private LB by 0.0002 but there was the chance that it could have increased me to Gold !!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Feb01f9d371f4a49be62093d5d2b638bd%2Fmeta.png?generation=1598057575239798&amp;alt=media\" alt=\"\"></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 981245,
          "author_name": "DarkCube",
          "author_url": "",
          "post_date": "2020-08-22T10:11:42.923000",
          "content": "<p>Thank you so much for investing your time into replying to such a newbie I really appreciate it. Due to the MASSIVE information spread about this field, it's really hard to get yourself around. If it wasn't for people like you, new comers won't have any chance keeping up. Especially if they were self-taught like me. Thank you so much</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1253672,
          "author_name": "Jacopo Repossi",
          "author_url": "",
          "post_date": "2021-03-26T23:35:26.560000",
          "content": "<blockquote>\n  <p>I simulated 1000 private leaderboards.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> may I ask you how one could simulate private leaderboards?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 977172,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2020-08-19T10:35:08.297000",
      "content": "<p>Thanks a lot for sharing and it is really good explanation for oof ensemble!</p>\n<p>May I ask something:<br>\nSome Kagglers save their folds to <code>folds.csv</code> and just load <code>folds.csv</code> for more experiments.</p>\n<p>If I use same seed(<code>e.g. random_seed = 42</code>), is it same folds result?</p>\n<p><code>new_folds = sklearn.model_selection.KFold(n_splits = 5, shuffle = True, random_seed = 42)</code><br>\n<code>fold.csv == new_folds</code></p>\n<p>Thanks again <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 977791,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-19T17:53:07.827000",
          "content": "<p>Yes. If two different computers use the same <code>random_seed = 42</code> then they will use the same folds. This is how working on a Kaggle team does it. Every team member uses the same seed on their own computer. Then you can compare everyone's CV score and use everyone's OOF to ensemble your final submission.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 979216,
          "author_name": "Heroseo",
          "author_url": "",
          "post_date": "2020-08-20T17:40:45.370000",
          "content": "<p>Thanks a lot for reply <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> ! <br>\nI want to use what I have learned from you in the next competition. :) </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 976954,
      "author_name": "esprit",
      "author_url": "",
      "post_date": "2020-08-19T07:37:31.457000",
      "content": "<p>What clarity!<br>\nI knew that intuitively, but I wasn't confident enough in my methods to implement it. I want to take advantage of this in the future.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 976589,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2020-08-19T01:07:29.990000",
      "content": "<p>Simple yet very efficient, thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , for me it seems obvious now, but before this competition I never saved my model's OOF predictions, after seeing your public notebook this is a practice that I will keep from now on!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 976604,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-19T01:19:06.400000",
          "content": "<p>Yes. For the past 2 months, I save every model's (1) OOF (2) Submission CSV (3) Model weights. These things always come in handy later.</p>\n<p>For example, maybe we want to build a model that reuses 3 models' weights then build a new model using those three backbones, concatenate the outputs and train a new MLP on top of that.</p>\n<p>Or maybe we want to extract embeddings from old model weight backbones, then use XGB to build a head, etc etc.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 976583,
      "author_name": "M&M",
      "author_url": "",
      "post_date": "2020-08-19T00:52:34.967000",
      "content": "<p>Could you clarify why you do a for loop instead of fitting a logistic regression or other optimization approaches?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 976598,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-19T01:15:16.573000",
          "content": "<p>I've tried logistic regression, linear regression and XGB. For some reason <code>hill climbing</code> achieves a higher ensemble CV. </p>\n<p>(Not sure why. Maybe I'm doing the other methods wrong).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 976688,
          "author_name": "M&M",
          "author_url": "",
          "post_date": "2020-08-19T03:30:10.147000",
          "content": "<p>Very interesting. I see many top kagglers use the same approach as yours, so there might be some underlying reasons.</p>\n<p>By the way thank you for all of your hard work! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 983070,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-24T03:04:21.437000",
          "content": "<p>Note that my forward selection notebook <a href=\"https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private\" target=\"_blank\">here</a> uses <code>TOL = 0.0003</code>. I'm still unsure what is the best selection for this variable. (This is CV AUC threshold increase for adding new models and I didn't want to clutter the ensemble with unnecessary additional models). Also I use <code>DUPLICATES = False</code>. That means that after it adds a model, it won't add that model again.</p>\n<p>But now that I think about it, it would be better to allow adding the same model again because that is equivalent to adding more weight to that model's coefficient. (Kind of how logistic regression or linear regression adjusts weights to find optimal).</p>\n<p>So, FYI, if you use <code>TOL = 0</code> and <code>DUPLICATES = True</code> in my public notebook, the CV gets up to 0.9507 and the private LB gets up to 0.9433! Therefore if you try to compare high climbing to logistic regression or linear regression, those are the CV and LB, you need to beat.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 976448,
      "author_name": "Waylon Wu",
      "author_url": "",
      "post_date": "2020-08-18T21:25:41.197000",
      "content": "<blockquote>\n  <p>Also every model should use the same folds. </p>\n</blockquote>\n<p>Is it necessary to have the same folds? I guess we can have different folds OOF csv files. Will it cause leak?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 976454,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-18T21:38:55.690000",
          "content": "<p>It won't cause a leak but you will get artificial CV gains. If one OOF uses folds A, B, C, D, E. Then every train image in fold A is predicted using 80% of train (B, C, D, E). Let's call one image in fold A, <code>img_123</code>. If a second OOF uses folds F, G, H, I, J. Then <code>img_123</code> is in fold G and it gets predicted with the 80% of train in F, H, I, J. </p>\n<p>The 80% that predicted <code>img_123</code> in OOF_1 is different than the 80% that predicted <code>img_123</code> in OOF_2. Therefore you will most certainly get an increase when you average OOF_1 and OOF_2. So using different folds gives increases even when the two models are similar and not diverse.</p>\n<p>However when you blend SUB_1 and SUB_2 you will not get an increase because both are already similar. In conclusion, using different folds gives increases in CV which do not give increases in SUB blending.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 976461,
          "author_name": "Waylon Wu",
          "author_url": "",
          "post_date": "2020-08-18T21:53:53.423000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I can see the reason why CV score increases. But I think SUB_1 and SUB_2 will be slightly different because models are trained on different folds of data. So my feeling is that if we run different folds OOF ensemble, CV will increase and LB will increase small or no increase.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 976638,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-19T02:12:07.977000",
          "content": "<p>SUB_1 and SUB2_1 may be slightly different. But if their combination increases public LB, it is a different reason than combining OOF_1 and OOF_2. Therefore the weights that optimize the increase of OOF_1 plus OOF_2 isn't necessarily the same weights that would optimize SUB_1 plus SUB_2.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 976391,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2020-08-18T20:18:19.303000",
      "content": "<p>This is great. As a beginner, ensembling has always puzzled me, What are more creative ways to do ensembling?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 976412,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-18T20:34:40.413000",
          "content": "<p>What we're doing here is similar to linear regression or logistic regression. We are using the OOF as \"train data\" where each OOF is one column feature and we are predicting target. Since our model is linear, we don't worry above overfitting (which seldom occurs for linear models).</p>\n<p>Instead of linear regression, we can use any model as our \"ensemble model\". For example we could train an XGB using the OOF to predict target. Also we can add additional features. For example, some columns can be the OOF and other columns can be meta data like age, sex, site, etc.</p>\n<p>Essentially this is stacking.</p>\n<pre><code>X_train[:,0] = oof_1.predictions\nX_train[:,1] = oof_2.predictions\nX_train[:,2] = oof_3.predictions\nX_train[:,3] = oof.age\nX_train[:,4] = oof.sex\ny_train = oof.target\n\nmodel.fit(X_train,y_train)\n\nX_test[:,0] = sub_1.target\nX_test[:,1] = sub_2.target\nX_test[:,2] = sub_3.target\nX_test[:,3] = sub.age\nX_test[:,4] = sub.sex\n\nkaggle_sub = model.predict(X_test)\n</code></pre>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 976413,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2020-08-18T20:35:03.787000",
          "content": "<p><code>Afterward we have predictions for every training image. This complete set of predictions is called OOF,</code> i think it should be test images</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 976416,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-18T20:38:07.037000",
          "content": "<p>No. Fold 1, Fold 2, Fold 3, Fold 4, Fold 5 are each 20% of the train images. So when we build a model with Fold 2, 3, 4, 5 and predict Fold 1. We now have predictions for the 20% of train images in Fold 1.</p>\n<p>After we do this five times, we have predictions for 100% of train images. These are called OOF</p>\n<p>Additionally, each fold model also makes predictions on the test images. After averaging these 5 sets of predictions, these are called SUB</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 976370,
      "author_name": "Santiago Viquez",
      "author_url": "",
      "post_date": "2020-08-18T20:13:25.350000",
      "content": "<p>Gonna bookmark this! thanks for the detailed explanation.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 976311,
      "author_name": "Mehul Sampat",
      "author_url": "",
      "post_date": "2020-08-18T19:25:47.447000",
      "content": "<p>thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> ! <br>\nI used your starter notebook and had 12 models. <br>\ni did ensembling using another starter kernel. <br>\nthis is my final ensemble with 6 models (Private LB = 0.9413). <br>\n<a href=\"https://www.kaggle.com/mpsampat/final-simple-oof-ensembling-methods-6-models\" target=\"_blank\">https://www.kaggle.com/mpsampat/final-simple-oof-ensembling-methods-6-models</a></p>\n<p>I had a hard time to find the optimal subset of models. Your starter notebook will be very handy in learning this task! <br>\nthanks again for sharing this and other notebooks! <br>\nCheers! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 985887,
      "author_name": "UjjwalJain",
      "author_url": "",
      "post_date": "2020-08-26T04:53:38.763000",
      "content": "<blockquote>\n  <p>For each of the 5 fold models above, we predict the test images. Therefore we have 5 predictions for each test image.</p>\n</blockquote>\n<p>Amazing article. I keep revisiting this again &amp; again. I was wondering how there are 5 predictions for the test image since the 5 different folds will be a part of single model? So essentially, one model will predict one output for a single test image? I understand for different models the prediction of test image will vary but how is it varying for single model (and different folds)? Do we take the average for test image output for different models or for different folds of the same model?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 986903,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-26T21:34:31.537000",
          "content": "<p>In 5 KFold, you actually build 5 models. All models use the same architecture and hyperparameters so they will perform similarily, but each is trained on a different 80% of data. (Their 4 folds).</p>\n<p>So after performing 5 KFold, we have 5 models and each model predicts all the test images. Therefore each test image has 5 predictions (where each pred is between 0 and 1). We take the simple average of these 5 numbers as the prediction for test.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 987068,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-27T00:48:28.893000",
          "content": "<p><a href=\"https://www.kaggle.com/ujjwal29jain\" target=\"_blank\">@ujjwal29jain</a> You can think about CV KFold like a random forest. After building a random forest, you call it \"one model\" but really, the forest is many models each trained with a different 80% of the data and then ensembled together.</p>\n<p>Similarily, KFold is K models each trained with a different 80% of the data and then ensembled together. (And both help prevent overfitting and generalize better to unseen data).</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1110842,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2020-12-13T06:07:41.173000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Dear Chris, was revisiting your old posts, and as usual, I have noticed that for such classification problems (and even the recent Cassava Comp), I generally build 5 folds and average out the predictions towards the end when I inference. Usually, this produces better results - but is there any statistical justification for this. Appreciate it if you can link me to one article that explains so…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 976706,
      "author_name": "datasaurus",
      "author_url": "",
      "post_date": "2020-08-19T03:49:00.227000",
      "content": "<p>Thanks for sharing Chris! Probably the clearest explanation of this I’ve seen.</p>\n<p>Is it necessary to use the same folds you used to train your models or can this CV scheme be independent of the initial scheme?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 976758,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-19T04:34:36.367000",
          "content": "<p>You should always use the same folds. Otherwise if you ensemble models using different folds, you will always see a CV increase but the LB may not increase.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 976947,
          "author_name": "Manjesh Gupta",
          "author_url": "",
          "post_date": "2020-08-19T07:30:46.203000",
          "content": "<p>Chris, can you please elaborate this a bit further? </p>\n<p>Because, in this competition I blended different models using oof (each having a different seed value, merged oof files using image id) and it seems to have increased both CV and LB(public and private).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 981999,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-23T00:02:14.210000",
          "content": "<p><a href=\"https://www.kaggle.com/manjeshg03\" target=\"_blank\">@manjeshg03</a> Using different seeds does add variety to your models. However it is hard to compare the models. For example if you have one model with CV 0.920 using seed=42 and another with CV 0.922 with seed=123, which is better? We don't know.</p>\n<p>Also if you ensemble two identical models where one uses seed=42 and the another uses seed=123, the ensemble CV score will always increase. This is because a 5 fold model only uses 80% of train data to predict the OOF. When you ensemble two different seeds, the result are predictions that uses more than 80% (most likely around 95%) of train data for each OOF prediction.</p>\n<p>However when you ensemble the submission.csv files of two identical models where one uses seed=42 and the another uses seed=123, you won't automatically have an increase. Because each model already uses the full 100% train to make the submission file (as result of blending the 5 folds together). So, the CV will always increase but the LB will <strong>not</strong> always increase. This makes using different seeds a confusing situation. (Not necessarily a bad situation, but certainly a confusing situation to evaluate CV LB).</p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 984837,
          "author_name": "Manjesh Gupta",
          "author_url": "",
          "post_date": "2020-08-25T10:20:51.067000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Got the point. Thanks a lot!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 976329,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2020-08-18T19:42:00.883000",
      "content": "<p>Nice brute force algorithm. I will definitely give it a try to improve my optimization techniques. How would you do this for OOF predictions for models of different seeds?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 976639,
          "author_name": "Sam Klein",
          "author_url": "",
          "post_date": "2020-08-19T02:12:34.507000",
          "content": "<p>I would suggest training a NN on the OOF predictions, adding a small amount of noise as data augmentation. </p>\n<p>As Chris has pointed out below - if you train exactly the same architecture twice, with different seeds each time, then the OOF predictions of the two trained models will be different, but their submissions will be similar. </p>\n<p>I think this can be overcome using noise because I found, with my own submissions, that I could almost perfectly distinguish the combined OOF predictions ([oof_1,oof_2,…,oof_n]) from the combined submissions ([sub_1,sub_2,…,sub_n]). Meaning that the distributions were slightly different. This isn't surprising because each OOF prediction is made by a single model, while each submission prediction comes from a combination of models. However, I found that I could partially avoid this issue by using a small amount of gaussian noise as augmentation while training a NN. With a small amount of noise the two sets could not be distinguished very well. Performing TTA with this noise and using stratified cross validation I trained a model that combined my OOF predictions slightly better than a straightforward linear blend in terms of CV score.</p>\n<p>What I am saying is that if the right amount of noise is added (something with less variance than that of the model that is having noise added to it) then the combined OOF predictions are difficult to distinguish from the combined submissions. As this was the case when using the same seed every time, I imagine it will also be the case when you do not have the same seed. And, in your case, it would also ensure that any of the OOF predictions that are similar become difficult to distinguish, removing their ability to provide artificial gains in the combined OOF predictions.</p>\n<p>I hope that makes sense, and it would be interesting to know if it works for you!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2919597,
      "author_name": "Mohammad Sadat Hossain",
      "author_url": "",
      "post_date": "2024-07-13T03:10:49.583000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, this was a nice discussion. But I was wondering maybe if you could give some advice in context of code competitions, like the <a href=\"https://www.kaggle.com/competitions/isic-2024-challenge\" target=\"_blank\">current ISIC competition</a>. Here, we do not have the entire test set, and hence need to do the full inference within the limit of 9 hours. </p>\n<p>Say, I have trained using 5-fold cross validation. So essentially I have 5 models to use for inference. Which one should I choose? Infering using all of them and then averaging just takes too much time.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2017546,
      "author_name": "Myo Min Htet(wnp)🇲🇲",
      "author_url": "",
      "post_date": "2022-11-04T23:54:03.187000",
      "content": "<p>The best concept I've ever seen. And I always learn new things from you.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1871067,
      "author_name": "Matoo",
      "author_url": "",
      "post_date": "2022-07-26T04:03:15.060000",
      "content": "<p>Thank you for sharing your Hill Climbing approach.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1822540,
      "author_name": "highscoreman",
      "author_url": "",
      "post_date": "2022-06-16T13:40:08.877000",
      "content": "<p>thanks for sharing. hill climbing,good job.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1729854,
      "author_name": "CHIRAGhj",
      "author_url": "",
      "post_date": "2022-03-20T15:53:22.940000",
      "content": "<p>Thankyou helped a lot</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1654624,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-01-18T17:11:02.800000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1118700,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-19T10:11:24.077000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 976390,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T20:18:06.273000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 976409,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T20:26:26.717000",
          "content": "",
          "votes": 7,
          "replies": []
        },
        {
          "id": 977235,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-19T11:13:17.920000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1946594,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-19T23:39:58.900000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1947541,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-09-20T14:14:43.423000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 987263,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-27T05:58:34.753000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 986044,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-26T06:57:01.327000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 983744,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-24T14:54:53.317000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 976322,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T19:35:36.740000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 976303,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T19:22:13.703000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1797454,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-22T01:26:52.757000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "976295": "One thing I learned at Kaggle is how to ensemble models using OOF files. This was very helpful in Melanoma Competition. I would like to share a simple procedure below called `hill climbing` and I posted a starter notebook [here][1].  \n\n# Cross Validation\nThe acronym CV refers to cross validation. We start with the full training dataset picture below to the far left. Next, we divide it into 5 subsets, called  `Fold 1, Fold 2, Fold 3, Fold 4, Fold 5`. Then we train 5 models. We train our first model using data from Folds 2-5 and predict Fold 1. Next we train model 2 using Folds 1, 3, 4, 5 and predict 2. Next 1, 2, 4, 5 and predict 3, etc etc.\n\nAfterward we have predictions for every training image. This compete set of predictions is called OOF, \"out of fold\" predictions. It is a good practice to save these predictions for every model you build during a competition as `oof.csv`.\n\nThe CV score (or OOF AUC) is then calculated with `OOF_AUC = roc_auc_score( train.target, oof.prediction)`. And this is the best indicator of how your model performs. It is a better indicator than public LB.\n\nDuring the competition, every model should use the same folds. This is accomplished by using the same seed with `sklearn.model_selection.KFold(n_splits = 5, shuffle = True, random_seed = 42)`\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F79668c5b80f69cbe78b5ee91deb6798b%2Fkfold.png?generation=1597776677559528&alt=media)\n\n# Submission Files\nFor each of the 5 fold models above, we predict the test images. Therefore we have 5 predictions for each test image. We take the average of these 5 sets of predictions and this is our `submission.csv` file that we submit to Kaggle. When you submit this to Kaggle, your LB score should be similar to your CV score.\n\n# Model Ensemble\nNow say that you build 2 models (that means that you did 5 KFold twice). You now have `oof_1.csv`, `oof_2.csv`, `sub_1.csv`, and `sub_2.csv`. How do we blend the two models?\n\nWe find the weight `w` such that `w * oof_1.predictions + (1-w) * oof_2.predictions` has the largest AUC.\n\n     all = []\n     for w in [0.00, 0.01, 0.02, ..., 0.98, 0.99, 1.00]:\n         ensemble_pred = w * oof_1.predictions + (1-w) * oof_2.predictions\n         ensemble_auc = roc_auc_score( oof.target , ensemble_pred )\n         all.append( ensemble_auc )\n     best_weight = np.argmax( all ) / 100.\n\nThen our submission to kaggle will be\n\n     kaggle_sub = best_weight * sub_1.target + (1-best_weight) * sub_2.target\n\n# Starter Notebook\nAfter weeks working on a competition, we will have more than 2 models. So we need a more sophisticated approach than a single for-loop. \n\nThe simplest approach is `hill climbing` (or `forward selection`). Start with the one model that has highest CV score. Next iterate through all your additional models and find the one model that combines with the first model to generate the highest two model ensemble CV score. Then search for the best third model. Repeat until ensemble CV does not increase anymore.\n\nI posted a starter notebook [here][1] which includes all my `oof.csv` and `sub.csv` for this Melanoma comp. There are 39 models. Forward selection chooses 8 of them and the resultant ensemble has OOF CV 0.950, Public LB 0.958, and Private LB 0.942\n\n[1]: https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private",
    "976302": "Kagglers refer to this as `hill climbing` just as a side note if others see this phrase somewhere else.",
    "979176": "Thanks for sharing this! Good explanation for CV!\n\nAfter reading some discussions in the past few days after the competition end, I gradually realized that what more people need is not the top solutions, but the basic ML knowledge, just like this posts.\n\nI see many people still don't understand why they should not trust public LB score (especially in this comp). But I don’t know how to explain to them in a simple and understandable way.\n\nI think your post can be a good explanation for it as well. Thanks again!",
    "1654315": "Thanks @cdeotte, your Hill Climbing approach is a great trick. I'll for sure try it out in my next competition.\nI have small doubt, in the post you've linked you say \n>During the competition, every model should use the same folds. \n\nDoes that really matter when your oof predictions have all the train samples anyway? I specifically ask in a context of ensembling models as a team, where different members worked with different splits.",
    "988805": "Hi @cdeotte, have you seen this paper \"On Cross-Validation and Stacking: Building seemingly predictive models on random data\"?\nhttps://www.kdd.org/exploration_files/v12-02-4-UR-Perlich.pdf \n\nI'd be interested to know what you and other Kagglers think about it",
    "979325": "Hi Chris, the explanation is very clear but I have a question and after reading the post a couple of times I can't figure out the answer on my own. Talking about the submission file, you say\n> For each of the 5 fold models above, we predict the test images. Therefore we have 5 predictions for each test image. We take the average of these 5 sets of predictions and this is our submission.csv file that we submit to Kaggle. When you submit this to Kaggle, your LB score should be similar to your CV score.\n\nand this makes perfect sense, but why do you never train a model using all train data once you have a high CV score for that model?  \nIs it because you would not know how it performs when trained with an additional fold, making all the work useless and destroying the ensembling process?  \nI am asking because so far, when I perform cross-validation, I always fit the model on all train data with the best parameters to make the test predictions and so your way of doing it by averaging is completely new to me.\nThank you!",
    "976417": "The strategy I usually go for is finding a set of models that give the best CV when averaged together.\n\nDo you have any idea how your ensemble performs (especially on the leaderboard) when giving equal weights to every model ? \nWhat about using the 8  (or any number of) models that have the best CV when averging the predictions together ? \n\nThis would be nice to see how much hill climbing actually helps.",
    "998817": "Suppose we are using a model and I have to choose the best hyperparameters for that model. Now, will it be right that for each hyperparameter I do a new cv and note down the best parameters? By doing this am I introducing any leakage?",
    "997407": "Thanks for sharing. I do follow these approaches on Stacking while working on the projects.\nHave you tried Voting Classifiers as well?",
    "988133": "A good explaination, was very helpful.",
    "987504": "Good one",
    "983330": "Very crisp, clear & to the point explanation. I am a novice at kaggle and I am learning amazing stuff reading your notebook & discussion. Thanks!",
    "983068": "Thanks for sharing great knowledge about ensembling! \n\nI didn't know hill climbing approach to ensemble the model and helped me a lot! ",
    "979255": "Because of hardware limitations I have been trying this....\n\nfor each epoch in 1:30\ntrain for 100 steps; in each  step use a batch of n images that is a random sample of the train set with augmentations (resampled at each step)\nvalidate for 30 steps; in each step use a batch of n images that is a random sample of the train set without augmentations (resampled at each step)\n\nn is 16 32 64 depending on image size\n\nDifferent runs for the same model got consistent scores between \"this cv score\" and private set... not with public set thas was a bit lower...\n\nchanging the set of images at each step may result in underfitting? it certainly did not overfit",
    "977172": "Thanks a lot for sharing and it is really good explanation for oof ensemble!\n\nMay I ask something:\nSome Kagglers save their folds to `folds.csv` and just load `folds.csv` for more experiments.\n\nIf I use same seed(`e.g. random_seed = 42`), is it same folds result?\n\n`new_folds = sklearn.model_selection.KFold(n_splits = 5, shuffle = True, random_seed = 42)`\n`fold.csv == new_folds`\n\nThanks again @cdeotte.",
    "976954": "What clarity!\nI knew that intuitively, but I wasn't confident enough in my methods to implement it. I want to take advantage of this in the future.",
    "976589": "Simple yet very efficient, thanks @cdeotte , for me it seems obvious now, but before this competition I never saved my model's OOF predictions, after seeing your public notebook this is a practice that I will keep from now on!",
    "976583": "Could you clarify why you do a for loop instead of fitting a logistic regression or other optimization approaches?",
    "976448": "> Also every model should use the same folds. \n\nIs it necessary to have the same folds? I guess we can have different folds OOF csv files. Will it cause leak?",
    "976391": "This is great. As a beginner, ensembling has always puzzled me, What are more creative ways to do ensembling?",
    "976370": "Gonna bookmark this! thanks for the detailed explanation.",
    "976311": "thanks @cdeotte ! \nI used your starter notebook and had 12 models. \ni did ensembling using another starter kernel. \nthis is my final ensemble with 6 models (Private LB = 0.9413). \nhttps://www.kaggle.com/mpsampat/final-simple-oof-ensembling-methods-6-models\n\nI had a hard time to find the optimal subset of models. Your starter notebook will be very handy in learning this task! \nthanks again for sharing this and other notebooks! \nCheers! ",
    "985887": "> For each of the 5 fold models above, we predict the test images. Therefore we have 5 predictions for each test image.\n\n\nAmazing article. I keep revisiting this again & again. I was wondering how there are 5 predictions for the test image since the 5 different folds will be a part of single model? So essentially, one model will predict one output for a single test image? I understand for different models the prediction of test image will vary but how is it varying for single model (and different folds)? Do we take the average for test image output for different models or for different folds of the same model?",
    "976706": "Thanks for sharing Chris! Probably the clearest explanation of this I’ve seen.\n\nIs it necessary to use the same folds you used to train your models or can this CV scheme be independent of the initial scheme?",
    "976329": "Nice brute force algorithm. I will definitely give it a try to improve my optimization techniques. How would you do this for OOF predictions for models of different seeds?",
    "2919597": "@cdeotte, this was a nice discussion. But I was wondering maybe if you could give some advice in context of code competitions, like the [current ISIC competition](https://www.kaggle.com/competitions/isic-2024-challenge). Here, we do not have the entire test set, and hence need to do the full inference within the limit of 9 hours. \n\nSay, I have trained using 5-fold cross validation. So essentially I have 5 models to use for inference. Which one should I choose? Infering using all of them and then averaging just takes too much time.",
    "2017546": "The best concept I've ever seen. And I always learn new things from you.",
    "1871067": "Thank you for sharing your Hill Climbing approach.",
    "1822540": "thanks for sharing. hill climbing,good job.",
    "1729854": "Thankyou helped a lot",
    "1654624": "Thanks for sharing this!\nIts been really helpful.",
    "1118700": "Great post! Crisp and Clear! Thanks for sharing",
    "976390": "Are you by finding these optimal weights not overfitting the validation set?",
    "1946594": "",
    "987263": "Great stuff. Thanks for explaining ",
    "986044": "nice explanation \nThanks for sharing",
    "983744": "thanks for sharing this!",
    "976322": "Thanks sir",
    "976303": "This the light!! Thanks!",
    "1797454": "Thanks for sharing, extremely helpful!!"
  }
}