{
  "id": 91328,
  "title": "Early stopping",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/91328",
  "author_name": "CPMP",
  "post_date": "2019-05-03T09:33:24.214000",
  "votes": 52,
  "comment_count": 68,
  "views": 0,
  "content": "<p>I keep reading that early stopping leads to overfitting.  This is not what I observed in past competitions, and it is not what I am observing here either.  Also, using it or not does not change my CV or my LB much.  </p>\n\n<p>But the large number of people mentioning it here made me think.  I suspect that the cause of overfititng is not early stopping but rather the large number of features.  If one uses 400 features for training data with 4k samples, then the risk of overfititng is extremely high.  If you use 14 features like me then the risk is way smaller.</p>\n\n<p>With many features you need more regularization, and enforcing the same number of trees per fold is one form of regularization indeed.</p>\n\n<p>I welcome contradiction or different viewpoints because I am not sure I understand why we have such different experience</p>\n\n<p>Edit: <a href=\"/philippsinger\">@philippsinger</a> made me test some hypothesis.  Results with running a set of train/test split experiments show that there is  no correlation between the length of the validation fold and the number of trees with early stopping. None.</p>\n\n<p>It also confirms that there is a correlation with the difference of mean. The higher the validation mean minus the training fold mean, the more trees. See the attached picture.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/527025/13152/cor.png\" alt=\"cor\"></p>",
  "messages": [
    {
      "id": 526566,
      "postDate": "2019-05-03T09:33:24.213Z",
      "content": "<p>I keep reading that early stopping leads to overfitting.  This is not what I observed in past competitions, and it is not what I am observing here either.  Also, using it or not does not change my CV or my LB much.  </p>\n\n<p>But the large number of people mentioning it here made me think.  I suspect that the cause of overfititng is not early stopping but rather the large number of features.  If one uses 400 features for training data with 4k samples, then the risk of overfititng is extremely high.  If you use 14 features like me then the risk is way smaller.</p>\n\n<p>With many features you need more regularization, and enforcing the same number of trees per fold is one form of regularization indeed.</p>\n\n<p>I welcome contradiction or different viewpoints because I am not sure I understand why we have such different experience</p>\n\n<p>Edit: <a href=\"/philippsinger\">@philippsinger</a> made me test some hypothesis.  Results with running a set of train/test split experiments show that there is  no correlation between the length of the validation fold and the number of trees with early stopping. None.</p>\n\n<p>It also confirms that there is a correlation with the difference of mean. The higher the validation mean minus the training fold mean, the more trees. See the attached picture.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/527025/13152/cor.png\" alt=\"cor\"></p>",
      "rawMarkdown": "I keep reading that early stopping leads to overfitting.  This is not what I observed in past competitions, and it is not what I am observing here either.  Also, using it or not does not change my CV or my LB much.  \n\nBut the large number of people mentioning it here made me think.  I suspect that the cause of overfititng is not early stopping but rather the large number of features.  If one uses 400 features for training data with 4k samples, then the risk of overfititng is extremely high.  If you use 14 features like me then the risk is way smaller.\n\nWith many features you need more regularization, and enforcing the same number of trees per fold is one form of regularization indeed.\n\nI welcome contradiction or different viewpoints because I am not sure I understand why we have such different experience\n\nEdit: @philippsinger made me test some hypothesis.  Results with running a set of train/test split experiments show that there is  no correlation between the length of the validation fold and the number of trees with early stopping. None.\n\nIt also confirms that there is a correlation with the difference of mean. The higher the validation mean minus the training fold mean, the more trees. See the attached picture.\n\n![cor](https://storage.googleapis.com/kaggle-forum-message-attachments/527025/13152/cor.png)\n",
      "votes": 52
    },
    {
      "id": 527168,
      "postDate": "2019-05-04T18:52:59.227Z",
      "content": "<p>In my experience using early_stopping per fold overfits a bit each fold individually and it can be a problem if you are going to stack models and train a level 2 predictor.  So using early_stop combined with a crossvalidation strategy like lgb.cv or xgb.cv functions is the best way to avoid overfit. Of course, overfit can be relative, just loading the data and taking a look at it is enough to start overfitting :) </p>",
      "rawMarkdown": "In my experience using early_stopping per fold overfits a bit each fold individually and it can be a problem if you are going to stack models and train a level 2 predictor.  So using early_stop combined with a crossvalidation strategy like lgb.cv or xgb.cv functions is the best way to avoid overfit. Of course, overfit can be relative, just loading the data and taking a look at it is enough to start overfitting :) ",
      "votes": 11,
      "replies": [
        {
          "id": 527305,
          "postDate": "2019-05-05T04:26:33.183Z",
          "content": "<p>Right, someone  else also mentioned impact on stacking.  I'll check next time Im' stacking.  But for some reason, stacking wasn't effective in recent competitions I entered ;)</p>",
          "rawMarkdown": "Right, someone  else also mentioned impact on stacking.  I'll check next time Im' stacking.  But for some reason, stacking wasn't effective in recent competitions I entered ;)",
          "votes": 2
        },
        {
          "id": 528304,
          "postDate": "2019-05-07T13:01:57.147Z",
          "content": "<p>123</p>",
          "rawMarkdown": "123",
          "votes": -6
        },
        {
          "id": 528305,
          "postDate": "2019-05-07T13:02:05.377Z",
          "content": "<p>15254</p>",
          "rawMarkdown": "15254",
          "votes": -6
        },
        {
          "id": 528699,
          "postDate": "2019-05-08T12:01:13.620Z",
          "content": "<p><a href=\"/titericz\">@titericz</a> Giba,\nHow do you use lgb.cv to avoid overfit?\ndo you use it on entire train set  and then use the resulting num_rounds for each of the folds?</p>\n\n<p>Thanks!</p>",
          "rawMarkdown": "@titericz Giba,\nHow do you use lgb.cv to avoid overfit?\ndo you use it on entire train set  and then use the resulting num_rounds for each of the folds?\n\nThanks!",
          "votes": 1
        },
        {
          "id": 528738,
          "postDate": "2019-05-08T13:23:08.727Z",
          "content": "<p><a href=\"/tpthegreat\">@tpthegreat</a> Amir, I believe he implied using two layers of folds. The outer folds to get OOF predictions and inner folds (with lgb.cv) for early stopping. If there are N1 outer folds and N2 inner folds then there are N1*N2 models to train. lgb.cv here is just a convenient implementation of this inner loop of folds.</p>",
          "rawMarkdown": "@tpthegreat Amir, I believe he implied using two layers of folds. The outer folds to get OOF predictions and inner folds (with lgb.cv) for early stopping. If there are N1 outer folds and N2 inner folds then there are N1*N2 models to train. lgb.cv here is just a convenient implementation of this inner loop of folds.",
          "votes": 1
        },
        {
          "id": 534769,
          "postDate": "2019-05-21T19:53:34.927Z",
          "content": "<p>When we use the underscore sign <code>_</code> it starts a italic formating that goes until next <code>_</code>. Some who wants to avoid this can use the ` (<a href=\"https://en.wikipedia.org/wiki/Grave_accent\">Grave accent</a>) around the variable he uses underscore.</p>",
          "rawMarkdown": "When we use the underscore sign `_` it starts a italic formating that goes until next `_`. Some who wants to avoid this can use the \\` ([Grave accent](https://en.wikipedia.org/wiki/Grave_accent)) around the variable he uses underscore."
        }
      ]
    },
    {
      "id": 527163,
      "postDate": "2019-05-04T18:35:23.123Z",
      "content": "<p>Interesting discussion. I'll add my 2 cents. </p>\n\n<p>1st cent. Is early stopping an overfitting? Technically, of course. With early stopping we are using our validation fold for directly tuning <code>num_rounds</code> parameter. Imagine we have 5 folds A, B, C, D, E - when making OOF predictions for fold A we are training on the remaining folds B, C, D, E and with early stopping we'll get the best possible <code>num_rounds</code> value for fold A. And we make predictions for fold A with this best <code>num_rounds</code>. And then we use the same <code>num_rounds</code> for predicting test set - obviously for test set this won't be \"best possible\" <code>num_rounds</code>. So OOF predictions for fold A will be slightly \"better\" (=overfitted) than for test set. And the same goes for all other folds - we predict all OOF folds with <code>num_rounds</code> trained specifically for that fold.</p>\n\n<p>2nd cent. Is that a problem? Actually no. At least mostly. I'm using OOF with early stopping very often and as <a href=\"/cpmpml\">@cpmpml</a> already mentioned - nothing bad happens. CV OOF scores are quite similar with test scores. Why? In typical setup when making predictions for test data, we are averaging predictions made for each fold. And averaging usually increases score, what in turn compensates slightly worse predictions for each fold. So CV OOF predictions are little bit overfitted, but test predictions are made better by averaging several fold predictions - so both scores could turn out to be quite similar in practice. </p>",
      "rawMarkdown": "Interesting discussion. I'll add my 2 cents. \n\n1st cent. Is early stopping an overfitting? Technically, of course. With early stopping we are using our validation fold for directly tuning `num_rounds` parameter. Imagine we have 5 folds A, B, C, D, E - when making OOF predictions for fold A we are training on the remaining folds B, C, D, E and with early stopping we'll get the best possible `num_rounds` value for fold A. And we make predictions for fold A with this best `num_rounds`. And then we use the same `num_rounds` for predicting test set - obviously for test set this won't be \"best possible\" `num_rounds`. So OOF predictions for fold A will be slightly \"better\" (=overfitted) than for test set. And the same goes for all other folds - we predict all OOF folds with `num_rounds` trained specifically for that fold.\n\n2nd cent. Is that a problem? Actually no. At least mostly. I'm using OOF with early stopping very often and as @cpmpml already mentioned - nothing bad happens. CV OOF scores are quite similar with test scores. Why? In typical setup when making predictions for test data, we are averaging predictions made for each fold. And averaging usually increases score, what in turn compensates slightly worse predictions for each fold. So CV OOF predictions are little bit overfitted, but test predictions are made better by averaging several fold predictions - so both scores could turn out to be quite similar in practice. ",
      "votes": 7
    },
    {
      "id": 526594,
      "postDate": "2019-05-03T11:00:59.043Z",
      "content": "<p>Early stopping is a problem if you have a bias of earthquake length in your validation set that you stop on, for example when you do 1-EQ-OOF CV strategy. What happens is that the booster usually starts with the mean of the training data and the more iterations you train on the more it deviates from it potentially if the mean on validation data is different (longer or shorter). So for different folds you will use quite different number of boosting rounds which can lead to quite some overfitting. You can check then that test predictions are very different in each fold, blending them helps of course, might still introduce some biases though. With shuffled validation it's much less of a problem.</p>",
      "rawMarkdown": "Early stopping is a problem if you have a bias of earthquake length in your validation set that you stop on, for example when you do 1-EQ-OOF CV strategy. What happens is that the booster usually starts with the mean of the training data and the more iterations you train on the more it deviates from it potentially if the mean on validation data is different (longer or shorter). So for different folds you will use quite different number of boosting rounds which can lead to quite some overfitting. You can check then that test predictions are very different in each fold, blending them helps of course, might still introduce some biases though. With shuffled validation it's much less of a problem.",
      "votes": 7,
      "replies": [
        {
          "id": 526627,
          "postDate": "2019-05-03T12:33:48.930Z",
          "content": "<p>Thanks.</p>\n\n<p>I don't get your point.  Sure, early stopping will result in different number of trees per fold.  Why does it lead to overfitting?  What does the number of trees has to do with quality of prediction?</p>",
          "rawMarkdown": "Thanks.\n\nI don't get your point.  Sure, early stopping will result in different number of trees per fold.  Why does it lead to overfitting?  What does the number of trees has to do with quality of prediction?"
        },
        {
          "id": 526655,
          "postDate": "2019-05-03T13:44:25.483Z",
          "content": "<p>Because the optimal number of trees is very different dependent on how long the earthquake (i.e. the mean of the eq is) and you basically \"leak\" this information as you optimize towards it. </p>",
          "rawMarkdown": "Because the optimal number of trees is very different dependent on how long the earthquake (i.e. the mean of the eq is) and you basically \"leak\" this information as you optimize towards it. ",
          "votes": 1
        },
        {
          "id": 526822,
          "postDate": "2019-05-03T21:09:35.867Z",
          "content": "<blockquote>\n  <p>What happens is that the booster usually starts with the mean of the training data and the more iterations you train on the more it deviates from it potentially if the mean on validation data is different (longer or shorter).</p>\n</blockquote>\n\n<p>I disagree, prediction does not depend on validation data.  </p>",
          "rawMarkdown": "&gt;  What happens is that the booster usually starts with the mean of the training data and the more iterations you train on the more it deviates from it potentially if the mean on validation data is different (longer or shorter).\n\nI disagree, prediction does not depend on validation data.  "
        },
        {
          "id": 527025,
          "postDate": "2019-05-04T11:57:31.880Z",
          "content": "<p>I checked <a href=\"/philippsinger\">@philippsinger</a> hypothesis by running a large number of fold training.  It confirms what I suspected: there is absolutely no correlation between the length of the validation fold and the number of trees with early stopping.  None.  </p>\n\n<p>It also confirms that there is a correlation with the difference of mean. However the correlation is not that number of trees increases with the absolute difference of mean.  In reality, the higher the validation mean minus the training fold mean, the more trees.  See the attached picture.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/527025/13152/cor.png\" alt=\"cor\"></p>\n\n<p>Disclaimer: this my hold only for my folds.</p>",
          "rawMarkdown": "I checked @philippsinger hypothesis by running a large number of fold training.  It confirms what I suspected: there is absolutely no correlation between the length of the validation fold and the number of trees with early stopping.  None.  \n\nIt also confirms that there is a correlation with the difference of mean. However the correlation is not that number of trees increases with the absolute difference of mean.  In reality, the higher the validation mean minus the training fold mean, the more trees.  See the attached picture.\n\n![cor](https://storage.googleapis.com/kaggle-forum-message-attachments/527025/13152/cor.png)\n\nDisclaimer: this my hold only for my folds.",
          "votes": 2
        },
        {
          "id": 527034,
          "postDate": "2019-05-04T12:32:33.183Z",
          "content": "<p>In 1-EQ-OOF the length of validation fold is proportional to its mean.</p>",
          "rawMarkdown": "In 1-EQ-OOF the length of validation fold is proportional to its mean.",
          "votes": 1
        },
        {
          "id": 527035,
          "postDate": "2019-05-04T12:38:48.920Z",
          "content": "<p>How are your validation folds looking like? What I mean is not the absolute length of the validation set if you mix multiple EQs, but rather if you put a single EQ in and the respective length of that EQ. And as <a href=\"/mykper\">@mykper</a> says, this correlates with the mean of the validation set which is also what you are observing (as you then leave this EQ out of train).</p>",
          "rawMarkdown": "How are your validation folds looking like? What I mean is not the absolute length of the validation set if you mix multiple EQs, but rather if you put a single EQ in and the respective length of that EQ. And as @mykper says, this correlates with the mean of the validation set which is also what you are observing (as you then leave this EQ out of train)."
        },
        {
          "id": 527065,
          "postDate": "2019-05-04T13:51:50.073Z",
          "content": "<blockquote>\n  <p>In 1-EQ-OOF the length of validation fold is proportional to its mean.</p>\n</blockquote>\n\n<p>Then I disclosed I am not using that ;)  I won't say more on what I do.</p>",
          "rawMarkdown": "&gt; In 1-EQ-OOF the length of validation fold is proportional to its mean.\n\nThen I disclosed I am not using that ;)  I won't say more on what I do."
        },
        {
          "id": 527074,
          "postDate": "2019-05-04T14:22:43.400Z",
          "content": "<p>I got downvoted for saying I disagree with this:</p>\n\n<p>&gt; What happens is that the booster usually starts with the mean of the training data and the more iterations you train on the more it deviates from it potentially if the mean on validation data is different (longer or shorter).</p>\n\n<p>I checked it.  The figure below shows this is not true, at least for my way of creating folds.  x is the number of iterations, while y is the test prediction mean minus train prediction mean.   We see that more trains does not mean larger deviation.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/527074/13154/cor1.png\" alt=\"cor\"></p>",
          "rawMarkdown": "I got downvoted for saying I disagree with this:\n\n&gt; What happens is that the booster usually starts with the mean of the training data and the more iterations you train on the more it deviates from it potentially if the mean on validation data is different (longer or shorter).\n\nI checked it.  The figure below shows this is not true, at least for my way of creating folds.  x is the number of iterations, while y is the test prediction mean minus train prediction mean.   We see that more trains does not mean larger deviation.\n\n![cor](https://storage.googleapis.com/kaggle-forum-message-attachments/527074/13154/cor1.png)",
          "votes": 1
        },
        {
          "id": 527103,
          "postDate": "2019-05-04T15:27:54.120Z",
          "content": "<p>One thing that can explain why we don't understand each others.  Given train and val are a split of all train data, their target means are inversely correlated.  Therefore, if an outcome is correlated with train fold mean, then it is also correlated with validation fold mean.  This does not mean that the second correlation is a causation.  </p>\n\n<p>To find out who is right we should split a subset of train so that train fold target mean and val fold target mean aren't inversely correlate anymore and look at what is most correlated with test pred mean. </p>",
          "rawMarkdown": "One thing that can explain why we don't understand each others.  Given train and val are a split of all train data, their target means are inversely correlated.  Therefore, if an outcome is correlated with train fold mean, then it is also correlated with validation fold mean.  This does not mean that the second correlation is a causation.  \n\nTo find out who is right we should split a subset of train so that train fold target mean and val fold target mean aren't inversely correlate anymore and look at what is most correlated with test pred mean. ",
          "votes": 1
        },
        {
          "id": 527256,
          "postDate": "2019-05-05T00:02:57.777Z",
          "content": "<blockquote>\n  <p>One thing that can explain why we don't understand each others. Given train and val are a split of all train data, their target means are inversely correlated. Therefore, if an outcome is correlated with train fold mean, then it is also correlated with validation fold mean. This does not mean that the second correlation is a causation.</p>\n</blockquote>\n\n<p>I believe that you found what we had expected all along. The training should be done on a set with a similar mean (for TTF) as the test set. This mean can be deducted from the published papers. </p>",
          "rawMarkdown": "&gt; One thing that can explain why we don't understand each others. Given train and val are a split of all train data, their target means are inversely correlated. Therefore, if an outcome is correlated with train fold mean, then it is also correlated with validation fold mean. This does not mean that the second correlation is a causation.\n\nI believe that you found what we had expected all along. The training should be done on a set with a similar mean (for TTF) as the test set. This mean can be deducted from the published papers. "
        }
      ]
    },
    {
      "id": 526840,
      "postDate": "2019-05-03T22:39:40.620Z",
      "content": "<p>Using early stopping does lead to overfitting if considered from the stacking point of view. In case your stacking second stage model has some OOF predictions which were obtained with early stopping on itself, it will overfit on such predictions (targets are leaked into features...). Not sure if the magnitude of this effect justifies considering it. Thank you for initiating meaningful discussion!</p>",
      "rawMarkdown": "Using early stopping does lead to overfitting if considered from the stacking point of view. In case your stacking second stage model has some OOF predictions which were obtained with early stopping on itself, it will overfit on such predictions (targets are leaked into features...). Not sure if the magnitude of this effect justifies considering it. Thank you for initiating meaningful discussion!",
      "votes": 5,
      "replies": [
        {
          "id": 526879,
          "postDate": "2019-05-04T02:25:47.427Z",
          "content": "<p>This is very interesting.  I did not think of consequences on stacking.</p>\n\n<p>I still don't get how target is leaked with early stopping.  I have good experience with stacking on models with early stopping.  But I'll try without to check what you say, next time stacking is relevant.</p>",
          "rawMarkdown": "This is very interesting.  I did not think of consequences on stacking.\n\nI still don't get how target is leaked with early stopping.  I have good experience with stacking on models with early stopping.  But I'll try without to check what you say, next time stacking is relevant.",
          "votes": 3
        },
        {
          "id": 527263,
          "postDate": "2019-05-05T00:45:27.130Z",
          "content": "<p>Hi <a href=\"/zaharch\">@zaharch</a> , I'm new to the concept of \"stacking\". My understanding of stacking is we put 2 training set together, such as <code>vstack</code>. And the use the trained model to re-train on this stacked dataset. </p>\n\n<p>I assume maybe the pre-trained model was trained on one of the training set, but the new stacking trainingset  has a lot of new samples as well. Will this not kill overfitting?</p>",
          "rawMarkdown": "Hi @zaharch , I'm new to the concept of \"stacking\". My understanding of stacking is we put 2 training set together, such as `vstack`. And the use the trained model to re-train on this stacked dataset. \n\nI assume maybe the pre-trained model was trained on one of the training set, but the new stacking trainingset  has a lot of new samples as well. Will this not kill overfitting?"
        },
        {
          "id": 527457,
          "postDate": "2019-05-05T15:17:50.460Z",
          "content": "<p>Stacking means using out of fold predictions of models to train a new model.  See <a href=\"http://blog.kaggle.com/2016/12/27/a-kagglers-guide-to-model-stacking-in-practice/\">http://blog.kaggle.com/2016/12/27/a-kagglers-guide-to-model-stacking-in-practice/</a></p>",
          "rawMarkdown": "Stacking means using out of fold predictions of models to train a new model.  See http://blog.kaggle.com/2016/12/27/a-kagglers-guide-to-model-stacking-in-practice/",
          "votes": 1
        }
      ]
    },
    {
      "id": 526580,
      "postDate": "2019-05-03T10:07:32.717Z",
      "content": "<p>100% agree - when you see a blind fold score worse as the iterations chug along - overfitting is being displayed right before your eyes!  If early stopping doesn't improve your scores it means there isn't enough data or the features aren't great.  I am not up to uncle's level but early stopping has almost always improved my scores for both xgb and lgb and that is why I don't even consider not using it!</p>",
      "rawMarkdown": "100% agree - when you see a blind fold score worse as the iterations chug along - overfitting is being displayed right before your eyes!  If early stopping doesn't improve your scores it means there isn't enough data or the features aren't great.  I am not up to uncle's level but early stopping has almost always improved my scores for both xgb and lgb and that is why I don't even consider not using it!",
      "votes": 5,
      "replies": [
        {
          "id": 526583,
          "postDate": "2019-05-03T10:10:49.730Z",
          "content": "<p>Thanks.  Maybe people not using ES purposely stop <em>before</em> the point where early stopping would stop.  </p>",
          "rawMarkdown": "Thanks.  Maybe people not using ES purposely stop *before* the point where early stopping would stop.  ",
          "votes": 1
        },
        {
          "id": 526586,
          "postDate": "2019-05-03T10:24:30.900Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 526591,
          "postDate": "2019-05-03T10:50:51.957Z",
          "content": "<p>Early stopping stops where the validation metric(s) are best.  If you want to include train metric(s) in it then you should add train to the validation data.  GBMs like XGBoost, LightGBM, etc. have early stopping argument to their training api.</p>",
          "rawMarkdown": "Early stopping stops where the validation metric(s) are best.  If you want to include train metric(s) in it then you should add train to the validation data.  GBMs like XGBoost, LightGBM, etc. have early stopping argument to their training api.",
          "votes": 2
        },
        {
          "id": 526595,
          "postDate": "2019-05-03T11:02:49.327Z",
          "content": "<p>Btw, what is the best practice - use only validation set or both validation and train for early stopping?\nDoes it depend on a problem?</p>\n\n<p>Question update. It seems like the question is not really correct as early stopping can be implemented differently. For example, LightGBM stops when at least one metric of one set does not improve during last  <code>early_stopping_round</code>. XGBOOST uses last metric and last set for early stopping.\nSo let my question be about LightGBM.</p>",
          "rawMarkdown": "Btw, what is the best practice - use only validation set or both validation and train for early stopping?\nDoes it depend on a problem?\n\nQuestion update. It seems like the question is not really correct as early stopping can be implemented differently. For example, LightGBM stops when at least one metric of one set does not improve during last  `early_stopping_round`. XGBOOST uses last metric and last set for early stopping.\nSo let my question be about LightGBM."
        },
        {
          "id": 526597,
          "postDate": "2019-05-03T11:05:03.957Z",
          "content": "<p>If you use training you will usually never stop :)</p>",
          "rawMarkdown": "If you use training you will usually never stop :)"
        },
        {
          "id": 526600,
          "postDate": "2019-05-03T11:07:56.530Z",
          "content": "<p>I disagree, if the metric is not the loss function then you can get early stopping using train alone.  to your point, chances are that validation metric will degrade way before train metric degrades.</p>",
          "rawMarkdown": "I disagree, if the metric is not the loss function then you can get early stopping using train alone.  to your point, chances are that validation metric will degrade way before train metric degrades.",
          "votes": 2
        },
        {
          "id": 526712,
          "postDate": "2019-05-03T15:19:57.277Z",
          "content": "<p>I asked about early stopping based either on validation or validation + train.</p>",
          "rawMarkdown": "I asked about early stopping based either on validation or validation + train."
        },
        {
          "id": 526716,
          "postDate": "2019-05-03T15:31:51.553Z",
          "content": "<blockquote>\n  <p>I asked about early stopping based either on validation or validation + train.</p>\n</blockquote>\n\n<p>Yes, and it probably makes no difference as valid metric will degrade sooner than train metric.  lgb early stopping stops as soon as one valid metric degrades.</p>",
          "rawMarkdown": "&gt; I asked about early stopping based either on validation or validation + train.\n\nYes, and it probably makes no difference as valid metric will degrade sooner than train metric.  lgb early stopping stops as soon as one valid metric degrades.\n\n",
          "votes": 2
        }
      ]
    },
    {
      "id": 526702,
      "postDate": "2019-05-03T15:04:05.647Z",
      "content": "<p>I am blown away impressed that you can do this with just 14 features.  Also the comments about regularization, that helps me understand the model tuning better.</p>",
      "rawMarkdown": "I am blown away impressed that you can do this with just 14 features.  Also the comments about regularization, that helps me understand the model tuning better.",
      "votes": 6
    },
    {
      "id": 526636,
      "postDate": "2019-05-03T12:47:14.030Z",
      "content": "<p>It does lead to overfitting as any other form of \"peaking\" at out-of-fold data, in the end it is about not to do it too much.\nI guess in your setup the easiest way to observe it would be to increase number of folds. I believe with 100-200 folds you will already see how CV improves and with the extreme case of leave-one-out (#folds = size of training data) CV will look amazing even with a handful of features.</p>",
      "rawMarkdown": "It does lead to overfitting as any other form of \"peaking\" at out-of-fold data, in the end it is about not to do it too much.\nI guess in your setup the easiest way to observe it would be to increase number of folds. I believe with 100-200 folds you will already see how CV improves and with the extreme case of leave-one-out (#folds = size of training data) CV will look amazing even with a handful of features.",
      "votes": 3,
      "replies": [
        {
          "id": 526645,
          "postDate": "2019-05-03T13:21:54.843Z",
          "content": "<p>Sure, early stopping leads to more optimistic CV score.  If this is your definition of overfitting then I know why we disagree.</p>\n\n<p>Overfitting means that the model does not work well on unseen data.  It has nothing to do with how good it is on training data.  </p>\n\n<p>To follow your example, will using more folds degrade test data score?  Not sure it will here.  In general it does not.  </p>",
          "rawMarkdown": "Sure, early stopping leads to more optimistic CV score.  If this is your definition of overfitting then I know why we disagree.\n\nOverfitting means that the model does not work well on unseen data.  It has nothing to do with how good it is on training data.  \n\nTo follow your example, will using more folds degrade test data score?  Not sure it will here.  In general it does not.  "
        },
        {
          "id": 526649,
          "postDate": "2019-05-03T13:32:26.213Z",
          "content": "<p>Anyway, thanks for the answer, I agree that the extreme leave one out case would not be good with ES. </p>",
          "rawMarkdown": "Anyway, thanks for the answer, I agree that the extreme leave one out case would not be good with ES. "
        },
        {
          "id": 528515,
          "postDate": "2019-05-08T02:06:49.840Z",
          "content": "<p>Yep,  it seems that the \"over-optimistic\" CV results are being confused with \"over-fitted\" model.</p>",
          "rawMarkdown": "Yep,  it seems that the \"over-optimistic\" CV results are being confused with \"over-fitted\" model.",
          "votes": 3
        },
        {
          "id": 528534,
          "postDate": "2019-05-08T03:21:46.110Z",
          "content": "<p>Correct. Overfitting on single leave-out fold may lead to underfitting OR overfitting of the model in terms of generalisation. </p>",
          "rawMarkdown": "Correct. Overfitting on single leave-out fold may lead to underfitting OR overfitting of the model in terms of generalisation. "
        },
        {
          "id": 528586,
          "postDate": "2019-05-08T05:53:30.850Z",
          "content": "<p><a href=\"/sibmike\">@sibmike</a> You are spot on.</p>",
          "rawMarkdown": "@sibmike You are spot on."
        }
      ]
    },
    {
      "id": 542666,
      "postDate": "2019-06-04T03:09:09.367Z",
      "content": "<p>Thanks to your advice, I was able to pick a best submission and become an expert!</p>",
      "rawMarkdown": "Thanks to your advice, I was able to pick a best submission and become an expert!",
      "votes": 1
    },
    {
      "id": 526917,
      "postDate": "2019-05-04T05:16:25.330Z",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> thank you for sharing your observations of Early stopping and the affect to CV and LB.  The explaination of select numbers of features and parameter tuning help me learn a lot.</p>",
      "rawMarkdown": "@cpmpml thank you for sharing your observations of Early stopping and the affect to CV and LB.  The explaination of select numbers of features and parameter tuning help me learn a lot.",
      "votes": 1
    },
    {
      "id": 526619,
      "postDate": "2019-05-03T12:23:25.020Z",
      "content": "<p>I think it's very useful if you shuffle the data and use it there myself. If you however split by earthquakes and your first quake only consists of a few samples that you validate and early stop on, you get a model that fits those few samples perfectly, which is problematic imo. If you regularize in some other way (less depth, less iterations, ...) you'll get a more realistic score of how your model actually performs, early stopping might give you an overoptimistic cv here and lead to worse results. So i think it's dependent on how your validation data that you stop on looks like.</p>",
      "rawMarkdown": "I think it's very useful if you shuffle the data and use it there myself. If you however split by earthquakes and your first quake only consists of a few samples that you validate and early stop on, you get a model that fits those few samples perfectly, which is problematic imo. If you regularize in some other way (less depth, less iterations, ...) you'll get a more realistic score of how your model actually performs, early stopping might give you an overoptimistic cv here and lead to worse results. So i think it's dependent on how your validation data that you stop on looks like.",
      "votes": 2,
      "replies": [
        {
          "id": 526625,
          "postDate": "2019-05-03T12:29:57.940Z",
          "content": "<p>When you use early stopping then you average all the fold model prediction for the test data.  That each model 'overfits' to its validation fold is not an issue as their mean is what matters.  I put quotes as I don't get why a model would overfit because of early stopping.  The model isn't trained on the validation fold data by definition.</p>\n\n<p>It seems you are more worried about the validation score than the quality of predictions on test data here.  I'm only interested in the latter.</p>",
          "rawMarkdown": "When you use early stopping then you average all the fold model prediction for the test data.  That each model 'overfits' to its validation fold is not an issue as their mean is what matters.  I put quotes as I don't get why a model would overfit because of early stopping.  The model isn't trained on the validation fold data by definition.\n\nIt seems you are more worried about the validation score than the quality of predictions on test data here.  I'm only interested in the latter.",
          "votes": 3
        },
        {
          "id": 526637,
          "postDate": "2019-05-03T12:49:58.370Z",
          "content": "<p>It's not trained on the validation data, but you still use information from the validation data in the training procedure. If your validation data is small (= not representative), you'll average models that are not generalizing too well and hope that by averaging, the result is still good. Also if you just average, then the model which stopped on small quakes (=less information) will have the same weight as a model which had more val data to evaluate performance, which introduces unwanted bias. \nMaybe overfitting was the wrong term, could of course also lead to underfitting models. But who knows, my gut feeling says early stopping with quake-wise split is a bad idea, but personally, i switched to shuffing with stopping and get much better lb, so i might be completely wrong here.</p>",
          "rawMarkdown": "It's not trained on the validation data, but you still use information from the validation data in the training procedure. If your validation data is small (= not representative), you'll average models that are not generalizing too well and hope that by averaging, the result is still good. Also if you just average, then the model which stopped on small quakes (=less information) will have the same weight as a model which had more val data to evaluate performance, which introduces unwanted bias. \nMaybe overfitting was the wrong term, could of course also lead to underfitting models. But who knows, my gut feeling says early stopping with quake-wise split is a bad idea, but personally, i switched to shuffing with stopping and get much better lb, so i might be completely wrong here.",
          "votes": 2
        },
        {
          "id": 526650,
          "postDate": "2019-05-03T13:32:53.200Z",
          "content": "<p>Thanks, you made me think of possible improvements.</p>",
          "rawMarkdown": "Thanks, you made me think of possible improvements.",
          "votes": 1
        },
        {
          "id": 526656,
          "postDate": "2019-05-03T13:45:02.993Z",
          "content": "<blockquote>\n  <p>you still use information from the validation data in the training procedure.</p>\n</blockquote>\n\n<p>Sure.  Isn't the same true when you don't use early stopping?  How do you decide how many iterations to run without looking at validation data?</p>",
          "rawMarkdown": "&gt; you still use information from the validation data in the training procedure.\n\nSure.  Isn't the same true when you don't use early stopping?  How do you decide how many iterations to run without looking at validation data?",
          "votes": 1
        },
        {
          "id": 526662,
          "postDate": "2019-05-03T14:01:20.247Z",
          "content": "<p>Yes it is, i just think it's less problematic to tune hyperparameters where you evaluate perfomance on all samples, than to stop on a small validation set which has a unique distribution . I preferred to set the number of iterations high enough to let models converge and use shallow models to deal with overfitting. But i tried different things and don't have much experience with gradient boosing models,  so you probably know better what works and what doesn't :)</p>",
          "rawMarkdown": "Yes it is, i just think it's less problematic to tune hyperparameters where you evaluate perfomance on all samples, than to stop on a small validation set which has a unique distribution . I preferred to set the number of iterations high enough to let models converge and use shallow models to deal with overfitting. But i tried different things and don't have much experience with gradient boosing models,  so you probably know better what works and what doesn't :)",
          "votes": 1
        },
        {
          "id": 526669,
          "postDate": "2019-05-03T14:12:44.947Z",
          "content": "<blockquote>\n  <p>where you evaluate perfomance on all samples</p>\n</blockquote>\n\n<p>When you use ES you also evaluate performance on all samples when you average scores over all folds. ;)</p>\n\n<p>Anyway, my goal was not to convince people to use early stopping.  It was to understand why people say it overfits.</p>",
          "rawMarkdown": "&gt; where you evaluate perfomance on all samples\n\nWhen you use ES you also evaluate performance on all samples when you average scores over all folds. ;)\n\nAnyway, my goal was not to convince people to use early stopping.  It was to understand why people say it overfits.",
          "votes": 1
        },
        {
          "id": 526848,
          "postDate": "2019-05-03T23:18:28.363Z",
          "content": "<p>I agree with Sven on the underfitting part. Depending on your validation set your model might stop learning too early. At least that the feeling I had too.Should be easy to test actually by flooring the number of tree while keep using ES.</p>",
          "rawMarkdown": "I agree with Sven on the underfitting part. Depending on your validation set your model might stop learning too early. At least that the feeling I had too.Should be easy to test actually by flooring the number of tree while keep using ES.",
          "votes": 1
        },
        {
          "id": 526894,
          "postDate": "2019-05-04T03:24:00.343Z",
          "content": "<p>I guess the confusion might be coming from different viewpoints of the problem. \nWhen you use early stopping with a model that fits a bad validation fold well, you overfit on that validation data and the resulting model won't generalize well. From the the training data point of view, it's probably just underfitting because it stopped training too soon. Hope that makes sense ^^</p>",
          "rawMarkdown": "I guess the confusion might be coming from different viewpoints of the problem. \nWhen you use early stopping with a model that fits a bad validation fold well, you overfit on that validation data and the resulting model won't generalize well. From the the training data point of view, it's probably just underfitting because it stopped training too soon. Hope that makes sense ^^",
          "votes": 1
        },
        {
          "id": 526898,
          "postDate": "2019-05-04T03:56:33.400Z",
          "content": "<p>you are basically saying that early stopping can lead to underfitting, right?</p>\n\n<p>Then you can have interesting arguments with those why say the opposite here ;)</p>",
          "rawMarkdown": "you are basically saying that early stopping can lead to underfitting, right?\n\nThen you can have interesting arguments with those why say the opposite here ;)"
        },
        {
          "id": 527045,
          "postDate": "2019-05-04T13:03:41.203Z",
          "content": "<p>Well i let's say \"believe\" to stay conservative since i can't prove anything, that you'll get a model that fits the validation data too well (=overfitting to validation data), and that at the same time, doesn't fit the training data well enough (=underfitting to the training data). So what is this? Overfitting? Underfitting? Both? None? Leaks and magic? Idk. My point was more that it's probably not good, but some people like you seem to have good results with it. You can also reduce the bad effects by e.g. using multiple quakes in one validation fold, so that the probability of stopping at a bad timing gets reduced, but this comes at the cost of having less training samples, i think it's hard to say what's really best. </p>",
          "rawMarkdown": "Well i let's say \"believe\" to stay conservative since i can't prove anything, that you'll get a model that fits the validation data too well (=overfitting to validation data), and that at the same time, doesn't fit the training data well enough (=underfitting to the training data). So what is this? Overfitting? Underfitting? Both? None? Leaks and magic? Idk. My point was more that it's probably not good, but some people like you seem to have good results with it. You can also reduce the bad effects by e.g. using multiple quakes in one validation fold, so that the probability of stopping at a bad timing gets reduced, but this comes at the cost of having less training samples, i think it's hard to say what's really best. "
        }
      ]
    },
    {
      "id": 526882,
      "postDate": "2019-05-04T02:28:38.390Z",
      "content": "<p>Early stopping never leads to overfitting. It is the way to choose the best validation accuracy over training accuracy also like choosing just right model. If we trained with more and more samples then overfitting on training set is possible. So for overcome these problems we divide features into some batches like  max_features. </p>",
      "rawMarkdown": "Early stopping never leads to overfitting. It is the way to choose the best validation accuracy over training accuracy also like choosing just right model. If we trained with more and more samples then overfitting on training set is possible. So for overcome these problems we divide features into some batches like  max_features. ",
      "votes": 1
    },
    {
      "id": 531071,
      "postDate": "2019-05-14T08:34:23.367Z",
      "content": "<p>I found this definition of overfitting on Kaggle site (<a href=\"https://www.kaggle.com/docs/competitions\">https://www.kaggle.com/docs/competitions</a> ):</p>\n\n<blockquote>\n  <p>It’s very easy to overfit a model, creating something that performs very well on the public leaderboard, but very badly on the private one. This is called overfitting.</p>\n</blockquote>\n\n<p>While I agree that early stopping leads to optimistic CV values, I still don't think it leads to overfitting as defined above.</p>",
      "rawMarkdown": "I found this definition of overfitting on Kaggle site (https://www.kaggle.com/docs/competitions ):\n&gt; It’s very easy to overfit a model, creating something that performs very well on the public leaderboard, but very badly on the private one. This is called overfitting.\n\nWhile I agree that early stopping leads to optimistic CV values, I still don't think it leads to overfitting as defined above.",
      "replies": [
        {
          "id": 531141,
          "postDate": "2019-05-14T11:32:12.377Z",
          "content": "<p>You can overfit to anything. Your definition above is \"overfitting to public leaderboard\". You can also overfit to your CV or whatever else.</p>",
          "rawMarkdown": "You can overfit to anything. Your definition above is \"overfitting to public leaderboard\". You can also overfit to your CV or whatever else.",
          "votes": 1
        },
        {
          "id": 531143,
          "postDate": "2019-05-14T11:32:54.320Z",
          "content": "<p>I like this \"Kaggler\" definition.</p>\n\n<p>Overfitting: \"when your Public LB position is good and your Private LB position is terrible\"</p>\n\n<p>I cannot wait to see how many places I drop ;)</p>",
          "rawMarkdown": "I like this \"Kaggler\" definition.\n\nOverfitting: \"when your Public LB position is good and your Private LB position is terrible\"\n\nI cannot wait to see how many places I drop ;)"
        },
        {
          "id": 531169,
          "postDate": "2019-05-14T12:30:28.483Z",
          "content": "<blockquote>\n  <p>Your definition above </p>\n</blockquote>\n\n<p>Not my definition.</p>",
          "rawMarkdown": "&gt; Your definition above \n\nNot my definition.",
          "votes": 1
        },
        {
          "id": 531174,
          "postDate": "2019-05-14T12:33:48.273Z",
          "content": "<blockquote>\n  <p>I cannot wait to see how many places I drop ;)</p>\n</blockquote>\n\n<p>I can also drop.  There will be a huge shakeup here.</p>\n\n<p>Another nice read about overfitting: <a href=\"http://blog.kaggle.com/2012/07/06/the-dangers-of-overfitting-psychopathy-post-mortem/\">http://blog.kaggle.com/2012/07/06/the-dangers-of-overfitting-psychopathy-post-mortem/</a></p>",
          "rawMarkdown": "&gt; I cannot wait to see how many places I drop ;)\n\nI can also drop.  There will be a huge shakeup here.\n\nAnother nice read about overfitting: http://blog.kaggle.com/2012/07/06/the-dangers-of-overfitting-psychopathy-post-mortem/\n",
          "votes": 2
        },
        {
          "id": 531177,
          "postDate": "2019-05-14T12:37:22.813Z",
          "content": "<p>In the future I would like LANL to get back to us to see if the best models are good at not overfitting this particular experiment or if they they generalize well for future experiments.  Hopefully they will give us some feedback.</p>",
          "rawMarkdown": "In the future I would like LANL to get back to us to see if the best models are good at not overfitting this particular experiment or if they they generalize well for future experiments.  Hopefully they will give us some feedback.",
          "votes": 4
        }
      ]
    },
    {
      "id": 528028,
      "postDate": "2019-05-06T22:32:04.740Z",
      "content": "<p>I am absolutely looking forward to try this again.</p>",
      "rawMarkdown": "I am absolutely looking forward to try this again."
    },
    {
      "id": 527566,
      "postDate": "2019-05-05T20:32:08.357Z",
      "content": "<p>interesting</p>",
      "rawMarkdown": "interesting"
    },
    {
      "id": 527166,
      "postDate": "2019-05-04T18:45:32.770Z",
      "content": "<p>I think early stopping is not effecting the model, as much as its effecting the human who is learning about how to prevent overfitting. When I'm dealing with training, I usually observe the generalization gap up to around 2-3x times further than early stopping would train. For someone who is learning how to deal with overfitting and using early stopping, there is no clear difference between a validation loss that increases exponentially after the optimum, and one that increases by a fixed delta and then remains constant. In this circumstance its kind of easy to see how one can get stuck so long as early stopping is enabled. </p>",
      "rawMarkdown": "I think early stopping is not effecting the model, as much as its effecting the human who is learning about how to prevent overfitting. When I'm dealing with training, I usually observe the generalization gap up to around 2-3x times further than early stopping would train. For someone who is learning how to deal with overfitting and using early stopping, there is no clear difference between a validation loss that increases exponentially after the optimum, and one that increases by a fixed delta and then remains constant. In this circumstance its kind of easy to see how one can get stuck so long as early stopping is enabled. \n\n"
    },
    {
      "id": 526641,
      "postDate": "2019-05-03T13:02:42.470Z",
      "content": "<p>Can I ask you how many observations you have in your training set ? Thank you in advance.</p>",
      "rawMarkdown": "Can I ask you how many observations you have in your training set ? Thank you in advance.",
      "replies": [
        {
          "id": 526677,
          "postDate": "2019-05-03T14:19:29.560Z",
          "content": "<p>You can ask but I won't answer before competition end ;)</p>",
          "rawMarkdown": "You can ask but I won't answer before competition end ;)",
          "votes": 6
        },
        {
          "id": 526681,
          "postDate": "2019-05-03T14:24:48.710Z",
          "content": "<p>hahaha, nice ! looking forward knowing your solution then :p </p>",
          "rawMarkdown": "hahaha, nice ! looking forward knowing your solution then :p ",
          "votes": 3
        }
      ]
    },
    {
      "id": 527003,
      "postDate": "2019-05-04T10:57:24.503Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true
    },
    {
      "id": 526770,
      "postDate": "2019-05-03T18:03:02.353Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 526821,
          "postDate": "2019-05-03T21:08:21.343Z",
          "content": "<blockquote>\n  <p>A compromise is to train on the training dataset but to stop training at the point when performance on a validation dataset starts to degrade</p>\n</blockquote>\n\n<p>That's the definition of early stopping ;)</p>",
          "rawMarkdown": "&gt; A compromise is to train on the training dataset but to stop training at the point when performance on a validation dataset starts to degrade\n\nThat's the definition of early stopping ;)",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 527168,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2019-05-04T18:52:59.227000",
      "content": "<p>In my experience using early_stopping per fold overfits a bit each fold individually and it can be a problem if you are going to stack models and train a level 2 predictor.  So using early_stop combined with a crossvalidation strategy like lgb.cv or xgb.cv functions is the best way to avoid overfit. Of course, overfit can be relative, just loading the data and taking a look at it is enough to start overfitting :) </p>",
      "votes": 11,
      "replies": [
        {
          "id": 527305,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-05T04:26:33.183000",
          "content": "<p>Right, someone  else also mentioned impact on stacking.  I'll check next time Im' stacking.  But for some reason, stacking wasn't effective in recent competitions I entered ;)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 528304,
          "author_name": "John Smith",
          "author_url": "",
          "post_date": "2019-05-07T13:01:57.147000",
          "content": "<p>123</p>",
          "votes": -6,
          "replies": []
        },
        {
          "id": 528305,
          "author_name": "John Smith",
          "author_url": "",
          "post_date": "2019-05-07T13:02:05.377000",
          "content": "<p>15254</p>",
          "votes": -6,
          "replies": []
        },
        {
          "id": 528699,
          "author_name": "AmirH",
          "author_url": "",
          "post_date": "2019-05-08T12:01:13.620000",
          "content": "<p><a href=\"/titericz\">@titericz</a> Giba,\nHow do you use lgb.cv to avoid overfit?\ndo you use it on entire train set  and then use the resulting num_rounds for each of the folds?</p>\n\n<p>Thanks!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 528738,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2019-05-08T13:23:08.727000",
          "content": "<p><a href=\"/tpthegreat\">@tpthegreat</a> Amir, I believe he implied using two layers of folds. The outer folds to get OOF predictions and inner folds (with lgb.cv) for early stopping. If there are N1 outer folds and N2 inner folds then there are N1*N2 models to train. lgb.cv here is just a convenient implementation of this inner loop of folds.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 534769,
          "author_name": "Alex V B",
          "author_url": "",
          "post_date": "2019-05-21T19:53:34.927000",
          "content": "<p>When we use the underscore sign <code>_</code> it starts a italic formating that goes until next <code>_</code>. Some who wants to avoid this can use the ` (<a href=\"https://en.wikipedia.org/wiki/Grave_accent\">Grave accent</a>) around the variable he uses underscore.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 527163,
      "author_name": "alijs",
      "author_url": "",
      "post_date": "2019-05-04T18:35:23.123000",
      "content": "<p>Interesting discussion. I'll add my 2 cents. </p>\n\n<p>1st cent. Is early stopping an overfitting? Technically, of course. With early stopping we are using our validation fold for directly tuning <code>num_rounds</code> parameter. Imagine we have 5 folds A, B, C, D, E - when making OOF predictions for fold A we are training on the remaining folds B, C, D, E and with early stopping we'll get the best possible <code>num_rounds</code> value for fold A. And we make predictions for fold A with this best <code>num_rounds</code>. And then we use the same <code>num_rounds</code> for predicting test set - obviously for test set this won't be \"best possible\" <code>num_rounds</code>. So OOF predictions for fold A will be slightly \"better\" (=overfitted) than for test set. And the same goes for all other folds - we predict all OOF folds with <code>num_rounds</code> trained specifically for that fold.</p>\n\n<p>2nd cent. Is that a problem? Actually no. At least mostly. I'm using OOF with early stopping very often and as <a href=\"/cpmpml\">@cpmpml</a> already mentioned - nothing bad happens. CV OOF scores are quite similar with test scores. Why? In typical setup when making predictions for test data, we are averaging predictions made for each fold. And averaging usually increases score, what in turn compensates slightly worse predictions for each fold. So CV OOF predictions are little bit overfitted, but test predictions are made better by averaging several fold predictions - so both scores could turn out to be quite similar in practice. </p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 526594,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2019-05-03T11:00:59.043000",
      "content": "<p>Early stopping is a problem if you have a bias of earthquake length in your validation set that you stop on, for example when you do 1-EQ-OOF CV strategy. What happens is that the booster usually starts with the mean of the training data and the more iterations you train on the more it deviates from it potentially if the mean on validation data is different (longer or shorter). So for different folds you will use quite different number of boosting rounds which can lead to quite some overfitting. You can check then that test predictions are very different in each fold, blending them helps of course, might still introduce some biases though. With shuffled validation it's much less of a problem.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 526627,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-03T12:33:48.930000",
          "content": "<p>Thanks.</p>\n\n<p>I don't get your point.  Sure, early stopping will result in different number of trees per fold.  Why does it lead to overfitting?  What does the number of trees has to do with quality of prediction?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 526655,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-05-03T13:44:25.483000",
          "content": "<p>Because the optimal number of trees is very different dependent on how long the earthquake (i.e. the mean of the eq is) and you basically \"leak\" this information as you optimize towards it. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 526822,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-03T21:09:35.867000",
          "content": "<blockquote>\n  <p>What happens is that the booster usually starts with the mean of the training data and the more iterations you train on the more it deviates from it potentially if the mean on validation data is different (longer or shorter).</p>\n</blockquote>\n\n<p>I disagree, prediction does not depend on validation data.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 527025,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-04T11:57:31.880000",
          "content": "<p>I checked <a href=\"/philippsinger\">@philippsinger</a> hypothesis by running a large number of fold training.  It confirms what I suspected: there is absolutely no correlation between the length of the validation fold and the number of trees with early stopping.  None.  </p>\n\n<p>It also confirms that there is a correlation with the difference of mean. However the correlation is not that number of trees increases with the absolute difference of mean.  In reality, the higher the validation mean minus the training fold mean, the more trees.  See the attached picture.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/527025/13152/cor.png\" alt=\"cor\"></p>\n\n<p>Disclaimer: this my hold only for my folds.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 527034,
          "author_name": "mykper",
          "author_url": "",
          "post_date": "2019-05-04T12:32:33.183000",
          "content": "<p>In 1-EQ-OOF the length of validation fold is proportional to its mean.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 527035,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-05-04T12:38:48.920000",
          "content": "<p>How are your validation folds looking like? What I mean is not the absolute length of the validation set if you mix multiple EQs, but rather if you put a single EQ in and the respective length of that EQ. And as <a href=\"/mykper\">@mykper</a> says, this correlates with the mean of the validation set which is also what you are observing (as you then leave this EQ out of train).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 527065,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-04T13:51:50.073000",
          "content": "<blockquote>\n  <p>In 1-EQ-OOF the length of validation fold is proportional to its mean.</p>\n</blockquote>\n\n<p>Then I disclosed I am not using that ;)  I won't say more on what I do.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 527074,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-04T14:22:43.400000",
          "content": "<p>I got downvoted for saying I disagree with this:</p>\n\n<p>&gt; What happens is that the booster usually starts with the mean of the training data and the more iterations you train on the more it deviates from it potentially if the mean on validation data is different (longer or shorter).</p>\n\n<p>I checked it.  The figure below shows this is not true, at least for my way of creating folds.  x is the number of iterations, while y is the test prediction mean minus train prediction mean.   We see that more trains does not mean larger deviation.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/527074/13154/cor1.png\" alt=\"cor\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 527103,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-04T15:27:54.120000",
          "content": "<p>One thing that can explain why we don't understand each others.  Given train and val are a split of all train data, their target means are inversely correlated.  Therefore, if an outcome is correlated with train fold mean, then it is also correlated with validation fold mean.  This does not mean that the second correlation is a causation.  </p>\n\n<p>To find out who is right we should split a subset of train so that train fold target mean and val fold target mean aren't inversely correlate anymore and look at what is most correlated with test pred mean. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 527256,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-05-05T00:02:57.777000",
          "content": "<blockquote>\n  <p>One thing that can explain why we don't understand each others. Given train and val are a split of all train data, their target means are inversely correlated. Therefore, if an outcome is correlated with train fold mean, then it is also correlated with validation fold mean. This does not mean that the second correlation is a causation.</p>\n</blockquote>\n\n<p>I believe that you found what we had expected all along. The training should be done on a set with a similar mean (for TTF) as the test set. This mean can be deducted from the published papers. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 526840,
      "author_name": "nosound",
      "author_url": "",
      "post_date": "2019-05-03T22:39:40.620000",
      "content": "<p>Using early stopping does lead to overfitting if considered from the stacking point of view. In case your stacking second stage model has some OOF predictions which were obtained with early stopping on itself, it will overfit on such predictions (targets are leaked into features...). Not sure if the magnitude of this effect justifies considering it. Thank you for initiating meaningful discussion!</p>",
      "votes": 5,
      "replies": [
        {
          "id": 526879,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-04T02:25:47.427000",
          "content": "<p>This is very interesting.  I did not think of consequences on stacking.</p>\n\n<p>I still don't get how target is leaked with early stopping.  I have good experience with stacking on models with early stopping.  But I'll try without to check what you say, next time stacking is relevant.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 527263,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-05T00:45:27.130000",
          "content": "<p>Hi <a href=\"/zaharch\">@zaharch</a> , I'm new to the concept of \"stacking\". My understanding of stacking is we put 2 training set together, such as <code>vstack</code>. And the use the trained model to re-train on this stacked dataset. </p>\n\n<p>I assume maybe the pre-trained model was trained on one of the training set, but the new stacking trainingset  has a lot of new samples as well. Will this not kill overfitting?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 527457,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-05T15:17:50.460000",
          "content": "<p>Stacking means using out of fold predictions of models to train a new model.  See <a href=\"http://blog.kaggle.com/2016/12/27/a-kagglers-guide-to-model-stacking-in-practice/\">http://blog.kaggle.com/2016/12/27/a-kagglers-guide-to-model-stacking-in-practice/</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 526580,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2019-05-03T10:07:32.717000",
      "content": "<p>100% agree - when you see a blind fold score worse as the iterations chug along - overfitting is being displayed right before your eyes!  If early stopping doesn't improve your scores it means there isn't enough data or the features aren't great.  I am not up to uncle's level but early stopping has almost always improved my scores for both xgb and lgb and that is why I don't even consider not using it!</p>",
      "votes": 5,
      "replies": [
        {
          "id": 526583,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-03T10:10:49.730000",
          "content": "<p>Thanks.  Maybe people not using ES purposely stop <em>before</em> the point where early stopping would stop.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 526586,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-03T10:24:30.900000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 526591,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-03T10:50:51.957000",
          "content": "<p>Early stopping stops where the validation metric(s) are best.  If you want to include train metric(s) in it then you should add train to the validation data.  GBMs like XGBoost, LightGBM, etc. have early stopping argument to their training api.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 526595,
          "author_name": "Alexander Firsov",
          "author_url": "",
          "post_date": "2019-05-03T11:02:49.327000",
          "content": "<p>Btw, what is the best practice - use only validation set or both validation and train for early stopping?\nDoes it depend on a problem?</p>\n\n<p>Question update. It seems like the question is not really correct as early stopping can be implemented differently. For example, LightGBM stops when at least one metric of one set does not improve during last  <code>early_stopping_round</code>. XGBOOST uses last metric and last set for early stopping.\nSo let my question be about LightGBM.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 526597,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-05-03T11:05:03.957000",
          "content": "<p>If you use training you will usually never stop :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 526600,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-03T11:07:56.530000",
          "content": "<p>I disagree, if the metric is not the loss function then you can get early stopping using train alone.  to your point, chances are that validation metric will degrade way before train metric degrades.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 526712,
          "author_name": "Alexander Firsov",
          "author_url": "",
          "post_date": "2019-05-03T15:19:57.277000",
          "content": "<p>I asked about early stopping based either on validation or validation + train.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 526716,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-03T15:31:51.553000",
          "content": "<blockquote>\n  <p>I asked about early stopping based either on validation or validation + train.</p>\n</blockquote>\n\n<p>Yes, and it probably makes no difference as valid metric will degrade sooner than train metric.  lgb early stopping stops as soon as one valid metric degrades.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 526702,
      "author_name": "Vettejeep",
      "author_url": "",
      "post_date": "2019-05-03T15:04:05.647000",
      "content": "<p>I am blown away impressed that you can do this with just 14 features.  Also the comments about regularization, that helps me understand the model tuning better.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 526636,
      "author_name": "dott",
      "author_url": "",
      "post_date": "2019-05-03T12:47:14.030000",
      "content": "<p>It does lead to overfitting as any other form of \"peaking\" at out-of-fold data, in the end it is about not to do it too much.\nI guess in your setup the easiest way to observe it would be to increase number of folds. I believe with 100-200 folds you will already see how CV improves and with the extreme case of leave-one-out (#folds = size of training data) CV will look amazing even with a handful of features.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 526645,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-03T13:21:54.843000",
          "content": "<p>Sure, early stopping leads to more optimistic CV score.  If this is your definition of overfitting then I know why we disagree.</p>\n\n<p>Overfitting means that the model does not work well on unseen data.  It has nothing to do with how good it is on training data.  </p>\n\n<p>To follow your example, will using more folds degrade test data score?  Not sure it will here.  In general it does not.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 526649,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-03T13:32:26.213000",
          "content": "<p>Anyway, thanks for the answer, I agree that the extreme leave one out case would not be good with ES. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 528515,
          "author_name": "sibmike",
          "author_url": "",
          "post_date": "2019-05-08T02:06:49.840000",
          "content": "<p>Yep,  it seems that the \"over-optimistic\" CV results are being confused with \"over-fitted\" model.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 528534,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-05-08T03:21:46.110000",
          "content": "<p>Correct. Overfitting on single leave-out fold may lead to underfitting OR overfitting of the model in terms of generalisation. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 528586,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-08T05:53:30.850000",
          "content": "<p><a href=\"/sibmike\">@sibmike</a> You are spot on.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 542666,
      "author_name": "Yuko Ishizaki",
      "author_url": "",
      "post_date": "2019-06-04T03:09:09.367000",
      "content": "<p>Thanks to your advice, I was able to pick a best submission and become an expert!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 526917,
      "author_name": "Beans",
      "author_url": "",
      "post_date": "2019-05-04T05:16:25.330000",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> thank you for sharing your observations of Early stopping and the affect to CV and LB.  The explaination of select numbers of features and parameter tuning help me learn a lot.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 526619,
      "author_name": "Sven Hinderer",
      "author_url": "",
      "post_date": "2019-05-03T12:23:25.020000",
      "content": "<p>I think it's very useful if you shuffle the data and use it there myself. If you however split by earthquakes and your first quake only consists of a few samples that you validate and early stop on, you get a model that fits those few samples perfectly, which is problematic imo. If you regularize in some other way (less depth, less iterations, ...) you'll get a more realistic score of how your model actually performs, early stopping might give you an overoptimistic cv here and lead to worse results. So i think it's dependent on how your validation data that you stop on looks like.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 526625,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-03T12:29:57.940000",
          "content": "<p>When you use early stopping then you average all the fold model prediction for the test data.  That each model 'overfits' to its validation fold is not an issue as their mean is what matters.  I put quotes as I don't get why a model would overfit because of early stopping.  The model isn't trained on the validation fold data by definition.</p>\n\n<p>It seems you are more worried about the validation score than the quality of predictions on test data here.  I'm only interested in the latter.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 526637,
          "author_name": "Sven Hinderer",
          "author_url": "",
          "post_date": "2019-05-03T12:49:58.370000",
          "content": "<p>It's not trained on the validation data, but you still use information from the validation data in the training procedure. If your validation data is small (= not representative), you'll average models that are not generalizing too well and hope that by averaging, the result is still good. Also if you just average, then the model which stopped on small quakes (=less information) will have the same weight as a model which had more val data to evaluate performance, which introduces unwanted bias. \nMaybe overfitting was the wrong term, could of course also lead to underfitting models. But who knows, my gut feeling says early stopping with quake-wise split is a bad idea, but personally, i switched to shuffing with stopping and get much better lb, so i might be completely wrong here.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 526650,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-03T13:32:53.200000",
          "content": "<p>Thanks, you made me think of possible improvements.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 526656,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-03T13:45:02.993000",
          "content": "<blockquote>\n  <p>you still use information from the validation data in the training procedure.</p>\n</blockquote>\n\n<p>Sure.  Isn't the same true when you don't use early stopping?  How do you decide how many iterations to run without looking at validation data?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 526662,
          "author_name": "Sven Hinderer",
          "author_url": "",
          "post_date": "2019-05-03T14:01:20.247000",
          "content": "<p>Yes it is, i just think it's less problematic to tune hyperparameters where you evaluate perfomance on all samples, than to stop on a small validation set which has a unique distribution . I preferred to set the number of iterations high enough to let models converge and use shallow models to deal with overfitting. But i tried different things and don't have much experience with gradient boosing models,  so you probably know better what works and what doesn't :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 526669,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-03T14:12:44.947000",
          "content": "<blockquote>\n  <p>where you evaluate perfomance on all samples</p>\n</blockquote>\n\n<p>When you use ES you also evaluate performance on all samples when you average scores over all folds. ;)</p>\n\n<p>Anyway, my goal was not to convince people to use early stopping.  It was to understand why people say it overfits.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 526848,
          "author_name": "Antoine",
          "author_url": "",
          "post_date": "2019-05-03T23:18:28.363000",
          "content": "<p>I agree with Sven on the underfitting part. Depending on your validation set your model might stop learning too early. At least that the feeling I had too.Should be easy to test actually by flooring the number of tree while keep using ES.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 526894,
          "author_name": "Sven Hinderer",
          "author_url": "",
          "post_date": "2019-05-04T03:24:00.343000",
          "content": "<p>I guess the confusion might be coming from different viewpoints of the problem. \nWhen you use early stopping with a model that fits a bad validation fold well, you overfit on that validation data and the resulting model won't generalize well. From the the training data point of view, it's probably just underfitting because it stopped training too soon. Hope that makes sense ^^</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 526898,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-04T03:56:33.400000",
          "content": "<p>you are basically saying that early stopping can lead to underfitting, right?</p>\n\n<p>Then you can have interesting arguments with those why say the opposite here ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 527045,
          "author_name": "Sven Hinderer",
          "author_url": "",
          "post_date": "2019-05-04T13:03:41.203000",
          "content": "<p>Well i let's say \"believe\" to stay conservative since i can't prove anything, that you'll get a model that fits the validation data too well (=overfitting to validation data), and that at the same time, doesn't fit the training data well enough (=underfitting to the training data). So what is this? Overfitting? Underfitting? Both? None? Leaks and magic? Idk. My point was more that it's probably not good, but some people like you seem to have good results with it. You can also reduce the bad effects by e.g. using multiple quakes in one validation fold, so that the probability of stopping at a bad timing gets reduced, but this comes at the cost of having less training samples, i think it's hard to say what's really best. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 526882,
      "author_name": "Himanshu Soni",
      "author_url": "",
      "post_date": "2019-05-04T02:28:38.390000",
      "content": "<p>Early stopping never leads to overfitting. It is the way to choose the best validation accuracy over training accuracy also like choosing just right model. If we trained with more and more samples then overfitting on training set is possible. So for overcome these problems we divide features into some batches like  max_features. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 531071,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-05-14T08:34:23.367000",
      "content": "<p>I found this definition of overfitting on Kaggle site (<a href=\"https://www.kaggle.com/docs/competitions\">https://www.kaggle.com/docs/competitions</a> ):</p>\n\n<blockquote>\n  <p>It’s very easy to overfit a model, creating something that performs very well on the public leaderboard, but very badly on the private one. This is called overfitting.</p>\n</blockquote>\n\n<p>While I agree that early stopping leads to optimistic CV values, I still don't think it leads to overfitting as defined above.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 531141,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-05-14T11:32:12.377000",
          "content": "<p>You can overfit to anything. Your definition above is \"overfitting to public leaderboard\". You can also overfit to your CV or whatever else.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 531143,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-05-14T11:32:54.320000",
          "content": "<p>I like this \"Kaggler\" definition.</p>\n\n<p>Overfitting: \"when your Public LB position is good and your Private LB position is terrible\"</p>\n\n<p>I cannot wait to see how many places I drop ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 531169,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-14T12:30:28.483000",
          "content": "<blockquote>\n  <p>Your definition above </p>\n</blockquote>\n\n<p>Not my definition.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 531174,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-14T12:33:48.273000",
          "content": "<blockquote>\n  <p>I cannot wait to see how many places I drop ;)</p>\n</blockquote>\n\n<p>I can also drop.  There will be a huge shakeup here.</p>\n\n<p>Another nice read about overfitting: <a href=\"http://blog.kaggle.com/2012/07/06/the-dangers-of-overfitting-psychopathy-post-mortem/\">http://blog.kaggle.com/2012/07/06/the-dangers-of-overfitting-psychopathy-post-mortem/</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 531177,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-05-14T12:37:22.813000",
          "content": "<p>In the future I would like LANL to get back to us to see if the best models are good at not overfitting this particular experiment or if they they generalize well for future experiments.  Hopefully they will give us some feedback.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 528028,
      "author_name": "Vagish Anand",
      "author_url": "",
      "post_date": "2019-05-06T22:32:04.740000",
      "content": "<p>I am absolutely looking forward to try this again.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 527566,
      "author_name": "Jeff",
      "author_url": "",
      "post_date": "2019-05-05T20:32:08.357000",
      "content": "<p>interesting</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 527166,
      "author_name": "Vannak",
      "author_url": "",
      "post_date": "2019-05-04T18:45:32.770000",
      "content": "<p>I think early stopping is not effecting the model, as much as its effecting the human who is learning about how to prevent overfitting. When I'm dealing with training, I usually observe the generalization gap up to around 2-3x times further than early stopping would train. For someone who is learning how to deal with overfitting and using early stopping, there is no clear difference between a validation loss that increases exponentially after the optimum, and one that increases by a fixed delta and then remains constant. In this circumstance its kind of easy to see how one can get stuck so long as early stopping is enabled. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 526641,
      "author_name": "propower",
      "author_url": "",
      "post_date": "2019-05-03T13:02:42.470000",
      "content": "<p>Can I ask you how many observations you have in your training set ? Thank you in advance.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 526677,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-03T14:19:29.560000",
          "content": "<p>You can ask but I won't answer before competition end ;)</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 526681,
          "author_name": "propower",
          "author_url": "",
          "post_date": "2019-05-03T14:24:48.710000",
          "content": "<p>hahaha, nice ! looking forward knowing your solution then :p </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 527003,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-04T10:57:24.503000",
      "content": "",
      "votes": -2,
      "replies": []
    },
    {
      "id": 526770,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-03T18:03:02.353000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 526821,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-03T21:08:21.343000",
          "content": "<blockquote>\n  <p>A compromise is to train on the training dataset but to stop training at the point when performance on a validation dataset starts to degrade</p>\n</blockquote>\n\n<p>That's the definition of early stopping ;)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "526566": "I keep reading that early stopping leads to overfitting.  This is not what I observed in past competitions, and it is not what I am observing here either.  Also, using it or not does not change my CV or my LB much.  \n\nBut the large number of people mentioning it here made me think.  I suspect that the cause of overfititng is not early stopping but rather the large number of features.  If one uses 400 features for training data with 4k samples, then the risk of overfititng is extremely high.  If you use 14 features like me then the risk is way smaller.\n\nWith many features you need more regularization, and enforcing the same number of trees per fold is one form of regularization indeed.\n\nI welcome contradiction or different viewpoints because I am not sure I understand why we have such different experience\n\nEdit: @philippsinger made me test some hypothesis.  Results with running a set of train/test split experiments show that there is  no correlation between the length of the validation fold and the number of trees with early stopping. None.\n\nIt also confirms that there is a correlation with the difference of mean. The higher the validation mean minus the training fold mean, the more trees. See the attached picture.\n\n![cor](https://storage.googleapis.com/kaggle-forum-message-attachments/527025/13152/cor.png)\n",
    "527168": "In my experience using early_stopping per fold overfits a bit each fold individually and it can be a problem if you are going to stack models and train a level 2 predictor.  So using early_stop combined with a crossvalidation strategy like lgb.cv or xgb.cv functions is the best way to avoid overfit. Of course, overfit can be relative, just loading the data and taking a look at it is enough to start overfitting :) ",
    "527163": "Interesting discussion. I'll add my 2 cents. \n\n1st cent. Is early stopping an overfitting? Technically, of course. With early stopping we are using our validation fold for directly tuning `num_rounds` parameter. Imagine we have 5 folds A, B, C, D, E - when making OOF predictions for fold A we are training on the remaining folds B, C, D, E and with early stopping we'll get the best possible `num_rounds` value for fold A. And we make predictions for fold A with this best `num_rounds`. And then we use the same `num_rounds` for predicting test set - obviously for test set this won't be \"best possible\" `num_rounds`. So OOF predictions for fold A will be slightly \"better\" (=overfitted) than for test set. And the same goes for all other folds - we predict all OOF folds with `num_rounds` trained specifically for that fold.\n\n2nd cent. Is that a problem? Actually no. At least mostly. I'm using OOF with early stopping very often and as @cpmpml already mentioned - nothing bad happens. CV OOF scores are quite similar with test scores. Why? In typical setup when making predictions for test data, we are averaging predictions made for each fold. And averaging usually increases score, what in turn compensates slightly worse predictions for each fold. So CV OOF predictions are little bit overfitted, but test predictions are made better by averaging several fold predictions - so both scores could turn out to be quite similar in practice. ",
    "526594": "Early stopping is a problem if you have a bias of earthquake length in your validation set that you stop on, for example when you do 1-EQ-OOF CV strategy. What happens is that the booster usually starts with the mean of the training data and the more iterations you train on the more it deviates from it potentially if the mean on validation data is different (longer or shorter). So for different folds you will use quite different number of boosting rounds which can lead to quite some overfitting. You can check then that test predictions are very different in each fold, blending them helps of course, might still introduce some biases though. With shuffled validation it's much less of a problem.",
    "526840": "Using early stopping does lead to overfitting if considered from the stacking point of view. In case your stacking second stage model has some OOF predictions which were obtained with early stopping on itself, it will overfit on such predictions (targets are leaked into features...). Not sure if the magnitude of this effect justifies considering it. Thank you for initiating meaningful discussion!",
    "526580": "100% agree - when you see a blind fold score worse as the iterations chug along - overfitting is being displayed right before your eyes!  If early stopping doesn't improve your scores it means there isn't enough data or the features aren't great.  I am not up to uncle's level but early stopping has almost always improved my scores for both xgb and lgb and that is why I don't even consider not using it!",
    "526702": "I am blown away impressed that you can do this with just 14 features.  Also the comments about regularization, that helps me understand the model tuning better.",
    "526636": "It does lead to overfitting as any other form of \"peaking\" at out-of-fold data, in the end it is about not to do it too much.\nI guess in your setup the easiest way to observe it would be to increase number of folds. I believe with 100-200 folds you will already see how CV improves and with the extreme case of leave-one-out (#folds = size of training data) CV will look amazing even with a handful of features.",
    "542666": "Thanks to your advice, I was able to pick a best submission and become an expert!",
    "526917": "@cpmpml thank you for sharing your observations of Early stopping and the affect to CV and LB.  The explaination of select numbers of features and parameter tuning help me learn a lot.",
    "526619": "I think it's very useful if you shuffle the data and use it there myself. If you however split by earthquakes and your first quake only consists of a few samples that you validate and early stop on, you get a model that fits those few samples perfectly, which is problematic imo. If you regularize in some other way (less depth, less iterations, ...) you'll get a more realistic score of how your model actually performs, early stopping might give you an overoptimistic cv here and lead to worse results. So i think it's dependent on how your validation data that you stop on looks like.",
    "526882": "Early stopping never leads to overfitting. It is the way to choose the best validation accuracy over training accuracy also like choosing just right model. If we trained with more and more samples then overfitting on training set is possible. So for overcome these problems we divide features into some batches like  max_features. ",
    "531071": "I found this definition of overfitting on Kaggle site (https://www.kaggle.com/docs/competitions ):\n&gt; It’s very easy to overfit a model, creating something that performs very well on the public leaderboard, but very badly on the private one. This is called overfitting.\n\nWhile I agree that early stopping leads to optimistic CV values, I still don't think it leads to overfitting as defined above.",
    "528028": "I am absolutely looking forward to try this again.",
    "527566": "interesting",
    "527166": "I think early stopping is not effecting the model, as much as its effecting the human who is learning about how to prevent overfitting. When I'm dealing with training, I usually observe the generalization gap up to around 2-3x times further than early stopping would train. For someone who is learning how to deal with overfitting and using early stopping, there is no clear difference between a validation loss that increases exponentially after the optimum, and one that increases by a fixed delta and then remains constant. In this circumstance its kind of easy to see how one can get stuck so long as early stopping is enabled. \n\n",
    "526641": "Can I ask you how many observations you have in your training set ? Thank you in advance.",
    "527003": "",
    "526770": ""
  }
}