{
  "id": 174865,
  "title": "Early Stopping observations",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/174865",
  "author_name": "",
  "post_date": "2020-08-15T21:01:19.700233900Z",
  "votes": 1,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I use Early Stopping, but obviously I always wonder \"What if I allowed it to go one more Epoch?\".  What I have done to put me at ease, is every now and then I'll run like up to 40 epochs and then observe \"Did the model ever improve after where I would have set my stop?\".  And typically its no, or if it did improve it was insignificant.  From this, I have set my stopping to be 5, but sometimes I'll set it to 8, and I really don't feel like I am leaving any points on the table.</p>\n<p>My LR decreases if no improvement in 3 Epochs.  So when I have my stopping set to 8, it allows me to have 2 LR drops and then some before I call it quits, that is what I like about setting it to 8 sometimes.</p>\n<p>What about you?  What are your thoughts on Early Stopping?</p>",
  "messages": [
    {
      "id": "971738",
      "postDate": "08/15/2020 21:01:19",
      "content": "<p>I use Early Stopping, but obviously I always wonder \"What if I allowed it to go one more Epoch?\".  What I have done to put me at ease, is every now and then I'll run like up to 40 epochs and then observe \"Did the model ever improve after where I would have set my stop?\".  And typically its no, or if it did improve it was insignificant.  From this, I have set my stopping to be 5, but sometimes I'll set it to 8, and I really don't feel like I am leaving any points on the table.</p>\n<p>My LR decreases if no improvement in 3 Epochs.  So when I have my stopping set to 8, it allows me to have 2 LR drops and then some before I call it quits, that is what I like about setting it to 8 sometimes.</p>\n<p>What about you?  What are your thoughts on Early Stopping?</p>",
      "rawMarkdown": "I use Early Stopping, but obviously I always wonder \"What if I allowed it to go one more Epoch?\".  What I have done to put me at ease, is every now and then I'll run like up to 40 epochs and then observe \"Did the model ever improve after where I would have set my stop?\".  And typically its no, or if it did improve it was insignificant.  From this, I have set my stopping to be 5, but sometimes I'll set it to 8, and I really don't feel like I am leaving any points on the table.\n\nMy LR decreases if no improvement in 3 Epochs.  So when I have my stopping set to 8, it allows me to have 2 LR drops and then some before I call it quits, that is what I like about setting it to 8 sometimes.\n\nWhat about you?  What are your thoughts on Early Stopping?",
      "votes": null
    },
    {
      "id": "971868",
      "postDate": "08/16/2020 02:53:55",
      "content": "<p>An alternative to early stopping is <code>tf.keras.callbacks.ModelCheckpoint()</code>. Then you don't need to choose a <code>patience</code> variable. You just set your training to a high number of epochs and the best metric fold is saved. Then when training stops, you load the weights from the best fold.</p>\n<p>Of course the downside is  training for all those extra epochs after AUC plateaus or starts to decrease. So in this comp where each epoch takes a while it may not be best, but in other comps where epochs are fast, i prefer this.</p>\n<p>(Then after you train a few models, you learn to set the max number of epochs efficiently).</p>",
      "rawMarkdown": "An alternative to early stopping is `tf.keras.callbacks.ModelCheckpoint()`. Then you don't need to choose a `patience` variable. You just set your training to a high number of epochs and the best metric fold is saved. Then when training stops, you load the weights from the best fold.\n  \nOf course the downside is  training for all those extra epochs after AUC plateaus or starts to decrease. So in this comp where each epoch takes a while it may not be best, but in other comps where epochs are fast, i prefer this.\n\n(Then after you train a few models, you learn to set the max number of epochs efficiently).",
      "votes": null
    },
    {
      "id": "971884",
      "postDate": "08/16/2020 03:34:00",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> yes i do that too! I save the model each time a better metric is achieved.  I use the same name per fold so it just overwrites the last best.  And yes, to your point the epochs can take some time.  </p>",
      "rawMarkdown": "cdeotte yes i do that too! I save the model each time a better metric is achieved.  I use the same name per fold so it just overwrites the last best.  And yes, to your point the epochs can take some time.",
      "votes": null
    },
    {
      "id": "971885",
      "postDate": "08/16/2020 03:37:58",
      "content": "<p>I started with fixed amount of epoch approach first, in past week decided to use early stopping. Usually starting with big steps of learning and when loss hits the plateau decreasing the learning rate gets me out of it, but you usually can find optima if you watch some of your models training…</p>",
      "rawMarkdown": "I started with fixed amount of epoch approach first, in past week decided to use early stopping. Usually starting with big steps of learning and when loss hits the plateau decreasing the learning rate gets me out of it, but you usually can find optima if you watch some of your models training...",
      "votes": null
    },
    {
      "id": "977963",
      "postDate": "08/19/2020 20:25:02",
      "content": "<p>Early stopping is rarely a good idea as it overfits to validation set and overestimates your CV. I usually only use it to get a feeling of good scheduling, and then always switch to fixed epochs.</p>",
      "rawMarkdown": "Early stopping is rarely a good idea as it overfits to validation set and overestimates your CV. I usually only use it to get a feeling of good scheduling, and then always switch to fixed epochs.",
      "votes": null
    },
    {
      "id": "977970",
      "postDate": "08/19/2020 20:32:09",
      "content": "<p>I used a fixed number of epochs and tuned the rest so that the last epoch is the one with best CV.</p>",
      "rawMarkdown": "I used a fixed number of epochs and tuned the rest so that the last epoch is the one with best CV.",
      "votes": null
    },
    {
      "id": "978016",
      "postDate": "08/19/2020 21:31:06",
      "content": "<p>Would you please elaborate on this, I don't know how you can tune training so that the last epoch is the one with the best CV? I thought by nature this training was stochastic</p>",
      "rawMarkdown": "Would you please elaborate on this, I don't know how you can tune training so that the last epoch is the one with the best CV? I thought by nature this training was stochastic",
      "votes": null
    },
    {
      "id": "979130",
      "postDate": "08/20/2020 16:32:01",
      "content": "<p>Interesting strategy. I always thought of not using early stopping as a lost opportunity. When doing it, you can benefit from some diversity when combining test predictions from different folds, where additional diversity comes from having a different number of epochs for each base model. But I also see how early stopping can overfit each validation fold, leading to an overoptimistic CV. In other words, you overestimate your CV but you can also get a gain on LB. I am guessing that the trade-off between the two effects might be very different depending on model/data.</p>",
      "rawMarkdown": "Interesting strategy. I always thought of not using early stopping as a lost opportunity. When doing it, you can benefit from some diversity when combining test predictions from different folds, where additional diversity comes from having a different number of epochs for each base model. But I also see how early stopping can overfit each validation fold, leading to an overoptimistic CV. In other words, you overestimate your CV but you can also get a gain on LB. I am guessing that the trade-off between the two effects might be very different depending on model/data.",
      "votes": null
    },
    {
      "id": "979155",
      "postDate": "08/20/2020 16:57:59",
      "content": "<p>I am not sure I understand your question.  When you tune your model you should compute CV score epoch per epoch. Then you can see the effect of your tuning on CV per epoch.</p>",
      "rawMarkdown": "I am not sure I understand your question.  When you tune your model you should compute CV score epoch per epoch. Then you can see the effect of your tuning on CV per epoch.",
      "votes": null
    },
    {
      "id": "979519",
      "postDate": "08/20/2020 22:59:12",
      "content": "<p>Sorry for not being clear, and sorry to turn your small comment into all this text, but I would like to understand this.</p>\n<p>Typically when I perform cross validation I train one model at a time and calculate my overall CV score once using the final model from each fold. What I thought you meant was that you tuned so that the best score for each fold was in the last epoch. But if I now understand you correctly what I should do is save the validation scores for the entire history of each fold, calculate the overall CV score per epoch, and then tune so that the last epoch is the one with the best CV?</p>\n<p>Now that I have thought about this I have a couple of other questions.</p>\n<p>For CV score do you use the AUC in this case? Selecting anything with this method seems unreliable to me because the AUC is only a semi-proper score. Also the validation history isn't calculated using TTA and so we only have a shaky estimate of our scores during training.</p>\n<p>Also, do you select the number of epochs that you will use ahead of time and then tune to that? If that is the case I do not understand why this method makes sense because instead of tuning to have the best CV score, you tune to have the best CV score in the last epoch. I would assume that the latter CV score would be lower than the former, and I am not sure why models trained in this way would generalise better.</p>\n<p>Thanks if you read this, and hopefully it makes sense.</p>",
      "rawMarkdown": "Sorry for not being clear, and sorry to turn your small comment into all this text, but I would like to understand this.\n\nTypically when I perform cross validation I train one model at a time and calculate my overall CV score once using the final model from each fold. What I thought you meant was that you tuned so that the best score for each fold was in the last epoch. But if I now understand you correctly what I should do is save the validation scores for the entire history of each fold, calculate the overall CV score per epoch, and then tune so that the last epoch is the one with the best CV?\n\nNow that I have thought about this I have a couple of other questions.\n\nFor CV score do you use the AUC in this case? Selecting anything with this method seems unreliable to me because the AUC is only a semi-proper score. Also the validation history isn't calculated using TTA and so we only have a shaky estimate of our scores during training.\n\nAlso, do you select the number of epochs that you will use ahead of time and then tune to that? If that is the case I do not understand why this method makes sense because instead of tuning to have the best CV score, you tune to have the best CV score in the last epoch. I would assume that the latter CV score would be lower than the former, and I am not sure why models trained in this way would generalise better.\n\nThanks if you read this, and hopefully it makes sense.",
      "votes": null
    },
    {
      "id": "979927",
      "postDate": "08/21/2020 07:54:25",
      "content": "<p>You are right, there is a potential to gain diversity, but it heavily decreases the reliability of your CV usually. If blending different epochs is helpful, I would rather go the direction of fixed checkpoint ensembling. And what I always try to do if the complexity allows it is to blend different seeds of the same model and fold to even further strengthen my CV.</p>",
      "rawMarkdown": "You are right, there is a potential to gain diversity, but it heavily decreases the reliability of your CV usually. If blending different epochs is helpful, I would rather go the direction of fixed checkpoint ensembling. And what I always try to do if the complexity allows it is to blend different seeds of the same model and fold to even further strengthen my CV.",
      "votes": null
    },
    {
      "id": "980120",
      "postDate": "08/21/2020 10:40:38",
      "content": "<p>I save validation predictions (out of fold) after each epoch for each fold.  Once all folds are trained I can then get out of fold predictions for all training data, for every epoch.  Then, after training I compute the roc-auc for each epoch.  This is how you can select which epoch is best.</p>\n<p>Selecting best epoch fold by fold is called early stopping, and it can lead to overfitting to the validation data.  Using the same epoch for all folds is safer when you have overfitting issues, which is the case here.</p>\n<p>I use TTA for validation too.  What you do for validation should be as close as possible to what you do for test predictions.  I agree that without TTA validation scores are too noisy.</p>\n<p>The goal iI set was not to get the best CV.  The goal was to train models without overfitting to validation data.</p>",
      "rawMarkdown": "I save validation predictions (out of fold) after each epoch for each fold.  Once all folds are trained I can then get out of fold predictions for all training data, for every epoch.  Then, after training I compute the roc-auc for each epoch.  This is how you can select which epoch is best.\n\nSelecting best epoch fold by fold is called early stopping, and it can lead to overfitting to the validation data.  Using the same epoch for all folds is safer when you have overfitting issues, which is the case here.\n\nI use TTA for validation too.  What you do for validation should be as close as possible to what you do for test predictions.  I agree that without TTA validation scores are too noisy.\n\nThe goal iI set was not to get the best CV.  The goal was to train models without overfitting to validation data.",
      "votes": null
    },
    {
      "id": "980986",
      "postDate": "08/22/2020 04:45:07",
      "content": "<p>Thank you for replying, that is really helpful and the method makes sense to me now. I will have to give this a try.</p>",
      "rawMarkdown": "Thank you for replying, that is really helpful and the method makes sense to me now. I will have to give this a try.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 971868,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/16/2020 02:53:55",
      "content": "<p>An alternative to early stopping is <code>tf.keras.callbacks.ModelCheckpoint()</code>. Then you don't need to choose a <code>patience</code> variable. You just set your training to a high number of epochs and the best metric fold is saved. Then when training stops, you load the weights from the best fold.</p>\n<p>Of course the downside is  training for all those extra epochs after AUC plateaus or starts to decrease. So in this comp where each epoch takes a while it may not be best, but in other comps where epochs are fast, i prefer this.</p>\n<p>(Then after you train a few models, you learn to set the max number of epochs efficiently).</p>",
      "votes": null,
      "replies": [
        {
          "id": 971884,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/16/2020 03:34:00",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> yes i do that too! I save the model each time a better metric is achieved.  I use the same name per fold so it just overwrites the last best.  And yes, to your point the epochs can take some time.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 971885,
      "author_name": "datafan07",
      "author_url": "",
      "post_date": "08/16/2020 03:37:58",
      "content": "<p>I started with fixed amount of epoch approach first, in past week decided to use early stopping. Usually starting with big steps of learning and when loss hits the plateau decreasing the learning rate gets me out of it, but you usually can find optima if you watch some of your models training…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 977963,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "08/19/2020 20:25:02",
      "content": "<p>Early stopping is rarely a good idea as it overfits to validation set and overestimates your CV. I usually only use it to get a feeling of good scheduling, and then always switch to fixed epochs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 979130,
          "author_name": "kozodoi",
          "author_url": "",
          "post_date": "08/20/2020 16:32:01",
          "content": "<p>Interesting strategy. I always thought of not using early stopping as a lost opportunity. When doing it, you can benefit from some diversity when combining test predictions from different folds, where additional diversity comes from having a different number of epochs for each base model. But I also see how early stopping can overfit each validation fold, leading to an overoptimistic CV. In other words, you overestimate your CV but you can also get a gain on LB. I am guessing that the trade-off between the two effects might be very different depending on model/data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979927,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "08/21/2020 07:54:25",
          "content": "<p>You are right, there is a potential to gain diversity, but it heavily decreases the reliability of your CV usually. If blending different epochs is helpful, I would rather go the direction of fixed checkpoint ensembling. And what I always try to do if the complexity allows it is to blend different seeds of the same model and fold to even further strengthen my CV.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 977970,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/19/2020 20:32:09",
      "content": "<p>I used a fixed number of epochs and tuned the rest so that the last epoch is the one with best CV.</p>",
      "votes": null,
      "replies": [
        {
          "id": 978016,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "08/19/2020 21:31:06",
          "content": "<p>Would you please elaborate on this, I don't know how you can tune training so that the last epoch is the one with the best CV? I thought by nature this training was stochastic</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979155,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/20/2020 16:57:59",
          "content": "<p>I am not sure I understand your question.  When you tune your model you should compute CV score epoch per epoch. Then you can see the effect of your tuning on CV per epoch.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979519,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "08/20/2020 22:59:12",
          "content": "<p>Sorry for not being clear, and sorry to turn your small comment into all this text, but I would like to understand this.</p>\n<p>Typically when I perform cross validation I train one model at a time and calculate my overall CV score once using the final model from each fold. What I thought you meant was that you tuned so that the best score for each fold was in the last epoch. But if I now understand you correctly what I should do is save the validation scores for the entire history of each fold, calculate the overall CV score per epoch, and then tune so that the last epoch is the one with the best CV?</p>\n<p>Now that I have thought about this I have a couple of other questions.</p>\n<p>For CV score do you use the AUC in this case? Selecting anything with this method seems unreliable to me because the AUC is only a semi-proper score. Also the validation history isn't calculated using TTA and so we only have a shaky estimate of our scores during training.</p>\n<p>Also, do you select the number of epochs that you will use ahead of time and then tune to that? If that is the case I do not understand why this method makes sense because instead of tuning to have the best CV score, you tune to have the best CV score in the last epoch. I would assume that the latter CV score would be lower than the former, and I am not sure why models trained in this way would generalise better.</p>\n<p>Thanks if you read this, and hopefully it makes sense.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 980120,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/21/2020 10:40:38",
          "content": "<p>I save validation predictions (out of fold) after each epoch for each fold.  Once all folds are trained I can then get out of fold predictions for all training data, for every epoch.  Then, after training I compute the roc-auc for each epoch.  This is how you can select which epoch is best.</p>\n<p>Selecting best epoch fold by fold is called early stopping, and it can lead to overfitting to the validation data.  Using the same epoch for all folds is safer when you have overfitting issues, which is the case here.</p>\n<p>I use TTA for validation too.  What you do for validation should be as close as possible to what you do for test predictions.  I agree that without TTA validation scores are too noisy.</p>\n<p>The goal iI set was not to get the best CV.  The goal was to train models without overfitting to validation data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 980986,
          "author_name": "samklein",
          "author_url": "",
          "post_date": "08/22/2020 04:45:07",
          "content": "<p>Thank you for replying, that is really helpful and the method makes sense to me now. I will have to give this a try.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "971738": "I use Early Stopping, but obviously I always wonder \"What if I allowed it to go one more Epoch?\".  What I have done to put me at ease, is every now and then I'll run like up to 40 epochs and then observe \"Did the model ever improve after where I would have set my stop?\".  And typically its no, or if it did improve it was insignificant.  From this, I have set my stopping to be 5, but sometimes I'll set it to 8, and I really don't feel like I am leaving any points on the table.\n\nMy LR decreases if no improvement in 3 Epochs.  So when I have my stopping set to 8, it allows me to have 2 LR drops and then some before I call it quits, that is what I like about setting it to 8 sometimes.\n\nWhat about you?  What are your thoughts on Early Stopping?",
    "971868": "An alternative to early stopping is `tf.keras.callbacks.ModelCheckpoint()`. Then you don't need to choose a `patience` variable. You just set your training to a high number of epochs and the best metric fold is saved. Then when training stops, you load the weights from the best fold.\n  \nOf course the downside is  training for all those extra epochs after AUC plateaus or starts to decrease. So in this comp where each epoch takes a while it may not be best, but in other comps where epochs are fast, i prefer this.\n\n(Then after you train a few models, you learn to set the max number of epochs efficiently).",
    "971884": "cdeotte yes i do that too! I save the model each time a better metric is achieved.  I use the same name per fold so it just overwrites the last best.  And yes, to your point the epochs can take some time.",
    "971885": "I started with fixed amount of epoch approach first, in past week decided to use early stopping. Usually starting with big steps of learning and when loss hits the plateau decreasing the learning rate gets me out of it, but you usually can find optima if you watch some of your models training...",
    "977963": "Early stopping is rarely a good idea as it overfits to validation set and overestimates your CV. I usually only use it to get a feeling of good scheduling, and then always switch to fixed epochs.",
    "977970": "I used a fixed number of epochs and tuned the rest so that the last epoch is the one with best CV.",
    "978016": "Would you please elaborate on this, I don't know how you can tune training so that the last epoch is the one with the best CV? I thought by nature this training was stochastic",
    "979130": "Interesting strategy. I always thought of not using early stopping as a lost opportunity. When doing it, you can benefit from some diversity when combining test predictions from different folds, where additional diversity comes from having a different number of epochs for each base model. But I also see how early stopping can overfit each validation fold, leading to an overoptimistic CV. In other words, you overestimate your CV but you can also get a gain on LB. I am guessing that the trade-off between the two effects might be very different depending on model/data.",
    "979155": "I am not sure I understand your question.  When you tune your model you should compute CV score epoch per epoch. Then you can see the effect of your tuning on CV per epoch.",
    "979519": "Sorry for not being clear, and sorry to turn your small comment into all this text, but I would like to understand this.\n\nTypically when I perform cross validation I train one model at a time and calculate my overall CV score once using the final model from each fold. What I thought you meant was that you tuned so that the best score for each fold was in the last epoch. But if I now understand you correctly what I should do is save the validation scores for the entire history of each fold, calculate the overall CV score per epoch, and then tune so that the last epoch is the one with the best CV?\n\nNow that I have thought about this I have a couple of other questions.\n\nFor CV score do you use the AUC in this case? Selecting anything with this method seems unreliable to me because the AUC is only a semi-proper score. Also the validation history isn't calculated using TTA and so we only have a shaky estimate of our scores during training.\n\nAlso, do you select the number of epochs that you will use ahead of time and then tune to that? If that is the case I do not understand why this method makes sense because instead of tuning to have the best CV score, you tune to have the best CV score in the last epoch. I would assume that the latter CV score would be lower than the former, and I am not sure why models trained in this way would generalise better.\n\nThanks if you read this, and hopefully it makes sense.",
    "979927": "You are right, there is a potential to gain diversity, but it heavily decreases the reliability of your CV usually. If blending different epochs is helpful, I would rather go the direction of fixed checkpoint ensembling. And what I always try to do if the complexity allows it is to blend different seeds of the same model and fold to even further strengthen my CV.",
    "980120": "I save validation predictions (out of fold) after each epoch for each fold.  Once all folds are trained I can then get out of fold predictions for all training data, for every epoch.  Then, after training I compute the roc-auc for each epoch.  This is how you can select which epoch is best.\n\nSelecting best epoch fold by fold is called early stopping, and it can lead to overfitting to the validation data.  Using the same epoch for all folds is safer when you have overfitting issues, which is the case here.\n\nI use TTA for validation too.  What you do for validation should be as close as possible to what you do for test predictions.  I agree that without TTA validation scores are too noisy.\n\nThe goal iI set was not to get the best CV.  The goal was to train models without overfitting to validation data.",
    "980986": "Thank you for replying, that is really helpful and the method makes sense to me now. I will have to give this a try."
  },
  "source": "meta"
}