{
  "id": 390756,
  "title": "Early Stopping and Cross Validation Interplay",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/390756",
  "author_name": "",
  "post_date": "2023-02-27T05:16:10.937480700Z",
  "votes": 18,
  "comment_count": 7,
  "views": 0,
  "content": "<p>When I cross-validate (CV) a model and use the validation set for <code>early_stopping</code>, each fold of CV often have a different number of early-stopped <code>n_estimators</code> due to different validation sets being used. If you are in the same case, what model do you use for inference? What is the best practice in your opinion? I'm thinking of the following options</p>\n<ol>\n<li><strong>Randomly choosing one of the fold's model</strong>, and thereby assuming that all folds are relatively the same. By experiments, I've found that this is not true, and the folds' scores vary by 0.001-0.003.</li>\n<li><strong>Using all folds' models and averaging out their output to get final predictions</strong>. I notice that this method gives a small bump in LB compared to (1).</li>\n<li><strong>Retraining the model with <code>n_estimators=X</code> and all train data, then using it for inference</strong>. I feel like this is the best approach since we're utilizing all train data without long prediction time (like option 2). However I am not sure what <code>X</code> should be - average of all folds, a random fold, etc.?</li>\n<li><strong>Not using CV for determining <code>n_estimators</code></strong> (suggested in <a href=\"https://datascience.stackexchange.com/questions/74351/what-is-the-proper-way-to-use-early-stopping-with-cross-validation\" target=\"_blank\">this Stack Exchange thread</a>). Instead, use CV for only for tuning other hyper-params. After other params have been fixed, re-train the model on the majority of train data and keep out a small subset as validation for early stopping, and use this final model for inference.</li>\n</ol>",
  "messages": [
    {
      "id": "2160928",
      "postDate": "02/27/2023 05:16:10",
      "content": "<p>When I cross-validate (CV) a model and use the validation set for <code>early_stopping</code>, each fold of CV often have a different number of early-stopped <code>n_estimators</code> due to different validation sets being used. If you are in the same case, what model do you use for inference? What is the best practice in your opinion? I'm thinking of the following options</p>\n<ol>\n<li><strong>Randomly choosing one of the fold's model</strong>, and thereby assuming that all folds are relatively the same. By experiments, I've found that this is not true, and the folds' scores vary by 0.001-0.003.</li>\n<li><strong>Using all folds' models and averaging out their output to get final predictions</strong>. I notice that this method gives a small bump in LB compared to (1).</li>\n<li><strong>Retraining the model with <code>n_estimators=X</code> and all train data, then using it for inference</strong>. I feel like this is the best approach since we're utilizing all train data without long prediction time (like option 2). However I am not sure what <code>X</code> should be - average of all folds, a random fold, etc.?</li>\n<li><strong>Not using CV for determining <code>n_estimators</code></strong> (suggested in <a href=\"https://datascience.stackexchange.com/questions/74351/what-is-the-proper-way-to-use-early-stopping-with-cross-validation\" target=\"_blank\">this Stack Exchange thread</a>). Instead, use CV for only for tuning other hyper-params. After other params have been fixed, re-train the model on the majority of train data and keep out a small subset as validation for early stopping, and use this final model for inference.</li>\n</ol>",
      "rawMarkdown": "When I cross-validate (CV) a model and use the validation set for `early_stopping`, each fold of CV often have a different number of early-stopped `n_estimators` due to different validation sets being used. If you are in the same case, what model do you use for inference? What is the best practice in your opinion? I'm thinking of the following options\n1. **Randomly choosing one of the fold's model**, and thereby assuming that all folds are relatively the same. By experiments, I've found that this is not true, and the folds' scores vary by 0.001-0.003.\n2. **Using all folds' models and averaging out their output to get final predictions**. I notice that this method gives a small bump in LB compared to (1).\n3. **Retraining the model with `n_estimators=X` and all train data, then using it for inference**. I feel like this is the best approach since we're utilizing all train data without long prediction time (like option 2). However I am not sure what `X` should be - average of all folds, a random fold, etc.?\n4. **Not using CV for determining `n_estimators`** (suggested in [this Stack Exchange thread](https://datascience.stackexchange.com/questions/74351/what-is-the-proper-way-to-use-early-stopping-with-cross-validation)). Instead, use CV for only for tuning other hyper-params. After other params have been fixed, re-train the model on the majority of train data and keep out a small subset as validation for early stopping, and use this final model for inference.",
      "votes": null
    },
    {
      "id": "2160953",
      "postDate": "02/27/2023 05:53:01",
      "content": "<p><a href=\"https://www.kaggle.com/hoangnguyen719\" target=\"_blank\">@hoangnguyen719</a>  <br>\nI am not sure about point 4 but I can confirm your point 1 and 2. I've also had the same experience while doing CV.</p>\n<p>Point 3 is an interesting approach and I think that you should take a naive approach to determine n_esrimators.</p>\n<p>But thanks for bringing this up. </p>",
      "rawMarkdown": "hoangnguyen719  \nI am not sure about point 4 but I can confirm your point 1 and 2. I've also had the same experience while doing CV.\n\nPoint 3 is an interesting approach and I think that you should take a naive approach to determine n_esrimators.\n\nBut thanks for bringing this up.",
      "votes": null
    },
    {
      "id": "2161831",
      "postDate": "02/27/2023 19:17:01",
      "content": "<p>I think saving <code>n_estimators</code> for each fold_question and then retraining on all data using <code>n_estimtors = np.mean(estimators[q])</code> or the <code>median</code> is a good strategy</p>",
      "rawMarkdown": "I think saving `n_estimators` for each fold_question and then retraining on all data using `n_estimtors = np.mean(estimators[q])` or the `median` is a good strategy",
      "votes": null
    },
    {
      "id": "2162150",
      "postDate": "02/28/2023 03:04:40",
      "content": "<p>Thanks Reacher! That's also the first thing I thought of to calculate <code>X</code> in option (3) above, but I have not found a good logic to back it. For example, some of my folds' <code>n_estimators</code> are twice as many as other folds', in which case I don't feel that mean is the best option.<br>\nDo you do additional testing after calculating such means? Do you see any differences when the inter-fold differences are large?</p>",
      "rawMarkdown": "Thanks Reacher! That's also the first thing I thought of to calculate `X` in option (3) above, but I have not found a good logic to back it. For example, some of my folds' `n_estimators` are twice as many as other folds', in which case I don't feel that mean is the best option.\nDo you do additional testing after calculating such means? Do you see any differences when the inter-fold differences are large?",
      "votes": null
    },
    {
      "id": "2162853",
      "postDate": "02/28/2023 12:55:02",
      "content": "<p>I haven't checked the variance between folds tbh , since I'm focusing more on FE but you're right if variance is too high choosing mean or median isn't appropriate!</p>",
      "rawMarkdown": "I haven't checked the variance between folds tbh , since I'm focusing more on FE but you're right if variance is too high choosing mean or median isn't appropriate!",
      "votes": null
    },
    {
      "id": "2235600",
      "postDate": "04/26/2023 07:34:04",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hoangnguyen719\" target=\"_blank\">@hoangnguyen719</a> , I believe you should use the largest <code>n_estimators</code>. </p>\n<p>Let me elaborate my reason: if your model early stopped at 100 rounds, this means that your model stopped to learn at 100 rounds, and if you force your model to learn 200 rounds, the accuracy of your model should not reduced (on the condition that the learning rate is not too big so that the global optimal point was not jumped during the learning process)</p>",
      "rawMarkdown": "Hi @hoangnguyen719 , I believe you should use the largest `n_estimators`. \n\nLet me elaborate my reason: if your model early stopped at 100 rounds, this means that your model stopped to learn at 100 rounds, and if you force your model to learn 200 rounds, the accuracy of your model should not reduced (on the condition that the learning rate is not too big so that the global optimal point was not jumped during the learning process)",
      "votes": null
    },
    {
      "id": "2235603",
      "postDate": "04/26/2023 07:39:37",
      "content": "<p>I hoped I made myself clear</p>",
      "rawMarkdown": "I hoped I made myself clear",
      "votes": null
    },
    {
      "id": "2235611",
      "postDate": "04/26/2023 07:51:32",
      "content": "<p>you can take a look at the example below. The model stopped to learn at No.182 round, but ask the model to learn more rounds do not harm the accuracy. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7285296%2F52a4af62650961305bedb748ce65e007%2FCapture.PNG?generation=1682495391935741&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "you can take a look at the example below. The model stopped to learn at No.182 round, but ask the model to learn more rounds do not harm the accuracy. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7285296%2F52a4af62650961305bedb748ce65e007%2FCapture.PNG?generation=1682495391935741&alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2160953,
      "author_name": "surajdengale",
      "author_url": "",
      "post_date": "02/27/2023 05:53:01",
      "content": "<p><a href=\"https://www.kaggle.com/hoangnguyen719\" target=\"_blank\">@hoangnguyen719</a>  <br>\nI am not sure about point 4 but I can confirm your point 1 and 2. I've also had the same experience while doing CV.</p>\n<p>Point 3 is an interesting approach and I think that you should take a naive approach to determine n_esrimators.</p>\n<p>But thanks for bringing this up. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2161831,
      "author_name": "ihebch",
      "author_url": "",
      "post_date": "02/27/2023 19:17:01",
      "content": "<p>I think saving <code>n_estimators</code> for each fold_question and then retraining on all data using <code>n_estimtors = np.mean(estimators[q])</code> or the <code>median</code> is a good strategy</p>",
      "votes": null,
      "replies": [
        {
          "id": 2162150,
          "author_name": "hoangnguyen719",
          "author_url": "",
          "post_date": "02/28/2023 03:04:40",
          "content": "<p>Thanks Reacher! That's also the first thing I thought of to calculate <code>X</code> in option (3) above, but I have not found a good logic to back it. For example, some of my folds' <code>n_estimators</code> are twice as many as other folds', in which case I don't feel that mean is the best option.<br>\nDo you do additional testing after calculating such means? Do you see any differences when the inter-fold differences are large?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2162853,
              "author_name": "ihebch",
              "author_url": "",
              "post_date": "02/28/2023 12:55:02",
              "content": "<p>I haven't checked the variance between folds tbh , since I'm focusing more on FE but you're right if variance is too high choosing mean or median isn't appropriate!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2235600,
      "author_name": "zhangyue325",
      "author_url": "",
      "post_date": "04/26/2023 07:34:04",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hoangnguyen719\" target=\"_blank\">@hoangnguyen719</a> , I believe you should use the largest <code>n_estimators</code>. </p>\n<p>Let me elaborate my reason: if your model early stopped at 100 rounds, this means that your model stopped to learn at 100 rounds, and if you force your model to learn 200 rounds, the accuracy of your model should not reduced (on the condition that the learning rate is not too big so that the global optimal point was not jumped during the learning process)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2235603,
          "author_name": "zhangyue325",
          "author_url": "",
          "post_date": "04/26/2023 07:39:37",
          "content": "<p>I hoped I made myself clear</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2235611,
          "author_name": "zhangyue325",
          "author_url": "",
          "post_date": "04/26/2023 07:51:32",
          "content": "<p>you can take a look at the example below. The model stopped to learn at No.182 round, but ask the model to learn more rounds do not harm the accuracy. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7285296%2F52a4af62650961305bedb748ce65e007%2FCapture.PNG?generation=1682495391935741&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2160928": "When I cross-validate (CV) a model and use the validation set for `early_stopping`, each fold of CV often have a different number of early-stopped `n_estimators` due to different validation sets being used. If you are in the same case, what model do you use for inference? What is the best practice in your opinion? I'm thinking of the following options\n1. **Randomly choosing one of the fold's model**, and thereby assuming that all folds are relatively the same. By experiments, I've found that this is not true, and the folds' scores vary by 0.001-0.003.\n2. **Using all folds' models and averaging out their output to get final predictions**. I notice that this method gives a small bump in LB compared to (1).\n3. **Retraining the model with `n_estimators=X` and all train data, then using it for inference**. I feel like this is the best approach since we're utilizing all train data without long prediction time (like option 2). However I am not sure what `X` should be - average of all folds, a random fold, etc.?\n4. **Not using CV for determining `n_estimators`** (suggested in [this Stack Exchange thread](https://datascience.stackexchange.com/questions/74351/what-is-the-proper-way-to-use-early-stopping-with-cross-validation)). Instead, use CV for only for tuning other hyper-params. After other params have been fixed, re-train the model on the majority of train data and keep out a small subset as validation for early stopping, and use this final model for inference.",
    "2160953": "hoangnguyen719  \nI am not sure about point 4 but I can confirm your point 1 and 2. I've also had the same experience while doing CV.\n\nPoint 3 is an interesting approach and I think that you should take a naive approach to determine n_esrimators.\n\nBut thanks for bringing this up.",
    "2161831": "I think saving `n_estimators` for each fold_question and then retraining on all data using `n_estimtors = np.mean(estimators[q])` or the `median` is a good strategy",
    "2162150": "Thanks Reacher! That's also the first thing I thought of to calculate `X` in option (3) above, but I have not found a good logic to back it. For example, some of my folds' `n_estimators` are twice as many as other folds', in which case I don't feel that mean is the best option.\nDo you do additional testing after calculating such means? Do you see any differences when the inter-fold differences are large?",
    "2162853": "I haven't checked the variance between folds tbh , since I'm focusing more on FE but you're right if variance is too high choosing mean or median isn't appropriate!",
    "2235600": "Hi @hoangnguyen719 , I believe you should use the largest `n_estimators`. \n\nLet me elaborate my reason: if your model early stopped at 100 rounds, this means that your model stopped to learn at 100 rounds, and if you force your model to learn 200 rounds, the accuracy of your model should not reduced (on the condition that the learning rate is not too big so that the global optimal point was not jumped during the learning process)",
    "2235603": "I hoped I made myself clear",
    "2235611": "you can take a look at the example below. The model stopped to learn at No.182 round, but ask the model to learn more rounds do not harm the accuracy. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7285296%2F52a4af62650961305bedb748ce65e007%2FCapture.PNG?generation=1682495391935741&alt=media)"
  },
  "source": "meta"
}