{
  "id": 400989,
  "title": "Hyper-parameter Tuning Strategy (for Tree Boosting)?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/400989",
  "author_name": "",
  "post_date": "2023-04-11T08:24:52.173288100Z",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I have a few questions regarding hyper-parameter tuning for XGBoost/Tree Boosting in general</p>\n<h3>1. <code>n_estimators</code></h3>\n<p>From my perspective, I think there are two ways to find an appropriate <code>n_estimators</code>:</p>\n<ol>\n<li><strong>Directly tuning <code>n_estimators</code></strong>: by fitting on the entire training set and <em>not</em> using early stopping.</li>\n<li><strong>Using early-stopping and letting the model choose its own <code>n_estimators</code></strong>: we'll set an arbitrarily high <code>n_estimators</code>, and keep 10-20% of training data for validation/early stopping. Probably also tune <code>early_stopping_rounds</code> if you can.</li>\n</ol>\n<p>Did you try either of them, or any other method? What do you find work better? I'm more inclined towards option (2) because in this competition we're using a large number of features and therefore there is a high chance of overfitting.</p>\n<h3>2. Important hyper-parameters?</h3>\n<p>If you don't mind sharing, what hyper-parameters have been the most crucial in your tuning? For me they are <code>learning_rate</code>, <code>colsample_bytree</code> and <code>alpha</code>. I also tried tuning <code>learning_rate</code> and <code>lambda</code> but it didn't make much of a difference.</p>",
  "messages": [
    {
      "id": "2217876",
      "postDate": "04/11/2023 08:24:52",
      "content": "<p>I have a few questions regarding hyper-parameter tuning for XGBoost/Tree Boosting in general</p>\n<h3>1. <code>n_estimators</code></h3>\n<p>From my perspective, I think there are two ways to find an appropriate <code>n_estimators</code>:</p>\n<ol>\n<li><strong>Directly tuning <code>n_estimators</code></strong>: by fitting on the entire training set and <em>not</em> using early stopping.</li>\n<li><strong>Using early-stopping and letting the model choose its own <code>n_estimators</code></strong>: we'll set an arbitrarily high <code>n_estimators</code>, and keep 10-20% of training data for validation/early stopping. Probably also tune <code>early_stopping_rounds</code> if you can.</li>\n</ol>\n<p>Did you try either of them, or any other method? What do you find work better? I'm more inclined towards option (2) because in this competition we're using a large number of features and therefore there is a high chance of overfitting.</p>\n<h3>2. Important hyper-parameters?</h3>\n<p>If you don't mind sharing, what hyper-parameters have been the most crucial in your tuning? For me they are <code>learning_rate</code>, <code>colsample_bytree</code> and <code>alpha</code>. I also tried tuning <code>learning_rate</code> and <code>lambda</code> but it didn't make much of a difference.</p>",
      "rawMarkdown": "I have a few questions regarding hyper-parameter tuning for XGBoost/Tree Boosting in general\n### 1. `n_estimators`\nFrom my perspective, I think there are two ways to find an appropriate `n_estimators`:\n1. **Directly tuning `n_estimators`**: by fitting on the entire training set and *not* using early stopping.\n2. **Using early-stopping and letting the model choose its own `n_estimators`**: we'll set an arbitrarily high `n_estimators`, and keep 10-20% of training data for validation/early stopping. Probably also tune `early_stopping_rounds` if you can.\n\nDid you try either of them, or any other method? What do you find work better? I'm more inclined towards option (2) because in this competition we're using a large number of features and therefore there is a high chance of overfitting.\n\n### 2. Important hyper-parameters?\nIf you don't mind sharing, what hyper-parameters have been the most crucial in your tuning? For me they are `learning_rate`, `colsample_bytree` and `alpha`. I also tried tuning `learning_rate` and `lambda` but it didn't make much of a difference.",
      "votes": null
    },
    {
      "id": "2289330",
      "postDate": "06/06/2023 04:58:30",
      "content": "<p>What were your conclusions about hyper-parameter tuning <a href=\"https://www.kaggle.com/hoangnguyen719\" target=\"_blank\">@hoangnguyen719</a> ? I used optuna, optimized n_estimators, learning_rate, colsample_bytree…. but when I do, I get a worse LB 😣</p>",
      "rawMarkdown": "What were your conclusions about hyper-parameter tuning @hoangnguyen719 ? I used optuna, optimized n_estimators, learning_rate, colsample_bytree.... but when I do, I get a worse LB 😣",
      "votes": null
    },
    {
      "id": "2289521",
      "postDate": "06/06/2023 07:57:38",
      "content": "<p>Welcome in \"Art and Science of Hyperparameter Optimizations\" 😁</p>",
      "rawMarkdown": "Welcome in \"Art and Science of Hyperparameter Optimizations\" 😁",
      "votes": null
    },
    {
      "id": "2290406",
      "postDate": "06/06/2023 19:04:40",
      "content": "<p>haha thanks for welcoming me, I see you just scaled to the first place, congratulations 🎉<br>\n<a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> would you recommend me some topic to learn about this? or some other tool I could look at? <br>\nGreetings to Poland , I was there last year, I loved it… </p>",
      "rawMarkdown": "haha thanks for welcoming me, I see you just scaled to the first place, congratulations 🎉\n@remekkinas would you recommend me some topic to learn about this? or some other tool I could look at? \nGreetings to Poland , I was there last year, I loved it...",
      "votes": null
    },
    {
      "id": "2290443",
      "postDate": "06/06/2023 19:38:50",
      "content": "<p>Thank you for greetings :)</p>\n<p>I use this book: <a href=\"https://link.springer.com/book/10.1007/978-981-19-5170-1\" target=\"_blank\">https://link.springer.com/book/10.1007/978-981-19-5170-1</a> <br>\nYou can download it for free now.</p>",
      "rawMarkdown": "Thank you for greetings :)\n\nI use this book: https://link.springer.com/book/10.1007/978-981-19-5170-1 \nYou can download it for free now.",
      "votes": null
    },
    {
      "id": "2290539",
      "postDate": "06/06/2023 23:19:47",
      "content": "<p>Not sure if this is relevant to your experience, but I tried optimizing model hyperparameters for each questions. Although the f1 score was boosted to various degree, the overall macro f1, which is the metric of this competition, was not improved at all. It's a bit tricky than individual model tuning IMHO.</p>",
      "rawMarkdown": "Not sure if this is relevant to your experience, but I tried optimizing model hyperparameters for each questions. Although the f1 score was boosted to various degree, the overall macro f1, which is the metric of this competition, was not improved at all. It's a bit tricky than individual model tuning IMHO.",
      "votes": null
    },
    {
      "id": "2290554",
      "postDate": "06/07/2023 00:01:19",
      "content": "<p>dziękuję bardzo! <br>\nalready downloaded, thanks Remek</p>",
      "rawMarkdown": "dziękuję bardzo! \nalready downloaded, thanks Remek",
      "votes": null
    },
    {
      "id": "2290556",
      "postDate": "06/07/2023 00:03:00",
      "content": "<p>Do you have any idea why that happens? exactly that's why I asked, I thought I was going to boost the overall LB and the result was the opposite, final result decreased 🙁</p>",
      "rawMarkdown": "Do you have any idea why that happens? exactly that's why I asked, I thought I was going to boost the overall LB and the result was the opposite, final result decreased 🙁",
      "votes": null
    },
    {
      "id": "2294204",
      "postDate": "06/09/2023 21:10:54",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/409497\" target=\"_blank\">Here</a> has a pretty good discussion on the comparison between individual f1 score versus overall f1 score. The conclusion is a bit counter-intuitive, but its example illustrates why it is tricky to optimize model for each question.</p>\n<p>Note that for each question, the macro f1 score is the unweighted average of f1 for correct (1) and wrong (0) classes. The final f1 for all 18 questions, hence, is the same unweighted average of f1 scores of correct and wrong classes. I think the crucial part here is to boost the model performance for both classes in order to really see the improvement of the final macro f1.</p>",
      "rawMarkdown": "[Here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/409497) has a pretty good discussion on the comparison between individual f1 score versus overall f1 score. The conclusion is a bit counter-intuitive, but its example illustrates why it is tricky to optimize model for each question.\n\nNote that for each question, the macro f1 score is the unweighted average of f1 for correct (1) and wrong (0) classes. The final f1 for all 18 questions, hence, is the same unweighted average of f1 scores of correct and wrong classes. I think the crucial part here is to boost the model performance for both classes in order to really see the improvement of the final macro f1.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2289330,
      "author_name": "javihm77",
      "author_url": "",
      "post_date": "06/06/2023 04:58:30",
      "content": "<p>What were your conclusions about hyper-parameter tuning <a href=\"https://www.kaggle.com/hoangnguyen719\" target=\"_blank\">@hoangnguyen719</a> ? I used optuna, optimized n_estimators, learning_rate, colsample_bytree…. but when I do, I get a worse LB 😣</p>",
      "votes": null,
      "replies": [
        {
          "id": 2289521,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "06/06/2023 07:57:38",
          "content": "<p>Welcome in \"Art and Science of Hyperparameter Optimizations\" 😁</p>",
          "votes": null,
          "replies": [
            {
              "id": 2290406,
              "author_name": "javihm77",
              "author_url": "",
              "post_date": "06/06/2023 19:04:40",
              "content": "<p>haha thanks for welcoming me, I see you just scaled to the first place, congratulations 🎉<br>\n<a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> would you recommend me some topic to learn about this? or some other tool I could look at? <br>\nGreetings to Poland , I was there last year, I loved it… </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2290443,
                  "author_name": "remekkinas",
                  "author_url": "",
                  "post_date": "06/06/2023 19:38:50",
                  "content": "<p>Thank you for greetings :)</p>\n<p>I use this book: <a href=\"https://link.springer.com/book/10.1007/978-981-19-5170-1\" target=\"_blank\">https://link.springer.com/book/10.1007/978-981-19-5170-1</a> <br>\nYou can download it for free now.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2290554,
                      "author_name": "javihm77",
                      "author_url": "",
                      "post_date": "06/07/2023 00:01:19",
                      "content": "<p>dziękuję bardzo! <br>\nalready downloaded, thanks Remek</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 2290539,
          "author_name": "randwlkr",
          "author_url": "",
          "post_date": "06/06/2023 23:19:47",
          "content": "<p>Not sure if this is relevant to your experience, but I tried optimizing model hyperparameters for each questions. Although the f1 score was boosted to various degree, the overall macro f1, which is the metric of this competition, was not improved at all. It's a bit tricky than individual model tuning IMHO.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2290556,
              "author_name": "javihm77",
              "author_url": "",
              "post_date": "06/07/2023 00:03:00",
              "content": "<p>Do you have any idea why that happens? exactly that's why I asked, I thought I was going to boost the overall LB and the result was the opposite, final result decreased 🙁</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2294204,
                  "author_name": "randwlkr",
                  "author_url": "",
                  "post_date": "06/09/2023 21:10:54",
                  "content": "<p><a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/409497\" target=\"_blank\">Here</a> has a pretty good discussion on the comparison between individual f1 score versus overall f1 score. The conclusion is a bit counter-intuitive, but its example illustrates why it is tricky to optimize model for each question.</p>\n<p>Note that for each question, the macro f1 score is the unweighted average of f1 for correct (1) and wrong (0) classes. The final f1 for all 18 questions, hence, is the same unweighted average of f1 scores of correct and wrong classes. I think the crucial part here is to boost the model performance for both classes in order to really see the improvement of the final macro f1.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2217876": "I have a few questions regarding hyper-parameter tuning for XGBoost/Tree Boosting in general\n### 1. `n_estimators`\nFrom my perspective, I think there are two ways to find an appropriate `n_estimators`:\n1. **Directly tuning `n_estimators`**: by fitting on the entire training set and *not* using early stopping.\n2. **Using early-stopping and letting the model choose its own `n_estimators`**: we'll set an arbitrarily high `n_estimators`, and keep 10-20% of training data for validation/early stopping. Probably also tune `early_stopping_rounds` if you can.\n\nDid you try either of them, or any other method? What do you find work better? I'm more inclined towards option (2) because in this competition we're using a large number of features and therefore there is a high chance of overfitting.\n\n### 2. Important hyper-parameters?\nIf you don't mind sharing, what hyper-parameters have been the most crucial in your tuning? For me they are `learning_rate`, `colsample_bytree` and `alpha`. I also tried tuning `learning_rate` and `lambda` but it didn't make much of a difference.",
    "2289330": "What were your conclusions about hyper-parameter tuning @hoangnguyen719 ? I used optuna, optimized n_estimators, learning_rate, colsample_bytree.... but when I do, I get a worse LB 😣",
    "2289521": "Welcome in \"Art and Science of Hyperparameter Optimizations\" 😁",
    "2290406": "haha thanks for welcoming me, I see you just scaled to the first place, congratulations 🎉\n@remekkinas would you recommend me some topic to learn about this? or some other tool I could look at? \nGreetings to Poland , I was there last year, I loved it...",
    "2290443": "Thank you for greetings :)\n\nI use this book: https://link.springer.com/book/10.1007/978-981-19-5170-1 \nYou can download it for free now.",
    "2290539": "Not sure if this is relevant to your experience, but I tried optimizing model hyperparameters for each questions. Although the f1 score was boosted to various degree, the overall macro f1, which is the metric of this competition, was not improved at all. It's a bit tricky than individual model tuning IMHO.",
    "2290554": "dziękuję bardzo! \nalready downloaded, thanks Remek",
    "2290556": "Do you have any idea why that happens? exactly that's why I asked, I thought I was going to boost the overall LB and the result was the opposite, final result decreased 🙁",
    "2294204": "[Here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/409497) has a pretty good discussion on the comparison between individual f1 score versus overall f1 score. The conclusion is a bit counter-intuitive, but its example illustrates why it is tricky to optimize model for each question.\n\nNote that for each question, the macro f1 score is the unweighted average of f1 for correct (1) and wrong (0) classes. The final f1 for all 18 questions, hence, is the same unweighted average of f1 scores of correct and wrong classes. I think the crucial part here is to boost the model performance for both classes in order to really see the improvement of the final macro f1."
  },
  "source": "meta"
}