{
  "id": 550074,
  "title": "The Optimisers Curse",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/550074",
  "author_name": "",
  "post_date": "2024-12-05T09:44:29.808207300Z",
  "votes": 20,
  "comment_count": 9,
  "views": 0,
  "content": "<p>During this competition one of the things that I focused on was hyperparameter optimisation. While analysing the hyperparameters of the trials I came across an interesting statistical phenomenon called <code>the optimiser’s curse</code>. Basically what it says is that when there is noise in the data the optimiser can overestimate the value or the importance of certain hyperparameters by sheer luck. One way to mitigate this is not to optimise the hyperparameters that are not correlated with the validation score because adding them introduces extra randomness. </p>\n<p>Plot below could be used to study the relationship between the individual hyperparameters and the validation score.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F225499%2Fbdee903e1dc7294046400744aed2852a%2Fslice_plot.png?generation=1733391524963030&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>The <code>learning_rate</code> shows a clear relationship with the objective value.</li>\n<li>For <code>lambda_1</code> and <code>lambda_2</code> one can observe some cluster forming at the top right region (although a weak one). And given that there is quit a bit of noise in the data I would choose not to leave these out.</li>\n<li>The rest of the hyperparameters do not really correlate with the objective function.</li>\n</ul>\n<p>Based on these observations one can argue to just optimise <code>leaning_rate</code>, <code>lambda_1</code>, <code>lambda_2</code>.  </p>\n<p>Let me know what you think.</p>\n<p>This is a nice <a href=\"https://www.youtube.com/watch?v=vC9sAD-ymhk&amp;t=181s\" target=\"_blank\">youtube</a> link where this phenomenon is explained.</p>",
  "messages": [
    {
      "id": "3064159",
      "postDate": "12/05/2024 09:44:29",
      "content": "<p>During this competition one of the things that I focused on was hyperparameter optimisation. While analysing the hyperparameters of the trials I came across an interesting statistical phenomenon called <code>the optimiser’s curse</code>. Basically what it says is that when there is noise in the data the optimiser can overestimate the value or the importance of certain hyperparameters by sheer luck. One way to mitigate this is not to optimise the hyperparameters that are not correlated with the validation score because adding them introduces extra randomness. </p>\n<p>Plot below could be used to study the relationship between the individual hyperparameters and the validation score.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F225499%2Fbdee903e1dc7294046400744aed2852a%2Fslice_plot.png?generation=1733391524963030&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>The <code>learning_rate</code> shows a clear relationship with the objective value.</li>\n<li>For <code>lambda_1</code> and <code>lambda_2</code> one can observe some cluster forming at the top right region (although a weak one). And given that there is quit a bit of noise in the data I would choose not to leave these out.</li>\n<li>The rest of the hyperparameters do not really correlate with the objective function.</li>\n</ul>\n<p>Based on these observations one can argue to just optimise <code>leaning_rate</code>, <code>lambda_1</code>, <code>lambda_2</code>.  </p>\n<p>Let me know what you think.</p>\n<p>This is a nice <a href=\"https://www.youtube.com/watch?v=vC9sAD-ymhk&amp;t=181s\" target=\"_blank\">youtube</a> link where this phenomenon is explained.</p>",
      "rawMarkdown": "During this competition one of the things that I focused on was hyperparameter optimisation. While analysing the hyperparameters of the trials I came across an interesting statistical phenomenon called `the optimiser’s curse`. Basically what it says is that when there is noise in the data the optimiser can overestimate the value or the importance of certain hyperparameters by sheer luck. One way to mitigate this is not to optimise the hyperparameters that are not correlated with the validation score because adding them introduces extra randomness. \n\nPlot below could be used to study the relationship between the individual hyperparameters and the validation score.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F225499%2Fbdee903e1dc7294046400744aed2852a%2Fslice_plot.png?generation=1733391524963030&alt=media)\n\n-  The `learning_rate` shows a clear relationship with the objective value.\n- For `lambda_1` and `lambda_2` one can observe some cluster forming at the top right region (although a weak one). And given that there is quit a bit of noise in the data I would choose not to leave these out.\n- The rest of the hyperparameters do not really correlate with the objective function.\n\nBased on these observations one can argue to just optimise `leaning_rate`, `lambda_1`, `lambda_2`.  \n\nLet me know what you think.\n\nThis is a nice [youtube](https://www.youtube.com/watch?v=vC9sAD-ymhk&t=181s) link where this phenomenon is explained.",
      "votes": null
    },
    {
      "id": "3065411",
      "postDate": "12/06/2024 19:04:16",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/wti200\" target=\"_blank\">@wti200</a>,</p>\n<p>Thank you for the link. Nice explanation.</p>\n<p>With your plots, I see a relation between objective function and feature fraction and bagging fraction too. </p>\n<p>And with this small dataset, it's interesting to refit the 5 or 10 best trials with severel seeds to unsure the best trial is not a lucky one . </p>\n<p>In the following exmaple, let's imagine we have detected the best trial by using optuna with random_seed 42 (in my kfolds for example). Then I have sorted trials by score :</p>\n<table>\n<thead>\n<tr>\n<th>seed</th>\n<th>42</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>trial n°1</td>\n<td>.49</td>\n</tr>\n<tr>\n<td>trial n°2</td>\n<td>.48</td>\n</tr>\n<tr>\n<td>trial n°3</td>\n<td>.475</td>\n</tr>\n<tr>\n<td>trial n°4</td>\n<td>.47</td>\n</tr>\n<tr>\n<td>trial n°5</td>\n<td>.46</td>\n</tr>\n</tbody>\n</table>\n<p>Next, I do some other CV with those 5 trials, but with 2 other seeds (65 and 88) : </p>\n<table>\n<thead>\n<tr>\n<th>seed</th>\n<th>42</th>\n<th>65</th>\n<th>88</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>trial n°1</td>\n<td>.49</td>\n<td>.45</td>\n<td>.40</td>\n</tr>\n<tr>\n<td>trial n°2</td>\n<td>.48</td>\n<td>.475</td>\n<td>.49</td>\n</tr>\n<tr>\n<td>trial n°3</td>\n<td>.475</td>\n<td>.46</td>\n<td>.5</td>\n</tr>\n<tr>\n<td>trial n°4</td>\n<td>.47</td>\n<td>.4</td>\n<td>.43</td>\n</tr>\n<tr>\n<td>trial n°5</td>\n<td>.46</td>\n<td>.4</td>\n<td>.5</td>\n</tr>\n</tbody>\n</table>\n<p>We can see that trial n°2 was not the best with random state 42, but is better than trial n°1 for those 2 other random states. <br>\nAnd CV score with trial n°2 has a lower variance than with trial n°1 : in this case, I will choose hyperparams of trial n°2 !</p>",
      "rawMarkdown": "Hi @wti200,\n\nThank you for the link. Nice explanation.\n\nWith your plots, I see a relation between objective function and feature fraction and bagging fraction too. \n\nAnd with this small dataset, it's interesting to refit the 5 or 10 best trials with severel seeds to unsure the best trial is not a lucky one . \n\n\nIn the following exmaple, let's imagine we have detected the best trial by using optuna with random_seed 42 (in my kfolds for example). Then I have sorted trials by score :\n| seed | 42 | \n| --- | --- |\n| trial n°1 | .49 |   \n| trial n°2 | .48 |  \n| trial n°3 | .475 | \n| trial n°4 | .47 |  \n| trial n°5 | .46 |  \n\nNext, I do some other CV with those 5 trials, but with 2 other seeds (65 and 88) : \n\n| seed | 42 | 65 | 88 |\n| --- | --- | --- | --- |\n| trial n°1 | .49 | .45 | .40 |  \n| trial n°2 | .48 | .475 | .49 |  \n| trial n°3 | .475 | .46 | .5  | \n| trial n°4 | .47 | .4 | .43  | \n| trial n°5 | .46 | .4 | .5  | \n\nWe can see that trial n°2 was not the best with random state 42, but is better than trial n°1 for those 2 other random states. \nAnd CV score with trial n°2 has a lower variance than with trial n°1 : in this case, I will choose hyperparams of trial n°2 !",
      "votes": null
    },
    {
      "id": "3065479",
      "postDate": "12/06/2024 21:31:40",
      "content": "<p><a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">@adaubas</a> interesting, thanks for sharing. Optimal hyperparameters aside, do the overall scatterplots between each hyperparam vs. objective function follow the same pattern as the ones produced by <a href=\"https://www.kaggle.com/wti200\" target=\"_blank\">@wti200</a>, even after experimenting with different seeds?</p>",
      "rawMarkdown": "adaubas interesting, thanks for sharing. Optimal hyperparameters aside, do the overall scatterplots between each hyperparam vs. objective function follow the same pattern as the ones produced by @wti200, even after experimenting with different seeds?",
      "votes": null
    },
    {
      "id": "3065499",
      "postDate": "12/06/2024 22:29:14",
      "content": "<p>I tested different seeds <em>after</em> generating the resulting scatterplots from the optuna trials.<br>\nAnd no, my scatterplots don't have same patterns as bellow : for example, I set <strong>n_estimators</strong> using early stopping, not with optuna.</p>",
      "rawMarkdown": "I tested different seeds *after* generating the resulting scatterplots from the optuna trials.\nAnd no, my scatterplots don't have same patterns as bellow : for example, I set **n_estimators** using early stopping, not with optuna.",
      "votes": null
    },
    {
      "id": "3065702",
      "postDate": "12/07/2024 06:58:03",
      "content": "<p>off topic, this is a good point I have just been thinking about with this small dataset, whether if one should set <strong>n_estimators</strong> with early stopping and get virtually better cv results or set it through optuna and be sure to not be overfitting the cv…</p>",
      "rawMarkdown": "off topic, this is a good point I have just been thinking about with this small dataset, whether if one should set **n_estimators** with early stopping and get virtually better cv results or set it through optuna and be sure to not be overfitting the cv…",
      "votes": null
    },
    {
      "id": "3065742",
      "postDate": "12/07/2024 08:28:23",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dib\" target=\"_blank\">@dib</a>,<br>\nTake a look to a recent discussion about \"early stopping or not\" in another competition a few days ago :  <a href=\"https://www.kaggle.com/competitions/playground-series-s4e12/discussion/549788\" target=\"_blank\">https://www.kaggle.com/competitions/playground-series-s4e12/discussion/549788</a></p>\n<p>I believe there is no better way than early stopping to estimate the best n_estimators… Even in our current competition ; that's my opinion.</p>",
      "rawMarkdown": "Hi @dib,\nTake a look to a recent discussion about \"early stopping or not\" in another competition a few days ago :  https://www.kaggle.com/competitions/playground-series-s4e12/discussion/549788\n\nI believe there is no better way than early stopping to estimate the best n_estimators... Even in our current competition ; that's my opinion.",
      "votes": null
    },
    {
      "id": "3066043",
      "postDate": "12/07/2024 15:29:52",
      "content": "<p>very good discussion post, thanks for sharing <a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">@adaubas</a>. I agree that cv+es might be the best option most of the times, it is just that this time with these small dataset and metric based on hard labels makes me rethink it.  Let's see if the future winners of the competition disclose how they set n_estimators</p>",
      "rawMarkdown": "very good discussion post, thanks for sharing @adaubas. I agree that cv+es might be the best option most of the times, it is just that this time with these small dataset and metric based on hard labels makes me rethink it.  Let's see if the future winners of the competition disclose how they set n_estimators",
      "votes": null
    },
    {
      "id": "3067053",
      "postDate": "12/08/2024 19:56:31",
      "content": "<p>Thanks for pointing out that youtube video! The noise in hyperparameter fitting can 9ndeed make it an unsatisfying process…</p>\n<p>fyi, I've been fitting with XGB models using a very limited set of hyperparameters: colsample_bytree, learning_rate, max_depth, n_estimators, and subsample. Doing a 2d grid search of learning_rate and n_estimators, and keeping the others fixed. To see and reduce variation in this process, the same grids were done on 5 shuffled XGB models and combined. There's a clear band of the highest test scores along the higher-learning-rates-with-fewer-estimators diagonal.</p>",
      "rawMarkdown": "Thanks for pointing out that youtube video! The noise in hyperparameter fitting can 9ndeed make it an unsatisfying process...\n\nfyi, I've been fitting with XGB models using a very limited set of hyperparameters: colsample_bytree, learning_rate, max_depth, n_estimators, and subsample. Doing a 2d grid search of learning_rate and n_estimators, and keeping the others fixed. To see and reduce variation in this process, the same grids were done on 5 shuffled XGB models and combined. There's a clear band of the highest test scores along the higher-learning-rates-with-fewer-estimators diagonal.",
      "votes": null
    },
    {
      "id": "3068876",
      "postDate": "12/10/2024 18:48:59",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">@adaubas</a>, sorry for the late response. </p>\n<p>My hunch is that the random seed is just another hyperparameter that does not correlate with the objective function. And optimising it comes with the risk of overfitting. I do believe that generating predictions with multiple seeds can introduce additional variance in the individual models, which may improve the overall ensemble performance. </p>",
      "rawMarkdown": "Hi @adaubas, sorry for the late response. \n\nMy hunch is that the random seed is just another hyperparameter that does not correlate with the objective function. And optimising it comes with the risk of overfitting. I do believe that generating predictions with multiple seeds can introduce additional variance in the individual models, which may improve the overall ensemble performance.",
      "votes": null
    },
    {
      "id": "3068879",
      "postDate": "12/10/2024 18:49:52",
      "content": "<p><a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">@adaubas</a> thanks for the link!</p>",
      "rawMarkdown": "adaubas thanks for the link!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3065411,
      "author_name": "adaubas",
      "author_url": "",
      "post_date": "12/06/2024 19:04:16",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/wti200\" target=\"_blank\">@wti200</a>,</p>\n<p>Thank you for the link. Nice explanation.</p>\n<p>With your plots, I see a relation between objective function and feature fraction and bagging fraction too. </p>\n<p>And with this small dataset, it's interesting to refit the 5 or 10 best trials with severel seeds to unsure the best trial is not a lucky one . </p>\n<p>In the following exmaple, let's imagine we have detected the best trial by using optuna with random_seed 42 (in my kfolds for example). Then I have sorted trials by score :</p>\n<table>\n<thead>\n<tr>\n<th>seed</th>\n<th>42</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>trial n°1</td>\n<td>.49</td>\n</tr>\n<tr>\n<td>trial n°2</td>\n<td>.48</td>\n</tr>\n<tr>\n<td>trial n°3</td>\n<td>.475</td>\n</tr>\n<tr>\n<td>trial n°4</td>\n<td>.47</td>\n</tr>\n<tr>\n<td>trial n°5</td>\n<td>.46</td>\n</tr>\n</tbody>\n</table>\n<p>Next, I do some other CV with those 5 trials, but with 2 other seeds (65 and 88) : </p>\n<table>\n<thead>\n<tr>\n<th>seed</th>\n<th>42</th>\n<th>65</th>\n<th>88</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>trial n°1</td>\n<td>.49</td>\n<td>.45</td>\n<td>.40</td>\n</tr>\n<tr>\n<td>trial n°2</td>\n<td>.48</td>\n<td>.475</td>\n<td>.49</td>\n</tr>\n<tr>\n<td>trial n°3</td>\n<td>.475</td>\n<td>.46</td>\n<td>.5</td>\n</tr>\n<tr>\n<td>trial n°4</td>\n<td>.47</td>\n<td>.4</td>\n<td>.43</td>\n</tr>\n<tr>\n<td>trial n°5</td>\n<td>.46</td>\n<td>.4</td>\n<td>.5</td>\n</tr>\n</tbody>\n</table>\n<p>We can see that trial n°2 was not the best with random state 42, but is better than trial n°1 for those 2 other random states. <br>\nAnd CV score with trial n°2 has a lower variance than with trial n°1 : in this case, I will choose hyperparams of trial n°2 !</p>",
      "votes": null,
      "replies": [
        {
          "id": 3065479,
          "author_name": "tztang",
          "author_url": "",
          "post_date": "12/06/2024 21:31:40",
          "content": "<p><a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">@adaubas</a> interesting, thanks for sharing. Optimal hyperparameters aside, do the overall scatterplots between each hyperparam vs. objective function follow the same pattern as the ones produced by <a href=\"https://www.kaggle.com/wti200\" target=\"_blank\">@wti200</a>, even after experimenting with different seeds?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3065499,
              "author_name": "adaubas",
              "author_url": "",
              "post_date": "12/06/2024 22:29:14",
              "content": "<p>I tested different seeds <em>after</em> generating the resulting scatterplots from the optuna trials.<br>\nAnd no, my scatterplots don't have same patterns as bellow : for example, I set <strong>n_estimators</strong> using early stopping, not with optuna.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3065702,
                  "author_name": "diegoiglesias",
                  "author_url": "",
                  "post_date": "12/07/2024 06:58:03",
                  "content": "<p>off topic, this is a good point I have just been thinking about with this small dataset, whether if one should set <strong>n_estimators</strong> with early stopping and get virtually better cv results or set it through optuna and be sure to not be overfitting the cv…</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3065742,
                      "author_name": "adaubas",
                      "author_url": "",
                      "post_date": "12/07/2024 08:28:23",
                      "content": "<p>Hi <a href=\"https://www.kaggle.com/dib\" target=\"_blank\">@dib</a>,<br>\nTake a look to a recent discussion about \"early stopping or not\" in another competition a few days ago :  <a href=\"https://www.kaggle.com/competitions/playground-series-s4e12/discussion/549788\" target=\"_blank\">https://www.kaggle.com/competitions/playground-series-s4e12/discussion/549788</a></p>\n<p>I believe there is no better way than early stopping to estimate the best n_estimators… Even in our current competition ; that's my opinion.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3066043,
                          "author_name": "diegoiglesias",
                          "author_url": "",
                          "post_date": "12/07/2024 15:29:52",
                          "content": "<p>very good discussion post, thanks for sharing <a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">@adaubas</a>. I agree that cv+es might be the best option most of the times, it is just that this time with these small dataset and metric based on hard labels makes me rethink it.  Let's see if the future winners of the competition disclose how they set n_estimators</p>",
                          "votes": null,
                          "replies": []
                        },
                        {
                          "id": 3068879,
                          "author_name": "wti200",
                          "author_url": "",
                          "post_date": "12/10/2024 18:49:52",
                          "content": "<p><a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">@adaubas</a> thanks for the link!</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 3068876,
          "author_name": "wti200",
          "author_url": "",
          "post_date": "12/10/2024 18:48:59",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">@adaubas</a>, sorry for the late response. </p>\n<p>My hunch is that the random seed is just another hyperparameter that does not correlate with the objective function. And optimising it comes with the risk of overfitting. I do believe that generating predictions with multiple seeds can introduce additional variance in the individual models, which may improve the overall ensemble performance. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3067053,
      "author_name": "dan3dewey",
      "author_url": "",
      "post_date": "12/08/2024 19:56:31",
      "content": "<p>Thanks for pointing out that youtube video! The noise in hyperparameter fitting can 9ndeed make it an unsatisfying process…</p>\n<p>fyi, I've been fitting with XGB models using a very limited set of hyperparameters: colsample_bytree, learning_rate, max_depth, n_estimators, and subsample. Doing a 2d grid search of learning_rate and n_estimators, and keeping the others fixed. To see and reduce variation in this process, the same grids were done on 5 shuffled XGB models and combined. There's a clear band of the highest test scores along the higher-learning-rates-with-fewer-estimators diagonal.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3064159": "During this competition one of the things that I focused on was hyperparameter optimisation. While analysing the hyperparameters of the trials I came across an interesting statistical phenomenon called `the optimiser’s curse`. Basically what it says is that when there is noise in the data the optimiser can overestimate the value or the importance of certain hyperparameters by sheer luck. One way to mitigate this is not to optimise the hyperparameters that are not correlated with the validation score because adding them introduces extra randomness. \n\nPlot below could be used to study the relationship between the individual hyperparameters and the validation score.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F225499%2Fbdee903e1dc7294046400744aed2852a%2Fslice_plot.png?generation=1733391524963030&alt=media)\n\n-  The `learning_rate` shows a clear relationship with the objective value.\n- For `lambda_1` and `lambda_2` one can observe some cluster forming at the top right region (although a weak one). And given that there is quit a bit of noise in the data I would choose not to leave these out.\n- The rest of the hyperparameters do not really correlate with the objective function.\n\nBased on these observations one can argue to just optimise `leaning_rate`, `lambda_1`, `lambda_2`.  \n\nLet me know what you think.\n\nThis is a nice [youtube](https://www.youtube.com/watch?v=vC9sAD-ymhk&t=181s) link where this phenomenon is explained.",
    "3065411": "Hi @wti200,\n\nThank you for the link. Nice explanation.\n\nWith your plots, I see a relation between objective function and feature fraction and bagging fraction too. \n\nAnd with this small dataset, it's interesting to refit the 5 or 10 best trials with severel seeds to unsure the best trial is not a lucky one . \n\n\nIn the following exmaple, let's imagine we have detected the best trial by using optuna with random_seed 42 (in my kfolds for example). Then I have sorted trials by score :\n| seed | 42 | \n| --- | --- |\n| trial n°1 | .49 |   \n| trial n°2 | .48 |  \n| trial n°3 | .475 | \n| trial n°4 | .47 |  \n| trial n°5 | .46 |  \n\nNext, I do some other CV with those 5 trials, but with 2 other seeds (65 and 88) : \n\n| seed | 42 | 65 | 88 |\n| --- | --- | --- | --- |\n| trial n°1 | .49 | .45 | .40 |  \n| trial n°2 | .48 | .475 | .49 |  \n| trial n°3 | .475 | .46 | .5  | \n| trial n°4 | .47 | .4 | .43  | \n| trial n°5 | .46 | .4 | .5  | \n\nWe can see that trial n°2 was not the best with random state 42, but is better than trial n°1 for those 2 other random states. \nAnd CV score with trial n°2 has a lower variance than with trial n°1 : in this case, I will choose hyperparams of trial n°2 !",
    "3065479": "adaubas interesting, thanks for sharing. Optimal hyperparameters aside, do the overall scatterplots between each hyperparam vs. objective function follow the same pattern as the ones produced by @wti200, even after experimenting with different seeds?",
    "3065499": "I tested different seeds *after* generating the resulting scatterplots from the optuna trials.\nAnd no, my scatterplots don't have same patterns as bellow : for example, I set **n_estimators** using early stopping, not with optuna.",
    "3065702": "off topic, this is a good point I have just been thinking about with this small dataset, whether if one should set **n_estimators** with early stopping and get virtually better cv results or set it through optuna and be sure to not be overfitting the cv…",
    "3065742": "Hi @dib,\nTake a look to a recent discussion about \"early stopping or not\" in another competition a few days ago :  https://www.kaggle.com/competitions/playground-series-s4e12/discussion/549788\n\nI believe there is no better way than early stopping to estimate the best n_estimators... Even in our current competition ; that's my opinion.",
    "3066043": "very good discussion post, thanks for sharing @adaubas. I agree that cv+es might be the best option most of the times, it is just that this time with these small dataset and metric based on hard labels makes me rethink it.  Let's see if the future winners of the competition disclose how they set n_estimators",
    "3067053": "Thanks for pointing out that youtube video! The noise in hyperparameter fitting can 9ndeed make it an unsatisfying process...\n\nfyi, I've been fitting with XGB models using a very limited set of hyperparameters: colsample_bytree, learning_rate, max_depth, n_estimators, and subsample. Doing a 2d grid search of learning_rate and n_estimators, and keeping the others fixed. To see and reduce variation in this process, the same grids were done on 5 shuffled XGB models and combined. There's a clear band of the highest test scores along the higher-learning-rates-with-fewer-estimators diagonal.",
    "3068876": "Hi @adaubas, sorry for the late response. \n\nMy hunch is that the random seed is just another hyperparameter that does not correlate with the objective function. And optimising it comes with the risk of overfitting. I do believe that generating predictions with multiple seeds can introduce additional variance in the individual models, which may improve the overall ensemble performance.",
    "3068879": "adaubas thanks for the link!"
  },
  "source": "meta"
}