{
  "id": 500014,
  "title": "Strategies for hyperparameter tuning in large dataset competitions",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/500014",
  "author_name": "",
  "post_date": "2024-05-03T22:50:08.541424Z",
  "votes": 10,
  "comment_count": 12,
  "views": 0,
  "content": "<p>In this competition, I'm facing challenges with hyperparameter tuning for LightGBM and CatBoost due to memory constraints using <strong>GridSearch</strong>. I am exploring alternative methods, such as tracking changes manually with ML platforms like <em>neptune.ai</em> or <em>MLflow</em>.</p>\n<p>Could anyone share their strategies or tips for effectively tuning hyperparameters when handling large datasets?</p>\n<ul>\n<li>Are there particular tools or techniques that are found to be effective for managing memory while tuning?</li>\n<li>How do you balance computational resources and tuning granularity?</li>\n</ul>\n<p>Any insights or experiences would be greatly appreciated!</p>",
  "messages": [
    {
      "id": "2791902",
      "postDate": "05/03/2024 22:50:08",
      "content": "<p>In this competition, I'm facing challenges with hyperparameter tuning for LightGBM and CatBoost due to memory constraints using <strong>GridSearch</strong>. I am exploring alternative methods, such as tracking changes manually with ML platforms like <em>neptune.ai</em> or <em>MLflow</em>.</p>\n<p>Could anyone share their strategies or tips for effectively tuning hyperparameters when handling large datasets?</p>\n<ul>\n<li>Are there particular tools or techniques that are found to be effective for managing memory while tuning?</li>\n<li>How do you balance computational resources and tuning granularity?</li>\n</ul>\n<p>Any insights or experiences would be greatly appreciated!</p>",
      "rawMarkdown": "In this competition, I'm facing challenges with hyperparameter tuning for LightGBM and CatBoost due to memory constraints using **GridSearch**. I am exploring alternative methods, such as tracking changes manually with ML platforms like *neptune.ai* or *MLflow*.\n\nCould anyone share their strategies or tips for effectively tuning hyperparameters when handling large datasets?\n\n- Are there particular tools or techniques that are found to be effective for managing memory while tuning?\n- How do you balance computational resources and tuning granularity?\n\nAny insights or experiences would be greatly appreciated!",
      "votes": null
    },
    {
      "id": "2792236",
      "postDate": "05/04/2024 06:13:45",
      "content": "<p>Optuna is a better tuning option than the conventional grid-search in my opinion. Just be careful in specifying the search space <a href=\"https://www.kaggle.com/faithk7u\" target=\"_blank\">@faithk7u</a> </p>",
      "rawMarkdown": "Optuna is a better tuning option than the conventional grid-search in my opinion. Just be careful in specifying the search space @faithk7u",
      "votes": null
    },
    {
      "id": "2793831",
      "postDate": "05/05/2024 02:10:24",
      "content": "<p>Optuna &amp; hyperopt but Optuna is most efficient way because its uses bayesian , still I am also facing same issue as you!!</p>",
      "rawMarkdown": "Optuna & hyperopt but Optuna is most efficient way because its uses bayesian , still I am also facing same issue as you!!",
      "votes": null
    },
    {
      "id": "2794446",
      "postDate": "05/05/2024 09:47:04",
      "content": "<p>It depends how do you test the parameters set during optimization, maybe you use 5 folds and the memory is not enough, try 3 folds, 1 fold…<br>\nUse a 70-80% subsample of the train set - not the entire one</p>",
      "rawMarkdown": "It depends how do you test the parameters set during optimization, maybe you use 5 folds and the memory is not enough, try 3 folds, 1 fold...\nUse a 70-80% subsample of the train set - not the entire one",
      "votes": null
    },
    {
      "id": "2796982",
      "postDate": "05/06/2024 13:48:32",
      "content": "<p>GridSearch is not a good strategy here IMO. The feature space is too large and there are too many samples to take into account. As many have mentioned already, Optuna with cross validation and Tree-structured Parzen Estimatos (TPE) should get you relatively quickly to where you need to be. </p>",
      "rawMarkdown": "GridSearch is not a good strategy here IMO. The feature space is too large and there are too many samples to take into account. As many have mentioned already, Optuna with cross validation and Tree-structured Parzen Estimatos (TPE) should get you relatively quickly to where you need to be.",
      "votes": null
    },
    {
      "id": "2799337",
      "postDate": "05/07/2024 17:49:05",
      "content": "<p>Thanks for the tip. Will try out Optuna. </p>",
      "rawMarkdown": "Thanks for the tip. Will try out Optuna.",
      "votes": null
    },
    {
      "id": "2800008",
      "postDate": "05/08/2024 04:55:12",
      "content": "<p>I tried it, but the final score seems to have deteriorated</p>",
      "rawMarkdown": "I tried it, but the final score seems to have deteriorated",
      "votes": null
    },
    {
      "id": "2802582",
      "postDate": "05/09/2024 05:25:58",
      "content": "<p>thanks a lot</p>",
      "rawMarkdown": "thanks a lot",
      "votes": null
    },
    {
      "id": "2804748",
      "postDate": "05/10/2024 07:26:38",
      "content": "<p>I think you can start by conducting a round of feature selection, followed by a small but wide range of grid search for parameter tuning. Then, based on this foundation, you can proceed with further parameter tuning, such as using packages like Optuna.😀</p>",
      "rawMarkdown": "I think you can start by conducting a round of feature selection, followed by a small but wide range of grid search for parameter tuning. Then, based on this foundation, you can proceed with further parameter tuning, such as using packages like Optuna.😀",
      "votes": null
    },
    {
      "id": "2806106",
      "postDate": "05/10/2024 21:46:54",
      "content": "<p>Thanks for the tips! <a href=\"https://www.kaggle.com/roger92\" target=\"_blank\">@roger92</a> How do you usually select features? I saw many selected features based on the feature importance by LightGBM or CatBoost, but wondering if there are other methods that one can adopt (variance/correlation) and thinking we should eliminate additional features that are highly correlated with the others. </p>\n<p>Will try out the randomized followed by the grid search method as suggested. Thanks!</p>",
      "rawMarkdown": "Thanks for the tips! @roger92 How do you usually select features? I saw many selected features based on the feature importance by LightGBM or CatBoost, but wondering if there are other methods that one can adopt (variance/correlation) and thinking we should eliminate additional features that are highly correlated with the others. \n\nWill try out the randomized followed by the grid search method as suggested. Thanks!",
      "votes": null
    },
    {
      "id": "2807530",
      "postDate": "05/11/2024 17:48:29",
      "content": "<p>Here is Optuna tuning with integrated stability metric <a href=\"https://www.kaggle.com/code/eu1234/optuna-with-integrated-stability-metric\" target=\"_blank\">https://www.kaggle.com/code/eu1234/optuna-with-integrated-stability-metric</a></p>",
      "rawMarkdown": "Here is Optuna tuning with integrated stability metric https://www.kaggle.com/code/eu1234/optuna-with-integrated-stability-metric",
      "votes": null
    },
    {
      "id": "2808080",
      "postDate": "05/12/2024 04:34:42",
      "content": "<p>I would like to know how the effect is after everyone uses Optuna? Do you have any feedback?</p>",
      "rawMarkdown": "I would like to know how the effect is after everyone uses Optuna? Do you have any feedback?",
      "votes": null
    },
    {
      "id": "2813176",
      "postDate": "05/14/2024 16:02:40",
      "content": "<p>Optuna is my go-to grid-search and I have found it more efficient that most others. I would suggest running it for fewer trials (about 20-25 trials) at a time and then continuing the search without defining a new objective if you need to tune further. Another approach I usually take is run optuna for 20 or 25 trials and that will give you a good idea of smaller search space. Then define a new objective for smaller search space to get the best results.</p>",
      "rawMarkdown": "Optuna is my go-to grid-search and I have found it more efficient that most others. I would suggest running it for fewer trials (about 20-25 trials) at a time and then continuing the search without defining a new objective if you need to tune further. Another approach I usually take is run optuna for 20 or 25 trials and that will give you a good idea of smaller search space. Then define a new objective for smaller search space to get the best results.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2792236,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "05/04/2024 06:13:45",
      "content": "<p>Optuna is a better tuning option than the conventional grid-search in my opinion. Just be careful in specifying the search space <a href=\"https://www.kaggle.com/faithk7u\" target=\"_blank\">@faithk7u</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 2799337,
          "author_name": "faithk7u",
          "author_url": "",
          "post_date": "05/07/2024 17:49:05",
          "content": "<p>Thanks for the tip. Will try out Optuna. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2800008,
              "author_name": "majiaqi111",
              "author_url": "",
              "post_date": "05/08/2024 04:55:12",
              "content": "<p>I tried it, but the final score seems to have deteriorated</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2793831,
      "author_name": "kunduruanil",
      "author_url": "",
      "post_date": "05/05/2024 02:10:24",
      "content": "<p>Optuna &amp; hyperopt but Optuna is most efficient way because its uses bayesian , still I am also facing same issue as you!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2794446,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "05/05/2024 09:47:04",
      "content": "<p>It depends how do you test the parameters set during optimization, maybe you use 5 folds and the memory is not enough, try 3 folds, 1 fold…<br>\nUse a 70-80% subsample of the train set - not the entire one</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2796982,
      "author_name": "eduardastefanescu",
      "author_url": "",
      "post_date": "05/06/2024 13:48:32",
      "content": "<p>GridSearch is not a good strategy here IMO. The feature space is too large and there are too many samples to take into account. As many have mentioned already, Optuna with cross validation and Tree-structured Parzen Estimatos (TPE) should get you relatively quickly to where you need to be. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2802582,
      "author_name": "smartgpu",
      "author_url": "",
      "post_date": "05/09/2024 05:25:58",
      "content": "<p>thanks a lot</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2804748,
      "author_name": "roger92",
      "author_url": "",
      "post_date": "05/10/2024 07:26:38",
      "content": "<p>I think you can start by conducting a round of feature selection, followed by a small but wide range of grid search for parameter tuning. Then, based on this foundation, you can proceed with further parameter tuning, such as using packages like Optuna.😀</p>",
      "votes": null,
      "replies": [
        {
          "id": 2806106,
          "author_name": "faithk7u",
          "author_url": "",
          "post_date": "05/10/2024 21:46:54",
          "content": "<p>Thanks for the tips! <a href=\"https://www.kaggle.com/roger92\" target=\"_blank\">@roger92</a> How do you usually select features? I saw many selected features based on the feature importance by LightGBM or CatBoost, but wondering if there are other methods that one can adopt (variance/correlation) and thinking we should eliminate additional features that are highly correlated with the others. </p>\n<p>Will try out the randomized followed by the grid search method as suggested. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2807530,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "05/11/2024 17:48:29",
      "content": "<p>Here is Optuna tuning with integrated stability metric <a href=\"https://www.kaggle.com/code/eu1234/optuna-with-integrated-stability-metric\" target=\"_blank\">https://www.kaggle.com/code/eu1234/optuna-with-integrated-stability-metric</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2808080,
      "author_name": "archerpan",
      "author_url": "",
      "post_date": "05/12/2024 04:34:42",
      "content": "<p>I would like to know how the effect is after everyone uses Optuna? Do you have any feedback?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2813176,
      "author_name": "varuniraothumsi",
      "author_url": "",
      "post_date": "05/14/2024 16:02:40",
      "content": "<p>Optuna is my go-to grid-search and I have found it more efficient that most others. I would suggest running it for fewer trials (about 20-25 trials) at a time and then continuing the search without defining a new objective if you need to tune further. Another approach I usually take is run optuna for 20 or 25 trials and that will give you a good idea of smaller search space. Then define a new objective for smaller search space to get the best results.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2791902": "In this competition, I'm facing challenges with hyperparameter tuning for LightGBM and CatBoost due to memory constraints using **GridSearch**. I am exploring alternative methods, such as tracking changes manually with ML platforms like *neptune.ai* or *MLflow*.\n\nCould anyone share their strategies or tips for effectively tuning hyperparameters when handling large datasets?\n\n- Are there particular tools or techniques that are found to be effective for managing memory while tuning?\n- How do you balance computational resources and tuning granularity?\n\nAny insights or experiences would be greatly appreciated!",
    "2792236": "Optuna is a better tuning option than the conventional grid-search in my opinion. Just be careful in specifying the search space @faithk7u",
    "2793831": "Optuna & hyperopt but Optuna is most efficient way because its uses bayesian , still I am also facing same issue as you!!",
    "2794446": "It depends how do you test the parameters set during optimization, maybe you use 5 folds and the memory is not enough, try 3 folds, 1 fold...\nUse a 70-80% subsample of the train set - not the entire one",
    "2796982": "GridSearch is not a good strategy here IMO. The feature space is too large and there are too many samples to take into account. As many have mentioned already, Optuna with cross validation and Tree-structured Parzen Estimatos (TPE) should get you relatively quickly to where you need to be.",
    "2799337": "Thanks for the tip. Will try out Optuna.",
    "2800008": "I tried it, but the final score seems to have deteriorated",
    "2802582": "thanks a lot",
    "2804748": "I think you can start by conducting a round of feature selection, followed by a small but wide range of grid search for parameter tuning. Then, based on this foundation, you can proceed with further parameter tuning, such as using packages like Optuna.😀",
    "2806106": "Thanks for the tips! @roger92 How do you usually select features? I saw many selected features based on the feature importance by LightGBM or CatBoost, but wondering if there are other methods that one can adopt (variance/correlation) and thinking we should eliminate additional features that are highly correlated with the others. \n\nWill try out the randomized followed by the grid search method as suggested. Thanks!",
    "2807530": "Here is Optuna tuning with integrated stability metric https://www.kaggle.com/code/eu1234/optuna-with-integrated-stability-metric",
    "2808080": "I would like to know how the effect is after everyone uses Optuna? Do you have any feedback?",
    "2813176": "Optuna is my go-to grid-search and I have found it more efficient that most others. I would suggest running it for fewer trials (about 20-25 trials) at a time and then continuing the search without defining a new objective if you need to tune further. Another approach I usually take is run optuna for 20 or 25 trials and that will give you a good idea of smaller search space. Then define a new objective for smaller search space to get the best results."
  },
  "source": "meta"
}