{
  "id": 206100,
  "title": "LGBM Parameter Tuning",
  "url": "/competitions/riiid-test-answer-prediction/discussion/206100",
  "author_name": "",
  "post_date": "2020-12-23T08:16:16.833556Z",
  "votes": 15,
  "comment_count": 22,
  "views": 0,
  "content": "<p>I have a baseline LGBM model which is giving me around 0.782 on CV and 0.776 on the Public LB. I am using Loop feature engineering and a total of 18 features. I saw a few kernels having 44-47 features which is giving 0.775 on LB but their CV is slightly overfitting.</p>\n<p>I just wanted to know that is the parameter tuning of LGBM model playing a crucial part? Can the tuning increase the score by 0.5-1%? </p>\n<p>Can any of the top guys using LGBM please confirm?</p>",
  "messages": [
    {
      "id": "1123438",
      "postDate": "12/23/2020 08:16:16",
      "content": "<p>I have a baseline LGBM model which is giving me around 0.782 on CV and 0.776 on the Public LB. I am using Loop feature engineering and a total of 18 features. I saw a few kernels having 44-47 features which is giving 0.775 on LB but their CV is slightly overfitting.</p>\n<p>I just wanted to know that is the parameter tuning of LGBM model playing a crucial part? Can the tuning increase the score by 0.5-1%? </p>\n<p>Can any of the top guys using LGBM please confirm?</p>",
      "rawMarkdown": "I have a baseline LGBM model which is giving me around 0.782 on CV and 0.776 on the Public LB. I am using Loop feature engineering and a total of 18 features. I saw a few kernels having 44-47 features which is giving 0.775 on LB but their CV is slightly overfitting.\n\nI just wanted to know that is the parameter tuning of LGBM model playing a crucial part? Can the tuning increase the score by 0.5-1%? \n\nCan any of the top guys using LGBM please confirm?",
      "votes": null
    },
    {
      "id": "1123626",
      "postDate": "12/23/2020 11:34:35",
      "content": "<p>For my case learning rate search increased local validation from 0.773 -&gt; 0.785 , my local validation seems stable so I think LB score won't be far from it.</p>",
      "rawMarkdown": "For my case learning rate search increased local validation from 0.773 -> 0.785 , my local validation seems stable so I think LB score won't be far from it.",
      "votes": null
    },
    {
      "id": "1123707",
      "postDate": "12/23/2020 12:49:07",
      "content": "<p>Wow. That's impressive. Could you share some more details about how you searched the learning rate? </p>",
      "rawMarkdown": "Wow. That's impressive. Could you share some more details about how you searched the learning rate?",
      "votes": null
    },
    {
      "id": "1123719",
      "postDate": "12/23/2020 13:06:39",
      "content": "<p>It depends on your features, but in general,  lightgbm parameter do not play a critical role in any competition. I believe this competition will be no exception.</p>\n<p>I have a collection of lightgbm parameters from previous top solutions in my repository. I hope this will be of some help.</p>\n<p><a href=\"https://github.com/nyanp/nyaggle/blob/master/nyaggle/hyper_parameters/lightgbm.py\" target=\"_blank\">https://github.com/nyanp/nyaggle/blob/master/nyaggle/hyper_parameters/lightgbm.py</a></p>\n<ul>\n<li>In all solutions, at least one of <code>max_depth</code> or <code>num_leaves</code> has been adjusted. This is one of the most important parameter of lightgbm.</li>\n<li>Many solutions make <code>learning_rate</code> smaller than the default (0.1). This will increase the accuracy a bit in most cases (but will increase the learning and inference time instead).</li>\n</ul>",
      "rawMarkdown": "It depends on your features, but in general,  lightgbm parameter do not play a critical role in any competition. I believe this competition will be no exception.\n\nI have a collection of lightgbm parameters from previous top solutions in my repository. I hope this will be of some help.\n\nhttps://github.com/nyanp/nyaggle/blob/master/nyaggle/hyper_parameters/lightgbm.py\n\n- In all solutions, at least one of `max_depth` or `num_leaves` has been adjusted. This is one of the most important parameter of lightgbm.\n- Many solutions make `learning_rate` smaller than the default (0.1). This will increase the accuracy a bit in most cases (but will increase the learning and inference time instead).",
      "votes": null
    },
    {
      "id": "1124408",
      "postDate": "12/23/2020 23:03:06",
      "content": "<p>This super helpful. Thanks for sharing the list of winning parameters.</p>",
      "rawMarkdown": "This super helpful. Thanks for sharing the list of winning parameters.",
      "votes": null
    },
    {
      "id": "1124412",
      "postDate": "12/23/2020 23:06:18",
      "content": "<p><a href=\"https://www.kaggle.com/higepon\" target=\"_blank\">@higepon</a> Pretty straightforward, a grid search with twenty possible lr values</p>",
      "rawMarkdown": "higepon Pretty straightforward, a grid search with twenty possible lr values",
      "votes": null
    },
    {
      "id": "1124420",
      "postDate": "12/23/2020 23:22:52",
      "content": "<p>Thank you!<br>\nOkay it sounds time consuming if you did it with all the training data.  Any tips to make it faster?</p>",
      "rawMarkdown": "Thank you!\nOkay it sounds time consuming if you did it with all the training data.  Any tips to make it faster?",
      "votes": null
    },
    {
      "id": "1124488",
      "postDate": "12/24/2020 01:42:38",
      "content": "<p>Please don't tune learning rate. Great documentation of Laurae, which is included in the official documentation of lightgbm (<a href=\"https://sites.google.com/view/lauraepp/parameters\" target=\"_blank\">https://sites.google.com/view/lauraepp/parameters</a>) says<br>\n<code>\nOnce your learning rate is fixed, do not change it.\nIt is not a good practice to consider the learning rate as a hyperparameter to tune.\nLearning rate should be tuned according to your training speed and performance tradeoff.\nDo not let an optimizer tune it. One must not expect to see an overfitting learning rate of 0.0202048.\n</code><br>\nPlease also remember <br>\n\"The smaller the learning rate is, the better the performance is.<br>\nThe smaller the learning rate is, the more time-consuming training is.\"</p>",
      "rawMarkdown": "Please don't tune learning rate. Great documentation of Laurae, which is included in the official documentation of lightgbm (https://sites.google.com/view/lauraepp/parameters) says\n``\nOnce your learning rate is fixed, do not change it.\nIt is not a good practice to consider the learning rate as a hyperparameter to tune.\nLearning rate should be tuned according to your training speed and performance tradeoff.\nDo not let an optimizer tune it. One must not expect to see an overfitting learning rate of 0.0202048.\n``\nPlease also remember \n\"The smaller the learning rate is, the better the performance is.\nThe smaller the learning rate is, the more time-consuming training is.\"",
      "votes": null
    },
    {
      "id": "1124708",
      "postDate": "12/24/2020 06:43:43",
      "content": "<p>Please let me piggy back on this question because it's related.<br>\nI have LGBM model trained with full train data. It task N hours to train. <br>\nIf you do parameter search with grid search or optuna type of search, it may require M trials of train. IIUC it would take N * M hours to get the best param which may be expensive for some cases.</p>\n<p>Is there any way to mitigate that? eg) Maybe train with X % of data is sufficient for the search?</p>\n<p>Thank you in advance,<br>\nhigepon</p>",
      "rawMarkdown": "Please let me piggy back on this question because it's related.\nI have LGBM model trained with full train data. It task N hours to train. \nIf you do parameter search with grid search or optuna type of search, it may require M trials of train. IIUC it would take N * M hours to get the best param which may be expensive for some cases.\n\nIs there any way to mitigate that? eg) Maybe train with X % of data is sufficient for the search?\n\nThank you in advance,\nhigepon",
      "votes": null
    },
    {
      "id": "1124788",
      "postDate": "12/24/2020 07:51:31",
      "content": "<p>Yes, it's possible to only use X % of data for training, but the optimal parameter may be somewhat different from the one when the model is trained using 100% of data. <br>\nIf N is not too big, tuning only <code>num_leaves</code> using for loop is also a possible option. as nyanp says, other parameters are generally not so important. </p>",
      "rawMarkdown": "Yes, it's possible to only use X % of data for training, but the optimal parameter may be somewhat different from the one when the model is trained using 100% of data. \nIf N is not too big, tuning only `num_leaves` using for loop is also a possible option. as nyanp says, other parameters are generally not so important.",
      "votes": null
    },
    {
      "id": "1124837",
      "postDate": "12/24/2020 08:19:12",
      "content": "<p>Thank you mamas for your reply.<br>\nThat makes sense. I'll tune only num_leaves for now.</p>\n<p>Much appreciated.</p>",
      "rawMarkdown": "Thank you mamas for your reply.\nThat makes sense. I'll tune only num_leaves for now.\n\nMuch appreciated.",
      "votes": null
    },
    {
      "id": "1124884",
      "postDate": "12/24/2020 08:55:48",
      "content": "<p>Thank you for sharing your great insight, <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> . That's pretty useful. </p>\n<p>Now I am wondering, if a guy achieved a good result with less features than others, is that mean the features are outstanding? <br>\nIn <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/204801\" target=\"_blank\">this discussion</a>, one kaggler says <code>0.786 Single LGB model with about 14-16 features.</code>. On the other hand, another kaggler says <code>Single LGB model with about 50 features, CV:0.784 and LB:0.787</code>. </p>\n<p>I didn't know how it happens.<br>\nSo it might be happening because of the nature of their own features? Not because of hyper parameter, am I right? <br>\nOr, sometimes even having less features performs well with LGBM? </p>",
      "rawMarkdown": "Thank you for sharing your great insight, @nyanpn . That's pretty useful. \n\nNow I am wondering, if a guy achieved a good result with less features than others, is that mean the features are outstanding? \nIn [this discussion](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/204801), one kaggler says `0.786 Single LGB model with about 14-16 features.`. On the other hand, another kaggler says `Single LGB model with about 50 features, CV:0.784 and LB:0.787`. \n\nI didn't know how it happens.\nSo it might be happening because of the nature of their own features? Not because of hyper parameter, am I right? \nOr, sometimes even having less features performs well with LGBM?",
      "votes": null
    },
    {
      "id": "1125159",
      "postDate": "12/24/2020 12:40:50",
      "content": "<p>This may be due to the difference of the features, or simply due to different CV methods. Rather than comparing the number of features with others, I recommend that you focus on your own features.</p>",
      "rawMarkdown": "This may be due to the difference of the features, or simply due to different CV methods. Rather than comparing the number of features with others, I recommend that you focus on your own features.",
      "votes": null
    },
    {
      "id": "1125733",
      "postDate": "12/25/2020 02:48:42",
      "content": "<p>Currently I use 69 features, and the latest model has 86 features (haven't submitted). How bad my features are. Each submission cost me about $11 to train the model on AWS host with 256Gi memory for about 10 hours.</p>",
      "rawMarkdown": "Currently I use 69 features, and the latest model has 86 features (haven't submitted). How bad my features are. Each submission cost me about $11 to train the model on AWS host with 256Gi memory for about 10 hours.",
      "votes": null
    },
    {
      "id": "1125736",
      "postDate": "12/25/2020 02:50:45",
      "content": "<p>lol this is insane <a href=\"https://www.kaggle.com/wuwenmin\" target=\"_blank\">@wuwenmin</a> wish I had that money to train more data, you think you can make it work in inference pipeline with 69 feature ?</p>",
      "rawMarkdown": "lol this is insane @wuwenmin wish I had that money to train more data, you think you can make it work in inference pipeline with 69 feature ?",
      "votes": null
    },
    {
      "id": "1125791",
      "postDate": "12/25/2020 04:30:04",
      "content": "<p>Yup, my current best LB score is the one with 69 features. And the submission running time takes &lt; 40 mins. I will share a notebook this Sunday about the implementation detail. The basic ideas are:</p>\n<ol>\n<li><p>store the user features in {u_id: UserFeats} and question features in {q_id: QuesFeats} format to increase the cache hit rate when looking up. Many public notebooks store the user features in separate dicts which decreases the cache hit rate when looking up.</p></li>\n<li><p>extract the features with a list of functions. Following is my code snippet to extract features of each row in test_df. The feature extraction functions can be generated during initialization. Stupid if/else will also slow down your codes.</p></li>\n</ol>\n<pre><code>def _get_row_fvs(self, row: Series, u_feats: UserFeats, q_feats: QuesFeats) -&gt; List[float]:\n        cache = (...) # stats shared by multiple features\n        return [\n            fn(u_feats, q_feats, row, cache) for fn in self._feat_extract_fns\n        ]\n</code></pre>",
      "rawMarkdown": "Yup, my current best LB score is the one with 69 features. And the submission running time takes < 40 mins. I will share a notebook this Sunday about the implementation detail. The basic ideas are:\n\n1. store the user features in {u_id: UserFeats} and question features in {q_id: QuesFeats} format to increase the cache hit rate when looking up. Many public notebooks store the user features in separate dicts which decreases the cache hit rate when looking up.\n\n2. extract the features with a list of functions. Following is my code snippet to extract features of each row in test_df. The feature extraction functions can be generated during initialization. Stupid if/else will also slow down your codes.\n\n```Python\ndef _get_row_fvs(self, row: Series, u_feats: UserFeats, q_feats: QuesFeats) -> List[float]:\n        cache = (...) # stats shared by multiple features\n        return [\n            fn(u_feats, q_feats, row, cache) for fn in self._feat_extract_fns\n        ]\n```",
      "votes": null
    },
    {
      "id": "1125792",
      "postDate": "12/25/2020 04:30:32",
      "content": "<blockquote>\n  <p>In this discussion, one kaggler says 0.786 Single LGB model with about 14-16 features.. On the other hand, another kaggler says Single LGB model with about 50 features, CV:0.784 and LB:0.787.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/kokitanisaka\" target=\"_blank\">@kokitanisaka</a> I guess you are talking about our model. To answer your question, yes. Our model has very powerful features, so with less number of features we are getting a decent score. Some people like to call it magic features 😂</p>",
      "rawMarkdown": "> In this discussion, one kaggler says 0.786 Single LGB model with about 14-16 features.. On the other hand, another kaggler says Single LGB model with about 50 features, CV:0.784 and LB:0.787.\n\n@kokitanisaka I guess you are talking about our model. To answer your question, yes. Our model has very powerful features, so with less number of features we are getting a decent score. Some people like to call it magic features 😂",
      "votes": null
    },
    {
      "id": "1125817",
      "postDate": "12/25/2020 05:14:45",
      "content": "<p>Here's the <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206529\" target=\"_blank\">duscussion</a> about the more details. </p>",
      "rawMarkdown": "Here's the [duscussion](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206529) about the more details.",
      "votes": null
    },
    {
      "id": "1125849",
      "postDate": "12/25/2020 06:04:07",
      "content": "<p><a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a> <br>\nI'm looking forward to see your magic features after the competition has finished. 😄</p>",
      "rawMarkdown": "nikhilmishradev \nI'm looking forward to see your magic features after the competition has finished. 😄",
      "votes": null
    },
    {
      "id": "1127109",
      "postDate": "12/26/2020 08:50:06",
      "content": "<p>According to the official website LGBM <a href=\"https://lightgbm.readthedocs.io/en/latest/Parameters-Tuning.html#for-better-accuracy\" target=\"_blank\">here</a><br>\n<strong>For Better Accuracy</strong></p>\n<ul>\n<li>Use large max_bin (may be slower)</li>\n<li>Use small learning_rate with large num_iterations</li>\n<li>Use large num_leaves (may cause over-fitting)</li>\n<li>Use bigger training data</li>\n<li>Try dart</li>\n</ul>",
      "rawMarkdown": "According to the official website LGBM [here](https://lightgbm.readthedocs.io/en/latest/Parameters-Tuning.html#for-better-accuracy)\n**For Better Accuracy**\n- Use large max_bin (may be slower)\n- Use small learning_rate with large num_iterations\n- Use large num_leaves (may cause over-fitting)\n- Use bigger training data\n- Try dart",
      "votes": null
    },
    {
      "id": "1128388",
      "postDate": "12/27/2020 12:00:39",
      "content": "<p>I tried your mentions on your public notebook. But on using dart the execution time is too high. Its more than 9hrs I guess and the runtime gets cancelled.</p>",
      "rawMarkdown": "I tried your mentions on your public notebook. But on using dart the execution time is too high. Its more than 9hrs I guess and the runtime gets cancelled.",
      "votes": null
    },
    {
      "id": "1128419",
      "postDate": "12/27/2020 12:35:47",
      "content": "<p>optuna has a special integration for lgb that is very easy to use. <br>\nCheck this out: <br>\n<a href=\"https://medium.com/optuna/lightgbm-tuner-new-optuna-integration-for-hyperparameter-optimization-8b7095e99258\" target=\"_blank\">https://medium.com/optuna/lightgbm-tuner-new-optuna-integration-for-hyperparameter-optimization-8b7095e99258</a></p>\n<p>As I don't want to waste too much time in hyperparameter tunning, that's what I am using (using a smaller sample to go faster)</p>",
      "rawMarkdown": "optuna has a special integration for lgb that is very easy to use. \nCheck this out: \n[https://medium.com/optuna/lightgbm-tuner-new-optuna-integration-for-hyperparameter-optimization-8b7095e99258](https://medium.com/optuna/lightgbm-tuner-new-optuna-integration-for-hyperparameter-optimization-8b7095e99258)\n\nAs I don't want to waste too much time in hyperparameter tunning, that's what I am using (using a smaller sample to go faster)",
      "votes": null
    },
    {
      "id": "1128490",
      "postDate": "12/27/2020 13:33:10",
      "content": "<p>sure <a href=\"https://www.kaggle.com/kokitanisaka\" target=\"_blank\">@kokitanisaka</a>  :)</p>",
      "rawMarkdown": "sure @kokitanisaka  :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1123626,
      "author_name": "abdessalemboukil",
      "author_url": "",
      "post_date": "12/23/2020 11:34:35",
      "content": "<p>For my case learning rate search increased local validation from 0.773 -&gt; 0.785 , my local validation seems stable so I think LB score won't be far from it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1123707,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/23/2020 12:49:07",
          "content": "<p>Wow. That's impressive. Could you share some more details about how you searched the learning rate? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1124412,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "12/23/2020 23:06:18",
          "content": "<p><a href=\"https://www.kaggle.com/higepon\" target=\"_blank\">@higepon</a> Pretty straightforward, a grid search with twenty possible lr values</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1124420,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/23/2020 23:22:52",
          "content": "<p>Thank you!<br>\nOkay it sounds time consuming if you did it with all the training data.  Any tips to make it faster?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1124488,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "12/24/2020 01:42:38",
          "content": "<p>Please don't tune learning rate. Great documentation of Laurae, which is included in the official documentation of lightgbm (<a href=\"https://sites.google.com/view/lauraepp/parameters\" target=\"_blank\">https://sites.google.com/view/lauraepp/parameters</a>) says<br>\n<code>\nOnce your learning rate is fixed, do not change it.\nIt is not a good practice to consider the learning rate as a hyperparameter to tune.\nLearning rate should be tuned according to your training speed and performance tradeoff.\nDo not let an optimizer tune it. One must not expect to see an overfitting learning rate of 0.0202048.\n</code><br>\nPlease also remember <br>\n\"The smaller the learning rate is, the better the performance is.<br>\nThe smaller the learning rate is, the more time-consuming training is.\"</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1123719,
      "author_name": "nyanpn",
      "author_url": "",
      "post_date": "12/23/2020 13:06:39",
      "content": "<p>It depends on your features, but in general,  lightgbm parameter do not play a critical role in any competition. I believe this competition will be no exception.</p>\n<p>I have a collection of lightgbm parameters from previous top solutions in my repository. I hope this will be of some help.</p>\n<p><a href=\"https://github.com/nyanp/nyaggle/blob/master/nyaggle/hyper_parameters/lightgbm.py\" target=\"_blank\">https://github.com/nyanp/nyaggle/blob/master/nyaggle/hyper_parameters/lightgbm.py</a></p>\n<ul>\n<li>In all solutions, at least one of <code>max_depth</code> or <code>num_leaves</code> has been adjusted. This is one of the most important parameter of lightgbm.</li>\n<li>Many solutions make <code>learning_rate</code> smaller than the default (0.1). This will increase the accuracy a bit in most cases (but will increase the learning and inference time instead).</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1124408,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/23/2020 23:03:06",
          "content": "<p>This super helpful. Thanks for sharing the list of winning parameters.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1124884,
          "author_name": "kokitanisaka",
          "author_url": "",
          "post_date": "12/24/2020 08:55:48",
          "content": "<p>Thank you for sharing your great insight, <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> . That's pretty useful. </p>\n<p>Now I am wondering, if a guy achieved a good result with less features than others, is that mean the features are outstanding? <br>\nIn <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/204801\" target=\"_blank\">this discussion</a>, one kaggler says <code>0.786 Single LGB model with about 14-16 features.</code>. On the other hand, another kaggler says <code>Single LGB model with about 50 features, CV:0.784 and LB:0.787</code>. </p>\n<p>I didn't know how it happens.<br>\nSo it might be happening because of the nature of their own features? Not because of hyper parameter, am I right? <br>\nOr, sometimes even having less features performs well with LGBM? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125159,
          "author_name": "nyanpn",
          "author_url": "",
          "post_date": "12/24/2020 12:40:50",
          "content": "<p>This may be due to the difference of the features, or simply due to different CV methods. Rather than comparing the number of features with others, I recommend that you focus on your own features.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125733,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "12/25/2020 02:48:42",
          "content": "<p>Currently I use 69 features, and the latest model has 86 features (haven't submitted). How bad my features are. Each submission cost me about $11 to train the model on AWS host with 256Gi memory for about 10 hours.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125736,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "12/25/2020 02:50:45",
          "content": "<p>lol this is insane <a href=\"https://www.kaggle.com/wuwenmin\" target=\"_blank\">@wuwenmin</a> wish I had that money to train more data, you think you can make it work in inference pipeline with 69 feature ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125791,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "12/25/2020 04:30:04",
          "content": "<p>Yup, my current best LB score is the one with 69 features. And the submission running time takes &lt; 40 mins. I will share a notebook this Sunday about the implementation detail. The basic ideas are:</p>\n<ol>\n<li><p>store the user features in {u_id: UserFeats} and question features in {q_id: QuesFeats} format to increase the cache hit rate when looking up. Many public notebooks store the user features in separate dicts which decreases the cache hit rate when looking up.</p></li>\n<li><p>extract the features with a list of functions. Following is my code snippet to extract features of each row in test_df. The feature extraction functions can be generated during initialization. Stupid if/else will also slow down your codes.</p></li>\n</ol>\n<pre><code>def _get_row_fvs(self, row: Series, u_feats: UserFeats, q_feats: QuesFeats) -&gt; List[float]:\n        cache = (...) # stats shared by multiple features\n        return [\n            fn(u_feats, q_feats, row, cache) for fn in self._feat_extract_fns\n        ]\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125792,
          "author_name": "nikhilmishradev",
          "author_url": "",
          "post_date": "12/25/2020 04:30:32",
          "content": "<blockquote>\n  <p>In this discussion, one kaggler says 0.786 Single LGB model with about 14-16 features.. On the other hand, another kaggler says Single LGB model with about 50 features, CV:0.784 and LB:0.787.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/kokitanisaka\" target=\"_blank\">@kokitanisaka</a> I guess you are talking about our model. To answer your question, yes. Our model has very powerful features, so with less number of features we are getting a decent score. Some people like to call it magic features 😂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125817,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "12/25/2020 05:14:45",
          "content": "<p>Here's the <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206529\" target=\"_blank\">duscussion</a> about the more details. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125849,
          "author_name": "kokitanisaka",
          "author_url": "",
          "post_date": "12/25/2020 06:04:07",
          "content": "<p><a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a> <br>\nI'm looking forward to see your magic features after the competition has finished. 😄</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1128490,
          "author_name": "nikhilmishradev",
          "author_url": "",
          "post_date": "12/27/2020 13:33:10",
          "content": "<p>sure <a href=\"https://www.kaggle.com/kokitanisaka\" target=\"_blank\">@kokitanisaka</a>  :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1124708,
      "author_name": "higepon",
      "author_url": "",
      "post_date": "12/24/2020 06:43:43",
      "content": "<p>Please let me piggy back on this question because it's related.<br>\nI have LGBM model trained with full train data. It task N hours to train. <br>\nIf you do parameter search with grid search or optuna type of search, it may require M trials of train. IIUC it would take N * M hours to get the best param which may be expensive for some cases.</p>\n<p>Is there any way to mitigate that? eg) Maybe train with X % of data is sufficient for the search?</p>\n<p>Thank you in advance,<br>\nhigepon</p>",
      "votes": null,
      "replies": [
        {
          "id": 1124788,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "12/24/2020 07:51:31",
          "content": "<p>Yes, it's possible to only use X % of data for training, but the optimal parameter may be somewhat different from the one when the model is trained using 100% of data. <br>\nIf N is not too big, tuning only <code>num_leaves</code> using for loop is also a possible option. as nyanp says, other parameters are generally not so important. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1124837,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/24/2020 08:19:12",
          "content": "<p>Thank you mamas for your reply.<br>\nThat makes sense. I'll tune only num_leaves for now.</p>\n<p>Much appreciated.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1127109,
      "author_name": "ammarnassanalhajali",
      "author_url": "",
      "post_date": "12/26/2020 08:50:06",
      "content": "<p>According to the official website LGBM <a href=\"https://lightgbm.readthedocs.io/en/latest/Parameters-Tuning.html#for-better-accuracy\" target=\"_blank\">here</a><br>\n<strong>For Better Accuracy</strong></p>\n<ul>\n<li>Use large max_bin (may be slower)</li>\n<li>Use small learning_rate with large num_iterations</li>\n<li>Use large num_leaves (may cause over-fitting)</li>\n<li>Use bigger training data</li>\n<li>Try dart</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1128388,
          "author_name": "jagadish13",
          "author_url": "",
          "post_date": "12/27/2020 12:00:39",
          "content": "<p>I tried your mentions on your public notebook. But on using dart the execution time is too high. Its more than 9hrs I guess and the runtime gets cancelled.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1128419,
      "author_name": "bowaka",
      "author_url": "",
      "post_date": "12/27/2020 12:35:47",
      "content": "<p>optuna has a special integration for lgb that is very easy to use. <br>\nCheck this out: <br>\n<a href=\"https://medium.com/optuna/lightgbm-tuner-new-optuna-integration-for-hyperparameter-optimization-8b7095e99258\" target=\"_blank\">https://medium.com/optuna/lightgbm-tuner-new-optuna-integration-for-hyperparameter-optimization-8b7095e99258</a></p>\n<p>As I don't want to waste too much time in hyperparameter tunning, that's what I am using (using a smaller sample to go faster)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1123438": "I have a baseline LGBM model which is giving me around 0.782 on CV and 0.776 on the Public LB. I am using Loop feature engineering and a total of 18 features. I saw a few kernels having 44-47 features which is giving 0.775 on LB but their CV is slightly overfitting.\n\nI just wanted to know that is the parameter tuning of LGBM model playing a crucial part? Can the tuning increase the score by 0.5-1%? \n\nCan any of the top guys using LGBM please confirm?",
    "1123626": "For my case learning rate search increased local validation from 0.773 -> 0.785 , my local validation seems stable so I think LB score won't be far from it.",
    "1123707": "Wow. That's impressive. Could you share some more details about how you searched the learning rate?",
    "1123719": "It depends on your features, but in general,  lightgbm parameter do not play a critical role in any competition. I believe this competition will be no exception.\n\nI have a collection of lightgbm parameters from previous top solutions in my repository. I hope this will be of some help.\n\nhttps://github.com/nyanp/nyaggle/blob/master/nyaggle/hyper_parameters/lightgbm.py\n\n- In all solutions, at least one of `max_depth` or `num_leaves` has been adjusted. This is one of the most important parameter of lightgbm.\n- Many solutions make `learning_rate` smaller than the default (0.1). This will increase the accuracy a bit in most cases (but will increase the learning and inference time instead).",
    "1124408": "This super helpful. Thanks for sharing the list of winning parameters.",
    "1124412": "higepon Pretty straightforward, a grid search with twenty possible lr values",
    "1124420": "Thank you!\nOkay it sounds time consuming if you did it with all the training data.  Any tips to make it faster?",
    "1124488": "Please don't tune learning rate. Great documentation of Laurae, which is included in the official documentation of lightgbm (https://sites.google.com/view/lauraepp/parameters) says\n``\nOnce your learning rate is fixed, do not change it.\nIt is not a good practice to consider the learning rate as a hyperparameter to tune.\nLearning rate should be tuned according to your training speed and performance tradeoff.\nDo not let an optimizer tune it. One must not expect to see an overfitting learning rate of 0.0202048.\n``\nPlease also remember \n\"The smaller the learning rate is, the better the performance is.\nThe smaller the learning rate is, the more time-consuming training is.\"",
    "1124708": "Please let me piggy back on this question because it's related.\nI have LGBM model trained with full train data. It task N hours to train. \nIf you do parameter search with grid search or optuna type of search, it may require M trials of train. IIUC it would take N * M hours to get the best param which may be expensive for some cases.\n\nIs there any way to mitigate that? eg) Maybe train with X % of data is sufficient for the search?\n\nThank you in advance,\nhigepon",
    "1124788": "Yes, it's possible to only use X % of data for training, but the optimal parameter may be somewhat different from the one when the model is trained using 100% of data. \nIf N is not too big, tuning only `num_leaves` using for loop is also a possible option. as nyanp says, other parameters are generally not so important.",
    "1124837": "Thank you mamas for your reply.\nThat makes sense. I'll tune only num_leaves for now.\n\nMuch appreciated.",
    "1124884": "Thank you for sharing your great insight, @nyanpn . That's pretty useful. \n\nNow I am wondering, if a guy achieved a good result with less features than others, is that mean the features are outstanding? \nIn [this discussion](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/204801), one kaggler says `0.786 Single LGB model with about 14-16 features.`. On the other hand, another kaggler says `Single LGB model with about 50 features, CV:0.784 and LB:0.787`. \n\nI didn't know how it happens.\nSo it might be happening because of the nature of their own features? Not because of hyper parameter, am I right? \nOr, sometimes even having less features performs well with LGBM?",
    "1125159": "This may be due to the difference of the features, or simply due to different CV methods. Rather than comparing the number of features with others, I recommend that you focus on your own features.",
    "1125733": "Currently I use 69 features, and the latest model has 86 features (haven't submitted). How bad my features are. Each submission cost me about $11 to train the model on AWS host with 256Gi memory for about 10 hours.",
    "1125736": "lol this is insane @wuwenmin wish I had that money to train more data, you think you can make it work in inference pipeline with 69 feature ?",
    "1125791": "Yup, my current best LB score is the one with 69 features. And the submission running time takes < 40 mins. I will share a notebook this Sunday about the implementation detail. The basic ideas are:\n\n1. store the user features in {u_id: UserFeats} and question features in {q_id: QuesFeats} format to increase the cache hit rate when looking up. Many public notebooks store the user features in separate dicts which decreases the cache hit rate when looking up.\n\n2. extract the features with a list of functions. Following is my code snippet to extract features of each row in test_df. The feature extraction functions can be generated during initialization. Stupid if/else will also slow down your codes.\n\n```Python\ndef _get_row_fvs(self, row: Series, u_feats: UserFeats, q_feats: QuesFeats) -> List[float]:\n        cache = (...) # stats shared by multiple features\n        return [\n            fn(u_feats, q_feats, row, cache) for fn in self._feat_extract_fns\n        ]\n```",
    "1125792": "> In this discussion, one kaggler says 0.786 Single LGB model with about 14-16 features.. On the other hand, another kaggler says Single LGB model with about 50 features, CV:0.784 and LB:0.787.\n\n@kokitanisaka I guess you are talking about our model. To answer your question, yes. Our model has very powerful features, so with less number of features we are getting a decent score. Some people like to call it magic features 😂",
    "1125817": "Here's the [duscussion](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206529) about the more details.",
    "1125849": "nikhilmishradev \nI'm looking forward to see your magic features after the competition has finished. 😄",
    "1127109": "According to the official website LGBM [here](https://lightgbm.readthedocs.io/en/latest/Parameters-Tuning.html#for-better-accuracy)\n**For Better Accuracy**\n- Use large max_bin (may be slower)\n- Use small learning_rate with large num_iterations\n- Use large num_leaves (may cause over-fitting)\n- Use bigger training data\n- Try dart",
    "1128388": "I tried your mentions on your public notebook. But on using dart the execution time is too high. Its more than 9hrs I guess and the runtime gets cancelled.",
    "1128419": "optuna has a special integration for lgb that is very easy to use. \nCheck this out: \n[https://medium.com/optuna/lightgbm-tuner-new-optuna-integration-for-hyperparameter-optimization-8b7095e99258](https://medium.com/optuna/lightgbm-tuner-new-optuna-integration-for-hyperparameter-optimization-8b7095e99258)\n\nAs I don't want to waste too much time in hyperparameter tunning, that's what I am using (using a smaller sample to go faster)",
    "1128490": "sure @kokitanisaka  :)"
  },
  "source": "meta"
}