{
  "id": 93129,
  "title": "Using XGB, better CV gets better LB, but LBG not, why?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/93129",
  "author_name": "",
  "post_date": "2019-05-23T13:24:14.980497700Z",
  "votes": null,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I have been tuning parameters based on <a href=\"https://www.kaggle.com/vettejeep/masters-final-project-model-lb-1-392\">Vettejeep's kernel</a> and discovered something incomprehensible. </p>\n\n<p>I tuned the parameters one by one and observed the connection with CV and LB. When I use XGB, tuning parameters like \"learning_rate\" and \"max_depth\" can give me a better CV and LB score. However when I use LGB,  there appears to be almost no correlation between CV and LB, parameters with low CV score gets high score in LB.</p>\n\n<p>In addition, I found great CV by tuning \"num_features\" always got bad LB, both using XGB and LGB. These phenomena distressed me because without decrease in LB I can't be convinced these parameters are better.</p>\n\n<p>Has anyone encountered a similar situation? I'm not familiar with principles of XGB and LGB, can anyone explain these phenomenon? Thanks!</p>",
  "messages": [
    {
      "id": "535804",
      "postDate": "05/23/2019 13:24:14",
      "content": "<p>I have been tuning parameters based on <a href=\"https://www.kaggle.com/vettejeep/masters-final-project-model-lb-1-392\">Vettejeep's kernel</a> and discovered something incomprehensible. </p>\n\n<p>I tuned the parameters one by one and observed the connection with CV and LB. When I use XGB, tuning parameters like \"learning_rate\" and \"max_depth\" can give me a better CV and LB score. However when I use LGB,  there appears to be almost no correlation between CV and LB, parameters with low CV score gets high score in LB.</p>\n\n<p>In addition, I found great CV by tuning \"num_features\" always got bad LB, both using XGB and LGB. These phenomena distressed me because without decrease in LB I can't be convinced these parameters are better.</p>\n\n<p>Has anyone encountered a similar situation? I'm not familiar with principles of XGB and LGB, can anyone explain these phenomenon? Thanks!</p>",
      "rawMarkdown": "I have been tuning parameters based on [Vettejeep's kernel](https://www.kaggle.com/vettejeep/masters-final-project-model-lb-1-392) and discovered something incomprehensible. \n\nI tuned the parameters one by one and observed the connection with CV and LB. When I use XGB, tuning parameters like \"learning\\_rate\" and \"max\\_depth\" can give me a better CV and LB score. However when I use LGB,  there appears to be almost no correlation between CV and LB, parameters with low CV score gets high score in LB.\n\nIn addition, I found great CV by tuning \"num\\_features\" always got bad LB, both using XGB and LGB. These phenomena distressed me because without decrease in LB I can't be convinced these parameters are better.\n\nHas anyone encountered a similar situation? I'm not familiar with principles of XGB and LGB, can anyone explain these phenomenon? Thanks!",
      "votes": null
    },
    {
      "id": "535819",
      "postDate": "05/23/2019 13:43:31",
      "content": "<p>I always found xgb easier to tune than lgb.  I can't comment on your case as I did not use the code you start from, but I'm not that surprised.</p>",
      "rawMarkdown": "I always found xgb easier to tune than lgb.  I can't comment on your case as I did not use the code you start from, but I'm not that surprised.",
      "votes": null
    },
    {
      "id": "535835",
      "postDate": "05/23/2019 13:54:13",
      "content": "<p><a href=\"/wanliyu\">@wanliyu</a> you may find this article helpful <a href=\"https://towardsdatascience.com/catboost-vs-light-gbm-vs-xgboost-5f93620723db\">https://towardsdatascience.com/catboost-vs-light-gbm-vs-xgboost-5f93620723db</a></p>\n\n<p></p>\n\n<p>another topic in this competition you may find relevant.<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89909#latest-520304\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89909#latest-520304</a> .</p>",
      "rawMarkdown": "wanliyu you may find this article helpful https://towardsdatascience.com/catboost-vs-light-gbm-vs-xgboost-5f93620723db\n\n![](https://cdn-images-1.medium.com/max/2400/1*w05Hg2QZ5ioDi2OXdCCMiw.png)\n\nanother topic in this competition you may find relevant.https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89909#latest-520304 .",
      "votes": null
    },
    {
      "id": "535847",
      "postDate": "05/23/2019 14:10:35",
      "content": "<p>This article is crap.  Parameters are tuned for overfitting.</p>\n\n<p>For instance having both max depth 50 and min child weight 1 for xgb is just stupid.  You can see it with a very clear overfit to training data. And this large depth results in way too long running times.   </p>",
      "rawMarkdown": "This article is crap.  Parameters are tuned for overfitting.\n\nFor instance having both max depth 50 and min child weight 1 for xgb is just stupid.  You can see it with a very clear overfit to training data. And this large depth results in way too long running times.",
      "votes": null
    },
    {
      "id": "535874",
      "postDate": "05/23/2019 15:05:05",
      "content": "<p>I have experienced the similar issues. With 65 features and  1.89 for 8 fold cv my LB is 1.55 but with 200 features and 1.98  cv my LB is under 1.4.</p>",
      "rawMarkdown": "I have experienced the similar issues. With 65 features and  1.89 for 8 fold cv my LB is 1.55 but with 200 features and 1.98  cv my LB is under 1.4.",
      "votes": null
    },
    {
      "id": "535878",
      "postDate": "05/23/2019 15:11:17",
      "content": "<p>Do you have the same opinion for people who use <code>max_depth= -1</code>? I ask, because -1 depth means there is no limit, i.e. <code>max_depth = +infinity</code>. Is it considered bad practice to have depth of -1?</p>",
      "rawMarkdown": "Do you have the same opinion for people who use `max_depth= -1`? I ask, because -1 depth means there is no limit, i.e. `max_depth = +infinity`. Is it considered bad practice to have depth of -1?",
      "votes": null
    },
    {
      "id": "535882",
      "postDate": "05/23/2019 15:13:36",
      "content": "<p>It depends if you have other parameters for regularization.  If you use max_depth -1 with min child weight 1 then it is the worst possible case.  Both parameters are the main parameters to use to limit model complexity.</p>",
      "rawMarkdown": "It depends if you have other parameters for regularization.  If you use max_depth -1 with min child weight 1 then it is the worst possible case.  Both parameters are the main parameters to use to limit model complexity.",
      "votes": null
    },
    {
      "id": "535883",
      "postDate": "05/23/2019 15:14:40",
      "content": "<p>For lgb replace max depth by num leaves as the main complexity controlling parameter.</p>",
      "rawMarkdown": "For lgb replace max depth by num leaves as the main complexity controlling parameter.",
      "votes": null
    },
    {
      "id": "535887",
      "postDate": "05/23/2019 15:21:03",
      "content": "<p>Why do you think num leaves is the main complexity controlling parameter in lgbm, instead of max depth?</p>",
      "rawMarkdown": "Why do you think num leaves is the main complexity controlling parameter in lgbm, instead of max depth?",
      "votes": null
    },
    {
      "id": "535891",
      "postDate": "05/23/2019 15:24:06",
      "content": "<p>Because it is written in lgb documentation, see <a href=\"https://lightgbm.readthedocs.io/en/latest/Parameters-Tuning.html\">https://lightgbm.readthedocs.io/en/latest/Parameters-Tuning.html</a>.  </p>",
      "rawMarkdown": "Because it is written in lgb documentation, see https://lightgbm.readthedocs.io/en/latest/Parameters-Tuning.html.",
      "votes": null
    },
    {
      "id": "535900",
      "postDate": "05/23/2019 15:38:04",
      "content": "<p><a href=\"/returnofsputnik\">@returnofsputnik</a> lightgbm grows the tree with a \"leaf-wise\" approach, as opposed to xgboost which utilises a \"depth-wise\" approach. This is the reason why num_leaves should be the parameter to tune for lightgbm</p>",
      "rawMarkdown": "returnofsputnik lightgbm grows the tree with a \"leaf-wise\" approach, as opposed to xgboost which utilises a \"depth-wise\" approach. This is the reason why num_leaves should be the parameter to tune for lightgbm",
      "votes": null
    },
    {
      "id": "535921",
      "postDate": "05/23/2019 16:19:05",
      "content": "<p>I used LGB configured as follow:</p>\n\n<p><code>model = lgb.LGBMRegressor(\n        num_leaves=10, <br>\n        min_child_samples=9, <br>\n        objective='gamma', <br>\n        max_depth= 4, <br>\n        learning_rate= 0.001,\n        boosting_type= \"gbdt\", \n        metric= 'mae',\n        reg_alpha=1,\n        reg_lambda =1, \n        verbosity= -1,\n        random_state= 11, \n        n_estimators = 50_000, \n        n_jobs = -1)</code></p>\n\n<p><code>model.fit(X_tr, \n              y_tr, \n              eval_set=[(X_tr, y_tr), (X_val, y_val)], \n              eval_metric='mae',\n              verbose=10000, \n              early_stopping_rounds=500,\n              )</code></p>\n\n<p>I select parameters after doing some test, and now i do not change them until i have a mature model, i explain:</p>\n\n<blockquote>\n  <p><code>objective='gamma'</code></p>\n</blockquote>\n\n<p>Gamma distribution is sometimes used to model the waiting times in statistical process (ex queue theory).\nHere I do not kwow if all the hipothesys for a gamma distributition are true, but i have assumed (all here is an absumption :-)</p>\n\n<p>I have a little hardware, so i work with 4k of samples, and less than 30 features, so i have found this trade off:</p>\n\n<blockquote>\n  <p><code>num_leaves=10, <br>\n     min_child_samples=9, <br>\n     max_depth= 4,</code>           </p>\n</blockquote>\n\n<p>this parameters are control of complexity: Little trees, but not with a high \nmin_child_samples, because with my features, i have find that \nthere are some situation that are predicted better.</p>\n\n<blockquote>\n  <p><code>learning_rate= 0.001,</code></p>\n</blockquote>\n\n<p>you have to tune for the accuracy of the model</p>\n\n<blockquote>\n  <p><code>reg_alpha=1,\n      reg_lambda =1,</code> </p>\n</blockquote>\n\n<p>regularization, i find that this value are acceptable, i think i'll tune them at the end</p>\n\n<blockquote>\n  <p><code>early_stopping_rounds=500,</code></p>\n</blockquote>\n\n<p>very good (and fast) in my opinion!</p>\n\n<p>I hope it is useful.</p>",
      "rawMarkdown": "I used LGB configured as follow:\n\n`model = lgb.LGBMRegressor(\n        num_leaves=10,         \n        min_child_samples=9,   \n        objective='gamma',      \n        max_depth= 4,           \n        learning_rate= 0.001,\n        boosting_type= \"gbdt\", \n        metric= 'mae',\n        reg_alpha=1,\n        reg_lambda =1, \n        verbosity= -1,\n        random_state= 11, \n        n_estimators = 50_000, \n        n_jobs = -1)`\n\n`model.fit(X_tr, \n              y_tr, \n              eval_set=[(X_tr, y_tr), (X_val, y_val)], \n              eval_metric='mae',\n              verbose=10000, \n              early_stopping_rounds=500,\n              )`\n\n\nI select parameters after doing some test, and now i do not change them until i have a mature model, i explain:\n\n&gt; `objective='gamma'`\n\nGamma distribution is sometimes used to model the waiting times in statistical process (ex queue theory).\nHere I do not kwow if all the hipothesys for a gamma distributition are true, but i have assumed (all here is an absumption :-)\n\nI have a little hardware, so i work with 4k of samples, and less than 30 features, so i have found this trade off:\n&gt; `num_leaves=10,         \n   min_child_samples=9,    \n   max_depth= 4,`           \n\nthis parameters are control of complexity: Little trees, but not with a high \nmin_child_samples, because with my features, i have find that \nthere are some situation that are predicted better.\n\n&gt; `learning_rate= 0.001,`\n\nyou have to tune for the accuracy of the model\n\n&gt; `reg_alpha=1,\n    reg_lambda =1,` \n\nregularization, i find that this value are acceptable, i think i'll tune them at the end\n\n&gt; `early_stopping_rounds=500,`\n\nvery good (and fast) in my opinion!\n\nI hope it is useful.",
      "votes": null
    },
    {
      "id": "536107",
      "postDate": "05/24/2019 01:06:37",
      "content": "<p>It is useful, thank you very much.</p>",
      "rawMarkdown": "It is useful, thank you very much.",
      "votes": null
    },
    {
      "id": "536846",
      "postDate": "05/25/2019 12:11:08",
      "content": "<p>For reference, the LGB num_leaves parameters scales at 2**max_depth. So the LGB equivalent of max_depth=50 in XGB would be num_leaves over 1.125e15! Both their docs, and common sense, indicate that num_leaves should be much, much lower than this to avoid overfitting. Not to mention the runtime and RAM usage...</p>",
      "rawMarkdown": "For reference, the LGB num_leaves parameters scales at 2**max_depth. So the LGB equivalent of max_depth=50 in XGB would be num_leaves over 1.125e15! Both their docs, and common sense, indicate that num_leaves should be much, much lower than this to avoid overfitting. Not to mention the runtime and RAM usage...",
      "votes": null
    },
    {
      "id": "536865",
      "postDate": "05/25/2019 13:44:43",
      "content": "<p>Number of leaves cannot exceed the number of samples, therefore a very high limit means no limit.</p>",
      "rawMarkdown": "Number of leaves cannot exceed the number of samples, therefore a very high limit means no limit.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 535819,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/23/2019 13:43:31",
      "content": "<p>I always found xgb easier to tune than lgb.  I can't comment on your case as I did not use the code you start from, but I'm not that surprised.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 535835,
      "author_name": "cyberia",
      "author_url": "",
      "post_date": "05/23/2019 13:54:13",
      "content": "<p><a href=\"/wanliyu\">@wanliyu</a> you may find this article helpful <a href=\"https://towardsdatascience.com/catboost-vs-light-gbm-vs-xgboost-5f93620723db\">https://towardsdatascience.com/catboost-vs-light-gbm-vs-xgboost-5f93620723db</a></p>\n\n<p></p>\n\n<p>another topic in this competition you may find relevant.<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89909#latest-520304\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89909#latest-520304</a> .</p>",
      "votes": null,
      "replies": [
        {
          "id": 535847,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/23/2019 14:10:35",
          "content": "<p>This article is crap.  Parameters are tuned for overfitting.</p>\n\n<p>For instance having both max depth 50 and min child weight 1 for xgb is just stupid.  You can see it with a very clear overfit to training data. And this large depth results in way too long running times.   </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 535878,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "05/23/2019 15:11:17",
          "content": "<p>Do you have the same opinion for people who use <code>max_depth= -1</code>? I ask, because -1 depth means there is no limit, i.e. <code>max_depth = +infinity</code>. Is it considered bad practice to have depth of -1?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 535882,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/23/2019 15:13:36",
          "content": "<p>It depends if you have other parameters for regularization.  If you use max_depth -1 with min child weight 1 then it is the worst possible case.  Both parameters are the main parameters to use to limit model complexity.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 535883,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/23/2019 15:14:40",
          "content": "<p>For lgb replace max depth by num leaves as the main complexity controlling parameter.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 535887,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "05/23/2019 15:21:03",
          "content": "<p>Why do you think num leaves is the main complexity controlling parameter in lgbm, instead of max depth?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 535891,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/23/2019 15:24:06",
          "content": "<p>Because it is written in lgb documentation, see <a href=\"https://lightgbm.readthedocs.io/en/latest/Parameters-Tuning.html\">https://lightgbm.readthedocs.io/en/latest/Parameters-Tuning.html</a>.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 535900,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "05/23/2019 15:38:04",
          "content": "<p><a href=\"/returnofsputnik\">@returnofsputnik</a> lightgbm grows the tree with a \"leaf-wise\" approach, as opposed to xgboost which utilises a \"depth-wise\" approach. This is the reason why num_leaves should be the parameter to tune for lightgbm</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 536846,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "05/25/2019 12:11:08",
          "content": "<p>For reference, the LGB num_leaves parameters scales at 2**max_depth. So the LGB equivalent of max_depth=50 in XGB would be num_leaves over 1.125e15! Both their docs, and common sense, indicate that num_leaves should be much, much lower than this to avoid overfitting. Not to mention the runtime and RAM usage...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 536865,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/25/2019 13:44:43",
          "content": "<p>Number of leaves cannot exceed the number of samples, therefore a very high limit means no limit.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 535874,
      "author_name": "siavrez",
      "author_url": "",
      "post_date": "05/23/2019 15:05:05",
      "content": "<p>I have experienced the similar issues. With 65 features and  1.89 for 8 fold cv my LB is 1.55 but with 200 features and 1.98  cv my LB is under 1.4.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 535921,
      "author_name": "bellofilippo",
      "author_url": "",
      "post_date": "05/23/2019 16:19:05",
      "content": "<p>I used LGB configured as follow:</p>\n\n<p><code>model = lgb.LGBMRegressor(\n        num_leaves=10, <br>\n        min_child_samples=9, <br>\n        objective='gamma', <br>\n        max_depth= 4, <br>\n        learning_rate= 0.001,\n        boosting_type= \"gbdt\", \n        metric= 'mae',\n        reg_alpha=1,\n        reg_lambda =1, \n        verbosity= -1,\n        random_state= 11, \n        n_estimators = 50_000, \n        n_jobs = -1)</code></p>\n\n<p><code>model.fit(X_tr, \n              y_tr, \n              eval_set=[(X_tr, y_tr), (X_val, y_val)], \n              eval_metric='mae',\n              verbose=10000, \n              early_stopping_rounds=500,\n              )</code></p>\n\n<p>I select parameters after doing some test, and now i do not change them until i have a mature model, i explain:</p>\n\n<blockquote>\n  <p><code>objective='gamma'</code></p>\n</blockquote>\n\n<p>Gamma distribution is sometimes used to model the waiting times in statistical process (ex queue theory).\nHere I do not kwow if all the hipothesys for a gamma distributition are true, but i have assumed (all here is an absumption :-)</p>\n\n<p>I have a little hardware, so i work with 4k of samples, and less than 30 features, so i have found this trade off:</p>\n\n<blockquote>\n  <p><code>num_leaves=10, <br>\n     min_child_samples=9, <br>\n     max_depth= 4,</code>           </p>\n</blockquote>\n\n<p>this parameters are control of complexity: Little trees, but not with a high \nmin_child_samples, because with my features, i have find that \nthere are some situation that are predicted better.</p>\n\n<blockquote>\n  <p><code>learning_rate= 0.001,</code></p>\n</blockquote>\n\n<p>you have to tune for the accuracy of the model</p>\n\n<blockquote>\n  <p><code>reg_alpha=1,\n      reg_lambda =1,</code> </p>\n</blockquote>\n\n<p>regularization, i find that this value are acceptable, i think i'll tune them at the end</p>\n\n<blockquote>\n  <p><code>early_stopping_rounds=500,</code></p>\n</blockquote>\n\n<p>very good (and fast) in my opinion!</p>\n\n<p>I hope it is useful.</p>",
      "votes": null,
      "replies": [
        {
          "id": 536107,
          "author_name": "wanliyu",
          "author_url": "",
          "post_date": "05/24/2019 01:06:37",
          "content": "<p>It is useful, thank you very much.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "535804": "I have been tuning parameters based on [Vettejeep's kernel](https://www.kaggle.com/vettejeep/masters-final-project-model-lb-1-392) and discovered something incomprehensible. \n\nI tuned the parameters one by one and observed the connection with CV and LB. When I use XGB, tuning parameters like \"learning\\_rate\" and \"max\\_depth\" can give me a better CV and LB score. However when I use LGB,  there appears to be almost no correlation between CV and LB, parameters with low CV score gets high score in LB.\n\nIn addition, I found great CV by tuning \"num\\_features\" always got bad LB, both using XGB and LGB. These phenomena distressed me because without decrease in LB I can't be convinced these parameters are better.\n\nHas anyone encountered a similar situation? I'm not familiar with principles of XGB and LGB, can anyone explain these phenomenon? Thanks!",
    "535819": "I always found xgb easier to tune than lgb.  I can't comment on your case as I did not use the code you start from, but I'm not that surprised.",
    "535835": "wanliyu you may find this article helpful https://towardsdatascience.com/catboost-vs-light-gbm-vs-xgboost-5f93620723db\n\n![](https://cdn-images-1.medium.com/max/2400/1*w05Hg2QZ5ioDi2OXdCCMiw.png)\n\nanother topic in this competition you may find relevant.https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89909#latest-520304 .",
    "535847": "This article is crap.  Parameters are tuned for overfitting.\n\nFor instance having both max depth 50 and min child weight 1 for xgb is just stupid.  You can see it with a very clear overfit to training data. And this large depth results in way too long running times.",
    "535874": "I have experienced the similar issues. With 65 features and  1.89 for 8 fold cv my LB is 1.55 but with 200 features and 1.98  cv my LB is under 1.4.",
    "535878": "Do you have the same opinion for people who use `max_depth= -1`? I ask, because -1 depth means there is no limit, i.e. `max_depth = +infinity`. Is it considered bad practice to have depth of -1?",
    "535882": "It depends if you have other parameters for regularization.  If you use max_depth -1 with min child weight 1 then it is the worst possible case.  Both parameters are the main parameters to use to limit model complexity.",
    "535883": "For lgb replace max depth by num leaves as the main complexity controlling parameter.",
    "535887": "Why do you think num leaves is the main complexity controlling parameter in lgbm, instead of max depth?",
    "535891": "Because it is written in lgb documentation, see https://lightgbm.readthedocs.io/en/latest/Parameters-Tuning.html.",
    "535900": "returnofsputnik lightgbm grows the tree with a \"leaf-wise\" approach, as opposed to xgboost which utilises a \"depth-wise\" approach. This is the reason why num_leaves should be the parameter to tune for lightgbm",
    "535921": "I used LGB configured as follow:\n\n`model = lgb.LGBMRegressor(\n        num_leaves=10,         \n        min_child_samples=9,   \n        objective='gamma',      \n        max_depth= 4,           \n        learning_rate= 0.001,\n        boosting_type= \"gbdt\", \n        metric= 'mae',\n        reg_alpha=1,\n        reg_lambda =1, \n        verbosity= -1,\n        random_state= 11, \n        n_estimators = 50_000, \n        n_jobs = -1)`\n\n`model.fit(X_tr, \n              y_tr, \n              eval_set=[(X_tr, y_tr), (X_val, y_val)], \n              eval_metric='mae',\n              verbose=10000, \n              early_stopping_rounds=500,\n              )`\n\n\nI select parameters after doing some test, and now i do not change them until i have a mature model, i explain:\n\n&gt; `objective='gamma'`\n\nGamma distribution is sometimes used to model the waiting times in statistical process (ex queue theory).\nHere I do not kwow if all the hipothesys for a gamma distributition are true, but i have assumed (all here is an absumption :-)\n\nI have a little hardware, so i work with 4k of samples, and less than 30 features, so i have found this trade off:\n&gt; `num_leaves=10,         \n   min_child_samples=9,    \n   max_depth= 4,`           \n\nthis parameters are control of complexity: Little trees, but not with a high \nmin_child_samples, because with my features, i have find that \nthere are some situation that are predicted better.\n\n&gt; `learning_rate= 0.001,`\n\nyou have to tune for the accuracy of the model\n\n&gt; `reg_alpha=1,\n    reg_lambda =1,` \n\nregularization, i find that this value are acceptable, i think i'll tune them at the end\n\n&gt; `early_stopping_rounds=500,`\n\nvery good (and fast) in my opinion!\n\nI hope it is useful.",
    "536107": "It is useful, thank you very much.",
    "536846": "For reference, the LGB num_leaves parameters scales at 2**max_depth. So the LGB equivalent of max_depth=50 in XGB would be num_leaves over 1.125e15! Both their docs, and common sense, indicate that num_leaves should be much, much lower than this to avoid overfitting. Not to mention the runtime and RAM usage...",
    "536865": "Number of leaves cannot exceed the number of samples, therefore a very high limit means no limit."
  },
  "source": "meta"
}