{
  "id": 89994,
  "title": "LightGBM hyperparameter search",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/89994",
  "author_name": "",
  "post_date": "2019-04-19T11:42:56.019470200Z",
  "votes": 24,
  "comment_count": 15,
  "views": 0,
  "content": "<p>After finding <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89513\">this discussion</a> from <a href=\"/ogrellier\">@ogrellier</a>, I realized that there is room to play with the hyperparameters when using LightGBM if you decide to change the objective function.</p>\n\n<p>My current stats using LightGBM are the following:</p>\n\n<p>| objective | 5-fold CV (mean) | 5-fold CV (std) | Public LB |\n| ----------- | -------------------- | ------------------ | ------------| \n| regression |2.0613         | 0.0705       | 1.557 |\n| huber |2.0232  | 0.0765       | 1.521 |\n| fair   |2.0336          | 0.0750       | 1.503 |\n| gamma |2.0276          | 0.0693 | 1.508 |\n| mae  |2.0283          | 0.0807       | 1.535 |</p>\n\n<p>Essentially, I am performing a random search of the following grid:\n<code>python\nparam_grid = {\n    'num_leaves': list(range(8, 92, 4)),\n    'min_data_in_leaf': [10, 20, 40, 60, 100],\n    'max_depth': [3, 4, 5, 6, 8, 12, 16, -1],\n    'learning_rate': [0.1, 0.05, 0.01, 0.005],\n    'bagging_freq': [3, 4, 5, 6, 7],\n    'bagging_fraction': np.linspace(0.6, 0.95, 10),\n    'reg_alpha': np.linspace(0.1, 0.95, 10),\n    'reg_lambda': np.linspace(0.1, 0.95, 10)\n}\n</code>\nwhile keeping the following parameters fixed:\n<code>python\nfixed_params = {\n    'objective': 'huber',\n    'boosting': 'gbdt',\n    'verbosity': -1,\n    'random_seed': 19,\n    'n_estimators': 50000,\n    'metric': 'mae',\n    'bagging_seed': 11\n}\n</code>\nAs an example, after randomly sampling the grid 1500 times with <code>huber</code> as <code>objective</code>, I got the following fancy plot:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/519652/12995/huber.png\" alt=\"Performance plot\"></p>\n\n<p>I plotted in yellow the simulation that I got the best mean CV score, which resulted in the following parameters:\n<code>\nbest_params_lgb_huber = {\n    \"objective\": \"huber\",\n    \"boosting\": \"gbdt\",\n    \"verbosity\": -1,\n    \"num_leaves\": 12,\n    \"min_data_in_leaf\": 40,\n    \"max_depth\": 8,\n    \"learning_rate\": 0.005,\n    \"bagging_freq\": 4,\n    \"bagging_fraction\": 0.6,\n    \"bagging_seed\": 11,\n    \"random_seed\": 19,\n    \"metric\": \"mae\",\n    \"reg_alpha\": 0.47777777777777775,\n    \"reg_lambda\": 0.47777777777777775\n}\n</code>\nKeep in mind that I did not explore all potential combinations. However, it seems that there won't be much value in searching further.</p>\n\n<p>I'd be glad to know your experience on this. Did you get similar parameters?</p>",
  "messages": [
    {
      "id": "519652",
      "postDate": "04/19/2019 11:42:56",
      "content": "<p>After finding <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89513\">this discussion</a> from <a href=\"/ogrellier\">@ogrellier</a>, I realized that there is room to play with the hyperparameters when using LightGBM if you decide to change the objective function.</p>\n\n<p>My current stats using LightGBM are the following:</p>\n\n<p>| objective | 5-fold CV (mean) | 5-fold CV (std) | Public LB |\n| ----------- | -------------------- | ------------------ | ------------| \n| regression |2.0613         | 0.0705       | 1.557 |\n| huber |2.0232  | 0.0765       | 1.521 |\n| fair   |2.0336          | 0.0750       | 1.503 |\n| gamma |2.0276          | 0.0693 | 1.508 |\n| mae  |2.0283          | 0.0807       | 1.535 |</p>\n\n<p>Essentially, I am performing a random search of the following grid:\n<code>python\nparam_grid = {\n    'num_leaves': list(range(8, 92, 4)),\n    'min_data_in_leaf': [10, 20, 40, 60, 100],\n    'max_depth': [3, 4, 5, 6, 8, 12, 16, -1],\n    'learning_rate': [0.1, 0.05, 0.01, 0.005],\n    'bagging_freq': [3, 4, 5, 6, 7],\n    'bagging_fraction': np.linspace(0.6, 0.95, 10),\n    'reg_alpha': np.linspace(0.1, 0.95, 10),\n    'reg_lambda': np.linspace(0.1, 0.95, 10)\n}\n</code>\nwhile keeping the following parameters fixed:\n<code>python\nfixed_params = {\n    'objective': 'huber',\n    'boosting': 'gbdt',\n    'verbosity': -1,\n    'random_seed': 19,\n    'n_estimators': 50000,\n    'metric': 'mae',\n    'bagging_seed': 11\n}\n</code>\nAs an example, after randomly sampling the grid 1500 times with <code>huber</code> as <code>objective</code>, I got the following fancy plot:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/519652/12995/huber.png\" alt=\"Performance plot\"></p>\n\n<p>I plotted in yellow the simulation that I got the best mean CV score, which resulted in the following parameters:\n<code>\nbest_params_lgb_huber = {\n    \"objective\": \"huber\",\n    \"boosting\": \"gbdt\",\n    \"verbosity\": -1,\n    \"num_leaves\": 12,\n    \"min_data_in_leaf\": 40,\n    \"max_depth\": 8,\n    \"learning_rate\": 0.005,\n    \"bagging_freq\": 4,\n    \"bagging_fraction\": 0.6,\n    \"bagging_seed\": 11,\n    \"random_seed\": 19,\n    \"metric\": \"mae\",\n    \"reg_alpha\": 0.47777777777777775,\n    \"reg_lambda\": 0.47777777777777775\n}\n</code>\nKeep in mind that I did not explore all potential combinations. However, it seems that there won't be much value in searching further.</p>\n\n<p>I'd be glad to know your experience on this. Did you get similar parameters?</p>",
      "rawMarkdown": "After finding [this discussion](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89513) from @ogrellier, I realized that there is room to play with the hyperparameters when using LightGBM if you decide to change the objective function.\n\nMy current stats using LightGBM are the following:\n\n| objective | 5-fold CV (mean) | 5-fold CV (std) | Public LB |\n| ----------- | -------------------- | ------------------ | ------------| \n| regression |2.0613         | 0.0705       | 1.557 |\n| huber |2.0232  | 0.0765       | 1.521 |\n| fair   |2.0336          | 0.0750       | 1.503 |\n| gamma |2.0276          | 0.0693 | 1.508 |\n| mae  |2.0283          | 0.0807       | 1.535 |\n\nEssentially, I am performing a random search of the following grid:\n```python\nparam_grid = {\n    'num_leaves': list(range(8, 92, 4)),\n    'min_data_in_leaf': [10, 20, 40, 60, 100],\n    'max_depth': [3, 4, 5, 6, 8, 12, 16, -1],\n    'learning_rate': [0.1, 0.05, 0.01, 0.005],\n    'bagging_freq': [3, 4, 5, 6, 7],\n    'bagging_fraction': np.linspace(0.6, 0.95, 10),\n    'reg_alpha': np.linspace(0.1, 0.95, 10),\n    'reg_lambda': np.linspace(0.1, 0.95, 10)\n}\n```\nwhile keeping the following parameters fixed:\n```python\nfixed_params = {\n    'objective': 'huber',\n    'boosting': 'gbdt',\n    'verbosity': -1,\n    'random_seed': 19,\n    'n_estimators': 50000,\n    'metric': 'mae',\n    'bagging_seed': 11\n}\n```\nAs an example, after randomly sampling the grid 1500 times with `huber` as `objective`, I got the following fancy plot:\n![Performance plot](https://storage.googleapis.com/kaggle-forum-message-attachments/519652/12995/huber.png)\n\nI plotted in yellow the simulation that I got the best mean CV score, which resulted in the following parameters:\n```\nbest_params_lgb_huber = {\n    \"objective\": \"huber\",\n    \"boosting\": \"gbdt\",\n    \"verbosity\": -1,\n    \"num_leaves\": 12,\n    \"min_data_in_leaf\": 40,\n    \"max_depth\": 8,\n    \"learning_rate\": 0.005,\n    \"bagging_freq\": 4,\n    \"bagging_fraction\": 0.6,\n    \"bagging_seed\": 11,\n    \"random_seed\": 19,\n    \"metric\": \"mae\",\n    \"reg_alpha\": 0.47777777777777775,\n    \"reg_lambda\": 0.47777777777777775\n}\n```\nKeep in mind that I did not explore all potential combinations. However, it seems that there won't be much value in searching further.\n\nI'd be glad to know your experience on this. Did you get similar parameters?",
      "votes": null
    },
    {
      "id": "519668",
      "postDate": "04/19/2019 12:23:31",
      "content": "<p>Run your parameters on my setup, got CV 2.0615. The parameters I use give 2.0625, yours are marginally better. I use gamma objective, standard set of features, and some fancy CV strategy. </p>\n\n<p>Thanks for sharing.</p>",
      "rawMarkdown": "Run your parameters on my setup, got CV 2.0615. The parameters I use give 2.0625, yours are marginally better. I use gamma objective, standard set of features, and some fancy CV strategy. \n\nThanks for sharing.",
      "votes": null
    },
    {
      "id": "520106",
      "postDate": "04/20/2019 07:17:06",
      "content": "<p>Small changes such as the number of features, the scaling, or if you include the segments with ttf=0 may change the value a bit.</p>",
      "rawMarkdown": "Small changes such as the number of features, the scaling, or if you include the segments with ttf=0 may change the value a bit.",
      "votes": null
    },
    {
      "id": "520228",
      "postDate": "04/20/2019 13:25:07",
      "content": "<p>Hi, I am a newby. I have used Andrews script for datas and the followig configuration \nof LGBMRegressor and obtain an LB score of 1.484. I post some intersting part of code. <br>\nThere is a test split of 10% for validating model, and the other 90% subdivided in 5 Kfolds\nnote i used 'huber' and colsample_bytree with 0.11 only! In this way I think I prevent some noise\non datas..\nIf you have any suggest for upgrading the model..</p>\n\n<p>X_train_KFold, X_test_model, y_train_KFold, y_test_model = train_test_split(X_train_scaled, y_train, test_size=0.10, random_state=2)</p>\n\n<p>n_fold = 5\nn_repeats = 1</p>\n\n<p>folds=RepeatedKFold(n_splits=n_fold, n_repeats=n_repeats, random_state=42)</p>\n\n<p>for fold_, (trn_idx, val_idx) in enumerate(folds.split(X_train_KFold, y_train_KFold.values)):\n    .\n    .</p>\n\n<pre><code>model = lgb.LGBMRegressor(\n    num_leaves=15, \n    objective='huber',\n    max_depth= 4,\n    learning_rate= 0.01,\n    boosting_type= \"gbdt\",\n    colsample_bytree= 0.11, \n    subsample= 0.7,\n    subsample_freq = 45, \n    metric= 'mae',\n    reg_alpha=15, \n    verbosity= -1,\n    random_state= 0, \n    n_estimators = 500000, \n    n_jobs = -1)\nmodel.fit(X_tr, \n          y_tr, \n          eval_set=[(X_tr, y_tr), (X_val, y_val)], \n          eval_metric='mae',\n          verbose=10000, \n          early_stopping_rounds=250)\n\n.\n.\n</code></pre>\n\n<p>Thanks an bye!</p>",
      "rawMarkdown": "Hi, I am a newby. I have used Andrews script for datas and the followig configuration \nof LGBMRegressor and obtain an LB score of 1.484. I post some intersting part of code.   \nThere is a test split of 10% for validating model, and the other 90% subdivided in 5 Kfolds\nnote i used 'huber' and colsample_bytree with 0.11 only! In this way I think I prevent some noise\non datas..\nIf you have any suggest for upgrading the model..\n\nX_train_KFold, X_test_model, y_train_KFold, y_test_model = train_test_split(X_train_scaled, y_train, test_size=0.10, random_state=2)\n\nn_fold = 5\nn_repeats = 1\n\nfolds=RepeatedKFold(n_splits=n_fold, n_repeats=n_repeats, random_state=42)\n\nfor fold_, (trn_idx, val_idx) in enumerate(folds.split(X_train_KFold, y_train_KFold.values)):\n    .\n    .\n\n    model = lgb.LGBMRegressor(\n        num_leaves=15, \n        objective='huber',\n        max_depth= 4,\n        learning_rate= 0.01,\n        boosting_type= \"gbdt\",\n        colsample_bytree= 0.11, \n        subsample= 0.7,\n        subsample_freq = 45, \n        metric= 'mae',\n        reg_alpha=15, \n        verbosity= -1,\n        random_state= 0, \n        n_estimators = 500000, \n        n_jobs = -1)\n    model.fit(X_tr, \n              y_tr, \n              eval_set=[(X_tr, y_tr), (X_val, y_val)], \n              eval_metric='mae',\n              verbose=10000, \n              early_stopping_rounds=250)\n    \n    .\n    .\n\nThanks an bye!",
      "votes": null
    },
    {
      "id": "520253",
      "postDate": "04/20/2019 15:01:32",
      "content": "<p>Nice!</p>",
      "rawMarkdown": "Nice!",
      "votes": null
    },
    {
      "id": "520284",
      "postDate": "04/20/2019 16:16:58",
      "content": "<p>Optinal hyperparameters depend heavily on the features (type and number) and clearly the cv split</p>",
      "rawMarkdown": "Optinal hyperparameters depend heavily on the features (type and number) and clearly the cv split",
      "votes": null
    },
    {
      "id": "520306",
      "postDate": "04/20/2019 16:59:17",
      "content": "<p>That's right. I am mostly using Andrew's features. However, it's interesting to see that changing the objective does not have a significative impact, and that most of the different parameters obtain a similar CV (mean and std).</p>",
      "rawMarkdown": "That's right. I am mostly using Andrew's features. However, it's interesting to see that changing the objective does not have a significative impact, and that most of the different parameters obtain a similar CV (mean and std).",
      "votes": null
    },
    {
      "id": "520310",
      "postDate": "04/20/2019 17:03:58",
      "content": "<p>I wonder, how do you choose the 10% of samples that you set aside to validate the model? Isn't it a bit redundant since you are already using a k-fold to train and validate?</p>",
      "rawMarkdown": "I wonder, how do you choose the 10% of samples that you set aside to validate the model? Isn't it a bit redundant since you are already using a k-fold to train and validate?",
      "votes": null
    },
    {
      "id": "520354",
      "postDate": "04/20/2019 19:30:45",
      "content": "<p>In my opinion test datas have not the same distribution of train datas, so i use the split for remove some \"redundant\" training samples, and not only for validate model (as you say i already use a k-fold ). I wonder too, but in this way algorithms works better with this hyperparameters and my datas. My previous test without the 10% split have an LB &gt;= 1511. Probably there are other solutions and  this is: \"Fortuna del principiante\". :-).</p>",
      "rawMarkdown": "In my opinion test datas have not the same distribution of train datas, so i use the split for remove some \"redundant\" training samples, and not only for validate model (as you say i already use a k-fold ). I wonder too, but in this way algorithms works better with this hyperparameters and my datas. My previous test without the 10% split have an LB &gt;= 1511. Probably there are other solutions and  this is: \"Fortuna del principiante\". :-).",
      "votes": null
    },
    {
      "id": "520359",
      "postDate": "04/20/2019 19:54:07",
      "content": "<p><code>best_params = {'learning_rate': 0.002518228110909547, \n               'feature_fraction': 0.1996819351525092, \n               'num_leaves': 229, \n               'min_data_in_leaf': 41, \n               'max_depth': 10, \n               'reg_alpha': 22.568189541698203, \n               'reg_lambda': 21.955827524975568, \n               'subsample': 0.6516461691622116, \n               'colsample_bytree': 0.7348101977117862, \n               'min_child_weight': 48, \n               'min_split_gain': 6, \n               'top_rate': 0.12399670121120156, \n               'other_rate': 0.21293675994060435,\n               'objective': 'regression', \n               'metric': 'l1', \n               'verbosity': -1, \n               'n_jobs': -1, \n               'boosting_type': 'goss', \n               'task': 'train',\n              }</code></p>\n\n<p>9 find validation CV=1.9809, LB=1.596. I have around 400-500 features so your optimized parameter should be different depending on your feature engineering. I don't know why my model always have a bad LB score but I think I will put my faith in CV considering that the public LB is only tested on 13% of the data.</p>\n\n<p>Another tip is that you could use optuna package to do parameter tuning, my best params comes out of 1000 trials of different parameters.</p>",
      "rawMarkdown": "`best_params = {'learning_rate': 0.002518228110909547, \n               'feature_fraction': 0.1996819351525092, \n               'num_leaves': 229, \n               'min_data_in_leaf': 41, \n               'max_depth': 10, \n               'reg_alpha': 22.568189541698203, \n               'reg_lambda': 21.955827524975568, \n               'subsample': 0.6516461691622116, \n               'colsample_bytree': 0.7348101977117862, \n               'min_child_weight': 48, \n               'min_split_gain': 6, \n               'top_rate': 0.12399670121120156, \n               'other_rate': 0.21293675994060435,\n               'objective': 'regression', \n               'metric': 'l1', \n               'verbosity': -1, \n               'n_jobs': -1, \n               'boosting_type': 'goss', \n               'task': 'train',\n              }`\n\n9 find validation CV=1.9809, LB=1.596. I have around 400-500 features so your optimized parameter should be different depending on your feature engineering. I don't know why my model always have a bad LB score but I think I will put my faith in CV considering that the public LB is only tested on 13% of the data.\n\nAnother tip is that you could use optuna package to do parameter tuning, my best params comes out of 1000 trials of different parameters.",
      "votes": null
    },
    {
      "id": "520602",
      "postDate": "04/21/2019 11:47:32",
      "content": "<p>Thanks for sharing! Indeed, I would put my faith on the CV rather than the LB score.\nI didn't know about <em>optuna</em>. I'll surely take a look. </p>",
      "rawMarkdown": "Thanks for sharing! Indeed, I would put my faith on the CV rather than the LB score.\nI didn't know about *optuna*. I'll surely take a look.",
      "votes": null
    },
    {
      "id": "520986",
      "postDate": "04/22/2019 06:05:56",
      "content": "<p>Thanks. Any suggestion of correction is welcome :)</p>",
      "rawMarkdown": "Thanks. Any suggestion of correction is welcome :)",
      "votes": null
    },
    {
      "id": "521669",
      "postDate": "04/23/2019 08:27:23",
      "content": "<p>I have experienced from my test that a best CV values have a better LB score in the LB range 1500 - 1515. Probably this is a good range for a working solution also in the final score, what do you think? \nFor example with a CV of 2.018583641491404 using 170 features, 6 Kfold repeated 2 times (with no other splits :-) ) and  LGBRegressor i have an LB of 1504\nmodel = lgb.LGBMRegressor(\n        num_leaves=6, \n        objective='huber',\n        max_depth= 4,\n        learning_rate= 0.01,\n        boosting_type= \"gbdt\",\n        colsample_bytree= 0.45, <br>\n        subsample= 0.80,\n        subsample_freq = 5, <br>\n        bagging_seed= 11,\n        metric= 'mae',\n        reg_alpha= 1, \n        reg_lambda = 0.45, #0.45\n        verbosity= -1,\n        random_state= 0, \n        n_estimators = 500000, \n        n_jobs = -1) </p>\n\n<p>Of course configuration depends heavily from the number and type of features; i try to mix them in a balanced (in my opinion) way, from statistic features, FFT, Sta_Lta, quantiles etc.. \nProbably a good balance start with features selection, the problem is what features i have to select? Visualizing distribution of datas in the training set and in the test set and discover if they have completely different values... what do you think? how do you suggest to select features?\nthx and bye!</p>",
      "rawMarkdown": "I have experienced from my test that a best CV values have a better LB score in the LB range 1500 - 1515. Probably this is a good range for a working solution also in the final score, what do you think? \nFor example with a CV of 2.018583641491404 using 170 features, 6 Kfold repeated 2 times (with no other splits :-) ) and  LGBRegressor i have an LB of 1504\nmodel = lgb.LGBMRegressor(\n        num_leaves=6, \n        objective='huber',\n        max_depth= 4,\n        learning_rate= 0.01,\n        boosting_type= \"gbdt\",\n        colsample_bytree= 0.45,  \n        subsample= 0.80,\n        subsample_freq = 5,  \n        bagging_seed= 11,\n        metric= 'mae',\n        reg_alpha= 1, \n        reg_lambda = 0.45, #0.45\n        verbosity= -1,\n        random_state= 0, \n        n_estimators = 500000, \n        n_jobs = -1) \n\nOf course configuration depends heavily from the number and type of features; i try to mix them in a balanced (in my opinion) way, from statistic features, FFT, Sta_Lta, quantiles etc.. \nProbably a good balance start with features selection, the problem is what features i have to select? Visualizing distribution of datas in the training set and in the test set and discover if they have completely different values... what do you think? how do you suggest to select features?\nthx and bye!",
      "votes": null
    },
    {
      "id": "521752",
      "postDate": "04/23/2019 11:17:08",
      "content": "<p>I've created a streamlined automated parameter-tuning kernel that might be of help - you can get the code here:\n<a href=\"https://www.kaggle.com/bigironsphere/parameter-tuning-in-one-function-with-hyperopt\">https://www.kaggle.com/bigironsphere/parameter-tuning-in-one-function-with-hyperopt</a></p>",
      "rawMarkdown": "I've created a streamlined automated parameter-tuning kernel that might be of help - you can get the code here:\nhttps://www.kaggle.com/bigironsphere/parameter-tuning-in-one-function-with-hyperopt",
      "votes": null
    },
    {
      "id": "521778",
      "postDate": "04/23/2019 12:05:08",
      "content": "<p>Thanks for the toolbox. I'll surely check it out :)</p>",
      "rawMarkdown": "Thanks for the toolbox. I'll surely check it out :)",
      "votes": null
    },
    {
      "id": "533798",
      "postDate": "05/20/2019 04:18:04",
      "content": "<p>I'm also using optuna for hyperparameter tuning.\nHere is a d<a href=\"https://optuna.readthedocs.io/en/stable/tutorial/first.html\">ocument</a> and there is a <a href=\"https://github.com/pfnet/optuna/blob/master/examples/lightgbm_simple.py\">simple example for lightGBM</a> :)</p>",
      "rawMarkdown": "I'm also using optuna for hyperparameter tuning.\nHere is a d[ocument](https://optuna.readthedocs.io/en/stable/tutorial/first.html) and there is a [simple example for lightGBM](https://github.com/pfnet/optuna/blob/master/examples/lightgbm_simple.py) :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 519668,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "04/19/2019 12:23:31",
      "content": "<p>Run your parameters on my setup, got CV 2.0615. The parameters I use give 2.0625, yours are marginally better. I use gamma objective, standard set of features, and some fancy CV strategy. </p>\n\n<p>Thanks for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 520106,
          "author_name": "ricarddelgado",
          "author_url": "",
          "post_date": "04/20/2019 07:17:06",
          "content": "<p>Small changes such as the number of features, the scaling, or if you include the segments with ttf=0 may change the value a bit.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 520228,
      "author_name": "bellofilippo",
      "author_url": "",
      "post_date": "04/20/2019 13:25:07",
      "content": "<p>Hi, I am a newby. I have used Andrews script for datas and the followig configuration \nof LGBMRegressor and obtain an LB score of 1.484. I post some intersting part of code. <br>\nThere is a test split of 10% for validating model, and the other 90% subdivided in 5 Kfolds\nnote i used 'huber' and colsample_bytree with 0.11 only! In this way I think I prevent some noise\non datas..\nIf you have any suggest for upgrading the model..</p>\n\n<p>X_train_KFold, X_test_model, y_train_KFold, y_test_model = train_test_split(X_train_scaled, y_train, test_size=0.10, random_state=2)</p>\n\n<p>n_fold = 5\nn_repeats = 1</p>\n\n<p>folds=RepeatedKFold(n_splits=n_fold, n_repeats=n_repeats, random_state=42)</p>\n\n<p>for fold_, (trn_idx, val_idx) in enumerate(folds.split(X_train_KFold, y_train_KFold.values)):\n    .\n    .</p>\n\n<pre><code>model = lgb.LGBMRegressor(\n    num_leaves=15, \n    objective='huber',\n    max_depth= 4,\n    learning_rate= 0.01,\n    boosting_type= \"gbdt\",\n    colsample_bytree= 0.11, \n    subsample= 0.7,\n    subsample_freq = 45, \n    metric= 'mae',\n    reg_alpha=15, \n    verbosity= -1,\n    random_state= 0, \n    n_estimators = 500000, \n    n_jobs = -1)\nmodel.fit(X_tr, \n          y_tr, \n          eval_set=[(X_tr, y_tr), (X_val, y_val)], \n          eval_metric='mae',\n          verbose=10000, \n          early_stopping_rounds=250)\n\n.\n.\n</code></pre>\n\n<p>Thanks an bye!</p>",
      "votes": null,
      "replies": [
        {
          "id": 520310,
          "author_name": "ricarddelgado",
          "author_url": "",
          "post_date": "04/20/2019 17:03:58",
          "content": "<p>I wonder, how do you choose the 10% of samples that you set aside to validate the model? Isn't it a bit redundant since you are already using a k-fold to train and validate?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 520354,
          "author_name": "bellofilippo",
          "author_url": "",
          "post_date": "04/20/2019 19:30:45",
          "content": "<p>In my opinion test datas have not the same distribution of train datas, so i use the split for remove some \"redundant\" training samples, and not only for validate model (as you say i already use a k-fold ). I wonder too, but in this way algorithms works better with this hyperparameters and my datas. My previous test without the 10% split have an LB &gt;= 1511. Probably there are other solutions and  this is: \"Fortuna del principiante\". :-).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 520253,
      "author_name": "hemang999",
      "author_url": "",
      "post_date": "04/20/2019 15:01:32",
      "content": "<p>Nice!</p>",
      "votes": null,
      "replies": [
        {
          "id": 520986,
          "author_name": "ricarddelgado",
          "author_url": "",
          "post_date": "04/22/2019 06:05:56",
          "content": "<p>Thanks. Any suggestion of correction is welcome :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 520284,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "04/20/2019 16:16:58",
      "content": "<p>Optinal hyperparameters depend heavily on the features (type and number) and clearly the cv split</p>",
      "votes": null,
      "replies": [
        {
          "id": 520306,
          "author_name": "ricarddelgado",
          "author_url": "",
          "post_date": "04/20/2019 16:59:17",
          "content": "<p>That's right. I am mostly using Andrew's features. However, it's interesting to see that changing the objective does not have a significative impact, and that most of the different parameters obtain a similar CV (mean and std).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 520359,
      "author_name": "cjinny",
      "author_url": "",
      "post_date": "04/20/2019 19:54:07",
      "content": "<p><code>best_params = {'learning_rate': 0.002518228110909547, \n               'feature_fraction': 0.1996819351525092, \n               'num_leaves': 229, \n               'min_data_in_leaf': 41, \n               'max_depth': 10, \n               'reg_alpha': 22.568189541698203, \n               'reg_lambda': 21.955827524975568, \n               'subsample': 0.6516461691622116, \n               'colsample_bytree': 0.7348101977117862, \n               'min_child_weight': 48, \n               'min_split_gain': 6, \n               'top_rate': 0.12399670121120156, \n               'other_rate': 0.21293675994060435,\n               'objective': 'regression', \n               'metric': 'l1', \n               'verbosity': -1, \n               'n_jobs': -1, \n               'boosting_type': 'goss', \n               'task': 'train',\n              }</code></p>\n\n<p>9 find validation CV=1.9809, LB=1.596. I have around 400-500 features so your optimized parameter should be different depending on your feature engineering. I don't know why my model always have a bad LB score but I think I will put my faith in CV considering that the public LB is only tested on 13% of the data.</p>\n\n<p>Another tip is that you could use optuna package to do parameter tuning, my best params comes out of 1000 trials of different parameters.</p>",
      "votes": null,
      "replies": [
        {
          "id": 520602,
          "author_name": "ricarddelgado",
          "author_url": "",
          "post_date": "04/21/2019 11:47:32",
          "content": "<p>Thanks for sharing! Indeed, I would put my faith on the CV rather than the LB score.\nI didn't know about <em>optuna</em>. I'll surely take a look. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 533798,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "05/20/2019 04:18:04",
          "content": "<p>I'm also using optuna for hyperparameter tuning.\nHere is a d<a href=\"https://optuna.readthedocs.io/en/stable/tutorial/first.html\">ocument</a> and there is a <a href=\"https://github.com/pfnet/optuna/blob/master/examples/lightgbm_simple.py\">simple example for lightGBM</a> :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 521669,
      "author_name": "bellofilippo",
      "author_url": "",
      "post_date": "04/23/2019 08:27:23",
      "content": "<p>I have experienced from my test that a best CV values have a better LB score in the LB range 1500 - 1515. Probably this is a good range for a working solution also in the final score, what do you think? \nFor example with a CV of 2.018583641491404 using 170 features, 6 Kfold repeated 2 times (with no other splits :-) ) and  LGBRegressor i have an LB of 1504\nmodel = lgb.LGBMRegressor(\n        num_leaves=6, \n        objective='huber',\n        max_depth= 4,\n        learning_rate= 0.01,\n        boosting_type= \"gbdt\",\n        colsample_bytree= 0.45, <br>\n        subsample= 0.80,\n        subsample_freq = 5, <br>\n        bagging_seed= 11,\n        metric= 'mae',\n        reg_alpha= 1, \n        reg_lambda = 0.45, #0.45\n        verbosity= -1,\n        random_state= 0, \n        n_estimators = 500000, \n        n_jobs = -1) </p>\n\n<p>Of course configuration depends heavily from the number and type of features; i try to mix them in a balanced (in my opinion) way, from statistic features, FFT, Sta_Lta, quantiles etc.. \nProbably a good balance start with features selection, the problem is what features i have to select? Visualizing distribution of datas in the training set and in the test set and discover if they have completely different values... what do you think? how do you suggest to select features?\nthx and bye!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 521752,
      "author_name": "bigironsphere",
      "author_url": "",
      "post_date": "04/23/2019 11:17:08",
      "content": "<p>I've created a streamlined automated parameter-tuning kernel that might be of help - you can get the code here:\n<a href=\"https://www.kaggle.com/bigironsphere/parameter-tuning-in-one-function-with-hyperopt\">https://www.kaggle.com/bigironsphere/parameter-tuning-in-one-function-with-hyperopt</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 521778,
          "author_name": "ricarddelgado",
          "author_url": "",
          "post_date": "04/23/2019 12:05:08",
          "content": "<p>Thanks for the toolbox. I'll surely check it out :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "519652": "After finding [this discussion](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89513) from @ogrellier, I realized that there is room to play with the hyperparameters when using LightGBM if you decide to change the objective function.\n\nMy current stats using LightGBM are the following:\n\n| objective | 5-fold CV (mean) | 5-fold CV (std) | Public LB |\n| ----------- | -------------------- | ------------------ | ------------| \n| regression |2.0613         | 0.0705       | 1.557 |\n| huber |2.0232  | 0.0765       | 1.521 |\n| fair   |2.0336          | 0.0750       | 1.503 |\n| gamma |2.0276          | 0.0693 | 1.508 |\n| mae  |2.0283          | 0.0807       | 1.535 |\n\nEssentially, I am performing a random search of the following grid:\n```python\nparam_grid = {\n    'num_leaves': list(range(8, 92, 4)),\n    'min_data_in_leaf': [10, 20, 40, 60, 100],\n    'max_depth': [3, 4, 5, 6, 8, 12, 16, -1],\n    'learning_rate': [0.1, 0.05, 0.01, 0.005],\n    'bagging_freq': [3, 4, 5, 6, 7],\n    'bagging_fraction': np.linspace(0.6, 0.95, 10),\n    'reg_alpha': np.linspace(0.1, 0.95, 10),\n    'reg_lambda': np.linspace(0.1, 0.95, 10)\n}\n```\nwhile keeping the following parameters fixed:\n```python\nfixed_params = {\n    'objective': 'huber',\n    'boosting': 'gbdt',\n    'verbosity': -1,\n    'random_seed': 19,\n    'n_estimators': 50000,\n    'metric': 'mae',\n    'bagging_seed': 11\n}\n```\nAs an example, after randomly sampling the grid 1500 times with `huber` as `objective`, I got the following fancy plot:\n![Performance plot](https://storage.googleapis.com/kaggle-forum-message-attachments/519652/12995/huber.png)\n\nI plotted in yellow the simulation that I got the best mean CV score, which resulted in the following parameters:\n```\nbest_params_lgb_huber = {\n    \"objective\": \"huber\",\n    \"boosting\": \"gbdt\",\n    \"verbosity\": -1,\n    \"num_leaves\": 12,\n    \"min_data_in_leaf\": 40,\n    \"max_depth\": 8,\n    \"learning_rate\": 0.005,\n    \"bagging_freq\": 4,\n    \"bagging_fraction\": 0.6,\n    \"bagging_seed\": 11,\n    \"random_seed\": 19,\n    \"metric\": \"mae\",\n    \"reg_alpha\": 0.47777777777777775,\n    \"reg_lambda\": 0.47777777777777775\n}\n```\nKeep in mind that I did not explore all potential combinations. However, it seems that there won't be much value in searching further.\n\nI'd be glad to know your experience on this. Did you get similar parameters?",
    "519668": "Run your parameters on my setup, got CV 2.0615. The parameters I use give 2.0625, yours are marginally better. I use gamma objective, standard set of features, and some fancy CV strategy. \n\nThanks for sharing.",
    "520106": "Small changes such as the number of features, the scaling, or if you include the segments with ttf=0 may change the value a bit.",
    "520228": "Hi, I am a newby. I have used Andrews script for datas and the followig configuration \nof LGBMRegressor and obtain an LB score of 1.484. I post some intersting part of code.   \nThere is a test split of 10% for validating model, and the other 90% subdivided in 5 Kfolds\nnote i used 'huber' and colsample_bytree with 0.11 only! In this way I think I prevent some noise\non datas..\nIf you have any suggest for upgrading the model..\n\nX_train_KFold, X_test_model, y_train_KFold, y_test_model = train_test_split(X_train_scaled, y_train, test_size=0.10, random_state=2)\n\nn_fold = 5\nn_repeats = 1\n\nfolds=RepeatedKFold(n_splits=n_fold, n_repeats=n_repeats, random_state=42)\n\nfor fold_, (trn_idx, val_idx) in enumerate(folds.split(X_train_KFold, y_train_KFold.values)):\n    .\n    .\n\n    model = lgb.LGBMRegressor(\n        num_leaves=15, \n        objective='huber',\n        max_depth= 4,\n        learning_rate= 0.01,\n        boosting_type= \"gbdt\",\n        colsample_bytree= 0.11, \n        subsample= 0.7,\n        subsample_freq = 45, \n        metric= 'mae',\n        reg_alpha=15, \n        verbosity= -1,\n        random_state= 0, \n        n_estimators = 500000, \n        n_jobs = -1)\n    model.fit(X_tr, \n              y_tr, \n              eval_set=[(X_tr, y_tr), (X_val, y_val)], \n              eval_metric='mae',\n              verbose=10000, \n              early_stopping_rounds=250)\n    \n    .\n    .\n\nThanks an bye!",
    "520253": "Nice!",
    "520284": "Optinal hyperparameters depend heavily on the features (type and number) and clearly the cv split",
    "520306": "That's right. I am mostly using Andrew's features. However, it's interesting to see that changing the objective does not have a significative impact, and that most of the different parameters obtain a similar CV (mean and std).",
    "520310": "I wonder, how do you choose the 10% of samples that you set aside to validate the model? Isn't it a bit redundant since you are already using a k-fold to train and validate?",
    "520354": "In my opinion test datas have not the same distribution of train datas, so i use the split for remove some \"redundant\" training samples, and not only for validate model (as you say i already use a k-fold ). I wonder too, but in this way algorithms works better with this hyperparameters and my datas. My previous test without the 10% split have an LB &gt;= 1511. Probably there are other solutions and  this is: \"Fortuna del principiante\". :-).",
    "520359": "`best_params = {'learning_rate': 0.002518228110909547, \n               'feature_fraction': 0.1996819351525092, \n               'num_leaves': 229, \n               'min_data_in_leaf': 41, \n               'max_depth': 10, \n               'reg_alpha': 22.568189541698203, \n               'reg_lambda': 21.955827524975568, \n               'subsample': 0.6516461691622116, \n               'colsample_bytree': 0.7348101977117862, \n               'min_child_weight': 48, \n               'min_split_gain': 6, \n               'top_rate': 0.12399670121120156, \n               'other_rate': 0.21293675994060435,\n               'objective': 'regression', \n               'metric': 'l1', \n               'verbosity': -1, \n               'n_jobs': -1, \n               'boosting_type': 'goss', \n               'task': 'train',\n              }`\n\n9 find validation CV=1.9809, LB=1.596. I have around 400-500 features so your optimized parameter should be different depending on your feature engineering. I don't know why my model always have a bad LB score but I think I will put my faith in CV considering that the public LB is only tested on 13% of the data.\n\nAnother tip is that you could use optuna package to do parameter tuning, my best params comes out of 1000 trials of different parameters.",
    "520602": "Thanks for sharing! Indeed, I would put my faith on the CV rather than the LB score.\nI didn't know about *optuna*. I'll surely take a look.",
    "520986": "Thanks. Any suggestion of correction is welcome :)",
    "521669": "I have experienced from my test that a best CV values have a better LB score in the LB range 1500 - 1515. Probably this is a good range for a working solution also in the final score, what do you think? \nFor example with a CV of 2.018583641491404 using 170 features, 6 Kfold repeated 2 times (with no other splits :-) ) and  LGBRegressor i have an LB of 1504\nmodel = lgb.LGBMRegressor(\n        num_leaves=6, \n        objective='huber',\n        max_depth= 4,\n        learning_rate= 0.01,\n        boosting_type= \"gbdt\",\n        colsample_bytree= 0.45,  \n        subsample= 0.80,\n        subsample_freq = 5,  \n        bagging_seed= 11,\n        metric= 'mae',\n        reg_alpha= 1, \n        reg_lambda = 0.45, #0.45\n        verbosity= -1,\n        random_state= 0, \n        n_estimators = 500000, \n        n_jobs = -1) \n\nOf course configuration depends heavily from the number and type of features; i try to mix them in a balanced (in my opinion) way, from statistic features, FFT, Sta_Lta, quantiles etc.. \nProbably a good balance start with features selection, the problem is what features i have to select? Visualizing distribution of datas in the training set and in the test set and discover if they have completely different values... what do you think? how do you suggest to select features?\nthx and bye!",
    "521752": "I've created a streamlined automated parameter-tuning kernel that might be of help - you can get the code here:\nhttps://www.kaggle.com/bigironsphere/parameter-tuning-in-one-function-with-hyperopt",
    "521778": "Thanks for the toolbox. I'll surely check it out :)",
    "533798": "I'm also using optuna for hyperparameter tuning.\nHere is a d[ocument](https://optuna.readthedocs.io/en/stable/tutorial/first.html) and there is a [simple example for lightGBM](https://github.com/pfnet/optuna/blob/master/examples/lightgbm_simple.py) :)"
  },
  "source": "meta"
}