{
  "id": 89513,
  "title": "Simple way to boost your score ?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/89513",
  "author_name": "",
  "post_date": "2019-04-15T07:30:37.055439800Z",
  "votes": 67,
  "comment_count": 36,
  "views": 0,
  "content": "<p>I'd like to thank @inversion for his <a href=\"https://www.kaggle.com/inversion/basic-feature-benchmark\">benchmark kernel</a> that rang a bell in my empty brain (it's till echoing though ...)</p>\n\n<p>For those of you using LightGBM, here are the results of a simple experiment:</p>\n\n<p>| objective | 10-fold CV | Public LB\n| --- | --- | --- |\n| regression | 2.1675 | 1.589 | \n| huber | 2.1597 | not tested | \n| fair | 2.1463 | not tested | \n| gamma | 2.1208 |  1.548 | </p>\n\n<p>I'm really far down the LB so it may not work for everyone but it's simple enough to test.</p>\n\n<p>In any case it's very interesting to blend predictions from models with different objectives :)</p>\n\n<p>Good luck.</p>",
  "messages": [
    {
      "id": "516911",
      "postDate": "04/15/2019 07:30:37",
      "content": "<p>I'd like to thank @inversion for his <a href=\"https://www.kaggle.com/inversion/basic-feature-benchmark\">benchmark kernel</a> that rang a bell in my empty brain (it's till echoing though ...)</p>\n\n<p>For those of you using LightGBM, here are the results of a simple experiment:</p>\n\n<p>| objective | 10-fold CV | Public LB\n| --- | --- | --- |\n| regression | 2.1675 | 1.589 | \n| huber | 2.1597 | not tested | \n| fair | 2.1463 | not tested | \n| gamma | 2.1208 |  1.548 | </p>\n\n<p>I'm really far down the LB so it may not work for everyone but it's simple enough to test.</p>\n\n<p>In any case it's very interesting to blend predictions from models with different objectives :)</p>\n\n<p>Good luck.</p>",
      "rawMarkdown": "I'd like to thank @inversion for his [benchmark kernel](https://www.kaggle.com/inversion/basic-feature-benchmark) that rang a bell in my empty brain (it's till echoing though ...)\n\nFor those of you using LightGBM, here are the results of a simple experiment:\n\n| objective | 10-fold CV | Public LB\n| --- | --- | --- |\n| regression | 2.1675 | 1.589 | \n| huber | 2.1597 | not tested | \n| fair | 2.1463 | not tested | \n| gamma | 2.1208 |  1.548 | \n\nI'm really far down the LB so it may not work for everyone but it's simple enough to test.\n\nIn any case it's very interesting to blend predictions from models with different objectives :)\n\nGood luck.",
      "votes": null
    },
    {
      "id": "516971",
      "postDate": "04/15/2019 09:28:53",
      "content": "<p>Nice!</p>",
      "rawMarkdown": "Nice!",
      "votes": null
    },
    {
      "id": "516981",
      "postDate": "04/15/2019 09:56:19",
      "content": "<p>Thanks <a href=\"/ogrellier\">@ogrellier</a> !</p>",
      "rawMarkdown": "Thanks @ogrellier !",
      "votes": null
    },
    {
      "id": "517015",
      "postDate": "04/15/2019 11:31:59",
      "content": "<p>Thanks Olivier. I tried, but my CV with \"regression\" is still the best, so is my LB. Can anyone else please check this?</p>",
      "rawMarkdown": "Thanks Olivier. I tried, but my CV with \"regression\" is still the best, so is my LB. Can anyone else please check this?",
      "votes": null
    },
    {
      "id": "517025",
      "postDate": "04/15/2019 11:43:28",
      "content": "<p>I've been using lightgbm with objective=\"mae\".</p>",
      "rawMarkdown": "I've been using lightgbm with objective=\"mae\".",
      "votes": null
    },
    {
      "id": "517028",
      "postDate": "04/15/2019 11:52:50",
      "content": "<p>Thanks <a href=\"/dslate\">@dslate</a>, totally forgot to test MAE. </p>\n\n<p>My CV score with objective \"mae\" is still higher than with gamma regression but lower than MSE.</p>",
      "rawMarkdown": "Thanks @dslate, totally forgot to test MAE. \n\nMy CV score with objective \"mae\" is still higher than with gamma regression but lower than MSE.",
      "votes": null
    },
    {
      "id": "517029",
      "postDate": "04/15/2019 11:53:57",
      "content": "<p>Thanks for your feedback <a href=\"/khahuras\">@khahuras</a>, I suspect other hyper parameters play a role here.</p>",
      "rawMarkdown": "Thanks for your feedback @khahuras, I suspect other hyper parameters play a role here.",
      "votes": null
    },
    {
      "id": "517141",
      "postDate": "04/15/2019 15:34:03",
      "content": "<p>Thanks Olivier 👍</p>",
      "rawMarkdown": "Thanks Olivier 👍",
      "votes": null
    },
    {
      "id": "517178",
      "postDate": "04/15/2019 17:14:34",
      "content": "<p>That is a great observation! I did the experiment as well with 5-fold CV and shuffling</p>\n\n<p>| objective | 5-fold CV (mean) | 5-fold CV (std) | Public LB |\n| ----------- | -------------------- | ------------------ | ------------| \n| regression                   |2.0613         | 0.0705      | 1.557 |\n| huber        |<strong>2.0232</strong>  | 0.0765       | 1.521 |\n| fair            |2.0336          | 0.0750       | <strong>1.503</strong> |\n| gamma     |2.0276          | <strong>0.0693</strong> | 1.508 |\n| mae           |2.0283          | 0.0807       | 1.535 |</p>\n\n<p><em>Huber</em> seemed to give me the best CV mean score and <em>gamma</em> the best std. However, the difference in the std seem very little to me and just just be chance. My best LB score with LightGBM  is <em>fair</em>.</p>",
      "rawMarkdown": "That is a great observation! I did the experiment as well with 5-fold CV and shuffling\n\n| objective | 5-fold CV (mean) | 5-fold CV (std) | Public LB |\n| ----------- | -------------------- | ------------------ | ------------| \n| regression                   |2.0613         | 0.0705      | 1.557 |\n| huber        |**2.0232**  | 0.0765       | 1.521 |\n| fair            |2.0336          | 0.0750       | **1.503** |\n| gamma     |2.0276          | **0.0693** | 1.508 |\n| mae           |2.0283          | 0.0807       | 1.535 |\n\n*Huber* seemed to give me the best CV mean score and *gamma* the best std. However, the difference in the std seem very little to me and just just be chance. My best LB score with LightGBM  is *fair*.",
      "votes": null
    },
    {
      "id": "517416",
      "postDate": "04/16/2019 02:08:43",
      "content": "<p>Thanks Oliver. I've tried 3 objectives; regression, gamma, and huber.</p>\n\n<p>| Obj | 10 fold CV  | Public LB |\n| --- | --- | --- |\n| regression | 2.180 | 1.621 |\n| gamma | 2.136 | 1.518 |\n| huber | 2.182 |1.560|</p>\n\n<p>In my environment, the best model is \"gamma\" objective, and there are little difference between \"regression\" and \"huber\" on local CV.\nHyper parameter is based on this kernel: <a href=\"https://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction\">https://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction</a></p>",
      "rawMarkdown": "Thanks Oliver. I've tried 3 objectives; regression, gamma, and huber.\n\n| Obj | 10 fold CV  | Public LB |\n| --- | --- | --- |\n| regression | 2.180 | 1.621 |\n| gamma | 2.136 | 1.518 |\n| huber | 2.182 |1.560|\n\nIn my environment, the best model is \"gamma\" objective, and there are little difference between \"regression\" and \"huber\" on local CV.\nHyper parameter is based on this kernel: https://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction",
      "votes": null
    },
    {
      "id": "517457",
      "postDate": "04/16/2019 03:39:16",
      "content": "<p>change regression_l1 to gamma,  LB 1.527 to 1.486   </p>",
      "rawMarkdown": "change regression_l1 to gamma,  LB 1.527 to 1.486",
      "votes": null
    },
    {
      "id": "517474",
      "postDate": "04/16/2019 04:20:47",
      "content": "<p>May I ask what CV are you using (shuffle k-fold or leave-1-quake-out?), and does your 1.486 sub have better CV? In my case gamma reg is worse in both CV and LB.</p>",
      "rawMarkdown": "May I ask what CV are you using (shuffle k-fold or leave-1-quake-out?), and does your 1.486 sub have better CV? In my case gamma reg is worse in both CV and LB.",
      "votes": null
    },
    {
      "id": "517502",
      "postDate": "04/16/2019 05:38:56",
      "content": "<p>Can you add a column with the subs mean predicted ttf? </p>",
      "rawMarkdown": "Can you add a column with the subs mean predicted ttf?",
      "votes": null
    },
    {
      "id": "517507",
      "postDate": "04/16/2019 05:50:57",
      "content": "<p><a href=\"/ricarddelgado\">@ricarddelgado</a>, when I searched for the right objective I initially thought Huber would be best and this is what you find. However the target seems to have a gamma distribution... </p>",
      "rawMarkdown": "ricarddelgado, when I searched for the right objective I initially thought Huber would be best and this is what you find. However the target seems to have a gamma distribution...",
      "votes": null
    },
    {
      "id": "517508",
      "postDate": "04/16/2019 05:52:28",
      "content": "<p>I used shuffle k-fold, and I got a worse cv  from 2.015 to 2.037</p>",
      "rawMarkdown": "I used shuffle k-fold, and I got a worse cv  from 2.015 to 2.037",
      "votes": null
    },
    {
      "id": "517529",
      "postDate": "04/16/2019 06:28:46",
      "content": "<p>I'll do some hyperparameter optimization to see how far I can go with the others :) </p>",
      "rawMarkdown": "I'll do some hyperparameter optimization to see how far I can go with the others :)",
      "votes": null
    },
    {
      "id": "517531",
      "postDate": "04/16/2019 06:34:57",
      "content": "<p>The countdown? </p>\n\n<p>Or the duration of inter-quake intervals?</p>",
      "rawMarkdown": "The countdown? \n\nOr the duration of inter-quake intervals?",
      "votes": null
    },
    {
      "id": "517572",
      "postDate": "04/16/2019 08:03:41",
      "content": "<p>Olivier, congratulate you with Discussion GM Title. Thank you for interest insights which help us a lot on the way. Good luck in new journeys )</p>",
      "rawMarkdown": "Olivier, congratulate you with Discussion GM Title. Thank you for interest insights which help us a lot on the way. Good luck in new journeys )",
      "votes": null
    },
    {
      "id": "517586",
      "postDate": "04/16/2019 08:28:53",
      "content": "<p>Here, the performance is on the ttf, so, the countdown.</p>",
      "rawMarkdown": "Here, the performance is on the ttf, so, the countdown.",
      "votes": null
    },
    {
      "id": "517591",
      "postDate": "04/16/2019 08:34:00",
      "content": "<p>I tested both mae and rmse to train the model, rmse gives me a better cv than mae (from 1.9 to 1.8 mae) but not submitted yet </p>",
      "rawMarkdown": "I tested both mae and rmse to train the model, rmse gives me a better cv than mae (from 1.9 to 1.8 mae) but not submitted yet",
      "votes": null
    },
    {
      "id": "517624",
      "postDate": "04/16/2019 09:26:52",
      "content": "<p>Thanks <a href=\"/alexanderkireev\">@alexanderkireev</a> for your kind words ! Good luck with the competition.</p>",
      "rawMarkdown": "Thanks @alexanderkireev for your kind words ! Good luck with the competition.",
      "votes": null
    },
    {
      "id": "517954",
      "postDate": "04/16/2019 17:36:34",
      "content": "<p>Thanks @Olivier. I applied Kolmogorov-Smirnov test and filtered features . Shuffled  huber : CV = 2.025  , LB=1.516</p>",
      "rawMarkdown": "Thanks @Olivier. I applied Kolmogorov-Smirnov test and filtered features . Shuffled  huber : CV = 2.025  , LB=1.516",
      "votes": null
    },
    {
      "id": "517962",
      "postDate": "04/16/2019 17:45:46",
      "content": "<p>Cramér-von Mises?</p>",
      "rawMarkdown": "Cramér-von Mises?",
      "votes": null
    },
    {
      "id": "517993",
      "postDate": "04/16/2019 18:32:14",
      "content": "<p>2 Sample Kolmogorov-Smirnov :  *scipy.stats.ks_2samp*</p>\n\n<p>for whom wants to know more:\n<a href=\"https://stats.stackexchange.com/questions/201434/2-sample-kolmogorov-smirnov-vs-anderson-darling-vs-cramer-von-mises\">https://stats.stackexchange.com/questions/201434/2-sample-kolmogorov-smirnov-vs-anderson-darling-vs-cramer-von-mises</a></p>",
      "rawMarkdown": "2 Sample Kolmogorov-Smirnov :  *scipy.stats.ks_2samp*\n\nfor whom wants to know more:\nhttps://stats.stackexchange.com/questions/201434/2-sample-kolmogorov-smirnov-vs-anderson-darling-vs-cramer-von-mises",
      "votes": null
    },
    {
      "id": "518440",
      "postDate": "04/17/2019 07:50:12",
      "content": "<p>in my experience, trying kinds of objectives always bring surprise :)</p>",
      "rawMarkdown": "in my experience, trying kinds of objectives always bring surprise :)",
      "votes": null
    },
    {
      "id": "518510",
      "postDate": "04/17/2019 11:47:13",
      "content": "<p>There are even stranger things going on using different objectives with shuffling. </p>\n\n<p>I need a few submissions to check out local results but it seems that using MAE and shuffling would lead to the best CV score (by a lot). I assume the model overfits but I need to check that out on the LB (though 13% is quite small any sort of solid conclusions). </p>\n\n<p>Without shuffling Gamma is still best for me.</p>",
      "rawMarkdown": "There are even stranger things going on using different objectives with shuffling. \n\nI need a few submissions to check out local results but it seems that using MAE and shuffling would lead to the best CV score (by a lot). I assume the model overfits but I need to check that out on the LB (though 13% is quite small any sort of solid conclusions). \n\nWithout shuffling Gamma is still best for me.",
      "votes": null
    },
    {
      "id": "518633",
      "postDate": "04/17/2019 14:43:10",
      "content": "<p>Thanks a lot for sharing it.</p>",
      "rawMarkdown": "Thanks a lot for sharing it.",
      "votes": null
    },
    {
      "id": "518759",
      "postDate": "04/17/2019 18:42:26",
      "content": "<p>What do you mean by shuffling? does it make sense to shuffle this type of data?</p>",
      "rawMarkdown": "What do you mean by shuffling? does it make sense to shuffle this type of data?",
      "votes": null
    },
    {
      "id": "518782",
      "postDate": "04/17/2019 19:38:55",
      "content": "<p><a href=\"/mhviraf\">@mhviraf</a>, with shuffling the model learns from all types of earthquakes at the same time. Without shuffling you're able to see where earthquakes don't generalize. I believe both ways are useful.</p>",
      "rawMarkdown": "mhviraf, with shuffling the model learns from all types of earthquakes at the same time. Without shuffling you're able to see where earthquakes don't generalize. I believe both ways are useful.",
      "votes": null
    },
    {
      "id": "518804",
      "postDate": "04/17/2019 20:52:39",
      "content": "<p>I am confused about shuffling. Are you talking about shuffling ~4100 training segments we have? or shuffling rows of each of those segments? </p>",
      "rawMarkdown": "I am confused about shuffling. Are you talking about shuffling ~4100 training segments we have? or shuffling rows of each of those segments?",
      "votes": null
    },
    {
      "id": "519085",
      "postDate": "04/18/2019 10:44:10",
      "content": "<p><a href=\"/mhviraf\">@mhviraf</a>, sorry about any confusion. I'm talking about shuffling the training segments or if you prefer setting the shuffle parameter of KFold to True.</p>\n\n<p>I don't think shuffling the samples in each segment would make any sense and I did not try :)</p>",
      "rawMarkdown": "mhviraf, sorry about any confusion. I'm talking about shuffling the training segments or if you prefer setting the shuffle parameter of KFold to True.\n\nI don't think shuffling the samples in each segment would make any sense and I did not try :)",
      "votes": null
    },
    {
      "id": "519132",
      "postDate": "04/18/2019 12:18:07",
      "content": "<p>In case someone wants to change the objective in xgboost, pasting some objectives implementations below. All of them give similar scores for me, and better than default MSE. All three functions are approximations for MAE. Please tell me if you get interesting results with it :)</p>\n\n<p><code>\n1. sqrt(1 + (x/h)**2)\n2. ln(cosh(x))\n3. log(exp(-x) + exp(x)) \n</code>\nimplementations\n```\ndef huberobj(preds, dtrain):\n    d = preds - dtrain.get_label()\n    h = 1\n    scale = 1 + (d / h) ** 2\n    scale_sqrt = np.sqrt(scale)\n    grad = d / scale_sqrt / h\n    hess = 1 / scale / scale_sqrt / h\n    return grad, hess</p>\n\n<p>def logcoshobj(preds, dtrain):\n    labels = dtrain.get_label()\n    grad = np.tanh(preds - labels)\n    hess = 1.0 - grad*grad\n    return grad, hess</p>\n\n<p>def logexpexp(preds, dtrain):\n    labels = dtrain.get_label()\n    x = preds - labels\n    grad = (np.exp(2.0*x) - 1) / (np.exp(2.0*x) + 1)\n    hess = (4.0*np.exp(2.0*x)) / (np.exp(2.0*x) + 1)**2 \n    return grad, hess</p>\n\n<p>model = xgb.train(...,  obj=logcoshobj)</p>\n\n<p>```</p>",
      "rawMarkdown": "In case someone wants to change the objective in xgboost, pasting some objectives implementations below. All of them give similar scores for me, and better than default MSE. All three functions are approximations for MAE. Please tell me if you get interesting results with it :)\n\n```\n1. sqrt(1 + (x/h)**2)\n2. ln(cosh(x))\n3. log(exp(-x) + exp(x)) \n```\nimplementations\n```\ndef huberobj(preds, dtrain):\n    d = preds - dtrain.get_label()\n    h = 1\n    scale = 1 + (d / h) ** 2\n    scale_sqrt = np.sqrt(scale)\n    grad = d / scale_sqrt / h\n    hess = 1 / scale / scale_sqrt / h\n    return grad, hess\n\ndef logcoshobj(preds, dtrain):\n    labels = dtrain.get_label()\n    grad = np.tanh(preds - labels)\n    hess = 1.0 - grad*grad\n    return grad, hess\n\ndef logexpexp(preds, dtrain):\n    labels = dtrain.get_label()\n    x = preds - labels\n    grad = (np.exp(2.0*x) - 1) / (np.exp(2.0*x) + 1)\n    hess = (4.0*np.exp(2.0*x)) / (np.exp(2.0*x) + 1)**2 \n    return grad, hess\n\nmodel = xgb.train(...,  obj=logcoshobj)\n\n```",
      "votes": null
    },
    {
      "id": "519237",
      "postDate": "04/18/2019 15:29:32",
      "content": "<p>I tried again huber, mae, regression, gamma with different n_folds and different hyper-parameters  . but didn't change CV strategy yet (e,g, cv based on quake id).  I think only features to use and better cv strategy will make difference and improve the result among different lgb models. what do you think ?</p>",
      "rawMarkdown": "I tried again huber, mae, regression, gamma with different n_folds and different hyper-parameters  . but didn't change CV strategy yet (e,g, cv based on quake id).  I think only features to use and better cv strategy will make difference and improve the result among different lgb models. what do you think ?",
      "votes": null
    },
    {
      "id": "519244",
      "postDate": "04/18/2019 15:44:25",
      "content": "<p>Thanks. That's why I was confused too. </p>",
      "rawMarkdown": "Thanks. That's why I was confused too.",
      "votes": null
    },
    {
      "id": "519586",
      "postDate": "04/19/2019 08:33:40",
      "content": "<p>ah, distributions matter normally</p>",
      "rawMarkdown": "ah, distributions matter normally",
      "votes": null
    },
    {
      "id": "519628",
      "postDate": "04/19/2019 10:17:55",
      "content": "<p>Yes setting shuffle=True into Kfold gave me a way better CV, but it performed very badly on the LB. I am trying to explore which is the best splitting option</p>",
      "rawMarkdown": "Yes setting shuffle=True into Kfold gave me a way better CV, but it performed very badly on the LB. I am trying to explore which is the best splitting option",
      "votes": null
    },
    {
      "id": "522456",
      "postDate": "04/24/2019 13:00:27",
      "content": "<p>For those using XGBoost, the improvement using a 5-fold CV with following parameters:\nparams = {\n    \"learning_rate\": 0.001,\n    \"max_depth\": 3,\n    \"n_estimators\": 10000,\n    \"min_child_weight\": 4,\n    \"colsample_bytree\": 1,\n    \"subsample\": 0.8,\n    \"nthread\": 12,\n    \"random_state\": 42,\n}, \nthe improvement recorded up on using reg:gamma as objective is minimal. \n<strong>|  objective    | 5-fold CV score |\n|  reg:linear   |       2.027358      |\n| reg:gamma |       2.026449     |</strong></p>",
      "rawMarkdown": "For those using XGBoost, the improvement using a 5-fold CV with following parameters:\nparams = {\n    \"learning_rate\": 0.001,\n    \"max_depth\": 3,\n    \"n_estimators\": 10000,\n    \"min_child_weight\": 4,\n    \"colsample_bytree\": 1,\n    \"subsample\": 0.8,\n    \"nthread\": 12,\n    \"random_state\": 42,\n}, \nthe improvement recorded up on using reg:gamma as objective is minimal. \n**|  objective    | 5-fold CV score |\n|  reg:linear   |       2.027358      |\n| reg:gamma |       2.026449     |**",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 516971,
      "author_name": "darbin",
      "author_url": "",
      "post_date": "04/15/2019 09:28:53",
      "content": "<p>Nice!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 516981,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "04/15/2019 09:56:19",
      "content": "<p>Thanks <a href=\"/ogrellier\">@ogrellier</a> !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 517015,
      "author_name": "khahuras",
      "author_url": "",
      "post_date": "04/15/2019 11:31:59",
      "content": "<p>Thanks Olivier. I tried, but my CV with \"regression\" is still the best, so is my LB. Can anyone else please check this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 517029,
          "author_name": "ogrellier",
          "author_url": "",
          "post_date": "04/15/2019 11:53:57",
          "content": "<p>Thanks for your feedback <a href=\"/khahuras\">@khahuras</a>, I suspect other hyper parameters play a role here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 519237,
          "author_name": "arashnic",
          "author_url": "",
          "post_date": "04/18/2019 15:29:32",
          "content": "<p>I tried again huber, mae, regression, gamma with different n_folds and different hyper-parameters  . but didn't change CV strategy yet (e,g, cv based on quake id).  I think only features to use and better cv strategy will make difference and improve the result among different lgb models. what do you think ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 517025,
      "author_name": "dslate",
      "author_url": "",
      "post_date": "04/15/2019 11:43:28",
      "content": "<p>I've been using lightgbm with objective=\"mae\".</p>",
      "votes": null,
      "replies": [
        {
          "id": 517028,
          "author_name": "ogrellier",
          "author_url": "",
          "post_date": "04/15/2019 11:52:50",
          "content": "<p>Thanks <a href=\"/dslate\">@dslate</a>, totally forgot to test MAE. </p>\n\n<p>My CV score with objective \"mae\" is still higher than with gamma regression but lower than MSE.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 517591,
          "author_name": "kerzer",
          "author_url": "",
          "post_date": "04/16/2019 08:34:00",
          "content": "<p>I tested both mae and rmse to train the model, rmse gives me a better cv than mae (from 1.9 to 1.8 mae) but not submitted yet </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 517141,
      "author_name": "salehaf",
      "author_url": "",
      "post_date": "04/15/2019 15:34:03",
      "content": "<p>Thanks Olivier 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 517178,
      "author_name": "ricarddelgado",
      "author_url": "",
      "post_date": "04/15/2019 17:14:34",
      "content": "<p>That is a great observation! I did the experiment as well with 5-fold CV and shuffling</p>\n\n<p>| objective | 5-fold CV (mean) | 5-fold CV (std) | Public LB |\n| ----------- | -------------------- | ------------------ | ------------| \n| regression                   |2.0613         | 0.0705      | 1.557 |\n| huber        |<strong>2.0232</strong>  | 0.0765       | 1.521 |\n| fair            |2.0336          | 0.0750       | <strong>1.503</strong> |\n| gamma     |2.0276          | <strong>0.0693</strong> | 1.508 |\n| mae           |2.0283          | 0.0807       | 1.535 |</p>\n\n<p><em>Huber</em> seemed to give me the best CV mean score and <em>gamma</em> the best std. However, the difference in the std seem very little to me and just just be chance. My best LB score with LightGBM  is <em>fair</em>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 517507,
          "author_name": "ogrellier",
          "author_url": "",
          "post_date": "04/16/2019 05:50:57",
          "content": "<p><a href=\"/ricarddelgado\">@ricarddelgado</a>, when I searched for the right objective I initially thought Huber would be best and this is what you find. However the target seems to have a gamma distribution... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 517529,
          "author_name": "ricarddelgado",
          "author_url": "",
          "post_date": "04/16/2019 06:28:46",
          "content": "<p>I'll do some hyperparameter optimization to see how far I can go with the others :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 517531,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "04/16/2019 06:34:57",
          "content": "<p>The countdown? </p>\n\n<p>Or the duration of inter-quake intervals?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 517586,
          "author_name": "ricarddelgado",
          "author_url": "",
          "post_date": "04/16/2019 08:28:53",
          "content": "<p>Here, the performance is on the ttf, so, the countdown.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 517416,
      "author_name": "xxbxyae",
      "author_url": "",
      "post_date": "04/16/2019 02:08:43",
      "content": "<p>Thanks Oliver. I've tried 3 objectives; regression, gamma, and huber.</p>\n\n<p>| Obj | 10 fold CV  | Public LB |\n| --- | --- | --- |\n| regression | 2.180 | 1.621 |\n| gamma | 2.136 | 1.518 |\n| huber | 2.182 |1.560|</p>\n\n<p>In my environment, the best model is \"gamma\" objective, and there are little difference between \"regression\" and \"huber\" on local CV.\nHyper parameter is based on this kernel: <a href=\"https://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction\">https://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 517457,
      "author_name": "marcuslin",
      "author_url": "",
      "post_date": "04/16/2019 03:39:16",
      "content": "<p>change regression_l1 to gamma,  LB 1.527 to 1.486   </p>",
      "votes": null,
      "replies": [
        {
          "id": 517474,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "04/16/2019 04:20:47",
          "content": "<p>May I ask what CV are you using (shuffle k-fold or leave-1-quake-out?), and does your 1.486 sub have better CV? In my case gamma reg is worse in both CV and LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 517508,
          "author_name": "marcuslin",
          "author_url": "",
          "post_date": "04/16/2019 05:52:28",
          "content": "<p>I used shuffle k-fold, and I got a worse cv  from 2.015 to 2.037</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 517502,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "04/16/2019 05:38:56",
      "content": "<p>Can you add a column with the subs mean predicted ttf? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 517572,
      "author_name": "alexanderkireev",
      "author_url": "",
      "post_date": "04/16/2019 08:03:41",
      "content": "<p>Olivier, congratulate you with Discussion GM Title. Thank you for interest insights which help us a lot on the way. Good luck in new journeys )</p>",
      "votes": null,
      "replies": [
        {
          "id": 517624,
          "author_name": "ogrellier",
          "author_url": "",
          "post_date": "04/16/2019 09:26:52",
          "content": "<p>Thanks <a href=\"/alexanderkireev\">@alexanderkireev</a> for your kind words ! Good luck with the competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 517954,
      "author_name": "arashnic",
      "author_url": "",
      "post_date": "04/16/2019 17:36:34",
      "content": "<p>Thanks @Olivier. I applied Kolmogorov-Smirnov test and filtered features . Shuffled  huber : CV = 2.025  , LB=1.516</p>",
      "votes": null,
      "replies": [
        {
          "id": 517962,
          "author_name": "glimmung",
          "author_url": "",
          "post_date": "04/16/2019 17:45:46",
          "content": "<p>Cramér-von Mises?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 517993,
          "author_name": "arashnic",
          "author_url": "",
          "post_date": "04/16/2019 18:32:14",
          "content": "<p>2 Sample Kolmogorov-Smirnov :  *scipy.stats.ks_2samp*</p>\n\n<p>for whom wants to know more:\n<a href=\"https://stats.stackexchange.com/questions/201434/2-sample-kolmogorov-smirnov-vs-anderson-darling-vs-cramer-von-mises\">https://stats.stackexchange.com/questions/201434/2-sample-kolmogorov-smirnov-vs-anderson-darling-vs-cramer-von-mises</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 518440,
      "author_name": "senkin13",
      "author_url": "",
      "post_date": "04/17/2019 07:50:12",
      "content": "<p>in my experience, trying kinds of objectives always bring surprise :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 518510,
      "author_name": "ogrellier",
      "author_url": "",
      "post_date": "04/17/2019 11:47:13",
      "content": "<p>There are even stranger things going on using different objectives with shuffling. </p>\n\n<p>I need a few submissions to check out local results but it seems that using MAE and shuffling would lead to the best CV score (by a lot). I assume the model overfits but I need to check that out on the LB (though 13% is quite small any sort of solid conclusions). </p>\n\n<p>Without shuffling Gamma is still best for me.</p>",
      "votes": null,
      "replies": [
        {
          "id": 518759,
          "author_name": "mhviraf",
          "author_url": "",
          "post_date": "04/17/2019 18:42:26",
          "content": "<p>What do you mean by shuffling? does it make sense to shuffle this type of data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 518782,
          "author_name": "ogrellier",
          "author_url": "",
          "post_date": "04/17/2019 19:38:55",
          "content": "<p><a href=\"/mhviraf\">@mhviraf</a>, with shuffling the model learns from all types of earthquakes at the same time. Without shuffling you're able to see where earthquakes don't generalize. I believe both ways are useful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 518804,
          "author_name": "mhviraf",
          "author_url": "",
          "post_date": "04/17/2019 20:52:39",
          "content": "<p>I am confused about shuffling. Are you talking about shuffling ~4100 training segments we have? or shuffling rows of each of those segments? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 519085,
          "author_name": "ogrellier",
          "author_url": "",
          "post_date": "04/18/2019 10:44:10",
          "content": "<p><a href=\"/mhviraf\">@mhviraf</a>, sorry about any confusion. I'm talking about shuffling the training segments or if you prefer setting the shuffle parameter of KFold to True.</p>\n\n<p>I don't think shuffling the samples in each segment would make any sense and I did not try :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 519244,
          "author_name": "mhviraf",
          "author_url": "",
          "post_date": "04/18/2019 15:44:25",
          "content": "<p>Thanks. That's why I was confused too. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 519628,
          "author_name": "flemeille",
          "author_url": "",
          "post_date": "04/19/2019 10:17:55",
          "content": "<p>Yes setting shuffle=True into Kfold gave me a way better CV, but it performed very badly on the LB. I am trying to explore which is the best splitting option</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 518633,
      "author_name": "rhhridoy",
      "author_url": "",
      "post_date": "04/17/2019 14:43:10",
      "content": "<p>Thanks a lot for sharing it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 519132,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "04/18/2019 12:18:07",
      "content": "<p>In case someone wants to change the objective in xgboost, pasting some objectives implementations below. All of them give similar scores for me, and better than default MSE. All three functions are approximations for MAE. Please tell me if you get interesting results with it :)</p>\n\n<p><code>\n1. sqrt(1 + (x/h)**2)\n2. ln(cosh(x))\n3. log(exp(-x) + exp(x)) \n</code>\nimplementations\n```\ndef huberobj(preds, dtrain):\n    d = preds - dtrain.get_label()\n    h = 1\n    scale = 1 + (d / h) ** 2\n    scale_sqrt = np.sqrt(scale)\n    grad = d / scale_sqrt / h\n    hess = 1 / scale / scale_sqrt / h\n    return grad, hess</p>\n\n<p>def logcoshobj(preds, dtrain):\n    labels = dtrain.get_label()\n    grad = np.tanh(preds - labels)\n    hess = 1.0 - grad*grad\n    return grad, hess</p>\n\n<p>def logexpexp(preds, dtrain):\n    labels = dtrain.get_label()\n    x = preds - labels\n    grad = (np.exp(2.0*x) - 1) / (np.exp(2.0*x) + 1)\n    hess = (4.0*np.exp(2.0*x)) / (np.exp(2.0*x) + 1)**2 \n    return grad, hess</p>\n\n<p>model = xgb.train(...,  obj=logcoshobj)</p>\n\n<p>```</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 519586,
      "author_name": "anonemaus",
      "author_url": "",
      "post_date": "04/19/2019 08:33:40",
      "content": "<p>ah, distributions matter normally</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 522456,
      "author_name": "arpitr07",
      "author_url": "",
      "post_date": "04/24/2019 13:00:27",
      "content": "<p>For those using XGBoost, the improvement using a 5-fold CV with following parameters:\nparams = {\n    \"learning_rate\": 0.001,\n    \"max_depth\": 3,\n    \"n_estimators\": 10000,\n    \"min_child_weight\": 4,\n    \"colsample_bytree\": 1,\n    \"subsample\": 0.8,\n    \"nthread\": 12,\n    \"random_state\": 42,\n}, \nthe improvement recorded up on using reg:gamma as objective is minimal. \n<strong>|  objective    | 5-fold CV score |\n|  reg:linear   |       2.027358      |\n| reg:gamma |       2.026449     |</strong></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "516911": "I'd like to thank @inversion for his [benchmark kernel](https://www.kaggle.com/inversion/basic-feature-benchmark) that rang a bell in my empty brain (it's till echoing though ...)\n\nFor those of you using LightGBM, here are the results of a simple experiment:\n\n| objective | 10-fold CV | Public LB\n| --- | --- | --- |\n| regression | 2.1675 | 1.589 | \n| huber | 2.1597 | not tested | \n| fair | 2.1463 | not tested | \n| gamma | 2.1208 |  1.548 | \n\nI'm really far down the LB so it may not work for everyone but it's simple enough to test.\n\nIn any case it's very interesting to blend predictions from models with different objectives :)\n\nGood luck.",
    "516971": "Nice!",
    "516981": "Thanks @ogrellier !",
    "517015": "Thanks Olivier. I tried, but my CV with \"regression\" is still the best, so is my LB. Can anyone else please check this?",
    "517025": "I've been using lightgbm with objective=\"mae\".",
    "517028": "Thanks @dslate, totally forgot to test MAE. \n\nMy CV score with objective \"mae\" is still higher than with gamma regression but lower than MSE.",
    "517029": "Thanks for your feedback @khahuras, I suspect other hyper parameters play a role here.",
    "517141": "Thanks Olivier 👍",
    "517178": "That is a great observation! I did the experiment as well with 5-fold CV and shuffling\n\n| objective | 5-fold CV (mean) | 5-fold CV (std) | Public LB |\n| ----------- | -------------------- | ------------------ | ------------| \n| regression                   |2.0613         | 0.0705      | 1.557 |\n| huber        |**2.0232**  | 0.0765       | 1.521 |\n| fair            |2.0336          | 0.0750       | **1.503** |\n| gamma     |2.0276          | **0.0693** | 1.508 |\n| mae           |2.0283          | 0.0807       | 1.535 |\n\n*Huber* seemed to give me the best CV mean score and *gamma* the best std. However, the difference in the std seem very little to me and just just be chance. My best LB score with LightGBM  is *fair*.",
    "517416": "Thanks Oliver. I've tried 3 objectives; regression, gamma, and huber.\n\n| Obj | 10 fold CV  | Public LB |\n| --- | --- | --- |\n| regression | 2.180 | 1.621 |\n| gamma | 2.136 | 1.518 |\n| huber | 2.182 |1.560|\n\nIn my environment, the best model is \"gamma\" objective, and there are little difference between \"regression\" and \"huber\" on local CV.\nHyper parameter is based on this kernel: https://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction",
    "517457": "change regression_l1 to gamma,  LB 1.527 to 1.486",
    "517474": "May I ask what CV are you using (shuffle k-fold or leave-1-quake-out?), and does your 1.486 sub have better CV? In my case gamma reg is worse in both CV and LB.",
    "517502": "Can you add a column with the subs mean predicted ttf?",
    "517507": "ricarddelgado, when I searched for the right objective I initially thought Huber would be best and this is what you find. However the target seems to have a gamma distribution...",
    "517508": "I used shuffle k-fold, and I got a worse cv  from 2.015 to 2.037",
    "517529": "I'll do some hyperparameter optimization to see how far I can go with the others :)",
    "517531": "The countdown? \n\nOr the duration of inter-quake intervals?",
    "517572": "Olivier, congratulate you with Discussion GM Title. Thank you for interest insights which help us a lot on the way. Good luck in new journeys )",
    "517586": "Here, the performance is on the ttf, so, the countdown.",
    "517591": "I tested both mae and rmse to train the model, rmse gives me a better cv than mae (from 1.9 to 1.8 mae) but not submitted yet",
    "517624": "Thanks @alexanderkireev for your kind words ! Good luck with the competition.",
    "517954": "Thanks @Olivier. I applied Kolmogorov-Smirnov test and filtered features . Shuffled  huber : CV = 2.025  , LB=1.516",
    "517962": "Cramér-von Mises?",
    "517993": "2 Sample Kolmogorov-Smirnov :  *scipy.stats.ks_2samp*\n\nfor whom wants to know more:\nhttps://stats.stackexchange.com/questions/201434/2-sample-kolmogorov-smirnov-vs-anderson-darling-vs-cramer-von-mises",
    "518440": "in my experience, trying kinds of objectives always bring surprise :)",
    "518510": "There are even stranger things going on using different objectives with shuffling. \n\nI need a few submissions to check out local results but it seems that using MAE and shuffling would lead to the best CV score (by a lot). I assume the model overfits but I need to check that out on the LB (though 13% is quite small any sort of solid conclusions). \n\nWithout shuffling Gamma is still best for me.",
    "518633": "Thanks a lot for sharing it.",
    "518759": "What do you mean by shuffling? does it make sense to shuffle this type of data?",
    "518782": "mhviraf, with shuffling the model learns from all types of earthquakes at the same time. Without shuffling you're able to see where earthquakes don't generalize. I believe both ways are useful.",
    "518804": "I am confused about shuffling. Are you talking about shuffling ~4100 training segments we have? or shuffling rows of each of those segments?",
    "519085": "mhviraf, sorry about any confusion. I'm talking about shuffling the training segments or if you prefer setting the shuffle parameter of KFold to True.\n\nI don't think shuffling the samples in each segment would make any sense and I did not try :)",
    "519132": "In case someone wants to change the objective in xgboost, pasting some objectives implementations below. All of them give similar scores for me, and better than default MSE. All three functions are approximations for MAE. Please tell me if you get interesting results with it :)\n\n```\n1. sqrt(1 + (x/h)**2)\n2. ln(cosh(x))\n3. log(exp(-x) + exp(x)) \n```\nimplementations\n```\ndef huberobj(preds, dtrain):\n    d = preds - dtrain.get_label()\n    h = 1\n    scale = 1 + (d / h) ** 2\n    scale_sqrt = np.sqrt(scale)\n    grad = d / scale_sqrt / h\n    hess = 1 / scale / scale_sqrt / h\n    return grad, hess\n\ndef logcoshobj(preds, dtrain):\n    labels = dtrain.get_label()\n    grad = np.tanh(preds - labels)\n    hess = 1.0 - grad*grad\n    return grad, hess\n\ndef logexpexp(preds, dtrain):\n    labels = dtrain.get_label()\n    x = preds - labels\n    grad = (np.exp(2.0*x) - 1) / (np.exp(2.0*x) + 1)\n    hess = (4.0*np.exp(2.0*x)) / (np.exp(2.0*x) + 1)**2 \n    return grad, hess\n\nmodel = xgb.train(...,  obj=logcoshobj)\n\n```",
    "519237": "I tried again huber, mae, regression, gamma with different n_folds and different hyper-parameters  . but didn't change CV strategy yet (e,g, cv based on quake id).  I think only features to use and better cv strategy will make difference and improve the result among different lgb models. what do you think ?",
    "519244": "Thanks. That's why I was confused too.",
    "519586": "ah, distributions matter normally",
    "519628": "Yes setting shuffle=True into Kfold gave me a way better CV, but it performed very badly on the LB. I am trying to explore which is the best splitting option",
    "522456": "For those using XGBoost, the improvement using a 5-fold CV with following parameters:\nparams = {\n    \"learning_rate\": 0.001,\n    \"max_depth\": 3,\n    \"n_estimators\": 10000,\n    \"min_child_weight\": 4,\n    \"colsample_bytree\": 1,\n    \"subsample\": 0.8,\n    \"nthread\": 12,\n    \"random_state\": 42,\n}, \nthe improvement recorded up on using reg:gamma as objective is minimal. \n**|  objective    | 5-fold CV score |\n|  reg:linear   |       2.027358      |\n| reg:gamma |       2.026449     |**"
  },
  "source": "meta"
}