{
  "id": 91500,
  "title": "Quake-wise early stopping - leak explained",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/91500",
  "author_name": "",
  "post_date": "2019-05-05T18:10:56.040968200Z",
  "votes": 9,
  "comment_count": 16,
  "views": 0,
  "content": "<p>The discussion about using leave-one-quake-out cross validation and early stopping started in <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125#latest-527210\">this thread</a> when <a href=\"/cpmpml\">@cpmpml</a> attached his oof plot. I was intrigued that he is using early stopping in quake-wise CV and getting good results. At least when I tried it, my LB score was &gt; 1.6 due to pre-mature stopping resulting from leakage. Others like <a href=\"/philippsinger\">@philippsinger</a> shared my frustration. Even <a href=\"/cpmpml\">@cpmpml</a> started doubting himself and started <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91328#latest-527457\">this thread about early stopping</a> where many had very good arguments about leakage, but <a href=\"/cpmpml\">@cpmpml</a> almost could swear his model is not leaking.</p>\n\n<p>So I investigated the issue and figured out the reason for this discrepancy. Those who are suffering from leak when using this strategy are probably using catboost (at least me). Others are using LGB. \nIt turned out that catboost initialize the boosting with zeros. In small quakes, the error in the validation set reaches a minimum very fast before convergence because the values start increasing from zero. However, LGB has a parameter <code>boost_from_average</code>, which is true by default. Therefore, LGB doesn't have the same problem.</p>\n\n<p>To test, set <code>boost_from_average</code> to false, and the CV score will drop ~0.2 due to leakage. I couldn't find a parameter in catboost that controls this behavior.</p>\n\n<p>So to conclude, if you want to use quake-wise CV, switch to LGB. That doesn't exclude the possibility of leakage, but at least it won't be as severe as we have experienced. \nI don't know yet how that will affect my LB score, but fingers crossed. </p>\n\n<p>Update: after trying quake-wise CV with early stopping in LGB (boostfromaverage = true), the leak is still there. Not as it was with catboost, but it is still severely under/over-fitting. \nSo I am sticking with shuffling, at least my models stop optimally in this setting.</p>",
  "messages": [
    {
      "id": "527521",
      "postDate": "05/05/2019 18:10:56",
      "content": "<p>The discussion about using leave-one-quake-out cross validation and early stopping started in <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125#latest-527210\">this thread</a> when <a href=\"/cpmpml\">@cpmpml</a> attached his oof plot. I was intrigued that he is using early stopping in quake-wise CV and getting good results. At least when I tried it, my LB score was &gt; 1.6 due to pre-mature stopping resulting from leakage. Others like <a href=\"/philippsinger\">@philippsinger</a> shared my frustration. Even <a href=\"/cpmpml\">@cpmpml</a> started doubting himself and started <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91328#latest-527457\">this thread about early stopping</a> where many had very good arguments about leakage, but <a href=\"/cpmpml\">@cpmpml</a> almost could swear his model is not leaking.</p>\n\n<p>So I investigated the issue and figured out the reason for this discrepancy. Those who are suffering from leak when using this strategy are probably using catboost (at least me). Others are using LGB. \nIt turned out that catboost initialize the boosting with zeros. In small quakes, the error in the validation set reaches a minimum very fast before convergence because the values start increasing from zero. However, LGB has a parameter <code>boost_from_average</code>, which is true by default. Therefore, LGB doesn't have the same problem.</p>\n\n<p>To test, set <code>boost_from_average</code> to false, and the CV score will drop ~0.2 due to leakage. I couldn't find a parameter in catboost that controls this behavior.</p>\n\n<p>So to conclude, if you want to use quake-wise CV, switch to LGB. That doesn't exclude the possibility of leakage, but at least it won't be as severe as we have experienced. \nI don't know yet how that will affect my LB score, but fingers crossed. </p>\n\n<p>Update: after trying quake-wise CV with early stopping in LGB (boostfromaverage = true), the leak is still there. Not as it was with catboost, but it is still severely under/over-fitting. \nSo I am sticking with shuffling, at least my models stop optimally in this setting.</p>",
      "rawMarkdown": "The discussion about using leave-one-quake-out cross validation and early stopping started in [this thread](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125#latest-527210) when @cpmpml attached his oof plot. I was intrigued that he is using early stopping in quake-wise CV and getting good results. At least when I tried it, my LB score was &gt; 1.6 due to pre-mature stopping resulting from leakage. Others like @philippsinger shared my frustration. Even @cpmpml started doubting himself and started [this thread about early stopping](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91328#latest-527457) where many had very good arguments about leakage, but @cpmpml almost could swear his model is not leaking.\n\nSo I investigated the issue and figured out the reason for this discrepancy. Those who are suffering from leak when using this strategy are probably using catboost (at least me). Others are using LGB. \nIt turned out that catboost initialize the boosting with zeros. In small quakes, the error in the validation set reaches a minimum very fast before convergence because the values start increasing from zero. However, LGB has a parameter `boost_from_average`, which is true by default. Therefore, LGB doesn't have the same problem.\n\nTo test, set `boost_from_average` to false, and the CV score will drop ~0.2 due to leakage. I couldn't find a parameter in catboost that controls this behavior.\n\nSo to conclude, if you want to use quake-wise CV, switch to LGB. That doesn't exclude the possibility of leakage, but at least it won't be as severe as we have experienced. \nI don't know yet how that will affect my LB score, but fingers crossed. \n\nUpdate: after trying quake-wise CV with early stopping in LGB (boostfromaverage = true), the leak is still there. Not as it was with catboost, but it is still severely under/over-fitting. \nSo I am sticking with shuffling, at least my models stop optimally in this setting.",
      "votes": null
    },
    {
      "id": "527522",
      "postDate": "05/05/2019 18:16:31",
      "content": "<blockquote>\n  <p>he is using early stopping in quake-wise CV </p>\n</blockquote>\n\n<p>I am not doing that.  I even said it explicitly in the discussion you cite ;)</p>",
      "rawMarkdown": "&gt; he is using early stopping in quake-wise CV \n\nI am not doing that.  I even said it explicitly in the discussion you cite ;)",
      "votes": null
    },
    {
      "id": "527525",
      "postDate": "05/05/2019 18:21:17",
      "content": "<p>Sorry for misunderstanding then.  I guess this was only my assumption because I could only reproduce a similar pattern to your oof with early stopping.</p>",
      "rawMarkdown": "Sorry for misunderstanding then.  I guess this was only my assumption because I could only reproduce a similar pattern to your oof with early stopping.",
      "votes": null
    },
    {
      "id": "527531",
      "postDate": "05/05/2019 18:28:48",
      "content": "<p>What is right in your post is that lgb has many parameters, and some can have a significant effect like the one you single out.  </p>",
      "rawMarkdown": "What is right in your post is that lgb has many parameters, and some can have a significant effect like the one you single out.",
      "votes": null
    },
    {
      "id": "527533",
      "postDate": "05/05/2019 18:37:54",
      "content": "<p>Everything is right about my post, except my assumption about your secret fancy strategy ;)</p>",
      "rawMarkdown": "Everything is right about my post, except my assumption about your secret fancy strategy ;)",
      "votes": null
    },
    {
      "id": "527568",
      "postDate": "05/05/2019 20:40:30",
      "content": "<p>Very interesting. Thanks for posting this insight. Is it safe to say that XGBoost would exhibit the same behavior as catboost? Quick google search found xgboost doesn't have <code>boost_from_average</code> implemented yet but is being discussed on their github: <a href=\"https://github.com/dmlc/xgboost/issues/4321\">https://github.com/dmlc/xgboost/issues/4321</a></p>",
      "rawMarkdown": "Very interesting. Thanks for posting this insight. Is it safe to say that XGBoost would exhibit the same behavior as catboost? Quick google search found xgboost doesn't have `boost_from_average` implemented yet but is being discussed on their github: https://github.com/dmlc/xgboost/issues/4321",
      "votes": null
    },
    {
      "id": "527572",
      "postDate": "05/05/2019 20:58:14",
      "content": "<p>I haven't tried XGBoost in this competition yet. I'll post an update when I try it.</p>",
      "rawMarkdown": "I haven't tried XGBoost in this competition yet. I'll post an update when I try it.",
      "votes": null
    },
    {
      "id": "527583",
      "postDate": "05/05/2019 21:46:30",
      "content": "<p>You can use base_score parameter to starts from whatever value you want with XGBoost.  Setting it to target average gives you the behavior you discuss.</p>",
      "rawMarkdown": "You can use base_score parameter to starts from whatever value you want with XGBoost.  Setting it to target average gives you the behavior you discuss.",
      "votes": null
    },
    {
      "id": "527584",
      "postDate": "05/05/2019 21:49:19",
      "content": "<p>It's all about the mean... What is funny BTW is that <code>boost_from_average</code> behaves differently between l1 and gamma loss.</p>",
      "rawMarkdown": "It's all about the mean... What is funny BTW is that `boost_from_average` behaves differently between l1 and gamma loss.",
      "votes": null
    },
    {
      "id": "528240",
      "postDate": "05/07/2019 10:32:12",
      "content": "<p>Ehm... help? Setting <code>'boost_from_average': 'false',</code> in the parameters causes the best iteration for each fold to be iteration 1, and overall CV jumps to 5.679</p>",
      "rawMarkdown": "Ehm... help? Setting `'boost_from_average': 'false',` in the parameters causes the best iteration for each fold to be iteration 1, and overall CV jumps to 5.679",
      "votes": null
    },
    {
      "id": "528247",
      "postDate": "05/07/2019 10:44:45",
      "content": "<p>That's odd. It will stop prematurely but not like that.</p>",
      "rawMarkdown": "That's odd. It will stop prematurely but not like that.",
      "votes": null
    },
    {
      "id": "528249",
      "postDate": "05/07/2019 10:50:34",
      "content": "<p>Update: after trying quake-wise CV with early stopping in LGB (boost_from_average = true), the leak is still there. Not as it was with catboost, but it is still severely under/over-fitting. \n So I am sticking with shuffling, at least my models stop optimally in this setting.</p>",
      "rawMarkdown": "Update: after trying quake-wise CV with early stopping in LGB (boost_from_average = true), the leak is still there. Not as it was with catboost, but it is still severely under/over-fitting. \n So I am sticking with shuffling, at least my models stop optimally in this setting.",
      "votes": null
    },
    {
      "id": "528252",
      "postDate": "05/07/2019 11:08:19",
      "content": "<p>I have the same with objective mae.</p>",
      "rawMarkdown": "I have the same with objective mae.",
      "votes": null
    },
    {
      "id": "528282",
      "postDate": "05/07/2019 12:20:15",
      "content": "<p>Well, how do we avoid this error? I am also using <code>objective=mean_absolute_error</code>. Are you using <code>gamma</code> loss instead of <code>mae</code> ? Or maybe, the parameters need to be retuned.</p>",
      "rawMarkdown": "Well, how do we avoid this error? I am also using `objective=mean_absolute_error`. Are you using `gamma` loss instead of `mae` ? Or maybe, the parameters need to be retuned.",
      "votes": null
    },
    {
      "id": "528287",
      "postDate": "05/07/2019 12:27:13",
      "content": "<p>&gt; Well, how do we avoid this error? </p>\n\n<p>Don't use mae with <code>boost_from_average</code> set to False ;)</p>\n\n<p>I'll say more after competition end.</p>",
      "rawMarkdown": "&gt; Well, how do we avoid this error? \n\nDon't use mae with `boost_from_average` set to False ;)\n\nI'll say more after competition end.",
      "votes": null
    },
    {
      "id": "530197",
      "postDate": "05/12/2019 04:26:31",
      "content": "<p>I just tried xgboost with similar results as catboost (Much better CV score than LGB but an abysmal LB score). I need to look more into this <code>base_score</code> parameter that <a href=\"/cpmpml\">@cpmpml</a> speaks of....</p>",
      "rawMarkdown": "I just tried xgboost with similar results as catboost (Much better CV score than LGB but an abysmal LB score). I need to look more into this `base_score` parameter that @cpmpml speaks of....",
      "votes": null
    },
    {
      "id": "530271",
      "postDate": "05/12/2019 10:36:56",
      "content": "<p>I have started with xgboost and so far my CV is higher than with lgb.  I haven't submitted yet.</p>",
      "rawMarkdown": "I have started with xgboost and so far my CV is higher than with lgb.  I haven't submitted yet.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 527522,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/05/2019 18:16:31",
      "content": "<blockquote>\n  <p>he is using early stopping in quake-wise CV </p>\n</blockquote>\n\n<p>I am not doing that.  I even said it explicitly in the discussion you cite ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 527525,
          "author_name": "amjad85",
          "author_url": "",
          "post_date": "05/05/2019 18:21:17",
          "content": "<p>Sorry for misunderstanding then.  I guess this was only my assumption because I could only reproduce a similar pattern to your oof with early stopping.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 527531,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/05/2019 18:28:48",
          "content": "<p>What is right in your post is that lgb has many parameters, and some can have a significant effect like the one you single out.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 527533,
          "author_name": "amjad85",
          "author_url": "",
          "post_date": "05/05/2019 18:37:54",
          "content": "<p>Everything is right about my post, except my assumption about your secret fancy strategy ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 527568,
      "author_name": "robikscube",
      "author_url": "",
      "post_date": "05/05/2019 20:40:30",
      "content": "<p>Very interesting. Thanks for posting this insight. Is it safe to say that XGBoost would exhibit the same behavior as catboost? Quick google search found xgboost doesn't have <code>boost_from_average</code> implemented yet but is being discussed on their github: <a href=\"https://github.com/dmlc/xgboost/issues/4321\">https://github.com/dmlc/xgboost/issues/4321</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 527572,
          "author_name": "amjad85",
          "author_url": "",
          "post_date": "05/05/2019 20:58:14",
          "content": "<p>I haven't tried XGBoost in this competition yet. I'll post an update when I try it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 527583,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/05/2019 21:46:30",
          "content": "<p>You can use base_score parameter to starts from whatever value you want with XGBoost.  Setting it to target average gives you the behavior you discuss.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 530197,
          "author_name": "robikscube",
          "author_url": "",
          "post_date": "05/12/2019 04:26:31",
          "content": "<p>I just tried xgboost with similar results as catboost (Much better CV score than LGB but an abysmal LB score). I need to look more into this <code>base_score</code> parameter that <a href=\"/cpmpml\">@cpmpml</a> speaks of....</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 530271,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/12/2019 10:36:56",
          "content": "<p>I have started with xgboost and so far my CV is higher than with lgb.  I haven't submitted yet.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 527584,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "05/05/2019 21:49:19",
      "content": "<p>It's all about the mean... What is funny BTW is that <code>boost_from_average</code> behaves differently between l1 and gamma loss.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 528240,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "05/07/2019 10:32:12",
      "content": "<p>Ehm... help? Setting <code>'boost_from_average': 'false',</code> in the parameters causes the best iteration for each fold to be iteration 1, and overall CV jumps to 5.679</p>",
      "votes": null,
      "replies": [
        {
          "id": 528247,
          "author_name": "amjad85",
          "author_url": "",
          "post_date": "05/07/2019 10:44:45",
          "content": "<p>That's odd. It will stop prematurely but not like that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 528252,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/07/2019 11:08:19",
          "content": "<p>I have the same with objective mae.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 528282,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "05/07/2019 12:20:15",
          "content": "<p>Well, how do we avoid this error? I am also using <code>objective=mean_absolute_error</code>. Are you using <code>gamma</code> loss instead of <code>mae</code> ? Or maybe, the parameters need to be retuned.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 528287,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/07/2019 12:27:13",
          "content": "<p>&gt; Well, how do we avoid this error? </p>\n\n<p>Don't use mae with <code>boost_from_average</code> set to False ;)</p>\n\n<p>I'll say more after competition end.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 528249,
      "author_name": "amjad85",
      "author_url": "",
      "post_date": "05/07/2019 10:50:34",
      "content": "<p>Update: after trying quake-wise CV with early stopping in LGB (boost_from_average = true), the leak is still there. Not as it was with catboost, but it is still severely under/over-fitting. \n So I am sticking with shuffling, at least my models stop optimally in this setting.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "527521": "The discussion about using leave-one-quake-out cross validation and early stopping started in [this thread](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125#latest-527210) when @cpmpml attached his oof plot. I was intrigued that he is using early stopping in quake-wise CV and getting good results. At least when I tried it, my LB score was &gt; 1.6 due to pre-mature stopping resulting from leakage. Others like @philippsinger shared my frustration. Even @cpmpml started doubting himself and started [this thread about early stopping](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91328#latest-527457) where many had very good arguments about leakage, but @cpmpml almost could swear his model is not leaking.\n\nSo I investigated the issue and figured out the reason for this discrepancy. Those who are suffering from leak when using this strategy are probably using catboost (at least me). Others are using LGB. \nIt turned out that catboost initialize the boosting with zeros. In small quakes, the error in the validation set reaches a minimum very fast before convergence because the values start increasing from zero. However, LGB has a parameter `boost_from_average`, which is true by default. Therefore, LGB doesn't have the same problem.\n\nTo test, set `boost_from_average` to false, and the CV score will drop ~0.2 due to leakage. I couldn't find a parameter in catboost that controls this behavior.\n\nSo to conclude, if you want to use quake-wise CV, switch to LGB. That doesn't exclude the possibility of leakage, but at least it won't be as severe as we have experienced. \nI don't know yet how that will affect my LB score, but fingers crossed. \n\nUpdate: after trying quake-wise CV with early stopping in LGB (boostfromaverage = true), the leak is still there. Not as it was with catboost, but it is still severely under/over-fitting. \nSo I am sticking with shuffling, at least my models stop optimally in this setting.",
    "527522": "&gt; he is using early stopping in quake-wise CV \n\nI am not doing that.  I even said it explicitly in the discussion you cite ;)",
    "527525": "Sorry for misunderstanding then.  I guess this was only my assumption because I could only reproduce a similar pattern to your oof with early stopping.",
    "527531": "What is right in your post is that lgb has many parameters, and some can have a significant effect like the one you single out.",
    "527533": "Everything is right about my post, except my assumption about your secret fancy strategy ;)",
    "527568": "Very interesting. Thanks for posting this insight. Is it safe to say that XGBoost would exhibit the same behavior as catboost? Quick google search found xgboost doesn't have `boost_from_average` implemented yet but is being discussed on their github: https://github.com/dmlc/xgboost/issues/4321",
    "527572": "I haven't tried XGBoost in this competition yet. I'll post an update when I try it.",
    "527583": "You can use base_score parameter to starts from whatever value you want with XGBoost.  Setting it to target average gives you the behavior you discuss.",
    "527584": "It's all about the mean... What is funny BTW is that `boost_from_average` behaves differently between l1 and gamma loss.",
    "528240": "Ehm... help? Setting `'boost_from_average': 'false',` in the parameters causes the best iteration for each fold to be iteration 1, and overall CV jumps to 5.679",
    "528247": "That's odd. It will stop prematurely but not like that.",
    "528249": "Update: after trying quake-wise CV with early stopping in LGB (boost_from_average = true), the leak is still there. Not as it was with catboost, but it is still severely under/over-fitting. \n So I am sticking with shuffling, at least my models stop optimally in this setting.",
    "528252": "I have the same with objective mae.",
    "528282": "Well, how do we avoid this error? I am also using `objective=mean_absolute_error`. Are you using `gamma` loss instead of `mae` ? Or maybe, the parameters need to be retuned.",
    "528287": "&gt; Well, how do we avoid this error? \n\nDon't use mae with `boost_from_average` set to False ;)\n\nI'll say more after competition end.",
    "530197": "I just tried xgboost with similar results as catboost (Much better CV score than LGB but an abysmal LB score). I need to look more into this `base_score` parameter that @cpmpml speaks of....",
    "530271": "I have started with xgboost and so far my CV is higher than with lgb.  I haven't submitted yet."
  },
  "source": "meta"
}