{
  "id": 337609,
  "title": "Which score should we trust more, CV or LB?",
  "url": "/competitions/amex-default-prediction/discussion/337609",
  "author_name": "Jianhong Zhu",
  "post_date": "2022-07-16T19:40:34.724000",
  "votes": 18,
  "comment_count": 17,
  "views": 0,
  "content": "<p>I have done some experiments on my LGBM model, and I found that even my CV score got improved but the LB score got worse. Does anyone have the same issue? Which score should we trust more, CV or LB? Thank you!<br>\nHere's my score:<br>\nCV 0.7964 ---- LB 0.799<br>\nCV 0.7972 ---- LB 0.798<br>\nCV 0.7976 ---- LB 0.798</p>",
  "messages": [
    {
      "id": 1858410,
      "postDate": "2022-07-16T22:09:16.697Z",
      "content": "<p>In the end, you have 2 final submissions.</p>\n<p>My strategy has always been: first sub - best CV model. second sub - best CV model/2 + best LB model/2.</p>\n<p>Looking back at my previous competitions, 8 times best CV was best on private LB, 2 times the average was better. And only once best public LB model was best on private.</p>",
      "rawMarkdown": "In the end, you have 2 final submissions.\n\nMy strategy has always been: first sub - best CV model. second sub - best CV model/2 + best LB model/2.\n\nLooking back at my previous competitions, 8 times best CV was best on private LB, 2 times the average was better. And only once best public LB model was best on private.\n",
      "votes": 23,
      "replies": [
        {
          "id": 1858476,
          "postDate": "2022-07-17T00:25:06.397Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> ! So if my CV gets better and the LB get worse simultaneously, it does not necessarily mean that the model is overfitting?</p>",
          "rawMarkdown": "Thanks @raddar ! So if my CV gets better and the LB get worse simultaneously, it does not necessarily mean that the model is overfitting?",
          "votes": 1
        },
        {
          "id": 1858698,
          "postDate": "2022-07-17T06:23:27.027Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> for sharing your submission strategy which I've never tried! Hope there won't be a big shake up this time 😃</p>",
          "rawMarkdown": "Thank you @raddar for sharing your submission strategy which I've never tried! Hope there won't be a big shake up this time 😃"
        },
        {
          "id": 1858801,
          "postDate": "2022-07-17T08:07:40.377Z",
          "content": "<p><a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a> that is the very definition of public LB overfitting. If you are only looking at things that work on public LB, you are 99% to be overfitting on LB. </p>\n<p>CV increase with public LB decrease should be expected ~ half of the time if CV increases are only slightly incremental.</p>",
          "rawMarkdown": "@mohammadrahmati that is the very definition of public LB overfitting. If you are only looking at things that work on public LB, you are 99% to be overfitting on LB. \n\nCV increase with public LB decrease should be expected ~ half of the time if CV increases are only slightly incremental.\n",
          "votes": 4
        },
        {
          "id": 1858803,
          "postDate": "2022-07-17T08:10:16.713Z",
          "content": "<p>And in general, you should avoid using term \"model is overfitting\". Models are overfitting to train set all the time, we still use them. It's the overfitting on LB that matters, and this does not have to do much with models, but more about what decisions we make :)</p>",
          "rawMarkdown": "And in general, you should avoid using term \"model is overfitting\". Models are overfitting to train set all the time, we still use them. It's the overfitting on LB that matters, and this does not have to do much with models, but more about what decisions we make :)",
          "votes": 10
        },
        {
          "id": 1858989,
          "postDate": "2022-07-17T10:24:22.033Z",
          "content": "<p>thnk you <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> ! I appreciate your help 🙏</p>",
          "rawMarkdown": "thnk you @raddar ! I appreciate your help 🙏"
        }
      ]
    },
    {
      "id": 1858296,
      "postDate": "2022-07-16T19:40:34.723Z",
      "content": "<p>I have done some experiments on my LGBM model, and I found that even my CV score got improved but the LB score got worse. Does anyone have the same issue? Which score should we trust more, CV or LB? Thank you!<br>\nHere's my score:<br>\nCV 0.7964 ---- LB 0.799<br>\nCV 0.7972 ---- LB 0.798<br>\nCV 0.7976 ---- LB 0.798</p>",
      "rawMarkdown": "I have done some experiments on my LGBM model, and I found that even my CV score got improved but the LB score got worse. Does anyone have the same issue? Which score should we trust more, CV or LB? Thank you!\nHere's my score:\nCV 0.7964 ---- LB 0.799\nCV 0.7972 ---- LB 0.798\nCV 0.7976 ---- LB 0.798\n",
      "votes": 18
    },
    {
      "id": 1858406,
      "postDate": "2022-07-16T22:05:25.187Z",
      "content": "<p>So far my CV/LB has been almost, if not exactly directionally consistent (with LB a bit under .001 or so higher, though of course it's hard to tell with the 4th decimal missing -- my current CV is fold avg ~.7992). I believe others have reported a similar relationship. It's possible we've been lucky with our CV fold setups, but it's worth double-checking how you're calculating your CV and other factors that may be at play. Is your CV score for an ensemble where there might be some subtle leakage? Did you engineer features whose distributions are different between train and test? Those might look great on CV but be mismatched to test.   </p>\n<p>Beyond the point about public LB size made by Tilii, another advantage of the public LB is that the data is later in time than CV and closer to the private LB time. I don't expect huge data drift caused by time for this problem, but even a little (inevitable?) drift could be tricky to deal with. Where I land is wanting to see everything improve to be confident -- I only make submissions when CV improves, and an LB decline would make me hesitate. Also, the classic best CV + best public LB final submission selection seems like a decent hedge, as often.  </p>",
      "rawMarkdown": "So far my CV/LB has been almost, if not exactly directionally consistent (with LB a bit under .001 or so higher, though of course it's hard to tell with the 4th decimal missing -- my current CV is fold avg ~.7992). I believe others have reported a similar relationship. It's possible we've been lucky with our CV fold setups, but it's worth double-checking how you're calculating your CV and other factors that may be at play. Is your CV score for an ensemble where there might be some subtle leakage? Did you engineer features whose distributions are different between train and test? Those might look great on CV but be mismatched to test.   \n\nBeyond the point about public LB size made by Tilii, another advantage of the public LB is that the data is later in time than CV and closer to the private LB time. I don't expect huge data drift caused by time for this problem, but even a little (inevitable?) drift could be tricky to deal with. Where I land is wanting to see everything improve to be confident -- I only make submissions when CV improves, and an LB decline would make me hesitate. Also, the classic best CV + best public LB final submission selection seems like a decent hedge, as often.  ",
      "votes": 6,
      "replies": [
        {
          "id": 1858693,
          "postDate": "2022-07-17T06:17:33.633Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> for reminding me of double checking my CV and feature engineering! So far I am just using the stratified 5 fold cv that everyone is using. I am thinking if there's a way to improve our cv strategy. Btw, LB seems to be more reliable this time, but who knows.  </p>",
          "rawMarkdown": "Thank you @aquatic for reminding me of double checking my CV and feature engineering! So far I am just using the stratified 5 fold cv that everyone is using. I am thinking if there's a way to improve our cv strategy. Btw, LB seems to be more reliable this time, but who knows.  "
        }
      ]
    },
    {
      "id": 1858380,
      "postDate": "2022-07-16T21:32:53.623Z",
      "content": "<p>My understanding is that a submission is on the right track when a CV frameworks mirrors an LB score, and with every increase in CV the same delta would appear in LB. That is a clear signal that the solution and the CV framework is meaningful. So maybe an answer to your question is that they go hand in hand? Either way its a good experience and every dataset is special in some way or the other! Good luck :)  </p>",
      "rawMarkdown": "My understanding is that a submission is on the right track when a CV frameworks mirrors an LB score, and with every increase in CV the same delta would appear in LB. That is a clear signal that the solution and the CV framework is meaningful. So maybe an answer to your question is that they go hand in hand? Either way its a good experience and every dataset is special in some way or the other! Good luck :)  ",
      "votes": 3
    },
    {
      "id": 1858321,
      "postDate": "2022-07-16T19:59:10.883Z",
      "content": "<p>The same question likely has come up in every single competition. A general answer is that CV scores are more reliable because we have control over data preparation and prediction methods. That said, we don't know whether your models employ internal fold consistency and how exactly they were made.</p>\n<p>It should be mentioned that a public leaderboard that is based on 50% of test data - like in this competition - is more trustworthy than a leaderboard based on 10% of test data. I think for most models the <code>[CV-LB]</code> difference will end up being small even when private leaderboard is released.</p>",
      "rawMarkdown": "The same question likely has come up in every single competition. A general answer is that CV scores are more reliable because we have control over data preparation and prediction methods. That said, we don't know whether your models employ internal fold consistency and how exactly they were made.\n\nIt should be mentioned that a public leaderboard that is based on 50% of test data - like in this competition - is more trustworthy than a leaderboard based on 10% of test data. I think for most models the `[CV-LB]` difference will end up being small even when private leaderboard is released.",
      "votes": 3,
      "replies": [
        {
          "id": 1858680,
          "postDate": "2022-07-17T06:04:33.810Z",
          "content": "<p>LB based on 50% of test data is exactly what I am concerned about, meaning my model is not generalizing well in the test data even CV score is improved. I gotta put more effort on the data and cv strategy than the model itself. Thank you <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> </p>",
          "rawMarkdown": "LB based on 50% of test data is exactly what I am concerned about, meaning my model is not generalizing well in the test data even CV score is improved. I gotta put more effort on the data and cv strategy than the model itself. Thank you @tilii7 "
        }
      ]
    },
    {
      "id": 1864465,
      "postDate": "2022-07-21T04:27:57.407Z",
      "content": "<p>I have trust issues. I don't trust either one. But a bunch of CV scores across different seeds? I start to trust. Since that's impossible on the LB I'd only use it to verify. In other words I only use the LB to go back to NOT trusting, lol. </p>",
      "rawMarkdown": "I have trust issues. I don't trust either one. But a bunch of CV scores across different seeds? I start to trust. Since that's impossible on the LB I'd only use it to verify. In other words I only use the LB to go back to NOT trusting, lol. ",
      "votes": 1
    },
    {
      "id": 1862675,
      "postDate": "2022-07-19T22:22:12.003Z",
      "content": "<p>I think we should choose to trust cv.<br>\nFor me, when my cv score is lower than 0.800, the correspondence between cv and LB is always good; when cv starts to be higher than 0.800, this correspondence becomes less stable. I think my models have fallen into a local optimum.</p>",
      "rawMarkdown": "I think we should choose to trust cv.\nFor me, when my cv score is lower than 0.800, the correspondence between cv and LB is always good; when cv starts to be higher than 0.800, this correspondence becomes less stable. I think my models have fallen into a local optimum.",
      "votes": 1
    },
    {
      "id": 1862047,
      "postDate": "2022-07-19T12:25:36.853Z",
      "content": "<p>In my experience, CV was always more adequate than LB. Also, it's hard to track LB in this competition because of rounding. So, trust your cv :)</p>",
      "rawMarkdown": "In my experience, CV was always more adequate than LB. Also, it's hard to track LB in this competition because of rounding. So, trust your cv :)",
      "votes": 1
    },
    {
      "id": 1864824,
      "postDate": "2022-07-21T10:23:20.880Z",
      "content": "<p>CV and LB, both are important.<br>\nThe public LB has almost as many custoemr_IDs as Train data.<br>\nI think it is important to aim to create a model where CV and LB match.</p>",
      "rawMarkdown": "CV and LB, both are important.\nThe public LB has almost as many custoemr_IDs as Train data.\nI think it is important to aim to create a model where CV and LB match.",
      "votes": 2
    },
    {
      "id": 1863202,
      "postDate": "2022-07-20T08:00:50.153Z",
      "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> Thanks for your sharing, Kaggle Grand Master!</p>",
      "rawMarkdown": "@raddar Thanks for your sharing, Kaggle Grand Master!"
    },
    {
      "id": 1858997,
      "postDate": "2022-07-17T10:32:07.690Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1858410,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2022-07-16T22:09:16.697000",
      "content": "<p>In the end, you have 2 final submissions.</p>\n<p>My strategy has always been: first sub - best CV model. second sub - best CV model/2 + best LB model/2.</p>\n<p>Looking back at my previous competitions, 8 times best CV was best on private LB, 2 times the average was better. And only once best public LB model was best on private.</p>",
      "votes": 23,
      "replies": [
        {
          "id": 1858476,
          "author_name": "1110Ra",
          "author_url": "",
          "post_date": "2022-07-17T00:25:06.397000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> ! So if my CV gets better and the LB get worse simultaneously, it does not necessarily mean that the model is overfitting?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1858698,
          "author_name": "Jianhong Zhu",
          "author_url": "",
          "post_date": "2022-07-17T06:23:27.027000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> for sharing your submission strategy which I've never tried! Hope there won't be a big shake up this time 😃</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1858801,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-07-17T08:07:40.377000",
          "content": "<p><a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a> that is the very definition of public LB overfitting. If you are only looking at things that work on public LB, you are 99% to be overfitting on LB. </p>\n<p>CV increase with public LB decrease should be expected ~ half of the time if CV increases are only slightly incremental.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1858803,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-07-17T08:10:16.713000",
          "content": "<p>And in general, you should avoid using term \"model is overfitting\". Models are overfitting to train set all the time, we still use them. It's the overfitting on LB that matters, and this does not have to do much with models, but more about what decisions we make :)</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 1858989,
          "author_name": "1110Ra",
          "author_url": "",
          "post_date": "2022-07-17T10:24:22.033000",
          "content": "<p>thnk you <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> ! I appreciate your help 🙏</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1858406,
      "author_name": "Joe Eddy",
      "author_url": "",
      "post_date": "2022-07-16T22:05:25.187000",
      "content": "<p>So far my CV/LB has been almost, if not exactly directionally consistent (with LB a bit under .001 or so higher, though of course it's hard to tell with the 4th decimal missing -- my current CV is fold avg ~.7992). I believe others have reported a similar relationship. It's possible we've been lucky with our CV fold setups, but it's worth double-checking how you're calculating your CV and other factors that may be at play. Is your CV score for an ensemble where there might be some subtle leakage? Did you engineer features whose distributions are different between train and test? Those might look great on CV but be mismatched to test.   </p>\n<p>Beyond the point about public LB size made by Tilii, another advantage of the public LB is that the data is later in time than CV and closer to the private LB time. I don't expect huge data drift caused by time for this problem, but even a little (inevitable?) drift could be tricky to deal with. Where I land is wanting to see everything improve to be confident -- I only make submissions when CV improves, and an LB decline would make me hesitate. Also, the classic best CV + best public LB final submission selection seems like a decent hedge, as often.  </p>",
      "votes": 6,
      "replies": [
        {
          "id": 1858693,
          "author_name": "Jianhong Zhu",
          "author_url": "",
          "post_date": "2022-07-17T06:17:33.633000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> for reminding me of double checking my CV and feature engineering! So far I am just using the stratified 5 fold cv that everyone is using. I am thinking if there's a way to improve our cv strategy. Btw, LB seems to be more reliable this time, but who knows.  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1858380,
      "author_name": "Leo",
      "author_url": "",
      "post_date": "2022-07-16T21:32:53.623000",
      "content": "<p>My understanding is that a submission is on the right track when a CV frameworks mirrors an LB score, and with every increase in CV the same delta would appear in LB. That is a clear signal that the solution and the CV framework is meaningful. So maybe an answer to your question is that they go hand in hand? Either way its a good experience and every dataset is special in some way or the other! Good luck :)  </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1858321,
      "author_name": "Tilii",
      "author_url": "",
      "post_date": "2022-07-16T19:59:10.883000",
      "content": "<p>The same question likely has come up in every single competition. A general answer is that CV scores are more reliable because we have control over data preparation and prediction methods. That said, we don't know whether your models employ internal fold consistency and how exactly they were made.</p>\n<p>It should be mentioned that a public leaderboard that is based on 50% of test data - like in this competition - is more trustworthy than a leaderboard based on 10% of test data. I think for most models the <code>[CV-LB]</code> difference will end up being small even when private leaderboard is released.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1858680,
          "author_name": "Jianhong Zhu",
          "author_url": "",
          "post_date": "2022-07-17T06:04:33.810000",
          "content": "<p>LB based on 50% of test data is exactly what I am concerned about, meaning my model is not generalizing well in the test data even CV score is improved. I gotta put more effort on the data and cv strategy than the model itself. Thank you <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1864465,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2022-07-21T04:27:57.407000",
      "content": "<p>I have trust issues. I don't trust either one. But a bunch of CV scores across different seeds? I start to trust. Since that's impossible on the LB I'd only use it to verify. In other words I only use the LB to go back to NOT trusting, lol. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1862675,
      "author_name": "Mengfei Li",
      "author_url": "",
      "post_date": "2022-07-19T22:22:12.003000",
      "content": "<p>I think we should choose to trust cv.<br>\nFor me, when my cv score is lower than 0.800, the correspondence between cv and LB is always good; when cv starts to be higher than 0.800, this correspondence becomes less stable. I think my models have fallen into a local optimum.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1862047,
      "author_name": "Sasha Turutin",
      "author_url": "",
      "post_date": "2022-07-19T12:25:36.853000",
      "content": "<p>In my experience, CV was always more adequate than LB. Also, it's hard to track LB in this competition because of rounding. So, trust your cv :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1864824,
      "author_name": "zakopuro",
      "author_url": "",
      "post_date": "2022-07-21T10:23:20.880000",
      "content": "<p>CV and LB, both are important.<br>\nThe public LB has almost as many custoemr_IDs as Train data.<br>\nI think it is important to aim to create a model where CV and LB match.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1863202,
      "author_name": "Kakushi",
      "author_url": "",
      "post_date": "2022-07-20T08:00:50.153000",
      "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> Thanks for your sharing, Kaggle Grand Master!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1858997,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-17T10:32:07.690000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1858410": "In the end, you have 2 final submissions.\n\nMy strategy has always been: first sub - best CV model. second sub - best CV model/2 + best LB model/2.\n\nLooking back at my previous competitions, 8 times best CV was best on private LB, 2 times the average was better. And only once best public LB model was best on private.\n",
    "1858296": "I have done some experiments on my LGBM model, and I found that even my CV score got improved but the LB score got worse. Does anyone have the same issue? Which score should we trust more, CV or LB? Thank you!\nHere's my score:\nCV 0.7964 ---- LB 0.799\nCV 0.7972 ---- LB 0.798\nCV 0.7976 ---- LB 0.798\n",
    "1858406": "So far my CV/LB has been almost, if not exactly directionally consistent (with LB a bit under .001 or so higher, though of course it's hard to tell with the 4th decimal missing -- my current CV is fold avg ~.7992). I believe others have reported a similar relationship. It's possible we've been lucky with our CV fold setups, but it's worth double-checking how you're calculating your CV and other factors that may be at play. Is your CV score for an ensemble where there might be some subtle leakage? Did you engineer features whose distributions are different between train and test? Those might look great on CV but be mismatched to test.   \n\nBeyond the point about public LB size made by Tilii, another advantage of the public LB is that the data is later in time than CV and closer to the private LB time. I don't expect huge data drift caused by time for this problem, but even a little (inevitable?) drift could be tricky to deal with. Where I land is wanting to see everything improve to be confident -- I only make submissions when CV improves, and an LB decline would make me hesitate. Also, the classic best CV + best public LB final submission selection seems like a decent hedge, as often.  ",
    "1858380": "My understanding is that a submission is on the right track when a CV frameworks mirrors an LB score, and with every increase in CV the same delta would appear in LB. That is a clear signal that the solution and the CV framework is meaningful. So maybe an answer to your question is that they go hand in hand? Either way its a good experience and every dataset is special in some way or the other! Good luck :)  ",
    "1858321": "The same question likely has come up in every single competition. A general answer is that CV scores are more reliable because we have control over data preparation and prediction methods. That said, we don't know whether your models employ internal fold consistency and how exactly they were made.\n\nIt should be mentioned that a public leaderboard that is based on 50% of test data - like in this competition - is more trustworthy than a leaderboard based on 10% of test data. I think for most models the `[CV-LB]` difference will end up being small even when private leaderboard is released.",
    "1864465": "I have trust issues. I don't trust either one. But a bunch of CV scores across different seeds? I start to trust. Since that's impossible on the LB I'd only use it to verify. In other words I only use the LB to go back to NOT trusting, lol. ",
    "1862675": "I think we should choose to trust cv.\nFor me, when my cv score is lower than 0.800, the correspondence between cv and LB is always good; when cv starts to be higher than 0.800, this correspondence becomes less stable. I think my models have fallen into a local optimum.",
    "1862047": "In my experience, CV was always more adequate than LB. Also, it's hard to track LB in this competition because of rounding. So, trust your cv :)",
    "1864824": "CV and LB, both are important.\nThe public LB has almost as many custoemr_IDs as Train data.\nI think it is important to aim to create a model where CV and LB match.",
    "1863202": "@raddar Thanks for your sharing, Kaggle Grand Master!",
    "1858997": ""
  }
}