{
  "id": 490762,
  "title": "Why do LB scores vary so much?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/490762",
  "author_name": "",
  "post_date": "2024-04-03T12:45:41.244435400Z",
  "votes": 12,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hello Kagglers, </p>\n<p>Based on many high scoring and well documented public notebooks, I created a week ago a notebook that scores 0.573. A week after, I download the notebook and import it on a new notebooks and save it. The model parameters, seed and data preprocessing pipelines are exactly the same. However, the new notebook, scored 0.564 - 0.009 less than the original one. Do you also face such fluctuations on your scores?</p>\n<p>I wouldn't bother if it was 0.009 up, but it got worse!</p>",
  "messages": [
    {
      "id": "2732982",
      "postDate": "04/03/2024 12:45:41",
      "content": "<p>Hello Kagglers, </p>\n<p>Based on many high scoring and well documented public notebooks, I created a week ago a notebook that scores 0.573. A week after, I download the notebook and import it on a new notebooks and save it. The model parameters, seed and data preprocessing pipelines are exactly the same. However, the new notebook, scored 0.564 - 0.009 less than the original one. Do you also face such fluctuations on your scores?</p>\n<p>I wouldn't bother if it was 0.009 up, but it got worse!</p>",
      "rawMarkdown": "Hello Kagglers, \n\nBased on many high scoring and well documented public notebooks, I created a week ago a notebook that scores 0.573. A week after, I download the notebook and import it on a new notebooks and save it. The model parameters, seed and data preprocessing pipelines are exactly the same. However, the new notebook, scored 0.564 - 0.009 less than the original one. Do you also face such fluctuations on your scores?\n\nI wouldn't bother if it was 0.009 up, but it got worse!",
      "votes": null
    },
    {
      "id": "2733000",
      "postDate": "04/03/2024 13:06:22",
      "content": "<p><a href=\"https://www.kaggle.com/andreasbis\" target=\"_blank\">@andreasbis</a> I also faced the same issue-  I have tried the same experiment 2-3 times and face the same problem as you, with a bigger fluctuation. The fluctuation in scores is high in my opinion and local CV and LB are not in sync at all. <br>\nThanks for raising the issue!</p>",
      "rawMarkdown": "andreasbis I also faced the same issue-  I have tried the same experiment 2-3 times and face the same problem as you, with a bigger fluctuation. The fluctuation in scores is high in my opinion and local CV and LB are not in sync at all. \nThanks for raising the issue!",
      "votes": null
    },
    {
      "id": "2733038",
      "postDate": "04/03/2024 13:32:27",
      "content": "<p>May I ask if you do not have a fixed random seed?<br>\n`import random#提供了一些用于生成随机数的函数</p>\n<h1>设置随机种子,保证模型可以复现</h1>\n<p>def seed_everything(seed):<br>\n    np.random.seed(seed)#numpy的随机种子<br>\n    random.seed(seed)#python内置的随机种子<br>\nseed_everything(cfg.seed)`</p>",
      "rawMarkdown": "May I ask if you do not have a fixed random seed?\n`import random#提供了一些用于生成随机数的函数\n#设置随机种子,保证模型可以复现\ndef seed_everything(seed):\n    np.random.seed(seed)#numpy的随机种子\n    random.seed(seed)#python内置的随机种子\nseed_everything(cfg.seed)`",
      "votes": null
    },
    {
      "id": "2733048",
      "postDate": "04/03/2024 13:37:12",
      "content": "<p>The truth is I only use seed (42) for the LightGBM model. But from my experience at least, a seed_everything function does not help at all.</p>",
      "rawMarkdown": "The truth is I only use seed (42) for the LightGBM model. But from my experience at least, a seed_everything function does not help at all.",
      "votes": null
    },
    {
      "id": "2733051",
      "postDate": "04/03/2024 13:39:12",
      "content": "<p>I have saved the model objects as joblib files and import them. I have seeded everything too. As a simple experiment, I took a public kernel and resubmitted just to validate it and got a different score from the kernel itself. I thought this was my problem alone but find others facing this as well. I hypothesize the below-</p>\n<ol>\n<li>Probably our CV strategy is collectively incorrect </li>\n<li>Submitted files are perhaps accessing different datasets of the same fraction as public LB</li>\n<li>The public kernel I took perhaps has a seed problem (need to check this further)</li>\n</ol>\n<p>This competition is certainly not as straight forward as it looks- I have learnt it thoroughly from experience <a href=\"https://www.kaggle.com/andreasbis\" target=\"_blank\">@andreasbis</a> <a href=\"https://www.kaggle.com/yunsuxiaozi\" target=\"_blank\">@yunsuxiaozi</a> </p>",
      "rawMarkdown": "I have saved the model objects as joblib files and import them. I have seeded everything too. As a simple experiment, I took a public kernel and resubmitted just to validate it and got a different score from the kernel itself. I thought this was my problem alone but find others facing this as well. I hypothesize the below-\n1. Probably our CV strategy is collectively incorrect \n2. Submitted files are perhaps accessing different datasets of the same fraction as public LB\n3. The public kernel I took perhaps has a seed problem (need to check this further)\n\nThis competition is certainly not as straight forward as it looks- I have learnt it thoroughly from experience @andreasbis @yunsuxiaozi",
      "votes": null
    },
    {
      "id": "2733064",
      "postDate": "04/03/2024 13:44:10",
      "content": "<p>Thus, considering your findings, we should evaluate our work from CV and not public LB?</p>",
      "rawMarkdown": "Thus, considering your findings, we should evaluate our work from CV and not public LB?",
      "votes": null
    },
    {
      "id": "2733162",
      "postDate": "04/03/2024 14:23:09",
      "content": "<p>In this case, it becomes even more important to use CV score as an evaluator for model performance <a href=\"https://www.kaggle.com/andreasbis\" target=\"_blank\">@andreasbis</a> </p>",
      "rawMarkdown": "In this case, it becomes even more important to use CV score as an evaluator for model performance @andreasbis",
      "votes": null
    },
    {
      "id": "2733201",
      "postDate": "04/03/2024 14:41:33",
      "content": "<p>This issue in the <code>lightgbm</code> may be on the same topic.<br>\n<a href=\"https://github.com/microsoft/LightGBM/issues/2659\" target=\"_blank\">https://github.com/microsoft/LightGBM/issues/2659</a></p>",
      "rawMarkdown": "This issue in the `lightgbm` may be on the same topic.\nhttps://github.com/microsoft/LightGBM/issues/2659",
      "votes": null
    },
    {
      "id": "2733344",
      "postDate": "04/03/2024 15:38:21",
      "content": "<blockquote>\n  <p>I wouldn't bother if it was 0.009 up</p>\n</blockquote>\n<p>If you wouldn't worry for 0.009 up, then you shouldn't worry for 0.009 down 😉</p>",
      "rawMarkdown": ">I wouldn't bother if it was 0.009 up\n\nIf you wouldn't worry for 0.009 up, then you shouldn't worry for 0.009 down 😉",
      "votes": null
    },
    {
      "id": "2733600",
      "postDate": "04/03/2024 17:52:25",
      "content": "<p>Besides the random seed issue discussed in other comments, perhaps because this metric is extremely noisy and random on truly unknown data that over time will diverge from training data.</p>\n<p>Even setting aside metric hacks, I know of no real fix for time-decay error other than, well, using the time variable. Which if I'm reading right (just browsing this competition so far), that is intended to be difficult if not impossible.</p>\n<p>Anyways, back a bit more on topic, probably you want to find model updates that improve both CV and LB scores, and diligently throw out adding features that only improve one or the other. And you also probably want to try a bunch of seeds and runs and take the score as the average minus the stddev or something similar. Higher variance models is probably a bad sign, lower variance models is probably a very good sign in this competition.</p>",
      "rawMarkdown": "Besides the random seed issue discussed in other comments, perhaps because this metric is extremely noisy and random on truly unknown data that over time will diverge from training data.\n\nEven setting aside metric hacks, I know of no real fix for time-decay error other than, well, using the time variable. Which if I'm reading right (just browsing this competition so far), that is intended to be difficult if not impossible.\n\nAnyways, back a bit more on topic, probably you want to find model updates that improve both CV and LB scores, and diligently throw out adding features that only improve one or the other. And you also probably want to try a bunch of seeds and runs and take the score as the average minus the stddev or something similar. Higher variance models is probably a bad sign, lower variance models is probably a very good sign in this competition.",
      "votes": null
    },
    {
      "id": "2733616",
      "postDate": "04/03/2024 18:04:35",
      "content": "<blockquote>\n  <p>Higher variance models is probably a bad sign, lower variance models is probably a very good sign in this competition.</p>\n</blockquote>\n<p>I will look more into that. Thanks!</p>",
      "rawMarkdown": ">Higher variance models is probably a bad sign, lower variance models is probably a very good sign in this competition.\n\nI will look more into that. Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2733000,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "04/03/2024 13:06:22",
      "content": "<p><a href=\"https://www.kaggle.com/andreasbis\" target=\"_blank\">@andreasbis</a> I also faced the same issue-  I have tried the same experiment 2-3 times and face the same problem as you, with a bigger fluctuation. The fluctuation in scores is high in my opinion and local CV and LB are not in sync at all. <br>\nThanks for raising the issue!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2733038,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "04/03/2024 13:32:27",
      "content": "<p>May I ask if you do not have a fixed random seed?<br>\n`import random#提供了一些用于生成随机数的函数</p>\n<h1>设置随机种子,保证模型可以复现</h1>\n<p>def seed_everything(seed):<br>\n    np.random.seed(seed)#numpy的随机种子<br>\n    random.seed(seed)#python内置的随机种子<br>\nseed_everything(cfg.seed)`</p>",
      "votes": null,
      "replies": [
        {
          "id": 2733048,
          "author_name": "andreasbis",
          "author_url": "",
          "post_date": "04/03/2024 13:37:12",
          "content": "<p>The truth is I only use seed (42) for the LightGBM model. But from my experience at least, a seed_everything function does not help at all.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2733051,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "04/03/2024 13:39:12",
          "content": "<p>I have saved the model objects as joblib files and import them. I have seeded everything too. As a simple experiment, I took a public kernel and resubmitted just to validate it and got a different score from the kernel itself. I thought this was my problem alone but find others facing this as well. I hypothesize the below-</p>\n<ol>\n<li>Probably our CV strategy is collectively incorrect </li>\n<li>Submitted files are perhaps accessing different datasets of the same fraction as public LB</li>\n<li>The public kernel I took perhaps has a seed problem (need to check this further)</li>\n</ol>\n<p>This competition is certainly not as straight forward as it looks- I have learnt it thoroughly from experience <a href=\"https://www.kaggle.com/andreasbis\" target=\"_blank\">@andreasbis</a> <a href=\"https://www.kaggle.com/yunsuxiaozi\" target=\"_blank\">@yunsuxiaozi</a> </p>",
          "votes": null,
          "replies": [
            {
              "id": 2733064,
              "author_name": "andreasbis",
              "author_url": "",
              "post_date": "04/03/2024 13:44:10",
              "content": "<p>Thus, considering your findings, we should evaluate our work from CV and not public LB?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2733162,
                  "author_name": "ravi20076",
                  "author_url": "",
                  "post_date": "04/03/2024 14:23:09",
                  "content": "<p>In this case, it becomes even more important to use CV score as an evaluator for model performance <a href=\"https://www.kaggle.com/andreasbis\" target=\"_blank\">@andreasbis</a> </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2733201,
      "author_name": "hirokiueno",
      "author_url": "",
      "post_date": "04/03/2024 14:41:33",
      "content": "<p>This issue in the <code>lightgbm</code> may be on the same topic.<br>\n<a href=\"https://github.com/microsoft/LightGBM/issues/2659\" target=\"_blank\">https://github.com/microsoft/LightGBM/issues/2659</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2733344,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "04/03/2024 15:38:21",
      "content": "<blockquote>\n  <p>I wouldn't bother if it was 0.009 up</p>\n</blockquote>\n<p>If you wouldn't worry for 0.009 up, then you shouldn't worry for 0.009 down 😉</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2733600,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "04/03/2024 17:52:25",
      "content": "<p>Besides the random seed issue discussed in other comments, perhaps because this metric is extremely noisy and random on truly unknown data that over time will diverge from training data.</p>\n<p>Even setting aside metric hacks, I know of no real fix for time-decay error other than, well, using the time variable. Which if I'm reading right (just browsing this competition so far), that is intended to be difficult if not impossible.</p>\n<p>Anyways, back a bit more on topic, probably you want to find model updates that improve both CV and LB scores, and diligently throw out adding features that only improve one or the other. And you also probably want to try a bunch of seeds and runs and take the score as the average minus the stddev or something similar. Higher variance models is probably a bad sign, lower variance models is probably a very good sign in this competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2733616,
          "author_name": "andreasbis",
          "author_url": "",
          "post_date": "04/03/2024 18:04:35",
          "content": "<blockquote>\n  <p>Higher variance models is probably a bad sign, lower variance models is probably a very good sign in this competition.</p>\n</blockquote>\n<p>I will look more into that. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2732982": "Hello Kagglers, \n\nBased on many high scoring and well documented public notebooks, I created a week ago a notebook that scores 0.573. A week after, I download the notebook and import it on a new notebooks and save it. The model parameters, seed and data preprocessing pipelines are exactly the same. However, the new notebook, scored 0.564 - 0.009 less than the original one. Do you also face such fluctuations on your scores?\n\nI wouldn't bother if it was 0.009 up, but it got worse!",
    "2733000": "andreasbis I also faced the same issue-  I have tried the same experiment 2-3 times and face the same problem as you, with a bigger fluctuation. The fluctuation in scores is high in my opinion and local CV and LB are not in sync at all. \nThanks for raising the issue!",
    "2733038": "May I ask if you do not have a fixed random seed?\n`import random#提供了一些用于生成随机数的函数\n#设置随机种子,保证模型可以复现\ndef seed_everything(seed):\n    np.random.seed(seed)#numpy的随机种子\n    random.seed(seed)#python内置的随机种子\nseed_everything(cfg.seed)`",
    "2733048": "The truth is I only use seed (42) for the LightGBM model. But from my experience at least, a seed_everything function does not help at all.",
    "2733051": "I have saved the model objects as joblib files and import them. I have seeded everything too. As a simple experiment, I took a public kernel and resubmitted just to validate it and got a different score from the kernel itself. I thought this was my problem alone but find others facing this as well. I hypothesize the below-\n1. Probably our CV strategy is collectively incorrect \n2. Submitted files are perhaps accessing different datasets of the same fraction as public LB\n3. The public kernel I took perhaps has a seed problem (need to check this further)\n\nThis competition is certainly not as straight forward as it looks- I have learnt it thoroughly from experience @andreasbis @yunsuxiaozi",
    "2733064": "Thus, considering your findings, we should evaluate our work from CV and not public LB?",
    "2733162": "In this case, it becomes even more important to use CV score as an evaluator for model performance @andreasbis",
    "2733201": "This issue in the `lightgbm` may be on the same topic.\nhttps://github.com/microsoft/LightGBM/issues/2659",
    "2733344": ">I wouldn't bother if it was 0.009 up\n\nIf you wouldn't worry for 0.009 up, then you shouldn't worry for 0.009 down 😉",
    "2733600": "Besides the random seed issue discussed in other comments, perhaps because this metric is extremely noisy and random on truly unknown data that over time will diverge from training data.\n\nEven setting aside metric hacks, I know of no real fix for time-decay error other than, well, using the time variable. Which if I'm reading right (just browsing this competition so far), that is intended to be difficult if not impossible.\n\nAnyways, back a bit more on topic, probably you want to find model updates that improve both CV and LB scores, and diligently throw out adding features that only improve one or the other. And you also probably want to try a bunch of seeds and runs and take the score as the average minus the stddev or something similar. Higher variance models is probably a bad sign, lower variance models is probably a very good sign in this competition.",
    "2733616": ">Higher variance models is probably a bad sign, lower variance models is probably a very good sign in this competition.\n\nI will look more into that. Thanks!"
  },
  "source": "meta"
}