{
  "id": 488466,
  "title": "Instability of Public LB and Issues with Time-Series CV",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/488466",
  "author_name": "",
  "post_date": "2024-04-02T14:12:36.660854600Z",
  "votes": 37,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I've encountered two issues:</p>\n<h2>1. The Public LB is unstable, and no correlation with local CV can be found.</h2>\n<p>This competition adopt a special evaluation metric focused on stability, which might be the reason, but the Public LB score is quite unstable.</p>\n<p>For instance, even when the local CV score increases, the Public LB score often decreases, showing no strong positive correlation. <br>\nMinor parameter changes can also lead to significant fluctuations in the Public LB.</p>\n<p>Additionally, submitting <strong>the exact same code</strong> can result in varying scores. I've submitted a publicl notebook as a baseline multiple times and observed variations from 0.568 to 0.570. <br>\nThis might be due to the lack of reproducibility when using GPU for calculations in LightGBM, but regardless, it indicates that the Public LB is unstable.</p>\n<p>I have no idea how to address this issue.</p>\n<h2>2. Should Cross Validation consider time series data?</h2>\n<p>Many public notebooks simply split the data into 5-fold without considering the time-series.<br>\nThis can lead to evaluating past data using future data, which generally isn't appropriate due to the leakage problem. <br>\nIdeally, future data should be used as the validation set, and past data as the training set.<br>\nI tried this approach, but the Public LB score unchanged, so I was confused.</p>\n<h2>Summary</h2>\n<p>For these reasons, I'm uncertain about how to proceed in this competition. <br>\nI think it's risky that we can't find a correlation between the public LB and local CV, which means<br>\nthat many of the currently high-scoring public notebooks might be overfitting to the public LB and could shake for Private LB.</p>\n<p>Does anyone have any opinion?</p>",
  "messages": [
    {
      "id": "2729018",
      "postDate": "04/02/2024 14:12:36",
      "content": "<p>I've encountered two issues:</p>\n<h2>1. The Public LB is unstable, and no correlation with local CV can be found.</h2>\n<p>This competition adopt a special evaluation metric focused on stability, which might be the reason, but the Public LB score is quite unstable.</p>\n<p>For instance, even when the local CV score increases, the Public LB score often decreases, showing no strong positive correlation. <br>\nMinor parameter changes can also lead to significant fluctuations in the Public LB.</p>\n<p>Additionally, submitting <strong>the exact same code</strong> can result in varying scores. I've submitted a publicl notebook as a baseline multiple times and observed variations from 0.568 to 0.570. <br>\nThis might be due to the lack of reproducibility when using GPU for calculations in LightGBM, but regardless, it indicates that the Public LB is unstable.</p>\n<p>I have no idea how to address this issue.</p>\n<h2>2. Should Cross Validation consider time series data?</h2>\n<p>Many public notebooks simply split the data into 5-fold without considering the time-series.<br>\nThis can lead to evaluating past data using future data, which generally isn't appropriate due to the leakage problem. <br>\nIdeally, future data should be used as the validation set, and past data as the training set.<br>\nI tried this approach, but the Public LB score unchanged, so I was confused.</p>\n<h2>Summary</h2>\n<p>For these reasons, I'm uncertain about how to proceed in this competition. <br>\nI think it's risky that we can't find a correlation between the public LB and local CV, which means<br>\nthat many of the currently high-scoring public notebooks might be overfitting to the public LB and could shake for Private LB.</p>\n<p>Does anyone have any opinion?</p>",
      "rawMarkdown": "I've encountered two issues:\n\n## 1. The Public LB is unstable, and no correlation with local CV can be found.\nThis competition adopt a special evaluation metric focused on stability, which might be the reason, but the Public LB score is quite unstable.\n\nFor instance, even when the local CV score increases, the Public LB score often decreases, showing no strong positive correlation. \nMinor parameter changes can also lead to significant fluctuations in the Public LB.\n\nAdditionally, submitting **the exact same code** can result in varying scores. I've submitted a publicl notebook as a baseline multiple times and observed variations from 0.568 to 0.570. \nThis might be due to the lack of reproducibility when using GPU for calculations in LightGBM, but regardless, it indicates that the Public LB is unstable.\n\nI have no idea how to address this issue.\n\n## 2. Should Cross Validation consider time series data?\n\nMany public notebooks simply split the data into 5-fold without considering the time-series.\nThis can lead to evaluating past data using future data, which generally isn't appropriate due to the leakage problem. \nIdeally, future data should be used as the validation set, and past data as the training set.\nI tried this approach, but the Public LB score unchanged, so I was confused.\n\n## Summary\nFor these reasons, I'm uncertain about how to proceed in this competition. \nI think it's risky that we can't find a correlation between the public LB and local CV, which means\nthat many of the currently high-scoring public notebooks might be overfitting to the public LB and could shake for Private LB.\n\nDoes anyone have any opinion?",
      "votes": null
    },
    {
      "id": "2729076",
      "postDate": "04/02/2024 14:38:49",
      "content": "<p><a href=\"https://www.kaggle.com/tetsuro731\" target=\"_blank\">@tetsuro731</a> I totally agree with this too.</p>\n<p>I also accidently did your experiment lately and have experienced code result fluctuations in the same range. <strong>Is everyone's public leaderboard the same?</strong> <br>\nI also concur on the shakeup potential and CV strategy too. I think building a good CV is of prime importance here and I doubt if this stratified grouped K-fold lives up to the task. Perhaps this thwarts past future surrogacy and this could be a problem in the private evaluation. </p>",
      "rawMarkdown": "tetsuro731 I totally agree with this too.\n\nI also accidently did your experiment lately and have experienced code result fluctuations in the same range. **Is everyone's public leaderboard the same?** \nI also concur on the shakeup potential and CV strategy too. I think building a good CV is of prime importance here and I doubt if this stratified grouped K-fold lives up to the task. Perhaps this thwarts past future surrogacy and this could be a problem in the private evaluation.",
      "votes": null
    },
    {
      "id": "2729255",
      "postDate": "04/02/2024 16:10:01",
      "content": "<p><a href=\"https://www.kaggle.com/tetsuro731\" target=\"_blank\">@tetsuro731</a> Due to the fact that I could not find a correlation between the leaderboard and local tests, I lost interest in the task.<br>\nI come in periodically to see if anyone has a solution.</p>\n<p>Due to the fact that we do not understand what the test sample represents, perhaps everyone is standing still.</p>\n<p>It’s not interesting to come up with new cool features when they locally give an increase, but not lb</p>",
      "rawMarkdown": "tetsuro731 Due to the fact that I could not find a correlation between the leaderboard and local tests, I lost interest in the task.\nI come in periodically to see if anyone has a solution.\n\nDue to the fact that we do not understand what the test sample represents, perhaps everyone is standing still.\n\nIt’s not interesting to come up with new cool features when they locally give an increase, but not lb",
      "votes": null
    },
    {
      "id": "2750193",
      "postDate": "04/13/2024 14:04:09",
      "content": "<blockquote>\n  <p>It’s not interesting to come up with new cool features when they locally give an increase, but not lb</p>\n</blockquote>\n<p>That totally makes sense to me, but I'd like to discover and figure out better approach for this competition.</p>",
      "rawMarkdown": ">It’s not interesting to come up with new cool features when they locally give an increase, but not lb\n\nThat totally makes sense to me, but I'd like to discover and figure out better approach for this competition.",
      "votes": null
    },
    {
      "id": "2750194",
      "postDate": "04/13/2024 14:06:53",
      "content": "<blockquote>\n  <p>Is everyone's public leaderboard the same?</p>\n</blockquote>\n<p>Good point.<br>\nI guess there is fluctuation even if using exactly the same code, which could be the problem.<br>\nOn the other hand, you know, final score will be calculated by the other private test data so creating robust model could be the better way to win the competition in my understanding.</p>",
      "rawMarkdown": ">Is everyone's public leaderboard the same?\n\nGood point.\nI guess there is fluctuation even if using exactly the same code, which could be the problem.\nOn the other hand, you know, final score will be calculated by the other private test data so creating robust model could be the better way to win the competition in my understanding.",
      "votes": null
    },
    {
      "id": "2752894",
      "postDate": "04/15/2024 07:54:11",
      "content": "<p>Hi,<br>\nsplit public - private leaderboard is fixed and same for every Kaggler. So difference in LB score is only caused by different model, not data.</p>",
      "rawMarkdown": "Hi,\nsplit public - private leaderboard is fixed and same for every Kaggler. So difference in LB score is only caused by different model, not data.",
      "votes": null
    },
    {
      "id": "2753406",
      "postDate": "04/15/2024 13:23:41",
      "content": "<blockquote>\n  <p>Additionally, submitting the exact same code can result in varying scores. I've submitted a publicl notebook as a baseline multiple times and observed variations from 0.568 to 0.570.</p>\n</blockquote>\n<p>I've also noticed this issue. I think this is caused by the randomness of the GPU version of CatBoost.  I am now starting to save the model each time it runs for future use.</p>",
      "rawMarkdown": ">Additionally, submitting the exact same code can result in varying scores. I've submitted a publicl notebook as a baseline multiple times and observed variations from 0.568 to 0.570.\n\nI've also noticed this issue. I think this is caused by the randomness of the GPU version of CatBoost.  I am now starting to save the model each time it runs for future use.",
      "votes": null
    },
    {
      "id": "2753675",
      "postDate": "04/15/2024 16:00:16",
      "content": "<p>That is exactly what is this competition about, a way how to create new features that would be able to take into account instabilities and the same for a model. </p>",
      "rawMarkdown": "That is exactly what is this competition about, a way how to create new features that would be able to take into account instabilities and the same for a model.",
      "votes": null
    },
    {
      "id": "2756105",
      "postDate": "04/16/2024 21:29:53",
      "content": "<p>I think using a time series split may give more accurate loss and and metric estimates but at the same time you're not using the most up to date data for training (just validation), which could be a detriment when testing on LB data. I wonder if it's better to use time series splitting for hyperparameter tuning and stratified k-fold for actual training for LB.</p>",
      "rawMarkdown": "I think using a time series split may give more accurate loss and and metric estimates but at the same time you're not using the most up to date data for training (just validation), which could be a detriment when testing on LB data. I wonder if it's better to use time series splitting for hyperparameter tuning and stratified k-fold for actual training for LB.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2729076,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "04/02/2024 14:38:49",
      "content": "<p><a href=\"https://www.kaggle.com/tetsuro731\" target=\"_blank\">@tetsuro731</a> I totally agree with this too.</p>\n<p>I also accidently did your experiment lately and have experienced code result fluctuations in the same range. <strong>Is everyone's public leaderboard the same?</strong> <br>\nI also concur on the shakeup potential and CV strategy too. I think building a good CV is of prime importance here and I doubt if this stratified grouped K-fold lives up to the task. Perhaps this thwarts past future surrogacy and this could be a problem in the private evaluation. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2750194,
          "author_name": "tetsuro731",
          "author_url": "",
          "post_date": "04/13/2024 14:06:53",
          "content": "<blockquote>\n  <p>Is everyone's public leaderboard the same?</p>\n</blockquote>\n<p>Good point.<br>\nI guess there is fluctuation even if using exactly the same code, which could be the problem.<br>\nOn the other hand, you know, final score will be calculated by the other private test data so creating robust model could be the better way to win the competition in my understanding.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2752894,
              "author_name": "tomasjeline2",
              "author_url": "",
              "post_date": "04/15/2024 07:54:11",
              "content": "<p>Hi,<br>\nsplit public - private leaderboard is fixed and same for every Kaggler. So difference in LB score is only caused by different model, not data.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2729255,
      "author_name": "dima1992",
      "author_url": "",
      "post_date": "04/02/2024 16:10:01",
      "content": "<p><a href=\"https://www.kaggle.com/tetsuro731\" target=\"_blank\">@tetsuro731</a> Due to the fact that I could not find a correlation between the leaderboard and local tests, I lost interest in the task.<br>\nI come in periodically to see if anyone has a solution.</p>\n<p>Due to the fact that we do not understand what the test sample represents, perhaps everyone is standing still.</p>\n<p>It’s not interesting to come up with new cool features when they locally give an increase, but not lb</p>",
      "votes": null,
      "replies": [
        {
          "id": 2750193,
          "author_name": "tetsuro731",
          "author_url": "",
          "post_date": "04/13/2024 14:04:09",
          "content": "<blockquote>\n  <p>It’s not interesting to come up with new cool features when they locally give an increase, but not lb</p>\n</blockquote>\n<p>That totally makes sense to me, but I'd like to discover and figure out better approach for this competition.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2753675,
              "author_name": "jetakow",
              "author_url": "",
              "post_date": "04/15/2024 16:00:16",
              "content": "<p>That is exactly what is this competition about, a way how to create new features that would be able to take into account instabilities and the same for a model. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2753406,
      "author_name": "bruceqdu",
      "author_url": "",
      "post_date": "04/15/2024 13:23:41",
      "content": "<blockquote>\n  <p>Additionally, submitting the exact same code can result in varying scores. I've submitted a publicl notebook as a baseline multiple times and observed variations from 0.568 to 0.570.</p>\n</blockquote>\n<p>I've also noticed this issue. I think this is caused by the randomness of the GPU version of CatBoost.  I am now starting to save the model each time it runs for future use.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2756105,
      "author_name": "caelhasse",
      "author_url": "",
      "post_date": "04/16/2024 21:29:53",
      "content": "<p>I think using a time series split may give more accurate loss and and metric estimates but at the same time you're not using the most up to date data for training (just validation), which could be a detriment when testing on LB data. I wonder if it's better to use time series splitting for hyperparameter tuning and stratified k-fold for actual training for LB.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2729018": "I've encountered two issues:\n\n## 1. The Public LB is unstable, and no correlation with local CV can be found.\nThis competition adopt a special evaluation metric focused on stability, which might be the reason, but the Public LB score is quite unstable.\n\nFor instance, even when the local CV score increases, the Public LB score often decreases, showing no strong positive correlation. \nMinor parameter changes can also lead to significant fluctuations in the Public LB.\n\nAdditionally, submitting **the exact same code** can result in varying scores. I've submitted a publicl notebook as a baseline multiple times and observed variations from 0.568 to 0.570. \nThis might be due to the lack of reproducibility when using GPU for calculations in LightGBM, but regardless, it indicates that the Public LB is unstable.\n\nI have no idea how to address this issue.\n\n## 2. Should Cross Validation consider time series data?\n\nMany public notebooks simply split the data into 5-fold without considering the time-series.\nThis can lead to evaluating past data using future data, which generally isn't appropriate due to the leakage problem. \nIdeally, future data should be used as the validation set, and past data as the training set.\nI tried this approach, but the Public LB score unchanged, so I was confused.\n\n## Summary\nFor these reasons, I'm uncertain about how to proceed in this competition. \nI think it's risky that we can't find a correlation between the public LB and local CV, which means\nthat many of the currently high-scoring public notebooks might be overfitting to the public LB and could shake for Private LB.\n\nDoes anyone have any opinion?",
    "2729076": "tetsuro731 I totally agree with this too.\n\nI also accidently did your experiment lately and have experienced code result fluctuations in the same range. **Is everyone's public leaderboard the same?** \nI also concur on the shakeup potential and CV strategy too. I think building a good CV is of prime importance here and I doubt if this stratified grouped K-fold lives up to the task. Perhaps this thwarts past future surrogacy and this could be a problem in the private evaluation.",
    "2729255": "tetsuro731 Due to the fact that I could not find a correlation between the leaderboard and local tests, I lost interest in the task.\nI come in periodically to see if anyone has a solution.\n\nDue to the fact that we do not understand what the test sample represents, perhaps everyone is standing still.\n\nIt’s not interesting to come up with new cool features when they locally give an increase, but not lb",
    "2750193": ">It’s not interesting to come up with new cool features when they locally give an increase, but not lb\n\nThat totally makes sense to me, but I'd like to discover and figure out better approach for this competition.",
    "2750194": ">Is everyone's public leaderboard the same?\n\nGood point.\nI guess there is fluctuation even if using exactly the same code, which could be the problem.\nOn the other hand, you know, final score will be calculated by the other private test data so creating robust model could be the better way to win the competition in my understanding.",
    "2752894": "Hi,\nsplit public - private leaderboard is fixed and same for every Kaggler. So difference in LB score is only caused by different model, not data.",
    "2753406": ">Additionally, submitting the exact same code can result in varying scores. I've submitted a publicl notebook as a baseline multiple times and observed variations from 0.568 to 0.570.\n\nI've also noticed this issue. I think this is caused by the randomness of the GPU version of CatBoost.  I am now starting to save the model each time it runs for future use.",
    "2753675": "That is exactly what is this competition about, a way how to create new features that would be able to take into account instabilities and the same for a model.",
    "2756105": "I think using a time series split may give more accurate loss and and metric estimates but at the same time you're not using the most up to date data for training (just validation), which could be a detriment when testing on LB data. I wonder if it's better to use time series splitting for hyperparameter tuning and stratified k-fold for actual training for LB."
  },
  "source": "meta"
}