{
  "id": 473765,
  "title": "Covid-19 Gonna make us dance again",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/473765",
  "author_name": "",
  "post_date": "2024-02-06T02:32:27.282531400Z",
  "votes": 44,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi, </p>\n<p>I've just started looking into the data for this competition and it's both frightening  and exciting (especially number of features to select from).<br>\nI've been doing finance and ecltv modeling work for quite a while and the covid19 data made us dance all the time.</p>\n<p>This is how the dance usually starts:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F8f5e0d4db7e67a5b8be34154a5aec045%2Fdate_decision.png?generation=1707186463417991&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fda6b05fb71bb02fe6a68127d575385a1%2Fmonth.png?generation=1707186477264626&amp;alt=media\"></p>\n<p>The drop of observations starting from March/April 2020 makes modeling tricky. Btw, can you imagine it’s been 4 years already since then. I am curious why this period was selected by the host and not something more recent. </p>",
  "messages": [
    {
      "id": "2637939",
      "postDate": "02/06/2024 02:32:27",
      "content": "<p>Hi, </p>\n<p>I've just started looking into the data for this competition and it's both frightening  and exciting (especially number of features to select from).<br>\nI've been doing finance and ecltv modeling work for quite a while and the covid19 data made us dance all the time.</p>\n<p>This is how the dance usually starts:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F8f5e0d4db7e67a5b8be34154a5aec045%2Fdate_decision.png?generation=1707186463417991&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fda6b05fb71bb02fe6a68127d575385a1%2Fmonth.png?generation=1707186477264626&amp;alt=media\"></p>\n<p>The drop of observations starting from March/April 2020 makes modeling tricky. Btw, can you imagine it’s been 4 years already since then. I am curious why this period was selected by the host and not something more recent. </p>",
      "rawMarkdown": "Hi, \n\nI've just started looking into the data for this competition and it's both frightening  and exciting (especially number of features to select from).\nI've been doing finance and ecltv modeling work for quite a while and the covid19 data made us dance all the time.\n\nThis is how the dance usually starts:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F8f5e0d4db7e67a5b8be34154a5aec045%2Fdate_decision.png?generation=1707186463417991&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fda6b05fb71bb02fe6a68127d575385a1%2Fmonth.png?generation=1707186477264626&alt=media)\n\nThe drop of observations starting from March/April 2020 makes modeling tricky. Btw, can you imagine it’s been 4 years already since then. I am curious why this period was selected by the host and not something more recent.",
      "votes": null
    },
    {
      "id": "2638056",
      "postDate": "02/06/2024 04:26:38",
      "content": "<p>I almost thought we were about to go through another pandemic after reading the title of this discussion, until i opened the discussions page, please change your title mate😭</p>",
      "rawMarkdown": "I almost thought we were about to go through another pandemic after reading the title of this discussion, until i opened the discussions page, please change your title mate😭",
      "votes": null
    },
    {
      "id": "2638631",
      "postDate": "02/06/2024 12:11:55",
      "content": "<p>Since they are looking for scorecards 'stable' for much longer than a year, they probably wanted to keep data for more years for evaluation.</p>",
      "rawMarkdown": "Since they are looking for scorecards 'stable' for much longer than a year, they probably wanted to keep data for more years for evaluation.",
      "votes": null
    },
    {
      "id": "2638692",
      "postDate": "02/06/2024 12:40:11",
      "content": "<p>Did they reveal how the public / private are split ? over time ?</p>",
      "rawMarkdown": "Did they reveal how the public / private are split ? over time ?",
      "votes": null
    },
    {
      "id": "2638764",
      "postDate": "02/06/2024 13:22:48",
      "content": "<p>Hi Sergey,<br>\nlet me briefly comment:</p>\n<ul>\n<li><p>since we want to focus on stability of performance, our evaluation sample covers quite long time period, and therefore we also used quite old data. Moreover, we measure risk quality of client, for that we need to observe several payments/installments, which cause also delay (we can't use data for example from last month)</p></li>\n<li><p>it's fair to say that covid had huge impact on financial industry and Home Credit was not exception. There were changes in quality of incoming population, out internal processes (e.g. shift between offline and online sales channels), sales productivity (lockdowns and other restrictions).</p></li>\n<li><p>we wanted to include into public dataset both pre-covid and covid period, so you can observe the difference in data.</p></li>\n</ul>",
      "rawMarkdown": "Hi Sergey,\nlet me briefly comment:\n\n- since we want to focus on stability of performance, our evaluation sample covers quite long time period, and therefore we also used quite old data. Moreover, we measure risk quality of client, for that we need to observe several payments/installments, which cause also delay (we can't use data for example from last month)\n\n- it's fair to say that covid had huge impact on financial industry and Home Credit was not exception. There were changes in quality of incoming population, out internal processes (e.g. shift between offline and online sales channels), sales productivity (lockdowns and other restrictions).\n\n- we wanted to include into public dataset both pre-covid and covid period, so you can observe the difference in data.",
      "votes": null
    },
    {
      "id": "2638827",
      "postDate": "02/06/2024 13:56:36",
      "content": "<p><a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> thank you for your prompt feedback. </p>\n<ul>\n<li><p>That makes a perfect sense for me that you wait on some features to age in order to collect them. It is a common practice. Then, I'd like to ask and address <a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a> question, what's the public/private time period split?</p></li>\n<li><p>My intention is to build stable and reliable model which will generalize well. So our goals are align here. Though I am worried about the \"stability\" optimizing this part of the metric:<br>\n<strong>$$-0.5 \\cdot std(\\text{residuals})$$</strong></p></li>\n</ul>\n<p>So what we are trying to achieve here - both accurate predictions and small variance. If the test public/private split encompasses period of 2022 till the end of <code>covid-19</code> only , then we will just do everything possible to overfit to <code>covid-19</code> period by introducing dummy pandemic features and adding some <code>covid-19</code> data. As a result, the generalization properties and overall \"stability\" of the models will be questionable.</p>\n<p>I really hope for the best, and I want a final solution to benefit Home Credit and its customers.</p>",
      "rawMarkdown": "tomasjeline2 thank you for your prompt feedback. \n\n- That makes a perfect sense for me that you wait on some features to age in order to collect them. It is a common practice. Then, I'd like to ask and address @lucasmorin question, what's the public/private time period split?\n\n- My intention is to build stable and reliable model which will generalize well. So our goals are align here. Though I am worried about the \"stability\" optimizing this part of the metric:\n**$$-0.5 \\cdot std(\\text{residuals})$$**\n\nSo what we are trying to achieve here - both accurate predictions and small variance. If the test public/private split encompasses period of 2022 till the end of `covid-19` only , then we will just do everything possible to overfit to `covid-19` period by introducing dummy pandemic features and adding some `covid-19` data. As a result, the generalization properties and overall \"stability\" of the models will be questionable.\n\nI really hope for the best, and I want a final solution to benefit Home Credit and its customers.",
      "votes": null
    },
    {
      "id": "2638923",
      "postDate": "02/06/2024 15:03:35",
      "content": "<p>No, we didn't reveal that and we won't plan on doing it.</p>",
      "rawMarkdown": "No, we didn't reveal that and we won't plan on doing it.",
      "votes": null
    },
    {
      "id": "2638942",
      "postDate": "02/06/2024 15:16:41",
      "content": "<blockquote>\n  <p>My intention is to build stable and reliable model which will generalize well. </p>\n</blockquote>\n<p>That is what we would like to see in this competition :)</p>\n<blockquote>\n  <p>If the test public/private split encompasses period of 2022 till the end of covid-19 only , then we will just do everything possible to overfit to covid-19 period by introducing dummy pandemic features and adding some covid-19 data. As a result, the generalization properties and overall \"stability\" of the models will be questionable.</p>\n</blockquote>\n<p>We won't reveal how exactly the public/private split was done, it is part of the challenge. I can tell you however the following</p>\n<ol>\n<li>The private sample contains some covid and some post-covid data.</li>\n<li>The public sample contains some covid and some post-covid data.</li>\n</ol>\n<p>For now we won't be providing more details on the split. </p>",
      "rawMarkdown": ">My intention is to build stable and reliable model which will generalize well. \n\nThat is what we would like to see in this competition :)\n\n>If the test public/private split encompasses period of 2022 till the end of covid-19 only , then we will just do everything possible to overfit to covid-19 period by introducing dummy pandemic features and adding some covid-19 data. As a result, the generalization properties and overall \"stability\" of the models will be questionable.\n\nWe won't reveal how exactly the public/private split was done, it is part of the challenge. I can tell you however the following\n1. The private sample contains some covid and some post-covid data.\n1. The public sample contains some covid and some post-covid data.\n\nFor now we won't be providing more details on the split.",
      "votes": null
    },
    {
      "id": "2658021",
      "postDate": "02/18/2024 23:08:13",
      "content": "<p>Hi Daniel, what's the split time point of covid and post-covid?</p>",
      "rawMarkdown": "Hi Daniel, what's the split time point of covid and post-covid?",
      "votes": null
    },
    {
      "id": "2658035",
      "postDate": "02/18/2024 23:19:54",
      "content": "<p>Assuming MONTH is the origination date, will the private or public sample contain collateral originating in 2021,2022 or 2023? It is stated in the competition that MONTH should be used for aggregation purposes..however, many are using it for prediction purposes (a problem!!)</p>",
      "rawMarkdown": "Assuming MONTH is the origination date, will the private or public sample contain collateral originating in 2021,2022 or 2023? It is stated in the competition that MONTH should be used for aggregation purposes..however, many are using it for prediction purposes (a problem!!)",
      "votes": null
    },
    {
      "id": "2658266",
      "postDate": "02/19/2024 05:22:09",
      "content": "<p>I was thinking of submitting two models for end evaluation. One low-key overfitting the covid part. The other pre-covid </p>",
      "rawMarkdown": "I was thinking of submitting two models for end evaluation. One low-key overfitting the covid part. The other pre-covid",
      "votes": null
    },
    {
      "id": "2658618",
      "postDate": "02/19/2024 09:44:11",
      "content": "<p>Sean, I agree with you. It's one of the topics we are discussing internally…</p>",
      "rawMarkdown": "Sean, I agree with you. It's one of the topics we are discussing internally...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2638056,
      "author_name": "varunmanojgupta",
      "author_url": "",
      "post_date": "02/06/2024 04:26:38",
      "content": "<p>I almost thought we were about to go through another pandemic after reading the title of this discussion, until i opened the discussions page, please change your title mate😭</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2638631,
      "author_name": "kwatrakartik",
      "author_url": "",
      "post_date": "02/06/2024 12:11:55",
      "content": "<p>Since they are looking for scorecards 'stable' for much longer than a year, they probably wanted to keep data for more years for evaluation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2638692,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "02/06/2024 12:40:11",
      "content": "<p>Did they reveal how the public / private are split ? over time ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2638923,
          "author_name": "jetakow",
          "author_url": "",
          "post_date": "02/06/2024 15:03:35",
          "content": "<p>No, we didn't reveal that and we won't plan on doing it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2638764,
      "author_name": "tomasjeline2",
      "author_url": "",
      "post_date": "02/06/2024 13:22:48",
      "content": "<p>Hi Sergey,<br>\nlet me briefly comment:</p>\n<ul>\n<li><p>since we want to focus on stability of performance, our evaluation sample covers quite long time period, and therefore we also used quite old data. Moreover, we measure risk quality of client, for that we need to observe several payments/installments, which cause also delay (we can't use data for example from last month)</p></li>\n<li><p>it's fair to say that covid had huge impact on financial industry and Home Credit was not exception. There were changes in quality of incoming population, out internal processes (e.g. shift between offline and online sales channels), sales productivity (lockdowns and other restrictions).</p></li>\n<li><p>we wanted to include into public dataset both pre-covid and covid period, so you can observe the difference in data.</p></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2638827,
          "author_name": "sergiosaharovskiy",
          "author_url": "",
          "post_date": "02/06/2024 13:56:36",
          "content": "<p><a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> thank you for your prompt feedback. </p>\n<ul>\n<li><p>That makes a perfect sense for me that you wait on some features to age in order to collect them. It is a common practice. Then, I'd like to ask and address <a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a> question, what's the public/private time period split?</p></li>\n<li><p>My intention is to build stable and reliable model which will generalize well. So our goals are align here. Though I am worried about the \"stability\" optimizing this part of the metric:<br>\n<strong>$$-0.5 \\cdot std(\\text{residuals})$$</strong></p></li>\n</ul>\n<p>So what we are trying to achieve here - both accurate predictions and small variance. If the test public/private split encompasses period of 2022 till the end of <code>covid-19</code> only , then we will just do everything possible to overfit to <code>covid-19</code> period by introducing dummy pandemic features and adding some <code>covid-19</code> data. As a result, the generalization properties and overall \"stability\" of the models will be questionable.</p>\n<p>I really hope for the best, and I want a final solution to benefit Home Credit and its customers.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2638942,
              "author_name": "jetakow",
              "author_url": "",
              "post_date": "02/06/2024 15:16:41",
              "content": "<blockquote>\n  <p>My intention is to build stable and reliable model which will generalize well. </p>\n</blockquote>\n<p>That is what we would like to see in this competition :)</p>\n<blockquote>\n  <p>If the test public/private split encompasses period of 2022 till the end of covid-19 only , then we will just do everything possible to overfit to covid-19 period by introducing dummy pandemic features and adding some covid-19 data. As a result, the generalization properties and overall \"stability\" of the models will be questionable.</p>\n</blockquote>\n<p>We won't reveal how exactly the public/private split was done, it is part of the challenge. I can tell you however the following</p>\n<ol>\n<li>The private sample contains some covid and some post-covid data.</li>\n<li>The public sample contains some covid and some post-covid data.</li>\n</ol>\n<p>For now we won't be providing more details on the split. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2658021,
                  "author_name": "makeli",
                  "author_url": "",
                  "post_date": "02/18/2024 23:08:13",
                  "content": "<p>Hi Daniel, what's the split time point of covid and post-covid?</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2658035,
                  "author_name": "seanmacmaghnusa",
                  "author_url": "",
                  "post_date": "02/18/2024 23:19:54",
                  "content": "<p>Assuming MONTH is the origination date, will the private or public sample contain collateral originating in 2021,2022 or 2023? It is stated in the competition that MONTH should be used for aggregation purposes..however, many are using it for prediction purposes (a problem!!)</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2658618,
                      "author_name": "tomasjeline2",
                      "author_url": "",
                      "post_date": "02/19/2024 09:44:11",
                      "content": "<p>Sean, I agree with you. It's one of the topics we are discussing internally…</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2658266,
      "author_name": "skaarface",
      "author_url": "",
      "post_date": "02/19/2024 05:22:09",
      "content": "<p>I was thinking of submitting two models for end evaluation. One low-key overfitting the covid part. The other pre-covid </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2637939": "Hi, \n\nI've just started looking into the data for this competition and it's both frightening  and exciting (especially number of features to select from).\nI've been doing finance and ecltv modeling work for quite a while and the covid19 data made us dance all the time.\n\nThis is how the dance usually starts:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F8f5e0d4db7e67a5b8be34154a5aec045%2Fdate_decision.png?generation=1707186463417991&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fda6b05fb71bb02fe6a68127d575385a1%2Fmonth.png?generation=1707186477264626&alt=media)\n\nThe drop of observations starting from March/April 2020 makes modeling tricky. Btw, can you imagine it’s been 4 years already since then. I am curious why this period was selected by the host and not something more recent.",
    "2638056": "I almost thought we were about to go through another pandemic after reading the title of this discussion, until i opened the discussions page, please change your title mate😭",
    "2638631": "Since they are looking for scorecards 'stable' for much longer than a year, they probably wanted to keep data for more years for evaluation.",
    "2638692": "Did they reveal how the public / private are split ? over time ?",
    "2638764": "Hi Sergey,\nlet me briefly comment:\n\n- since we want to focus on stability of performance, our evaluation sample covers quite long time period, and therefore we also used quite old data. Moreover, we measure risk quality of client, for that we need to observe several payments/installments, which cause also delay (we can't use data for example from last month)\n\n- it's fair to say that covid had huge impact on financial industry and Home Credit was not exception. There were changes in quality of incoming population, out internal processes (e.g. shift between offline and online sales channels), sales productivity (lockdowns and other restrictions).\n\n- we wanted to include into public dataset both pre-covid and covid period, so you can observe the difference in data.",
    "2638827": "tomasjeline2 thank you for your prompt feedback. \n\n- That makes a perfect sense for me that you wait on some features to age in order to collect them. It is a common practice. Then, I'd like to ask and address @lucasmorin question, what's the public/private time period split?\n\n- My intention is to build stable and reliable model which will generalize well. So our goals are align here. Though I am worried about the \"stability\" optimizing this part of the metric:\n**$$-0.5 \\cdot std(\\text{residuals})$$**\n\nSo what we are trying to achieve here - both accurate predictions and small variance. If the test public/private split encompasses period of 2022 till the end of `covid-19` only , then we will just do everything possible to overfit to `covid-19` period by introducing dummy pandemic features and adding some `covid-19` data. As a result, the generalization properties and overall \"stability\" of the models will be questionable.\n\nI really hope for the best, and I want a final solution to benefit Home Credit and its customers.",
    "2638923": "No, we didn't reveal that and we won't plan on doing it.",
    "2638942": ">My intention is to build stable and reliable model which will generalize well. \n\nThat is what we would like to see in this competition :)\n\n>If the test public/private split encompasses period of 2022 till the end of covid-19 only , then we will just do everything possible to overfit to covid-19 period by introducing dummy pandemic features and adding some covid-19 data. As a result, the generalization properties and overall \"stability\" of the models will be questionable.\n\nWe won't reveal how exactly the public/private split was done, it is part of the challenge. I can tell you however the following\n1. The private sample contains some covid and some post-covid data.\n1. The public sample contains some covid and some post-covid data.\n\nFor now we won't be providing more details on the split.",
    "2658021": "Hi Daniel, what's the split time point of covid and post-covid?",
    "2658035": "Assuming MONTH is the origination date, will the private or public sample contain collateral originating in 2021,2022 or 2023? It is stated in the competition that MONTH should be used for aggregation purposes..however, many are using it for prediction purposes (a problem!!)",
    "2658266": "I was thinking of submitting two models for end evaluation. One low-key overfitting the covid part. The other pre-covid",
    "2658618": "Sean, I agree with you. It's one of the topics we are discussing internally..."
  },
  "source": "meta"
}