{
  "id": 477075,
  "title": "Beware that the null count has a trend!",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/477075",
  "author_name": "",
  "post_date": "2024-02-14T15:35:41.707553100Z",
  "votes": 30,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I am looking at the null count per week. There are quite a lot of features <br>\n1) have 100% null rate in the first week<br>\n2) have 100% null rate in the last week<br>\n3) have a consecutive 100% null rate across multiple weeks.</p>\n<p>I only looked into the 2 static tables. Here are some features might have a trendy null rate:</p>\n<pre><code>[', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ']\n</code></pre>\n<p>For example <code>amtinstpaidbefduel24m_4187115A</code> (Number of instalments paid before due date in the last 24 months.)<br>\nhas 100% null rate in the first 10 weeks! and the null ratio is gradually decreasing over weeks:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2F9b3c50646386aad6511871bcb8ff40d7%2FScreenshot%20from%202024-02-14%2023-28-56.png?generation=1707924557438357&amp;alt=media\"></p>\n<p><code>clientscnt_136L</code> is opposite:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2F8ef1c742aa3f5384d59c8fedbb5da159%2FScreenshot%20from%202024-02-14%2023-31-17.png?generation=1707924686407129&amp;alt=media\"></p>\n<p>Some ideas:<br>\n1) for features that have consecutive null in the last few weeks, we should drop them completely.<br>\n2) for features that have consecutive null in the first few weeks, we can either drop them or find some ways to impute them (might be from other tables)<br>\n3) Only train on the latest few weeks that have fewer null issues…</p>",
  "messages": [
    {
      "id": "2652154",
      "postDate": "02/14/2024 15:35:41",
      "content": "<p>I am looking at the null count per week. There are quite a lot of features <br>\n1) have 100% null rate in the first week<br>\n2) have 100% null rate in the last week<br>\n3) have a consecutive 100% null rate across multiple weeks.</p>\n<p>I only looked into the 2 static tables. Here are some features might have a trendy null rate:</p>\n<pre><code>[', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ', ']\n</code></pre>\n<p>For example <code>amtinstpaidbefduel24m_4187115A</code> (Number of instalments paid before due date in the last 24 months.)<br>\nhas 100% null rate in the first 10 weeks! and the null ratio is gradually decreasing over weeks:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2F9b3c50646386aad6511871bcb8ff40d7%2FScreenshot%20from%202024-02-14%2023-28-56.png?generation=1707924557438357&amp;alt=media\"></p>\n<p><code>clientscnt_136L</code> is opposite:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2F8ef1c742aa3f5384d59c8fedbb5da159%2FScreenshot%20from%202024-02-14%2023-31-17.png?generation=1707924686407129&amp;alt=media\"></p>\n<p>Some ideas:<br>\n1) for features that have consecutive null in the last few weeks, we should drop them completely.<br>\n2) for features that have consecutive null in the first few weeks, we can either drop them or find some ways to impute them (might be from other tables)<br>\n3) Only train on the latest few weeks that have fewer null issues…</p>",
      "rawMarkdown": "I am looking at the null count per week. There are quite a lot of features \n1) have 100% null rate in the first week\n2) have 100% null rate in the last week\n3) have a consecutive 100% null rate across multiple weeks.\n\nI only looked into the 2 static tables. Here are some features might have a trendy null rate:\n```\n['amtinstpaidbefduel24m_4187115A', 'avgdbddpdlast3m_4187120P', 'avgdbdtollast24m_4525197P', 'avglnamtstart24m_4525187A', 'avgoutstandbalancel6m_4187114A', 'avgpmtlast12m_4525200A', 'clientscnt_136L', 'datelastinstal40dpd_247D', 'dtlastpmtallstes_4499206D', 'interestrategrace_34L', 'maxannuity_4075009A', 'maxdbddpdtollast6m_4187119P', 'maxlnamtstart6m_4525199A', 'maxoutstandbalancel12m_4187113A', 'maxpmtlast3m_4525190A', 'mindbdtollast24m_4525191P', 'numinstlswithdpd5_4187116L', 'numinstmatpaidtearly2d_4499204L', 'numinstpaid_4499208L', 'numinstpaidearly3dest_4493216L', 'numinstpaidearly5dest_4493211L', 'numinstpaidearly5dobd_4499205L', 'numinstpaidearlyest_4493214L', 'numinstpaidlastcontr_4325080L', 'numinstregularpaidest_4493210L', 'numinsttopaygrest_4493213L', 'numinstunpaidmaxest_4493212L', 'payvacationpostpone_4187118D', 'sumoutstandtotalest_4493215A', 'totinstallast1m_4525188A', 'validfrom_1069D', 'assignmentdate_238D', 'assignmentdate_4527235D', 'assignmentdate_4955616D', 'birthdate_574D', 'contractssum_5085716L', 'pmtaverage_3A', 'pmtaverage_4527227A', 'pmtaverage_4955615A', 'pmtcount_4527229L', 'pmtcount_4955617L', 'pmtcount_693L', 'pmtscount_423L', 'pmtssum_45A', 'requesttype_4525192L', 'responsedate_1012D', 'responsedate_4527233D', 'responsedate_4917613D']\n```\n\nFor example `amtinstpaidbefduel24m_4187115A` (Number of instalments paid before due date in the last 24 months.)\nhas 100% null rate in the first 10 weeks! and the null ratio is gradually decreasing over weeks:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2F9b3c50646386aad6511871bcb8ff40d7%2FScreenshot%20from%202024-02-14%2023-28-56.png?generation=1707924557438357&alt=media)\n\n`clientscnt_136L` is opposite:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2F8ef1c742aa3f5384d59c8fedbb5da159%2FScreenshot%20from%202024-02-14%2023-31-17.png?generation=1707924686407129&alt=media)\n\n\nSome ideas:\n1) for features that have consecutive null in the last few weeks, we should drop them completely.\n2) for features that have consecutive null in the first few weeks, we can either drop them or find some ways to impute them (might be from other tables)\n3) Only train on the latest few weeks that have fewer null issues...",
      "votes": null
    },
    {
      "id": "2653085",
      "postDate": "02/15/2024 07:25:13",
      "content": "<p>I tried to remove features with 100% null rate in the last 2 weeks here<br>\n<a href=\"https://www.kaggle.com/code/kingychiu/dropping-deprecated-features\" target=\"_blank\">https://www.kaggle.com/code/kingychiu/dropping-deprecated-features</a></p>\n<p>but the lb changed from 0.559 (original) to 0.558 (dropped)…</p>\n<p>I guess the null rate is something we need to guess or prompt in the test set…</p>\n<p>The data quality is far lower than what a usual Kaggle competition is. Need to do some more EDA.</p>",
      "rawMarkdown": "I tried to remove features with 100% null rate in the last 2 weeks here\nhttps://www.kaggle.com/code/kingychiu/dropping-deprecated-features\n\nbut the lb changed from 0.559 (original) to 0.558 (dropped)...\n\nI guess the null rate is something we need to guess or prompt in the test set...\n\nThe data quality is far lower than what a usual Kaggle competition is. Need to do some more EDA.",
      "votes": null
    },
    {
      "id": "2653103",
      "postDate": "02/15/2024 07:51:02",
      "content": "<p><a href=\"https://www.kaggle.com/kingychiu\" target=\"_blank\">@kingychiu</a> are those trends there in depth=0 tables as well ? </p>",
      "rawMarkdown": "kingychiu are those trends there in depth=0 tables as well ?",
      "votes": null
    },
    {
      "id": "2653105",
      "postDate": "02/15/2024 07:58:52",
      "content": "<p>Yes. I am only working with the 0-depth table. Don't want to be too overwhelmed haha…</p>",
      "rawMarkdown": "Yes. I am only working with the 0-depth table. Don't want to be too overwhelmed haha...",
      "votes": null
    },
    {
      "id": "2654801",
      "postDate": "02/16/2024 12:45:07",
      "content": "<p>Interesting! thanks for sharing!</p>\n<p>In case it helps you,  I have published a nb showing missing values from all tables <a href=\"https://www.kaggle.com/code/imeintanis/home-credit-2024-eda-missing-values\" target=\"_blank\">link</a> . </p>\n<p>eg. in <code>clientscnt_136L</code> only 417 rows out of 1.5m contain num values.. <br>\nhowever rest of  <code>clientscnt_*</code> seem better.  </p>\n<p>Any suggestions on better visualization are welcome! </p>",
      "rawMarkdown": "Interesting! thanks for sharing!\n\nIn case it helps you,  I have published a nb showing missing values from all tables [link](https://www.kaggle.com/code/imeintanis/home-credit-2024-eda-missing-values) . \n\neg. in `clientscnt_136L` only 417 rows out of 1.5m contain num values.. \nhowever rest of  `clientscnt_*` seem better.  \n\nAny suggestions on better visualization are welcome!",
      "votes": null
    },
    {
      "id": "2654823",
      "postDate": "02/16/2024 13:06:49",
      "content": "<p>Why not just average the missing values from a reasonable window.  The approach might can be susceptible to variance. But I do think on certain reasonable features that tend to have low variance and follows a trend, itwould make it beneficial </p>",
      "rawMarkdown": "Why not just average the missing values from a reasonable window.  The approach might can be susceptible to variance. But I do think on certain reasonable features that tend to have low variance and follows a trend, itwould make it beneficial",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2653085,
      "author_name": "kingychiu",
      "author_url": "",
      "post_date": "02/15/2024 07:25:13",
      "content": "<p>I tried to remove features with 100% null rate in the last 2 weeks here<br>\n<a href=\"https://www.kaggle.com/code/kingychiu/dropping-deprecated-features\" target=\"_blank\">https://www.kaggle.com/code/kingychiu/dropping-deprecated-features</a></p>\n<p>but the lb changed from 0.559 (original) to 0.558 (dropped)…</p>\n<p>I guess the null rate is something we need to guess or prompt in the test set…</p>\n<p>The data quality is far lower than what a usual Kaggle competition is. Need to do some more EDA.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2653103,
      "author_name": "davutpolat",
      "author_url": "",
      "post_date": "02/15/2024 07:51:02",
      "content": "<p><a href=\"https://www.kaggle.com/kingychiu\" target=\"_blank\">@kingychiu</a> are those trends there in depth=0 tables as well ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2653105,
          "author_name": "kingychiu",
          "author_url": "",
          "post_date": "02/15/2024 07:58:52",
          "content": "<p>Yes. I am only working with the 0-depth table. Don't want to be too overwhelmed haha…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2654801,
      "author_name": "imeintanis",
      "author_url": "",
      "post_date": "02/16/2024 12:45:07",
      "content": "<p>Interesting! thanks for sharing!</p>\n<p>In case it helps you,  I have published a nb showing missing values from all tables <a href=\"https://www.kaggle.com/code/imeintanis/home-credit-2024-eda-missing-values\" target=\"_blank\">link</a> . </p>\n<p>eg. in <code>clientscnt_136L</code> only 417 rows out of 1.5m contain num values.. <br>\nhowever rest of  <code>clientscnt_*</code> seem better.  </p>\n<p>Any suggestions on better visualization are welcome! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2654823,
      "author_name": "skaarface",
      "author_url": "",
      "post_date": "02/16/2024 13:06:49",
      "content": "<p>Why not just average the missing values from a reasonable window.  The approach might can be susceptible to variance. But I do think on certain reasonable features that tend to have low variance and follows a trend, itwould make it beneficial </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2652154": "I am looking at the null count per week. There are quite a lot of features \n1) have 100% null rate in the first week\n2) have 100% null rate in the last week\n3) have a consecutive 100% null rate across multiple weeks.\n\nI only looked into the 2 static tables. Here are some features might have a trendy null rate:\n```\n['amtinstpaidbefduel24m_4187115A', 'avgdbddpdlast3m_4187120P', 'avgdbdtollast24m_4525197P', 'avglnamtstart24m_4525187A', 'avgoutstandbalancel6m_4187114A', 'avgpmtlast12m_4525200A', 'clientscnt_136L', 'datelastinstal40dpd_247D', 'dtlastpmtallstes_4499206D', 'interestrategrace_34L', 'maxannuity_4075009A', 'maxdbddpdtollast6m_4187119P', 'maxlnamtstart6m_4525199A', 'maxoutstandbalancel12m_4187113A', 'maxpmtlast3m_4525190A', 'mindbdtollast24m_4525191P', 'numinstlswithdpd5_4187116L', 'numinstmatpaidtearly2d_4499204L', 'numinstpaid_4499208L', 'numinstpaidearly3dest_4493216L', 'numinstpaidearly5dest_4493211L', 'numinstpaidearly5dobd_4499205L', 'numinstpaidearlyest_4493214L', 'numinstpaidlastcontr_4325080L', 'numinstregularpaidest_4493210L', 'numinsttopaygrest_4493213L', 'numinstunpaidmaxest_4493212L', 'payvacationpostpone_4187118D', 'sumoutstandtotalest_4493215A', 'totinstallast1m_4525188A', 'validfrom_1069D', 'assignmentdate_238D', 'assignmentdate_4527235D', 'assignmentdate_4955616D', 'birthdate_574D', 'contractssum_5085716L', 'pmtaverage_3A', 'pmtaverage_4527227A', 'pmtaverage_4955615A', 'pmtcount_4527229L', 'pmtcount_4955617L', 'pmtcount_693L', 'pmtscount_423L', 'pmtssum_45A', 'requesttype_4525192L', 'responsedate_1012D', 'responsedate_4527233D', 'responsedate_4917613D']\n```\n\nFor example `amtinstpaidbefduel24m_4187115A` (Number of instalments paid before due date in the last 24 months.)\nhas 100% null rate in the first 10 weeks! and the null ratio is gradually decreasing over weeks:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2F9b3c50646386aad6511871bcb8ff40d7%2FScreenshot%20from%202024-02-14%2023-28-56.png?generation=1707924557438357&alt=media)\n\n`clientscnt_136L` is opposite:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2F8ef1c742aa3f5384d59c8fedbb5da159%2FScreenshot%20from%202024-02-14%2023-31-17.png?generation=1707924686407129&alt=media)\n\n\nSome ideas:\n1) for features that have consecutive null in the last few weeks, we should drop them completely.\n2) for features that have consecutive null in the first few weeks, we can either drop them or find some ways to impute them (might be from other tables)\n3) Only train on the latest few weeks that have fewer null issues...",
    "2653085": "I tried to remove features with 100% null rate in the last 2 weeks here\nhttps://www.kaggle.com/code/kingychiu/dropping-deprecated-features\n\nbut the lb changed from 0.559 (original) to 0.558 (dropped)...\n\nI guess the null rate is something we need to guess or prompt in the test set...\n\nThe data quality is far lower than what a usual Kaggle competition is. Need to do some more EDA.",
    "2653103": "kingychiu are those trends there in depth=0 tables as well ?",
    "2653105": "Yes. I am only working with the 0-depth table. Don't want to be too overwhelmed haha...",
    "2654801": "Interesting! thanks for sharing!\n\nIn case it helps you,  I have published a nb showing missing values from all tables [link](https://www.kaggle.com/code/imeintanis/home-credit-2024-eda-missing-values) . \n\neg. in `clientscnt_136L` only 417 rows out of 1.5m contain num values.. \nhowever rest of  `clientscnt_*` seem better.  \n\nAny suggestions on better visualization are welcome!",
    "2654823": "Why not just average the missing values from a reasonable window.  The approach might can be susceptible to variance. But I do think on certain reasonable features that tend to have low variance and follows a trend, itwould make it beneficial"
  },
  "source": "meta"
}