{
  "id": 478360,
  "title": "Has anyone managed to add Credit Bureau A data successfully?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/478360",
  "author_name": "",
  "post_date": "2024-02-20T12:41:37.661144100Z",
  "votes": 13,
  "comment_count": 15,
  "views": 0,
  "content": "<p>For me, adding Credit Bureau A data results in large improvement of VAL AUC (from 0.833 to 0.854), but also a large decrease in LB GINI (from 0.538 to 0.493). I show the journey of my models on the chart below.</p>\n<p>My validation framework is taking weeks until 60 to train, and since week 60 to validate. Before adding Credit Bureau A it had good correlation of Val scores &amp; Public LB.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F494641%2F5d2bbcf85d73681fba83d9a660ee34e4%2FLB%20GINI%20vs%20VAL%20AUC%20(2).png?generation=1708432678576006&amp;alt=media\"></p>",
  "messages": [
    {
      "id": "2660195",
      "postDate": "02/20/2024 12:41:37",
      "content": "<p>For me, adding Credit Bureau A data results in large improvement of VAL AUC (from 0.833 to 0.854), but also a large decrease in LB GINI (from 0.538 to 0.493). I show the journey of my models on the chart below.</p>\n<p>My validation framework is taking weeks until 60 to train, and since week 60 to validate. Before adding Credit Bureau A it had good correlation of Val scores &amp; Public LB.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F494641%2F5d2bbcf85d73681fba83d9a660ee34e4%2FLB%20GINI%20vs%20VAL%20AUC%20(2).png?generation=1708432678576006&amp;alt=media\"></p>",
      "rawMarkdown": "For me, adding Credit Bureau A data results in large improvement of VAL AUC (from 0.833 to 0.854), but also a large decrease in LB GINI (from 0.538 to 0.493). I show the journey of my models on the chart below.\n\nMy validation framework is taking weeks until 60 to train, and since week 60 to validate. Before adding Credit Bureau A it had good correlation of Val scores & Public LB.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F494641%2F5d2bbcf85d73681fba83d9a660ee34e4%2FLB%20GINI%20vs%20VAL%20AUC%20(2).png?generation=1708432678576006&alt=media)",
      "votes": null
    },
    {
      "id": "2660538",
      "postDate": "02/20/2024 16:55:02",
      "content": "<p>Hello, I'm not sure that your CV strategy is the best one for stability, even if you follow (if I understood) a kind of timeseries validation process (train with previous months and validation on the last months). Maybe the future new metric will prove your right 😉</p>",
      "rawMarkdown": "Hello, I'm not sure that your CV strategy is the best one for stability, even if you follow (if I understood) a kind of timeseries validation process (train with previous months and validation on the last months). Maybe the future new metric will prove your right 😉",
      "votes": null
    },
    {
      "id": "2660576",
      "postDate": "02/20/2024 17:18:45",
      "content": "<p>Nope. I spent a bit of time trying to figure out if there were any features in particular that were causing the overfit and tried dropping the ones I suspected when submitting, but I didn't find anything. That's hardly exhaustive or conclusive though. My guess for now is that this is the data they're primarily referring to when they say \"It's worth noting that some external data providers might not be available for future (test) evaluations, which is anticipated\". </p>",
      "rawMarkdown": "Nope. I spent a bit of time trying to figure out if there were any features in particular that were causing the overfit and tried dropping the ones I suspected when submitting, but I didn't find anything. That's hardly exhaustive or conclusive though. My guess for now is that this is the data they're primarily referring to when they say \"It's worth noting that some external data providers might not be available for future (test) evaluations, which is anticipated\".",
      "votes": null
    },
    {
      "id": "2660678",
      "postDate": "02/20/2024 18:55:26",
      "content": "<p>Could you ellaborate on how you loaded the corresponding df on kaggle? I kept getting memory errors, even after downcasting. In my local CV I observed increase similar to yours.</p>",
      "rawMarkdown": "Could you ellaborate on how you loaded the corresponding df on kaggle? I kept getting memory errors, even after downcasting. In my local CV I observed increase similar to yours.",
      "votes": null
    },
    {
      "id": "2660754",
      "postDate": "02/20/2024 19:59:27",
      "content": "<blockquote>\n  <p>\"It's worth noting that some external data providers might not be available for future (test) evaluations, which is anticipated\".</p>\n</blockquote>\n<p>This is my running hypothesis too.</p>",
      "rawMarkdown": "> \"It's worth noting that some external data providers might not be available for future (test) evaluations, which is anticipated\".\n\nThis is my running hypothesis too.",
      "votes": null
    },
    {
      "id": "2660757",
      "postDate": "02/20/2024 20:01:34",
      "content": "<p>I do a couple of tricks</p>\n<p>1/ I process this table in a separate script, saving intermediary results onto disk. This releases all memory after processing is finished.<br>\n2/ I map categorical values: there are quite a few categorical columns which have like 20-30 unique values but are hashed to something ridiculously big like <code>f3dhj32</code></p>",
      "rawMarkdown": "I do a couple of tricks\n\n1/ I process this table in a separate script, saving intermediary results onto disk. This releases all memory after processing is finished.\n2/ I map categorical values: there are quite a few categorical columns which have like 20-30 unique values but are hashed to something ridiculously big like `f3dhj32`",
      "votes": null
    },
    {
      "id": "2660760",
      "postDate": "02/20/2024 20:03:34",
      "content": "<p>My running hypothesis is that such data is simply unavailable for the future, as hosts quote here</p>\n<blockquote>\n  <p>\"It's worth noting that some external data providers might not be available for future (test) evaluations, which is anticipated\".</p>\n</blockquote>\n<p>So what happens is that LGBM uses quite a few splits to model a relationship, and then encounters mostly <code>nan</code> values during inference.</p>",
      "rawMarkdown": "My running hypothesis is that such data is simply unavailable for the future, as hosts quote here\n> \"It's worth noting that some external data providers might not be available for future (test) evaluations, which is anticipated\".\n\nSo what happens is that LGBM uses quite a few splits to model a relationship, and then encounters mostly `nan` values during inference.",
      "votes": null
    },
    {
      "id": "2660793",
      "postDate": "02/20/2024 20:18:47",
      "content": "<p>Hey I just confirmed that my notebook runs inference successfully when exchanging the test with train set. I think that is indicator that I don’t run into memory issues (what I suspected until now). So I guess I need to save the schema of train file as a pkl and than use that to convert the cols of test during inference. I suspect there are some columns where dtypes are differently  inferred compared to train when reading the parquet from test. Did you make similar experience?</p>\n<p>Edit: even after setting the dtypes exactly as the ones in the train df the error still persists. It only occurred with the train bureau a 1 table. If anyone else has similar error and manages to resolve it I would be thankful for letting me know.</p>",
      "rawMarkdown": "Hey I just confirmed that my notebook runs inference successfully when exchanging the test with train set. I think that is indicator that I don’t run into memory issues (what I suspected until now). So I guess I need to save the schema of train file as a pkl and than use that to convert the cols of test during inference. I suspect there are some columns where dtypes are differently  inferred compared to train when reading the parquet from test. Did you make similar experience?\n\nEdit: even after setting the dtypes exactly as the ones in the train df the error still persists. It only occurred with the train bureau a 1 table. If anyone else has similar error and manages to resolve it I would be thankful for letting me know.",
      "votes": null
    },
    {
      "id": "2661539",
      "postDate": "02/21/2024 10:21:28",
      "content": "<p>I remember I had some issues related to dtypes but decided to leave them altogether :) I haven't encountered them with Credit Bureau A data though.</p>",
      "rawMarkdown": "I remember I had some issues related to dtypes but decided to leave them altogether :) I haven't encountered them with Credit Bureau A data though.",
      "votes": null
    },
    {
      "id": "2661645",
      "postDate": "02/21/2024 11:56:28",
      "content": "<p><a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> I will be obliged to know why you decided with a hold out set involving the last 30+ weeks as the Dev-set. Why do you think this is better than a stratified group K-fold?</p>",
      "rawMarkdown": "narsil I will be obliged to know why you decided with a hold out set involving the last 30+ weeks as the Dev-set. Why do you think this is better than a stratified group K-fold?",
      "votes": null
    },
    {
      "id": "2661691",
      "postDate": "02/21/2024 12:41:06",
      "content": "<p>\"It's worth noting that some external data providers might not be available for future (test) evaluations, which is anticipated\".</p>\n<p>This can be tested: in submission, see what % of case_id's are in Credit Bureau A data, and if it is &lt; X,  crash.</p>\n<p>For comparison, in train data Credit Bureau A data contains .908 of all case_id's.</p>",
      "rawMarkdown": "\"It's worth noting that some external data providers might not be available for future (test) evaluations, which is anticipated\".\n\nThis can be tested: in submission, see what % of case_id's are in Credit Bureau A data, and if it is < X,  crash.\n\nFor comparison, in train data Credit Bureau A data contains .908 of all case_id's.",
      "votes": null
    },
    {
      "id": "2661704",
      "postDate": "02/21/2024 13:05:09",
      "content": "<p>i ran some tests, and test data for Credit Bureau A contains between .7 and .8 of all case_id's. So most of the cases do have this data available.</p>\n<p>But perhaps this data is missing in later weeks, causing downward slope in AUC and reduced metric. This can also be tested.</p>",
      "rawMarkdown": "i ran some tests, and test data for Credit Bureau A contains between .7 and .8 of all case_id's. So most of the cases do have this data available.\n\nBut perhaps this data is missing in later weeks, causing downward slope in AUC and reduced metric. This can also be tested.",
      "votes": null
    },
    {
      "id": "2661781",
      "postDate": "02/21/2024 13:57:10",
      "content": "<p>The question is also about the volume of the data - how many records in Credit Bureau AU per case id.</p>",
      "rawMarkdown": "The question is also about the volume of the data - how many records in Credit Bureau AU per case id.",
      "votes": null
    },
    {
      "id": "2661908",
      "postDate": "02/21/2024 15:27:32",
      "content": "<p>Excuse me what do you mean by leave them altogether?</p>",
      "rawMarkdown": "Excuse me what do you mean by leave them altogether?",
      "votes": null
    },
    {
      "id": "2662003",
      "postDate": "02/21/2024 16:30:00",
      "content": "<p>Don't care about them, don't use them for the time being, wait until a Public Notebook arrives with a solution ;)</p>",
      "rawMarkdown": "Don't care about them, don't use them for the time being, wait until a Public Notebook arrives with a solution ;)",
      "votes": null
    },
    {
      "id": "2664087",
      "postDate": "02/22/2024 19:26:35",
      "content": "<p>This is what I do next. For each file (chunk) I apply aggregation like min, and after concatenation I apply aggregation again because min(min(chunk_1), min(chunk_2), …, min(chunk_n)) = min(chunk_1 + chunk_2 + … + chunk_n). See in my <a href=\"https://www.kaggle.com/code/andreynesterov/home-credit-baseline-data\" target=\"_blank\">notebook </a></p>",
      "rawMarkdown": "This is what I do next. For each file (chunk) I apply aggregation like min, and after concatenation I apply aggregation again because min(min(chunk_1), min(chunk_2), ..., min(chunk_n)) = min(chunk_1 + chunk_2 + ... + chunk_n). See in my [notebook ](https://www.kaggle.com/code/andreynesterov/home-credit-baseline-data)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2660538,
      "author_name": "pourchot",
      "author_url": "",
      "post_date": "02/20/2024 16:55:02",
      "content": "<p>Hello, I'm not sure that your CV strategy is the best one for stability, even if you follow (if I understood) a kind of timeseries validation process (train with previous months and validation on the last months). Maybe the future new metric will prove your right 😉</p>",
      "votes": null,
      "replies": [
        {
          "id": 2660760,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "02/20/2024 20:03:34",
          "content": "<p>My running hypothesis is that such data is simply unavailable for the future, as hosts quote here</p>\n<blockquote>\n  <p>\"It's worth noting that some external data providers might not be available for future (test) evaluations, which is anticipated\".</p>\n</blockquote>\n<p>So what happens is that LGBM uses quite a few splits to model a relationship, and then encounters mostly <code>nan</code> values during inference.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2661691,
              "author_name": "ymatioun",
              "author_url": "",
              "post_date": "02/21/2024 12:41:06",
              "content": "<p>\"It's worth noting that some external data providers might not be available for future (test) evaluations, which is anticipated\".</p>\n<p>This can be tested: in submission, see what % of case_id's are in Credit Bureau A data, and if it is &lt; X,  crash.</p>\n<p>For comparison, in train data Credit Bureau A data contains .908 of all case_id's.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2661704,
                  "author_name": "ymatioun",
                  "author_url": "",
                  "post_date": "02/21/2024 13:05:09",
                  "content": "<p>i ran some tests, and test data for Credit Bureau A contains between .7 and .8 of all case_id's. So most of the cases do have this data available.</p>\n<p>But perhaps this data is missing in later weeks, causing downward slope in AUC and reduced metric. This can also be tested.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2661781,
                      "author_name": "narsil",
                      "author_url": "",
                      "post_date": "02/21/2024 13:57:10",
                      "content": "<p>The question is also about the volume of the data - how many records in Credit Bureau AU per case id.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2660576,
      "author_name": "brandenkmurray",
      "author_url": "",
      "post_date": "02/20/2024 17:18:45",
      "content": "<p>Nope. I spent a bit of time trying to figure out if there were any features in particular that were causing the overfit and tried dropping the ones I suspected when submitting, but I didn't find anything. That's hardly exhaustive or conclusive though. My guess for now is that this is the data they're primarily referring to when they say \"It's worth noting that some external data providers might not be available for future (test) evaluations, which is anticipated\". </p>",
      "votes": null,
      "replies": [
        {
          "id": 2660754,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "02/20/2024 19:59:27",
          "content": "<blockquote>\n  <p>\"It's worth noting that some external data providers might not be available for future (test) evaluations, which is anticipated\".</p>\n</blockquote>\n<p>This is my running hypothesis too.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2660678,
      "author_name": "simonveitner",
      "author_url": "",
      "post_date": "02/20/2024 18:55:26",
      "content": "<p>Could you ellaborate on how you loaded the corresponding df on kaggle? I kept getting memory errors, even after downcasting. In my local CV I observed increase similar to yours.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2660757,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "02/20/2024 20:01:34",
          "content": "<p>I do a couple of tricks</p>\n<p>1/ I process this table in a separate script, saving intermediary results onto disk. This releases all memory after processing is finished.<br>\n2/ I map categorical values: there are quite a few categorical columns which have like 20-30 unique values but are hashed to something ridiculously big like <code>f3dhj32</code></p>",
          "votes": null,
          "replies": [
            {
              "id": 2660793,
              "author_name": "simonveitner",
              "author_url": "",
              "post_date": "02/20/2024 20:18:47",
              "content": "<p>Hey I just confirmed that my notebook runs inference successfully when exchanging the test with train set. I think that is indicator that I don’t run into memory issues (what I suspected until now). So I guess I need to save the schema of train file as a pkl and than use that to convert the cols of test during inference. I suspect there are some columns where dtypes are differently  inferred compared to train when reading the parquet from test. Did you make similar experience?</p>\n<p>Edit: even after setting the dtypes exactly as the ones in the train df the error still persists. It only occurred with the train bureau a 1 table. If anyone else has similar error and manages to resolve it I would be thankful for letting me know.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2661539,
                  "author_name": "narsil",
                  "author_url": "",
                  "post_date": "02/21/2024 10:21:28",
                  "content": "<p>I remember I had some issues related to dtypes but decided to leave them altogether :) I haven't encountered them with Credit Bureau A data though.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2661908,
                      "author_name": "simonveitner",
                      "author_url": "",
                      "post_date": "02/21/2024 15:27:32",
                      "content": "<p>Excuse me what do you mean by leave them altogether?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2662003,
                          "author_name": "narsil",
                          "author_url": "",
                          "post_date": "02/21/2024 16:30:00",
                          "content": "<p>Don't care about them, don't use them for the time being, wait until a Public Notebook arrives with a solution ;)</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 2664087,
          "author_name": "andreynesterov",
          "author_url": "",
          "post_date": "02/22/2024 19:26:35",
          "content": "<p>This is what I do next. For each file (chunk) I apply aggregation like min, and after concatenation I apply aggregation again because min(min(chunk_1), min(chunk_2), …, min(chunk_n)) = min(chunk_1 + chunk_2 + … + chunk_n). See in my <a href=\"https://www.kaggle.com/code/andreynesterov/home-credit-baseline-data\" target=\"_blank\">notebook </a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2661645,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "02/21/2024 11:56:28",
      "content": "<p><a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> I will be obliged to know why you decided with a hold out set involving the last 30+ weeks as the Dev-set. Why do you think this is better than a stratified group K-fold?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2660195": "For me, adding Credit Bureau A data results in large improvement of VAL AUC (from 0.833 to 0.854), but also a large decrease in LB GINI (from 0.538 to 0.493). I show the journey of my models on the chart below.\n\nMy validation framework is taking weeks until 60 to train, and since week 60 to validate. Before adding Credit Bureau A it had good correlation of Val scores & Public LB.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F494641%2F5d2bbcf85d73681fba83d9a660ee34e4%2FLB%20GINI%20vs%20VAL%20AUC%20(2).png?generation=1708432678576006&alt=media)",
    "2660538": "Hello, I'm not sure that your CV strategy is the best one for stability, even if you follow (if I understood) a kind of timeseries validation process (train with previous months and validation on the last months). Maybe the future new metric will prove your right 😉",
    "2660576": "Nope. I spent a bit of time trying to figure out if there were any features in particular that were causing the overfit and tried dropping the ones I suspected when submitting, but I didn't find anything. That's hardly exhaustive or conclusive though. My guess for now is that this is the data they're primarily referring to when they say \"It's worth noting that some external data providers might not be available for future (test) evaluations, which is anticipated\".",
    "2660678": "Could you ellaborate on how you loaded the corresponding df on kaggle? I kept getting memory errors, even after downcasting. In my local CV I observed increase similar to yours.",
    "2660754": "> \"It's worth noting that some external data providers might not be available for future (test) evaluations, which is anticipated\".\n\nThis is my running hypothesis too.",
    "2660757": "I do a couple of tricks\n\n1/ I process this table in a separate script, saving intermediary results onto disk. This releases all memory after processing is finished.\n2/ I map categorical values: there are quite a few categorical columns which have like 20-30 unique values but are hashed to something ridiculously big like `f3dhj32`",
    "2660760": "My running hypothesis is that such data is simply unavailable for the future, as hosts quote here\n> \"It's worth noting that some external data providers might not be available for future (test) evaluations, which is anticipated\".\n\nSo what happens is that LGBM uses quite a few splits to model a relationship, and then encounters mostly `nan` values during inference.",
    "2660793": "Hey I just confirmed that my notebook runs inference successfully when exchanging the test with train set. I think that is indicator that I don’t run into memory issues (what I suspected until now). So I guess I need to save the schema of train file as a pkl and than use that to convert the cols of test during inference. I suspect there are some columns where dtypes are differently  inferred compared to train when reading the parquet from test. Did you make similar experience?\n\nEdit: even after setting the dtypes exactly as the ones in the train df the error still persists. It only occurred with the train bureau a 1 table. If anyone else has similar error and manages to resolve it I would be thankful for letting me know.",
    "2661539": "I remember I had some issues related to dtypes but decided to leave them altogether :) I haven't encountered them with Credit Bureau A data though.",
    "2661645": "narsil I will be obliged to know why you decided with a hold out set involving the last 30+ weeks as the Dev-set. Why do you think this is better than a stratified group K-fold?",
    "2661691": "\"It's worth noting that some external data providers might not be available for future (test) evaluations, which is anticipated\".\n\nThis can be tested: in submission, see what % of case_id's are in Credit Bureau A data, and if it is < X,  crash.\n\nFor comparison, in train data Credit Bureau A data contains .908 of all case_id's.",
    "2661704": "i ran some tests, and test data for Credit Bureau A contains between .7 and .8 of all case_id's. So most of the cases do have this data available.\n\nBut perhaps this data is missing in later weeks, causing downward slope in AUC and reduced metric. This can also be tested.",
    "2661781": "The question is also about the volume of the data - how many records in Credit Bureau AU per case id.",
    "2661908": "Excuse me what do you mean by leave them altogether?",
    "2662003": "Don't care about them, don't use them for the time being, wait until a Public Notebook arrives with a solution ;)",
    "2664087": "This is what I do next. For each file (chunk) I apply aggregation like min, and after concatenation I apply aggregation again because min(min(chunk_1), min(chunk_2), ..., min(chunk_n)) = min(chunk_1 + chunk_2 + ... + chunk_n). See in my [notebook ](https://www.kaggle.com/code/andreynesterov/home-credit-baseline-data)"
  },
  "source": "meta"
}