{
  "id": 482392,
  "title": "Understanding the dataset (depths, num_groupN)",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/482392",
  "author_name": "",
  "post_date": "2024-03-07T15:06:04.920465200Z",
  "votes": 9,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n<p>I am struggling to get a grasp and comprehension of the dataset. I find it very confusing and not explained well in the \"Overview\" section. Specifically, I don't understand the idea of the depths. What do the depths mean? What do you mean by historical records? There are multiple tables for depth=1 and multiple tables for depth=2. Does the num_group1 have meaning across the tables?</p>\n<p>For example let's see the file 'train_applprev_1_0.csv'. First of all, what does the table even represent? There is more than 2 millions of unique \"case_id\", but there are 20 unique \"num_group1\" values. It looks like the \"case_id\" values are categorized. But what do the categories represent?</p>\n<p>Another example is file \"train_person_1.csv\". What exactly does the num_group1 means in this table? Single loan application can have multiple rows and can be connected with several birthdates. How is that possible? It is standard practice for a loan to be issued to a single individual.</p>\n<p>Can anybody explain that to me? Thank you.</p>",
  "messages": [
    {
      "id": "2685988",
      "postDate": "03/07/2024 15:06:04",
      "content": "<p>Hello everyone,</p>\n<p>I am struggling to get a grasp and comprehension of the dataset. I find it very confusing and not explained well in the \"Overview\" section. Specifically, I don't understand the idea of the depths. What do the depths mean? What do you mean by historical records? There are multiple tables for depth=1 and multiple tables for depth=2. Does the num_group1 have meaning across the tables?</p>\n<p>For example let's see the file 'train_applprev_1_0.csv'. First of all, what does the table even represent? There is more than 2 millions of unique \"case_id\", but there are 20 unique \"num_group1\" values. It looks like the \"case_id\" values are categorized. But what do the categories represent?</p>\n<p>Another example is file \"train_person_1.csv\". What exactly does the num_group1 means in this table? Single loan application can have multiple rows and can be connected with several birthdates. How is that possible? It is standard practice for a loan to be issued to a single individual.</p>\n<p>Can anybody explain that to me? Thank you.</p>",
      "rawMarkdown": "Hello everyone,\n\nI am struggling to get a grasp and comprehension of the dataset. I find it very confusing and not explained well in the \"Overview\" section. Specifically, I don't understand the idea of the depths. What do the depths mean? What do you mean by historical records? There are multiple tables for depth=1 and multiple tables for depth=2. Does the num_group1 have meaning across the tables?\n\nFor example let's see the file 'train_applprev_1_0.csv'. First of all, what does the table even represent? There is more than 2 millions of unique \"case_id\", but there are 20 unique \"num_group1\" values. It looks like the \"case_id\" values are categorized. But what do the categories represent?\n\nAnother example is file \"train_person_1.csv\". What exactly does the num_group1 means in this table? Single loan application can have multiple rows and can be connected with several birthdates. How is that possible? It is standard practice for a loan to be issued to a single individual.\n\nCan anybody explain that to me? Thank you.",
      "votes": null
    },
    {
      "id": "2686046",
      "postDate": "03/07/2024 15:49:08",
      "content": "<p>This is apparently a table with the client's historical loans(train_applprev_1_0). One client historically could have many loans from 1..N. (num_group1,num_group2…num_groupN)<br>\nYou can calculate  different aggregation.<br>\nFor example: </p>\n<ul>\n<li>number of loans before the current application, </li>\n<li>maximum amount loan received</li>\n<li>maximum day overdue <br>\n….</li>\n</ul>",
      "rawMarkdown": "This is apparently a table with the client's historical loans(train_applprev_1_0). One client historically could have many loans from 1..N. (num_group1,num_group2...num_groupN)\nYou can calculate  different aggregation.\nFor example: \n- number of loans before the current application, \n- maximum amount loan received\n- maximum day overdue \n....",
      "votes": null
    },
    {
      "id": "2687721",
      "postDate": "03/08/2024 18:04:07",
      "content": "<p>it might help to think of it as a hierarchy, although it isn't really that way in the dataframes.  That's the logic of it though</p>\n<p>See if this example helps.  This is a depth 2 dataframe.</p>\n<p>case_id is just a normal index, then for each case ID, there can be multiple categories of credit account type which are categoricals in num_group1.   In turn, for each of the credit account types in num_group1, there can be multiple transactions that are in num_group2.  </p>\n<p>For example, num</p>\n<ul>\n<li>case_id: Unique identifier for a loan application.<ul>\n<li>num_group1 (Credit Account Type):<ul>\n<li>1: Credit Card</li>\n<li>2: Mortgage<ul>\n<li>num_group2 (Transaction Sequence within Account):<ul>\n<li>For Credit Card (num_group1 = 1):<ul>\n<li>1: January Transaction</li>\n<li>2: February Transaction</li></ul></li>\n<li>For Mortgage (num_group1 = 2):<ul>\n<li>1: January Payment</li>\n<li>2: February Payment</li></ul></li></ul></li></ul></li></ul></li></ul></li>\n</ul>",
      "rawMarkdown": "it might help to think of it as a hierarchy, although it isn't really that way in the dataframes.  That's the logic of it though\n\nSee if this example helps.  This is a depth 2 dataframe.\n\ncase_id is just a normal index, then for each case ID, there can be multiple categories of credit account type which are categoricals in num_group1.   In turn, for each of the credit account types in num_group1, there can be multiple transactions that are in num_group2.  \n\nFor example, num\n\n* case_id: Unique identifier for a loan application.\n\t* num_group1 (Credit Account Type):\n\t    * 1: Credit Card\n\t    * 2: Mortgage\n\t\t\t* num_group2 (Transaction Sequence within Account):\n\t\t\t\t* For Credit Card (num_group1 = 1):\n\t\t\t\t    * 1: January Transaction\n\t\t\t\t\t* 2: February Transaction\n\t\t\t\t* For Mortgage (num_group1 = 2):\n\t\t\t\t    * 1: January Payment\n\t\t\t\t    * 2: February Payment",
      "votes": null
    },
    {
      "id": "2746418",
      "postDate": "04/11/2024 07:51:24",
      "content": "<p>In train_applprev_1_0.csv,waht means the type of value of feature  status_219L? A,D,K,T…..</p>",
      "rawMarkdown": "In train_applprev_1_0.csv,waht means the type of value of feature  status_219L? A,D,K,T.....",
      "votes": null
    },
    {
      "id": "2752034",
      "postDate": "04/14/2024 17:10:38",
      "content": "<p>Hey man thanks you so much for your good explanation, I may have a little more question. <br>How is it different from the old credit risk model in term of stability (Credit risk model vs Credit risk model Stability). <br>Is the new model stable because we just have to write more robust code for long term use? is that it? <br>\nI just want to know why it call stability (I know it for longterm use whatsoever but I just don't see the different between the old model and the new one) Why don't just we use old model --&gt; change the input to fit this new data --&gt;  called it stability </p>",
      "rawMarkdown": "Hey man thanks you so much for your good explanation, I may have a little more question. <br>How is it different from the old credit risk model in term of stability (Credit risk model vs Credit risk model Stability). <br>Is the new model stable because we just have to write more robust code for long term use? is that it? \nI just want to know why it call stability (I know it for longterm use whatsoever but I just don't see the different between the old model and the new one) Why don't just we use old model --> change the input to fit this new data -->  called it stability",
      "votes": null
    },
    {
      "id": "2752913",
      "postDate": "04/15/2024 08:11:38",
      "content": "<p>There is different evaluation function. </p>\n<p>First competition use AUC or gini - goal was to have max \"average\" gini, and it didn't matter whether there were periods where model was underperforming.</p>\n<p>Instability in performance is issue in production, which we want to address in this competition. Besides gini there is penalization for instability in evaluation metric. Therefore, I suspect that trying to maximize gini without taking into account stability is not optimal strategy in this competition.</p>",
      "rawMarkdown": "There is different evaluation function. \n\nFirst competition use AUC or gini - goal was to have max \"average\" gini, and it didn't matter whether there were periods where model was underperforming.\n\nInstability in performance is issue in production, which we want to address in this competition. Besides gini there is penalization for instability in evaluation metric. Therefore, I suspect that trying to maximize gini without taking into account stability is not optimal strategy in this competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2686046,
      "author_name": "dima1992",
      "author_url": "",
      "post_date": "03/07/2024 15:49:08",
      "content": "<p>This is apparently a table with the client's historical loans(train_applprev_1_0). One client historically could have many loans from 1..N. (num_group1,num_group2…num_groupN)<br>\nYou can calculate  different aggregation.<br>\nFor example: </p>\n<ul>\n<li>number of loans before the current application, </li>\n<li>maximum amount loan received</li>\n<li>maximum day overdue <br>\n….</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2746418,
          "author_name": "sunyangd",
          "author_url": "",
          "post_date": "04/11/2024 07:51:24",
          "content": "<p>In train_applprev_1_0.csv,waht means the type of value of feature  status_219L? A,D,K,T…..</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2687721,
      "author_name": "jeremyboccabello",
      "author_url": "",
      "post_date": "03/08/2024 18:04:07",
      "content": "<p>it might help to think of it as a hierarchy, although it isn't really that way in the dataframes.  That's the logic of it though</p>\n<p>See if this example helps.  This is a depth 2 dataframe.</p>\n<p>case_id is just a normal index, then for each case ID, there can be multiple categories of credit account type which are categoricals in num_group1.   In turn, for each of the credit account types in num_group1, there can be multiple transactions that are in num_group2.  </p>\n<p>For example, num</p>\n<ul>\n<li>case_id: Unique identifier for a loan application.<ul>\n<li>num_group1 (Credit Account Type):<ul>\n<li>1: Credit Card</li>\n<li>2: Mortgage<ul>\n<li>num_group2 (Transaction Sequence within Account):<ul>\n<li>For Credit Card (num_group1 = 1):<ul>\n<li>1: January Transaction</li>\n<li>2: February Transaction</li></ul></li>\n<li>For Mortgage (num_group1 = 2):<ul>\n<li>1: January Payment</li>\n<li>2: February Payment</li></ul></li></ul></li></ul></li></ul></li></ul></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2752034,
          "author_name": "jackkitchai",
          "author_url": "",
          "post_date": "04/14/2024 17:10:38",
          "content": "<p>Hey man thanks you so much for your good explanation, I may have a little more question. <br>How is it different from the old credit risk model in term of stability (Credit risk model vs Credit risk model Stability). <br>Is the new model stable because we just have to write more robust code for long term use? is that it? <br>\nI just want to know why it call stability (I know it for longterm use whatsoever but I just don't see the different between the old model and the new one) Why don't just we use old model --&gt; change the input to fit this new data --&gt;  called it stability </p>",
          "votes": null,
          "replies": [
            {
              "id": 2752913,
              "author_name": "tomasjeline2",
              "author_url": "",
              "post_date": "04/15/2024 08:11:38",
              "content": "<p>There is different evaluation function. </p>\n<p>First competition use AUC or gini - goal was to have max \"average\" gini, and it didn't matter whether there were periods where model was underperforming.</p>\n<p>Instability in performance is issue in production, which we want to address in this competition. Besides gini there is penalization for instability in evaluation metric. Therefore, I suspect that trying to maximize gini without taking into account stability is not optimal strategy in this competition.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2685988": "Hello everyone,\n\nI am struggling to get a grasp and comprehension of the dataset. I find it very confusing and not explained well in the \"Overview\" section. Specifically, I don't understand the idea of the depths. What do the depths mean? What do you mean by historical records? There are multiple tables for depth=1 and multiple tables for depth=2. Does the num_group1 have meaning across the tables?\n\nFor example let's see the file 'train_applprev_1_0.csv'. First of all, what does the table even represent? There is more than 2 millions of unique \"case_id\", but there are 20 unique \"num_group1\" values. It looks like the \"case_id\" values are categorized. But what do the categories represent?\n\nAnother example is file \"train_person_1.csv\". What exactly does the num_group1 means in this table? Single loan application can have multiple rows and can be connected with several birthdates. How is that possible? It is standard practice for a loan to be issued to a single individual.\n\nCan anybody explain that to me? Thank you.",
    "2686046": "This is apparently a table with the client's historical loans(train_applprev_1_0). One client historically could have many loans from 1..N. (num_group1,num_group2...num_groupN)\nYou can calculate  different aggregation.\nFor example: \n- number of loans before the current application, \n- maximum amount loan received\n- maximum day overdue \n....",
    "2687721": "it might help to think of it as a hierarchy, although it isn't really that way in the dataframes.  That's the logic of it though\n\nSee if this example helps.  This is a depth 2 dataframe.\n\ncase_id is just a normal index, then for each case ID, there can be multiple categories of credit account type which are categoricals in num_group1.   In turn, for each of the credit account types in num_group1, there can be multiple transactions that are in num_group2.  \n\nFor example, num\n\n* case_id: Unique identifier for a loan application.\n\t* num_group1 (Credit Account Type):\n\t    * 1: Credit Card\n\t    * 2: Mortgage\n\t\t\t* num_group2 (Transaction Sequence within Account):\n\t\t\t\t* For Credit Card (num_group1 = 1):\n\t\t\t\t    * 1: January Transaction\n\t\t\t\t\t* 2: February Transaction\n\t\t\t\t* For Mortgage (num_group1 = 2):\n\t\t\t\t    * 1: January Payment\n\t\t\t\t    * 2: February Payment",
    "2746418": "In train_applprev_1_0.csv,waht means the type of value of feature  status_219L? A,D,K,T.....",
    "2752034": "Hey man thanks you so much for your good explanation, I may have a little more question. <br>How is it different from the old credit risk model in term of stability (Credit risk model vs Credit risk model Stability). <br>Is the new model stable because we just have to write more robust code for long term use? is that it? \nI just want to know why it call stability (I know it for longterm use whatsoever but I just don't see the different between the old model and the new one) Why don't just we use old model --> change the input to fit this new data -->  called it stability",
    "2752913": "There is different evaluation function. \n\nFirst competition use AUC or gini - goal was to have max \"average\" gini, and it didn't matter whether there were periods where model was underperforming.\n\nInstability in performance is issue in production, which we want to address in this competition. Besides gini there is penalization for instability in evaluation metric. Therefore, I suspect that trying to maximize gini without taking into account stability is not optimal strategy in this competition."
  },
  "source": "meta"
}