{
  "id": 490612,
  "title": "When does num_groupN change?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/490612",
  "author_name": "",
  "post_date": "2024-04-03T00:32:02.152313800Z",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>When do num_group1 and num_group2 change?</p>\n<p>I want to know exactly the standards by which they change.</p>",
  "messages": [
    {
      "id": "2731989",
      "postDate": "04/03/2024 00:32:02",
      "content": "<p>When do num_group1 and num_group2 change?</p>\n<p>I want to know exactly the standards by which they change.</p>",
      "rawMarkdown": "When do num_group1 and num_group2 change?\n\nI want to know exactly the standards by which they change.",
      "votes": null
    },
    {
      "id": "2732482",
      "postDate": "04/03/2024 06:57:29",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dirhq88\" target=\"_blank\">@dirhq88</a>,<br>\nI am not sure what you mean by change of num_group1 and num_group2 - could you explain more?</p>",
      "rawMarkdown": "Hi @dirhq88,\nI am not sure what you mean by change of num_group1 and num_group2 - could you explain more?",
      "votes": null
    },
    {
      "id": "2734465",
      "postDate": "04/04/2024 06:21:06",
      "content": "<p>I used the word 'change' because num_group means historical records of case_id.</p>\n<p>Does num_group1 increase by 1 when a major change in a particular case occurs and num_group2 increase by 1 when a minor change occurs?<br>\n(I know that the smaller the number of num_groupN, the more current condition.)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19837433%2F00dcb99e8be245bdf21a8f28a7ccbb2a%2F2024-04-04%20%202.34.31.png?generation=1712208893616616&amp;alt=media\"></p>\n<p>In the above example, the values of num_group1 in case_id = 6 are divided into 0, 1. And when num_group1 = 0, the value of num_group2 is divided into 0,1, and when num_group1 = 1, the value of num_group2 is divided into 0,1,2,3,4,5.</p>\n<p>In case_id=6, can it be seen that major changes occur 2 times and minor changes occur 2 and 5 times?</p>",
      "rawMarkdown": "I used the word 'change' because num_group means historical records of case_id.\n\nDoes num_group1 increase by 1 when a major change in a particular case occurs and num_group2 increase by 1 when a minor change occurs?\n(I know that the smaller the number of num_groupN, the more current condition.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19837433%2F00dcb99e8be245bdf21a8f28a7ccbb2a%2F2024-04-04%20%202.34.31.png?generation=1712208893616616&alt=media)\n\nIn the above example, the values of num_group1 in case_id = 6 are divided into 0, 1. And when num_group1 = 0, the value of num_group2 is divided into 0,1, and when num_group1 = 1, the value of num_group2 is divided into 0,1,2,3,4,5.\n\nIn case_id=6, can it be seen that major changes occur 2 times and minor changes occur 2 and 5 times?",
      "votes": null
    },
    {
      "id": "2734617",
      "postDate": "04/04/2024 07:54:45",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dirhq88\" target=\"_blank\">@dirhq88</a> </p>\n<p><code>because num_group means historical records of case_id</code>: this is not correct interpretation. Num_group is index for data with depth 1 or 2. Or to describe it differently, we have data with different granularity that are not aggregated on the level of case_id, for example previous applications, or contact references. It has nothing to do with changes in the data.</p>",
      "rawMarkdown": "Hi @dirhq88 \n\n`because num_group means historical records of case_id`: this is not correct interpretation. Num_group is index for data with depth 1 or 2. Or to describe it differently, we have data with different granularity that are not aggregated on the level of case_id, for example previous applications, or contact references. It has nothing to do with changes in the data.",
      "votes": null
    },
    {
      "id": "2736163",
      "postDate": "04/05/2024 04:13:58",
      "content": "<p>I have two question.</p>\n<p>\"num_group1 - This is an indexing column used for the historical records of case_id in both depth=1 and depth=2 tables.\"<br>\nWhat is the exact meaning of 'historical' here?</p>\n<p>\"Or to describe it differently, we have data with different granularity that are not aggregated on the level of case_id, for example previous applications, or contact references.\"<br>\nI don't understand your answer. Can you explain it in more detail?</p>",
      "rawMarkdown": "I have two question.\n\n\"num_group1 - This is an indexing column used for the historical records of case_id in both depth=1 and depth=2 tables.\"\nWhat is the exact meaning of 'historical' here?\n\n\"Or to describe it differently, we have data with different granularity that are not aggregated on the level of case_id, for example previous applications, or contact references.\"\nI don't understand your answer. Can you explain it in more detail?",
      "votes": null
    },
    {
      "id": "2736418",
      "postDate": "04/05/2024 07:24:26",
      "content": "<p>Hi,<br>\nokay, let me explain on hypothetical example:</p>\n<ul>\n<li>Tomas Jelinek applied for loan on 1.1.2024 - this is credit case, which has assigned unique case_id</li>\n<li>Tomas Jelinek is existing client in Home Credit, it means he had applications/loans with Home Credit before 1.1.2024, let's say 5 loans. Data describing parameters of those loans, their repayment history etc. are definitely valuable for credit scoring, therefore you have them in the sample. But they are not aggregated on level of case_id, there are 5 rows describing those 5 previous loans. To differentiate between those 5, you have to have some index - num_group_1, which will contain values 0,1,2,3,4…</li>\n<li>num_group_1 is not used only for previous applications, but for other data where we have several records per case_id, like contact references, records in credit registry, etc. What I want to say is there is different meaning for different tables &amp; attributes</li>\n<li>is some cases we might have even bigger detail, for example information about instalments for each previous application. Then you need num_group_2, for example let's say previous loan with num_group_1=0 have 3 instalments, then you will have 3 records with num_group_1 = 0 and num_group_2 = 0,1,2</li>\n</ul>",
      "rawMarkdown": "Hi,\nokay, let me explain on hypothetical example:\n\n- Tomas Jelinek applied for loan on 1.1.2024 - this is credit case, which has assigned unique case_id\n- Tomas Jelinek is existing client in Home Credit, it means he had applications/loans with Home Credit before 1.1.2024, let's say 5 loans. Data describing parameters of those loans, their repayment history etc. are definitely valuable for credit scoring, therefore you have them in the sample. But they are not aggregated on level of case_id, there are 5 rows describing those 5 previous loans. To differentiate between those 5, you have to have some index - num_group_1, which will contain values 0,1,2,3,4...\n- num_group_1 is not used only for previous applications, but for other data where we have several records per case_id, like contact references, records in credit registry, etc. What I want to say is there is different meaning for different tables & attributes\n- is some cases we might have even bigger detail, for example information about instalments for each previous application. Then you need num_group_2, for example let's say previous loan with num_group_1=0 have 3 instalments, then you will have 3 records with num_group_1 = 0 and num_group_2 = 0,1,2",
      "votes": null
    },
    {
      "id": "2736619",
      "postDate": "04/05/2024 09:48:00",
      "content": "<p>If it may help:<br>\nIn Depth=1 files - each case_id has a history of activities/features related to it - that is num_group1.<br>\nIn Depth=2 files - each num_group1 has a sub-history of activities/features related to them - that is num_group2<br>\nIn conclusion - you can groupby case_id and aggregate these values to have some representation of the sub-history in one single value. In the example notebook the host shared - they took the max value as aggregation method.</p>",
      "rawMarkdown": "If it may help:\nIn Depth=1 files - each case_id has a history of activities/features related to it - that is num_group1.\nIn Depth=2 files - each num_group1 has a sub-history of activities/features related to them - that is num_group2\nIn conclusion - you can groupby case_id and aggregate these values to have some representation of the sub-history in one single value. In the example notebook the host shared - they took the max value as aggregation method.",
      "votes": null
    },
    {
      "id": "2739898",
      "postDate": "04/07/2024 11:46:30",
      "content": "<p>Hi,I have a question:<br>\n     Are num_group1=2's records closer to the present than num_group1=1's, or they don't have such relations? </p>",
      "rawMarkdown": "Hi,I have a question:\n     Are num_group1=2's records closer to the present than num_group1=1's, or they don't have such relations?",
      "votes": null
    },
    {
      "id": "2741215",
      "postDate": "04/08/2024 07:48:45",
      "content": "<p>Hi,<br>\nyes, data should be sorted by num_group - timewise, position on application form,…</p>",
      "rawMarkdown": "Hi,\nyes, data should be sorted by num_group - timewise, position on application form,...",
      "votes": null
    },
    {
      "id": "2753193",
      "postDate": "04/15/2024 10:56:10",
      "content": "<p>I have 4 question.</p>\n<ol>\n<li>Do credit cases and loans mean the same thing?</li>\n</ol>\n<ul>\n<li>case_id: This is the unique identifier for each 'credit case'. You'll need this ID to join relevant tables to the base table.</li>\n</ul>\n<p>You mentioned that \"But they are not aggregated on level of case_id, there are 5 rows describing those 5 previous loans. To differentiate between those 5, you have to have some index - num_group_1, which will contain values 0,1,2,3,4…\". I don't understand what this mean.</p>\n<ol>\n<li><p>Since 5 previous loans also have their own case_id, the features of each case (5 previous loans) can be distinguished sufficiently by case_id, why do we need to distinguish the features through the num_group in one case_id?</p></li>\n<li><p>If the features of the previous cases are included in the sample, will the case_id of the previous cases be overwritten to the current(new) case_id?</p></li>\n<li><p>Is it possible to change the features of the application after submitting the loan application?</p></li>\n</ol>\n<p>Thank you so much :)</p>",
      "rawMarkdown": "I have 4 question.\n\n1. Do credit cases and loans mean the same thing?\n- case_id: This is the unique identifier for each 'credit case'. You'll need this ID to join relevant tables to the base table.\n\nYou mentioned that \"But they are not aggregated on level of case_id, there are 5 rows describing those 5 previous loans. To differentiate between those 5, you have to have some index - num_group_1, which will contain values 0,1,2,3,4…\". I don't understand what this mean.\n\n2. Since 5 previous loans also have their own case_id, the features of each case (5 previous loans) can be distinguished sufficiently by case_id, why do we need to distinguish the features through the num_group in one case_id?\n\n3. If the features of the previous cases are included in the sample, will the case_id of the previous cases be overwritten to the current(new) case_id?\n\n4. Is it possible to change the features of the application after submitting the loan application?\n\nThank you so much :)",
      "votes": null
    },
    {
      "id": "2753378",
      "postDate": "04/15/2024 13:10:18",
      "content": "<p>Hi,</p>\n<blockquote>\n  <p>Do credit cases and loans mean the same thing?</p>\n</blockquote>\n<p>Yes. More precisely credit case = loan application. But for Kaggle data, only applications for which we have target (info about default or non-default) are included, so in this case credit case = loan (client signed loan agreement, money were disbursed,…)</p>\n<p>ad 1.<br>\nYou need to link previous applications to current application using its case ID. Otherwise you don't have direct link.<br>\nHere is example: For case_id = 2, you have 2 records in applprev file, so for this credit case we have 2 previous loan applications, to distinguish them there is num_group1 index 0 or 1. Case_id = 2 is not id of those previous applications!<br>\nSimilarly case_id=3 has 1 previous application, case_id=6 has 3 previous applications… </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7718387%2F512f85c255915e60b65f2150d8cdf066%2Fnum_group_example.png?generation=1713185950174529&amp;alt=media\"></p>\n<p>ad 2.<br>\nLet's say my application case_id = 333 is in the datasample. And I had two previous applications, let's say case_id=222 and case_id = 111. Information about 222 and 111 are included in applprev dataset under case_id = 333, because those are historical data for client with application 333.<br>\nCase_id = 222 or case_id=111 might be also in the base table, if they are from the same time period which is covered by base table (if they are from different time period, they will not be in dataset)</p>\n<p>ad 3.<br>\nfeatures of applications are of course changing. What you have in dataset is snapshot as of time of application for particular case_id. So for example above, in applprev for case_id=333 you have info about loans/application 111 and 222 as of application date of 333.</p>",
      "rawMarkdown": "Hi,\n> Do credit cases and loans mean the same thing?\n\nYes. More precisely credit case = loan application. But for Kaggle data, only applications for which we have target (info about default or non-default) are included, so in this case credit case = loan (client signed loan agreement, money were disbursed,...)\n\nad 1.\nYou need to link previous applications to current application using its case ID. Otherwise you don't have direct link.\nHere is example: For case_id = 2, you have 2 records in applprev file, so for this credit case we have 2 previous loan applications, to distinguish them there is num_group1 index 0 or 1. Case_id = 2 is not id of those previous applications!\nSimilarly case_id=3 has 1 previous application, case_id=6 has 3 previous applications... \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7718387%2F512f85c255915e60b65f2150d8cdf066%2Fnum_group_example.png?generation=1713185950174529&alt=media)\n\nad 2.\nLet's say my application case_id = 333 is in the datasample. And I had two previous applications, let's say case_id=222 and case_id = 111. Information about 222 and 111 are included in applprev dataset under case_id = 333, because those are historical data for client with application 333.\nCase_id = 222 or case_id=111 might be also in the base table, if they are from the same time period which is covered by base table (if they are from different time period, they will not be in dataset)\n\nad 3.\nfeatures of applications are of course changing. What you have in dataset is snapshot as of time of application for particular case_id. So for example above, in applprev for case_id=333 you have info about loans/application 111 and 222 as of application date of 333.",
      "votes": null
    },
    {
      "id": "2778435",
      "postDate": "04/27/2024 06:56:50",
      "content": "<p>Thank you for your previous response! <br>\nI have a question regarding the \"applprev\" file, specifically about the skipped case IDs. For instance, there are no entries for case IDs 0 and 1. Does this omission mean that the data for these skipped case IDs have been incorporated into other existing case IDs under different 'num_groups' indexes? In other words, are these missing cases examples of previous applications whose updated information has been added to (or overwritten) in existing case IDs within the \"applprev\" file?</p>",
      "rawMarkdown": "Thank you for your previous response! \nI have a question regarding the \"applprev\" file, specifically about the skipped case IDs. For instance, there are no entries for case IDs 0 and 1. Does this omission mean that the data for these skipped case IDs have been incorporated into other existing case IDs under different 'num_groups' indexes? In other words, are these missing cases examples of previous applications whose updated information has been added to (or overwritten) in existing case IDs within the \"applprev\" file?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2732482,
      "author_name": "tomasjeline2",
      "author_url": "",
      "post_date": "04/03/2024 06:57:29",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dirhq88\" target=\"_blank\">@dirhq88</a>,<br>\nI am not sure what you mean by change of num_group1 and num_group2 - could you explain more?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2734465,
          "author_name": "dirhq88",
          "author_url": "",
          "post_date": "04/04/2024 06:21:06",
          "content": "<p>I used the word 'change' because num_group means historical records of case_id.</p>\n<p>Does num_group1 increase by 1 when a major change in a particular case occurs and num_group2 increase by 1 when a minor change occurs?<br>\n(I know that the smaller the number of num_groupN, the more current condition.)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19837433%2F00dcb99e8be245bdf21a8f28a7ccbb2a%2F2024-04-04%20%202.34.31.png?generation=1712208893616616&amp;alt=media\"></p>\n<p>In the above example, the values of num_group1 in case_id = 6 are divided into 0, 1. And when num_group1 = 0, the value of num_group2 is divided into 0,1, and when num_group1 = 1, the value of num_group2 is divided into 0,1,2,3,4,5.</p>\n<p>In case_id=6, can it be seen that major changes occur 2 times and minor changes occur 2 and 5 times?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2734617,
              "author_name": "tomasjeline2",
              "author_url": "",
              "post_date": "04/04/2024 07:54:45",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/dirhq88\" target=\"_blank\">@dirhq88</a> </p>\n<p><code>because num_group means historical records of case_id</code>: this is not correct interpretation. Num_group is index for data with depth 1 or 2. Or to describe it differently, we have data with different granularity that are not aggregated on the level of case_id, for example previous applications, or contact references. It has nothing to do with changes in the data.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2736163,
                  "author_name": "dirhq88",
                  "author_url": "",
                  "post_date": "04/05/2024 04:13:58",
                  "content": "<p>I have two question.</p>\n<p>\"num_group1 - This is an indexing column used for the historical records of case_id in both depth=1 and depth=2 tables.\"<br>\nWhat is the exact meaning of 'historical' here?</p>\n<p>\"Or to describe it differently, we have data with different granularity that are not aggregated on the level of case_id, for example previous applications, or contact references.\"<br>\nI don't understand your answer. Can you explain it in more detail?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2736418,
                      "author_name": "tomasjeline2",
                      "author_url": "",
                      "post_date": "04/05/2024 07:24:26",
                      "content": "<p>Hi,<br>\nokay, let me explain on hypothetical example:</p>\n<ul>\n<li>Tomas Jelinek applied for loan on 1.1.2024 - this is credit case, which has assigned unique case_id</li>\n<li>Tomas Jelinek is existing client in Home Credit, it means he had applications/loans with Home Credit before 1.1.2024, let's say 5 loans. Data describing parameters of those loans, their repayment history etc. are definitely valuable for credit scoring, therefore you have them in the sample. But they are not aggregated on level of case_id, there are 5 rows describing those 5 previous loans. To differentiate between those 5, you have to have some index - num_group_1, which will contain values 0,1,2,3,4…</li>\n<li>num_group_1 is not used only for previous applications, but for other data where we have several records per case_id, like contact references, records in credit registry, etc. What I want to say is there is different meaning for different tables &amp; attributes</li>\n<li>is some cases we might have even bigger detail, for example information about instalments for each previous application. Then you need num_group_2, for example let's say previous loan with num_group_1=0 have 3 instalments, then you will have 3 records with num_group_1 = 0 and num_group_2 = 0,1,2</li>\n</ul>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2739898,
                          "author_name": "redddddddddddd",
                          "author_url": "",
                          "post_date": "04/07/2024 11:46:30",
                          "content": "<p>Hi,I have a question:<br>\n     Are num_group1=2's records closer to the present than num_group1=1's, or they don't have such relations? </p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2741215,
                              "author_name": "tomasjeline2",
                              "author_url": "",
                              "post_date": "04/08/2024 07:48:45",
                              "content": "<p>Hi,<br>\nyes, data should be sorted by num_group - timewise, position on application form,…</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2753193,
                                  "author_name": "dirhq88",
                                  "author_url": "",
                                  "post_date": "04/15/2024 10:56:10",
                                  "content": "<p>I have 4 question.</p>\n<ol>\n<li>Do credit cases and loans mean the same thing?</li>\n</ol>\n<ul>\n<li>case_id: This is the unique identifier for each 'credit case'. You'll need this ID to join relevant tables to the base table.</li>\n</ul>\n<p>You mentioned that \"But they are not aggregated on level of case_id, there are 5 rows describing those 5 previous loans. To differentiate between those 5, you have to have some index - num_group_1, which will contain values 0,1,2,3,4…\". I don't understand what this mean.</p>\n<ol>\n<li><p>Since 5 previous loans also have their own case_id, the features of each case (5 previous loans) can be distinguished sufficiently by case_id, why do we need to distinguish the features through the num_group in one case_id?</p></li>\n<li><p>If the features of the previous cases are included in the sample, will the case_id of the previous cases be overwritten to the current(new) case_id?</p></li>\n<li><p>Is it possible to change the features of the application after submitting the loan application?</p></li>\n</ol>\n<p>Thank you so much :)</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 2753378,
                                      "author_name": "tomasjeline2",
                                      "author_url": "",
                                      "post_date": "04/15/2024 13:10:18",
                                      "content": "<p>Hi,</p>\n<blockquote>\n  <p>Do credit cases and loans mean the same thing?</p>\n</blockquote>\n<p>Yes. More precisely credit case = loan application. But for Kaggle data, only applications for which we have target (info about default or non-default) are included, so in this case credit case = loan (client signed loan agreement, money were disbursed,…)</p>\n<p>ad 1.<br>\nYou need to link previous applications to current application using its case ID. Otherwise you don't have direct link.<br>\nHere is example: For case_id = 2, you have 2 records in applprev file, so for this credit case we have 2 previous loan applications, to distinguish them there is num_group1 index 0 or 1. Case_id = 2 is not id of those previous applications!<br>\nSimilarly case_id=3 has 1 previous application, case_id=6 has 3 previous applications… </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7718387%2F512f85c255915e60b65f2150d8cdf066%2Fnum_group_example.png?generation=1713185950174529&amp;alt=media\"></p>\n<p>ad 2.<br>\nLet's say my application case_id = 333 is in the datasample. And I had two previous applications, let's say case_id=222 and case_id = 111. Information about 222 and 111 are included in applprev dataset under case_id = 333, because those are historical data for client with application 333.<br>\nCase_id = 222 or case_id=111 might be also in the base table, if they are from the same time period which is covered by base table (if they are from different time period, they will not be in dataset)</p>\n<p>ad 3.<br>\nfeatures of applications are of course changing. What you have in dataset is snapshot as of time of application for particular case_id. So for example above, in applprev for case_id=333 you have info about loans/application 111 and 222 as of application date of 333.</p>",
                                      "votes": null,
                                      "replies": [
                                        {
                                          "id": 2778435,
                                          "author_name": "dirhq88",
                                          "author_url": "",
                                          "post_date": "04/27/2024 06:56:50",
                                          "content": "<p>Thank you for your previous response! <br>\nI have a question regarding the \"applprev\" file, specifically about the skipped case IDs. For instance, there are no entries for case IDs 0 and 1. Does this omission mean that the data for these skipped case IDs have been incorporated into other existing case IDs under different 'num_groups' indexes? In other words, are these missing cases examples of previous applications whose updated information has been added to (or overwritten) in existing case IDs within the \"applprev\" file?</p>",
                                          "votes": null,
                                          "replies": []
                                        }
                                      ]
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2736619,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "04/05/2024 09:48:00",
      "content": "<p>If it may help:<br>\nIn Depth=1 files - each case_id has a history of activities/features related to it - that is num_group1.<br>\nIn Depth=2 files - each num_group1 has a sub-history of activities/features related to them - that is num_group2<br>\nIn conclusion - you can groupby case_id and aggregate these values to have some representation of the sub-history in one single value. In the example notebook the host shared - they took the max value as aggregation method.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2731989": "When do num_group1 and num_group2 change?\n\nI want to know exactly the standards by which they change.",
    "2732482": "Hi @dirhq88,\nI am not sure what you mean by change of num_group1 and num_group2 - could you explain more?",
    "2734465": "I used the word 'change' because num_group means historical records of case_id.\n\nDoes num_group1 increase by 1 when a major change in a particular case occurs and num_group2 increase by 1 when a minor change occurs?\n(I know that the smaller the number of num_groupN, the more current condition.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19837433%2F00dcb99e8be245bdf21a8f28a7ccbb2a%2F2024-04-04%20%202.34.31.png?generation=1712208893616616&alt=media)\n\nIn the above example, the values of num_group1 in case_id = 6 are divided into 0, 1. And when num_group1 = 0, the value of num_group2 is divided into 0,1, and when num_group1 = 1, the value of num_group2 is divided into 0,1,2,3,4,5.\n\nIn case_id=6, can it be seen that major changes occur 2 times and minor changes occur 2 and 5 times?",
    "2734617": "Hi @dirhq88 \n\n`because num_group means historical records of case_id`: this is not correct interpretation. Num_group is index for data with depth 1 or 2. Or to describe it differently, we have data with different granularity that are not aggregated on the level of case_id, for example previous applications, or contact references. It has nothing to do with changes in the data.",
    "2736163": "I have two question.\n\n\"num_group1 - This is an indexing column used for the historical records of case_id in both depth=1 and depth=2 tables.\"\nWhat is the exact meaning of 'historical' here?\n\n\"Or to describe it differently, we have data with different granularity that are not aggregated on the level of case_id, for example previous applications, or contact references.\"\nI don't understand your answer. Can you explain it in more detail?",
    "2736418": "Hi,\nokay, let me explain on hypothetical example:\n\n- Tomas Jelinek applied for loan on 1.1.2024 - this is credit case, which has assigned unique case_id\n- Tomas Jelinek is existing client in Home Credit, it means he had applications/loans with Home Credit before 1.1.2024, let's say 5 loans. Data describing parameters of those loans, their repayment history etc. are definitely valuable for credit scoring, therefore you have them in the sample. But they are not aggregated on level of case_id, there are 5 rows describing those 5 previous loans. To differentiate between those 5, you have to have some index - num_group_1, which will contain values 0,1,2,3,4...\n- num_group_1 is not used only for previous applications, but for other data where we have several records per case_id, like contact references, records in credit registry, etc. What I want to say is there is different meaning for different tables & attributes\n- is some cases we might have even bigger detail, for example information about instalments for each previous application. Then you need num_group_2, for example let's say previous loan with num_group_1=0 have 3 instalments, then you will have 3 records with num_group_1 = 0 and num_group_2 = 0,1,2",
    "2736619": "If it may help:\nIn Depth=1 files - each case_id has a history of activities/features related to it - that is num_group1.\nIn Depth=2 files - each num_group1 has a sub-history of activities/features related to them - that is num_group2\nIn conclusion - you can groupby case_id and aggregate these values to have some representation of the sub-history in one single value. In the example notebook the host shared - they took the max value as aggregation method.",
    "2739898": "Hi,I have a question:\n     Are num_group1=2's records closer to the present than num_group1=1's, or they don't have such relations?",
    "2741215": "Hi,\nyes, data should be sorted by num_group - timewise, position on application form,...",
    "2753193": "I have 4 question.\n\n1. Do credit cases and loans mean the same thing?\n- case_id: This is the unique identifier for each 'credit case'. You'll need this ID to join relevant tables to the base table.\n\nYou mentioned that \"But they are not aggregated on level of case_id, there are 5 rows describing those 5 previous loans. To differentiate between those 5, you have to have some index - num_group_1, which will contain values 0,1,2,3,4…\". I don't understand what this mean.\n\n2. Since 5 previous loans also have their own case_id, the features of each case (5 previous loans) can be distinguished sufficiently by case_id, why do we need to distinguish the features through the num_group in one case_id?\n\n3. If the features of the previous cases are included in the sample, will the case_id of the previous cases be overwritten to the current(new) case_id?\n\n4. Is it possible to change the features of the application after submitting the loan application?\n\nThank you so much :)",
    "2753378": "Hi,\n> Do credit cases and loans mean the same thing?\n\nYes. More precisely credit case = loan application. But for Kaggle data, only applications for which we have target (info about default or non-default) are included, so in this case credit case = loan (client signed loan agreement, money were disbursed,...)\n\nad 1.\nYou need to link previous applications to current application using its case ID. Otherwise you don't have direct link.\nHere is example: For case_id = 2, you have 2 records in applprev file, so for this credit case we have 2 previous loan applications, to distinguish them there is num_group1 index 0 or 1. Case_id = 2 is not id of those previous applications!\nSimilarly case_id=3 has 1 previous application, case_id=6 has 3 previous applications... \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7718387%2F512f85c255915e60b65f2150d8cdf066%2Fnum_group_example.png?generation=1713185950174529&alt=media)\n\nad 2.\nLet's say my application case_id = 333 is in the datasample. And I had two previous applications, let's say case_id=222 and case_id = 111. Information about 222 and 111 are included in applprev dataset under case_id = 333, because those are historical data for client with application 333.\nCase_id = 222 or case_id=111 might be also in the base table, if they are from the same time period which is covered by base table (if they are from different time period, they will not be in dataset)\n\nad 3.\nfeatures of applications are of course changing. What you have in dataset is snapshot as of time of application for particular case_id. So for example above, in applprev for case_id=333 you have info about loans/application 111 and 222 as of application date of 333.",
    "2778435": "Thank you for your previous response! \nI have a question regarding the \"applprev\" file, specifically about the skipped case IDs. For instance, there are no entries for case IDs 0 and 1. Does this omission mean that the data for these skipped case IDs have been incorporated into other existing case IDs under different 'num_groups' indexes? In other words, are these missing cases examples of previous applications whose updated information has been added to (or overwritten) in existing case IDs within the \"applprev\" file?"
  },
  "source": "meta"
}