{
  "id": 475373,
  "title": "What is num_group feature?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/475373",
  "author_name": "Cafelatte1",
  "post_date": "2024-02-08T07:19:58.042000",
  "votes": 14,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I don't understand about num_groups feature.</p>\n<p>Does this feature mean a period which measures some metrics on each case_ids?</p>",
  "messages": [
    {
      "id": 2642486,
      "postDate": "2024-02-08T07:19:58.043Z",
      "content": "<p>I don't understand about num_groups feature.</p>\n<p>Does this feature mean a period which measures some metrics on each case_ids?</p>",
      "rawMarkdown": "I don't understand about num_groups feature.\n\nDoes this feature mean a period which measures some metrics on each case_ids?",
      "votes": 14
    },
    {
      "id": 2642544,
      "postDate": "2024-02-08T08:24:14.750Z",
      "content": "<p>Num_groupN are indices, see in Data tab</p>\n<blockquote>\n  <p>num_group1 - This is an indexing column used for the historical records of case_id in both depth=1 and depth=2 tables.<br>\n  num_group2 - This is the second indexing column for depth=2 tables' historical records of case_id. The order of num_group1 and num_group2 is important and will be clarified in feature definitions.<br>\n  All other raw columns in the tables serve as predictors. Their definitions can be found in the file feature_definitions.csv. For depth=0 tables, predictors can be directly used as features. However, for tables with depth&gt;0, you may need to employ aggregation functions that will condense the historical records associated with each case_id into a single feature. In case num_group1 or num_group2 stands for person index (this is clear with predictor definitions) the zero index has special meaning. When num_groupN=0 it is the applicant (the person who applied for a loan).</p>\n</blockquote>",
      "rawMarkdown": "Num_groupN are indices, see in Data tab\n\n>num_group1 - This is an indexing column used for the historical records of case_id in both depth=1 and depth=2 tables.\nnum_group2 - This is the second indexing column for depth=2 tables' historical records of case_id. The order of num_group1 and num_group2 is important and will be clarified in feature definitions.\nAll other raw columns in the tables serve as predictors. Their definitions can be found in the file feature_definitions.csv. For depth=0 tables, predictors can be directly used as features. However, for tables with depth>0, you may need to employ aggregation functions that will condense the historical records associated with each case_id into a single feature. In case num_group1 or num_group2 stands for person index (this is clear with predictor definitions) the zero index has special meaning. When num_groupN=0 it is the applicant (the person who applied for a loan).\n\n",
      "replies": [
        {
          "id": 2642553,
          "postDate": "2024-02-08T08:42:32.327Z",
          "content": "<p>Then, is structure below right?</p>\n<p>A | B means that A is a source table and B is key for joining<br>\n-&gt; means joining</p>\n<p>If I want to find historical data related to a case_id,</p>\n<p>base_table | case_id -&gt; depth1_table | case_id &amp; num_group1 -&gt; depth2_table</p>",
          "rawMarkdown": "Then, is structure below right?\n\nA | B means that A is a source table and B is key for joining\n-> means joining\n\nIf I want to find historical data related to a case_id,\n\nbase_table | case_id -> depth1_table | case_id & num_group1 -> depth2_table",
          "replies": [
            {
              "id": 2642612,
              "postDate": "2024-02-08T09:33:14.017Z",
              "content": "<p>Not exactly. For all the joining you need only <code>case_id</code>. Assume we have a table of depth=1, there will be a feature sex which stands for sex of all the people the person who applied for loan stated on the application form. Then <code>num_group1</code>=0 implies that it is the sex is of the person who applied for the loan, next <code>num_group1</code>=1 is for the first person mentioned on the application form, <code>num_group1</code>=2 is for second and so on. Basically <code>num_group1</code> is an index for tables of depth=1. Does this answer your question?</p>",
              "rawMarkdown": "Not exactly. For all the joining you need only `case_id`. Assume we have a table of depth=1, there will be a feature sex which stands for sex of all the people the person who applied for loan stated on the application form. Then `num_group1`=0 implies that it is the sex is of the person who applied for the loan, next `num_group1`=1 is for the first person mentioned on the application form, `num_group1`=2 is for second and so on. Basically `num_group1` is an index for tables of depth=1. Does this answer your question?",
              "votes": 14
            },
            {
              "id": 2642775,
              "postDate": "2024-02-08T12:01:37.530Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2643900,
              "postDate": "2024-02-09T06:58:05.580Z",
              "content": "<p>Hmm.. Yes, helpful to understand it ! Thanks.</p>",
              "rawMarkdown": "Hmm.. Yes, helpful to understand it ! Thanks.",
              "votes": 1
            },
            {
              "id": 2644842,
              "postDate": "2024-02-09T17:56:39.390Z",
              "content": "<p>Hi Daniel, can I ask you another detail?<br>\nIf the table has depth=1 how the num_group1 can be related to the different people and not a temporal index as stated in the first message?<br>\nI'm not getting this point.</p>",
              "rawMarkdown": "Hi Daniel, can I ask you another detail?\nIf the table has depth=1 how the num_group1 can be related to the different people and not a temporal index as stated in the first message?\nI'm not getting this point.",
              "votes": 1
            },
            {
              "id": 2644990,
              "postDate": "2024-02-09T20:14:14.680Z",
              "content": "<p>We might produce a list of such features where the meaning of the index differs. I need to discuss this internally and I will get back you on Monday.</p>",
              "rawMarkdown": "We might produce a list of such features where the meaning of the index differs. I need to discuss this internally and I will get back you on Monday."
            },
            {
              "id": 2650369,
              "postDate": "2024-02-13T12:25:14.950Z",
              "content": "<p>Any news on that?</p>",
              "rawMarkdown": "Any news on that?"
            },
            {
              "id": 2654996,
              "postDate": "2024-02-16T16:22:15.463Z",
              "content": "<p>Hey Daniel, </p>\n<p>I'm also a bit confused on this one. If you look at the train_applprev_2.csv file which contains all of the num_group1 and num_group2 historical associations, there's one example that's nice to look at. </p>\n<p>If you filter to case_id = 2703453, you can see num_group1 ranges from 0 to 8. So does this mean that there's up to 9 people listed on this application form? </p>\n<p>Also within depth2, when num_group1 = 2, the num_group2 values range from 0 to 3, but the conts_type_509L and credacc_cards_status_52L values are all NULL. I'm trying to understand the value these rows have if these columns are all NULL and cacccardblochreas_147M = a55475b1.</p>\n<p>For the next observation when num_group1=3, it only has 2 records for num_group2. So it's a bit hard to follow the logic. </p>\n<p>Some clarity for case_id 2703453 would be super appreciated!</p>\n<p>Thanks,</p>\n<p>Mark</p>",
              "rawMarkdown": "Hey Daniel, \n\nI'm also a bit confused on this one. If you look at the train_applprev_2.csv file which contains all of the num_group1 and num_group2 historical associations, there's one example that's nice to look at. \n\nIf you filter to case_id = 2703453, you can see num_group1 ranges from 0 to 8. So does this mean that there's up to 9 people listed on this application form? \n\nAlso within depth2, when num_group1 = 2, the num_group2 values range from 0 to 3, but the conts_type_509L and credacc_cards_status_52L values are all NULL. I'm trying to understand the value these rows have if these columns are all NULL and cacccardblochreas_147M = a55475b1.\n\nFor the next observation when num_group1=3, it only has 2 records for num_group2. So it's a bit hard to follow the logic. \n\nSome clarity for case_id 2703453 would be super appreciated!\n\nThanks,\n\nMark",
              "votes": 3
            },
            {
              "id": 2655418,
              "postDate": "2024-02-16T20:51:54.710Z",
              "content": "<p>It looks like the num_group1 logic is answered in other threads. So in this case, the application for case_id 2703453 had 9 previous applications. But inside of each group, I'm not sure I understand the reasoning behind the rows where conts_type_509L and credacc_cards_status_52L are both NULL? Is it because cacccardblochreas_147M maps to other values?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F723687%2Ffe59a1f5aa68e893b3a892db4b564d24%2FScreenshot%202024-02-16%20154906.png?generation=1708116565676590&amp;alt=media\"></p>",
              "rawMarkdown": "It looks like the num_group1 logic is answered in other threads. So in this case, the application for case_id 2703453 had 9 previous applications. But inside of each group, I'm not sure I understand the reasoning behind the rows where conts_type_509L and credacc_cards_status_52L are both NULL? Is it because cacccardblochreas_147M maps to other values?\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F723687%2Ffe59a1f5aa68e893b3a892db4b564d24%2FScreenshot%202024-02-16%20154906.png?generation=1708116565676590&alt=media)",
              "votes": 1
            },
            {
              "id": 2697938,
              "postDate": "2024-03-15T07:06:03.463Z",
              "content": "<p>Hi ! mark pearl!  Do you solve the above problem? I get the same confusion.</p>",
              "rawMarkdown": "Hi ! mark pearl!  Do you solve the above problem? I get the same confusion."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2642544,
      "author_name": "Daniel Herman",
      "author_url": "",
      "post_date": "2024-02-08T08:24:14.750000",
      "content": "<p>Num_groupN are indices, see in Data tab</p>\n<blockquote>\n  <p>num_group1 - This is an indexing column used for the historical records of case_id in both depth=1 and depth=2 tables.<br>\n  num_group2 - This is the second indexing column for depth=2 tables' historical records of case_id. The order of num_group1 and num_group2 is important and will be clarified in feature definitions.<br>\n  All other raw columns in the tables serve as predictors. Their definitions can be found in the file feature_definitions.csv. For depth=0 tables, predictors can be directly used as features. However, for tables with depth&gt;0, you may need to employ aggregation functions that will condense the historical records associated with each case_id into a single feature. In case num_group1 or num_group2 stands for person index (this is clear with predictor definitions) the zero index has special meaning. When num_groupN=0 it is the applicant (the person who applied for a loan).</p>\n</blockquote>",
      "votes": 0,
      "replies": [
        {
          "id": 2642553,
          "author_name": "Cafelatte1",
          "author_url": "",
          "post_date": "2024-02-08T08:42:32.327000",
          "content": "<p>Then, is structure below right?</p>\n<p>A | B means that A is a source table and B is key for joining<br>\n-&gt; means joining</p>\n<p>If I want to find historical data related to a case_id,</p>\n<p>base_table | case_id -&gt; depth1_table | case_id &amp; num_group1 -&gt; depth2_table</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2642612,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-08T09:33:14.017000",
              "content": "<p>Not exactly. For all the joining you need only <code>case_id</code>. Assume we have a table of depth=1, there will be a feature sex which stands for sex of all the people the person who applied for loan stated on the application form. Then <code>num_group1</code>=0 implies that it is the sex is of the person who applied for the loan, next <code>num_group1</code>=1 is for the first person mentioned on the application form, <code>num_group1</code>=2 is for second and so on. Basically <code>num_group1</code> is an index for tables of depth=1. Does this answer your question?</p>",
              "votes": 14,
              "replies": []
            },
            {
              "id": 2642775,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-02-08T12:01:37.530000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2643900,
              "author_name": "Cafelatte1",
              "author_url": "",
              "post_date": "2024-02-09T06:58:05.580000",
              "content": "<p>Hmm.. Yes, helpful to understand it ! Thanks.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2644842,
              "author_name": "Julien94",
              "author_url": "",
              "post_date": "2024-02-09T17:56:39.390000",
              "content": "<p>Hi Daniel, can I ask you another detail?<br>\nIf the table has depth=1 how the num_group1 can be related to the different people and not a temporal index as stated in the first message?<br>\nI'm not getting this point.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2644990,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-09T20:14:14.680000",
              "content": "<p>We might produce a list of such features where the meaning of the index differs. I need to discuss this internally and I will get back you on Monday.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2650369,
              "author_name": "Julien94",
              "author_url": "",
              "post_date": "2024-02-13T12:25:14.950000",
              "content": "<p>Any news on that?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2654996,
              "author_name": "Mark Pearl",
              "author_url": "",
              "post_date": "2024-02-16T16:22:15.463000",
              "content": "<p>Hey Daniel, </p>\n<p>I'm also a bit confused on this one. If you look at the train_applprev_2.csv file which contains all of the num_group1 and num_group2 historical associations, there's one example that's nice to look at. </p>\n<p>If you filter to case_id = 2703453, you can see num_group1 ranges from 0 to 8. So does this mean that there's up to 9 people listed on this application form? </p>\n<p>Also within depth2, when num_group1 = 2, the num_group2 values range from 0 to 3, but the conts_type_509L and credacc_cards_status_52L values are all NULL. I'm trying to understand the value these rows have if these columns are all NULL and cacccardblochreas_147M = a55475b1.</p>\n<p>For the next observation when num_group1=3, it only has 2 records for num_group2. So it's a bit hard to follow the logic. </p>\n<p>Some clarity for case_id 2703453 would be super appreciated!</p>\n<p>Thanks,</p>\n<p>Mark</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2655418,
              "author_name": "Mark Pearl",
              "author_url": "",
              "post_date": "2024-02-16T20:51:54.710000",
              "content": "<p>It looks like the num_group1 logic is answered in other threads. So in this case, the application for case_id 2703453 had 9 previous applications. But inside of each group, I'm not sure I understand the reasoning behind the rows where conts_type_509L and credacc_cards_status_52L are both NULL? Is it because cacccardblochreas_147M maps to other values?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F723687%2Ffe59a1f5aa68e893b3a892db4b564d24%2FScreenshot%202024-02-16%20154906.png?generation=1708116565676590&amp;alt=media\"></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2697938,
              "author_name": "zpccccc",
              "author_url": "",
              "post_date": "2024-03-15T07:06:03.463000",
              "content": "<p>Hi ! mark pearl!  Do you solve the above problem? I get the same confusion.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2642486": "I don't understand about num_groups feature.\n\nDoes this feature mean a period which measures some metrics on each case_ids?",
    "2642544": "Num_groupN are indices, see in Data tab\n\n>num_group1 - This is an indexing column used for the historical records of case_id in both depth=1 and depth=2 tables.\nnum_group2 - This is the second indexing column for depth=2 tables' historical records of case_id. The order of num_group1 and num_group2 is important and will be clarified in feature definitions.\nAll other raw columns in the tables serve as predictors. Their definitions can be found in the file feature_definitions.csv. For depth=0 tables, predictors can be directly used as features. However, for tables with depth>0, you may need to employ aggregation functions that will condense the historical records associated with each case_id into a single feature. In case num_group1 or num_group2 stands for person index (this is clear with predictor definitions) the zero index has special meaning. When num_groupN=0 it is the applicant (the person who applied for a loan).\n\n"
  }
}