{
  "id": 476907,
  "title": "Confused about num_groupN",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/476907",
  "author_name": "Taichi Uemura",
  "post_date": "2024-02-14T02:25:40.237000",
  "votes": 16,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I am confused about the meaning of <code>num_groupN</code>. The columns of a <code>person_2</code> table are as follows.</p>\n<table>\n<thead>\n<tr>\n<th>Variable</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>\"addres_district_368M\"</td>\n<td>\"District of the person's address.\"</td>\n</tr>\n<tr>\n<td>\"addres_role_871L\"</td>\n<td>\"Role of person's address.\"</td>\n</tr>\n<tr>\n<td>\"addres_zip_823M\"</td>\n<td>\"Zip code of the address.\"</td>\n</tr>\n<tr>\n<td>\"conts_role_79M\"</td>\n<td>\"Type of contact role of a person.\"</td>\n</tr>\n<tr>\n<td>\"empls_economicalst_849M\"</td>\n<td>\"The economical status of the person (num_group1 - person, num_group2 - employment).\"</td>\n</tr>\n<tr>\n<td>\"empls_employedfrom_796D\"</td>\n<td>\"Start of employment (num_group1 - person, num_group2 - employment).\"</td>\n</tr>\n<tr>\n<td>\"empls_employer_name_740M\"</td>\n<td>\"Employer's name (num_group1 - person, num_group2 - employment).\"</td>\n</tr>\n<tr>\n<td>\"relatedpersons_role_762T\"</td>\n<td>\"Relationship type of a client's related person (num_group1 - person, num_group2 - related person).\"</td>\n</tr>\n</tbody>\n</table>\n<p>From the description of \"empls_economicalst_849M\" <code>num_group2</code> means employment, while from the description of \"relatedpersons_role_762T\" <code>num_group2</code> means related person. So the meaning of <code>num_group2</code> is not consistent across the table. How to understand <code>num_groupN</code> correctly?</p>",
  "messages": [
    {
      "id": 2651171,
      "postDate": "2024-02-14T02:25:40.237Z",
      "content": "<p>I am confused about the meaning of <code>num_groupN</code>. The columns of a <code>person_2</code> table are as follows.</p>\n<table>\n<thead>\n<tr>\n<th>Variable</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>\"addres_district_368M\"</td>\n<td>\"District of the person's address.\"</td>\n</tr>\n<tr>\n<td>\"addres_role_871L\"</td>\n<td>\"Role of person's address.\"</td>\n</tr>\n<tr>\n<td>\"addres_zip_823M\"</td>\n<td>\"Zip code of the address.\"</td>\n</tr>\n<tr>\n<td>\"conts_role_79M\"</td>\n<td>\"Type of contact role of a person.\"</td>\n</tr>\n<tr>\n<td>\"empls_economicalst_849M\"</td>\n<td>\"The economical status of the person (num_group1 - person, num_group2 - employment).\"</td>\n</tr>\n<tr>\n<td>\"empls_employedfrom_796D\"</td>\n<td>\"Start of employment (num_group1 - person, num_group2 - employment).\"</td>\n</tr>\n<tr>\n<td>\"empls_employer_name_740M\"</td>\n<td>\"Employer's name (num_group1 - person, num_group2 - employment).\"</td>\n</tr>\n<tr>\n<td>\"relatedpersons_role_762T\"</td>\n<td>\"Relationship type of a client's related person (num_group1 - person, num_group2 - related person).\"</td>\n</tr>\n</tbody>\n</table>\n<p>From the description of \"empls_economicalst_849M\" <code>num_group2</code> means employment, while from the description of \"relatedpersons_role_762T\" <code>num_group2</code> means related person. So the meaning of <code>num_group2</code> is not consistent across the table. How to understand <code>num_groupN</code> correctly?</p>",
      "rawMarkdown": "I am confused about the meaning of `num_groupN`. The columns of a `person_2` table are as follows.\n\n| Variable | Description |\n| --- | --- |\n| \"addres_district_368M\" | \"District of the person's address.\" |\n| \"addres_role_871L\" | \"Role of person's address.\" |\n| \"addres_zip_823M\" | \"Zip code of the address.\" |\n| \"conts_role_79M\" | \"Type of contact role of a person.\" |\n| \"empls_economicalst_849M\" | \"The economical status of the person (num_group1 - person, num_group2 - employment).\" |\n| \"empls_employedfrom_796D\" | \"Start of employment (num_group1 - person, num_group2 - employment).\" |\n| \"empls_employer_name_740M\" | \"Employer's name (num_group1 - person, num_group2 - employment).\" |\n| \"relatedpersons_role_762T\" | \"Relationship type of a client's related person (num_group1 - person, num_group2 - related person).\" |\n\nFrom the description of \"empls_economicalst_849M\" `num_group2` means employment, while from the description of \"relatedpersons_role_762T\" `num_group2` means related person. So the meaning of `num_group2` is not consistent across the table. How to understand `num_groupN` correctly?",
      "votes": 15
    },
    {
      "id": 2665136,
      "postDate": "2024-02-23T12:51:32.663Z",
      "content": "<p>same question, the meaning of num_group2 in this file is not clear <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> </p>",
      "rawMarkdown": "same question, the meaning of num_group2 in this file is not clear @jetakow ",
      "votes": 2
    },
    {
      "id": 2687856,
      "postDate": "2024-03-08T20:18:29.073Z",
      "content": "<p>My understanding from digging through the data it helps to think of this as data branches with the different meaning depending on the table (and column/variable). </p>\n<p>In case when there is no num_group1 or 2, data is directly related to a particular id, there is only <strong>one user_id -&gt; one data branch</strong>.</p>\n<p>In case where there is num_group1 only in the table (as in tax records, from my inspection), <strong>one user_id -&gt; multiple data branches</strong>. The value itself (in num_group1) might have a unique meaning or might be just an index variable, depending on column/table. Because in this case we have one to many relationship we will need to do some sort of data compression to compress all the branches related to user_id into one data branch (aka make it one user_id -&gt; one data branch).</p>\n<p>In case where there is num_group1 and num_group2, we have expansion of above. <strong>One user id -&gt; many num_group1 branches -&gt; many num_group2 branches</strong>. As above individual values in each num_group1 and num_group2 may have particular meaning or just be an index. But like above we will need to compress all of num_group2 values into one branch corresponding to particular user_id and num_group1 coombination. After we will need to do a second compression to compress all num_group1 corresponding to a particular user_id into a single data branch. </p>",
      "rawMarkdown": "My understanding from digging through the data it helps to think of this as data branches with the different meaning depending on the table (and column/variable). \n\nIn case when there is no num_group1 or 2, data is directly related to a particular id, there is only **one user_id -> one data branch**.\n\nIn case where there is num_group1 only in the table (as in tax records, from my inspection), **one user_id -> multiple data branches**. The value itself (in num_group1) might have a unique meaning or might be just an index variable, depending on column/table. Because in this case we have one to many relationship we will need to do some sort of data compression to compress all the branches related to user_id into one data branch (aka make it one user_id -> one data branch).\n\nIn case where there is num_group1 and num_group2, we have expansion of above. **One user id -> many num_group1 branches -> many num_group2 branches**. As above individual values in each num_group1 and num_group2 may have particular meaning or just be an index. But like above we will need to compress all of num_group2 values into one branch corresponding to particular user_id and num_group1 coombination. After we will need to do a second compression to compress all num_group1 corresponding to a particular user_id into a single data branch. ",
      "votes": 1
    },
    {
      "id": 2665853,
      "postDate": "2024-02-23T21:46:18.903Z",
      "content": "<p>If I won't respond on Monday, please ping me.</p>",
      "rawMarkdown": "If I won't respond on Monday, please ping me.",
      "replies": [
        {
          "id": 2670776,
          "postDate": "2024-02-27T05:52:31.467Z",
          "content": "<p>friendly reminder… <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> </p>",
          "rawMarkdown": "friendly reminder... @jetakow "
        },
        {
          "id": 2675296,
          "postDate": "2024-02-29T20:08:07.823Z",
          "content": "<p>friendly reminder</p>",
          "rawMarkdown": "friendly reminder"
        },
        {
          "id": 2675573,
          "postDate": "2024-03-01T02:19:55.403Z",
          "content": "<p>friendly reminder</p>",
          "rawMarkdown": "friendly reminder"
        },
        {
          "id": 2682938,
          "postDate": "2024-03-05T17:14:34.430Z",
          "content": "<blockquote>\n  <p>If I won't respond on Monday, please ping me.</p>\n</blockquote>\n<p>Hey Daniel, Friendly reminder. </p>",
          "rawMarkdown": "> If I won't respond on Monday, please ping me.\n\nHey Daniel, Friendly reminder. ",
          "replies": [
            {
              "id": 2685591,
              "postDate": "2024-03-07T09:38:31.320Z",
              "content": "<p>I apologize for late response. So the whole thing is rather simple. There are several groups of feature list you gave me.</p>\n<ol>\n<li>addres_district_368M, addres_role_871L, addres_zip_823M</li>\n<li>conts_role_79M</li>\n<li>empls_economicalst_849M, empls_employedfrom_796D, empls_employer_name_740M</li>\n<li>relatedpersons_role_762T</li>\n</ol>\n<p>They all have num_group1 as index of person, but 1) has num_group2 as addresses, 2) has num_group2 as contacts, 3) has num_group2 as employments and 4) has num_group2 as relatedPersons.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3772888%2Fc39d9573f01f0c98bfb2130bdecae0c1%2FScreenshot%202024-03-07%20at%2010.34.40.png?generation=1709804112035984&amp;alt=media\"></p>\n<p>When looking at the table we can infere/guess that \"a55475b1\" is probably a default value of some sort, or error value, or just filling value similar to None. Anyway, when considering values different than \"a55475b1\" we can see that the meaning of all the features you listed is not conflicting. When you are looking at groups with different meaning of num_group2 just think of them as independent values. It is the same as storing values in an array with different indices and instead of having many different arrays for every meaning of num_group2 we put it into one table.</p>\n<p><a href=\"https://www.kaggle.com/taichiuemura\" target=\"_blank\">@taichiuemura</a> let me know if it is clear now.</p>",
              "rawMarkdown": "I apologize for late response. So the whole thing is rather simple. There are several groups of feature list you gave me.\n1. addres_district_368M, addres_role_871L, addres_zip_823M\n2. conts_role_79M\n3. empls_economicalst_849M, empls_employedfrom_796D, empls_employer_name_740M\n4. relatedpersons_role_762T\n\nThey all have num_group1 as index of person, but 1) has num_group2 as addresses, 2) has num_group2 as contacts, 3) has num_group2 as employments and 4) has num_group2 as relatedPersons.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3772888%2Fc39d9573f01f0c98bfb2130bdecae0c1%2FScreenshot%202024-03-07%20at%2010.34.40.png?generation=1709804112035984&alt=media)\n\nWhen looking at the table we can infere/guess that \"a55475b1\" is probably a default value of some sort, or error value, or just filling value similar to None. Anyway, when considering values different than \"a55475b1\" we can see that the meaning of all the features you listed is not conflicting. When you are looking at groups with different meaning of num_group2 just think of them as independent values. It is the same as storing values in an array with different indices and instead of having many different arrays for every meaning of num_group2 we put it into one table.\n\n@taichiuemura let me know if it is clear now.",
              "votes": 4
            },
            {
              "id": 2686632,
              "postDate": "2024-03-08T00:30:40.677Z",
              "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> Thank you. That makes sense. So it's like the result of concatenation of tables with different columns. We have a few tables like</p>\n<table>\n<thead>\n<tr>\n<th>A</th>\n<th>B</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>\n<p>and</p>\n<table>\n<thead>\n<tr>\n<th>C</th>\n<th>D</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>2</td>\n<td>3</td>\n</tr>\n</tbody>\n</table>\n<p>and concatenate them into</p>\n<table>\n<thead>\n<tr>\n<th>A</th>\n<th>B</th>\n<th>C</th>\n<th>D</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>1</td>\n<td>null</td>\n<td>null</td>\n</tr>\n<tr>\n<td>null</td>\n<td>null</td>\n<td>2</td>\n<td>3</td>\n</tr>\n</tbody>\n</table>",
              "rawMarkdown": "@jetakow Thank you. That makes sense. So it's like the result of concatenation of tables with different columns. We have a few tables like\n\n| A | B |\n| --- | --- |\n| 0 | 1 |\n\nand\n\n| C | D |\n| --- | --- |\n| 2 | 3 |\n\nand concatenate them into\n\n| A | B | C | D |\n| --- | --- | --- | --- |\n| 0 | 1 | null | null|\n| null | null | 2 | 3 |\n"
            },
            {
              "id": 2687044,
              "postDate": "2024-03-08T08:13:46.183Z",
              "content": "<p>Thanks for the reply, but I'm still confused about it. </p>\n<p>(1) Does it mean, num_group1 and num_group2 have different meanings for different features, even if they are in the same table?<br>\n(2) Take feature \"addres_district_368M\" as an example, in 'feature_definitions.csv' there is no explanation of num_group1 and num_group2 for this feature. How do we interpret num_group1 and num_group2 for features<br>\nlike this? <br>\n(3) Take feature 'relatedpersons_role_762T' as another example, num_group1 means person, num_group2 means related person, then for case_id = 6, when (num_group1, num_group2) = (1, 1), it should mean person 1 and related person 1, but 'relatedpersons_role_762T' for this entry is 'OTHER RELATIVE'. My question is: In this case, '1' in num_group1 and num_group2 has different meaning as well? (They are not the same person?)<br>\n(4) For depth = 1 data, when should we interpret num_group1 as a (potentially) time-series index, and when should we interpret it as 'person'? </p>\n<p>Thanks for your clarification. </p>",
              "rawMarkdown": "Thanks for the reply, but I'm still confused about it. \n\n(1) Does it mean, num_group1 and num_group2 have different meanings for different features, even if they are in the same table?\n(2) Take feature \"addres_district_368M\" as an example, in 'feature_definitions.csv' there is no explanation of num_group1 and num_group2 for this feature. How do we interpret num_group1 and num_group2 for features\nlike this? \n(3) Take feature 'relatedpersons_role_762T' as another example, num_group1 means person, num_group2 means related person, then for case_id = 6, when (num_group1, num_group2) = (1, 1), it should mean person 1 and related person 1, but 'relatedpersons_role_762T' for this entry is 'OTHER RELATIVE'. My question is: In this case, '1' in num_group1 and num_group2 has different meaning as well? (They are not the same person?)\n(4) For depth = 1 data, when should we interpret num_group1 as a (potentially) time-series index, and when should we interpret it as 'person'? \n\nThanks for your clarification. ",
              "isDeleted": true
            },
            {
              "id": 2688694,
              "postDate": "2024-03-09T12:00:29.417Z",
              "content": "<p><a href=\"https://www.kaggle.com/minhengxiaoquant\" target=\"_blank\">@minhengxiaoquant</a> 1) yes 2) 3) I can provide explanation of remaining such features, but I will get to that on Monday or later. 4) it should be clear form the definitions provided, if not create a new post and tag me there </p>",
              "rawMarkdown": "@minhengxiaoquant 1) yes 2) 3) I can provide explanation of remaining such features, but I will get to that on Monday or later. 4) it should be clear form the definitions provided, if not create a new post and tag me there "
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2665136,
      "author_name": "claudiu",
      "author_url": "",
      "post_date": "2024-02-23T12:51:32.663000",
      "content": "<p>same question, the meaning of num_group2 in this file is not clear <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2687856,
      "author_name": "Sasha The Great",
      "author_url": "",
      "post_date": "2024-03-08T20:18:29.073000",
      "content": "<p>My understanding from digging through the data it helps to think of this as data branches with the different meaning depending on the table (and column/variable). </p>\n<p>In case when there is no num_group1 or 2, data is directly related to a particular id, there is only <strong>one user_id -&gt; one data branch</strong>.</p>\n<p>In case where there is num_group1 only in the table (as in tax records, from my inspection), <strong>one user_id -&gt; multiple data branches</strong>. The value itself (in num_group1) might have a unique meaning or might be just an index variable, depending on column/table. Because in this case we have one to many relationship we will need to do some sort of data compression to compress all the branches related to user_id into one data branch (aka make it one user_id -&gt; one data branch).</p>\n<p>In case where there is num_group1 and num_group2, we have expansion of above. <strong>One user id -&gt; many num_group1 branches -&gt; many num_group2 branches</strong>. As above individual values in each num_group1 and num_group2 may have particular meaning or just be an index. But like above we will need to compress all of num_group2 values into one branch corresponding to particular user_id and num_group1 coombination. After we will need to do a second compression to compress all num_group1 corresponding to a particular user_id into a single data branch. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2665853,
      "author_name": "Daniel Herman",
      "author_url": "",
      "post_date": "2024-02-23T21:46:18.903000",
      "content": "<p>If I won't respond on Monday, please ping me.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2670776,
          "author_name": "claudiu",
          "author_url": "",
          "post_date": "2024-02-27T05:52:31.467000",
          "content": "<p>friendly reminder… <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2675296,
          "author_name": "tianhe chen",
          "author_url": "",
          "post_date": "2024-02-29T20:08:07.823000",
          "content": "<p>friendly reminder</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2675573,
          "author_name": "FengSicheng",
          "author_url": "",
          "post_date": "2024-03-01T02:19:55.403000",
          "content": "<p>friendly reminder</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2682938,
          "author_name": "harshavardhan etcherla",
          "author_url": "",
          "post_date": "2024-03-05T17:14:34.430000",
          "content": "<blockquote>\n  <p>If I won't respond on Monday, please ping me.</p>\n</blockquote>\n<p>Hey Daniel, Friendly reminder. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2685591,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-03-07T09:38:31.320000",
              "content": "<p>I apologize for late response. So the whole thing is rather simple. There are several groups of feature list you gave me.</p>\n<ol>\n<li>addres_district_368M, addres_role_871L, addres_zip_823M</li>\n<li>conts_role_79M</li>\n<li>empls_economicalst_849M, empls_employedfrom_796D, empls_employer_name_740M</li>\n<li>relatedpersons_role_762T</li>\n</ol>\n<p>They all have num_group1 as index of person, but 1) has num_group2 as addresses, 2) has num_group2 as contacts, 3) has num_group2 as employments and 4) has num_group2 as relatedPersons.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3772888%2Fc39d9573f01f0c98bfb2130bdecae0c1%2FScreenshot%202024-03-07%20at%2010.34.40.png?generation=1709804112035984&amp;alt=media\"></p>\n<p>When looking at the table we can infere/guess that \"a55475b1\" is probably a default value of some sort, or error value, or just filling value similar to None. Anyway, when considering values different than \"a55475b1\" we can see that the meaning of all the features you listed is not conflicting. When you are looking at groups with different meaning of num_group2 just think of them as independent values. It is the same as storing values in an array with different indices and instead of having many different arrays for every meaning of num_group2 we put it into one table.</p>\n<p><a href=\"https://www.kaggle.com/taichiuemura\" target=\"_blank\">@taichiuemura</a> let me know if it is clear now.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2686632,
              "author_name": "Taichi Uemura",
              "author_url": "",
              "post_date": "2024-03-08T00:30:40.677000",
              "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> Thank you. That makes sense. So it's like the result of concatenation of tables with different columns. We have a few tables like</p>\n<table>\n<thead>\n<tr>\n<th>A</th>\n<th>B</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>\n<p>and</p>\n<table>\n<thead>\n<tr>\n<th>C</th>\n<th>D</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>2</td>\n<td>3</td>\n</tr>\n</tbody>\n</table>\n<p>and concatenate them into</p>\n<table>\n<thead>\n<tr>\n<th>A</th>\n<th>B</th>\n<th>C</th>\n<th>D</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>1</td>\n<td>null</td>\n<td>null</td>\n</tr>\n<tr>\n<td>null</td>\n<td>null</td>\n<td>2</td>\n<td>3</td>\n</tr>\n</tbody>\n</table>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2687044,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-03-08T08:13:46.183000",
              "content": "<p>Thanks for the reply, but I'm still confused about it. </p>\n<p>(1) Does it mean, num_group1 and num_group2 have different meanings for different features, even if they are in the same table?<br>\n(2) Take feature \"addres_district_368M\" as an example, in 'feature_definitions.csv' there is no explanation of num_group1 and num_group2 for this feature. How do we interpret num_group1 and num_group2 for features<br>\nlike this? <br>\n(3) Take feature 'relatedpersons_role_762T' as another example, num_group1 means person, num_group2 means related person, then for case_id = 6, when (num_group1, num_group2) = (1, 1), it should mean person 1 and related person 1, but 'relatedpersons_role_762T' for this entry is 'OTHER RELATIVE'. My question is: In this case, '1' in num_group1 and num_group2 has different meaning as well? (They are not the same person?)<br>\n(4) For depth = 1 data, when should we interpret num_group1 as a (potentially) time-series index, and when should we interpret it as 'person'? </p>\n<p>Thanks for your clarification. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2688694,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-03-09T12:00:29.417000",
              "content": "<p><a href=\"https://www.kaggle.com/minhengxiaoquant\" target=\"_blank\">@minhengxiaoquant</a> 1) yes 2) 3) I can provide explanation of remaining such features, but I will get to that on Monday or later. 4) it should be clear form the definitions provided, if not create a new post and tag me there </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2651171": "I am confused about the meaning of `num_groupN`. The columns of a `person_2` table are as follows.\n\n| Variable | Description |\n| --- | --- |\n| \"addres_district_368M\" | \"District of the person's address.\" |\n| \"addres_role_871L\" | \"Role of person's address.\" |\n| \"addres_zip_823M\" | \"Zip code of the address.\" |\n| \"conts_role_79M\" | \"Type of contact role of a person.\" |\n| \"empls_economicalst_849M\" | \"The economical status of the person (num_group1 - person, num_group2 - employment).\" |\n| \"empls_employedfrom_796D\" | \"Start of employment (num_group1 - person, num_group2 - employment).\" |\n| \"empls_employer_name_740M\" | \"Employer's name (num_group1 - person, num_group2 - employment).\" |\n| \"relatedpersons_role_762T\" | \"Relationship type of a client's related person (num_group1 - person, num_group2 - related person).\" |\n\nFrom the description of \"empls_economicalst_849M\" `num_group2` means employment, while from the description of \"relatedpersons_role_762T\" `num_group2` means related person. So the meaning of `num_group2` is not consistent across the table. How to understand `num_groupN` correctly?",
    "2665136": "same question, the meaning of num_group2 in this file is not clear @jetakow ",
    "2687856": "My understanding from digging through the data it helps to think of this as data branches with the different meaning depending on the table (and column/variable). \n\nIn case when there is no num_group1 or 2, data is directly related to a particular id, there is only **one user_id -> one data branch**.\n\nIn case where there is num_group1 only in the table (as in tax records, from my inspection), **one user_id -> multiple data branches**. The value itself (in num_group1) might have a unique meaning or might be just an index variable, depending on column/table. Because in this case we have one to many relationship we will need to do some sort of data compression to compress all the branches related to user_id into one data branch (aka make it one user_id -> one data branch).\n\nIn case where there is num_group1 and num_group2, we have expansion of above. **One user id -> many num_group1 branches -> many num_group2 branches**. As above individual values in each num_group1 and num_group2 may have particular meaning or just be an index. But like above we will need to compress all of num_group2 values into one branch corresponding to particular user_id and num_group1 coombination. After we will need to do a second compression to compress all num_group1 corresponding to a particular user_id into a single data branch. ",
    "2665853": "If I won't respond on Monday, please ping me."
  }
}