{
  "id": 485823,
  "title": "Clarification on num_group2 Interpretation for train_person_2",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/485823",
  "author_name": "",
  "post_date": "2024-03-22T09:08:44.674027600Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I'm diving into the train_person_2 dataset, specifically for case_id == 6, and trying to grasp the logic behind the num_group2 field. Based on my observations, the dataset seems to follow certain patterns:</p>\n<p>Address Information: There are 6 unique patterns.<br>\nContact Information: There are 3 unique patterns, represented by (P38_92_157, P7_147_157, a55475b1).<br>\nEmployment Information: There are 2 patterns, where the empls_economicalst_849M column has values (P164_110_33, a55475b1).<br>\nRelated Persons Information: There are 2 patterns (OTHER_RELATIVE, null).<br>\nFrom this, num_group2 appears to align with the number of rows based on the highest number of unique patterns (6 in this case), suggesting that it might be serving as an index for unique combinations across columns. However, I find it challenging to fully understand this interpretation, especially when comparing it to time-series data where num_group is more straightforward to interpret.</p>\n<p>I've come across similar topics, such as:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476907\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476907</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/475373\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/475373</a></li>\n</ul>\n<p>Even after reading through these discussions, I still find myself struggling to fully comprehend the role and interpretation of num_group2 in the context of the person data.</p>\n<p>Could someone please clarify if num_group2 should indeed be interpreted as an index for unique combinations across different columns? Or is there another interpretation or logic behind its assignment that I might be missing?</p>\n<p>Thank you in advance for your insights!</p>",
  "messages": [
    {
      "id": "2710451",
      "postDate": "03/22/2024 09:08:44",
      "content": "<p>Hi everyone,</p>\n<p>I'm diving into the train_person_2 dataset, specifically for case_id == 6, and trying to grasp the logic behind the num_group2 field. Based on my observations, the dataset seems to follow certain patterns:</p>\n<p>Address Information: There are 6 unique patterns.<br>\nContact Information: There are 3 unique patterns, represented by (P38_92_157, P7_147_157, a55475b1).<br>\nEmployment Information: There are 2 patterns, where the empls_economicalst_849M column has values (P164_110_33, a55475b1).<br>\nRelated Persons Information: There are 2 patterns (OTHER_RELATIVE, null).<br>\nFrom this, num_group2 appears to align with the number of rows based on the highest number of unique patterns (6 in this case), suggesting that it might be serving as an index for unique combinations across columns. However, I find it challenging to fully understand this interpretation, especially when comparing it to time-series data where num_group is more straightforward to interpret.</p>\n<p>I've come across similar topics, such as:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476907\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476907</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/475373\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/475373</a></li>\n</ul>\n<p>Even after reading through these discussions, I still find myself struggling to fully comprehend the role and interpretation of num_group2 in the context of the person data.</p>\n<p>Could someone please clarify if num_group2 should indeed be interpreted as an index for unique combinations across different columns? Or is there another interpretation or logic behind its assignment that I might be missing?</p>\n<p>Thank you in advance for your insights!</p>",
      "rawMarkdown": "Hi everyone,\n\nI'm diving into the train_person_2 dataset, specifically for case_id == 6, and trying to grasp the logic behind the num_group2 field. Based on my observations, the dataset seems to follow certain patterns:\n\nAddress Information: There are 6 unique patterns.\nContact Information: There are 3 unique patterns, represented by (P38_92_157, P7_147_157, a55475b1).\nEmployment Information: There are 2 patterns, where the empls_economicalst_849M column has values (P164_110_33, a55475b1).\nRelated Persons Information: There are 2 patterns (OTHER_RELATIVE, null).\nFrom this, num_group2 appears to align with the number of rows based on the highest number of unique patterns (6 in this case), suggesting that it might be serving as an index for unique combinations across columns. However, I find it challenging to fully understand this interpretation, especially when comparing it to time-series data where num_group is more straightforward to interpret.\n\nI've come across similar topics, such as:\n\n- https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476907\n- https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/475373\n\nEven after reading through these discussions, I still find myself struggling to fully comprehend the role and interpretation of num_group2 in the context of the person data.\n\nCould someone please clarify if num_group2 should indeed be interpreted as an index for unique combinations across different columns? Or is there another interpretation or logic behind its assignment that I might be missing?\n\nThank you in advance for your insights!",
      "votes": null
    },
    {
      "id": "2712589",
      "postDate": "03/23/2024 17:32:49",
      "content": "<p>I gave up on this table :)</p>\n<p>The only info i  can add: It seems the lower num_group2, the more information is specified / non-default values.</p>\n<p>If you want to see a messy case use 1336868. Some of the entries do not change at all but for num_group2</p>",
      "rawMarkdown": "I gave up on this table :)\n\nThe only info i  can add: It seems the lower num_group2, the more information is specified / non-default values.\n\nIf you want to see a messy case use 1336868. Some of the entries do not change at all but for num_group2",
      "votes": null
    },
    {
      "id": "2716244",
      "postDate": "03/25/2024 22:47:39",
      "content": "<p>In Depth=1 files each case_id has a history of activities/features related to it - that is num_group1. <br>\nIn Depth=2 files - each num_group1 has a sub-history of activities/features related to them - that is num_group2</p>",
      "rawMarkdown": "In Depth=1 files each case_id has a history of activities/features related to it - that is num_group1. \nIn Depth=2 files - each num_group1 has a sub-history of activities/features related to them - that is num_group2",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2712589,
      "author_name": "aahhammer",
      "author_url": "",
      "post_date": "03/23/2024 17:32:49",
      "content": "<p>I gave up on this table :)</p>\n<p>The only info i  can add: It seems the lower num_group2, the more information is specified / non-default values.</p>\n<p>If you want to see a messy case use 1336868. Some of the entries do not change at all but for num_group2</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2716244,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "03/25/2024 22:47:39",
      "content": "<p>In Depth=1 files each case_id has a history of activities/features related to it - that is num_group1. <br>\nIn Depth=2 files - each num_group1 has a sub-history of activities/features related to them - that is num_group2</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2710451": "Hi everyone,\n\nI'm diving into the train_person_2 dataset, specifically for case_id == 6, and trying to grasp the logic behind the num_group2 field. Based on my observations, the dataset seems to follow certain patterns:\n\nAddress Information: There are 6 unique patterns.\nContact Information: There are 3 unique patterns, represented by (P38_92_157, P7_147_157, a55475b1).\nEmployment Information: There are 2 patterns, where the empls_economicalst_849M column has values (P164_110_33, a55475b1).\nRelated Persons Information: There are 2 patterns (OTHER_RELATIVE, null).\nFrom this, num_group2 appears to align with the number of rows based on the highest number of unique patterns (6 in this case), suggesting that it might be serving as an index for unique combinations across columns. However, I find it challenging to fully understand this interpretation, especially when comparing it to time-series data where num_group is more straightforward to interpret.\n\nI've come across similar topics, such as:\n\n- https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476907\n- https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/475373\n\nEven after reading through these discussions, I still find myself struggling to fully comprehend the role and interpretation of num_group2 in the context of the person data.\n\nCould someone please clarify if num_group2 should indeed be interpreted as an index for unique combinations across different columns? Or is there another interpretation or logic behind its assignment that I might be missing?\n\nThank you in advance for your insights!",
    "2712589": "I gave up on this table :)\n\nThe only info i  can add: It seems the lower num_group2, the more information is specified / non-default values.\n\nIf you want to see a messy case use 1336868. Some of the entries do not change at all but for num_group2",
    "2716244": "In Depth=1 files each case_id has a history of activities/features related to it - that is num_group1. \nIn Depth=2 files - each num_group1 has a sub-history of activities/features related to them - that is num_group2"
  },
  "source": "meta"
}