{
  "id": 491974,
  "title": "how to add categarical values when we have duplicate case id's (ie: applprev data set of credit previous data)",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/491974",
  "author_name": "",
  "post_date": "2024-04-08T05:07:15.063791200Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>applprev data set of credit previous data has multiple case id's of previous credit tansational data which are categarical like 1. cacccardblochreas_147M (Card blocking reason) , 2. conts_type_509L (Person contact type in previous application), 3.credacc_cards_status_52L(Card status of the previous credit account).</p>\n<p>Data understanding : <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/kunduruanil/credit-risk-data-understanding</a></p>\n<p>Feature Extration : <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/kunduruanil/create-training-dataset</a></p>",
  "messages": [
    {
      "id": "2741039",
      "postDate": "04/08/2024 05:07:15",
      "content": "<p>applprev data set of credit previous data has multiple case id's of previous credit tansational data which are categarical like 1. cacccardblochreas_147M (Card blocking reason) , 2. conts_type_509L (Person contact type in previous application), 3.credacc_cards_status_52L(Card status of the previous credit account).</p>\n<p>Data understanding : <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/kunduruanil/credit-risk-data-understanding</a></p>\n<p>Feature Extration : <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/kunduruanil/create-training-dataset</a></p>",
      "rawMarkdown": "applprev data set of credit previous data has multiple case id's of previous credit tansational data which are categarical like 1. cacccardblochreas_147M (Card blocking reason) , 2. conts_type_509L (Person contact type in previous application), 3.credacc_cards_status_52L(Card status of the previous credit account).\n\nData understanding : [https://www.kaggle.com/code/kunduruanil/credit-risk-data-understanding](url)\n\nFeature Extration : [https://www.kaggle.com/code/kunduruanil/create-training-dataset](url)",
      "votes": null
    },
    {
      "id": "2741999",
      "postDate": "04/08/2024 17:31:16",
      "content": "<p>general procedure is to get a feature file from applprev data  such that every row if one case_id than to use join to add these features to the base file.<br>\nnow to get the feature file we need to aggregate using groupby over case_id. in there u can use various conditions like take the top most value or return True if certain value is present in the list, or count the occurrence of a specific value in the list, or u can use some other logic/condition of ur choice to aggregate the list to one value.<br>\nU can perform this multiple times to get multiple features from a single feature <br>\nmore u extract more information u get but also more noise. so u need get balance between which feature u are extracting and how many.</p>",
      "rawMarkdown": "general procedure is to get a feature file from applprev data  such that every row if one case_id than to use join to add these features to the base file.\nnow to get the feature file we need to aggregate using groupby over case_id. in there u can use various conditions like take the top most value or return True if certain value is present in the list, or count the occurrence of a specific value in the list, or u can use some other logic/condition of ur choice to aggregate the list to one value.\nU can perform this multiple times to get multiple features from a single feature \nmore u extract more information u get but also more noise. so u need get balance between which feature u are extracting and how many.",
      "votes": null
    },
    {
      "id": "2742190",
      "postDate": "04/08/2024 19:32:36",
      "content": "<p>You can calculate mode, count number of unique values, or perform one-hot encoding and then sum up values. </p>",
      "rawMarkdown": "You can calculate mode, count number of unique values, or perform one-hot encoding and then sum up values.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2741999,
      "author_name": "shreyas9181",
      "author_url": "",
      "post_date": "04/08/2024 17:31:16",
      "content": "<p>general procedure is to get a feature file from applprev data  such that every row if one case_id than to use join to add these features to the base file.<br>\nnow to get the feature file we need to aggregate using groupby over case_id. in there u can use various conditions like take the top most value or return True if certain value is present in the list, or count the occurrence of a specific value in the list, or u can use some other logic/condition of ur choice to aggregate the list to one value.<br>\nU can perform this multiple times to get multiple features from a single feature <br>\nmore u extract more information u get but also more noise. so u need get balance between which feature u are extracting and how many.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2742190,
      "author_name": "eivolkova",
      "author_url": "",
      "post_date": "04/08/2024 19:32:36",
      "content": "<p>You can calculate mode, count number of unique values, or perform one-hot encoding and then sum up values. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2741039": "applprev data set of credit previous data has multiple case id's of previous credit tansational data which are categarical like 1. cacccardblochreas_147M (Card blocking reason) , 2. conts_type_509L (Person contact type in previous application), 3.credacc_cards_status_52L(Card status of the previous credit account).\n\nData understanding : [https://www.kaggle.com/code/kunduruanil/credit-risk-data-understanding](url)\n\nFeature Extration : [https://www.kaggle.com/code/kunduruanil/create-training-dataset](url)",
    "2741999": "general procedure is to get a feature file from applprev data  such that every row if one case_id than to use join to add these features to the base file.\nnow to get the feature file we need to aggregate using groupby over case_id. in there u can use various conditions like take the top most value or return True if certain value is present in the list, or count the occurrence of a specific value in the list, or u can use some other logic/condition of ur choice to aggregate the list to one value.\nU can perform this multiple times to get multiple features from a single feature \nmore u extract more information u get but also more noise. so u need get balance between which feature u are extracting and how many.",
    "2742190": "You can calculate mode, count number of unique values, or perform one-hot encoding and then sum up values."
  },
  "source": "meta"
}