{
  "id": 496977,
  "title": "How do you handle strings?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/496977",
  "author_name": "",
  "post_date": "2024-04-23T05:26:59.880701100Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>In the data preprocessing step, the values ​​of many features are strings. For example, the value below the feature <em>classificationofcontr_400M</em> is a55475b1. How do you deal with it?</p>",
  "messages": [
    {
      "id": "2768971",
      "postDate": "04/23/2024 05:26:59",
      "content": "<p>In the data preprocessing step, the values ​​of many features are strings. For example, the value below the feature <em>classificationofcontr_400M</em> is a55475b1. How do you deal with it?</p>",
      "rawMarkdown": "In the data preprocessing step, the values ​​of many features are strings. For example, the value below the feature *classificationofcontr_400M* is a55475b1. How do you deal with it?",
      "votes": null
    },
    {
      "id": "2772010",
      "postDate": "04/24/2024 13:55:53",
      "content": "<p>感觉常规的编码效果不如不处理，模型自己有识别类别列的方法</p>",
      "rawMarkdown": "感觉常规的编码效果不如不处理，模型自己有识别类别列的方法",
      "votes": null
    },
    {
      "id": "2789587",
      "postDate": "05/02/2024 18:16:06",
      "content": "<p>You can:</p>\n<ol>\n<li>Encode them</li>\n<li>Use models that accept categoricals (LightGBM, XGB, CatBoost)</li>\n<li>The mix of the first two point splitting the cardinality at some point</li>\n</ol>",
      "rawMarkdown": "You can:\n1. Encode them\n2. Use models that accept categoricals (LightGBM, XGB, CatBoost)\n3. The mix of the first two point splitting the cardinality at some point",
      "votes": null
    },
    {
      "id": "2807650",
      "postDate": "05/11/2024 19:02:24",
      "content": "<p>For columns with strings with many categories like a person's name, I just dropped them. For columns with less categories, we can convert them into categorical variables. </p>",
      "rawMarkdown": "For columns with strings with many categories like a person's name, I just dropped them. For columns with less categories, we can convert them into categorical variables.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2772010,
      "author_name": "swufeleo",
      "author_url": "",
      "post_date": "04/24/2024 13:55:53",
      "content": "<p>感觉常规的编码效果不如不处理，模型自己有识别类别列的方法</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2789587,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "05/02/2024 18:16:06",
      "content": "<p>You can:</p>\n<ol>\n<li>Encode them</li>\n<li>Use models that accept categoricals (LightGBM, XGB, CatBoost)</li>\n<li>The mix of the first two point splitting the cardinality at some point</li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2807650,
      "author_name": "faithk7u",
      "author_url": "",
      "post_date": "05/11/2024 19:02:24",
      "content": "<p>For columns with strings with many categories like a person's name, I just dropped them. For columns with less categories, we can convert them into categorical variables. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2768971": "In the data preprocessing step, the values ​​of many features are strings. For example, the value below the feature *classificationofcontr_400M* is a55475b1. How do you deal with it?",
    "2772010": "感觉常规的编码效果不如不处理，模型自己有识别类别列的方法",
    "2789587": "You can:\n1. Encode them\n2. Use models that accept categoricals (LightGBM, XGB, CatBoost)\n3. The mix of the first two point splitting the cardinality at some point",
    "2807650": "For columns with strings with many categories like a person's name, I just dropped them. For columns with less categories, we can convert them into categorical variables."
  },
  "source": "meta"
}