{
  "id": 52265,
  "title": "Label Encoding",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/52265",
  "author_name": "WilsonSteven",
  "post_date": "2018-03-18T07:51:20.815000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi Everyone,</p>\n\n<p>I have a very naive question. Many kaggle masters use OOF/Noise added label encoding for high cardinality categorical features and i get the intuition behind it. I have a few questions:</p>\n\n<ol>\n<li>I have not seen anyone use this in the industry. Can please someone comment on it? If not, what is the preferred choice?</li>\n<li>I was reading up on handling high cardinality categorical features, does creating dummy features and using Multiple Correspondence Analysis sound like a good idea? </li>\n</ol>\n\n<p>Thanks</p>",
  "messages": [
    {
      "id": 297879,
      "postDate": "2018-03-18T07:51:20.817Z",
      "content": "<p>Hi Everyone,</p>\n\n<p>I have a very naive question. Many kaggle masters use OOF/Noise added label encoding for high cardinality categorical features and i get the intuition behind it. I have a few questions:</p>\n\n<ol>\n<li>I have not seen anyone use this in the industry. Can please someone comment on it? If not, what is the preferred choice?</li>\n<li>I was reading up on handling high cardinality categorical features, does creating dummy features and using Multiple Correspondence Analysis sound like a good idea? </li>\n</ol>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi Everyone,\n\nI have a very naive question. Many kaggle masters use OOF/Noise added label encoding for high cardinality categorical features and i get the intuition behind it. I have a few questions:\n\n1. I have not seen anyone use this in the industry. Can please someone comment on it? If not, what is the preferred choice?\n2. I was reading up on handling high cardinality categorical features, does creating dummy features and using Multiple Correspondence Analysis sound like a good idea? \n\nThanks"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "297879": "Hi Everyone,\n\nI have a very naive question. Many kaggle masters use OOF/Noise added label encoding for high cardinality categorical features and i get the intuition behind it. I have a few questions:\n\n1. I have not seen anyone use this in the industry. Can please someone comment on it? If not, what is the preferred choice?\n2. I was reading up on handling high cardinality categorical features, does creating dummy features and using Multiple Correspondence Analysis sound like a good idea? \n\nThanks"
  }
}