{
  "id": 25511,
  "title": "How to deal with the nominal attribute",
  "url": "/competitions/outbrain-click-prediction/discussion/25511",
  "author_name": "",
  "post_date": "2016-11-16T10:39:34.193Z",
  "votes": null,
  "comment_count": 3,
  "views": 509,
  "content": "<p>Hi, i m a new beginner of this competition. One of the largest problem i met is how to deal with the nominal attributes topic_id and category_id. They all have more than 1000+ classes.</p>\n\n<p>I try to use one-hot encoder on python. The small dataset is OK but when dealing with the large dataset it seems to be out of memory.</p>\n\n<p>Could you please give me some suggestion on how to deal with the feature? Thanks a lot.</p>",
  "messages": [
    {
      "id": "145020",
      "postDate": "11/16/2016 10:39:34",
      "content": "<p>Hi, i m a new beginner of this competition. One of the largest problem i met is how to deal with the nominal attributes topic_id and category_id. They all have more than 1000+ classes.</p>\n\n<p>I try to use one-hot encoder on python. The small dataset is OK but when dealing with the large dataset it seems to be out of memory.</p>\n\n<p>Could you please give me some suggestion on how to deal with the feature? Thanks a lot.</p>",
      "rawMarkdown": "Hi, i m a new beginner of this competition. One of the largest problem i met is how to deal with the nominal attributes topic_id and category_id. They all have more than 1000+ classes.\r\n\r\nI try to use one-hot encoder on python. The small dataset is OK but when dealing with the large dataset it seems to be out of memory.\r\n\r\nCould you please give me some suggestion on how to deal with the feature? Thanks a lot.",
      "votes": null
    },
    {
      "id": "145076",
      "postDate": "11/16/2016 17:09:14",
      "content": "<p>Hi Jiani,</p>\n\n<p>One idea when the number of classes are very high is to use <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html\">LabelEncoding</a> instead of one-hot encoding. Label encoders generally work well with tree based models. </p>\n\n<p>One another idea is to use <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.HashingVectorizer.html\">hashing vectorizer</a> which hashes the classes to numerical values within specified range which can then be used in modeling. </p>",
      "rawMarkdown": "Hi Jiani,\r\n\r\nOne idea when the number of classes are very high is to use [LabelEncoding][1] instead of one-hot encoding. Label encoders generally work well with tree based models. \r\n\r\nOne another idea is to use [hashing vectorizer][2] which hashes the classes to numerical values within specified range which can then be used in modeling. \r\n\r\n\r\n  [1]: http://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html\r\n  [2]: http://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.HashingVectorizer.html",
      "votes": null
    },
    {
      "id": "145390",
      "postDate": "11/18/2016 02:31:45",
      "content": "<p>Thank you SRK. I will try your advice.\n[quote=SRK;145076]</p>\n\n<p>Hi Jiani,</p>\n\n<p>One idea when the number of classes are very high is to use <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html\">LabelEncoding</a> instead of one-hot encoding. Label encoders generally work well with tree based models. </p>\n\n<p>One another idea is to use <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.HashingVectorizer.html\">hashing vectorizer</a> which hashes the classes to numerical values within specified range which can then be used in modeling. </p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Thank you SRK. I will try your advice.\r\n[quote=SRK;145076]\r\n\r\nHi Jiani,\r\n\r\nOne idea when the number of classes are very high is to use [LabelEncoding][1] instead of one-hot encoding. Label encoders generally work well with tree based models. \r\n\r\nOne another idea is to use [hashing vectorizer][2] which hashes the classes to numerical values within specified range which can then be used in modeling. \r\n\r\n\r\n  [1]: http://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html\r\n  [2]: http://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.HashingVectorizer.html\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "148492",
      "postDate": "12/05/2016 08:36:20",
      "content": "<p>[quote=SRK;145076]</p>\n\n<p>Hi Jiani,</p>\n\n<p>One idea when the number of classes are very high is to use <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html\">LabelEncoding</a> instead of one-hot encoding. Label encoders generally work well with tree based models. </p>\n\n<p>One another idea is to use <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.HashingVectorizer.html\">hashing vectorizer</a> which hashes the classes to numerical values within specified range which can then be used in modeling. </p>\n\n<p>[/quote]</p>\n\n<p>In this Outbrain competition where topics and categories are already in the format of numerical ID, how should I encode these two features? One-hot-encoding seems infeasible due to memory constraint. Can i use feature hashing to reduce the dimension of feature \"topic\" (currently 300) to a smaller dimension like 20? Does this technique greatly degrades the accuracy ? Thanks.</p>",
      "rawMarkdown": "[quote=SRK;145076]\r\n\r\nHi Jiani,\r\n\r\nOne idea when the number of classes are very high is to use [LabelEncoding][1] instead of one-hot encoding. Label encoders generally work well with tree based models. \r\n\r\nOne another idea is to use [hashing vectorizer][2] which hashes the classes to numerical values within specified range which can then be used in modeling. \r\n\r\n\r\n  [1]: http://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html\r\n  [2]: http://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.HashingVectorizer.html\r\n\r\n[/quote]\r\n\r\n\r\nIn this Outbrain competition where topics and categories are already in the format of numerical ID, how should I encode these two features? One-hot-encoding seems infeasible due to memory constraint. Can i use feature hashing to reduce the dimension of feature \"topic\" (currently 300) to a smaller dimension like 20? Does this technique greatly degrades the accuracy ? Thanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 145076,
      "author_name": "sudalairajkumar",
      "author_url": "",
      "post_date": "11/16/2016 17:09:14",
      "content": "<p>Hi Jiani,</p>\n\n<p>One idea when the number of classes are very high is to use <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html\">LabelEncoding</a> instead of one-hot encoding. Label encoders generally work well with tree based models. </p>\n\n<p>One another idea is to use <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.HashingVectorizer.html\">hashing vectorizer</a> which hashes the classes to numerical values within specified range which can then be used in modeling. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 145390,
      "author_name": "gujiani",
      "author_url": "",
      "post_date": "11/18/2016 02:31:45",
      "content": "<p>Thank you SRK. I will try your advice.\n[quote=SRK;145076]</p>\n\n<p>Hi Jiani,</p>\n\n<p>One idea when the number of classes are very high is to use <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html\">LabelEncoding</a> instead of one-hot encoding. Label encoders generally work well with tree based models. </p>\n\n<p>One another idea is to use <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.HashingVectorizer.html\">hashing vectorizer</a> which hashes the classes to numerical values within specified range which can then be used in modeling. </p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 148492,
      "author_name": "peixiang",
      "author_url": "",
      "post_date": "12/05/2016 08:36:20",
      "content": "<p>[quote=SRK;145076]</p>\n\n<p>Hi Jiani,</p>\n\n<p>One idea when the number of classes are very high is to use <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html\">LabelEncoding</a> instead of one-hot encoding. Label encoders generally work well with tree based models. </p>\n\n<p>One another idea is to use <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.HashingVectorizer.html\">hashing vectorizer</a> which hashes the classes to numerical values within specified range which can then be used in modeling. </p>\n\n<p>[/quote]</p>\n\n<p>In this Outbrain competition where topics and categories are already in the format of numerical ID, how should I encode these two features? One-hot-encoding seems infeasible due to memory constraint. Can i use feature hashing to reduce the dimension of feature \"topic\" (currently 300) to a smaller dimension like 20? Does this technique greatly degrades the accuracy ? Thanks.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "145020": "Hi, i m a new beginner of this competition. One of the largest problem i met is how to deal with the nominal attributes topic_id and category_id. They all have more than 1000+ classes.\r\n\r\nI try to use one-hot encoder on python. The small dataset is OK but when dealing with the large dataset it seems to be out of memory.\r\n\r\nCould you please give me some suggestion on how to deal with the feature? Thanks a lot.",
    "145076": "Hi Jiani,\r\n\r\nOne idea when the number of classes are very high is to use [LabelEncoding][1] instead of one-hot encoding. Label encoders generally work well with tree based models. \r\n\r\nOne another idea is to use [hashing vectorizer][2] which hashes the classes to numerical values within specified range which can then be used in modeling. \r\n\r\n\r\n  [1]: http://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html\r\n  [2]: http://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.HashingVectorizer.html",
    "145390": "Thank you SRK. I will try your advice.\r\n[quote=SRK;145076]\r\n\r\nHi Jiani,\r\n\r\nOne idea when the number of classes are very high is to use [LabelEncoding][1] instead of one-hot encoding. Label encoders generally work well with tree based models. \r\n\r\nOne another idea is to use [hashing vectorizer][2] which hashes the classes to numerical values within specified range which can then be used in modeling. \r\n\r\n\r\n  [1]: http://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html\r\n  [2]: http://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.HashingVectorizer.html\r\n\r\n[/quote]",
    "148492": "[quote=SRK;145076]\r\n\r\nHi Jiani,\r\n\r\nOne idea when the number of classes are very high is to use [LabelEncoding][1] instead of one-hot encoding. Label encoders generally work well with tree based models. \r\n\r\nOne another idea is to use [hashing vectorizer][2] which hashes the classes to numerical values within specified range which can then be used in modeling. \r\n\r\n\r\n  [1]: http://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html\r\n  [2]: http://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.HashingVectorizer.html\r\n\r\n[/quote]\r\n\r\n\r\nIn this Outbrain competition where topics and categories are already in the format of numerical ID, how should I encode these two features? One-hot-encoding seems infeasible due to memory constraint. Can i use feature hashing to reduce the dimension of feature \"topic\" (currently 300) to a smaller dimension like 20? Does this technique greatly degrades the accuracy ? Thanks."
  },
  "source": "meta"
}