{
  "id": 20440,
  "title": "Dummy all categorical variables?",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20440",
  "author_name": "",
  "post_date": "2016-04-26T07:52:29.697Z",
  "votes": null,
  "comment_count": 3,
  "views": 848,
  "content": "<p>In the book of variables. </p>\n\n<p>[1] &quot;date_time&quot;                 &quot;site_name&quot;                 &quot;posa_continent&quot;            &quot;user_location_country&quot;     &quot;user_location_region&quot; <br>\n[6] &quot;user_location_city&quot;        &quot;orig_destination_distance&quot; &quot;user_id&quot;                   &quot;is_mobile&quot;                 &quot;is_package&quot; <br>\n[11] &quot;channel&quot;                   &quot;srch_ci&quot;                   &quot;srch_co&quot;                   &quot;srch_adults_cnt&quot;           &quot;srch_children_cnt&quot; <br>\n[16] &quot;srch_rm_cnt&quot;               &quot;srch_destination_id&quot;       &quot;srch_destination_type_id&quot;  &quot;is_booking&quot;                &quot;cnt&quot; <br>\n[21] &quot;hotel_continent&quot;           &quot;hotel_country&quot;             &quot;hotel_market&quot;              &quot;hotel_cluster&quot;    </p>\n\n<p>Actually we have over 80% categorical variables presented in numeric format.</p>\n\n<p>And some of them have more than 100 categories. It's inefficient to set them into dummy variables. While numeric format will apparently create errors.</p>\n\n<p>Is anyone have a good solution on that? Or any pkg can handle such problem?</p>",
  "messages": [
    {
      "id": "116821",
      "postDate": "04/26/2016 07:52:29",
      "content": "<p>In the book of variables. </p>\n\n<p>[1] &quot;date_time&quot;                 &quot;site_name&quot;                 &quot;posa_continent&quot;            &quot;user_location_country&quot;     &quot;user_location_region&quot; <br>\n[6] &quot;user_location_city&quot;        &quot;orig_destination_distance&quot; &quot;user_id&quot;                   &quot;is_mobile&quot;                 &quot;is_package&quot; <br>\n[11] &quot;channel&quot;                   &quot;srch_ci&quot;                   &quot;srch_co&quot;                   &quot;srch_adults_cnt&quot;           &quot;srch_children_cnt&quot; <br>\n[16] &quot;srch_rm_cnt&quot;               &quot;srch_destination_id&quot;       &quot;srch_destination_type_id&quot;  &quot;is_booking&quot;                &quot;cnt&quot; <br>\n[21] &quot;hotel_continent&quot;           &quot;hotel_country&quot;             &quot;hotel_market&quot;              &quot;hotel_cluster&quot;    </p>\n\n<p>Actually we have over 80% categorical variables presented in numeric format.</p>\n\n<p>And some of them have more than 100 categories. It's inefficient to set them into dummy variables. While numeric format will apparently create errors.</p>\n\n<p>Is anyone have a good solution on that? Or any pkg can handle such problem?</p>",
      "rawMarkdown": "In the book of variables. \r\n\r\n\r\n[1] \"date_time\"                 \"site_name\"                 \"posa_continent\"            \"user_location_country\"     \"user_location_region\"     \r\n[6] \"user_location_city\"        \"orig_destination_distance\" \"user_id\"                   \"is_mobile\"                 \"is_package\"               \r\n[11] \"channel\"                   \"srch_ci\"                   \"srch_co\"                   \"srch_adults_cnt\"           \"srch_children_cnt\"        \r\n[16] \"srch_rm_cnt\"               \"srch_destination_id\"       \"srch_destination_type_id\"  \"is_booking\"                \"cnt\"                      \r\n[21] \"hotel_continent\"           \"hotel_country\"             \"hotel_market\"              \"hotel_cluster\"    \r\n\r\n\r\nActually we have over 80% categorical variables presented in numeric format.\r\n\r\nAnd some of them have more than 100 categories. It's inefficient to set them into dummy variables. While numeric format will apparently create errors.\r\n\r\nIs anyone have a good solution on that? Or any pkg can handle such problem?",
      "votes": null
    },
    {
      "id": "116846",
      "postDate": "04/26/2016 09:17:39",
      "content": "<p>You can try only to use the most frequent and &quot;other&quot; values in each column to dummy code. Also you can try to use a hashed model matrix.</p>",
      "rawMarkdown": "You can try only to use the most frequent and \"other\" values in each column to dummy code. Also you can try to use a hashed model matrix.",
      "votes": null
    },
    {
      "id": "117501",
      "postDate": "04/29/2016 12:51:01",
      "content": "<p>You can use scipy sparce matrix to store such huge and empty categorical matrices.\n<a href=\"http://docs.scipy.org/doc/scipy/reference/sparse.html\">http://docs.scipy.org/doc/scipy/reference/sparse.html</a></p>\n\n<p>Many packages such as numpy and pandas support this type of data.</p>",
      "rawMarkdown": "You can use scipy sparce matrix to store such huge and empty categorical matrices.\r\nhttp://docs.scipy.org/doc/scipy/reference/sparse.html\r\n\r\nMany packages such as numpy and pandas support this type of data.",
      "votes": null
    },
    {
      "id": "120451",
      "postDate": "05/18/2016 11:19:05",
      "content": "<p>@pixixi - What have you done eventually with regards to the dummy variables?gf</p>",
      "rawMarkdown": "pixixi - What have you done eventually with regards to the dummy variables?gf",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 116846,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/26/2016 09:17:39",
      "content": "<p>You can try only to use the most frequent and &quot;other&quot; values in each column to dummy code. Also you can try to use a hashed model matrix.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117501,
      "author_name": "bozhkov",
      "author_url": "",
      "post_date": "04/29/2016 12:51:01",
      "content": "<p>You can use scipy sparce matrix to store such huge and empty categorical matrices.\n<a href=\"http://docs.scipy.org/doc/scipy/reference/sparse.html\">http://docs.scipy.org/doc/scipy/reference/sparse.html</a></p>\n\n<p>Many packages such as numpy and pandas support this type of data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120451,
      "author_name": "dot277",
      "author_url": "",
      "post_date": "05/18/2016 11:19:05",
      "content": "<p>@pixixi - What have you done eventually with regards to the dummy variables?gf</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "116821": "In the book of variables. \r\n\r\n\r\n[1] \"date_time\"                 \"site_name\"                 \"posa_continent\"            \"user_location_country\"     \"user_location_region\"     \r\n[6] \"user_location_city\"        \"orig_destination_distance\" \"user_id\"                   \"is_mobile\"                 \"is_package\"               \r\n[11] \"channel\"                   \"srch_ci\"                   \"srch_co\"                   \"srch_adults_cnt\"           \"srch_children_cnt\"        \r\n[16] \"srch_rm_cnt\"               \"srch_destination_id\"       \"srch_destination_type_id\"  \"is_booking\"                \"cnt\"                      \r\n[21] \"hotel_continent\"           \"hotel_country\"             \"hotel_market\"              \"hotel_cluster\"    \r\n\r\n\r\nActually we have over 80% categorical variables presented in numeric format.\r\n\r\nAnd some of them have more than 100 categories. It's inefficient to set them into dummy variables. While numeric format will apparently create errors.\r\n\r\nIs anyone have a good solution on that? Or any pkg can handle such problem?",
    "116846": "You can try only to use the most frequent and \"other\" values in each column to dummy code. Also you can try to use a hashed model matrix.",
    "117501": "You can use scipy sparce matrix to store such huge and empty categorical matrices.\r\nhttp://docs.scipy.org/doc/scipy/reference/sparse.html\r\n\r\nMany packages such as numpy and pandas support this type of data.",
    "120451": "pixixi - What have you done eventually with regards to the dummy variables?gf"
  },
  "source": "meta"
}