{
  "id": 21039,
  "title": "How do you do feature selection here?",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21039",
  "author_name": "dot277",
  "post_date": "2016-05-18T07:20:16.747000",
  "votes": 0,
  "comment_count": 0,
  "views": 357,
  "content": "<p>Hi,\nI am trying to do the following for feature selection:</p>\n\n<p>1.\n I read the train file:</p>\n\n<pre><code>num_rows_to_read = 10000\ntrain_small = pd.read_csv(&quot;../input/train.csv, nrows=num_rows_to_read&quot;)\n</code></pre>\n\n<ol start=\"2\">\n<li>I change the type of the categorical features to 'category': </li>\n</ol>\n\n<p>non_categorial_features = ['orig_destination_distance','srch_adults_cnt',\n                                               'srch_children_cnt',\n                                               'srch_rm_cnt',\n                                              'cnt']</p>\n\n<pre><code>for categorical_feature in list(train_small.columns):\n    if categorical_feature not in non_categorial_features:\n        train_small[categorical_feature] = train_small[categorical_feature].astype('category')\n</code></pre>\n\n<ol start=\"3\">\n<li><p>I use one hot encoding:</p>\n\n<pre><code>       train_small_with_dummies = pd.get_dummies(train_small, sparse=True)\n</code></pre></li>\n</ol>\n\n<p>The problem is that the 3'rd part often get stuck, although i am using a strong machine.</p>\n\n<p>Thus, without the one hot encoding i can't to any feature selection, for determining the\nimportance of the features.</p>\n\n<p>What do you recommend?</p>\n\n<p>Thanks.</p>",
  "messages": [
    {
      "id": 120435,
      "postDate": "2016-05-18T07:20:16.747Z",
      "content": "<p>Hi,\nI am trying to do the following for feature selection:</p>\n\n<p>1.\n I read the train file:</p>\n\n<pre><code>num_rows_to_read = 10000\ntrain_small = pd.read_csv(&quot;../input/train.csv, nrows=num_rows_to_read&quot;)\n</code></pre>\n\n<ol start=\"2\">\n<li>I change the type of the categorical features to 'category': </li>\n</ol>\n\n<p>non_categorial_features = ['orig_destination_distance','srch_adults_cnt',\n                                               'srch_children_cnt',\n                                               'srch_rm_cnt',\n                                              'cnt']</p>\n\n<pre><code>for categorical_feature in list(train_small.columns):\n    if categorical_feature not in non_categorial_features:\n        train_small[categorical_feature] = train_small[categorical_feature].astype('category')\n</code></pre>\n\n<ol start=\"3\">\n<li><p>I use one hot encoding:</p>\n\n<pre><code>       train_small_with_dummies = pd.get_dummies(train_small, sparse=True)\n</code></pre></li>\n</ol>\n\n<p>The problem is that the 3'rd part often get stuck, although i am using a strong machine.</p>\n\n<p>Thus, without the one hot encoding i can't to any feature selection, for determining the\nimportance of the features.</p>\n\n<p>What do you recommend?</p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "Hi,\r\nI am trying to do the following for feature selection:\r\n\r\n1.\r\n I read the train file:\r\n\r\n    num_rows_to_read = 10000\r\n    train_small = pd.read_csv(\"../input/train.csv, nrows=num_rows_to_read\")\r\n\r\n2.  I change the type of the categorical features to 'category': \r\n\r\nnon_categorial_features = ['orig_destination_distance','srch_adults_cnt',\r\n                                               'srch_children_cnt',\r\n                                               'srch_rm_cnt',\r\n                                              'cnt']\r\n\r\n    for categorical_feature in list(train_small.columns):\r\n        if categorical_feature not in non_categorial_features:\r\n            train_small[categorical_feature] = train_small[categorical_feature].astype('category')\r\n\r\n3. I use one hot encoding:\r\n\r\n               train_small_with_dummies = pd.get_dummies(train_small, sparse=True)\r\n\r\n\r\nThe problem is that the 3'rd part often get stuck, although i am using a strong machine.\r\n\r\nThus, without the one hot encoding i can't to any feature selection, for determining the\r\nimportance of the features.\r\n\r\nWhat do you recommend?\r\n\r\nThanks.\r\n"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "120435": "Hi,\r\nI am trying to do the following for feature selection:\r\n\r\n1.\r\n I read the train file:\r\n\r\n    num_rows_to_read = 10000\r\n    train_small = pd.read_csv(\"../input/train.csv, nrows=num_rows_to_read\")\r\n\r\n2.  I change the type of the categorical features to 'category': \r\n\r\nnon_categorial_features = ['orig_destination_distance','srch_adults_cnt',\r\n                                               'srch_children_cnt',\r\n                                               'srch_rm_cnt',\r\n                                              'cnt']\r\n\r\n    for categorical_feature in list(train_small.columns):\r\n        if categorical_feature not in non_categorial_features:\r\n            train_small[categorical_feature] = train_small[categorical_feature].astype('category')\r\n\r\n3. I use one hot encoding:\r\n\r\n               train_small_with_dummies = pd.get_dummies(train_small, sparse=True)\r\n\r\n\r\nThe problem is that the 3'rd part often get stuck, although i am using a strong machine.\r\n\r\nThus, without the one hot encoding i can't to any feature selection, for determining the\r\nimportance of the features.\r\n\r\nWhat do you recommend?\r\n\r\nThanks.\r\n"
  }
}