{
  "id": 57094,
  "title": "How can we specify categorical features in XGB without one-hot encoding?",
  "url": "/competitions/avito-demand-prediction/discussion/57094",
  "author_name": "",
  "post_date": "2018-05-19T03:27:58.832163300Z",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi all,</p>\n\n<p>I wonder how can we specify categorical features in XGB (on Python) without using one-hot encoding? I noticed several kernels using XGB but didn't see any one-hot encoding, so how can XGB differentiate between numerical and categorical features?  To my knowledge, XGB does not support categorical features and it requires us to one-hot encode all categorical features.</p>",
  "messages": [
    {
      "id": "330520",
      "postDate": "05/19/2018 03:27:58",
      "content": "<p>Hi all,</p>\n\n<p>I wonder how can we specify categorical features in XGB (on Python) without using one-hot encoding? I noticed several kernels using XGB but didn't see any one-hot encoding, so how can XGB differentiate between numerical and categorical features?  To my knowledge, XGB does not support categorical features and it requires us to one-hot encode all categorical features.</p>",
      "rawMarkdown": "Hi all,\n\nI wonder how can we specify categorical features in XGB (on Python) without using one-hot encoding? I noticed several kernels using XGB but didn't see any one-hot encoding, so how can XGB differentiate between numerical and categorical features?  To my knowledge, XGB does not support categorical features and it requires us to one-hot encode all categorical features.",
      "votes": null
    },
    {
      "id": "330696",
      "postDate": "05/19/2018 13:32:00",
      "content": "<p>Hi @Kha, You can try these feature engineering as following link replace one-hot for categorical features. It is very excellent for your model performance and the finally score at LB.</p>\n\n<p><a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/55521\">https://www.kaggle.com/c/avito-demand-prediction/discussion/55521</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/57119\">https://www.kaggle.com/c/avito-demand-prediction/discussion/57119</a></p>",
      "rawMarkdown": "Hi @Kha, You can try these feature engineering as following link replace one-hot for categorical features. It is very excellent for your model performance and the finally score at LB.\n\nhttps://www.kaggle.com/c/avito-demand-prediction/discussion/55521\n\nhttps://www.kaggle.com/c/avito-demand-prediction/discussion/57119",
      "votes": null
    },
    {
      "id": "330701",
      "postDate": "05/19/2018 13:38:17",
      "content": "<p>Thanks DUO for your suggestion. I've skimmed through that link, but still not answer my original question. I notice some kernels proposed here using XGB (on Python) didn't apply any encoding method for categorical features. So I just wonder why XGB knows? Or the kernel authors forgot about this important issue?</p>",
      "rawMarkdown": "Thanks DUO for your suggestion. I've skimmed through that link, but still not answer my original question. I notice some kernels proposed here using XGB (on Python) didn't apply any encoding method for categorical features. So I just wonder why XGB knows? Or the kernel authors forgot about this important issue?",
      "votes": null
    },
    {
      "id": "330709",
      "postDate": "05/19/2018 14:02:25",
      "content": "<p>From my reading of xgboost documentation I didn't see any special handling of unordered categorical variables. In any case, many Tree algorithms will treat a categorical variable as ordered, which on the face of it seems bad.</p>\n\n<p>On the other hand, by splitting up your one categorical variable into a (large?) collection of one-hot encoded values (e.g. Booleans) will have taken a compact variable and turned it into a whole bunch of sparse variables.</p>\n\n<p><strong>The potential (and may be even likely result) is that a Tree algorithm will down play and even ignore the one-hot encoded features especially there are a lot of them.</strong> This is because Trees like variables that have high recall over those that have high precision but low recall.</p>\n\n<p>That is, a Tree tries to increase information gain on the data it's splitting at any given level. Unintuitively, it doesn't like to use sparse variables to do this.</p>\n\n<p>The upshot is you are in a real quandary. I suggest that you try it both ways to see which works better…. However, not using one-hot encoding might not be as bad as you think because a Tree might treat the categorical feature as a highly multimodal variable and make decent splits anyway.</p>\n\n<p>One can try a few other approaches:</p>\n\n<p>look at how the response variable responds to the categorical values and try to group them.\nFind another ML algorithm that works better with categorical features or with one-hot encoding and use that to train a submodel that just uses the categorical features. Then replace the categorical feature with a probability score. For instance, use a Logistic Regression on the hot-encoded values.\nTry to combine the categorical feature with some other features.\nBuild N xgboost classifiers, one for each category.\nThis may require playing around with the data a bit. Plotting the data may help you see patterns that you didn't know that were there.</p>\n\n<p>[added]</p>\n\n<p>Looks like that xgboost does only deal with numeric values:</p>\n\n<p><a href=\"https://cran.r-project.org/web/packages/xgboost/vignettes/discoverYourData.html\">Understand your dataset with XGBoost</a></p>\n\n<p>[added]</p>\n\n<p>An example of where I built N models, was when I built one for each major human language (English, French, etc) for web page classification.</p>",
      "rawMarkdown": "From my reading of xgboost documentation I didn't see any special handling of unordered categorical variables. In any case, many Tree algorithms will treat a categorical variable as ordered, which on the face of it seems bad.\n\nOn the other hand, by splitting up your one categorical variable into a (large?) collection of one-hot encoded values (e.g. Booleans) will have taken a compact variable and turned it into a whole bunch of sparse variables.\n\n**The potential (and may be even likely result) is that a Tree algorithm will down play and even ignore the one-hot encoded features especially there are a lot of them.** This is because Trees like variables that have high recall over those that have high precision but low recall.\n\nThat is, a Tree tries to increase information gain on the data it's splitting at any given level. Unintuitively, it doesn't like to use sparse variables to do this.\n\nThe upshot is you are in a real quandary. I suggest that you try it both ways to see which works better…. However, not using one-hot encoding might not be as bad as you think because a Tree might treat the categorical feature as a highly multimodal variable and make decent splits anyway.\n\nOne can try a few other approaches:\n\nlook at how the response variable responds to the categorical values and try to group them.\nFind another ML algorithm that works better with categorical features or with one-hot encoding and use that to train a submodel that just uses the categorical features. Then replace the categorical feature with a probability score. For instance, use a Logistic Regression on the hot-encoded values.\nTry to combine the categorical feature with some other features.\nBuild N xgboost classifiers, one for each category.\nThis may require playing around with the data a bit. Plotting the data may help you see patterns that you didn't know that were there.\n\n[added]\n\nLooks like that xgboost does only deal with numeric values:\n\n[Understand your dataset with XGBoost][1]\n\n[added]\n\nAn example of where I built N models, was when I built one for each major human language (English, French, etc) for web page classification.\n\n\n  [1]: https://cran.r-project.org/web/packages/xgboost/vignettes/discoverYourData.html",
      "votes": null
    },
    {
      "id": "330888",
      "postDate": "05/20/2018 01:11:13",
      "content": "<p>Really thanks for your help, DUO. Mark your thorough comment as my precious thing.</p>",
      "rawMarkdown": "Really thanks for your help, DUO. Mark your thorough comment as my precious thing.",
      "votes": null
    },
    {
      "id": "437016",
      "postDate": "12/11/2018 08:55:36",
      "content": "<p>how about label encoder?\nthis method just convert a category feature from string type into float type,e.g. \"big\":1,\"small\":2,\"middle\":3,etc.. an unordered list~\ni think the tree will train a piecewise function to fit these category feature...just guess.</p>",
      "rawMarkdown": "how about label encoder?\nthis method just convert a category feature from string type into float type,e.g. \"big\":1,\"small\":2,\"middle\":3,etc.. an unordered list~\ni think the tree will train a piecewise function to fit these category feature...just guess.",
      "votes": null
    },
    {
      "id": "443642",
      "postDate": "12/22/2018 02:40:16",
      "content": "<p>What do you mean by,  \"Then replace the categorical feature with a probability score. For instance, use a Logistic Regression on the hot-encoded values.\"? </p>\n\n<p>Currently I'm using the Random Forest Model to handle the one-hot-encoding. Is it required to use the Logistic Regression? My one-hot-encoded values are what I used as label and I'm trying to use it to replace the categorical feature in my data so that I can use it in the XGBRegressor model. Is that possible? </p>",
      "rawMarkdown": "What do you mean by,  \"Then replace the categorical feature with a probability score. For instance, use a Logistic Regression on the hot-encoded values.\"? \n\nCurrently I'm using the Random Forest Model to handle the one-hot-encoding. Is it required to use the Logistic Regression? My one-hot-encoded values are what I used as label and I'm trying to use it to replace the categorical feature in my data so that I can use it in the XGBRegressor model. Is that possible?",
      "votes": null
    },
    {
      "id": "484477",
      "postDate": "03/06/2019 04:21:30",
      "content": "<p>LightGBM has a parameter categorical_feature which takes a list of index of columns which are categorical.</p>",
      "rawMarkdown": "LightGBM has a parameter categorical_feature which takes a list of index of columns which are categorical.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 330696,
      "author_name": "classtag",
      "author_url": "",
      "post_date": "05/19/2018 13:32:00",
      "content": "<p>Hi @Kha, You can try these feature engineering as following link replace one-hot for categorical features. It is very excellent for your model performance and the finally score at LB.</p>\n\n<p><a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/55521\">https://www.kaggle.com/c/avito-demand-prediction/discussion/55521</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/57119\">https://www.kaggle.com/c/avito-demand-prediction/discussion/57119</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 330701,
      "author_name": "khahuras",
      "author_url": "",
      "post_date": "05/19/2018 13:38:17",
      "content": "<p>Thanks DUO for your suggestion. I've skimmed through that link, but still not answer my original question. I notice some kernels proposed here using XGB (on Python) didn't apply any encoding method for categorical features. So I just wonder why XGB knows? Or the kernel authors forgot about this important issue?</p>",
      "votes": null,
      "replies": [
        {
          "id": 330709,
          "author_name": "classtag",
          "author_url": "",
          "post_date": "05/19/2018 14:02:25",
          "content": "<p>From my reading of xgboost documentation I didn't see any special handling of unordered categorical variables. In any case, many Tree algorithms will treat a categorical variable as ordered, which on the face of it seems bad.</p>\n\n<p>On the other hand, by splitting up your one categorical variable into a (large?) collection of one-hot encoded values (e.g. Booleans) will have taken a compact variable and turned it into a whole bunch of sparse variables.</p>\n\n<p><strong>The potential (and may be even likely result) is that a Tree algorithm will down play and even ignore the one-hot encoded features especially there are a lot of them.</strong> This is because Trees like variables that have high recall over those that have high precision but low recall.</p>\n\n<p>That is, a Tree tries to increase information gain on the data it's splitting at any given level. Unintuitively, it doesn't like to use sparse variables to do this.</p>\n\n<p>The upshot is you are in a real quandary. I suggest that you try it both ways to see which works better…. However, not using one-hot encoding might not be as bad as you think because a Tree might treat the categorical feature as a highly multimodal variable and make decent splits anyway.</p>\n\n<p>One can try a few other approaches:</p>\n\n<p>look at how the response variable responds to the categorical values and try to group them.\nFind another ML algorithm that works better with categorical features or with one-hot encoding and use that to train a submodel that just uses the categorical features. Then replace the categorical feature with a probability score. For instance, use a Logistic Regression on the hot-encoded values.\nTry to combine the categorical feature with some other features.\nBuild N xgboost classifiers, one for each category.\nThis may require playing around with the data a bit. Plotting the data may help you see patterns that you didn't know that were there.</p>\n\n<p>[added]</p>\n\n<p>Looks like that xgboost does only deal with numeric values:</p>\n\n<p><a href=\"https://cran.r-project.org/web/packages/xgboost/vignettes/discoverYourData.html\">Understand your dataset with XGBoost</a></p>\n\n<p>[added]</p>\n\n<p>An example of where I built N models, was when I built one for each major human language (English, French, etc) for web page classification.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 330888,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "05/20/2018 01:11:13",
          "content": "<p>Really thanks for your help, DUO. Mark your thorough comment as my precious thing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443642,
          "author_name": "interestedmike",
          "author_url": "",
          "post_date": "12/22/2018 02:40:16",
          "content": "<p>What do you mean by,  \"Then replace the categorical feature with a probability score. For instance, use a Logistic Regression on the hot-encoded values.\"? </p>\n\n<p>Currently I'm using the Random Forest Model to handle the one-hot-encoding. Is it required to use the Logistic Regression? My one-hot-encoded values are what I used as label and I'm trying to use it to replace the categorical feature in my data so that I can use it in the XGBRegressor model. Is that possible? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 437016,
      "author_name": "novioleo",
      "author_url": "",
      "post_date": "12/11/2018 08:55:36",
      "content": "<p>how about label encoder?\nthis method just convert a category feature from string type into float type,e.g. \"big\":1,\"small\":2,\"middle\":3,etc.. an unordered list~\ni think the tree will train a piecewise function to fit these category feature...just guess.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 484477,
      "author_name": "kpriyanshu256",
      "author_url": "",
      "post_date": "03/06/2019 04:21:30",
      "content": "<p>LightGBM has a parameter categorical_feature which takes a list of index of columns which are categorical.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "330520": "Hi all,\n\nI wonder how can we specify categorical features in XGB (on Python) without using one-hot encoding? I noticed several kernels using XGB but didn't see any one-hot encoding, so how can XGB differentiate between numerical and categorical features?  To my knowledge, XGB does not support categorical features and it requires us to one-hot encode all categorical features.",
    "330696": "Hi @Kha, You can try these feature engineering as following link replace one-hot for categorical features. It is very excellent for your model performance and the finally score at LB.\n\nhttps://www.kaggle.com/c/avito-demand-prediction/discussion/55521\n\nhttps://www.kaggle.com/c/avito-demand-prediction/discussion/57119",
    "330701": "Thanks DUO for your suggestion. I've skimmed through that link, but still not answer my original question. I notice some kernels proposed here using XGB (on Python) didn't apply any encoding method for categorical features. So I just wonder why XGB knows? Or the kernel authors forgot about this important issue?",
    "330709": "From my reading of xgboost documentation I didn't see any special handling of unordered categorical variables. In any case, many Tree algorithms will treat a categorical variable as ordered, which on the face of it seems bad.\n\nOn the other hand, by splitting up your one categorical variable into a (large?) collection of one-hot encoded values (e.g. Booleans) will have taken a compact variable and turned it into a whole bunch of sparse variables.\n\n**The potential (and may be even likely result) is that a Tree algorithm will down play and even ignore the one-hot encoded features especially there are a lot of them.** This is because Trees like variables that have high recall over those that have high precision but low recall.\n\nThat is, a Tree tries to increase information gain on the data it's splitting at any given level. Unintuitively, it doesn't like to use sparse variables to do this.\n\nThe upshot is you are in a real quandary. I suggest that you try it both ways to see which works better…. However, not using one-hot encoding might not be as bad as you think because a Tree might treat the categorical feature as a highly multimodal variable and make decent splits anyway.\n\nOne can try a few other approaches:\n\nlook at how the response variable responds to the categorical values and try to group them.\nFind another ML algorithm that works better with categorical features or with one-hot encoding and use that to train a submodel that just uses the categorical features. Then replace the categorical feature with a probability score. For instance, use a Logistic Regression on the hot-encoded values.\nTry to combine the categorical feature with some other features.\nBuild N xgboost classifiers, one for each category.\nThis may require playing around with the data a bit. Plotting the data may help you see patterns that you didn't know that were there.\n\n[added]\n\nLooks like that xgboost does only deal with numeric values:\n\n[Understand your dataset with XGBoost][1]\n\n[added]\n\nAn example of where I built N models, was when I built one for each major human language (English, French, etc) for web page classification.\n\n\n  [1]: https://cran.r-project.org/web/packages/xgboost/vignettes/discoverYourData.html",
    "330888": "Really thanks for your help, DUO. Mark your thorough comment as my precious thing.",
    "437016": "how about label encoder?\nthis method just convert a category feature from string type into float type,e.g. \"big\":1,\"small\":2,\"middle\":3,etc.. an unordered list~\ni think the tree will train a piecewise function to fit these category feature...just guess.",
    "443642": "What do you mean by,  \"Then replace the categorical feature with a probability score. For instance, use a Logistic Regression on the hot-encoded values.\"? \n\nCurrently I'm using the Random Forest Model to handle the one-hot-encoding. Is it required to use the Logistic Regression? My one-hot-encoded values are what I used as label and I'm trying to use it to replace the categorical feature in my data so that I can use it in the XGBRegressor model. Is that possible?",
    "484477": "LightGBM has a parameter categorical_feature which takes a list of index of columns which are categorical."
  },
  "source": "meta"
}