{
  "id": 336911,
  "title": "Label Encoding is NOT essential for the LightGBM",
  "url": "/competitions/amex-default-prediction/discussion/336911",
  "author_name": "",
  "post_date": "2022-07-13T14:14:33.626719300Z",
  "votes": 14,
  "comment_count": 6,
  "views": 0,
  "content": "<h2>Overview</h2>\n<p>There are several methods to treat categorical features in LGBM.<br>\nAt the notebook which generate the current best score (0.799), categorical_features are specified as the arguments of LightGBM with Label Encoding.<br>\n<a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977</a></p>\n<p>However, I suppose one of them is enough to treat categorical features in LGBM.<br>\nIn other words, Label Encoding is not necessary if we specify categorical features.</p>\n<p>Then, I've checked the difference among the following methods:</p>\n<ol>\n<li>Label encoding + specify categorical_features</li>\n<li>Label encoding only</li>\n<li>Specify categorical_features only</li>\n</ol>\n<p>As a result, all methods lead the same results (same score).</p>\n<h2>More Detail Description</h2>\n<h3>1. Label encoding + specify categorical_features</h3>\n<ul>\n<li>Same as current best score notebook</li>\n</ul>\n<p>Example:</p>\n<pre><code>for cat_col in cat_features:\n    encoder = LabelEncoder()\n    train[cat_col] = encoder.fit_transform(train[cat_col])\n    test[cat_col] = encoder.transform(test[cat_col])\n    ...\n    lgb_train = lgb.Dataset(x_train, y_train, categorical_feature = cat_features)\n</code></pre>\n<h3>2. Label encoding only</h3>\n<ul>\n<li>Do not specify categorical features</li>\n</ul>\n<p>Example:</p>\n<pre><code>for cat_col in cat_features:\n    encoder = LabelEncoder()\n    train[cat_col] = encoder.fit_transform(train[cat_col])\n    test[cat_col] = encoder.transform(test[cat_col])\n    ...\n    lgb_train = lgb.Dataset(x_train, y_train)\n</code></pre>\n<h3>3. Specify categorical_features without label encoding</h3>\n<ul>\n<li>Categorical features are directly used</li>\n</ul>\n<pre><code>for cat_col in cat_features:\n    train[cat_col] = train[cat_col].astype('category')\n    test[cat_col] = test[cat_col].astype('category')\n    ...\n    lgb_train = lgb.Dataset(x_train, y_train, categorical_feature = cat_features)\n</code></pre>\n<h2>Note</h2>\n<p>Originally, GBDT can't treat categorical features.<br>\nTherefore, these features should be encoded to numerical values.<br>\nLabel encoding is good method to encode these values.<br>\nOn the other hand, we can specify the categorical features for LightGBM.<br>\nIf we specify these features, categorical values can be directly used as a features.</p>",
  "messages": [
    {
      "id": "1854221",
      "postDate": "07/13/2022 14:14:33",
      "content": "<h2>Overview</h2>\n<p>There are several methods to treat categorical features in LGBM.<br>\nAt the notebook which generate the current best score (0.799), categorical_features are specified as the arguments of LightGBM with Label Encoding.<br>\n<a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977</a></p>\n<p>However, I suppose one of them is enough to treat categorical features in LGBM.<br>\nIn other words, Label Encoding is not necessary if we specify categorical features.</p>\n<p>Then, I've checked the difference among the following methods:</p>\n<ol>\n<li>Label encoding + specify categorical_features</li>\n<li>Label encoding only</li>\n<li>Specify categorical_features only</li>\n</ol>\n<p>As a result, all methods lead the same results (same score).</p>\n<h2>More Detail Description</h2>\n<h3>1. Label encoding + specify categorical_features</h3>\n<ul>\n<li>Same as current best score notebook</li>\n</ul>\n<p>Example:</p>\n<pre><code>for cat_col in cat_features:\n    encoder = LabelEncoder()\n    train[cat_col] = encoder.fit_transform(train[cat_col])\n    test[cat_col] = encoder.transform(test[cat_col])\n    ...\n    lgb_train = lgb.Dataset(x_train, y_train, categorical_feature = cat_features)\n</code></pre>\n<h3>2. Label encoding only</h3>\n<ul>\n<li>Do not specify categorical features</li>\n</ul>\n<p>Example:</p>\n<pre><code>for cat_col in cat_features:\n    encoder = LabelEncoder()\n    train[cat_col] = encoder.fit_transform(train[cat_col])\n    test[cat_col] = encoder.transform(test[cat_col])\n    ...\n    lgb_train = lgb.Dataset(x_train, y_train)\n</code></pre>\n<h3>3. Specify categorical_features without label encoding</h3>\n<ul>\n<li>Categorical features are directly used</li>\n</ul>\n<pre><code>for cat_col in cat_features:\n    train[cat_col] = train[cat_col].astype('category')\n    test[cat_col] = test[cat_col].astype('category')\n    ...\n    lgb_train = lgb.Dataset(x_train, y_train, categorical_feature = cat_features)\n</code></pre>\n<h2>Note</h2>\n<p>Originally, GBDT can't treat categorical features.<br>\nTherefore, these features should be encoded to numerical values.<br>\nLabel encoding is good method to encode these values.<br>\nOn the other hand, we can specify the categorical features for LightGBM.<br>\nIf we specify these features, categorical values can be directly used as a features.</p>",
      "rawMarkdown": "## Overview\nThere are several methods to treat categorical features in LGBM.\nAt the notebook which generate the current best score (0.799), categorical_features are specified as the arguments of LightGBM with Label Encoding.\nhttps://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\n\nHowever, I suppose one of them is enough to treat categorical features in LGBM.\nIn other words, Label Encoding is not necessary if we specify categorical features.\n\nThen, I've checked the difference among the following methods:\n\n1. Label encoding + specify categorical_features\n2. Label encoding only\n3. Specify categorical_features only\n\nAs a result, all methods lead the same results (same score).\n\n\n## More Detail Description\n\n### 1. Label encoding + specify categorical_features\n- Same as current best score notebook\n\nExample:\n```\nfor cat_col in cat_features:\n    encoder = LabelEncoder()\n    train[cat_col] = encoder.fit_transform(train[cat_col])\n    test[cat_col] = encoder.transform(test[cat_col])\n    ...\n    lgb_train = lgb.Dataset(x_train, y_train, categorical_feature = cat_features)\n```\n\n### 2. Label encoding only\n- Do not specify categorical features\n\nExample:\n```\nfor cat_col in cat_features:\n    encoder = LabelEncoder()\n    train[cat_col] = encoder.fit_transform(train[cat_col])\n    test[cat_col] = encoder.transform(test[cat_col])\n    ...\n    lgb_train = lgb.Dataset(x_train, y_train)\n```\n\n### 3. Specify categorical_features without label encoding\n\n- Categorical features are directly used\n\n```\nfor cat_col in cat_features:\n    train[cat_col] = train[cat_col].astype('category')\n    test[cat_col] = test[cat_col].astype('category')\n    ...\n    lgb_train = lgb.Dataset(x_train, y_train, categorical_feature = cat_features)\n```\n\n\n## Note\nOriginally, GBDT can't treat categorical features.\nTherefore, these features should be encoded to numerical values.\nLabel encoding is good method to encode these values.\nOn the other hand, we can specify the categorical features for LightGBM.\nIf we specify these features, categorical values can be directly used as a features.",
      "votes": null
    },
    {
      "id": "1854386",
      "postDate": "07/13/2022 16:55:15",
      "content": "<p>nice work! thanks for sharing!</p>",
      "rawMarkdown": "nice work! thanks for sharing!",
      "votes": null
    },
    {
      "id": "1854391",
      "postDate": "07/13/2022 16:56:56",
      "content": "<p>Good work! This is quite informative, thanks for sharing!</p>",
      "rawMarkdown": "Good work! This is quite informative, thanks for sharing!",
      "votes": null
    },
    {
      "id": "1854516",
      "postDate": "07/13/2022 18:52:25",
      "content": "<p>If using raddar's clean data, I believe there's one and only one difference between #1 and #3.<br>\nRaddar used '-1' to encode np.nan from the original data.<br>\n1: Treats nan as its own category.<br>\n3: LightGBM issues warning about negative number, and forces to… nan! Depending on how the LightGBM model handles nan under the hood, this could get better or worse performance than #1. Since it matches the original data, it seems plausible that this could do very marginally better.</p>\n<p>2: I don't recommend this, if I understand correctly it now treats as a range, and it will make a big performance difference whether category 'really_good' was labeled between 'really_bad' and 'also_really_bad', or labeled to one end. Of course if you have endless cycles you could try all the permutations… But I assume that #1 and #3 are much smarter out of the box.</p>",
      "rawMarkdown": "If using raddar's clean data, I believe there's one and only one difference between #1 and #3.\nRaddar used '-1' to encode np.nan from the original data.\n1: Treats nan as its own category.\n3: LightGBM issues warning about negative number, and forces to... nan! Depending on how the LightGBM model handles nan under the hood, this could get better or worse performance than #1. Since it matches the original data, it seems plausible that this could do very marginally better.\n\n2: I don't recommend this, if I understand correctly it now treats as a range, and it will make a big performance difference whether category 'really_good' was labeled between 'really_bad' and 'also_really_bad', or labeled to one end. Of course if you have endless cycles you could try all the permutations... But I assume that #1 and #3 are much smarter out of the box.",
      "votes": null
    },
    {
      "id": "1856727",
      "postDate": "07/15/2022 14:46:11",
      "content": "<p>Thank you.<br>\nI'm glad to hear that.</p>",
      "rawMarkdown": "Thank you.\nI'm glad to hear that.",
      "votes": null
    },
    {
      "id": "1856729",
      "postDate": "07/15/2022 14:46:27",
      "content": "<p>Thank you so much!</p>",
      "rawMarkdown": "Thank you so much!",
      "votes": null
    },
    {
      "id": "1856743",
      "postDate": "07/15/2022 15:00:43",
      "content": "<p>Thank you for your comments.<br>\nRegarding 1 and 2, as you said, we should replace NaN to the different category at first which is not written in my example codes.<br>\nIf we don't replace the NaN value,  <code>train[cat_col].astype('category')</code> failed.</p>\n<p>As for 3, theoretically, decision tree can treat label encoded categorical values.<br>\nFor example, there are 3 label encoded categories which are 1, 2 and 3.<br>\nThe \"2\" category can be extracted by applying <code>value &lt; 2</code> and <code>2 &lt; value</code> blanches.</p>\n<p>Therefore, we can't conclude 3rd method is not recommended and I think it depends on the data.</p>",
      "rawMarkdown": "Thank you for your comments.\nRegarding 1 and 2, as you said, we should replace NaN to the different category at first which is not written in my example codes.\nIf we don't replace the NaN value,  `train[cat_col].astype('category')` failed.\n\nAs for 3, theoretically, decision tree can treat label encoded categorical values.\nFor example, there are 3 label encoded categories which are 1, 2 and 3.\nThe \"2\" category can be extracted by applying `value < 2` and `2 < value` blanches.\n\nTherefore, we can't conclude 3rd method is not recommended and I think it depends on the data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1854386,
      "author_name": "mohammadrahmati",
      "author_url": "",
      "post_date": "07/13/2022 16:55:15",
      "content": "<p>nice work! thanks for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1856727,
          "author_name": "tetsuro731",
          "author_url": "",
          "post_date": "07/15/2022 14:46:11",
          "content": "<p>Thank you.<br>\nI'm glad to hear that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1854391,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "07/13/2022 16:56:56",
      "content": "<p>Good work! This is quite informative, thanks for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1856729,
          "author_name": "tetsuro731",
          "author_url": "",
          "post_date": "07/15/2022 14:46:27",
          "content": "<p>Thank you so much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1854516,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "07/13/2022 18:52:25",
      "content": "<p>If using raddar's clean data, I believe there's one and only one difference between #1 and #3.<br>\nRaddar used '-1' to encode np.nan from the original data.<br>\n1: Treats nan as its own category.<br>\n3: LightGBM issues warning about negative number, and forces to… nan! Depending on how the LightGBM model handles nan under the hood, this could get better or worse performance than #1. Since it matches the original data, it seems plausible that this could do very marginally better.</p>\n<p>2: I don't recommend this, if I understand correctly it now treats as a range, and it will make a big performance difference whether category 'really_good' was labeled between 'really_bad' and 'also_really_bad', or labeled to one end. Of course if you have endless cycles you could try all the permutations… But I assume that #1 and #3 are much smarter out of the box.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1856743,
          "author_name": "tetsuro731",
          "author_url": "",
          "post_date": "07/15/2022 15:00:43",
          "content": "<p>Thank you for your comments.<br>\nRegarding 1 and 2, as you said, we should replace NaN to the different category at first which is not written in my example codes.<br>\nIf we don't replace the NaN value,  <code>train[cat_col].astype('category')</code> failed.</p>\n<p>As for 3, theoretically, decision tree can treat label encoded categorical values.<br>\nFor example, there are 3 label encoded categories which are 1, 2 and 3.<br>\nThe \"2\" category can be extracted by applying <code>value &lt; 2</code> and <code>2 &lt; value</code> blanches.</p>\n<p>Therefore, we can't conclude 3rd method is not recommended and I think it depends on the data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1854221": "## Overview\nThere are several methods to treat categorical features in LGBM.\nAt the notebook which generate the current best score (0.799), categorical_features are specified as the arguments of LightGBM with Label Encoding.\nhttps://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\n\nHowever, I suppose one of them is enough to treat categorical features in LGBM.\nIn other words, Label Encoding is not necessary if we specify categorical features.\n\nThen, I've checked the difference among the following methods:\n\n1. Label encoding + specify categorical_features\n2. Label encoding only\n3. Specify categorical_features only\n\nAs a result, all methods lead the same results (same score).\n\n\n## More Detail Description\n\n### 1. Label encoding + specify categorical_features\n- Same as current best score notebook\n\nExample:\n```\nfor cat_col in cat_features:\n    encoder = LabelEncoder()\n    train[cat_col] = encoder.fit_transform(train[cat_col])\n    test[cat_col] = encoder.transform(test[cat_col])\n    ...\n    lgb_train = lgb.Dataset(x_train, y_train, categorical_feature = cat_features)\n```\n\n### 2. Label encoding only\n- Do not specify categorical features\n\nExample:\n```\nfor cat_col in cat_features:\n    encoder = LabelEncoder()\n    train[cat_col] = encoder.fit_transform(train[cat_col])\n    test[cat_col] = encoder.transform(test[cat_col])\n    ...\n    lgb_train = lgb.Dataset(x_train, y_train)\n```\n\n### 3. Specify categorical_features without label encoding\n\n- Categorical features are directly used\n\n```\nfor cat_col in cat_features:\n    train[cat_col] = train[cat_col].astype('category')\n    test[cat_col] = test[cat_col].astype('category')\n    ...\n    lgb_train = lgb.Dataset(x_train, y_train, categorical_feature = cat_features)\n```\n\n\n## Note\nOriginally, GBDT can't treat categorical features.\nTherefore, these features should be encoded to numerical values.\nLabel encoding is good method to encode these values.\nOn the other hand, we can specify the categorical features for LightGBM.\nIf we specify these features, categorical values can be directly used as a features.",
    "1854386": "nice work! thanks for sharing!",
    "1854391": "Good work! This is quite informative, thanks for sharing!",
    "1854516": "If using raddar's clean data, I believe there's one and only one difference between #1 and #3.\nRaddar used '-1' to encode np.nan from the original data.\n1: Treats nan as its own category.\n3: LightGBM issues warning about negative number, and forces to... nan! Depending on how the LightGBM model handles nan under the hood, this could get better or worse performance than #1. Since it matches the original data, it seems plausible that this could do very marginally better.\n\n2: I don't recommend this, if I understand correctly it now treats as a range, and it will make a big performance difference whether category 'really_good' was labeled between 'really_bad' and 'also_really_bad', or labeled to one end. Of course if you have endless cycles you could try all the permutations... But I assume that #1 and #3 are much smarter out of the box.",
    "1856727": "Thank you.\nI'm glad to hear that.",
    "1856729": "Thank you so much!",
    "1856743": "Thank you for your comments.\nRegarding 1 and 2, as you said, we should replace NaN to the different category at first which is not written in my example codes.\nIf we don't replace the NaN value,  `train[cat_col].astype('category')` failed.\n\nAs for 3, theoretically, decision tree can treat label encoded categorical values.\nFor example, there are 3 label encoded categories which are 1, 2 and 3.\nThe \"2\" category can be extracted by applying `value < 2` and `2 < value` blanches.\n\nTherefore, we can't conclude 3rd method is not recommended and I think it depends on the data."
  },
  "source": "meta"
}