{
  "id": 331565,
  "title": "LabelEncoder vs OrdinalEncoder?",
  "url": "/competitions/amex-default-prediction/discussion/331565",
  "author_name": "",
  "post_date": "2022-06-17T23:32:05.625848300Z",
  "votes": 16,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Some public notebooks are using LabelEncoder to handle categorical encoding of the predictors. However, there is no option to handle unseen categories, which may pop up on the test set. I'm assuming the notebook would fail to execute in this case.</p>\n<p>The documentation of LabelEncoder says </p>\n<blockquote>\n  <p>This transformer should be used to encode target values, i.e. y, and not the input X.</p>\n</blockquote>\n<p><a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html\" target=\"_blank\">https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html</a></p>\n<p>OrdinalEncoder seems to be more suitable for this, it has additional parameters to handle unseen categories. Although the name does not sound intuitive to use for categories with no inherent ordering</p>\n<p><a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.OrdinalEncoder.html\" target=\"_blank\">https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.OrdinalEncoder.html</a></p>\n<p>Look forward to any feedback on this! </p>",
  "messages": [
    {
      "id": "1824040",
      "postDate": "06/17/2022 23:32:05",
      "content": "<p>Some public notebooks are using LabelEncoder to handle categorical encoding of the predictors. However, there is no option to handle unseen categories, which may pop up on the test set. I'm assuming the notebook would fail to execute in this case.</p>\n<p>The documentation of LabelEncoder says </p>\n<blockquote>\n  <p>This transformer should be used to encode target values, i.e. y, and not the input X.</p>\n</blockquote>\n<p><a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html\" target=\"_blank\">https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html</a></p>\n<p>OrdinalEncoder seems to be more suitable for this, it has additional parameters to handle unseen categories. Although the name does not sound intuitive to use for categories with no inherent ordering</p>\n<p><a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.OrdinalEncoder.html\" target=\"_blank\">https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.OrdinalEncoder.html</a></p>\n<p>Look forward to any feedback on this! </p>",
      "rawMarkdown": "Some public notebooks are using LabelEncoder to handle categorical encoding of the predictors. However, there is no option to handle unseen categories, which may pop up on the test set. I'm assuming the notebook would fail to execute in this case.\n\nThe documentation of LabelEncoder says \n> This transformer should be used to encode target values, i.e. y, and not the input X.\n\nhttps://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html\n\nOrdinalEncoder seems to be more suitable for this, it has additional parameters to handle unseen categories. Although the name does not sound intuitive to use for categories with no inherent ordering\n\nhttps://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.OrdinalEncoder.html\n\n\nLook forward to any feedback on this!",
      "votes": null
    },
    {
      "id": "1824250",
      "postDate": "06/18/2022 05:53:54",
      "content": "<p>Good comment! Great observation!</p>",
      "rawMarkdown": "Good comment! Great observation!",
      "votes": null
    },
    {
      "id": "1824391",
      "postDate": "06/18/2022 09:10:18",
      "content": "<p>What you're saying is correct, but it matters only if you do not know which categories will occur in the test data. This competition is special: We've got all test data in advance and we know that there won't be any unseen categories.</p>\n<p>If you're practicing for real life machine learning, use OrdinalEncoder. If you only want to win the competition, you can use either Encoder.</p>",
      "rawMarkdown": "What you're saying is correct, but it matters only if you do not know which categories will occur in the test data. This competition is special: We've got all test data in advance and we know that there won't be any unseen categories.\n\nIf you're practicing for real life machine learning, use OrdinalEncoder. If you only want to win the competition, you can use either Encoder.",
      "votes": null
    },
    {
      "id": "1825117",
      "postDate": "06/19/2022 00:39:20",
      "content": "<p>As its name implies, <code>LabelEncoder</code> is meant to encode labels (target values), not features. Using it on features is a kind of abuse, even though in many cases there are no harmful effects. </p>",
      "rawMarkdown": "As its name implies, `LabelEncoder` is meant to encode labels (target values), not features. Using it on features is a kind of abuse, even though in many cases there are no harmful effects.",
      "votes": null
    },
    {
      "id": "1825205",
      "postDate": "06/19/2022 04:55:25",
      "content": "<p>That makes complete sense! Thanks for your input :)</p>",
      "rawMarkdown": "That makes complete sense! Thanks for your input :)",
      "votes": null
    },
    {
      "id": "1825206",
      "postDate": "06/19/2022 04:58:33",
      "content": "<p>I think <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> hit the nail on the head as to why there would be no problem using it in this competition</p>",
      "rawMarkdown": "I think @ambrosm hit the nail on the head as to why there would be no problem using it in this competition",
      "votes": null
    },
    {
      "id": "1825207",
      "postDate": "06/19/2022 04:58:51",
      "content": "<p>Thank you sir!</p>",
      "rawMarkdown": "Thank you sir!",
      "votes": null
    },
    {
      "id": "1838946",
      "postDate": "07/01/2022 02:45:24",
      "content": "<p>AmbrosM already answered your question, so I would like to only add that models like XGBoost or LightGBM would work better with Label encoders compared to one hot encoding.</p>",
      "rawMarkdown": "AmbrosM already answered your question, so I would like to only add that models like XGBoost or LightGBM would work better with Label encoders compared to one hot encoding.",
      "votes": null
    },
    {
      "id": "1983156",
      "postDate": "10/11/2022 20:19:54",
      "content": "<p>hi <a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a> , could you clarify further why using LabelEncoder performs better than dummies in XGB/LGBM? I've been in a competition before where I got way better performance (F1-score was the metric) in some features OneHotEncoding them.</p>",
      "rawMarkdown": "hi @mohammadrahmati , could you clarify further why using LabelEncoder performs better than dummies in XGB/LGBM? I've been in a competition before where I got way better performance (F1-score was the metric) in some features OneHotEncoding them.",
      "votes": null
    },
    {
      "id": "2287490",
      "postDate": "06/04/2023 14:26:30",
      "content": "<p>I believe that to deal with missing values, <strong>imputation</strong> is more suitable.</p>",
      "rawMarkdown": "I believe that to deal with missing values, **imputation** is more suitable.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1824250,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "06/18/2022 05:53:54",
      "content": "<p>Good comment! Great observation!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1825207,
          "author_name": "illidan7",
          "author_url": "",
          "post_date": "06/19/2022 04:58:51",
          "content": "<p>Thank you sir!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1824391,
      "author_name": "ambrosm",
      "author_url": "",
      "post_date": "06/18/2022 09:10:18",
      "content": "<p>What you're saying is correct, but it matters only if you do not know which categories will occur in the test data. This competition is special: We've got all test data in advance and we know that there won't be any unseen categories.</p>\n<p>If you're practicing for real life machine learning, use OrdinalEncoder. If you only want to win the competition, you can use either Encoder.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1825205,
          "author_name": "illidan7",
          "author_url": "",
          "post_date": "06/19/2022 04:55:25",
          "content": "<p>That makes complete sense! Thanks for your input :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1825117,
      "author_name": "siukeitin",
      "author_url": "",
      "post_date": "06/19/2022 00:39:20",
      "content": "<p>As its name implies, <code>LabelEncoder</code> is meant to encode labels (target values), not features. Using it on features is a kind of abuse, even though in many cases there are no harmful effects. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1825206,
          "author_name": "illidan7",
          "author_url": "",
          "post_date": "06/19/2022 04:58:33",
          "content": "<p>I think <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> hit the nail on the head as to why there would be no problem using it in this competition</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1838946,
      "author_name": "mohammadrahmati",
      "author_url": "",
      "post_date": "07/01/2022 02:45:24",
      "content": "<p>AmbrosM already answered your question, so I would like to only add that models like XGBoost or LightGBM would work better with Label encoders compared to one hot encoding.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1983156,
          "author_name": "leandrodestefani",
          "author_url": "",
          "post_date": "10/11/2022 20:19:54",
          "content": "<p>hi <a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a> , could you clarify further why using LabelEncoder performs better than dummies in XGB/LGBM? I've been in a competition before where I got way better performance (F1-score was the metric) in some features OneHotEncoding them.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2287490,
      "author_name": "willhain",
      "author_url": "",
      "post_date": "06/04/2023 14:26:30",
      "content": "<p>I believe that to deal with missing values, <strong>imputation</strong> is more suitable.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1824040": "Some public notebooks are using LabelEncoder to handle categorical encoding of the predictors. However, there is no option to handle unseen categories, which may pop up on the test set. I'm assuming the notebook would fail to execute in this case.\n\nThe documentation of LabelEncoder says \n> This transformer should be used to encode target values, i.e. y, and not the input X.\n\nhttps://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html\n\nOrdinalEncoder seems to be more suitable for this, it has additional parameters to handle unseen categories. Although the name does not sound intuitive to use for categories with no inherent ordering\n\nhttps://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.OrdinalEncoder.html\n\n\nLook forward to any feedback on this!",
    "1824250": "Good comment! Great observation!",
    "1824391": "What you're saying is correct, but it matters only if you do not know which categories will occur in the test data. This competition is special: We've got all test data in advance and we know that there won't be any unseen categories.\n\nIf you're practicing for real life machine learning, use OrdinalEncoder. If you only want to win the competition, you can use either Encoder.",
    "1825117": "As its name implies, `LabelEncoder` is meant to encode labels (target values), not features. Using it on features is a kind of abuse, even though in many cases there are no harmful effects.",
    "1825205": "That makes complete sense! Thanks for your input :)",
    "1825206": "I think @ambrosm hit the nail on the head as to why there would be no problem using it in this competition",
    "1825207": "Thank you sir!",
    "1838946": "AmbrosM already answered your question, so I would like to only add that models like XGBoost or LightGBM would work better with Label encoders compared to one hot encoding.",
    "1983156": "hi @mohammadrahmati , could you clarify further why using LabelEncoder performs better than dummies in XGB/LGBM? I've been in a competition before where I got way better performance (F1-score was the metric) in some features OneHotEncoding them.",
    "2287490": "I believe that to deal with missing values, **imputation** is more suitable."
  },
  "source": "meta"
}