{
  "id": 59835,
  "title": "why lgbm get better cv score when I do not specify which columns is categorical feature",
  "url": "/competitions/avito-demand-prediction/discussion/59835",
  "author_name": "",
  "post_date": "2018-06-27T15:29:44.505768700Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>So if my lgbm input is numpy array, and also I specify the feature_name and categorical_features param in lgb.dataset, could it know which column in numpy array is categorical. I think it could do it if we keep all thing in an order, it will learn it from categoical feature in feature_name's index. </p>\n\n<p>But when I delete this params , I get better cv score, lgbm think all features columns as numerical . </p>\n\n<p>am I right ? and why ? thanks!</p>",
  "messages": [
    {
      "id": "348954",
      "postDate": "06/27/2018 15:29:44",
      "content": "<p>So if my lgbm input is numpy array, and also I specify the feature_name and categorical_features param in lgb.dataset, could it know which column in numpy array is categorical. I think it could do it if we keep all thing in an order, it will learn it from categoical feature in feature_name's index. </p>\n\n<p>But when I delete this params , I get better cv score, lgbm think all features columns as numerical . </p>\n\n<p>am I right ? and why ? thanks!</p>",
      "rawMarkdown": "So if my lgbm input is numpy array, and also I specify the feature_name and categorical_features param in lgb.dataset, could it know which column in numpy array is categorical. I think it could do it if we keep all thing in an order, it will learn it from categoical feature in feature_name's index. \n\nBut when I delete this params , I get better cv score, lgbm think all features columns as numerical . \n\nam I right ? and why ? thanks!",
      "votes": null
    },
    {
      "id": "349034",
      "postDate": "06/27/2018 17:29:49",
      "content": "<p>The same thing happened to me a few times.</p>\n\n<p>One thing to try out, though, is if it only affects your cv score or also the leaderboard score.</p>\n\n<p>If you have too many different categories (high cardinality), the model can have difficulties finding the differences between them if you force it to use specific categories. For low- to mid-cardinality features, specifying them as categories for LightGBM (also XGBoost and others) can help much more.</p>\n\n<p>In the end, like with many other parameters, it comes down to trial and error when you want to know which way works better.</p>",
      "rawMarkdown": "The same thing happened to me a few times.\n\nOne thing to try out, though, is if it only affects your cv score or also the leaderboard score.\n\nIf you have too many different categories (high cardinality), the model can have difficulties finding the differences between them if you force it to use specific categories. For low- to mid-cardinality features, specifying them as categories for LightGBM (also XGBoost and others) can help much more.\n\nIn the end, like with many other parameters, it comes down to trial and error when you want to know which way works better.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 349034,
      "author_name": "frankherfert",
      "author_url": "",
      "post_date": "06/27/2018 17:29:49",
      "content": "<p>The same thing happened to me a few times.</p>\n\n<p>One thing to try out, though, is if it only affects your cv score or also the leaderboard score.</p>\n\n<p>If you have too many different categories (high cardinality), the model can have difficulties finding the differences between them if you force it to use specific categories. For low- to mid-cardinality features, specifying them as categories for LightGBM (also XGBoost and others) can help much more.</p>\n\n<p>In the end, like with many other parameters, it comes down to trial and error when you want to know which way works better.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "348954": "So if my lgbm input is numpy array, and also I specify the feature_name and categorical_features param in lgb.dataset, could it know which column in numpy array is categorical. I think it could do it if we keep all thing in an order, it will learn it from categoical feature in feature_name's index. \n\nBut when I delete this params , I get better cv score, lgbm think all features columns as numerical . \n\nam I right ? and why ? thanks!",
    "349034": "The same thing happened to me a few times.\n\nOne thing to try out, though, is if it only affects your cv score or also the leaderboard score.\n\nIf you have too many different categories (high cardinality), the model can have difficulties finding the differences between them if you force it to use specific categories. For low- to mid-cardinality features, specifying them as categories for LightGBM (also XGBoost and others) can help much more.\n\nIn the end, like with many other parameters, it comes down to trial and error when you want to know which way works better."
  },
  "source": "meta"
}