{
  "id": 207307,
  "title": "Understanding LightGBM Feature Importance",
  "url": "/competitions/riiid-test-answer-prediction/discussion/207307",
  "author_name": "Manikanth Reddy",
  "post_date": "2020-12-29T06:19:42.471000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p><strong>tl;dr</strong>: By default LightGBM uses <strong>feature split</strong> for plotting feature importance. Check <strong>feature gain</strong> as well if you are using importance for selecting features. Feature split is bad if there is a combination of categorical and numerical features.</p>\n<hr>\n<p>I was going through <a href=\"https://www.kaggle.com/zyy2016/0-780-unoptimized-lgbm-interesting-features\" target=\"_blank\">this nice kernel</a>. It introduced the usage of <a href=\"https://www.microsoft.com/en-us/research/project/trueskill-ranking-system/\" target=\"_blank\">Microsoft True Skill Ranking System</a>. One thing I wanted to check was if True Skill features are adding value using feature importance plot. By default LightGBM uses feature split for calculating feature importance. </p>\n<h2>Feature Importance by Split</h2>\n<p><img src=\"https://i.imgur.com/pu2Y3KA.png\" alt=\"Feature Importance by Split\"></p>\n<ul>\n<li>Question ID is in the first place. This is due to the huge number of splits possible (17000+ unique values). Similar thing is also possible  for tag_1, tags_encoded. </li>\n<li>True Skill features are not in the bottom of the list. </li>\n<li>Some features can be very useful without too many splits: ex: std_deviation question accuracy, most_liked_guess_correct</li>\n<li>Some features may be leading to over-fitting due to too many splits: user_id</li>\n</ul>\n<h2>Feature Importance by Gain</h2>\n<p><img src=\"https://i.imgur.com/RVpcOYL.png\" alt=\"Feature Importance by Gain\"></p>\n<ul>\n<li>The first 2 features are from True Skill. It seems they are adding value. </li>\n<li>Remove the features that are not shown / very low gain in the gain plot.</li>\n</ul>\n<p>In summary: </p>\n<ul>\n<li>Try to use both split and gain feature importance plots. Remove the features that are not adding any value in both split and gain plots. They are only leading to overfitting. </li>\n<li>True Skill features are adding values over the common engineered features in Public notebooks. </li>\n</ul>\n<p>I have recently started using LightGBM. In case I misinterpreted the plots, please correct me. </p>\n<p>Acknowledgement: </p>\n<ul>\n<li><a href=\"https://www.kaggle.com/zyy2016/0-780-unoptimized-lgbm-interesting-features\" target=\"_blank\">[0.780]Unoptimized LGBM + Interesting Features</a></li>\n</ul>",
  "messages": [
    {
      "id": 1130576,
      "postDate": "2020-12-29T06:19:42.470Z",
      "content": "<p><strong>tl;dr</strong>: By default LightGBM uses <strong>feature split</strong> for plotting feature importance. Check <strong>feature gain</strong> as well if you are using importance for selecting features. Feature split is bad if there is a combination of categorical and numerical features.</p>\n<hr>\n<p>I was going through <a href=\"https://www.kaggle.com/zyy2016/0-780-unoptimized-lgbm-interesting-features\" target=\"_blank\">this nice kernel</a>. It introduced the usage of <a href=\"https://www.microsoft.com/en-us/research/project/trueskill-ranking-system/\" target=\"_blank\">Microsoft True Skill Ranking System</a>. One thing I wanted to check was if True Skill features are adding value using feature importance plot. By default LightGBM uses feature split for calculating feature importance. </p>\n<h2>Feature Importance by Split</h2>\n<p><img src=\"https://i.imgur.com/pu2Y3KA.png\" alt=\"Feature Importance by Split\"></p>\n<ul>\n<li>Question ID is in the first place. This is due to the huge number of splits possible (17000+ unique values). Similar thing is also possible  for tag_1, tags_encoded. </li>\n<li>True Skill features are not in the bottom of the list. </li>\n<li>Some features can be very useful without too many splits: ex: std_deviation question accuracy, most_liked_guess_correct</li>\n<li>Some features may be leading to over-fitting due to too many splits: user_id</li>\n</ul>\n<h2>Feature Importance by Gain</h2>\n<p><img src=\"https://i.imgur.com/RVpcOYL.png\" alt=\"Feature Importance by Gain\"></p>\n<ul>\n<li>The first 2 features are from True Skill. It seems they are adding value. </li>\n<li>Remove the features that are not shown / very low gain in the gain plot.</li>\n</ul>\n<p>In summary: </p>\n<ul>\n<li>Try to use both split and gain feature importance plots. Remove the features that are not adding any value in both split and gain plots. They are only leading to overfitting. </li>\n<li>True Skill features are adding values over the common engineered features in Public notebooks. </li>\n</ul>\n<p>I have recently started using LightGBM. In case I misinterpreted the plots, please correct me. </p>\n<p>Acknowledgement: </p>\n<ul>\n<li><a href=\"https://www.kaggle.com/zyy2016/0-780-unoptimized-lgbm-interesting-features\" target=\"_blank\">[0.780]Unoptimized LGBM + Interesting Features</a></li>\n</ul>",
      "rawMarkdown": "**tl;dr**: By default LightGBM uses **feature split** for plotting feature importance. Check **feature gain** as well if you are using importance for selecting features. Feature split is bad if there is a combination of categorical and numerical features.\n\n---\n\nI was going through [this nice kernel](https://www.kaggle.com/zyy2016/0-780-unoptimized-lgbm-interesting-features). It introduced the usage of [Microsoft True Skill Ranking System](https://www.microsoft.com/en-us/research/project/trueskill-ranking-system/). One thing I wanted to check was if True Skill features are adding value using feature importance plot. By default LightGBM uses feature split for calculating feature importance. \n\n## Feature Importance by Split\n![Feature Importance by Split](https://i.imgur.com/pu2Y3KA.png)\n\n- Question ID is in the first place. This is due to the huge number of splits possible (17000+ unique values). Similar thing is also possible  for tag_1, tags_encoded. \n- True Skill features are not in the bottom of the list. \n- Some features can be very useful without too many splits: ex: std_deviation question accuracy, most_liked_guess_correct\n- Some features may be leading to over-fitting due to too many splits: user_id\n\n## Feature Importance by Gain\n![Feature Importance by Gain](https://i.imgur.com/RVpcOYL.png)\n\n- The first 2 features are from True Skill. It seems they are adding value. \n- Remove the features that are not shown / very low gain in the gain plot.\n\nIn summary: \n- Try to use both split and gain feature importance plots. Remove the features that are not adding any value in both split and gain plots. They are only leading to overfitting. \n- True Skill features are adding values over the common engineered features in Public notebooks. \n\nI have recently started using LightGBM. In case I misinterpreted the plots, please correct me. \n\nAcknowledgement: \n- [[0.780]Unoptimized LGBM + Interesting Features](https://www.kaggle.com/zyy2016/0-780-unoptimized-lgbm-interesting-features)",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1130576": "**tl;dr**: By default LightGBM uses **feature split** for plotting feature importance. Check **feature gain** as well if you are using importance for selecting features. Feature split is bad if there is a combination of categorical and numerical features.\n\n---\n\nI was going through [this nice kernel](https://www.kaggle.com/zyy2016/0-780-unoptimized-lgbm-interesting-features). It introduced the usage of [Microsoft True Skill Ranking System](https://www.microsoft.com/en-us/research/project/trueskill-ranking-system/). One thing I wanted to check was if True Skill features are adding value using feature importance plot. By default LightGBM uses feature split for calculating feature importance. \n\n## Feature Importance by Split\n![Feature Importance by Split](https://i.imgur.com/pu2Y3KA.png)\n\n- Question ID is in the first place. This is due to the huge number of splits possible (17000+ unique values). Similar thing is also possible  for tag_1, tags_encoded. \n- True Skill features are not in the bottom of the list. \n- Some features can be very useful without too many splits: ex: std_deviation question accuracy, most_liked_guess_correct\n- Some features may be leading to over-fitting due to too many splits: user_id\n\n## Feature Importance by Gain\n![Feature Importance by Gain](https://i.imgur.com/RVpcOYL.png)\n\n- The first 2 features are from True Skill. It seems they are adding value. \n- Remove the features that are not shown / very low gain in the gain plot.\n\nIn summary: \n- Try to use both split and gain feature importance plots. Remove the features that are not adding any value in both split and gain plots. They are only leading to overfitting. \n- True Skill features are adding values over the common engineered features in Public notebooks. \n\nI have recently started using LightGBM. In case I misinterpreted the plots, please correct me. \n\nAcknowledgement: \n- [[0.780]Unoptimized LGBM + Interesting Features](https://www.kaggle.com/zyy2016/0-780-unoptimized-lgbm-interesting-features)"
  }
}