{
  "id": 54526,
  "title": "Feature Importance Technique",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54526",
  "author_name": "A.Barqawi",
  "post_date": "2018-04-14T11:46:21.451000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Many Kagglers asked about the features and which one to use.</p>\n\n<p>Lightgbm help you in the task by loading  from 7 to 12 features (<em>which is the average no. of features for good score based on discussions</em>) then at the end of training run following code:</p>\n\n<pre><code>import matplotlib.pyplot as plt\nprint('Plot feature importances...')\n #gbm is the model after training\nax = lgb.plot_importance(gbm, max_num_features=10)\nplt.show()\n</code></pre>\n\n<p>Based on the chart you can remove the least effective features and replace it with new ones in the next trial to increase score (to avoid having noise features or load data more than your RAM capabilities).</p>\n\n<p>Following are good sources in the competition for features to use, based on my trials nextclick have good impact on the training:</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">https://www.kaggle.com/nanomathias/feature-engineering-importance-testing</a></li>\n<li><a href=\"https://www.kaggle.com/tetyanayatsenko/prepare-data-form-features-find-their-importance\">https://www.kaggle.com/tetyanayatsenko/prepare-data-form-features-find-their-importance</a></li>\n<li><a href=\"https://www.kaggle.com/pranav84/talkingdata-with-breaking-bad-feature-engg\">https://www.kaggle.com/pranav84/talkingdata-with-breaking-bad-feature-engg</a></li>\n</ul>",
  "messages": [
    {
      "id": 314017,
      "postDate": "2018-04-14T11:46:21.450Z",
      "content": "<p>Many Kagglers asked about the features and which one to use.</p>\n\n<p>Lightgbm help you in the task by loading  from 7 to 12 features (<em>which is the average no. of features for good score based on discussions</em>) then at the end of training run following code:</p>\n\n<pre><code>import matplotlib.pyplot as plt\nprint('Plot feature importances...')\n #gbm is the model after training\nax = lgb.plot_importance(gbm, max_num_features=10)\nplt.show()\n</code></pre>\n\n<p>Based on the chart you can remove the least effective features and replace it with new ones in the next trial to increase score (to avoid having noise features or load data more than your RAM capabilities).</p>\n\n<p>Following are good sources in the competition for features to use, based on my trials nextclick have good impact on the training:</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">https://www.kaggle.com/nanomathias/feature-engineering-importance-testing</a></li>\n<li><a href=\"https://www.kaggle.com/tetyanayatsenko/prepare-data-form-features-find-their-importance\">https://www.kaggle.com/tetyanayatsenko/prepare-data-form-features-find-their-importance</a></li>\n<li><a href=\"https://www.kaggle.com/pranav84/talkingdata-with-breaking-bad-feature-engg\">https://www.kaggle.com/pranav84/talkingdata-with-breaking-bad-feature-engg</a></li>\n</ul>",
      "rawMarkdown": "Many Kagglers asked about the features and which one to use.\n\nLightgbm help you in the task by loading  from 7 to 12 features (*which is the average no. of features for good score based on discussions*) then at the end of training run following code:\n\n    import matplotlib.pyplot as plt\n    print('Plot feature importances...')\n     #gbm is the model after training\n    ax = lgb.plot_importance(gbm, max_num_features=10)\n    plt.show()\n\nBased on the chart you can remove the least effective features and replace it with new ones in the next trial to increase score (to avoid having noise features or load data more than your RAM capabilities).\n\nFollowing are good sources in the competition for features to use, based on my trials nextclick have good impact on the training:\n\n - https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\n - https://www.kaggle.com/tetyanayatsenko/prepare-data-form-features-find-their-importance\n - https://www.kaggle.com/pranav84/talkingdata-with-breaking-bad-feature-engg\n\n",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "314017": "Many Kagglers asked about the features and which one to use.\n\nLightgbm help you in the task by loading  from 7 to 12 features (*which is the average no. of features for good score based on discussions*) then at the end of training run following code:\n\n    import matplotlib.pyplot as plt\n    print('Plot feature importances...')\n     #gbm is the model after training\n    ax = lgb.plot_importance(gbm, max_num_features=10)\n    plt.show()\n\nBased on the chart you can remove the least effective features and replace it with new ones in the next trial to increase score (to avoid having noise features or load data more than your RAM capabilities).\n\nFollowing are good sources in the competition for features to use, based on my trials nextclick have good impact on the training:\n\n - https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\n - https://www.kaggle.com/tetyanayatsenko/prepare-data-form-features-find-their-importance\n - https://www.kaggle.com/pranav84/talkingdata-with-breaking-bad-feature-engg\n\n"
  }
}