{
  "id": 58918,
  "title": "How to find a important feature ?",
  "url": "/competitions/avito-demand-prediction/discussion/58918",
  "author_name": "Victor An",
  "post_date": "2018-06-15T10:08:36.710000",
  "votes": 0,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi guys:</p>\n\n<p>I'm a newer working on data scientist at kaggle, I know that there many common feature selection methods, but it really work well?</p>\n\n<p>So guys, would you like to share your method to check a feature is an important feature in your model.</p>\n\n<p>As I have shared a kernel for my approach to check a valuable feature in linear model and tree base model\n<a href=\"https://www.kaggle.com/classtag/review-your-feature-before-modeling\">https://www.kaggle.com/classtag/review-your-feature-before-modeling</a></p>\n\n<p>Looking for more efficient methods from you.</p>\n\n<p>Thanks.</p>",
  "messages": [
    {
      "id": 343528,
      "postDate": "2018-06-15T13:43:38.853Z",
      "content": "<p>I usually downsample data, run cv with different seeds, average the feature importances and review them. (Check the cv results as well as their importances). Xgb\\lgb and some of sklearn models has 'feature_importances_' attribute. You could make it auto by setting some threshold and filter those with feature importance greater than the threshold. But this way of feature selection is model-dependent. PCA\\SVD\\NMF are kinds of dimension reduction techniques, advantage: model independent, don't need to run models. disadvantage: different from feature selection, bad features might still be encoded into the condensed data, not interpretable after transformed. Hope it helps and please correct me if I'm wrong. Thanks.</p>",
      "rawMarkdown": "I usually downsample data, run cv with different seeds, average the feature importances and review them. (Check the cv results as well as their importances). Xgb\\lgb and some of sklearn models has 'feature_importances_' attribute. You could make it auto by setting some threshold and filter those with feature importance greater than the threshold. But this way of feature selection is model-dependent. PCA\\SVD\\NMF are kinds of dimension reduction techniques, advantage: model independent, don't need to run models. disadvantage: different from feature selection, bad features might still be encoded into the condensed data, not interpretable after transformed. Hope it helps and please correct me if I'm wrong. Thanks.",
      "votes": 3,
      "replies": [
        {
          "id": 343533,
          "postDate": "2018-06-15T13:50:25.057Z",
          "content": "<p>Yes, it is a better method, my consider is that we can't know the feature importances before the model trained, is there any quickly method?</p>",
          "rawMarkdown": "Yes, it is a better method, my consider is that we can't know the feature importances before the model trained, is there any quickly method?"
        },
        {
          "id": 343582,
          "postDate": "2018-06-15T16:00:26.743Z",
          "content": "<p>I think you should be looking for statistical feature selection methods like doing chi2 testing.\nHowever, IMHO, it might filter features that might be actually helpful to the models, which you wouldn't know if you haven't fit the models :))) Normally models like xgb\\lgb\\ann could dig the features really deep (e.g. this competition), where some features just turn out to be useful in deeper trees but rather hard to be identified with pure statistical methods. </p>\n\n<p>P.S. There are three kinds of feature selection to my knowledge: statistical testing based, modeling based, embed based (learning what features to be useful as part of learning).</p>",
          "rawMarkdown": "I think you should be looking for statistical feature selection methods like doing chi2 testing.\nHowever, IMHO, it might filter features that might be actually helpful to the models, which you wouldn't know if you haven't fit the models :))) Normally models like xgb\\lgb\\ann could dig the features really deep (e.g. this competition), where some features just turn out to be useful in deeper trees but rather hard to be identified with pure statistical methods. \n\nP.S. There are three kinds of feature selection to my knowledge: statistical testing based, modeling based, embed based (learning what features to be useful as part of learning).",
          "votes": 1
        }
      ]
    },
    {
      "id": 343448,
      "postDate": "2018-06-15T10:39:54.060Z",
      "content": "<p>PCA is one of the most popular techniques for dimensionality reduction, give it a try.</p>\n\n<pre><code>&gt;&gt;&gt; from sklearn.decomposition import PCA\n&gt;&gt;&gt; pca = PCA(n_components=10, svd_solver='full')\n&gt;&gt;&gt; pca.fit(df)\nPCA(copy=True, n_components=10, whiten=False)\n\n&gt;&gt;&gt; T = pca.transform(df)\n</code></pre>\n\n<p>Try varying different parameters, like n_components (top n features you would want to select) and check your model performance.</p>\n\n<p>Hope this helps.</p>\n\n<p>Kr\nPrashant</p>",
      "rawMarkdown": "PCA is one of the most popular techniques for dimensionality reduction, give it a try.\n\n    &gt;&gt;&gt; from sklearn.decomposition import PCA\n    &gt;&gt;&gt; pca = PCA(n_components=10, svd_solver='full')\n    &gt;&gt;&gt; pca.fit(df)\n    PCA(copy=True, n_components=10, whiten=False)\n    \n    &gt;&gt;&gt; T = pca.transform(df)\n\nTry varying different parameters, like n_components (top n features you would want to select) and check your model performance.\n\nHope this helps.\n\nKr\nPrashant",
      "votes": 1,
      "replies": [
        {
          "id": 343537,
          "postDate": "2018-06-15T14:00:10.483Z",
          "content": "<p>Thanks @Prashant I know PCA is a very efficient way to reduce feature dimensions.</p>\n\n<p>I hope learn more other efficient ways to verify which feature is benefit for our model.\nFor example, \nif I found one feature have higher correlation coefficient with our target value, it may has a good performance for linear model;\nYou'd better better to see the feature's gini or information gain of feature and target when you using a tree-base model.</p>\n\n<p>The above is just my personal point of view. I do not know whether it is correct or not.</p>",
          "rawMarkdown": "Thanks @Prashant I know PCA is a very efficient way to reduce feature dimensions.\n\nI hope learn more other efficient ways to verify which feature is benefit for our model.\nFor example, \nif I found one feature have higher correlation coefficient with our target value, it may has a good performance for linear model;\nYou'd better better to see the feature's gini or information gain of feature and target when you using a tree-base model.\n\nThe above is just my personal point of view. I do not know whether it is correct or not.\n"
        },
        {
          "id": 345091,
          "postDate": "2018-06-19T07:26:04.890Z",
          "content": "<p>Thanks for this info Prashant!</p>",
          "rawMarkdown": "Thanks for this info Prashant!"
        },
        {
          "id": 345293,
          "postDate": "2018-06-19T15:48:49.720Z",
          "content": "<p>No problem. Happy to help.</p>",
          "rawMarkdown": "No problem. Happy to help."
        }
      ]
    },
    {
      "id": 343433,
      "postDate": "2018-06-15T10:08:36.710Z",
      "content": "<p>Hi guys:</p>\n\n<p>I'm a newer working on data scientist at kaggle, I know that there many common feature selection methods, but it really work well?</p>\n\n<p>So guys, would you like to share your method to check a feature is an important feature in your model.</p>\n\n<p>As I have shared a kernel for my approach to check a valuable feature in linear model and tree base model\n<a href=\"https://www.kaggle.com/classtag/review-your-feature-before-modeling\">https://www.kaggle.com/classtag/review-your-feature-before-modeling</a></p>\n\n<p>Looking for more efficient methods from you.</p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "Hi guys:\n\nI'm a newer working on data scientist at kaggle, I know that there many common feature selection methods, but it really work well?\n\nSo guys, would you like to share your method to check a feature is an important feature in your model.\n\nAs I have shared a kernel for my approach to check a valuable feature in linear model and tree base model\nhttps://www.kaggle.com/classtag/review-your-feature-before-modeling\n\nLooking for more efficient methods from you.\n\nThanks.\n"
    },
    {
      "id": 345337,
      "postDate": "2018-06-19T17:39:40.777Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 343528,
      "author_name": "khyeh",
      "author_url": "",
      "post_date": "2018-06-15T13:43:38.853000",
      "content": "<p>I usually downsample data, run cv with different seeds, average the feature importances and review them. (Check the cv results as well as their importances). Xgb\\lgb and some of sklearn models has 'feature_importances_' attribute. You could make it auto by setting some threshold and filter those with feature importance greater than the threshold. But this way of feature selection is model-dependent. PCA\\SVD\\NMF are kinds of dimension reduction techniques, advantage: model independent, don't need to run models. disadvantage: different from feature selection, bad features might still be encoded into the condensed data, not interpretable after transformed. Hope it helps and please correct me if I'm wrong. Thanks.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 343533,
          "author_name": "Victor An",
          "author_url": "",
          "post_date": "2018-06-15T13:50:25.057000",
          "content": "<p>Yes, it is a better method, my consider is that we can't know the feature importances before the model trained, is there any quickly method?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 343582,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2018-06-15T16:00:26.743000",
          "content": "<p>I think you should be looking for statistical feature selection methods like doing chi2 testing.\nHowever, IMHO, it might filter features that might be actually helpful to the models, which you wouldn't know if you haven't fit the models :))) Normally models like xgb\\lgb\\ann could dig the features really deep (e.g. this competition), where some features just turn out to be useful in deeper trees but rather hard to be identified with pure statistical methods. </p>\n\n<p>P.S. There are three kinds of feature selection to my knowledge: statistical testing based, modeling based, embed based (learning what features to be useful as part of learning).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 343448,
      "author_name": "Prashant Patel",
      "author_url": "",
      "post_date": "2018-06-15T10:39:54.060000",
      "content": "<p>PCA is one of the most popular techniques for dimensionality reduction, give it a try.</p>\n\n<pre><code>&gt;&gt;&gt; from sklearn.decomposition import PCA\n&gt;&gt;&gt; pca = PCA(n_components=10, svd_solver='full')\n&gt;&gt;&gt; pca.fit(df)\nPCA(copy=True, n_components=10, whiten=False)\n\n&gt;&gt;&gt; T = pca.transform(df)\n</code></pre>\n\n<p>Try varying different parameters, like n_components (top n features you would want to select) and check your model performance.</p>\n\n<p>Hope this helps.</p>\n\n<p>Kr\nPrashant</p>",
      "votes": 1,
      "replies": [
        {
          "id": 343537,
          "author_name": "Victor An",
          "author_url": "",
          "post_date": "2018-06-15T14:00:10.483000",
          "content": "<p>Thanks @Prashant I know PCA is a very efficient way to reduce feature dimensions.</p>\n\n<p>I hope learn more other efficient ways to verify which feature is benefit for our model.\nFor example, \nif I found one feature have higher correlation coefficient with our target value, it may has a good performance for linear model;\nYou'd better better to see the feature's gini or information gain of feature and target when you using a tree-base model.</p>\n\n<p>The above is just my personal point of view. I do not know whether it is correct or not.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 345091,
          "author_name": "Mrhappysmile",
          "author_url": "",
          "post_date": "2018-06-19T07:26:04.890000",
          "content": "<p>Thanks for this info Prashant!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 345293,
          "author_name": "Prashant Patel",
          "author_url": "",
          "post_date": "2018-06-19T15:48:49.720000",
          "content": "<p>No problem. Happy to help.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 345337,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-19T17:39:40.777000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "343528": "I usually downsample data, run cv with different seeds, average the feature importances and review them. (Check the cv results as well as their importances). Xgb\\lgb and some of sklearn models has 'feature_importances_' attribute. You could make it auto by setting some threshold and filter those with feature importance greater than the threshold. But this way of feature selection is model-dependent. PCA\\SVD\\NMF are kinds of dimension reduction techniques, advantage: model independent, don't need to run models. disadvantage: different from feature selection, bad features might still be encoded into the condensed data, not interpretable after transformed. Hope it helps and please correct me if I'm wrong. Thanks.",
    "343448": "PCA is one of the most popular techniques for dimensionality reduction, give it a try.\n\n    &gt;&gt;&gt; from sklearn.decomposition import PCA\n    &gt;&gt;&gt; pca = PCA(n_components=10, svd_solver='full')\n    &gt;&gt;&gt; pca.fit(df)\n    PCA(copy=True, n_components=10, whiten=False)\n    \n    &gt;&gt;&gt; T = pca.transform(df)\n\nTry varying different parameters, like n_components (top n features you would want to select) and check your model performance.\n\nHope this helps.\n\nKr\nPrashant",
    "343433": "Hi guys:\n\nI'm a newer working on data scientist at kaggle, I know that there many common feature selection methods, but it really work well?\n\nSo guys, would you like to share your method to check a feature is an important feature in your model.\n\nAs I have shared a kernel for my approach to check a valuable feature in linear model and tree base model\nhttps://www.kaggle.com/classtag/review-your-feature-before-modeling\n\nLooking for more efficient methods from you.\n\nThanks.\n",
    "345337": ""
  }
}