{
  "id": 59771,
  "title": "How to identify if model catches signals or noises from features?",
  "url": "/competitions/avito-demand-prediction/discussion/59771",
  "author_name": "",
  "post_date": "2018-06-27T01:33:35.867992200Z",
  "votes": 9,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi folks,</p>\n\n<p>I made a model using LGBM with text features (tfidf and count) and I found it has ~10k features which have feature importance &gt; 0 but some of them only have very small importance score compared to others. </p>\n\n<p>So I started wondering is there a way to tell whether my model actually catches signals from those features with tiny importance scores? or it's just trying to fit the noise? </p>\n\n<p>I tried a simple way - I sampled some random numbers from uniform and normal distribution using numpy.random.uniform and numpy.random.normal and put them to my model along with other features. I expected one of them (or some of them) might be picked up by my model with some small feature importance score. Then I can throw away those features ranked below my random features since they contributed even less than a random variable. </p>\n\n<p>However, what I got from my experiment surprised me. My random variables got picked at top 10 from feature importance list (as the plot shows). What does this mean? Does this really mean those feature ranked below are not useful? Could anyone help me understand this?</p>\n\n<p>And I also want to ask how people usually do to identify signals vs. noises from large amount of features?</p>\n\n<p>Thank you.</p>",
  "messages": [
    {
      "id": "348590",
      "postDate": "06/27/2018 01:33:35",
      "content": "<p>Hi folks,</p>\n\n<p>I made a model using LGBM with text features (tfidf and count) and I found it has ~10k features which have feature importance &gt; 0 but some of them only have very small importance score compared to others. </p>\n\n<p>So I started wondering is there a way to tell whether my model actually catches signals from those features with tiny importance scores? or it's just trying to fit the noise? </p>\n\n<p>I tried a simple way - I sampled some random numbers from uniform and normal distribution using numpy.random.uniform and numpy.random.normal and put them to my model along with other features. I expected one of them (or some of them) might be picked up by my model with some small feature importance score. Then I can throw away those features ranked below my random features since they contributed even less than a random variable. </p>\n\n<p>However, what I got from my experiment surprised me. My random variables got picked at top 10 from feature importance list (as the plot shows). What does this mean? Does this really mean those feature ranked below are not useful? Could anyone help me understand this?</p>\n\n<p>And I also want to ask how people usually do to identify signals vs. noises from large amount of features?</p>\n\n<p>Thank you.</p>",
      "rawMarkdown": "Hi folks,\n\nI made a model using LGBM with text features (tfidf and count) and I found it has ~10k features which have feature importance &gt; 0 but some of them only have very small importance score compared to others. \n\nSo I started wondering is there a way to tell whether my model actually catches signals from those features with tiny importance scores? or it's just trying to fit the noise? \n\nI tried a simple way - I sampled some random numbers from uniform and normal distribution using numpy.random.uniform and numpy.random.normal and put them to my model along with other features. I expected one of them (or some of them) might be picked up by my model with some small feature importance score. Then I can throw away those features ranked below my random features since they contributed even less than a random variable. \n\nHowever, what I got from my experiment surprised me. My random variables got picked at top 10 from feature importance list (as the plot shows). What does this mean? Does this really mean those feature ranked below are not useful? Could anyone help me understand this?\n\nAnd I also want to ask how people usually do to identify signals vs. noises from large amount of features?\n\nThank you.",
      "votes": null
    },
    {
      "id": "348672",
      "postDate": "06/27/2018 05:24:01",
      "content": "<p>That is a very interesting approach. I dont have any additional insight but it was very smart of you to try injecting random noise and seeing where the model would fit that in terms of feature importance. Have you tried only matching features above that random noise threshold?</p>",
      "rawMarkdown": "That is a very interesting approach. I dont have any additional insight but it was very smart of you to try injecting random noise and seeing where the model would fit that in terms of feature importance. Have you tried only matching features above that random noise threshold?",
      "votes": null
    },
    {
      "id": "348713",
      "postDate": "06/27/2018 06:41:17",
      "content": "<p>The feature importance functions of LightGBM, XGBoost, Random Forests and most other common implementations are unfortunately not very thorough. If you train a LightGBM and a Random Forest model on one dataset to almost identical results, theycan have vastly different feature importances.</p>\n\n<p>You should take a look at shapley values which aim to help in understanding how and in what way features are important to the model. There is a good library here (<a href=\"https://github.com/slundberg/shap\">https://github.com/slundberg/shap</a>) which integrates with sklearn, xgb, lgb.\nFor better understandings, you should also listen to this podcast which goes into more details about the underlying ideas of shapley values (<a href=\"http://lineardigressions.com/episodes/2018/5/13/shap-shapley-values-in-machine-learning\">http://lineardigressions.com/episodes/2018/5/13/shap-shapley-values-in-machine-learning</a>).</p>",
      "rawMarkdown": "The feature importance functions of LightGBM, XGBoost, Random Forests and most other common implementations are unfortunately not very thorough. If you train a LightGBM and a Random Forest model on one dataset to almost identical results, theycan have vastly different feature importances.\n\nYou should take a look at shapley values which aim to help in understanding how and in what way features are important to the model. There is a good library here (https://github.com/slundberg/shap) which integrates with sklearn, xgb, lgb.\nFor better understandings, you should also listen to this podcast which goes into more details about the underlying ideas of shapley values (http://lineardigressions.com/episodes/2018/5/13/shap-shapley-values-in-machine-learning).",
      "votes": null
    },
    {
      "id": "348723",
      "postDate": "06/27/2018 06:55:32",
      "content": "<p>Shap is awesome. I think it will play a huge role in the Santander competition.</p>",
      "rawMarkdown": "Shap is awesome. I think it will play a huge role in the Santander competition.",
      "votes": null
    },
    {
      "id": "348743",
      "postDate": "06/27/2018 07:43:36",
      "content": "<p>I agree that \"feature importance\" of any boosting model is not good metric for feature selection. The one of the most useful approach (from my point of view) is to use any of the \"shuffling\" method, like Boruta, random target shuffling and random feature shuffling.</p>\n\n<p>You can find a lot information in the net by googling \"feature selection random shuffling\", below are the some starting points from previous competitions:</p>\n\n<p><a href=\"https://www.kaggle.com/tilii7/boruta-feature-elimination\">https://www.kaggle.com/tilii7/boruta-feature-elimination</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/41595#233852\">https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/41595#233852</a></p>\n\n<p><a href=\"https://www.kaggle.com/tilii7/features-we-don-t-need-no-stinking-features\">https://www.kaggle.com/tilii7/features-we-don-t-need-no-stinking-features</a></p>",
      "rawMarkdown": "I agree that \"feature importance\" of any boosting model is not good metric for feature selection. The one of the most useful approach (from my point of view) is to use any of the \"shuffling\" method, like Boruta, random target shuffling and random feature shuffling.\n\nYou can find a lot information in the net by googling \"feature selection random shuffling\", below are the some starting points from previous competitions:\n\nhttps://www.kaggle.com/tilii7/boruta-feature-elimination\n\nhttps://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/41595#233852\n\nhttps://www.kaggle.com/tilii7/features-we-don-t-need-no-stinking-features",
      "votes": null
    },
    {
      "id": "348747",
      "postDate": "06/27/2018 07:48:36",
      "content": "<p>Thank you for your awesome links,</p>",
      "rawMarkdown": "Thank you for your awesome links,",
      "votes": null
    },
    {
      "id": "349175",
      "postDate": "06/27/2018 22:23:27",
      "content": "<p>Hi ryches, I tried training with those factors above random variables and the validation error increased. So I didn't submit that model result to LB. </p>",
      "rawMarkdown": "Hi ryches, I tried training with those factors above random variables and the validation error increased. So I didn't submit that model result to LB.",
      "votes": null
    },
    {
      "id": "349182",
      "postDate": "06/27/2018 22:32:27",
      "content": "<p>Thanks Frank. This is helpful. I will try this maybe in the next competition. </p>",
      "rawMarkdown": "Thanks Frank. This is helpful. I will try this maybe in the next competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 348672,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "06/27/2018 05:24:01",
      "content": "<p>That is a very interesting approach. I dont have any additional insight but it was very smart of you to try injecting random noise and seeing where the model would fit that in terms of feature importance. Have you tried only matching features above that random noise threshold?</p>",
      "votes": null,
      "replies": [
        {
          "id": 349175,
          "author_name": "magicmango",
          "author_url": "",
          "post_date": "06/27/2018 22:23:27",
          "content": "<p>Hi ryches, I tried training with those factors above random variables and the validation error increased. So I didn't submit that model result to LB. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 348713,
      "author_name": "frankherfert",
      "author_url": "",
      "post_date": "06/27/2018 06:41:17",
      "content": "<p>The feature importance functions of LightGBM, XGBoost, Random Forests and most other common implementations are unfortunately not very thorough. If you train a LightGBM and a Random Forest model on one dataset to almost identical results, theycan have vastly different feature importances.</p>\n\n<p>You should take a look at shapley values which aim to help in understanding how and in what way features are important to the model. There is a good library here (<a href=\"https://github.com/slundberg/shap\">https://github.com/slundberg/shap</a>) which integrates with sklearn, xgb, lgb.\nFor better understandings, you should also listen to this podcast which goes into more details about the underlying ideas of shapley values (<a href=\"http://lineardigressions.com/episodes/2018/5/13/shap-shapley-values-in-machine-learning\">http://lineardigressions.com/episodes/2018/5/13/shap-shapley-values-in-machine-learning</a>).</p>",
      "votes": null,
      "replies": [
        {
          "id": 348723,
          "author_name": "yk1598",
          "author_url": "",
          "post_date": "06/27/2018 06:55:32",
          "content": "<p>Shap is awesome. I think it will play a huge role in the Santander competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 349182,
          "author_name": "magicmango",
          "author_url": "",
          "post_date": "06/27/2018 22:32:27",
          "content": "<p>Thanks Frank. This is helpful. I will try this maybe in the next competition. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 348743,
      "author_name": "kruegger",
      "author_url": "",
      "post_date": "06/27/2018 07:43:36",
      "content": "<p>I agree that \"feature importance\" of any boosting model is not good metric for feature selection. The one of the most useful approach (from my point of view) is to use any of the \"shuffling\" method, like Boruta, random target shuffling and random feature shuffling.</p>\n\n<p>You can find a lot information in the net by googling \"feature selection random shuffling\", below are the some starting points from previous competitions:</p>\n\n<p><a href=\"https://www.kaggle.com/tilii7/boruta-feature-elimination\">https://www.kaggle.com/tilii7/boruta-feature-elimination</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/41595#233852\">https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/41595#233852</a></p>\n\n<p><a href=\"https://www.kaggle.com/tilii7/features-we-don-t-need-no-stinking-features\">https://www.kaggle.com/tilii7/features-we-don-t-need-no-stinking-features</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 348747,
          "author_name": "liuhdsgoal",
          "author_url": "",
          "post_date": "06/27/2018 07:48:36",
          "content": "<p>Thank you for your awesome links,</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "348590": "Hi folks,\n\nI made a model using LGBM with text features (tfidf and count) and I found it has ~10k features which have feature importance &gt; 0 but some of them only have very small importance score compared to others. \n\nSo I started wondering is there a way to tell whether my model actually catches signals from those features with tiny importance scores? or it's just trying to fit the noise? \n\nI tried a simple way - I sampled some random numbers from uniform and normal distribution using numpy.random.uniform and numpy.random.normal and put them to my model along with other features. I expected one of them (or some of them) might be picked up by my model with some small feature importance score. Then I can throw away those features ranked below my random features since they contributed even less than a random variable. \n\nHowever, what I got from my experiment surprised me. My random variables got picked at top 10 from feature importance list (as the plot shows). What does this mean? Does this really mean those feature ranked below are not useful? Could anyone help me understand this?\n\nAnd I also want to ask how people usually do to identify signals vs. noises from large amount of features?\n\nThank you.",
    "348672": "That is a very interesting approach. I dont have any additional insight but it was very smart of you to try injecting random noise and seeing where the model would fit that in terms of feature importance. Have you tried only matching features above that random noise threshold?",
    "348713": "The feature importance functions of LightGBM, XGBoost, Random Forests and most other common implementations are unfortunately not very thorough. If you train a LightGBM and a Random Forest model on one dataset to almost identical results, theycan have vastly different feature importances.\n\nYou should take a look at shapley values which aim to help in understanding how and in what way features are important to the model. There is a good library here (https://github.com/slundberg/shap) which integrates with sklearn, xgb, lgb.\nFor better understandings, you should also listen to this podcast which goes into more details about the underlying ideas of shapley values (http://lineardigressions.com/episodes/2018/5/13/shap-shapley-values-in-machine-learning).",
    "348723": "Shap is awesome. I think it will play a huge role in the Santander competition.",
    "348743": "I agree that \"feature importance\" of any boosting model is not good metric for feature selection. The one of the most useful approach (from my point of view) is to use any of the \"shuffling\" method, like Boruta, random target shuffling and random feature shuffling.\n\nYou can find a lot information in the net by googling \"feature selection random shuffling\", below are the some starting points from previous competitions:\n\nhttps://www.kaggle.com/tilii7/boruta-feature-elimination\n\nhttps://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/41595#233852\n\nhttps://www.kaggle.com/tilii7/features-we-don-t-need-no-stinking-features",
    "348747": "Thank you for your awesome links,",
    "349175": "Hi ryches, I tried training with those factors above random variables and the validation error increased. So I didn't submit that model result to LB.",
    "349182": "Thanks Frank. This is helpful. I will try this maybe in the next competition."
  },
  "source": "meta"
}