{
  "id": 331131,
  "title": "Which is the right feature importance?",
  "url": "/competitions/amex-default-prediction/discussion/331131",
  "author_name": "AmbrosM",
  "post_date": "2022-06-15T21:12:23.507000",
  "votes": 210,
  "comment_count": 48,
  "views": 0,
  "content": "<p>Why do we evaluate feature importances? Because we want to select the good features and drop the bad ones. But which of the following three importances is the right one for this purpose?</p>\n<ol>\n<li>Split feature importance?</li>\n<li>Gain feature importance?</li>\n<li>Permutation feature importance?</li>\n</ol>\n<p>In this post I'll compare the three methods for feature selection and show that the last one is the right one.</p>\n<p>The former two methods are often seen in public gradient-boosting notebooks because people like to decorate their notebooks with a nice diagram, all the more as a single line of code suffices to generate the diagram. But the diagram alone doesn't give a better model! We need to interpret the importance values.</p>\n<p><strong>Split feature importance</strong> counts how often a feature is used in the model. The model I'm using for this experiment has 1200 trees, and split feature importances are between 0 and 1082. If a feature has importance = 0, this means that the feature is not used in the model at all. If we omit the feature, model quality doesn't change, but the training takes less time. If split feature importance is high, we know that the feature is used often in the model and that it improves the model's training score. </p>\n<p><strong>Gain feature importance</strong> is similar to split feature importance. It measures the gain which was achieved in all splits based on the feature, although the definition of this gain is hard to find in the documentation. The features of my model have importances between 0 and 3000000. Again, if a feature has importance = 0, this means that the feature is not used in the model at all.</p>\n<p>The following diagram shows that split and gain feature importance correlate quite well. The diagram is based on importance ranks, which are between 0 and 573 (my model has 574 features) - higher is better. The single dot at coordinates (30,30) represents 60 features which are not used in the model and have both importances equal to zero.</p>\n<p><img src=\"https://i.imgur.com/yS4eQXF.png\" alt=\"split-gain\"></p>\n<p>Split and gain importance are cheap byproducts of the training process, but they have a massive shortcoming: They are based on the training data rather than the validation data and consequently cannot measure generalization. Would you ever evaluate a model by its training score? Of course not! Every child learns that models must be scored on a validation dataset so that we can evaluate its generalization capability. The same principle holds for feature importance.</p>\n<p><strong>Permutation feature importance</strong> is a model inspection technique that can be used for any fitted estimator when the data is tabular. The permutation feature importance is defined as the decrease in a model's validation score when a single feature value is randomly shuffled. In contrast to split and gain importance, permutation feature importance can be positive or negative:</p>\n<ul>\n<li>Positive permutation feature importance means that the model profits from the feature and deteriorates if the feature is shuffled or removed.</li>\n<li>Zero permutation feature importance means that the model doesn't change if we shuffle or remove the feature.</li>\n<li>Negative permutation feature importance means that the model improves if the feature is removed.</li>\n</ul>\n<p>If a feature is complete noise, a decision tree can still use it to improve the training score, and the feature will get positive split/gain importances. Only permutation feature importance detects that such a feature doesn't generalize to the validation dataset.</p>\n<p>The following diagrams show how permutation feature importance compares to split and gain importance:</p>\n<ul>\n<li>The green dots represent features with positive permutation feature importance. These are the good features.</li>\n<li>The black dot at the left represents the 60 features which are not used in the model.</li>\n<li>The red dots represent the features with negative permutation feature importance. These are the bad features although they may have high split or gain importance. D_43_min and S_26_min (represented by the two rightmost red dots) belong to the top 40 features according to split importance, but the model gets better if we drop them.</li>\n</ul>\n<p><img src=\"https://i.imgur.com/CBWXsbu.png\" alt=\"split-gain-pfi\"></p>\n<p><strong>Conclusion:</strong> Although permutation feature importance is slower than the gradient boosters' built-in importance measures, it is much more useful because it measures generalization and distinguishes positive from negative importances. Features with negative importance should be dropped from the model.</p>\n<p>See <a href=\"https://scikit-learn.org/stable/auto_examples/inspection/plot_permutation_importance.html\" target=\"_blank\">here</a> for a demo of the same effect on the Titanic dataset.</p>",
  "messages": [
    {
      "id": 1821813,
      "postDate": "2022-06-15T21:12:23.507Z",
      "content": "<p>Why do we evaluate feature importances? Because we want to select the good features and drop the bad ones. But which of the following three importances is the right one for this purpose?</p>\n<ol>\n<li>Split feature importance?</li>\n<li>Gain feature importance?</li>\n<li>Permutation feature importance?</li>\n</ol>\n<p>In this post I'll compare the three methods for feature selection and show that the last one is the right one.</p>\n<p>The former two methods are often seen in public gradient-boosting notebooks because people like to decorate their notebooks with a nice diagram, all the more as a single line of code suffices to generate the diagram. But the diagram alone doesn't give a better model! We need to interpret the importance values.</p>\n<p><strong>Split feature importance</strong> counts how often a feature is used in the model. The model I'm using for this experiment has 1200 trees, and split feature importances are between 0 and 1082. If a feature has importance = 0, this means that the feature is not used in the model at all. If we omit the feature, model quality doesn't change, but the training takes less time. If split feature importance is high, we know that the feature is used often in the model and that it improves the model's training score. </p>\n<p><strong>Gain feature importance</strong> is similar to split feature importance. It measures the gain which was achieved in all splits based on the feature, although the definition of this gain is hard to find in the documentation. The features of my model have importances between 0 and 3000000. Again, if a feature has importance = 0, this means that the feature is not used in the model at all.</p>\n<p>The following diagram shows that split and gain feature importance correlate quite well. The diagram is based on importance ranks, which are between 0 and 573 (my model has 574 features) - higher is better. The single dot at coordinates (30,30) represents 60 features which are not used in the model and have both importances equal to zero.</p>\n<p><img src=\"https://i.imgur.com/yS4eQXF.png\" alt=\"split-gain\"></p>\n<p>Split and gain importance are cheap byproducts of the training process, but they have a massive shortcoming: They are based on the training data rather than the validation data and consequently cannot measure generalization. Would you ever evaluate a model by its training score? Of course not! Every child learns that models must be scored on a validation dataset so that we can evaluate its generalization capability. The same principle holds for feature importance.</p>\n<p><strong>Permutation feature importance</strong> is a model inspection technique that can be used for any fitted estimator when the data is tabular. The permutation feature importance is defined as the decrease in a model's validation score when a single feature value is randomly shuffled. In contrast to split and gain importance, permutation feature importance can be positive or negative:</p>\n<ul>\n<li>Positive permutation feature importance means that the model profits from the feature and deteriorates if the feature is shuffled or removed.</li>\n<li>Zero permutation feature importance means that the model doesn't change if we shuffle or remove the feature.</li>\n<li>Negative permutation feature importance means that the model improves if the feature is removed.</li>\n</ul>\n<p>If a feature is complete noise, a decision tree can still use it to improve the training score, and the feature will get positive split/gain importances. Only permutation feature importance detects that such a feature doesn't generalize to the validation dataset.</p>\n<p>The following diagrams show how permutation feature importance compares to split and gain importance:</p>\n<ul>\n<li>The green dots represent features with positive permutation feature importance. These are the good features.</li>\n<li>The black dot at the left represents the 60 features which are not used in the model.</li>\n<li>The red dots represent the features with negative permutation feature importance. These are the bad features although they may have high split or gain importance. D_43_min and S_26_min (represented by the two rightmost red dots) belong to the top 40 features according to split importance, but the model gets better if we drop them.</li>\n</ul>\n<p><img src=\"https://i.imgur.com/CBWXsbu.png\" alt=\"split-gain-pfi\"></p>\n<p><strong>Conclusion:</strong> Although permutation feature importance is slower than the gradient boosters' built-in importance measures, it is much more useful because it measures generalization and distinguishes positive from negative importances. Features with negative importance should be dropped from the model.</p>\n<p>See <a href=\"https://scikit-learn.org/stable/auto_examples/inspection/plot_permutation_importance.html\" target=\"_blank\">here</a> for a demo of the same effect on the Titanic dataset.</p>",
      "rawMarkdown": "Why do we evaluate feature importances? Because we want to select the good features and drop the bad ones. But which of the following three importances is the right one for this purpose?\n\n1. Split feature importance?\n2. Gain feature importance?\n3. Permutation feature importance?\n\nIn this post I'll compare the three methods for feature selection and show that the last one is the right one.\n\nThe former two methods are often seen in public gradient-boosting notebooks because people like to decorate their notebooks with a nice diagram, all the more as a single line of code suffices to generate the diagram. But the diagram alone doesn't give a better model! We need to interpret the importance values.\n\n**Split feature importance** counts how often a feature is used in the model. The model I'm using for this experiment has 1200 trees, and split feature importances are between 0 and 1082. If a feature has importance = 0, this means that the feature is not used in the model at all. If we omit the feature, model quality doesn't change, but the training takes less time. If split feature importance is high, we know that the feature is used often in the model and that it improves the model's training score. \n\n**Gain feature importance** is similar to split feature importance. It measures the gain which was achieved in all splits based on the feature, although the definition of this gain is hard to find in the documentation. The features of my model have importances between 0 and 3000000. Again, if a feature has importance = 0, this means that the feature is not used in the model at all.\n\nThe following diagram shows that split and gain feature importance correlate quite well. The diagram is based on importance ranks, which are between 0 and 573 (my model has 574 features) - higher is better. The single dot at coordinates (30,30) represents 60 features which are not used in the model and have both importances equal to zero.\n\n![split-gain](https://i.imgur.com/yS4eQXF.png)\n\nSplit and gain importance are cheap byproducts of the training process, but they have a massive shortcoming: They are based on the training data rather than the validation data and consequently cannot measure generalization. Would you ever evaluate a model by its training score? Of course not! Every child learns that models must be scored on a validation dataset so that we can evaluate its generalization capability. The same principle holds for feature importance.\n\n**Permutation feature importance** is a model inspection technique that can be used for any fitted estimator when the data is tabular. The permutation feature importance is defined as the decrease in a model's validation score when a single feature value is randomly shuffled. In contrast to split and gain importance, permutation feature importance can be positive or negative:\n- Positive permutation feature importance means that the model profits from the feature and deteriorates if the feature is shuffled or removed.\n- Zero permutation feature importance means that the model doesn't change if we shuffle or remove the feature.\n- Negative permutation feature importance means that the model improves if the feature is removed.\n\nIf a feature is complete noise, a decision tree can still use it to improve the training score, and the feature will get positive split/gain importances. Only permutation feature importance detects that such a feature doesn't generalize to the validation dataset.\n\nThe following diagrams show how permutation feature importance compares to split and gain importance:\n- The green dots represent features with positive permutation feature importance. These are the good features.\n- The black dot at the left represents the 60 features which are not used in the model.\n- The red dots represent the features with negative permutation feature importance. These are the bad features although they may have high split or gain importance. D_43_min and S_26_min (represented by the two rightmost red dots) belong to the top 40 features according to split importance, but the model gets better if we drop them.\n\n![split-gain-pfi](https://i.imgur.com/CBWXsbu.png)\n\n**Conclusion:** Although permutation feature importance is slower than the gradient boosters' built-in importance measures, it is much more useful because it measures generalization and distinguishes positive from negative importances. Features with negative importance should be dropped from the model.\n\nSee [here](https://scikit-learn.org/stable/auto_examples/inspection/plot_permutation_importance.html) for a demo of the same effect on the Titanic dataset.",
      "votes": 209
    },
    {
      "id": 1822040,
      "postDate": "2022-06-16T03:14:20.233Z",
      "content": "<p>I think you are very right, and your conclusion is very helpful. I have been using permutation feature importance for a while and I think it's a great tool.</p>\n<p>I would like to add one thing though: when you shuffle a feature you dont take in cosideration the interaction it has with other features that might still pose a problem for feature selection. <br>\nFor example, if you have two columns that are the same: You can permute them all day long your model will always use one of them and the score won't change. </p>",
      "rawMarkdown": "I think you are very right, and your conclusion is very helpful. I have been using permutation feature importance for a while and I think it's a great tool.\n\nI would like to add one thing though: when you shuffle a feature you dont take in cosideration the interaction it has with other features that might still pose a problem for feature selection. \nFor example, if you have two columns that are the same: You can permute them all day long your model will always use one of them and the score won't change. \n",
      "votes": 20,
      "replies": [
        {
          "id": 1825146,
          "postDate": "2022-06-19T02:25:08.750Z",
          "rawMarkdown": "",
          "votes": 3,
          "isDeleted": true
        },
        {
          "id": 1842765,
          "postDate": "2022-07-04T09:10:38.970Z",
          "content": "<p><a href=\"https://www.kaggle.com/yiwangnz\" target=\"_blank\">@yiwangnz</a> see it working. Besides being compute-intensive I see the task in figuring out for which threshold to group two features. This dataset has a high dimensionality after the regular feature engineering done by top scoring notebooks. It would not be easy to group two features from the same category without removing important correlations. </p>",
          "rawMarkdown": "@yiwangnz see it working. Besides being compute-intensive I see the task in figuring out for which threshold to group two features. This dataset has a high dimensionality after the regular feature engineering done by top scoring notebooks. It would not be easy to group two features from the same category without removing important correlations. "
        },
        {
          "id": 1863307,
          "postDate": "2022-07-20T09:36:03.200Z",
          "content": "<p>i think if score don't change means the feature doesn't work and we can drop one. When we evaluate the other one we can get the real feature importance, is that a problem?</p>",
          "rawMarkdown": "i think if score don't change means the feature doesn't work and we can drop one. When we evaluate the other one we can get the real feature importance, is that a problem?",
          "votes": 1
        }
      ]
    },
    {
      "id": 1822914,
      "postDate": "2022-06-16T20:17:58.190Z",
      "content": "<p>Contributing to the discussion, I think <strong>lofo</strong> (<a href=\"https://github.com/aerdem4/lofo-importance\" target=\"_blank\">https://github.com/aerdem4/lofo-importance</a>) is even a more robust approach for measuring feature importance, but it's computationally expensive as it needs to train a model for each feature in the dataset:</p>\n<blockquote>\n  <p>LOFO first evaluates the performance of the model with all the input features included, then iteratively removes one feature at a time, retrains the model, and evaluates its performance on a validation set.</p>\n</blockquote>\n<p>In the same package there also a <strong>fast-lofo</strong> (<a href=\"https://github.com/aerdem4/lofo-importance#flofo-importance\" target=\"_blank\">https://github.com/aerdem4/lofo-importance#flofo-importance</a>) method, that is basically a supercharged permutation importance, that deals with the correlated feature overestimation problem.</p>",
      "rawMarkdown": "Contributing to the discussion, I think **lofo** (https://github.com/aerdem4/lofo-importance) is even a more robust approach for measuring feature importance, but it's computationally expensive as it needs to train a model for each feature in the dataset:\n> LOFO first evaluates the performance of the model with all the input features included, then iteratively removes one feature at a time, retrains the model, and evaluates its performance on a validation set.\n\nIn the same package there also a **fast-lofo** (https://github.com/aerdem4/lofo-importance#flofo-importance) method, that is basically a supercharged permutation importance, that deals with the correlated feature overestimation problem.\n",
      "votes": 9,
      "replies": [
        {
          "id": 1844185,
          "postDate": "2022-07-05T11:59:46.507Z",
          "content": "<p>If you have an estimator which provides information about feature importance you can basically use recursive feature elimination from sklearn instead:<br>\n<a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.RFE.html\" target=\"_blank\">https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.RFE.html</a></p>",
          "rawMarkdown": "If you have an estimator which provides information about feature importance you can basically use recursive feature elimination from sklearn instead:\nhttps://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.RFE.html"
        }
      ]
    },
    {
      "id": 1885124,
      "postDate": "2022-08-04T22:59:34.070Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> for sharing! I implemented it and it helps a lot!</p>\n<p>I would just annotate the following for others working with tree based algorithms. In tree based algorithms the order of your features matters a lot for your model! So if you give [A, B, C] to xgboost you will get a different result compared to [B, A, C]. </p>\n<p>I noticed that my permutation importance changes a lot depending on the feature order! </p>\n<p>Therefore, if you are looking to find the \"general permutation importance\" of your features, independent of your model, fold or feature order, you might consider to perform several permutation importance valuations over different folds with different feature orders and average them. That helped me a lot! :)</p>",
      "rawMarkdown": "Thanks @ambrosm for sharing! I implemented it and it helps a lot!\n\nI would just annotate the following for others working with tree based algorithms. In tree based algorithms the order of your features matters a lot for your model! So if you give [A, B, C] to xgboost you will get a different result compared to [B, A, C]. \n\nI noticed that my permutation importance changes a lot depending on the feature order! \n\nTherefore, if you are looking to find the \"general permutation importance\" of your features, independent of your model, fold or feature order, you might consider to perform several permutation importance valuations over different folds with different feature orders and average them. That helped me a lot! :)",
      "votes": 5
    },
    {
      "id": 1880390,
      "postDate": "2022-08-01T17:08:30.933Z",
      "content": "<p>Nice and interesting post, here is my 2cts:</p>\n<p>I just think you need to be careful when saying :</p>\n<blockquote>\n  <p>Would you ever evaluate a model by its training score? Of course not! Every child learns that models must be scored on a validation dataset so that we can evaluate its generalization capability. The same principle holds for feature importance.</p>\n</blockquote>\n<p>In fact it is very debatable that a model could/should be explained by a validation set.</p>\n<p>Here you are looking at feature importance only to remove some and improve your global CV score. But that's only one specific use case.</p>\n<p>Feature importance and more broadly 'explainability' of a model are often used to explain what the model is doing and give trust to people relying on it.<br>\nWhen trying to explain a model, I think it makes sense to look at what it learnt during training and where the boundaries have been drawn, this only comes from the training data.</p>\n<p>In the end, I think it's not black and white because methods based on validation set will have other problems like data shifts, sampling representativity etc… </p>\n<p>Cheers!</p>",
      "rawMarkdown": "Nice and interesting post, here is my 2cts:\n\nI just think you need to be careful when saying :\n> Would you ever evaluate a model by its training score? Of course not! Every child learns that models must be scored on a validation dataset so that we can evaluate its generalization capability. The same principle holds for feature importance.\n\nIn fact it is very debatable that a model could/should be explained by a validation set.\n\nHere you are looking at feature importance only to remove some and improve your global CV score. But that's only one specific use case.\n\nFeature importance and more broadly 'explainability' of a model are often used to explain what the model is doing and give trust to people relying on it.\nWhen trying to explain a model, I think it makes sense to look at what it learnt during training and where the boundaries have been drawn, this only comes from the training data.\n\nIn the end, I think it's not black and white because methods based on validation set will have other problems like data shifts, sampling representativity etc... \n\nCheers!",
      "votes": 6,
      "replies": [
        {
          "id": 1880465,
          "postDate": "2022-08-01T18:37:29.360Z",
          "content": "<p>Good point, <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a>!</p>",
          "rawMarkdown": "Good point, @optimo!",
          "votes": 2
        },
        {
          "id": 1908886,
          "postDate": "2022-08-22T05:53:19.813Z",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> </p>\n<blockquote>\n  <p>\"<em>look at what it learnt during training and where the boundaries have been drawn</em>\"</p>\n</blockquote>\n<p>That is very relevant when it comes to tree based estimators. Such estimators cannot extrapolate (for example see the short notebook <a href=\"https://www.kaggle.com/code/carlmcbrideellis/extrapolation-do-not-stray-out-of-the-forest\" target=\"_blank\">\"Extrapolation: Do not stray out of the forest!\"</a>). This causes particular problems for permutation importance when they have been trained with correlated features; shuffling one of these features will create data-point pairs that are now either outside of the region that the estimator was trained on, or on the wrong side of a decision boundary, and the results for these points will become unreliable. There is an interesting (although somewhat heavy reading) paper on the subject: <a href=\"https://arxiv.org/pdf/1905.03151.pdf\" target=\"_blank\">\"<em>Unrestricted Permutation forces Extrapolation: Variable Importance Requires at least One More Model or There Is No Free Variable Importance</em>\"</a>.</p>\n<p>All the best,<br>\ncarl</p>",
          "rawMarkdown": "Dear @optimo \n\n> \"*look at what it learnt during training and where the boundaries have been drawn*\"\n\nThat is very relevant when it comes to tree based estimators. Such estimators cannot extrapolate (for example see the short notebook [\"Extrapolation: Do not stray out of the forest!\"](https://www.kaggle.com/code/carlmcbrideellis/extrapolation-do-not-stray-out-of-the-forest)). This causes particular problems for permutation importance when they have been trained with correlated features; shuffling one of these features will create data-point pairs that are now either outside of the region that the estimator was trained on, or on the wrong side of a decision boundary, and the results for these points will become unreliable. There is an interesting (although somewhat heavy reading) paper on the subject: [\"*Unrestricted Permutation forces Extrapolation: Variable Importance Requires at least One More Model or There Is No Free Variable Importance*\"](https://arxiv.org/pdf/1905.03151.pdf).\n\nAll the best,\ncarl",
          "votes": 2
        }
      ]
    },
    {
      "id": 1822610,
      "postDate": "2022-06-16T14:06:03.730Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> nice demonstration ! <br>\nThe only small limitations are a compatibility with scikit-learn API required for the model and the scoring that can not use (it seems) a custom metric but the scikit learn metrics. Amex scoring is so specific…<br>\nI noticed that you can get a gain when removing the worse score feature but if you do it again with the next one you may get a bad result. I guess the permutation can not take into account the cumulative result of 2 features deletion.</p>",
      "rawMarkdown": "Thank you @ambrosm nice demonstration ! \nThe only small limitations are a compatibility with scikit-learn API required for the model and the scoring that can not use (it seems) a custom metric but the scikit learn metrics. Amex scoring is so specific...\nI noticed that you can get a gain when removing the worse score feature but if you do it again with the next one you may get a bad result. I guess the permutation can not take into account the cumulative result of 2 features deletion.",
      "votes": 6,
      "replies": [
        {
          "id": 1857109,
          "postDate": "2022-07-15T20:51:09.903Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/pourchot\" target=\"_blank\">@pourchot</a> </p>\n<blockquote>\n  <p>I noticed that you can get a gain when removing the worse score feature but if you do it again with <strong>the next one</strong>…</p>\n</blockquote>\n<p>I did permutation and got a list of values corresponding to each feature. <br>\nNow, when we remove feature corresponding to worst score, we get a gain, If we now remove the <strong>Second Worse</strong> <br>\n feature from the list we may get a bad result.<br>\nDid you mean <strong>Second Worse</strong> when you tell <strong>next one</strong> or something else[like run permutation again and remove the worst]. </p>\n<p>Thanks,</p>",
          "rawMarkdown": "Hi @pourchot \n> I noticed that you can get a gain when removing the worse score feature but if you do it again with **the next one**...\n\nI did permutation and got a list of values corresponding to each feature. \nNow, when we remove feature corresponding to worst score, we get a gain, If we now remove the **Second Worse** \n feature from the list we may get a bad result.\nDid you mean **Second Worse** when you tell **next one** or something else[like run permutation again and remove the worst]. \n\nThanks,"
        }
      ]
    },
    {
      "id": 1826588,
      "postDate": "2022-06-20T13:19:52.497Z",
      "content": "<p>hi <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>, this is really great. From my own personal experience, the Permutation Feature Importance was great with text data where we wanted to identify the feature with right level of grams - usually the bigrams works better and the unigrams works the worst - Bigrams &gt; Trigrams &gt; Unigrams :)</p>",
      "rawMarkdown": "hi @ambrosm, this is really great. From my own personal experience, the Permutation Feature Importance was great with text data where we wanted to identify the feature with right level of grams - usually the bigrams works better and the unigrams works the worst - Bigrams > Trigrams > Unigrams :)\n",
      "votes": 4
    },
    {
      "id": 1992217,
      "postDate": "2022-10-17T15:25:22.097Z",
      "content": "<p>Greetings, AmbrosM!</p>\n<p>Thanks for sharing your knoledge!<br>\nAnswer my questions, please.</p>\n<ol>\n<li><p>Can we get a data leak (or something like data leak) if first we make permutation feature importance calculation on X,Y and then select some features and make model.fit(X[features],Y)?</p></li>\n<li><p>Can we make model.fit(X,Y) and permutation_importance(model, X, Y) on full data?<br>\nOr we should use train_test_split and then model.fit(Xtrain,Ytrain), permutation_importance(model, Xtest, Ytest)?</p></li>\n</ol>",
      "rawMarkdown": "Greetings, AmbrosM!\n\nThanks for sharing your knoledge!\nAnswer my questions, please.\n\n1. Can we get a data leak (or something like data leak) if first we make permutation feature importance calculation on X,Y and then select some features and make model.fit(X[features],Y)?\n\n2. Can we make model.fit(X,Y) and permutation_importance(model, X, Y) on full data?\nOr we should use train_test_split and then model.fit(Xtrain,Ytrain), permutation_importance(model, Xtest, Ytest)?",
      "votes": 1,
      "replies": [
        {
          "id": 1993771,
          "postDate": "2022-10-18T15:31:30.467Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/kaggledummie007\" target=\"_blank\">@kaggledummie007</a> </p>\n<p>Yes, it can be seen as something like a data leak: After you train on X_train and use X_val to calculate permutation feature importance and select features, the model will contain some leaked information about X_val. The situation is comparable to training on X_train and using X_val in Optuna to tune hyperparameters.  </p>\n<p>A possible solution is:</p>\n<ol>\n<li>Split the the data into three parts: X_train, X_val1, X_val2</li>\n<li>Train on X_train</li>\n<li>Evaluate permutation feature importance with <code>sklearn.inspection.permutation_importance</code> on X_val1 and select features</li>\n<li>Retrain on the selected features of X_train and X_val1; validate on X_val2</li>\n</ol>",
          "rawMarkdown": "Hi @kaggledummie007 \n\nYes, it can be seen as something like a data leak: After you train on X_train and use X_val to calculate permutation feature importance and select features, the model will contain some leaked information about X_val. The situation is comparable to training on X_train and using X_val in Optuna to tune hyperparameters.  \n\nA possible solution is:\n\n1. Split the the data into three parts: X_train, X_val1, X_val2\n2. Train on X_train\n3. Evaluate permutation feature importance with `sklearn.inspection.permutation_importance` on X_val1 and select features\n4. Retrain on the selected features of X_train and X_val1; validate on X_val2",
          "votes": 5
        },
        {
          "id": 1993910,
          "postDate": "2022-10-18T16:57:39.420Z",
          "content": "<p>Thank you so much!<br>\nI can't overestimate your sharing experience!</p>",
          "rawMarkdown": "Thank you so much!\nI can't overestimate your sharing experience!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1873352,
      "postDate": "2022-07-27T14:39:58.820Z",
      "content": "<p>Thanks for posting this! I had no idea that features being shuffled had an impact on the model accuracy. This is a great explainer post and really thoughtfully put together.</p>",
      "rawMarkdown": "Thanks for posting this! I had no idea that features being shuffled had an impact on the model accuracy. This is a great explainer post and really thoughtfully put together.",
      "votes": 1
    },
    {
      "id": 1861467,
      "postDate": "2022-07-19T03:44:00.640Z",
      "content": "<p><a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> Thanks for sharing. About the feature selection, how about using the boruta method before conducting the xgb  training? </p>",
      "rawMarkdown": " @ambrosm Thanks for sharing. About the feature selection, how about using the boruta method before conducting the xgb  training? ",
      "votes": 1,
      "replies": [
        {
          "id": 1861631,
          "postDate": "2022-07-19T06:18:33.513Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/alexyoyo\" target=\"_blank\">@alexyoyo</a> I don't know - I have no experience with Boruta.</p>",
          "rawMarkdown": "Hi @alexyoyo I don't know - I have no experience with Boruta.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1851459,
      "postDate": "2022-07-11T08:55:25.977Z",
      "content": "<p>Thanks a lot, super usefull information!</p>",
      "rawMarkdown": "Thanks a lot, super usefull information!",
      "votes": 1
    },
    {
      "id": 1827137,
      "postDate": "2022-06-20T21:27:18.627Z",
      "content": "<p>Looks great. very helpful</p>",
      "rawMarkdown": "Looks great. very helpful",
      "votes": 1
    },
    {
      "id": 1826142,
      "postDate": "2022-06-20T04:45:21.090Z",
      "content": "<p>This is super helpful. Great post!</p>",
      "rawMarkdown": "This is super helpful. Great post!",
      "votes": 1
    },
    {
      "id": 1824173,
      "postDate": "2022-06-18T03:56:47.733Z",
      "content": "<p>thx for sharing those insights!</p>",
      "rawMarkdown": "thx for sharing those insights!",
      "votes": 1
    },
    {
      "id": 1823801,
      "postDate": "2022-06-17T16:53:39.880Z",
      "content": "<p>Very much insight and good summery as always <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>",
      "rawMarkdown": "Very much insight and good summery as always @ambrosm ",
      "votes": 1
    },
    {
      "id": 1823540,
      "postDate": "2022-06-17T13:21:12.887Z",
      "content": "<p>Nice summary. Autoencoders are also a nice way to \"select\" (generate) important features out of the existing onces. Especially denosing autoencoders can be useful in this competition.</p>",
      "rawMarkdown": "Nice summary. Autoencoders are also a nice way to \"select\" (generate) important features out of the existing onces. Especially denosing autoencoders can be useful in this competition.",
      "votes": 1
    },
    {
      "id": 1822413,
      "postDate": "2022-06-16T10:55:27.583Z",
      "content": "<p>your posts always bring so many insights. Thank you <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>",
      "rawMarkdown": "your posts always bring so many insights. Thank you @ambrosm ",
      "votes": 1
    },
    {
      "id": 1821844,
      "postDate": "2022-06-15T21:24:37.580Z",
      "content": "<p>I always learn a lot with your posts, thanks <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>",
      "rawMarkdown": "I always learn a lot with your posts, thanks @ambrosm ",
      "votes": 1
    },
    {
      "id": 1927452,
      "postDate": "2022-09-05T16:21:21.790Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>. Great post! Thank you for sharing the knowledge. I have been using Mutual Information (MI) score to select important features since it assess each feature on both linear and non-linear linkage with output variable. It is computationally expensive but are there any other shortcomings in it? Please let me know.</p>",
      "rawMarkdown": "Hey @ambrosm. Great post! Thank you for sharing the knowledge. I have been using Mutual Information (MI) score to select important features since it assess each feature on both linear and non-linear linkage with output variable. It is computationally expensive but are there any other shortcomings in it? Please let me know."
    },
    {
      "id": 1892533,
      "postDate": "2022-08-10T06:17:26.210Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1841937,
      "postDate": "2022-07-03T15:03:02.543Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1825622,
      "postDate": "2022-06-19T14:59:04.757Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1825148,
      "postDate": "2022-06-19T02:28:50.487Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 1825175,
          "postDate": "2022-06-19T04:03:56.890Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yiwangnz\" target=\"_blank\">@yiwangnz</a> It seems that the gradient-boosting libraries provide these metrics only on the training set. So theoretically, you could measure these metrics on the validation set, but you would have to program them yourself. And this still wouldn't remedy the issue that they are always nonnegative.</p>",
          "rawMarkdown": "Hi @yiwangnz It seems that the gradient-boosting libraries provide these metrics only on the training set. So theoretically, you could measure these metrics on the validation set, but you would have to program them yourself. And this still wouldn't remedy the issue that they are always nonnegative.",
          "votes": 5
        }
      ]
    },
    {
      "id": 1877726,
      "postDate": "2022-07-31T01:08:51.050Z",
      "content": "<p>Very helpful. Thanks for sharing👍</p>",
      "rawMarkdown": "Very helpful. Thanks for sharing👍",
      "votes": 1
    },
    {
      "id": 1877294,
      "postDate": "2022-07-30T14:02:47.213Z",
      "content": "<p>Thanks for sharing this idea!</p>",
      "rawMarkdown": "Thanks for sharing this idea!",
      "votes": 1
    },
    {
      "id": 1842878,
      "postDate": "2022-07-04T11:13:12.867Z",
      "content": "<p>Solid advice. Thank you!</p>",
      "rawMarkdown": "Solid advice. Thank you!",
      "votes": 1
    },
    {
      "id": 1842686,
      "postDate": "2022-07-04T07:54:21.453Z",
      "content": "<p>Thanks. Very helpful !</p>",
      "rawMarkdown": "Thanks. Very helpful !",
      "votes": 1
    },
    {
      "id": 1842442,
      "postDate": "2022-07-04T02:27:28.627Z",
      "content": "<p>good thanks for sharing</p>",
      "rawMarkdown": "good thanks for sharing",
      "votes": 1
    },
    {
      "id": 1837750,
      "postDate": "2022-06-29T21:24:32.750Z",
      "content": "<p>Excellent post, thanks for sharing!</p>",
      "rawMarkdown": "Excellent post, thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1831244,
      "postDate": "2022-06-24T04:35:21.857Z",
      "content": "<p>great job!Thanks</p>",
      "rawMarkdown": "great job!Thanks",
      "votes": 1
    },
    {
      "id": 1829134,
      "postDate": "2022-06-22T10:47:53.717Z",
      "content": "<p>thanks, sharing.</p>",
      "rawMarkdown": "thanks, sharing.",
      "votes": 1
    },
    {
      "id": 1825638,
      "postDate": "2022-06-19T15:23:31.213Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 1824897,
      "postDate": "2022-06-18T17:53:40.110Z",
      "content": "<p>Thanks for the head-up. </p>",
      "rawMarkdown": "Thanks for the head-up. ",
      "votes": 1
    },
    {
      "id": 1823636,
      "postDate": "2022-06-17T14:40:31.270Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>",
      "rawMarkdown": "Thanks for sharing @ambrosm ",
      "votes": 1
    },
    {
      "id": 1822395,
      "postDate": "2022-06-16T10:32:36.960Z",
      "content": "<p>Thanks for sharing the insights.</p>",
      "rawMarkdown": "Thanks for sharing the insights.",
      "votes": 1
    },
    {
      "id": 1909895,
      "postDate": "2022-08-23T03:24:27.937Z",
      "content": "<p>Thanks for sharing! It helps a lot!</p>",
      "rawMarkdown": "Thanks for sharing! It helps a lot!",
      "votes": 2
    },
    {
      "id": 1821867,
      "postDate": "2022-06-15T21:35:04.473Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>",
      "rawMarkdown": "Thanks for sharing @ambrosm ",
      "votes": 2
    },
    {
      "id": 1899624,
      "postDate": "2022-08-15T11:57:36.300Z",
      "content": "<p>Very helpful. Thanks</p>",
      "rawMarkdown": "Very helpful. Thanks"
    }
  ],
  "comments": [
    {
      "id": 1822040,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-06-16T03:14:20.233000",
      "content": "<p>I think you are very right, and your conclusion is very helpful. I have been using permutation feature importance for a while and I think it's a great tool.</p>\n<p>I would like to add one thing though: when you shuffle a feature you dont take in cosideration the interaction it has with other features that might still pose a problem for feature selection. <br>\nFor example, if you have two columns that are the same: You can permute them all day long your model will always use one of them and the score won't change. </p>",
      "votes": 20,
      "replies": [
        {
          "id": 1825146,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-06-19T02:25:08.750000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1842765,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2022-07-04T09:10:38.970000",
          "content": "<p><a href=\"https://www.kaggle.com/yiwangnz\" target=\"_blank\">@yiwangnz</a> see it working. Besides being compute-intensive I see the task in figuring out for which threshold to group two features. This dataset has a high dimensionality after the regular feature engineering done by top scoring notebooks. It would not be easy to group two features from the same category without removing important correlations. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1863307,
          "author_name": "Dou Fan",
          "author_url": "",
          "post_date": "2022-07-20T09:36:03.200000",
          "content": "<p>i think if score don't change means the feature doesn't work and we can drop one. When we evaluate the other one we can get the real feature importance, is that a problem?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1822914,
      "author_name": "mavillan",
      "author_url": "",
      "post_date": "2022-06-16T20:17:58.190000",
      "content": "<p>Contributing to the discussion, I think <strong>lofo</strong> (<a href=\"https://github.com/aerdem4/lofo-importance\" target=\"_blank\">https://github.com/aerdem4/lofo-importance</a>) is even a more robust approach for measuring feature importance, but it's computationally expensive as it needs to train a model for each feature in the dataset:</p>\n<blockquote>\n  <p>LOFO first evaluates the performance of the model with all the input features included, then iteratively removes one feature at a time, retrains the model, and evaluates its performance on a validation set.</p>\n</blockquote>\n<p>In the same package there also a <strong>fast-lofo</strong> (<a href=\"https://github.com/aerdem4/lofo-importance#flofo-importance\" target=\"_blank\">https://github.com/aerdem4/lofo-importance#flofo-importance</a>) method, that is basically a supercharged permutation importance, that deals with the correlated feature overestimation problem.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 1844185,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2022-07-05T11:59:46.507000",
          "content": "<p>If you have an estimator which provides information about feature importance you can basically use recursive feature elimination from sklearn instead:<br>\n<a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.RFE.html\" target=\"_blank\">https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.RFE.html</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1885124,
      "author_name": "gzguevara",
      "author_url": "",
      "post_date": "2022-08-04T22:59:34.070000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> for sharing! I implemented it and it helps a lot!</p>\n<p>I would just annotate the following for others working with tree based algorithms. In tree based algorithms the order of your features matters a lot for your model! So if you give [A, B, C] to xgboost you will get a different result compared to [B, A, C]. </p>\n<p>I noticed that my permutation importance changes a lot depending on the feature order! </p>\n<p>Therefore, if you are looking to find the \"general permutation importance\" of your features, independent of your model, fold or feature order, you might consider to perform several permutation importance valuations over different folds with different feature orders and average them. That helped me a lot! :)</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1880390,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2022-08-01T17:08:30.933000",
      "content": "<p>Nice and interesting post, here is my 2cts:</p>\n<p>I just think you need to be careful when saying :</p>\n<blockquote>\n  <p>Would you ever evaluate a model by its training score? Of course not! Every child learns that models must be scored on a validation dataset so that we can evaluate its generalization capability. The same principle holds for feature importance.</p>\n</blockquote>\n<p>In fact it is very debatable that a model could/should be explained by a validation set.</p>\n<p>Here you are looking at feature importance only to remove some and improve your global CV score. But that's only one specific use case.</p>\n<p>Feature importance and more broadly 'explainability' of a model are often used to explain what the model is doing and give trust to people relying on it.<br>\nWhen trying to explain a model, I think it makes sense to look at what it learnt during training and where the boundaries have been drawn, this only comes from the training data.</p>\n<p>In the end, I think it's not black and white because methods based on validation set will have other problems like data shifts, sampling representativity etc… </p>\n<p>Cheers!</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1880465,
          "author_name": "AmbrosM",
          "author_url": "",
          "post_date": "2022-08-01T18:37:29.360000",
          "content": "<p>Good point, <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a>!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1908886,
          "author_name": "Carl McBride Ellis",
          "author_url": "",
          "post_date": "2022-08-22T05:53:19.813000",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> </p>\n<blockquote>\n  <p>\"<em>look at what it learnt during training and where the boundaries have been drawn</em>\"</p>\n</blockquote>\n<p>That is very relevant when it comes to tree based estimators. Such estimators cannot extrapolate (for example see the short notebook <a href=\"https://www.kaggle.com/code/carlmcbrideellis/extrapolation-do-not-stray-out-of-the-forest\" target=\"_blank\">\"Extrapolation: Do not stray out of the forest!\"</a>). This causes particular problems for permutation importance when they have been trained with correlated features; shuffling one of these features will create data-point pairs that are now either outside of the region that the estimator was trained on, or on the wrong side of a decision boundary, and the results for these points will become unreliable. There is an interesting (although somewhat heavy reading) paper on the subject: <a href=\"https://arxiv.org/pdf/1905.03151.pdf\" target=\"_blank\">\"<em>Unrestricted Permutation forces Extrapolation: Variable Importance Requires at least One More Model or There Is No Free Variable Importance</em>\"</a>.</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1822610,
      "author_name": "Laurent Pourchot",
      "author_url": "",
      "post_date": "2022-06-16T14:06:03.730000",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> nice demonstration ! <br>\nThe only small limitations are a compatibility with scikit-learn API required for the model and the scoring that can not use (it seems) a custom metric but the scikit learn metrics. Amex scoring is so specific…<br>\nI noticed that you can get a gain when removing the worse score feature but if you do it again with the next one you may get a bad result. I guess the permutation can not take into account the cumulative result of 2 features deletion.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1857109,
          "author_name": "AKR",
          "author_url": "",
          "post_date": "2022-07-15T20:51:09.903000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/pourchot\" target=\"_blank\">@pourchot</a> </p>\n<blockquote>\n  <p>I noticed that you can get a gain when removing the worse score feature but if you do it again with <strong>the next one</strong>…</p>\n</blockquote>\n<p>I did permutation and got a list of values corresponding to each feature. <br>\nNow, when we remove feature corresponding to worst score, we get a gain, If we now remove the <strong>Second Worse</strong> <br>\n feature from the list we may get a bad result.<br>\nDid you mean <strong>Second Worse</strong> when you tell <strong>next one</strong> or something else[like run permutation again and remove the worst]. </p>\n<p>Thanks,</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1826588,
      "author_name": "MuralidharanM",
      "author_url": "",
      "post_date": "2022-06-20T13:19:52.497000",
      "content": "<p>hi <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>, this is really great. From my own personal experience, the Permutation Feature Importance was great with text data where we wanted to identify the feature with right level of grams - usually the bigrams works better and the unigrams works the worst - Bigrams &gt; Trigrams &gt; Unigrams :)</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1992217,
      "author_name": "Oleg Khudyakov",
      "author_url": "",
      "post_date": "2022-10-17T15:25:22.097000",
      "content": "<p>Greetings, AmbrosM!</p>\n<p>Thanks for sharing your knoledge!<br>\nAnswer my questions, please.</p>\n<ol>\n<li><p>Can we get a data leak (or something like data leak) if first we make permutation feature importance calculation on X,Y and then select some features and make model.fit(X[features],Y)?</p></li>\n<li><p>Can we make model.fit(X,Y) and permutation_importance(model, X, Y) on full data?<br>\nOr we should use train_test_split and then model.fit(Xtrain,Ytrain), permutation_importance(model, Xtest, Ytest)?</p></li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 1993771,
          "author_name": "AmbrosM",
          "author_url": "",
          "post_date": "2022-10-18T15:31:30.467000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/kaggledummie007\" target=\"_blank\">@kaggledummie007</a> </p>\n<p>Yes, it can be seen as something like a data leak: After you train on X_train and use X_val to calculate permutation feature importance and select features, the model will contain some leaked information about X_val. The situation is comparable to training on X_train and using X_val in Optuna to tune hyperparameters.  </p>\n<p>A possible solution is:</p>\n<ol>\n<li>Split the the data into three parts: X_train, X_val1, X_val2</li>\n<li>Train on X_train</li>\n<li>Evaluate permutation feature importance with <code>sklearn.inspection.permutation_importance</code> on X_val1 and select features</li>\n<li>Retrain on the selected features of X_train and X_val1; validate on X_val2</li>\n</ol>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1993910,
          "author_name": "Oleg Khudyakov",
          "author_url": "",
          "post_date": "2022-10-18T16:57:39.420000",
          "content": "<p>Thank you so much!<br>\nI can't overestimate your sharing experience!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1873352,
      "author_name": "Kevin Kwan",
      "author_url": "",
      "post_date": "2022-07-27T14:39:58.820000",
      "content": "<p>Thanks for posting this! I had no idea that features being shuffled had an impact on the model accuracy. This is a great explainer post and really thoughtfully put together.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1861467,
      "author_name": "Alexyoyo",
      "author_url": "",
      "post_date": "2022-07-19T03:44:00.640000",
      "content": "<p><a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> Thanks for sharing. About the feature selection, how about using the boruta method before conducting the xgb  training? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1861631,
          "author_name": "AmbrosM",
          "author_url": "",
          "post_date": "2022-07-19T06:18:33.513000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/alexyoyo\" target=\"_blank\">@alexyoyo</a> I don't know - I have no experience with Boruta.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1851459,
      "author_name": "Oleg Khudyakov",
      "author_url": "",
      "post_date": "2022-07-11T08:55:25.977000",
      "content": "<p>Thanks a lot, super usefull information!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1827137,
      "author_name": "Muatasm Abdelsalam",
      "author_url": "",
      "post_date": "2022-06-20T21:27:18.627000",
      "content": "<p>Looks great. very helpful</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1826142,
      "author_name": "Eric Inman",
      "author_url": "",
      "post_date": "2022-06-20T04:45:21.090000",
      "content": "<p>This is super helpful. Great post!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1824173,
      "author_name": "qilin chen",
      "author_url": "",
      "post_date": "2022-06-18T03:56:47.733000",
      "content": "<p>thx for sharing those insights!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1823801,
      "author_name": "Gaju Ahmed",
      "author_url": "",
      "post_date": "2022-06-17T16:53:39.880000",
      "content": "<p>Very much insight and good summery as always <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1823540,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2022-06-17T13:21:12.887000",
      "content": "<p>Nice summary. Autoencoders are also a nice way to \"select\" (generate) important features out of the existing onces. Especially denosing autoencoders can be useful in this competition.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1822413,
      "author_name": "naiborhujosua",
      "author_url": "",
      "post_date": "2022-06-16T10:55:27.583000",
      "content": "<p>your posts always bring so many insights. Thank you <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1821844,
      "author_name": "mavillan",
      "author_url": "",
      "post_date": "2022-06-15T21:24:37.580000",
      "content": "<p>I always learn a lot with your posts, thanks <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1927452,
      "author_name": "Nishant Jairath",
      "author_url": "",
      "post_date": "2022-09-05T16:21:21.790000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>. Great post! Thank you for sharing the knowledge. I have been using Mutual Information (MI) score to select important features since it assess each feature on both linear and non-linear linkage with output variable. It is computationally expensive but are there any other shortcomings in it? Please let me know.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1892533,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-10T06:17:26.210000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1841937,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-03T15:03:02.543000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1825622,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-19T14:59:04.757000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1825148,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-19T02:28:50.487000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1825175,
          "author_name": "AmbrosM",
          "author_url": "",
          "post_date": "2022-06-19T04:03:56.890000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yiwangnz\" target=\"_blank\">@yiwangnz</a> It seems that the gradient-boosting libraries provide these metrics only on the training set. So theoretically, you could measure these metrics on the validation set, but you would have to program them yourself. And this still wouldn't remedy the issue that they are always nonnegative.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1877726,
      "author_name": "YASUKENN@Meiji_IMS_ND",
      "author_url": "",
      "post_date": "2022-07-31T01:08:51.050000",
      "content": "<p>Very helpful. Thanks for sharing👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1877294,
      "author_name": "ChuanhaoLi2022",
      "author_url": "",
      "post_date": "2022-07-30T14:02:47.213000",
      "content": "<p>Thanks for sharing this idea!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1842878,
      "author_name": "Sohail Ahmed",
      "author_url": "",
      "post_date": "2022-07-04T11:13:12.867000",
      "content": "<p>Solid advice. Thank you!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1842686,
      "author_name": "andy jennings",
      "author_url": "",
      "post_date": "2022-07-04T07:54:21.453000",
      "content": "<p>Thanks. Very helpful !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1842442,
      "author_name": "hgn1102",
      "author_url": "",
      "post_date": "2022-07-04T02:27:28.627000",
      "content": "<p>good thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1837750,
      "author_name": "Adam Wurdits",
      "author_url": "",
      "post_date": "2022-06-29T21:24:32.750000",
      "content": "<p>Excellent post, thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1831244,
      "author_name": "Kyrieshaohua",
      "author_url": "",
      "post_date": "2022-06-24T04:35:21.857000",
      "content": "<p>great job!Thanks</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1829134,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-22T10:47:53.717000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1825638,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-19T15:23:31.213000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1824897,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-18T17:53:40.110000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1823636,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-17T14:40:31.270000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1822395,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-16T10:32:36.960000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1909895,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-23T03:24:27.937000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1821867,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-15T21:35:04.473000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1899624,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-15T11:57:36.300000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1821813": "Why do we evaluate feature importances? Because we want to select the good features and drop the bad ones. But which of the following three importances is the right one for this purpose?\n\n1. Split feature importance?\n2. Gain feature importance?\n3. Permutation feature importance?\n\nIn this post I'll compare the three methods for feature selection and show that the last one is the right one.\n\nThe former two methods are often seen in public gradient-boosting notebooks because people like to decorate their notebooks with a nice diagram, all the more as a single line of code suffices to generate the diagram. But the diagram alone doesn't give a better model! We need to interpret the importance values.\n\n**Split feature importance** counts how often a feature is used in the model. The model I'm using for this experiment has 1200 trees, and split feature importances are between 0 and 1082. If a feature has importance = 0, this means that the feature is not used in the model at all. If we omit the feature, model quality doesn't change, but the training takes less time. If split feature importance is high, we know that the feature is used often in the model and that it improves the model's training score. \n\n**Gain feature importance** is similar to split feature importance. It measures the gain which was achieved in all splits based on the feature, although the definition of this gain is hard to find in the documentation. The features of my model have importances between 0 and 3000000. Again, if a feature has importance = 0, this means that the feature is not used in the model at all.\n\nThe following diagram shows that split and gain feature importance correlate quite well. The diagram is based on importance ranks, which are between 0 and 573 (my model has 574 features) - higher is better. The single dot at coordinates (30,30) represents 60 features which are not used in the model and have both importances equal to zero.\n\n![split-gain](https://i.imgur.com/yS4eQXF.png)\n\nSplit and gain importance are cheap byproducts of the training process, but they have a massive shortcoming: They are based on the training data rather than the validation data and consequently cannot measure generalization. Would you ever evaluate a model by its training score? Of course not! Every child learns that models must be scored on a validation dataset so that we can evaluate its generalization capability. The same principle holds for feature importance.\n\n**Permutation feature importance** is a model inspection technique that can be used for any fitted estimator when the data is tabular. The permutation feature importance is defined as the decrease in a model's validation score when a single feature value is randomly shuffled. In contrast to split and gain importance, permutation feature importance can be positive or negative:\n- Positive permutation feature importance means that the model profits from the feature and deteriorates if the feature is shuffled or removed.\n- Zero permutation feature importance means that the model doesn't change if we shuffle or remove the feature.\n- Negative permutation feature importance means that the model improves if the feature is removed.\n\nIf a feature is complete noise, a decision tree can still use it to improve the training score, and the feature will get positive split/gain importances. Only permutation feature importance detects that such a feature doesn't generalize to the validation dataset.\n\nThe following diagrams show how permutation feature importance compares to split and gain importance:\n- The green dots represent features with positive permutation feature importance. These are the good features.\n- The black dot at the left represents the 60 features which are not used in the model.\n- The red dots represent the features with negative permutation feature importance. These are the bad features although they may have high split or gain importance. D_43_min and S_26_min (represented by the two rightmost red dots) belong to the top 40 features according to split importance, but the model gets better if we drop them.\n\n![split-gain-pfi](https://i.imgur.com/CBWXsbu.png)\n\n**Conclusion:** Although permutation feature importance is slower than the gradient boosters' built-in importance measures, it is much more useful because it measures generalization and distinguishes positive from negative importances. Features with negative importance should be dropped from the model.\n\nSee [here](https://scikit-learn.org/stable/auto_examples/inspection/plot_permutation_importance.html) for a demo of the same effect on the Titanic dataset.",
    "1822040": "I think you are very right, and your conclusion is very helpful. I have been using permutation feature importance for a while and I think it's a great tool.\n\nI would like to add one thing though: when you shuffle a feature you dont take in cosideration the interaction it has with other features that might still pose a problem for feature selection. \nFor example, if you have two columns that are the same: You can permute them all day long your model will always use one of them and the score won't change. \n",
    "1822914": "Contributing to the discussion, I think **lofo** (https://github.com/aerdem4/lofo-importance) is even a more robust approach for measuring feature importance, but it's computationally expensive as it needs to train a model for each feature in the dataset:\n> LOFO first evaluates the performance of the model with all the input features included, then iteratively removes one feature at a time, retrains the model, and evaluates its performance on a validation set.\n\nIn the same package there also a **fast-lofo** (https://github.com/aerdem4/lofo-importance#flofo-importance) method, that is basically a supercharged permutation importance, that deals with the correlated feature overestimation problem.\n",
    "1885124": "Thanks @ambrosm for sharing! I implemented it and it helps a lot!\n\nI would just annotate the following for others working with tree based algorithms. In tree based algorithms the order of your features matters a lot for your model! So if you give [A, B, C] to xgboost you will get a different result compared to [B, A, C]. \n\nI noticed that my permutation importance changes a lot depending on the feature order! \n\nTherefore, if you are looking to find the \"general permutation importance\" of your features, independent of your model, fold or feature order, you might consider to perform several permutation importance valuations over different folds with different feature orders and average them. That helped me a lot! :)",
    "1880390": "Nice and interesting post, here is my 2cts:\n\nI just think you need to be careful when saying :\n> Would you ever evaluate a model by its training score? Of course not! Every child learns that models must be scored on a validation dataset so that we can evaluate its generalization capability. The same principle holds for feature importance.\n\nIn fact it is very debatable that a model could/should be explained by a validation set.\n\nHere you are looking at feature importance only to remove some and improve your global CV score. But that's only one specific use case.\n\nFeature importance and more broadly 'explainability' of a model are often used to explain what the model is doing and give trust to people relying on it.\nWhen trying to explain a model, I think it makes sense to look at what it learnt during training and where the boundaries have been drawn, this only comes from the training data.\n\nIn the end, I think it's not black and white because methods based on validation set will have other problems like data shifts, sampling representativity etc... \n\nCheers!",
    "1822610": "Thank you @ambrosm nice demonstration ! \nThe only small limitations are a compatibility with scikit-learn API required for the model and the scoring that can not use (it seems) a custom metric but the scikit learn metrics. Amex scoring is so specific...\nI noticed that you can get a gain when removing the worse score feature but if you do it again with the next one you may get a bad result. I guess the permutation can not take into account the cumulative result of 2 features deletion.",
    "1826588": "hi @ambrosm, this is really great. From my own personal experience, the Permutation Feature Importance was great with text data where we wanted to identify the feature with right level of grams - usually the bigrams works better and the unigrams works the worst - Bigrams > Trigrams > Unigrams :)\n",
    "1992217": "Greetings, AmbrosM!\n\nThanks for sharing your knoledge!\nAnswer my questions, please.\n\n1. Can we get a data leak (or something like data leak) if first we make permutation feature importance calculation on X,Y and then select some features and make model.fit(X[features],Y)?\n\n2. Can we make model.fit(X,Y) and permutation_importance(model, X, Y) on full data?\nOr we should use train_test_split and then model.fit(Xtrain,Ytrain), permutation_importance(model, Xtest, Ytest)?",
    "1873352": "Thanks for posting this! I had no idea that features being shuffled had an impact on the model accuracy. This is a great explainer post and really thoughtfully put together.",
    "1861467": " @ambrosm Thanks for sharing. About the feature selection, how about using the boruta method before conducting the xgb  training? ",
    "1851459": "Thanks a lot, super usefull information!",
    "1827137": "Looks great. very helpful",
    "1826142": "This is super helpful. Great post!",
    "1824173": "thx for sharing those insights!",
    "1823801": "Very much insight and good summery as always @ambrosm ",
    "1823540": "Nice summary. Autoencoders are also a nice way to \"select\" (generate) important features out of the existing onces. Especially denosing autoencoders can be useful in this competition.",
    "1822413": "your posts always bring so many insights. Thank you @ambrosm ",
    "1821844": "I always learn a lot with your posts, thanks @ambrosm ",
    "1927452": "Hey @ambrosm. Great post! Thank you for sharing the knowledge. I have been using Mutual Information (MI) score to select important features since it assess each feature on both linear and non-linear linkage with output variable. It is computationally expensive but are there any other shortcomings in it? Please let me know.",
    "1892533": "",
    "1841937": "",
    "1825622": "",
    "1825148": "",
    "1877726": "Very helpful. Thanks for sharing👍",
    "1877294": "Thanks for sharing this idea!",
    "1842878": "Solid advice. Thank you!",
    "1842686": "Thanks. Very helpful !",
    "1842442": "good thanks for sharing",
    "1837750": "Excellent post, thanks for sharing!",
    "1831244": "great job!Thanks",
    "1829134": "thanks, sharing.",
    "1825638": "Thanks for sharing.",
    "1824897": "Thanks for the head-up. ",
    "1823636": "Thanks for sharing @ambrosm ",
    "1822395": "Thanks for sharing the insights.",
    "1909895": "Thanks for sharing! It helps a lot!",
    "1821867": "Thanks for sharing @ambrosm ",
    "1899624": "Very helpful. Thanks"
  }
}