{
  "id": 332558,
  "title": "An explanation of the Boruta-Shap Feature Selection Technique",
  "url": "/competitions/amex-default-prediction/discussion/332558",
  "author_name": "Susnato Dhar",
  "post_date": "2022-06-22T09:22:01.518000",
  "votes": 19,
  "comment_count": 0,
  "views": 0,
  "content": "<p><strong>This is an explanation of the Boruta-Shap Feature Selection.</strong></p>\n<h2>Boruta</h2>\n<p>It is a pretty smart algorithm used for Feature Selection. (just like Permutation feature selection). It was developed as a package of R(<a href=\"https://www.jstatsoft.org/article/view/v036i11\" target=\"_blank\">here</a>). There are some simple steps required for this algorithm.</p>\n<p>: <code>In Boruta the Features compete with the randomized version of themselves</code>. For example, suppose we have a small dataset consisting of 3 features, (<code>age</code>, <code>height</code>, <code>weight</code>) with only 5 entries and we have to predict <code>income</code>.</p>\n<p><img src=\"https://i.postimg.cc/2SdHN5fL/1-Pgw-A0-J423hc-JNpym8o-M9-w.png\" alt=\"\"></p>\n<p>Now we need to create another DataFrame by shuffling each feature. We will call it <code>Shadow features</code> and then we will concatenate this with the main DataFrame, let's call it X_baruta.</p>\n<p><img src=\"https://i.postimg.cc/7PctgcVh/1-z-Stv-S-9-Gp-DEJJFZb0-Av-G4-Q.png\" alt=\"\"></p>\n<p>: Now we have to fit the model(Random Forest) on X_baruta and y. After training the model, we will take the highest feature importance recorded among the Shadow features and consider it as the threshold value. When the importance of a feature(from the main data frame) is higher than this threshold, we will call it a <code>hit</code>. <strong>The idea is that a feature is useful only if it’s capable of doing better than the best-randomized feature</strong>.  The result is, </p>\n<p><img src=\"https://i.postimg.cc/V6DfGtgf/1-bt-H1-CNb6l-Q62zak-Bx-YECYg.png\" alt=\"\"></p>\n<p>Here, the threshold is 14%(shadow_height). It seems that only <code>age</code> and <code>height</code> are useful for this run.<br>\nBut we can't trust this one run and determine the importance of the features, so we will run it multiple times. </p>\n<p>:  Result after 20 trials is, </p>\n<p><img src=\"https://i.postimg.cc/x8cpzk1p/1-Zly3-ZYol-DDsvn8x8-Me-Ieg.png\" alt=\"\"></p>\n<p>All features will have only two outcomes: <code>hit</code> or <code>not hit</code>, therefore we can perform the previous step several times and build a binomial distribution out of the features.</p>\n<p><img src=\"https://i.postimg.cc/qvznH8Z5/1-yq-AUl-Mt-PUi-Fyr8g-YLFag-TA.png\" alt=\"\"><br>\n<a href=\"https://towardsdatascience.com/boruta-explained-the-way-i-wish-someone-explained-it-to-me-4489d70e154a\" target=\"_blank\">Image and Table Source</a></p>\n<p>: We will keep the features which are in the Green region, we will drop those features in the Red region and we haven't come to a conclusion for features in the Blue region. We can run the algorithm for more trials to reach a conclusion about those indecisive features.</p>\n<h2>SHAP (SHapley Additive exPlanations)</h2>\n<p><a href=\"https://github.com/slundberg/shap\" target=\"_blank\">SHAP</a> is a unified approach to explain the output of any machine learning model. SHAP connects game theory with local explanations to create the only consistent and accurate explainer. We can use it to compute feature importance for each model. More about it <a href=\"https://christophm.github.io/interpretable-ml-book/shap.html\" target=\"_blank\">here</a>. </p>\n<p>We use this feature importance technique with Boruta algorithm to get a robust feature selection technique. </p>\n<h2>My Results &amp; Conclusion</h2>\n<p>I used Boruta-Shap(XGBOOST as the base model) and was really amazed by the results. <br>\nThe dataset I use consists of 1328 features. The baseline result of my XGBOOST model was 0.7897781200376643 (using all of the features). But after running this algorithm it only accepted 116 <a href=\"https://www.kaggle.com/datasets/susnato/shapxgboost\" target=\"_blank\">features</a>. I trained the XGBOOST model with the accepted features and with the same hyper-parameters again and I got 0.7893407118088636!</p>\n<p><strong>I retained  99.9446163% of the performance by using only  8.7349398% of the features!</strong></p>\n<p>The accepted features are <a href=\"https://www.kaggle.com/datasets/susnato/shapxgboost\" target=\"_blank\">here</a>.</p>\n<p>The implementation and the results can be found in this notebook: <a href=\"https://www.kaggle.com/code/susnato/amex-borutashap-feature-selection/notebook\" target=\"_blank\">AMEX - BorutaShap Feature Selection</a></p>",
  "messages": [
    {
      "id": 1829070,
      "postDate": "2022-06-22T09:22:01.520Z",
      "content": "<p><strong>This is an explanation of the Boruta-Shap Feature Selection.</strong></p>\n<h2>Boruta</h2>\n<p>It is a pretty smart algorithm used for Feature Selection. (just like Permutation feature selection). It was developed as a package of R(<a href=\"https://www.jstatsoft.org/article/view/v036i11\" target=\"_blank\">here</a>). There are some simple steps required for this algorithm.</p>\n<p>: <code>In Boruta the Features compete with the randomized version of themselves</code>. For example, suppose we have a small dataset consisting of 3 features, (<code>age</code>, <code>height</code>, <code>weight</code>) with only 5 entries and we have to predict <code>income</code>.</p>\n<p><img src=\"https://i.postimg.cc/2SdHN5fL/1-Pgw-A0-J423hc-JNpym8o-M9-w.png\" alt=\"\"></p>\n<p>Now we need to create another DataFrame by shuffling each feature. We will call it <code>Shadow features</code> and then we will concatenate this with the main DataFrame, let's call it X_baruta.</p>\n<p><img src=\"https://i.postimg.cc/7PctgcVh/1-z-Stv-S-9-Gp-DEJJFZb0-Av-G4-Q.png\" alt=\"\"></p>\n<p>: Now we have to fit the model(Random Forest) on X_baruta and y. After training the model, we will take the highest feature importance recorded among the Shadow features and consider it as the threshold value. When the importance of a feature(from the main data frame) is higher than this threshold, we will call it a <code>hit</code>. <strong>The idea is that a feature is useful only if it’s capable of doing better than the best-randomized feature</strong>.  The result is, </p>\n<p><img src=\"https://i.postimg.cc/V6DfGtgf/1-bt-H1-CNb6l-Q62zak-Bx-YECYg.png\" alt=\"\"></p>\n<p>Here, the threshold is 14%(shadow_height). It seems that only <code>age</code> and <code>height</code> are useful for this run.<br>\nBut we can't trust this one run and determine the importance of the features, so we will run it multiple times. </p>\n<p>:  Result after 20 trials is, </p>\n<p><img src=\"https://i.postimg.cc/x8cpzk1p/1-Zly3-ZYol-DDsvn8x8-Me-Ieg.png\" alt=\"\"></p>\n<p>All features will have only two outcomes: <code>hit</code> or <code>not hit</code>, therefore we can perform the previous step several times and build a binomial distribution out of the features.</p>\n<p><img src=\"https://i.postimg.cc/qvznH8Z5/1-yq-AUl-Mt-PUi-Fyr8g-YLFag-TA.png\" alt=\"\"><br>\n<a href=\"https://towardsdatascience.com/boruta-explained-the-way-i-wish-someone-explained-it-to-me-4489d70e154a\" target=\"_blank\">Image and Table Source</a></p>\n<p>: We will keep the features which are in the Green region, we will drop those features in the Red region and we haven't come to a conclusion for features in the Blue region. We can run the algorithm for more trials to reach a conclusion about those indecisive features.</p>\n<h2>SHAP (SHapley Additive exPlanations)</h2>\n<p><a href=\"https://github.com/slundberg/shap\" target=\"_blank\">SHAP</a> is a unified approach to explain the output of any machine learning model. SHAP connects game theory with local explanations to create the only consistent and accurate explainer. We can use it to compute feature importance for each model. More about it <a href=\"https://christophm.github.io/interpretable-ml-book/shap.html\" target=\"_blank\">here</a>. </p>\n<p>We use this feature importance technique with Boruta algorithm to get a robust feature selection technique. </p>\n<h2>My Results &amp; Conclusion</h2>\n<p>I used Boruta-Shap(XGBOOST as the base model) and was really amazed by the results. <br>\nThe dataset I use consists of 1328 features. The baseline result of my XGBOOST model was 0.7897781200376643 (using all of the features). But after running this algorithm it only accepted 116 <a href=\"https://www.kaggle.com/datasets/susnato/shapxgboost\" target=\"_blank\">features</a>. I trained the XGBOOST model with the accepted features and with the same hyper-parameters again and I got 0.7893407118088636!</p>\n<p><strong>I retained  99.9446163% of the performance by using only  8.7349398% of the features!</strong></p>\n<p>The accepted features are <a href=\"https://www.kaggle.com/datasets/susnato/shapxgboost\" target=\"_blank\">here</a>.</p>\n<p>The implementation and the results can be found in this notebook: <a href=\"https://www.kaggle.com/code/susnato/amex-borutashap-feature-selection/notebook\" target=\"_blank\">AMEX - BorutaShap Feature Selection</a></p>",
      "rawMarkdown": "**This is an explanation of the Boruta-Shap Feature Selection.**\n\n\n## Boruta\nIt is a pretty smart algorithm used for Feature Selection. (just like Permutation feature selection). It was developed as a package of R([here](https://www.jstatsoft.org/article/view/v036i11)). There are some simple steps required for this algorithm.\n\n<u>**Step 1(Shadow Features)**</u>: `In Boruta the Features compete with the randomized version of themselves`. For example, suppose we have a small dataset consisting of 3 features, (`age`, `height`, `weight`) with only 5 entries and we have to predict `income`.\n\n![](https://i.postimg.cc/2SdHN5fL/1-Pgw-A0-J423hc-JNpym8o-M9-w.png)\n\nNow we need to create another DataFrame by shuffling each feature. We will call it `Shadow features` and then we will concatenate this with the main DataFrame, let's call it X_baruta.\n\n![](https://i.postimg.cc/7PctgcVh/1-z-Stv-S-9-Gp-DEJJFZb0-Av-G4-Q.png)\n\n\n<u>**Step 2(Fitting the model)**</u>: Now we have to fit the model(Random Forest) on X_baruta and y. After training the model, we will take the highest feature importance recorded among the Shadow features and consider it as the threshold value. When the importance of a feature(from the main data frame) is higher than this threshold, we will call it a `hit`. **The idea is that a feature is useful only if it’s capable of doing better than the best-randomized feature**.  The result is, \n\n![](https://i.postimg.cc/V6DfGtgf/1-bt-H1-CNb6l-Q62zak-Bx-YECYg.png)\n\nHere, the threshold is 14%(shadow_height). It seems that only `age` and `height` are useful for this run.\nBut we can't trust this one run and determine the importance of the features, so we will run it multiple times. \n\n\n<u>**Step 3(Binomial Distribution)**</u>:  Result after 20 trials is, \n\n![](https://i.postimg.cc/x8cpzk1p/1-Zly3-ZYol-DDsvn8x8-Me-Ieg.png)\n\nAll features will have only two outcomes: `hit` or `not hit`, therefore we can perform the previous step several times and build a binomial distribution out of the features.\n\n![](https://i.postimg.cc/qvznH8Z5/1-yq-AUl-Mt-PUi-Fyr8g-YLFag-TA.png)\n[Image and Table Source](https://towardsdatascience.com/boruta-explained-the-way-i-wish-someone-explained-it-to-me-4489d70e154a)\n\n<u>**Step 4(Keeping/Dropping Features)**</u>: We will keep the features which are in the Green region, we will drop those features in the Red region and we haven't come to a conclusion for features in the Blue region. We can run the algorithm for more trials to reach a conclusion about those indecisive features.\n\n\n## SHAP (SHapley Additive exPlanations)\n\n[SHAP](https://github.com/slundberg/shap) is a unified approach to explain the output of any machine learning model. SHAP connects game theory with local explanations to create the only consistent and accurate explainer. We can use it to compute feature importance for each model. More about it [here](https://christophm.github.io/interpretable-ml-book/shap.html). \n\nWe use this feature importance technique with Boruta algorithm to get a robust feature selection technique. \n\n## My Results & Conclusion\n\nI used Boruta-Shap(XGBOOST as the base model) and was really amazed by the results. \nThe dataset I use consists of 1328 features. The baseline result of my XGBOOST model was 0.7897781200376643 (using all of the features). But after running this algorithm it only accepted 116 [features](https://www.kaggle.com/datasets/susnato/shapxgboost). I trained the XGBOOST model with the accepted features and with the same hyper-parameters again and I got 0.7893407118088636!\n\n**I retained  99.9446163% of the performance by using only  8.7349398% of the features!**\n\nThe accepted features are [here](https://www.kaggle.com/datasets/susnato/shapxgboost).\n\nThe implementation and the results can be found in this notebook: [AMEX - BorutaShap Feature Selection](https://www.kaggle.com/code/susnato/amex-borutashap-feature-selection/notebook)\n\n",
      "votes": 19
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1829070": "**This is an explanation of the Boruta-Shap Feature Selection.**\n\n\n## Boruta\nIt is a pretty smart algorithm used for Feature Selection. (just like Permutation feature selection). It was developed as a package of R([here](https://www.jstatsoft.org/article/view/v036i11)). There are some simple steps required for this algorithm.\n\n<u>**Step 1(Shadow Features)**</u>: `In Boruta the Features compete with the randomized version of themselves`. For example, suppose we have a small dataset consisting of 3 features, (`age`, `height`, `weight`) with only 5 entries and we have to predict `income`.\n\n![](https://i.postimg.cc/2SdHN5fL/1-Pgw-A0-J423hc-JNpym8o-M9-w.png)\n\nNow we need to create another DataFrame by shuffling each feature. We will call it `Shadow features` and then we will concatenate this with the main DataFrame, let's call it X_baruta.\n\n![](https://i.postimg.cc/7PctgcVh/1-z-Stv-S-9-Gp-DEJJFZb0-Av-G4-Q.png)\n\n\n<u>**Step 2(Fitting the model)**</u>: Now we have to fit the model(Random Forest) on X_baruta and y. After training the model, we will take the highest feature importance recorded among the Shadow features and consider it as the threshold value. When the importance of a feature(from the main data frame) is higher than this threshold, we will call it a `hit`. **The idea is that a feature is useful only if it’s capable of doing better than the best-randomized feature**.  The result is, \n\n![](https://i.postimg.cc/V6DfGtgf/1-bt-H1-CNb6l-Q62zak-Bx-YECYg.png)\n\nHere, the threshold is 14%(shadow_height). It seems that only `age` and `height` are useful for this run.\nBut we can't trust this one run and determine the importance of the features, so we will run it multiple times. \n\n\n<u>**Step 3(Binomial Distribution)**</u>:  Result after 20 trials is, \n\n![](https://i.postimg.cc/x8cpzk1p/1-Zly3-ZYol-DDsvn8x8-Me-Ieg.png)\n\nAll features will have only two outcomes: `hit` or `not hit`, therefore we can perform the previous step several times and build a binomial distribution out of the features.\n\n![](https://i.postimg.cc/qvznH8Z5/1-yq-AUl-Mt-PUi-Fyr8g-YLFag-TA.png)\n[Image and Table Source](https://towardsdatascience.com/boruta-explained-the-way-i-wish-someone-explained-it-to-me-4489d70e154a)\n\n<u>**Step 4(Keeping/Dropping Features)**</u>: We will keep the features which are in the Green region, we will drop those features in the Red region and we haven't come to a conclusion for features in the Blue region. We can run the algorithm for more trials to reach a conclusion about those indecisive features.\n\n\n## SHAP (SHapley Additive exPlanations)\n\n[SHAP](https://github.com/slundberg/shap) is a unified approach to explain the output of any machine learning model. SHAP connects game theory with local explanations to create the only consistent and accurate explainer. We can use it to compute feature importance for each model. More about it [here](https://christophm.github.io/interpretable-ml-book/shap.html). \n\nWe use this feature importance technique with Boruta algorithm to get a robust feature selection technique. \n\n## My Results & Conclusion\n\nI used Boruta-Shap(XGBOOST as the base model) and was really amazed by the results. \nThe dataset I use consists of 1328 features. The baseline result of my XGBOOST model was 0.7897781200376643 (using all of the features). But after running this algorithm it only accepted 116 [features](https://www.kaggle.com/datasets/susnato/shapxgboost). I trained the XGBOOST model with the accepted features and with the same hyper-parameters again and I got 0.7893407118088636!\n\n**I retained  99.9446163% of the performance by using only  8.7349398% of the features!**\n\nThe accepted features are [here](https://www.kaggle.com/datasets/susnato/shapxgboost).\n\nThe implementation and the results can be found in this notebook: [AMEX - BorutaShap Feature Selection](https://www.kaggle.com/code/susnato/amex-borutashap-feature-selection/notebook)\n\n"
  }
}