{
  "id": 206434,
  "title": "Boosting - Reduce features to increase score",
  "url": "/competitions/riiid-test-answer-prediction/discussion/206434",
  "author_name": "",
  "post_date": "2020-12-24T15:52:49.987061800Z",
  "votes": 24,
  "comment_count": 33,
  "views": 0,
  "content": "<p>Hey all !</p>\n<p>Quick advice for those playing around with boosting algorithms: try to reduce the number of features. This will allows better convergence of the model as well as possibility to fit more samples in RAM for training.</p>\n<p>I am personnaly using the <strong>shap</strong> library to compute the importance of my features, it allowed me to identify reduntants features and remove them, going from ~40 to 18 features(for now). <br>\nEven better: it looks like on a dummy train set of 1 Million data, my overall test score slightly improve by 0.002 (from 0.779X to 0.781X). Not much, but worth taken.</p>",
  "messages": [
    {
      "id": "1125300",
      "postDate": "12/24/2020 15:52:49",
      "content": "<p>Hey all !</p>\n<p>Quick advice for those playing around with boosting algorithms: try to reduce the number of features. This will allows better convergence of the model as well as possibility to fit more samples in RAM for training.</p>\n<p>I am personnaly using the <strong>shap</strong> library to compute the importance of my features, it allowed me to identify reduntants features and remove them, going from ~40 to 18 features(for now). <br>\nEven better: it looks like on a dummy train set of 1 Million data, my overall test score slightly improve by 0.002 (from 0.779X to 0.781X). Not much, but worth taken.</p>",
      "rawMarkdown": "Hey all !\n\nQuick advice for those playing around with boosting algorithms: try to reduce the number of features. This will allows better convergence of the model as well as possibility to fit more samples in RAM for training.\n\nI am personnaly using the **shap** library to compute the importance of my features, it allowed me to identify reduntants features and remove them, going from ~40 to 18 features(for now). \nEven better: it looks like on a dummy train set of 1 Million data, my overall test score slightly improve by 0.002 (from 0.779X to 0.781X). Not much, but worth taken.",
      "votes": null
    },
    {
      "id": "1125329",
      "postDate": "12/24/2020 16:20:19",
      "content": "<p>good idea ,I'll try when I create all features,now I use about 34 features.</p>",
      "rawMarkdown": "good idea ,I'll try when I create all features,now I use about 34 features.",
      "votes": null
    },
    {
      "id": "1125519",
      "postDate": "12/24/2020 19:15:25",
      "content": "<p>40 to 18, that's impressive. Thank you for sharing as usual, Bowaka!<br>\nI'm facing the memory issue. I keep adding features and they give me some improvement on LB. But in the end, the model can't fit on the memory for now… </p>",
      "rawMarkdown": "40 to 18, that's impressive. Thank you for sharing as usual, Bowaka!\nI'm facing the memory issue. I keep adding features and they give me some improvement on LB. But in the end, the model can't fit on the memory for now...",
      "votes": null
    },
    {
      "id": "1125521",
      "postDate": "12/24/2020 19:23:31",
      "content": "<p>I have a followup question, beginner's one. <br>\nWhen we use LGBM, we can easily acquire the importance of features with <code>.feature_importance()</code>. <br>\nJust getting rid of less important features doesn't work? <br>\n(I'm trying and now I'm observing degradation of validation AUC. Let's see the result after a while</p>",
      "rawMarkdown": "I have a followup question, beginner's one. \nWhen we use LGBM, we can easily acquire the importance of features with `.feature_importance()`. \nJust getting rid of less important features doesn't work? \n(I'm trying and now I'm observing degradation of validation AUC. Let's see the result after a while",
      "votes": null
    },
    {
      "id": "1125524",
      "postDate": "12/24/2020 19:27:13",
      "content": "<p>A notebook I found that uses the <strong>SHAP</strong> library to explain random forest features:  <a href=\"https://github.com/dataman-git/codes_for_articles/blob/master/Explain%20your%20model%20with%20the%20SHAP%20values%20for%20article.ipynb\" target=\"_blank\">THE NOTEBOOK</a></p>",
      "rawMarkdown": "A notebook I found that uses the **SHAP** library to explain random forest features:  [THE NOTEBOOK](https://github.com/dataman-git/codes_for_articles/blob/master/Explain%20your%20model%20with%20the%20SHAP%20values%20for%20article.ipynb)",
      "votes": null
    },
    {
      "id": "1125568",
      "postDate": "12/24/2020 20:18:05",
      "content": "<p><a href=\"https://www.kaggle.com/cast42/lightgbm-model-explained-by-shap\" target=\"_blank\">Here's another example</a> of applying <code>shap</code> on LGBM. This is amazing, thank you for introducing SHAP to us, Bowaka.<br>\nI'm trying to apply <code>shap</code> on my models, it takes so much time to calculate <code>shap_values</code>, haven't got the result yet. </p>",
      "rawMarkdown": "[Here's another example](https://www.kaggle.com/cast42/lightgbm-model-explained-by-shap) of applying `shap` on LGBM. This is amazing, thank you for introducing SHAP to us, Bowaka.\nI'm trying to apply `shap` on my models, it takes so much time to calculate `shap_values`, haven't got the result yet.",
      "votes": null
    },
    {
      "id": "1125570",
      "postDate": "12/24/2020 20:24:17",
      "content": "<p>Feature importances are not always reliable in that respect. For example high cardinality features like User_id will always have a high LGB feature importance but that does not necesarily mean that its good to use it since the model might just overfit this feature. :)</p>",
      "rawMarkdown": "Feature importances are not always reliable in that respect. For example high cardinality features like User_id will always have a high LGB feature importance but that does not necesarily mean that its good to use it since the model might just overfit this feature. :)",
      "votes": null
    },
    {
      "id": "1125579",
      "postDate": "12/24/2020 20:36:32",
      "content": "<p>Thank you for showing the convincing example, hrunic! <br>\nSo we should assess the model from various aspects and SHAP helps it. </p>",
      "rawMarkdown": "Thank you for showing the convincing example, hrunic! \nSo we should assess the model from various aspects and SHAP helps it.",
      "votes": null
    },
    {
      "id": "1125583",
      "postDate": "12/24/2020 20:42:29",
      "content": "<p>Great post. But, please remember reducing features sometimes degrades the performance after ensemble, even if single model performance improves. That is why I don't stick to single model performance.</p>",
      "rawMarkdown": "Great post. But, please remember reducing features sometimes degrades the performance after ensemble, even if single model performance improves. That is why I don't stick to single model performance.",
      "votes": null
    },
    {
      "id": "1125609",
      "postDate": "12/24/2020 21:16:02",
      "content": "<p>But many features would stop us from using more data to train! Is there a way around it ?</p>",
      "rawMarkdown": "But many features would stop us from using more data to train! Is there a way around it ?",
      "votes": null
    },
    {
      "id": "1125613",
      "postDate": "12/24/2020 21:18:23",
      "content": "<p>No. To gain something is always to lose something.<br>\nBut, I think column-subsampling and row-subsampling is a possible option to reduce training time when using LightGBM with many features. </p>",
      "rawMarkdown": "No. To gain something is always to lose something.\nBut, I think column-subsampling and row-subsampling is a possible option to reduce training time when using LightGBM with many features.",
      "votes": null
    },
    {
      "id": "1125636",
      "postDate": "12/24/2020 22:31:20",
      "content": "<p>sometimes drop less importance features get LB score down…</p>",
      "rawMarkdown": "sometimes drop less importance features get LB score down...",
      "votes": null
    },
    {
      "id": "1125730",
      "postDate": "12/25/2020 02:46:59",
      "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>: Agreed. can you please guide us on the possible range value for subsampling?</p>",
      "rawMarkdown": "mamasinkgs: Agreed. can you please guide us on the possible range value for subsampling?",
      "votes": null
    },
    {
      "id": "1125751",
      "postDate": "12/25/2020 03:10:38",
      "content": "<p>It's not true, more features just mean you need to rent an AWS EC2 with bigger memory. Currently, the memory I'm using is 256 Gi which is sufficient for about 100 features. And the training on the whole dataset takes 10 hours, so each submission cost me $10 …</p>",
      "rawMarkdown": "It's not true, more features just mean you need to rent an AWS EC2 with bigger memory. Currently, the memory I'm using is 256 Gi which is sufficient for about 100 features. And the training on the whole dataset takes 10 hours, so each submission cost me $10 ...",
      "votes": null
    },
    {
      "id": "1125754",
      "postDate": "12/25/2020 03:12:16",
      "content": "<blockquote>\n  <p>That is why I don't stick to single model performance.</p>\n</blockquote>\n<p>Do you mean you're using multiple models for bagging now?</p>",
      "rawMarkdown": ">That is why I don't stick to single model performance.\n\nDo you mean you're using multiple models for bagging now?",
      "votes": null
    },
    {
      "id": "1125762",
      "postDate": "12/25/2020 03:25:14",
      "content": "<p>It's very hard to say, because it does depends on the features you use. Please understand my answer is not specific to this competition, but a general guide from my GBDT experience. <br>\nI think setting bagging_fraction &lt; 0.4 often degrades the score. So, the possible range is [0.4, 1]. But if your machine is not strong and the number of cores are small (and especially when <code>force_row_wise</code> = True), I think it can't be avoided to set bagging_fraction &lt; 0.4.<br>\nFor the feature fraction, it's harder to say. In some competitions, I heard setting very small feature_fraction (e.g. &lt; 0.1) worked well. <br>\nThe typical setting is bagging_fracition = feature_fraction = 0.7, as Laurae's documentation <a href=\"https://sites.google.com/view/lauraepp/parameters\" target=\"_blank\">https://sites.google.com/view/lauraepp/parameters</a> says.</p>",
      "rawMarkdown": "It's very hard to say, because it does depends on the features you use. Please understand my answer is not specific to this competition, but a general guide from my GBDT experience. \nI think setting bagging_fraction < 0.4 often degrades the score. So, the possible range is [0.4, 1]. But if your machine is not strong and the number of cores are small (and especially when `force_row_wise` = True), I think it can't be avoided to set bagging_fraction < 0.4.\nFor the feature fraction, it's harder to say. In some competitions, I heard setting very small feature_fraction (e.g. < 0.1) worked well. \nThe typical setting is bagging_fracition = feature_fraction = 0.7, as Laurae's documentation https://sites.google.com/view/lauraepp/parameters says.",
      "votes": null
    },
    {
      "id": "1125783",
      "postDate": "12/25/2020 04:16:26",
      "content": "<p>It depends on the kind of features you have. If you have too many features that try to convey almost the same thing(can be thought of similar to multicollinearity), your model performance can be improved by feature selection. But if you have features, each of which give a different idea about the data to the model, removing them will only hurt your score. Also I would point out that constrained by the memory and huge data , most of us have very few features, less than 40 or even less 30 in many cases. In such cases, the chances of improvement while dropping a feature will be quite low, if you have engineered your features carefully.</p>\n<p>Feature selection is a good step, when you have more than 100s of features, or some extremely noisy/ leaking feature. For this competition I will suggest people to focus on feature engineering rather than feature selection. One good feature will give you more gain, than 5 noisy features removed.</p>",
      "rawMarkdown": "It depends on the kind of features you have. If you have too many features that try to convey almost the same thing(can be thought of similar to multicollinearity), your model performance can be improved by feature selection. But if you have features, each of which give a different idea about the data to the model, removing them will only hurt your score. Also I would point out that constrained by the memory and huge data , most of us have very few features, less than 40 or even less 30 in many cases. In such cases, the chances of improvement while dropping a feature will be quite low, if you have engineered your features carefully.\n\nFeature selection is a good step, when you have more than 100s of features, or some extremely noisy/ leaking feature. For this competition I will suggest people to focus on feature engineering rather than feature selection. One good feature will give you more gain, than 5 noisy features removed.",
      "votes": null
    },
    {
      "id": "1125919",
      "postDate": "12/25/2020 07:13:42",
      "content": "<p>Nice work Bowaka</p>",
      "rawMarkdown": "Nice work Bowaka",
      "votes": null
    },
    {
      "id": "1126068",
      "postDate": "12/25/2020 10:05:23",
      "content": "<p>how long did you calculate shap values ?I did't get a result after  8 hours ….</p>",
      "rawMarkdown": "how long did you calculate shap values ?I did't get a result after  8 hours ....",
      "votes": null
    },
    {
      "id": "1126090",
      "postDate": "12/25/2020 10:25:34",
      "content": "<p>Hi qiaqia, personnaly I used a subsample of 200 000 test rows to check feature importances, it takes ~10-15min with my vanilla lgb model</p>",
      "rawMarkdown": "Hi qiaqia, personnaly I used a subsample of 200 000 test rows to check feature importances, it takes ~10-15min with my vanilla lgb model",
      "votes": null
    },
    {
      "id": "1126095",
      "postDate": "12/25/2020 10:29:06",
      "content": "<p>Yes totally agree to that ! </p>\n<p>In my case, I have several features that were probably giving same information or nearly same information, and in that case reducing dimension is helping me a lot. I am probably missing some key features yet, but reducing them allows me to make test faster as I reduce also computationnal time.</p>",
      "rawMarkdown": "Yes totally agree to that ! \n\nIn my case, I have several features that were probably giving same information or nearly same information, and in that case reducing dimension is helping me a lot. I am probably missing some key features yet, but reducing them allows me to make test faster as I reduce also computationnal time.",
      "votes": null
    },
    {
      "id": "1126099",
      "postDate": "12/25/2020 10:31:18",
      "content": "<p>To have a quick overview, you can reduce the number of samples on which you calculate the shap values. I personnaly use a subsample of 200 000 rows to start checking… </p>",
      "rawMarkdown": "To have a quick overview, you can reduce the number of samples on which you calculate the shap values. I personnaly use a subsample of 200 000 rows to start checking...",
      "votes": null
    },
    {
      "id": "1126105",
      "postDate": "12/25/2020 10:40:14",
      "content": "<p>thanks, I used 2M dataset and 41 features ,reduce dataset may be useful.</p>\n<p>but train 26M and 41 features use 7.1G RAM,I'd like to create more features  before using shap library </p>",
      "rawMarkdown": "thanks, I used 2M dataset and 41 features ,reduce dataset may be useful.\n\nbut train 26M and 41 features use 7.1G RAM,I'd like to create more features  before using shap library",
      "votes": null
    },
    {
      "id": "1126106",
      "postDate": "12/25/2020 10:40:57",
      "content": "<p>I don't remember how is calculated feature importance for boosting models, but I would not rely much on it… On the other hand, the shap values are calculated based on game theory, calculating the contribution of each feature to the global prediction, so it shall give a much better overview of the feature importance.</p>\n<p>In general, the shap values are very hard to calculate due to the fact that it is suppose to try all different combination of features to calculate the different contributions, but in the case of tree-based models (like boosting), methods allow much faster calculations… BUT I don't manage to find back the original paper now ☹️</p>",
      "rawMarkdown": "I don't remember how is calculated feature importance for boosting models, but I would not rely much on it... On the other hand, the shap values are calculated based on game theory, calculating the contribution of each feature to the global prediction, so it shall give a much better overview of the feature importance.\n\nIn general, the shap values are very hard to calculate due to the fact that it is suppose to try all different combination of features to calculate the different contributions, but in the case of tree-based models (like boosting), methods allow much faster calculations... BUT I don't manage to find back the original paper now ☹️",
      "votes": null
    },
    {
      "id": "1126109",
      "postDate": "12/25/2020 10:48:17",
      "content": "<p>From local explanations to global understanding with explainable AI for trees<br>\nA Unified Approach to Interpreting Model Predictions<br>\ntwo shap original paper  I <a href=\"https://github.com/slundberg/shap\" target=\"_blank\">find </a></p>",
      "rawMarkdown": "From local explanations to global understanding with explainable AI for trees\nA Unified Approach to Interpreting Model Predictions\ntwo shap original paper  I [find ](https://github.com/slundberg/shap)",
      "votes": null
    },
    {
      "id": "1127090",
      "postDate": "12/26/2020 08:23:09",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a> </p>",
      "rawMarkdown": "Thanks for sharing @bowaka",
      "votes": null
    },
    {
      "id": "1128767",
      "postDate": "12/27/2020 18:02:29",
      "content": "<p>Nice work <a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a>, could you please share a bit more about how did you find redundant features based on Sharp? Currently, I'm using 85 features, and I achieved a little improvement by removing several features. Very interested in removing features in a more data-driven way:)</p>",
      "rawMarkdown": "Nice work @bowaka, could you please share a bit more about how did you find redundant features based on Sharp? Currently, I'm using 85 features, and I achieved a little improvement by removing several features. Very interested in removing features in a more data-driven way:)",
      "votes": null
    },
    {
      "id": "1128986",
      "postDate": "12/28/2020 00:52:01",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": null
    },
    {
      "id": "1128998",
      "postDate": "12/28/2020 01:23:24",
      "content": "<p>Thanks Bowaka for sharing. I just wanted to know, is there any thumb rule or a way to compute that given x numbers of rows and y number of features and additional analysis like Pearson correlation , we could remove a certain number of features.</p>",
      "rawMarkdown": "Thanks Bowaka for sharing. I just wanted to know, is there any thumb rule or a way to compute that given x numbers of rows and y number of features and additional analysis like Pearson correlation , we could remove a certain number of features.",
      "votes": null
    },
    {
      "id": "1129505",
      "postDate": "12/28/2020 11:49:39",
      "content": "<p>After ranking the SHAP value of features, how can I know the specific bound under which  features can be reduced?</p>",
      "rawMarkdown": "After ranking the SHAP value of features, how can I know the specific bound under which  features can be reduced?",
      "votes": null
    },
    {
      "id": "1129544",
      "postDate": "12/28/2020 12:26:29",
      "content": "<p>I personnaly remove one after the other until my CV starts to drop 🙂</p>",
      "rawMarkdown": "I personnaly remove one after the other until my CV starts to drop 🙂",
      "votes": null
    },
    {
      "id": "1129550",
      "postDate": "12/28/2020 12:30:06",
      "content": "<p>I personally compute the SHAP values on a small number of rows from my test set (100 000), and then try to remove the columns one by one on my main feature matrix by order of lower importance. When I see my CV scheme drop, I stop the process. </p>",
      "rawMarkdown": "I personally compute the SHAP values on a small number of rows from my test set (100 000), and then try to remove the columns one by one on my main feature matrix by order of lower importance. When I see my CV scheme drop, I stop the process.",
      "votes": null
    },
    {
      "id": "1129877",
      "postDate": "12/28/2020 15:57:59",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a>. If reducing features one by one, permutation importance maybe a better choice, calculating permutation importance of features is much faster than calculating SHAP values.</p>",
      "rawMarkdown": "Thanks, @bowaka. If reducing features one by one, permutation importance maybe a better choice, calculating permutation importance of features is much faster than calculating SHAP values.",
      "votes": null
    },
    {
      "id": "1130372",
      "postDate": "12/29/2020 01:14:37",
      "content": "<p><a href=\"https://www.kaggle.com/wuwenmin\" target=\"_blank\">@wuwenmin</a> <br>\nI'd like to know practical superiority of permutation importance other than other selection methods (original feature importance implemented by lgmb and null importance..). Shuffling each columns are intuitive to me, but not so many notebooks uses this method (I guess that's because perm imp takes lots of time than original implemented method)</p>",
      "rawMarkdown": "wuwenmin \nI'd like to know practical superiority of permutation importance other than other selection methods (original feature importance implemented by lgmb and null importance..). Shuffling each columns are intuitive to me, but not so many notebooks uses this method (I guess that's because perm imp takes lots of time than original implemented method)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1125329,
      "author_name": "yangxiaoshuai",
      "author_url": "",
      "post_date": "12/24/2020 16:20:19",
      "content": "<p>good idea ,I'll try when I create all features,now I use about 34 features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1125519,
      "author_name": "kokitanisaka",
      "author_url": "",
      "post_date": "12/24/2020 19:15:25",
      "content": "<p>40 to 18, that's impressive. Thank you for sharing as usual, Bowaka!<br>\nI'm facing the memory issue. I keep adding features and they give me some improvement on LB. But in the end, the model can't fit on the memory for now… </p>",
      "votes": null,
      "replies": [
        {
          "id": 1125521,
          "author_name": "kokitanisaka",
          "author_url": "",
          "post_date": "12/24/2020 19:23:31",
          "content": "<p>I have a followup question, beginner's one. <br>\nWhen we use LGBM, we can easily acquire the importance of features with <code>.feature_importance()</code>. <br>\nJust getting rid of less important features doesn't work? <br>\n(I'm trying and now I'm observing degradation of validation AUC. Let's see the result after a while</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125570,
          "author_name": "nicohrubec",
          "author_url": "",
          "post_date": "12/24/2020 20:24:17",
          "content": "<p>Feature importances are not always reliable in that respect. For example high cardinality features like User_id will always have a high LGB feature importance but that does not necesarily mean that its good to use it since the model might just overfit this feature. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125579,
          "author_name": "kokitanisaka",
          "author_url": "",
          "post_date": "12/24/2020 20:36:32",
          "content": "<p>Thank you for showing the convincing example, hrunic! <br>\nSo we should assess the model from various aspects and SHAP helps it. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125636,
          "author_name": "yangxiaoshuai",
          "author_url": "",
          "post_date": "12/24/2020 22:31:20",
          "content": "<p>sometimes drop less importance features get LB score down…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1126106,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "12/25/2020 10:40:57",
          "content": "<p>I don't remember how is calculated feature importance for boosting models, but I would not rely much on it… On the other hand, the shap values are calculated based on game theory, calculating the contribution of each feature to the global prediction, so it shall give a much better overview of the feature importance.</p>\n<p>In general, the shap values are very hard to calculate due to the fact that it is suppose to try all different combination of features to calculate the different contributions, but in the case of tree-based models (like boosting), methods allow much faster calculations… BUT I don't manage to find back the original paper now ☹️</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1126109,
          "author_name": "yangxiaoshuai",
          "author_url": "",
          "post_date": "12/25/2020 10:48:17",
          "content": "<p>From local explanations to global understanding with explainable AI for trees<br>\nA Unified Approach to Interpreting Model Predictions<br>\ntwo shap original paper  I <a href=\"https://github.com/slundberg/shap\" target=\"_blank\">find </a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1125524,
      "author_name": "abdessalemboukil",
      "author_url": "",
      "post_date": "12/24/2020 19:27:13",
      "content": "<p>A notebook I found that uses the <strong>SHAP</strong> library to explain random forest features:  <a href=\"https://github.com/dataman-git/codes_for_articles/blob/master/Explain%20your%20model%20with%20the%20SHAP%20values%20for%20article.ipynb\" target=\"_blank\">THE NOTEBOOK</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1125568,
          "author_name": "kokitanisaka",
          "author_url": "",
          "post_date": "12/24/2020 20:18:05",
          "content": "<p><a href=\"https://www.kaggle.com/cast42/lightgbm-model-explained-by-shap\" target=\"_blank\">Here's another example</a> of applying <code>shap</code> on LGBM. This is amazing, thank you for introducing SHAP to us, Bowaka.<br>\nI'm trying to apply <code>shap</code> on my models, it takes so much time to calculate <code>shap_values</code>, haven't got the result yet. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1126099,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "12/25/2020 10:31:18",
          "content": "<p>To have a quick overview, you can reduce the number of samples on which you calculate the shap values. I personnaly use a subsample of 200 000 rows to start checking… </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1125583,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "12/24/2020 20:42:29",
      "content": "<p>Great post. But, please remember reducing features sometimes degrades the performance after ensemble, even if single model performance improves. That is why I don't stick to single model performance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1125609,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "12/24/2020 21:16:02",
          "content": "<p>But many features would stop us from using more data to train! Is there a way around it ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125613,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "12/24/2020 21:18:23",
          "content": "<p>No. To gain something is always to lose something.<br>\nBut, I think column-subsampling and row-subsampling is a possible option to reduce training time when using LightGBM with many features. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125730,
          "author_name": "projdev",
          "author_url": "",
          "post_date": "12/25/2020 02:46:59",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>: Agreed. can you please guide us on the possible range value for subsampling?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125751,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "12/25/2020 03:10:38",
          "content": "<p>It's not true, more features just mean you need to rent an AWS EC2 with bigger memory. Currently, the memory I'm using is 256 Gi which is sufficient for about 100 features. And the training on the whole dataset takes 10 hours, so each submission cost me $10 …</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125754,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "12/25/2020 03:12:16",
          "content": "<blockquote>\n  <p>That is why I don't stick to single model performance.</p>\n</blockquote>\n<p>Do you mean you're using multiple models for bagging now?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125762,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "12/25/2020 03:25:14",
          "content": "<p>It's very hard to say, because it does depends on the features you use. Please understand my answer is not specific to this competition, but a general guide from my GBDT experience. <br>\nI think setting bagging_fraction &lt; 0.4 often degrades the score. So, the possible range is [0.4, 1]. But if your machine is not strong and the number of cores are small (and especially when <code>force_row_wise</code> = True), I think it can't be avoided to set bagging_fraction &lt; 0.4.<br>\nFor the feature fraction, it's harder to say. In some competitions, I heard setting very small feature_fraction (e.g. &lt; 0.1) worked well. <br>\nThe typical setting is bagging_fracition = feature_fraction = 0.7, as Laurae's documentation <a href=\"https://sites.google.com/view/lauraepp/parameters\" target=\"_blank\">https://sites.google.com/view/lauraepp/parameters</a> says.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1125783,
      "author_name": "nikhilmishradev",
      "author_url": "",
      "post_date": "12/25/2020 04:16:26",
      "content": "<p>It depends on the kind of features you have. If you have too many features that try to convey almost the same thing(can be thought of similar to multicollinearity), your model performance can be improved by feature selection. But if you have features, each of which give a different idea about the data to the model, removing them will only hurt your score. Also I would point out that constrained by the memory and huge data , most of us have very few features, less than 40 or even less 30 in many cases. In such cases, the chances of improvement while dropping a feature will be quite low, if you have engineered your features carefully.</p>\n<p>Feature selection is a good step, when you have more than 100s of features, or some extremely noisy/ leaking feature. For this competition I will suggest people to focus on feature engineering rather than feature selection. One good feature will give you more gain, than 5 noisy features removed.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1126095,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "12/25/2020 10:29:06",
          "content": "<p>Yes totally agree to that ! </p>\n<p>In my case, I have several features that were probably giving same information or nearly same information, and in that case reducing dimension is helping me a lot. I am probably missing some key features yet, but reducing them allows me to make test faster as I reduce also computationnal time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1125919,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "12/25/2020 07:13:42",
      "content": "<p>Nice work Bowaka</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1126068,
      "author_name": "yangxiaoshuai",
      "author_url": "",
      "post_date": "12/25/2020 10:05:23",
      "content": "<p>how long did you calculate shap values ?I did't get a result after  8 hours ….</p>",
      "votes": null,
      "replies": [
        {
          "id": 1126090,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "12/25/2020 10:25:34",
          "content": "<p>Hi qiaqia, personnaly I used a subsample of 200 000 test rows to check feature importances, it takes ~10-15min with my vanilla lgb model</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1126105,
          "author_name": "yangxiaoshuai",
          "author_url": "",
          "post_date": "12/25/2020 10:40:14",
          "content": "<p>thanks, I used 2M dataset and 41 features ,reduce dataset may be useful.</p>\n<p>but train 26M and 41 features use 7.1G RAM,I'd like to create more features  before using shap library </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1127090,
      "author_name": "saurabhshahane",
      "author_url": "",
      "post_date": "12/26/2020 08:23:09",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1128767,
      "author_name": "wuwenmin",
      "author_url": "",
      "post_date": "12/27/2020 18:02:29",
      "content": "<p>Nice work <a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a>, could you please share a bit more about how did you find redundant features based on Sharp? Currently, I'm using 85 features, and I achieved a little improvement by removing several features. Very interested in removing features in a more data-driven way:)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1129550,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "12/28/2020 12:30:06",
          "content": "<p>I personally compute the SHAP values on a small number of rows from my test set (100 000), and then try to remove the columns one by one on my main feature matrix by order of lower importance. When I see my CV scheme drop, I stop the process. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1129877,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "12/28/2020 15:57:59",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a>. If reducing features one by one, permutation importance maybe a better choice, calculating permutation importance of features is much faster than calculating SHAP values.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1130372,
          "author_name": "ant3ng",
          "author_url": "",
          "post_date": "12/29/2020 01:14:37",
          "content": "<p><a href=\"https://www.kaggle.com/wuwenmin\" target=\"_blank\">@wuwenmin</a> <br>\nI'd like to know practical superiority of permutation importance other than other selection methods (original feature importance implemented by lgmb and null importance..). Shuffling each columns are intuitive to me, but not so many notebooks uses this method (I guess that's because perm imp takes lots of time than original implemented method)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1128986,
      "author_name": "lokeshkum",
      "author_url": "",
      "post_date": "12/28/2020 00:52:01",
      "content": "<p>Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1128998,
      "author_name": "shrirangkulkarni",
      "author_url": "",
      "post_date": "12/28/2020 01:23:24",
      "content": "<p>Thanks Bowaka for sharing. I just wanted to know, is there any thumb rule or a way to compute that given x numbers of rows and y number of features and additional analysis like Pearson correlation , we could remove a certain number of features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1129505,
      "author_name": "xinlearning",
      "author_url": "",
      "post_date": "12/28/2020 11:49:39",
      "content": "<p>After ranking the SHAP value of features, how can I know the specific bound under which  features can be reduced?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1129544,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "12/28/2020 12:26:29",
          "content": "<p>I personnaly remove one after the other until my CV starts to drop 🙂</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1125300": "Hey all !\n\nQuick advice for those playing around with boosting algorithms: try to reduce the number of features. This will allows better convergence of the model as well as possibility to fit more samples in RAM for training.\n\nI am personnaly using the **shap** library to compute the importance of my features, it allowed me to identify reduntants features and remove them, going from ~40 to 18 features(for now). \nEven better: it looks like on a dummy train set of 1 Million data, my overall test score slightly improve by 0.002 (from 0.779X to 0.781X). Not much, but worth taken.",
    "1125329": "good idea ,I'll try when I create all features,now I use about 34 features.",
    "1125519": "40 to 18, that's impressive. Thank you for sharing as usual, Bowaka!\nI'm facing the memory issue. I keep adding features and they give me some improvement on LB. But in the end, the model can't fit on the memory for now...",
    "1125521": "I have a followup question, beginner's one. \nWhen we use LGBM, we can easily acquire the importance of features with `.feature_importance()`. \nJust getting rid of less important features doesn't work? \n(I'm trying and now I'm observing degradation of validation AUC. Let's see the result after a while",
    "1125524": "A notebook I found that uses the **SHAP** library to explain random forest features:  [THE NOTEBOOK](https://github.com/dataman-git/codes_for_articles/blob/master/Explain%20your%20model%20with%20the%20SHAP%20values%20for%20article.ipynb)",
    "1125568": "[Here's another example](https://www.kaggle.com/cast42/lightgbm-model-explained-by-shap) of applying `shap` on LGBM. This is amazing, thank you for introducing SHAP to us, Bowaka.\nI'm trying to apply `shap` on my models, it takes so much time to calculate `shap_values`, haven't got the result yet.",
    "1125570": "Feature importances are not always reliable in that respect. For example high cardinality features like User_id will always have a high LGB feature importance but that does not necesarily mean that its good to use it since the model might just overfit this feature. :)",
    "1125579": "Thank you for showing the convincing example, hrunic! \nSo we should assess the model from various aspects and SHAP helps it.",
    "1125583": "Great post. But, please remember reducing features sometimes degrades the performance after ensemble, even if single model performance improves. That is why I don't stick to single model performance.",
    "1125609": "But many features would stop us from using more data to train! Is there a way around it ?",
    "1125613": "No. To gain something is always to lose something.\nBut, I think column-subsampling and row-subsampling is a possible option to reduce training time when using LightGBM with many features.",
    "1125636": "sometimes drop less importance features get LB score down...",
    "1125730": "mamasinkgs: Agreed. can you please guide us on the possible range value for subsampling?",
    "1125751": "It's not true, more features just mean you need to rent an AWS EC2 with bigger memory. Currently, the memory I'm using is 256 Gi which is sufficient for about 100 features. And the training on the whole dataset takes 10 hours, so each submission cost me $10 ...",
    "1125754": ">That is why I don't stick to single model performance.\n\nDo you mean you're using multiple models for bagging now?",
    "1125762": "It's very hard to say, because it does depends on the features you use. Please understand my answer is not specific to this competition, but a general guide from my GBDT experience. \nI think setting bagging_fraction < 0.4 often degrades the score. So, the possible range is [0.4, 1]. But if your machine is not strong and the number of cores are small (and especially when `force_row_wise` = True), I think it can't be avoided to set bagging_fraction < 0.4.\nFor the feature fraction, it's harder to say. In some competitions, I heard setting very small feature_fraction (e.g. < 0.1) worked well. \nThe typical setting is bagging_fracition = feature_fraction = 0.7, as Laurae's documentation https://sites.google.com/view/lauraepp/parameters says.",
    "1125783": "It depends on the kind of features you have. If you have too many features that try to convey almost the same thing(can be thought of similar to multicollinearity), your model performance can be improved by feature selection. But if you have features, each of which give a different idea about the data to the model, removing them will only hurt your score. Also I would point out that constrained by the memory and huge data , most of us have very few features, less than 40 or even less 30 in many cases. In such cases, the chances of improvement while dropping a feature will be quite low, if you have engineered your features carefully.\n\nFeature selection is a good step, when you have more than 100s of features, or some extremely noisy/ leaking feature. For this competition I will suggest people to focus on feature engineering rather than feature selection. One good feature will give you more gain, than 5 noisy features removed.",
    "1125919": "Nice work Bowaka",
    "1126068": "how long did you calculate shap values ?I did't get a result after  8 hours ....",
    "1126090": "Hi qiaqia, personnaly I used a subsample of 200 000 test rows to check feature importances, it takes ~10-15min with my vanilla lgb model",
    "1126095": "Yes totally agree to that ! \n\nIn my case, I have several features that were probably giving same information or nearly same information, and in that case reducing dimension is helping me a lot. I am probably missing some key features yet, but reducing them allows me to make test faster as I reduce also computationnal time.",
    "1126099": "To have a quick overview, you can reduce the number of samples on which you calculate the shap values. I personnaly use a subsample of 200 000 rows to start checking...",
    "1126105": "thanks, I used 2M dataset and 41 features ,reduce dataset may be useful.\n\nbut train 26M and 41 features use 7.1G RAM,I'd like to create more features  before using shap library",
    "1126106": "I don't remember how is calculated feature importance for boosting models, but I would not rely much on it... On the other hand, the shap values are calculated based on game theory, calculating the contribution of each feature to the global prediction, so it shall give a much better overview of the feature importance.\n\nIn general, the shap values are very hard to calculate due to the fact that it is suppose to try all different combination of features to calculate the different contributions, but in the case of tree-based models (like boosting), methods allow much faster calculations... BUT I don't manage to find back the original paper now ☹️",
    "1126109": "From local explanations to global understanding with explainable AI for trees\nA Unified Approach to Interpreting Model Predictions\ntwo shap original paper  I [find ](https://github.com/slundberg/shap)",
    "1127090": "Thanks for sharing @bowaka",
    "1128767": "Nice work @bowaka, could you please share a bit more about how did you find redundant features based on Sharp? Currently, I'm using 85 features, and I achieved a little improvement by removing several features. Very interested in removing features in a more data-driven way:)",
    "1128986": "Thanks for sharing.",
    "1128998": "Thanks Bowaka for sharing. I just wanted to know, is there any thumb rule or a way to compute that given x numbers of rows and y number of features and additional analysis like Pearson correlation , we could remove a certain number of features.",
    "1129505": "After ranking the SHAP value of features, how can I know the specific bound under which  features can be reduced?",
    "1129544": "I personnaly remove one after the other until my CV starts to drop 🙂",
    "1129550": "I personally compute the SHAP values on a small number of rows from my test set (100 000), and then try to remove the columns one by one on my main feature matrix by order of lower importance. When I see my CV scheme drop, I stop the process.",
    "1129877": "Thanks, @bowaka. If reducing features one by one, permutation importance maybe a better choice, calculating permutation importance of features is much faster than calculating SHAP values.",
    "1130372": "wuwenmin \nI'd like to know practical superiority of permutation importance other than other selection methods (original feature importance implemented by lgmb and null importance..). Shuffling each columns are intuitive to me, but not so many notebooks uses this method (I guess that's because perm imp takes lots of time than original implemented method)"
  },
  "source": "meta"
}