{
  "id": 339071,
  "title": "Feature Selection approach : Base->Blowup->Base",
  "url": "/competitions/amex-default-prediction/discussion/339071",
  "author_name": "Noir.s",
  "post_date": "2022-07-23T06:09:35.533000",
  "votes": 17,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Simply running group by aggregations (min, max, mean, last etc.) on all available features quickly blows up the feature set and we soon run into memory issues and features being reused redundantly within the model. Have not had as much time as I would have liked and all I have built so far are XGB models.</p>\n<p>I've been using a fairly simplistic feature selection approach, which I don't have really have a name for, so I made up the title. It's basically some sort of a manual stepwise forward selection. But so far it has been effective in improving CV score while minimizing the number of features being used in my modeling dataset. </p>\n<hr>\n<p>BASE (469 features)</p>\n<hr>\n<p>CV 0.7937</p>\n<p>Started with a minimum set of features which I know has decent performance. I used the ones from this notebook (<a href=\"https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart</a>) which has 469 features</p>\n<p>Fit XGB model, removed least important features (split importance), made sure model performance did not degrade a whole lot</p>\n<hr>\n<p>BLOWUP (1486 features)</p>\n<hr>\n<p>CV 0.7939</p>\n<p>Then built a larger feature set incorporating different feature engineering ideas from the competition, blowing up the feature set to ~1500 features (<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/336557\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/336557</a>)</p>\n<p>Fit model, </p>\n<hr>\n<p>BASE (520 features)</p>\n<hr>\n<p>CV 0.7947</p>\n<p>Then brought in only the most important features to my base feature set. Creating my \"new\" base feature set</p>\n<hr>\n<p>REPEAT</p>\n<hr>\n<p>And so on….. Build a larger feature set with more feature engineering ideas. Bring back the most important features to the base dataset </p>\n<p>The disadvantage is that, the features being dropped do not get to interact with the new features in the larger dataset at each stage. Which could mean there is a lot of value being lost with this approach</p>\n<p>Look forward to any feedback on this and any new ideas you can share on how to go about feature selection. </p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": 1867308,
      "postDate": "2022-07-23T06:09:35.533Z",
      "content": "<p>Simply running group by aggregations (min, max, mean, last etc.) on all available features quickly blows up the feature set and we soon run into memory issues and features being reused redundantly within the model. Have not had as much time as I would have liked and all I have built so far are XGB models.</p>\n<p>I've been using a fairly simplistic feature selection approach, which I don't have really have a name for, so I made up the title. It's basically some sort of a manual stepwise forward selection. But so far it has been effective in improving CV score while minimizing the number of features being used in my modeling dataset. </p>\n<hr>\n<p>BASE (469 features)</p>\n<hr>\n<p>CV 0.7937</p>\n<p>Started with a minimum set of features which I know has decent performance. I used the ones from this notebook (<a href=\"https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart</a>) which has 469 features</p>\n<p>Fit XGB model, removed least important features (split importance), made sure model performance did not degrade a whole lot</p>\n<hr>\n<p>BLOWUP (1486 features)</p>\n<hr>\n<p>CV 0.7939</p>\n<p>Then built a larger feature set incorporating different feature engineering ideas from the competition, blowing up the feature set to ~1500 features (<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/336557\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/336557</a>)</p>\n<p>Fit model, </p>\n<hr>\n<p>BASE (520 features)</p>\n<hr>\n<p>CV 0.7947</p>\n<p>Then brought in only the most important features to my base feature set. Creating my \"new\" base feature set</p>\n<hr>\n<p>REPEAT</p>\n<hr>\n<p>And so on….. Build a larger feature set with more feature engineering ideas. Bring back the most important features to the base dataset </p>\n<p>The disadvantage is that, the features being dropped do not get to interact with the new features in the larger dataset at each stage. Which could mean there is a lot of value being lost with this approach</p>\n<p>Look forward to any feedback on this and any new ideas you can share on how to go about feature selection. </p>\n<p>Thanks!</p>",
      "rawMarkdown": "Simply running group by aggregations (min, max, mean, last etc.) on all available features quickly blows up the feature set and we soon run into memory issues and features being reused redundantly within the model. Have not had as much time as I would have liked and all I have built so far are XGB models.\n\nI've been using a fairly simplistic feature selection approach, which I don't have really have a name for, so I made up the title. It's basically some sort of a manual stepwise forward selection. But so far it has been effective in improving CV score while minimizing the number of features being used in my modeling dataset. \n\n_______________________________________________\nBASE (469 features)\n_______________________________________________\nCV 0.7937\n\nStarted with a minimum set of features which I know has decent performance. I used the ones from this notebook (https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart) which has 469 features\n\nFit XGB model, removed least important features (split importance), made sure model performance did not degrade a whole lot\n\n_______________________________________________\nBLOWUP (1486 features)\n_______________________________________________\nCV 0.7939\n\nThen built a larger feature set incorporating different feature engineering ideas from the competition, blowing up the feature set to ~1500 features (https://www.kaggle.com/competitions/amex-default-prediction/discussion/336557)\n\nFit model, \n\n_______________________________________________\nBASE (520 features)\n_______________________________________________\nCV 0.7947\n\nThen brought in only the most important features to my base feature set. Creating my \"new\" base feature set\n\n_______________________________________________\nREPEAT\n_______________________________________________\n\nAnd so on..... Build a larger feature set with more feature engineering ideas. Bring back the most important features to the base dataset \n\nThe disadvantage is that, the features being dropped do not get to interact with the new features in the larger dataset at each stage. Which could mean there is a lot of value being lost with this approach\n\nLook forward to any feedback on this and any new ideas you can share on how to go about feature selection. \n\nThanks!\n\n",
      "votes": 15
    },
    {
      "id": 1875334,
      "postDate": "2022-07-28T23:40:28.630Z",
      "content": "<p>Thanks for sharing!</p>\n<p>It sounds similar to my idea. I'm a bit paranoid about losing any true signal though, so my idea - not yet even tried - is to always keep the best for every base feature. Or even best 2.</p>\n<p>For example, try weighted average, mean, median: keep the best one on an individual column basis. (And maybe keep more if above some threshold, like top 300 or whatever.) Try min and max: keep the higher importance one. Try last minus first, last minus mean, last minus prior, the same but divide instead of subtract, again just keep the best 1 per base feature.</p>\n<p>Another more convoluted idea, the idea that less correlated models ensemble better: let's say we think up around 50 feature sets total to apply to our ~180 numeric columns. 9000 total. We do some method to get scores for everything. Then instead of taking the best scores and dropping the rest, perhaps we simply allow the higher scoring features to be included in more individual models? top 180 go in every model, next best 180*3 go in 1/3rd of the models, and the rest only go in one model each. 8 models of around 1400 features each. I'm doubtful but might try it and see. :)</p>",
      "rawMarkdown": "Thanks for sharing!\n\nIt sounds similar to my idea. I'm a bit paranoid about losing any true signal though, so my idea - not yet even tried - is to always keep the best for every base feature. Or even best 2.\n\nFor example, try weighted average, mean, median: keep the best one on an individual column basis. (And maybe keep more if above some threshold, like top 300 or whatever.) Try min and max: keep the higher importance one. Try last minus first, last minus mean, last minus prior, the same but divide instead of subtract, again just keep the best 1 per base feature.\n\nAnother more convoluted idea, the idea that less correlated models ensemble better: let's say we think up around 50 feature sets total to apply to our ~180 numeric columns. 9000 total. We do some method to get scores for everything. Then instead of taking the best scores and dropping the rest, perhaps we simply allow the higher scoring features to be included in more individual models? top 180 go in every model, next best 180*3 go in 1/3rd of the models, and the rest only go in one model each. 8 models of around 1400 features each. I'm doubtful but might try it and see. :)",
      "votes": 1
    },
    {
      "id": 1870351,
      "postDate": "2022-07-25T13:45:27.950Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/illidan7\" target=\"_blank\">@illidan7</a> , I tried your approach today. Here are my  results from xgb model without <code>dart</code>.</p>\n<blockquote>\n  <p>Base_features + feature_set1 :  score: <code>0.792373</code>, CV: <code>0.79303222093288</code>, pblb score: <code>0.796</code></p>\n</blockquote>\n<p>Found the New set of features using split_importance. Then optimized model again  on this new set and made prediction.</p>\n<blockquote>\n  <p>New_features: score: <code>0.792535</code>,  CV: <code>0.7922492740935698</code>,  pblb score: <code>0.794</code></p>\n</blockquote>\n<p>score is the amex_score on doing hyperparameter optimization. <br>\nWe see that, though score improves slightly but <code>CV</code> decreases and similarly <code>pblb_score</code> also decreases. </p>\n<p>I am not sure if this approach works. Did you face this issue?</p>",
      "rawMarkdown": "Hi @illidan7 , I tried your approach today. Here are my  results from xgb model without `dart`.\n\n> Base_features + feature_set1 :  score: `0.792373`, CV: `0.79303222093288`, pblb score: `0.796`\n\nFound the New set of features using split_importance. Then optimized model again  on this new set and made prediction.\n\n> New_features: score: `0.792535`,  CV: `0.7922492740935698`,  pblb score: `0.794`\n\nscore is the amex_score on doing hyperparameter optimization. \nWe see that, though score improves slightly but `CV` decreases and similarly `pblb_score ` also decreases. \n\nI am not sure if this approach works. Did you face this issue?",
      "votes": 1,
      "replies": [
        {
          "id": 1870477,
          "postDate": "2022-07-25T15:18:03.787Z",
          "content": "<p>So far, I have been able to improve score incrementally. My approach has been to peel from the top in terms of split importance from the \"Blowup\" dataset and try incorporating those to my base feature set. It does take some experimentation in terms of including/excluding features, and I try to treat the split importance as a rough guide and not absolute </p>\n<p>Are you using split importance averaged across all folds? </p>\n<p>Also, there is a lot of random behavior in terms of the public scoring metric<br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/336957\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/336957</a></p>\n<p>Just a nice to know, but I wouldn't spend too much time blaming the scoring and rather do our best with our CV process</p>",
          "rawMarkdown": "So far, I have been able to improve score incrementally. My approach has been to peel from the top in terms of split importance from the \"Blowup\" dataset and try incorporating those to my base feature set. It does take some experimentation in terms of including/excluding features, and I try to treat the split importance as a rough guide and not absolute \n\nAre you using split importance averaged across all folds? \n\nAlso, there is a lot of random behavior in terms of the public scoring metric\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/336957\n\nJust a nice to know, but I wouldn't spend too much time blaming the scoring and rather do our best with our CV process",
          "votes": 1
        },
        {
          "id": 1870492,
          "postDate": "2022-07-25T15:34:49.003Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/illidan7\" target=\"_blank\">@illidan7</a> , here is my approach. </p>\n<p>I do hyperparameter tuning on 80% of training set using optuna. <br>\nI run optuna for 10 trials. In all of these trials I store <code>split_importance</code> as columns to a table. At the end I pick top 3 best trials take the weighted mean of their feature importance based on trial score and then pick all features whose importance is greater than <code>10.0</code>. </p>\n<p>Next I use this as base and add new set to this base to repeat the process.<br>\nMaybe the new feature set is not important i.e. it is just some random noise. But the fact that it decreased my cv,lb makes it little unreliable (same behavior explained in <a href=\"https://scikit-learn.org/stable/auto_examples/inspection/plot_permutation_importance.html\" target=\"_blank\">here</a>). <br>\nSo if you added a new feature set, did the split_importance but it decreased you cv, then did you continued with adding next feature set , or just discarded that feature set and tried new feature set.</p>\n<p>Just a little confirmation: by <code>split_importance</code> you mean <code>model.get_score(importance_type='weight')</code>  of <code>xgb</code>.<br>\nref notebook: <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">this</a> of <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> .</p>",
          "rawMarkdown": "Hi @illidan7 , here is my approach. \n\nI do hyperparameter tuning on 80% of training set using optuna. \nI run optuna for 10 trials. In all of these trials I store `split_importance` as columns to a table. At the end I pick top 3 best trials take the weighted mean of their feature importance based on trial score and then pick all features whose importance is greater than `10.0`. \n\nNext I use this as base and add new set to this base to repeat the process.\nMaybe the new feature set is not important i.e. it is just some random noise. But the fact that it decreased my cv,lb makes it little unreliable (same behavior explained in [here](https://scikit-learn.org/stable/auto_examples/inspection/plot_permutation_importance.html)). \nSo if you added a new feature set, did the split_importance but it decreased you cv, then did you continued with adding next feature set , or just discarded that feature set and tried new feature set.\n\nJust a little confirmation: by `split_importance` you mean `model.get_score(importance_type='weight')`  of `xgb`.\nref notebook: [this](https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793) of @cdeotte ."
        },
        {
          "id": 1870513,
          "postDate": "2022-07-25T15:48:30.173Z",
          "content": "<p>So I guess it depends on the size of dataset you are using and also the number of boosting rounds, depth etc. of your trees, but I have not tried to include importance &gt; 10. I am trying to use split importance sparingly, so my assumption is that the variables way at the top are more reliable. In the example above, I have only brought in the top 100 or so features for example from this feature set (<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/336557\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/336557</a>) to my base</p>\n<p>If there are redundant features providing the same signal their importance will the shared and therefore dampened. These will take a little more digging to find on the rank ordering of split importance</p>\n<blockquote>\n  <p>Just a little confirmation: by split_importance you mean model.get_score(importance_type='weight')</p>\n</blockquote>\n<p>Yup this is the one</p>",
          "rawMarkdown": "So I guess it depends on the size of dataset you are using and also the number of boosting rounds, depth etc. of your trees, but I have not tried to include importance > 10. I am trying to use split importance sparingly, so my assumption is that the variables way at the top are more reliable. In the example above, I have only brought in the top 100 or so features for example from this feature set (https://www.kaggle.com/competitions/amex-default-prediction/discussion/336557) to my base\n\nIf there are redundant features providing the same signal their importance will the shared and therefore dampened. These will take a little more digging to find on the rank ordering of split importance\n\n> Just a little confirmation: by split_importance you mean model.get_score(importance_type='weight')\n\nYup this is the one",
          "votes": 3
        },
        {
          "id": 1870636,
          "postDate": "2022-07-25T17:22:53.587Z",
          "content": "<p>thanks, I will try what you said.👍</p>",
          "rawMarkdown": "thanks, I will try what you said.👍"
        }
      ]
    },
    {
      "id": 1869008,
      "postDate": "2022-07-24T12:08:00.543Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/illidan7\" target=\"_blank\">@illidan7</a> , <strong>Base-&gt;Blowup-&gt;Base</strong> this is really a good approach. One thing I want to ask you.<br>\nWe know <code>permutation importance</code> is really time consuming so we switched to <code>split importance</code>. <br>\nDo you use it with <code>dart</code> booster of <code>xgb</code>, because that would be still slow. <br>\nIf we remove <code>dart</code> parameter and find feature importance using <code>split importance</code>.<br>\nNext we use the filtered features and train with <code>dart</code> then that might be problematic. </p>\n<p>What is your approach. thanks : )</p>",
      "rawMarkdown": "Hey @illidan7 , **Base->Blowup->Base** this is really a good approach. One thing I want to ask you.\nWe know `permutation importance` is really time consuming so we switched to `split importance`. \nDo you use it with `dart` booster of `xgb`, because that would be still slow. \nIf we remove `dart` parameter and find feature importance using `split importance`.\nNext we use the filtered features and train with `dart` then that might be problematic. \n\nWhat is your approach. thanks : )",
      "votes": 1,
      "replies": [
        {
          "id": 1869240,
          "postDate": "2022-07-24T15:47:03.533Z",
          "content": "<p>In my approach so far, I have not used dart booster for XGB. I do not know much about dart but what I have gathered from other discussions is that it drops trees during the training process kind of like dropout for NNs. It might shake up the split importance ordering a little bit, but my intuition is that it would be largely be similar. I may be wrong here</p>\n<p>In any case, dart on XGB is so slow that it makes it unviable for me to try </p>",
          "rawMarkdown": "In my approach so far, I have not used dart booster for XGB. I do not know much about dart but what I have gathered from other discussions is that it drops trees during the training process kind of like dropout for NNs. It might shake up the split importance ordering a little bit, but my intuition is that it would be largely be similar. I may be wrong here\n\nIn any case, dart on XGB is so slow that it makes it unviable for me to try ",
          "votes": 1
        },
        {
          "id": 1869391,
          "postDate": "2022-07-24T17:54:35.470Z",
          "content": "<p>thanks for the reply. Glad to know that you were able to do feature selection without using dart👍</p>",
          "rawMarkdown": "thanks for the reply. Glad to know that you were able to do feature selection without using dart👍"
        }
      ]
    },
    {
      "id": 1868232,
      "postDate": "2022-07-23T20:15:49.693Z",
      "content": "<p>Thanks for sharing! The larger feature set includes the base feature set?</p>",
      "rawMarkdown": "Thanks for sharing! The larger feature set includes the base feature set?",
      "votes": 2,
      "replies": [
        {
          "id": 1868312,
          "postDate": "2022-07-23T22:51:02.480Z",
          "content": "<p>In my experiments, I have included the base feature set to allow the new features to compete with the strongest ones so far. But it may be good to try it without the base features as well. We need to be careful not to introduce too many redundant features though</p>",
          "rawMarkdown": "In my experiments, I have included the base feature set to allow the new features to compete with the strongest ones so far. But it may be good to try it without the base features as well. We need to be careful not to introduce too many redundant features though",
          "votes": 1
        }
      ]
    },
    {
      "id": 1867552,
      "postDate": "2022-07-23T10:26:20.100Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 1868204,
          "postDate": "2022-07-23T19:32:07.763Z",
          "content": "<p>It is true but this is also exactly why I used split importance because it is very easy to get as part of the training process. And I'm counting on it to be at least directionally accurate of which features are helping the model learn on the dataset. I think layering in other feature importance metrics would be a great value addition in the overall feature selection process</p>\n<p>Check out this notebook </p>\n<p><a href=\"https://www.kaggle.com/code/illidan7/amex-permutation-feature-importance\" target=\"_blank\">https://www.kaggle.com/code/illidan7/amex-permutation-feature-importance</a></p>",
          "rawMarkdown": "It is true but this is also exactly why I used split importance because it is very easy to get as part of the training process. And I'm counting on it to be at least directionally accurate of which features are helping the model learn on the dataset. I think layering in other feature importance metrics would be a great value addition in the overall feature selection process\n\nCheck out this notebook \n\nhttps://www.kaggle.com/code/illidan7/amex-permutation-feature-importance",
          "votes": 1
        }
      ]
    },
    {
      "id": 1867313,
      "postDate": "2022-07-23T06:12:36.457Z",
      "rawMarkdown": "",
      "votes": -3,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1875334,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2022-07-28T23:40:28.630000",
      "content": "<p>Thanks for sharing!</p>\n<p>It sounds similar to my idea. I'm a bit paranoid about losing any true signal though, so my idea - not yet even tried - is to always keep the best for every base feature. Or even best 2.</p>\n<p>For example, try weighted average, mean, median: keep the best one on an individual column basis. (And maybe keep more if above some threshold, like top 300 or whatever.) Try min and max: keep the higher importance one. Try last minus first, last minus mean, last minus prior, the same but divide instead of subtract, again just keep the best 1 per base feature.</p>\n<p>Another more convoluted idea, the idea that less correlated models ensemble better: let's say we think up around 50 feature sets total to apply to our ~180 numeric columns. 9000 total. We do some method to get scores for everything. Then instead of taking the best scores and dropping the rest, perhaps we simply allow the higher scoring features to be included in more individual models? top 180 go in every model, next best 180*3 go in 1/3rd of the models, and the rest only go in one model each. 8 models of around 1400 features each. I'm doubtful but might try it and see. :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1870351,
      "author_name": "AKR",
      "author_url": "",
      "post_date": "2022-07-25T13:45:27.950000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/illidan7\" target=\"_blank\">@illidan7</a> , I tried your approach today. Here are my  results from xgb model without <code>dart</code>.</p>\n<blockquote>\n  <p>Base_features + feature_set1 :  score: <code>0.792373</code>, CV: <code>0.79303222093288</code>, pblb score: <code>0.796</code></p>\n</blockquote>\n<p>Found the New set of features using split_importance. Then optimized model again  on this new set and made prediction.</p>\n<blockquote>\n  <p>New_features: score: <code>0.792535</code>,  CV: <code>0.7922492740935698</code>,  pblb score: <code>0.794</code></p>\n</blockquote>\n<p>score is the amex_score on doing hyperparameter optimization. <br>\nWe see that, though score improves slightly but <code>CV</code> decreases and similarly <code>pblb_score</code> also decreases. </p>\n<p>I am not sure if this approach works. Did you face this issue?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1870477,
          "author_name": "Noir.s",
          "author_url": "",
          "post_date": "2022-07-25T15:18:03.787000",
          "content": "<p>So far, I have been able to improve score incrementally. My approach has been to peel from the top in terms of split importance from the \"Blowup\" dataset and try incorporating those to my base feature set. It does take some experimentation in terms of including/excluding features, and I try to treat the split importance as a rough guide and not absolute </p>\n<p>Are you using split importance averaged across all folds? </p>\n<p>Also, there is a lot of random behavior in terms of the public scoring metric<br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/336957\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/336957</a></p>\n<p>Just a nice to know, but I wouldn't spend too much time blaming the scoring and rather do our best with our CV process</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1870492,
          "author_name": "AKR",
          "author_url": "",
          "post_date": "2022-07-25T15:34:49.003000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/illidan7\" target=\"_blank\">@illidan7</a> , here is my approach. </p>\n<p>I do hyperparameter tuning on 80% of training set using optuna. <br>\nI run optuna for 10 trials. In all of these trials I store <code>split_importance</code> as columns to a table. At the end I pick top 3 best trials take the weighted mean of their feature importance based on trial score and then pick all features whose importance is greater than <code>10.0</code>. </p>\n<p>Next I use this as base and add new set to this base to repeat the process.<br>\nMaybe the new feature set is not important i.e. it is just some random noise. But the fact that it decreased my cv,lb makes it little unreliable (same behavior explained in <a href=\"https://scikit-learn.org/stable/auto_examples/inspection/plot_permutation_importance.html\" target=\"_blank\">here</a>). <br>\nSo if you added a new feature set, did the split_importance but it decreased you cv, then did you continued with adding next feature set , or just discarded that feature set and tried new feature set.</p>\n<p>Just a little confirmation: by <code>split_importance</code> you mean <code>model.get_score(importance_type='weight')</code>  of <code>xgb</code>.<br>\nref notebook: <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">this</a> of <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> .</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1870513,
          "author_name": "Noir.s",
          "author_url": "",
          "post_date": "2022-07-25T15:48:30.173000",
          "content": "<p>So I guess it depends on the size of dataset you are using and also the number of boosting rounds, depth etc. of your trees, but I have not tried to include importance &gt; 10. I am trying to use split importance sparingly, so my assumption is that the variables way at the top are more reliable. In the example above, I have only brought in the top 100 or so features for example from this feature set (<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/336557\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/336557</a>) to my base</p>\n<p>If there are redundant features providing the same signal their importance will the shared and therefore dampened. These will take a little more digging to find on the rank ordering of split importance</p>\n<blockquote>\n  <p>Just a little confirmation: by split_importance you mean model.get_score(importance_type='weight')</p>\n</blockquote>\n<p>Yup this is the one</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1870636,
          "author_name": "AKR",
          "author_url": "",
          "post_date": "2022-07-25T17:22:53.587000",
          "content": "<p>thanks, I will try what you said.👍</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1869008,
      "author_name": "AKR",
      "author_url": "",
      "post_date": "2022-07-24T12:08:00.543000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/illidan7\" target=\"_blank\">@illidan7</a> , <strong>Base-&gt;Blowup-&gt;Base</strong> this is really a good approach. One thing I want to ask you.<br>\nWe know <code>permutation importance</code> is really time consuming so we switched to <code>split importance</code>. <br>\nDo you use it with <code>dart</code> booster of <code>xgb</code>, because that would be still slow. <br>\nIf we remove <code>dart</code> parameter and find feature importance using <code>split importance</code>.<br>\nNext we use the filtered features and train with <code>dart</code> then that might be problematic. </p>\n<p>What is your approach. thanks : )</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1869240,
          "author_name": "Noir.s",
          "author_url": "",
          "post_date": "2022-07-24T15:47:03.533000",
          "content": "<p>In my approach so far, I have not used dart booster for XGB. I do not know much about dart but what I have gathered from other discussions is that it drops trees during the training process kind of like dropout for NNs. It might shake up the split importance ordering a little bit, but my intuition is that it would be largely be similar. I may be wrong here</p>\n<p>In any case, dart on XGB is so slow that it makes it unviable for me to try </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1869391,
          "author_name": "AKR",
          "author_url": "",
          "post_date": "2022-07-24T17:54:35.470000",
          "content": "<p>thanks for the reply. Glad to know that you were able to do feature selection without using dart👍</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1868232,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "2022-07-23T20:15:49.693000",
      "content": "<p>Thanks for sharing! The larger feature set includes the base feature set?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1868312,
          "author_name": "Noir.s",
          "author_url": "",
          "post_date": "2022-07-23T22:51:02.480000",
          "content": "<p>In my experiments, I have included the base feature set to allow the new features to compete with the strongest ones so far. But it may be good to try it without the base features as well. We need to be careful not to introduce too many redundant features though</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1867552,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-23T10:26:20.100000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1868204,
          "author_name": "Noir.s",
          "author_url": "",
          "post_date": "2022-07-23T19:32:07.763000",
          "content": "<p>It is true but this is also exactly why I used split importance because it is very easy to get as part of the training process. And I'm counting on it to be at least directionally accurate of which features are helping the model learn on the dataset. I think layering in other feature importance metrics would be a great value addition in the overall feature selection process</p>\n<p>Check out this notebook </p>\n<p><a href=\"https://www.kaggle.com/code/illidan7/amex-permutation-feature-importance\" target=\"_blank\">https://www.kaggle.com/code/illidan7/amex-permutation-feature-importance</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1867313,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-23T06:12:36.457000",
      "content": "",
      "votes": -3,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1867308": "Simply running group by aggregations (min, max, mean, last etc.) on all available features quickly blows up the feature set and we soon run into memory issues and features being reused redundantly within the model. Have not had as much time as I would have liked and all I have built so far are XGB models.\n\nI've been using a fairly simplistic feature selection approach, which I don't have really have a name for, so I made up the title. It's basically some sort of a manual stepwise forward selection. But so far it has been effective in improving CV score while minimizing the number of features being used in my modeling dataset. \n\n_______________________________________________\nBASE (469 features)\n_______________________________________________\nCV 0.7937\n\nStarted with a minimum set of features which I know has decent performance. I used the ones from this notebook (https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart) which has 469 features\n\nFit XGB model, removed least important features (split importance), made sure model performance did not degrade a whole lot\n\n_______________________________________________\nBLOWUP (1486 features)\n_______________________________________________\nCV 0.7939\n\nThen built a larger feature set incorporating different feature engineering ideas from the competition, blowing up the feature set to ~1500 features (https://www.kaggle.com/competitions/amex-default-prediction/discussion/336557)\n\nFit model, \n\n_______________________________________________\nBASE (520 features)\n_______________________________________________\nCV 0.7947\n\nThen brought in only the most important features to my base feature set. Creating my \"new\" base feature set\n\n_______________________________________________\nREPEAT\n_______________________________________________\n\nAnd so on..... Build a larger feature set with more feature engineering ideas. Bring back the most important features to the base dataset \n\nThe disadvantage is that, the features being dropped do not get to interact with the new features in the larger dataset at each stage. Which could mean there is a lot of value being lost with this approach\n\nLook forward to any feedback on this and any new ideas you can share on how to go about feature selection. \n\nThanks!\n\n",
    "1875334": "Thanks for sharing!\n\nIt sounds similar to my idea. I'm a bit paranoid about losing any true signal though, so my idea - not yet even tried - is to always keep the best for every base feature. Or even best 2.\n\nFor example, try weighted average, mean, median: keep the best one on an individual column basis. (And maybe keep more if above some threshold, like top 300 or whatever.) Try min and max: keep the higher importance one. Try last minus first, last minus mean, last minus prior, the same but divide instead of subtract, again just keep the best 1 per base feature.\n\nAnother more convoluted idea, the idea that less correlated models ensemble better: let's say we think up around 50 feature sets total to apply to our ~180 numeric columns. 9000 total. We do some method to get scores for everything. Then instead of taking the best scores and dropping the rest, perhaps we simply allow the higher scoring features to be included in more individual models? top 180 go in every model, next best 180*3 go in 1/3rd of the models, and the rest only go in one model each. 8 models of around 1400 features each. I'm doubtful but might try it and see. :)",
    "1870351": "Hi @illidan7 , I tried your approach today. Here are my  results from xgb model without `dart`.\n\n> Base_features + feature_set1 :  score: `0.792373`, CV: `0.79303222093288`, pblb score: `0.796`\n\nFound the New set of features using split_importance. Then optimized model again  on this new set and made prediction.\n\n> New_features: score: `0.792535`,  CV: `0.7922492740935698`,  pblb score: `0.794`\n\nscore is the amex_score on doing hyperparameter optimization. \nWe see that, though score improves slightly but `CV` decreases and similarly `pblb_score ` also decreases. \n\nI am not sure if this approach works. Did you face this issue?",
    "1869008": "Hey @illidan7 , **Base->Blowup->Base** this is really a good approach. One thing I want to ask you.\nWe know `permutation importance` is really time consuming so we switched to `split importance`. \nDo you use it with `dart` booster of `xgb`, because that would be still slow. \nIf we remove `dart` parameter and find feature importance using `split importance`.\nNext we use the filtered features and train with `dart` then that might be problematic. \n\nWhat is your approach. thanks : )",
    "1868232": "Thanks for sharing! The larger feature set includes the base feature set?",
    "1867552": "",
    "1867313": ""
  }
}