{
  "id": 336145,
  "title": "Feature Selection",
  "url": "/competitions/amex-default-prediction/discussion/336145",
  "author_name": "",
  "post_date": "2022-07-09T15:28:27.125463400Z",
  "votes": 31,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hello Kagglers, </p>\n<p>I am curious if some of you have had luck with feature selection. I have spent quite some time trying to eliminate non-useful features but it doesn't seem to make my model better. Sometimes even when I removed the 0-split/gain features from my LGB model, cv got worse. Maybe it's due to instability of the metric? </p>",
  "messages": [
    {
      "id": "1849496",
      "postDate": "07/09/2022 15:28:27",
      "content": "<p>Hello Kagglers, </p>\n<p>I am curious if some of you have had luck with feature selection. I have spent quite some time trying to eliminate non-useful features but it doesn't seem to make my model better. Sometimes even when I removed the 0-split/gain features from my LGB model, cv got worse. Maybe it's due to instability of the metric? </p>",
      "rawMarkdown": "Hello Kagglers, \n\nI am curious if some of you have had luck with feature selection. I have spent quite some time trying to eliminate non-useful features but it doesn't seem to make my model better. Sometimes even when I removed the 0-split/gain features from my LGB model, cv got worse. Maybe it's due to instability of the metric?",
      "votes": null
    },
    {
      "id": "1849959",
      "postDate": "07/10/2022 00:48:42",
      "content": "<p>Do you use the <a href=\"https://www.kaggle.com/code/ogrellier/feature-selection-with-null-importances\" target=\"_blank\">one</a> and only rightful feature selection method?</p>",
      "rawMarkdown": "Do you use the [one](https://www.kaggle.com/code/ogrellier/feature-selection-with-null-importances) and only rightful feature selection method?",
      "votes": null
    },
    {
      "id": "1849966",
      "postDate": "07/10/2022 00:53:41",
      "content": "<p>I didn’t. I did try permutation importance and its variations. I believe they are in the same arena? </p>",
      "rawMarkdown": "I didn’t. I did try permutation importance and its variations. I believe they are in the same arena?",
      "votes": null
    },
    {
      "id": "1850205",
      "postDate": "07/10/2022 07:41:45",
      "content": "<p>If you are not worried about your RAM, I suggest to keep all the non-correlated features. LGBM will take care of non-useful ones.</p>",
      "rawMarkdown": "If you are not worried about your RAM, I suggest to keep all the non-correlated features. LGBM will take care of non-useful ones.",
      "votes": null
    },
    {
      "id": "1850602",
      "postDate": "07/10/2022 14:24:26",
      "content": "<p>Thanks for the feedback. Usually I'd like to build a lightweight model and get rid of non-useful features manually, and it usually works. The fact that removing 0-importance features impact cv makes me wonder if I can trust my cv at all. How frustrating.</p>",
      "rawMarkdown": "Thanks for the feedback. Usually I'd like to build a lightweight model and get rid of non-useful features manually, and it usually works. The fact that removing 0-importance features impact cv makes me wonder if I can trust my cv at all. How frustrating.",
      "votes": null
    },
    {
      "id": "1850644",
      "postDate": "07/10/2022 15:14:46",
      "content": "<p>gain and split are not good indicators for feature importance. you may want to use permutation. </p>",
      "rawMarkdown": "gain and split are not good indicators for feature importance. you may want to use permutation.",
      "votes": null
    },
    {
      "id": "1850687",
      "postDate": "07/10/2022 16:07:44",
      "content": "<p>Part of the reason why feature selection and engineering is difficult in this comp is because many models have large variation in their CV scores if you train the same exact model with a different seed. </p>\n<p>For example, if you run my XGB starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">here</a> 200 times with 200 different XGB seeds and the same 5 folds each time, the CV varies a lot. The mean overall CV is 0.7920 and the STD = 0.0012. And of course this is the same model each time <strong>with the same expected LB score</strong>. (Changing seed does not make a model better or worse).</p>\n<p>This means that if we <strong>add or remove</strong> a feature which has <strong>no effect</strong>, there is a 1 out of 6 chance that the CV score will boost <code>+0.0012</code> and 1 out of 44 chance that the CV score will boost <code>+0.0024</code> by pure random chance. This misleads us into thinking the model improved when the model did not change (and the feature we added or removed had no effect).</p>\n<p>In conclusion with regard to my XGB starter. We need to see a boost (or drop) of <code>0.0025</code> or more in the overall CV before we can be statistically 95% confident that the change in features made a difference.</p>\n<p>(Note that for other models such as LGBM, NN, XGB with different features, etc, you need to compute the STD of that overall CV to determine statistically significant threshold for those models).</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jul-2022/xgb_std.png\" alt=\"\"></p>",
      "rawMarkdown": "Part of the reason why feature selection and engineering is difficult in this comp is because many models have large variation in their CV scores if you train the same exact model with a different seed. \n\nFor example, if you run my XGB starter notebook [here][1] 200 times with 200 different XGB seeds and the same 5 folds each time, the CV varies a lot. The mean overall CV is 0.7920 and the STD = 0.0012. And of course this is the same model each time **with the same expected LB score**. (Changing seed does not make a model better or worse).\n\nThis means that if we **add or remove** a feature which has **no effect**, there is a 1 out of 6 chance that the CV score will boost `+0.0012` and 1 out of 44 chance that the CV score will boost `+0.0024` by pure random chance. This misleads us into thinking the model improved when the model did not change (and the feature we added or removed had no effect).\n\nIn conclusion with regard to my XGB starter. We need to see a boost (or drop) of `0.0025` or more in the overall CV before we can be statistically 95% confident that the change in features made a difference.\n\n(Note that for other models such as LGBM, NN, XGB with different features, etc, you need to compute the STD of that overall CV to determine statistically significant threshold for those models).\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jul-2022/xgb_std.png)\n\n[1]: https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793",
      "votes": null
    },
    {
      "id": "1850699",
      "postDate": "07/10/2022 16:21:50",
      "content": "<p>Nice experiment Chris! I knew there's instability in this metric but it didn't occur to me that we could set up an experiment and quantify it. Learned a lot! </p>",
      "rawMarkdown": "Nice experiment Chris! I knew there's instability in this metric but it didn't occur to me that we could set up an experiment and quantify it. Learned a lot!",
      "votes": null
    },
    {
      "id": "1851811",
      "postDate": "07/11/2022 14:52:00",
      "content": "<p>Based on this parameter 'colsample_bytree', doesn't it mean that you can get a better model by dropping the unimportant features and it is not just a RAM constraint issue? Let me know your thoughts</p>",
      "rawMarkdown": "Based on this parameter 'colsample_bytree', doesn't it mean that you can get a better model by dropping the unimportant features and it is not just a RAM constraint issue? Let me know your thoughts",
      "votes": null
    },
    {
      "id": "1851823",
      "postDate": "07/11/2022 15:01:35",
      "content": "<p>Thank you for sharing this <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !</p>\n<p>Definitely looking forward to implementing this and testing out since I am having trouble with feature selection</p>\n<p>My followup question is: <br>\nIs the feature set also dependent on the hyperparameters? If we do hyperparameter tuning at a later point in the pipeline, does this mean that there may have been better features that were discarded due to the baseline set of parameters used?</p>\n<p>My intuition says yes the parameters do affect the feature set selected but I'm not sure how to address the issue. Would love to know your thoughts on this!</p>",
      "rawMarkdown": "Thank you for sharing this @cdeotte !\n\nDefinitely looking forward to implementing this and testing out since I am having trouble with feature selection\n\nMy followup question is: \nIs the feature set also dependent on the hyperparameters? If we do hyperparameter tuning at a later point in the pipeline, does this mean that there may have been better features that were discarded due to the baseline set of parameters used?\n\nMy intuition says yes the parameters do affect the feature set selected but I'm not sure how to address the issue. Would love to know your thoughts on this!",
      "votes": null
    },
    {
      "id": "1851869",
      "postDate": "07/11/2022 15:34:42",
      "content": "<p>Thanks for open this thread. I am facing the same issue, and adding that permutation importance is too slow with that huge number of features. Doing it in a stepped way (i.e. train the model -&gt; calculate importances -&gt; remove the least useful features -&gt; train a model again -&gt; calculate importances, …) could work, or at least somehow worked for me in a couple of past submissions.</p>",
      "rawMarkdown": "Thanks for open this thread. I am facing the same issue, and adding that permutation importance is too slow with that huge number of features. Doing it in a stepped way (i.e. train the model -> calculate importances -> remove the least useful features -> train a model again -> calculate importances, ...) could work, or at least somehow worked for me in a couple of past submissions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1849959,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "07/10/2022 00:48:42",
      "content": "<p>Do you use the <a href=\"https://www.kaggle.com/code/ogrellier/feature-selection-with-null-importances\" target=\"_blank\">one</a> and only rightful feature selection method?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1849966,
          "author_name": "raphael1123",
          "author_url": "",
          "post_date": "07/10/2022 00:53:41",
          "content": "<p>I didn’t. I did try permutation importance and its variations. I believe they are in the same arena? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1850205,
      "author_name": "mohammadrahmati",
      "author_url": "",
      "post_date": "07/10/2022 07:41:45",
      "content": "<p>If you are not worried about your RAM, I suggest to keep all the non-correlated features. LGBM will take care of non-useful ones.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1850602,
          "author_name": "raphael1123",
          "author_url": "",
          "post_date": "07/10/2022 14:24:26",
          "content": "<p>Thanks for the feedback. Usually I'd like to build a lightweight model and get rid of non-useful features manually, and it usually works. The fact that removing 0-importance features impact cv makes me wonder if I can trust my cv at all. How frustrating.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1850644,
          "author_name": "mohammadrahmati",
          "author_url": "",
          "post_date": "07/10/2022 15:14:46",
          "content": "<p>gain and split are not good indicators for feature importance. you may want to use permutation. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1851811,
          "author_name": "illidan7",
          "author_url": "",
          "post_date": "07/11/2022 14:52:00",
          "content": "<p>Based on this parameter 'colsample_bytree', doesn't it mean that you can get a better model by dropping the unimportant features and it is not just a RAM constraint issue? Let me know your thoughts</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1851869,
          "author_name": "delai50",
          "author_url": "",
          "post_date": "07/11/2022 15:34:42",
          "content": "<p>Thanks for open this thread. I am facing the same issue, and adding that permutation importance is too slow with that huge number of features. Doing it in a stepped way (i.e. train the model -&gt; calculate importances -&gt; remove the least useful features -&gt; train a model again -&gt; calculate importances, …) could work, or at least somehow worked for me in a couple of past submissions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1850687,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/10/2022 16:07:44",
      "content": "<p>Part of the reason why feature selection and engineering is difficult in this comp is because many models have large variation in their CV scores if you train the same exact model with a different seed. </p>\n<p>For example, if you run my XGB starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">here</a> 200 times with 200 different XGB seeds and the same 5 folds each time, the CV varies a lot. The mean overall CV is 0.7920 and the STD = 0.0012. And of course this is the same model each time <strong>with the same expected LB score</strong>. (Changing seed does not make a model better or worse).</p>\n<p>This means that if we <strong>add or remove</strong> a feature which has <strong>no effect</strong>, there is a 1 out of 6 chance that the CV score will boost <code>+0.0012</code> and 1 out of 44 chance that the CV score will boost <code>+0.0024</code> by pure random chance. This misleads us into thinking the model improved when the model did not change (and the feature we added or removed had no effect).</p>\n<p>In conclusion with regard to my XGB starter. We need to see a boost (or drop) of <code>0.0025</code> or more in the overall CV before we can be statistically 95% confident that the change in features made a difference.</p>\n<p>(Note that for other models such as LGBM, NN, XGB with different features, etc, you need to compute the STD of that overall CV to determine statistically significant threshold for those models).</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jul-2022/xgb_std.png\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1850699,
          "author_name": "raphael1123",
          "author_url": "",
          "post_date": "07/10/2022 16:21:50",
          "content": "<p>Nice experiment Chris! I knew there's instability in this metric but it didn't occur to me that we could set up an experiment and quantify it. Learned a lot! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1851823,
          "author_name": "illidan7",
          "author_url": "",
          "post_date": "07/11/2022 15:01:35",
          "content": "<p>Thank you for sharing this <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !</p>\n<p>Definitely looking forward to implementing this and testing out since I am having trouble with feature selection</p>\n<p>My followup question is: <br>\nIs the feature set also dependent on the hyperparameters? If we do hyperparameter tuning at a later point in the pipeline, does this mean that there may have been better features that were discarded due to the baseline set of parameters used?</p>\n<p>My intuition says yes the parameters do affect the feature set selected but I'm not sure how to address the issue. Would love to know your thoughts on this!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1849496": "Hello Kagglers, \n\nI am curious if some of you have had luck with feature selection. I have spent quite some time trying to eliminate non-useful features but it doesn't seem to make my model better. Sometimes even when I removed the 0-split/gain features from my LGB model, cv got worse. Maybe it's due to instability of the metric?",
    "1849959": "Do you use the [one](https://www.kaggle.com/code/ogrellier/feature-selection-with-null-importances) and only rightful feature selection method?",
    "1849966": "I didn’t. I did try permutation importance and its variations. I believe they are in the same arena?",
    "1850205": "If you are not worried about your RAM, I suggest to keep all the non-correlated features. LGBM will take care of non-useful ones.",
    "1850602": "Thanks for the feedback. Usually I'd like to build a lightweight model and get rid of non-useful features manually, and it usually works. The fact that removing 0-importance features impact cv makes me wonder if I can trust my cv at all. How frustrating.",
    "1850644": "gain and split are not good indicators for feature importance. you may want to use permutation.",
    "1850687": "Part of the reason why feature selection and engineering is difficult in this comp is because many models have large variation in their CV scores if you train the same exact model with a different seed. \n\nFor example, if you run my XGB starter notebook [here][1] 200 times with 200 different XGB seeds and the same 5 folds each time, the CV varies a lot. The mean overall CV is 0.7920 and the STD = 0.0012. And of course this is the same model each time **with the same expected LB score**. (Changing seed does not make a model better or worse).\n\nThis means that if we **add or remove** a feature which has **no effect**, there is a 1 out of 6 chance that the CV score will boost `+0.0012` and 1 out of 44 chance that the CV score will boost `+0.0024` by pure random chance. This misleads us into thinking the model improved when the model did not change (and the feature we added or removed had no effect).\n\nIn conclusion with regard to my XGB starter. We need to see a boost (or drop) of `0.0025` or more in the overall CV before we can be statistically 95% confident that the change in features made a difference.\n\n(Note that for other models such as LGBM, NN, XGB with different features, etc, you need to compute the STD of that overall CV to determine statistically significant threshold for those models).\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jul-2022/xgb_std.png)\n\n[1]: https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793",
    "1850699": "Nice experiment Chris! I knew there's instability in this metric but it didn't occur to me that we could set up an experiment and quantify it. Learned a lot!",
    "1851811": "Based on this parameter 'colsample_bytree', doesn't it mean that you can get a better model by dropping the unimportant features and it is not just a RAM constraint issue? Let me know your thoughts",
    "1851823": "Thank you for sharing this @cdeotte !\n\nDefinitely looking forward to implementing this and testing out since I am having trouble with feature selection\n\nMy followup question is: \nIs the feature set also dependent on the hyperparameters? If we do hyperparameter tuning at a later point in the pipeline, does this mean that there may have been better features that were discarded due to the baseline set of parameters used?\n\nMy intuition says yes the parameters do affect the feature set selected but I'm not sure how to address the issue. Would love to know your thoughts on this!",
    "1851869": "Thanks for open this thread. I am facing the same issue, and adding that permutation importance is too slow with that huge number of features. Doing it in a stepped way (i.e. train the model -> calculate importances -> remove the least useful features -> train a model again -> calculate importances, ...) could work, or at least somehow worked for me in a couple of past submissions."
  },
  "source": "meta"
}