{
  "id": 406563,
  "title": "Problems with Permutation Importance",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/406563",
  "author_name": "",
  "post_date": "2023-05-02T19:38:36.094504200Z",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Just wondering if anyone has been able to use feature selection based on Permutation Importance.</p>\n<p>I've done a few runs using the old training set as the training set and the old test set as the test set, and cv increases by 0.002+, but the quality on the test degrades a bit in most cases.<br>\nI have tried 2 methods: </p>\n<ul>\n<li>compute feature importances using xgboost and then train xgboost on selected features</li>\n<li>compute feature importances using random forest and then train xgboost on selected features</li>\n</ul>\n<p>I also tried using BoostARoota, but that didn't help in my case either.</p>\n<p>Maybe I'm doing something wrong</p>",
  "messages": [
    {
      "id": "2243295",
      "postDate": "05/02/2023 19:38:36",
      "content": "<p>Just wondering if anyone has been able to use feature selection based on Permutation Importance.</p>\n<p>I've done a few runs using the old training set as the training set and the old test set as the test set, and cv increases by 0.002+, but the quality on the test degrades a bit in most cases.<br>\nI have tried 2 methods: </p>\n<ul>\n<li>compute feature importances using xgboost and then train xgboost on selected features</li>\n<li>compute feature importances using random forest and then train xgboost on selected features</li>\n</ul>\n<p>I also tried using BoostARoota, but that didn't help in my case either.</p>\n<p>Maybe I'm doing something wrong</p>",
      "rawMarkdown": "Just wondering if anyone has been able to use feature selection based on Permutation Importance.\n\nI've done a few runs using the old training set as the training set and the old test set as the test set, and cv increases by 0.002+, but the quality on the test degrades a bit in most cases.\nI have tried 2 methods: \n- compute feature importances using xgboost and then train xgboost on selected features\n- compute feature importances using random forest and then train xgboost on selected features\n\nI also tried using BoostARoota, but that didn't help in my case either.\n\nMaybe I'm doing something wrong",
      "votes": null
    },
    {
      "id": "2243411",
      "postDate": "05/02/2023 22:13:26",
      "content": "<p>I've been trying to figure this out too. I've computed the permutation importance of the new training set, and selected features accordingly but the CV barely increases. Moreover, the increase in CV score only happens when I remove a few features base on importance ranking, but if I try to remove more than a handful of features with negative importance, the CV score quickly degrades.</p>\n<p>I've implemented the permutation importance based on the macro F1 score. If you don't mind sharing, may I ask what metric did you use to evaluate the importance? </p>",
      "rawMarkdown": "I've been trying to figure this out too. I've computed the permutation importance of the new training set, and selected features accordingly but the CV barely increases. Moreover, the increase in CV score only happens when I remove a few features base on importance ranking, but if I try to remove more than a handful of features with negative importance, the CV score quickly degrades.\n\nI've implemented the permutation importance based on the macro F1 score. If you don't mind sharing, may I ask what metric did you use to evaluate the importance?",
      "votes": null
    },
    {
      "id": "2243627",
      "postDate": "05/03/2023 03:58:48",
      "content": "<p>May I ask how many features you guys use?  My original model with mass features generation has 1000/3000/6000 features, with some basic feature elimination it still has 400/1000/1000 features. Either case is just too cost-heavy for permutation importance.</p>\n<p>Also correct me if I'm wrong, permutation importance relies on CV to work, but according to my experiences CV and LB does not correlate so well when I do feature elimination. Feature elimination does help LB/CV boost a lot, they just do not get the boost synchronously. For my best LB score, local CV change ±0.0002 after I do basic feature elimination which I would consider as noise if I don't have some prior experience that it should work. Then it dramatically boosts LB by +0.002, which is certainly not noise. In such a case I doubt permutation importance would work with CV since it would also consider ±0.0002 as noise.</p>",
      "rawMarkdown": "May I ask how many features you guys use?  My original model with mass features generation has 1000/3000/6000 features, with some basic feature elimination it still has 400/1000/1000 features. Either case is just too cost-heavy for permutation importance.\n\nAlso correct me if I'm wrong, permutation importance relies on CV to work, but according to my experiences CV and LB does not correlate so well when I do feature elimination. Feature elimination does help LB/CV boost a lot, they just do not get the boost synchronously. For my best LB score, local CV change ±0.0002 after I do basic feature elimination which I would consider as noise if I don't have some prior experience that it should work. Then it dramatically boosts LB by +0.002, which is certainly not noise. In such a case I doubt permutation importance would work with CV since it would also consider ±0.0002 as noise.",
      "votes": null
    },
    {
      "id": "2243689",
      "postDate": "05/03/2023 05:11:39",
      "content": "<p>I used logloss</p>",
      "rawMarkdown": "I used logloss",
      "votes": null
    },
    {
      "id": "2243707",
      "postDate": "05/03/2023 05:31:53",
      "content": "<p>I used 1000/2500/5000 features (no selection) for my best cv-lb results</p>\n<p>Yes, indeed, the data looks very noisy(</p>",
      "rawMarkdown": "I used 1000/2500/5000 features (no selection) for my best cv-lb results\n\nYes, indeed, the data looks very noisy(",
      "votes": null
    },
    {
      "id": "2243721",
      "postDate": "05/03/2023 05:39:46",
      "content": "<p>Consider that eliminating feature also changes the way decision trees are grown, the boost in lb maybe is due to randomness. I've noticed +- 0.002 difference in macro F1 for different folds by simply shuffle around column orders.</p>",
      "rawMarkdown": "Consider that eliminating feature also changes the way decision trees are grown, the boost in lb maybe is due to randomness. I've noticed +- 0.002 difference in macro F1 for different folds by simply shuffle around column orders.",
      "votes": null
    },
    {
      "id": "2243739",
      "postDate": "05/03/2023 06:12:34",
      "content": "<p>According to my experience with applying basic feature elimination to different models, it would stably boost LB +0.001 with a ±0.001 randomness. There is a certain pattern which I can't give a proper explanation, like if I choose 1000 features then I always get the best LB score compared to 800 or 1200 features, even if the CV score of those different models show little differences around ±0.0002 randomly which I considered as noise.</p>\n<p>Honestly, I don't have any confidence in my elimination strategy, it most likely causes either leakage/overfitting according to other discussions. I believe the only reason it worked is that 6000 features are too much for a 20k dataset, and reducing it to around 1000 prevents the curse of dimensionality anyway. I checked both public notebooks/discussions and I see no common strategy for feature selection. People do talk about permutation importance, but I doubt it's too expensive to afford. My 400/1000/1000 features model takes 3 hours to run and I see no way I do permutation importance for every single feature. Maybe there is some hidden trick we would only learn when this competition ends.</p>",
      "rawMarkdown": "According to my experience with applying basic feature elimination to different models, it would stably boost LB +0.001 with a ±0.001 randomness. There is a certain pattern which I can't give a proper explanation, like if I choose 1000 features then I always get the best LB score compared to 800 or 1200 features, even if the CV score of those different models show little differences around ±0.0002 randomly which I considered as noise.\n\nHonestly, I don't have any confidence in my elimination strategy, it most likely causes either leakage/overfitting according to other discussions. I believe the only reason it worked is that 6000 features are too much for a 20k dataset, and reducing it to around 1000 prevents the curse of dimensionality anyway. I checked both public notebooks/discussions and I see no common strategy for feature selection. People do talk about permutation importance, but I doubt it's too expensive to afford. My 400/1000/1000 features model takes 3 hours to run and I see no way I do permutation importance for every single feature. Maybe there is some hidden trick we would only learn when this competition ends.",
      "votes": null
    },
    {
      "id": "2244093",
      "postDate": "05/03/2023 12:39:36",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/myppka\" target=\"_blank\">@myppka</a>,  I tried to keep the most important features, but it was improving my CV and deteriorating my LB score. It is probably overfitting the train data.</p>",
      "rawMarkdown": "Hi @myppka,  I tried to keep the most important features, but it was improving my CV and deteriorating my LB score. It is probably overfitting the train data.",
      "votes": null
    },
    {
      "id": "2244179",
      "postDate": "05/03/2023 13:46:59",
      "content": "<p>Yes, i have the same problem. Just after that i started experimenting with the old test set</p>",
      "rawMarkdown": "Yes, i have the same problem. Just after that i started experimenting with the old test set",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2243411,
      "author_name": "woprime",
      "author_url": "",
      "post_date": "05/02/2023 22:13:26",
      "content": "<p>I've been trying to figure this out too. I've computed the permutation importance of the new training set, and selected features accordingly but the CV barely increases. Moreover, the increase in CV score only happens when I remove a few features base on importance ranking, but if I try to remove more than a handful of features with negative importance, the CV score quickly degrades.</p>\n<p>I've implemented the permutation importance based on the macro F1 score. If you don't mind sharing, may I ask what metric did you use to evaluate the importance? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2243689,
          "author_name": "myppka",
          "author_url": "",
          "post_date": "05/03/2023 05:11:39",
          "content": "<p>I used logloss</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2243627,
      "author_name": "ataraxian",
      "author_url": "",
      "post_date": "05/03/2023 03:58:48",
      "content": "<p>May I ask how many features you guys use?  My original model with mass features generation has 1000/3000/6000 features, with some basic feature elimination it still has 400/1000/1000 features. Either case is just too cost-heavy for permutation importance.</p>\n<p>Also correct me if I'm wrong, permutation importance relies on CV to work, but according to my experiences CV and LB does not correlate so well when I do feature elimination. Feature elimination does help LB/CV boost a lot, they just do not get the boost synchronously. For my best LB score, local CV change ±0.0002 after I do basic feature elimination which I would consider as noise if I don't have some prior experience that it should work. Then it dramatically boosts LB by +0.002, which is certainly not noise. In such a case I doubt permutation importance would work with CV since it would also consider ±0.0002 as noise.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2243707,
          "author_name": "myppka",
          "author_url": "",
          "post_date": "05/03/2023 05:31:53",
          "content": "<p>I used 1000/2500/5000 features (no selection) for my best cv-lb results</p>\n<p>Yes, indeed, the data looks very noisy(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2243721,
          "author_name": "woprime",
          "author_url": "",
          "post_date": "05/03/2023 05:39:46",
          "content": "<p>Consider that eliminating feature also changes the way decision trees are grown, the boost in lb maybe is due to randomness. I've noticed +- 0.002 difference in macro F1 for different folds by simply shuffle around column orders.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2243739,
              "author_name": "ataraxian",
              "author_url": "",
              "post_date": "05/03/2023 06:12:34",
              "content": "<p>According to my experience with applying basic feature elimination to different models, it would stably boost LB +0.001 with a ±0.001 randomness. There is a certain pattern which I can't give a proper explanation, like if I choose 1000 features then I always get the best LB score compared to 800 or 1200 features, even if the CV score of those different models show little differences around ±0.0002 randomly which I considered as noise.</p>\n<p>Honestly, I don't have any confidence in my elimination strategy, it most likely causes either leakage/overfitting according to other discussions. I believe the only reason it worked is that 6000 features are too much for a 20k dataset, and reducing it to around 1000 prevents the curse of dimensionality anyway. I checked both public notebooks/discussions and I see no common strategy for feature selection. People do talk about permutation importance, but I doubt it's too expensive to afford. My 400/1000/1000 features model takes 3 hours to run and I see no way I do permutation importance for every single feature. Maybe there is some hidden trick we would only learn when this competition ends.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2244093,
      "author_name": "gehallak",
      "author_url": "",
      "post_date": "05/03/2023 12:39:36",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/myppka\" target=\"_blank\">@myppka</a>,  I tried to keep the most important features, but it was improving my CV and deteriorating my LB score. It is probably overfitting the train data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2244179,
          "author_name": "myppka",
          "author_url": "",
          "post_date": "05/03/2023 13:46:59",
          "content": "<p>Yes, i have the same problem. Just after that i started experimenting with the old test set</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2243295": "Just wondering if anyone has been able to use feature selection based on Permutation Importance.\n\nI've done a few runs using the old training set as the training set and the old test set as the test set, and cv increases by 0.002+, but the quality on the test degrades a bit in most cases.\nI have tried 2 methods: \n- compute feature importances using xgboost and then train xgboost on selected features\n- compute feature importances using random forest and then train xgboost on selected features\n\nI also tried using BoostARoota, but that didn't help in my case either.\n\nMaybe I'm doing something wrong",
    "2243411": "I've been trying to figure this out too. I've computed the permutation importance of the new training set, and selected features accordingly but the CV barely increases. Moreover, the increase in CV score only happens when I remove a few features base on importance ranking, but if I try to remove more than a handful of features with negative importance, the CV score quickly degrades.\n\nI've implemented the permutation importance based on the macro F1 score. If you don't mind sharing, may I ask what metric did you use to evaluate the importance?",
    "2243627": "May I ask how many features you guys use?  My original model with mass features generation has 1000/3000/6000 features, with some basic feature elimination it still has 400/1000/1000 features. Either case is just too cost-heavy for permutation importance.\n\nAlso correct me if I'm wrong, permutation importance relies on CV to work, but according to my experiences CV and LB does not correlate so well when I do feature elimination. Feature elimination does help LB/CV boost a lot, they just do not get the boost synchronously. For my best LB score, local CV change ±0.0002 after I do basic feature elimination which I would consider as noise if I don't have some prior experience that it should work. Then it dramatically boosts LB by +0.002, which is certainly not noise. In such a case I doubt permutation importance would work with CV since it would also consider ±0.0002 as noise.",
    "2243689": "I used logloss",
    "2243707": "I used 1000/2500/5000 features (no selection) for my best cv-lb results\n\nYes, indeed, the data looks very noisy(",
    "2243721": "Consider that eliminating feature also changes the way decision trees are grown, the boost in lb maybe is due to randomness. I've noticed +- 0.002 difference in macro F1 for different folds by simply shuffle around column orders.",
    "2243739": "According to my experience with applying basic feature elimination to different models, it would stably boost LB +0.001 with a ±0.001 randomness. There is a certain pattern which I can't give a proper explanation, like if I choose 1000 features then I always get the best LB score compared to 800 or 1200 features, even if the CV score of those different models show little differences around ±0.0002 randomly which I considered as noise.\n\nHonestly, I don't have any confidence in my elimination strategy, it most likely causes either leakage/overfitting according to other discussions. I believe the only reason it worked is that 6000 features are too much for a 20k dataset, and reducing it to around 1000 prevents the curse of dimensionality anyway. I checked both public notebooks/discussions and I see no common strategy for feature selection. People do talk about permutation importance, but I doubt it's too expensive to afford. My 400/1000/1000 features model takes 3 hours to run and I see no way I do permutation importance for every single feature. Maybe there is some hidden trick we would only learn when this competition ends.",
    "2244093": "Hi @myppka,  I tried to keep the most important features, but it was improving my CV and deteriorating my LB score. It is probably overfitting the train data.",
    "2244179": "Yes, i have the same problem. Just after that i started experimenting with the old test set"
  },
  "source": "meta"
}