{
  "id": 491286,
  "title": "How do you pick the important features",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/491286",
  "author_name": "",
  "post_date": "2024-04-05T10:14:57.904900900Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>In addition to dimensionality reduction, I think screening suitable features is an important aspect of enhancing robustness. My model has a relatively serious overfitting problem and I want to solve this problem by removing redundant features, but I don't know how to filter, according to the importance provided by the model?</p>",
  "messages": [
    {
      "id": "2736642",
      "postDate": "04/05/2024 10:14:57",
      "content": "<p>In addition to dimensionality reduction, I think screening suitable features is an important aspect of enhancing robustness. My model has a relatively serious overfitting problem and I want to solve this problem by removing redundant features, but I don't know how to filter, according to the importance provided by the model?</p>",
      "rawMarkdown": "In addition to dimensionality reduction, I think screening suitable features is an important aspect of enhancing robustness. My model has a relatively serious overfitting problem and I want to solve this problem by removing redundant features, but I don't know how to filter, according to the importance provided by the model?",
      "votes": null
    },
    {
      "id": "2737374",
      "postDate": "04/05/2024 18:26:11",
      "content": "<p>one way is to remove the features that had low importance. So u train your model, get the feature importance, get the list of features that are below certain threshold or some other criteria. remove those features from train and test, everywhere. then train the model again, then check if this removal helped or not. if ur train data frame is pandas u can use pop to remove. for polars you can del or filter. or you can try to select the important ones leaving the non important ones.</p>",
      "rawMarkdown": "one way is to remove the features that had low importance. So u train your model, get the feature importance, get the list of features that are below certain threshold or some other criteria. remove those features from train and test, everywhere. then train the model again, then check if this removal helped or not. if ur train data frame is pandas u can use pop to remove. for polars you can del or filter. or you can try to select the important ones leaving the non important ones.",
      "votes": null
    },
    {
      "id": "2738903",
      "postDate": "04/06/2024 17:32:59",
      "content": "<p>May be you can costruct model with one feature and if AUC is small you can remove feature</p>",
      "rawMarkdown": "May be you can costruct model with one feature and if AUC is small you can remove feature",
      "votes": null
    },
    {
      "id": "2738904",
      "postDate": "04/06/2024 17:35:15",
      "content": "<p>You can also read about: WEIGHT OF EVIDENCE (WOE) AND INFORMATION VALUE (IV). If IV is small you can remove feature</p>",
      "rawMarkdown": "You can also read about: WEIGHT OF EVIDENCE (WOE) AND INFORMATION VALUE (IV). If IV is small you can remove feature",
      "votes": null
    },
    {
      "id": "2748354",
      "postDate": "04/12/2024 12:04:10",
      "content": "<p>Congrats on 7th place tho <a href=\"https://www.kaggle.com/zivanwan\" target=\"_blank\">@zivanwan</a> !!</p>",
      "rawMarkdown": "Congrats on 7th place tho @zivanwan !!",
      "votes": null
    },
    {
      "id": "2749188",
      "postDate": "04/12/2024 22:29:13",
      "content": "<p>One way is to train <code>lasso regression</code> (i.e. add L1 regularization) and see which features it chooses. Another way is RFE (recursive feature elimination), we make a <code>for-loop</code> and remove 1 feature at a time to see if CV improves. If it does, we remove the feature and continue the <code>for-loop</code> to find other features to remove. A third way is permutation importance. We continually infer the validation in a <code>for-loop</code>, each time we randomly shuffle a feature and observe the effect on validation score.</p>",
      "rawMarkdown": "One way is to train `lasso regression` (i.e. add L1 regularization) and see which features it chooses. Another way is RFE (recursive feature elimination), we make a `for-loop` and remove 1 feature at a time to see if CV improves. If it does, we remove the feature and continue the `for-loop` to find other features to remove. A third way is permutation importance. We continually infer the validation in a `for-loop`, each time we randomly shuffle a feature and observe the effect on validation score.",
      "votes": null
    },
    {
      "id": "2758009",
      "postDate": "04/17/2024 20:50:46",
      "content": "<p>Hi Chris, Thank you for your helpful suggestions. <br>\nI'm new to lgbm and have a related question: we have different features to select, and we can do it by cross validation, but we also have different hyper parameters to set in lgbm (or other models), and we can also set hyper parameters by cross validation. <br>\nIn your projects, how do you balance between these two? Do we first select features with a fixed lgbm hyperparameters and then choose best hyper parameters, or are there other recommended methods? Or maybe we can first select features with simple methods (such as correlation value), and then tune the hyper parameters with fixed features? In practice, which will work better?<br>\nThank you!</p>",
      "rawMarkdown": "Hi Chris, Thank you for your helpful suggestions. \n\nI'm new to lgbm and have a related question: we have different features to select, and we can do it by cross validation, but we also have different hyper parameters to set in lgbm (or other models), and we can also set hyper parameters by cross validation. \nIn your projects, how do you balance between these two? Do we first select features with a fixed lgbm hyperparameters and then choose best hyper parameters, or are there other recommended methods? Or maybe we can first select features with simple methods (such as correlation value), and then tune the hyper parameters with fixed features? In practice, which will work better?\nThank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2737374,
      "author_name": "shreyas9181",
      "author_url": "",
      "post_date": "04/05/2024 18:26:11",
      "content": "<p>one way is to remove the features that had low importance. So u train your model, get the feature importance, get the list of features that are below certain threshold or some other criteria. remove those features from train and test, everywhere. then train the model again, then check if this removal helped or not. if ur train data frame is pandas u can use pop to remove. for polars you can del or filter. or you can try to select the important ones leaving the non important ones.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2738903,
      "author_name": "rinatbayanov",
      "author_url": "",
      "post_date": "04/06/2024 17:32:59",
      "content": "<p>May be you can costruct model with one feature and if AUC is small you can remove feature</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2738904,
      "author_name": "rinatbayanov",
      "author_url": "",
      "post_date": "04/06/2024 17:35:15",
      "content": "<p>You can also read about: WEIGHT OF EVIDENCE (WOE) AND INFORMATION VALUE (IV). If IV is small you can remove feature</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2748354,
      "author_name": "luciferisback",
      "author_url": "",
      "post_date": "04/12/2024 12:04:10",
      "content": "<p>Congrats on 7th place tho <a href=\"https://www.kaggle.com/zivanwan\" target=\"_blank\">@zivanwan</a> !!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2749188,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "04/12/2024 22:29:13",
      "content": "<p>One way is to train <code>lasso regression</code> (i.e. add L1 regularization) and see which features it chooses. Another way is RFE (recursive feature elimination), we make a <code>for-loop</code> and remove 1 feature at a time to see if CV improves. If it does, we remove the feature and continue the <code>for-loop</code> to find other features to remove. A third way is permutation importance. We continually infer the validation in a <code>for-loop</code>, each time we randomly shuffle a feature and observe the effect on validation score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2758009,
          "author_name": "ivyzang",
          "author_url": "",
          "post_date": "04/17/2024 20:50:46",
          "content": "<p>Hi Chris, Thank you for your helpful suggestions. <br>\nI'm new to lgbm and have a related question: we have different features to select, and we can do it by cross validation, but we also have different hyper parameters to set in lgbm (or other models), and we can also set hyper parameters by cross validation. <br>\nIn your projects, how do you balance between these two? Do we first select features with a fixed lgbm hyperparameters and then choose best hyper parameters, or are there other recommended methods? Or maybe we can first select features with simple methods (such as correlation value), and then tune the hyper parameters with fixed features? In practice, which will work better?<br>\nThank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2736642": "In addition to dimensionality reduction, I think screening suitable features is an important aspect of enhancing robustness. My model has a relatively serious overfitting problem and I want to solve this problem by removing redundant features, but I don't know how to filter, according to the importance provided by the model?",
    "2737374": "one way is to remove the features that had low importance. So u train your model, get the feature importance, get the list of features that are below certain threshold or some other criteria. remove those features from train and test, everywhere. then train the model again, then check if this removal helped or not. if ur train data frame is pandas u can use pop to remove. for polars you can del or filter. or you can try to select the important ones leaving the non important ones.",
    "2738903": "May be you can costruct model with one feature and if AUC is small you can remove feature",
    "2738904": "You can also read about: WEIGHT OF EVIDENCE (WOE) AND INFORMATION VALUE (IV). If IV is small you can remove feature",
    "2748354": "Congrats on 7th place tho @zivanwan !!",
    "2749188": "One way is to train `lasso regression` (i.e. add L1 regularization) and see which features it chooses. Another way is RFE (recursive feature elimination), we make a `for-loop` and remove 1 feature at a time to see if CV improves. If it does, we remove the feature and continue the `for-loop` to find other features to remove. A third way is permutation importance. We continually infer the validation in a `for-loop`, each time we randomly shuffle a feature and observe the effect on validation score.",
    "2758009": "Hi Chris, Thank you for your helpful suggestions. \n\nI'm new to lgbm and have a related question: we have different features to select, and we can do it by cross validation, but we also have different hyper parameters to set in lgbm (or other models), and we can also set hyper parameters by cross validation. \nIn your projects, how do you balance between these two? Do we first select features with a fixed lgbm hyperparameters and then choose best hyper parameters, or are there other recommended methods? Or maybe we can first select features with simple methods (such as correlation value), and then tune the hyper parameters with fixed features? In practice, which will work better?\nThank you!"
  },
  "source": "meta"
}