{
  "id": 339583,
  "title": "Feature Selection approaches",
  "url": "/competitions/amex-default-prediction/discussion/339583",
  "author_name": "AKR",
  "post_date": "2022-07-25T14:27:54.864000",
  "votes": 14,
  "comment_count": 0,
  "views": 0,
  "content": "<p>So far I have come to following conclusion.</p>\n<ul>\n<li>Models like lgb, xgb, catboost perform well </li>\n<li>Dart booster perform really well in this competition</li>\n<li>Lag features of <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a>  works</li>\n<li>Last - mean features of <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> works</li>\n</ul>\n<p>Using these things we can get pblb score of <code>0.799</code> . <br>\nBut to further increase score one of the important option is to do <strong>FEATURE ENGINEERING</strong>. </p>\n<p>I have created lots of features based on the statistical information/aggregation on different groups of features.  <br>\nNow the question is how to do FEATURE SELECTION, because using all the features at a time will blow up. </p>\n<p>Also I started with the <code>permutation_importance</code> as suggested by <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> in <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131\" target=\"_blank\">this</a> post. <br>\nI have done few experiments and indeed this improves cv as well as pblb score. But it is quite slow and wouldn't be possible to do it on all the features. Even picking one set of features finding perm_features then picking next and finding perm_features and repeating it will not be feasible. (Not atleast in my compute)</p>\n<p>So inspired by <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339071\" target=\"_blank\">this</a> post of <a href=\"https://www.kaggle.com/illidan7\" target=\"_blank\">@illidan7</a> , I tried the second best option i.e. <strong>split_importance</strong>. But in this both cv and LB score decreased for me. </p>\n<p>I think <strong>feature engineering</strong> and <strong>feature selection</strong> are the skills which will help win medal in this competition. <br>\nI will not ask about features people have created because that will not be fair ; ) </p>\n<hr>\n<p>1) What I want to know is the <strong>approach for feature selection</strong>. <br>\n2) Also is statistical and aggregation feature engineering enough for this competition or,<br>\n   Are people able to do some domain based feature engineering which improved the score. </p>\n<p>Apologies if some questions are naive. </p>",
  "messages": [
    {
      "id": 1870410,
      "postDate": "2022-07-25T14:27:54.863Z",
      "content": "<p>So far I have come to following conclusion.</p>\n<ul>\n<li>Models like lgb, xgb, catboost perform well </li>\n<li>Dart booster perform really well in this competition</li>\n<li>Lag features of <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a>  works</li>\n<li>Last - mean features of <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> works</li>\n</ul>\n<p>Using these things we can get pblb score of <code>0.799</code> . <br>\nBut to further increase score one of the important option is to do <strong>FEATURE ENGINEERING</strong>. </p>\n<p>I have created lots of features based on the statistical information/aggregation on different groups of features.  <br>\nNow the question is how to do FEATURE SELECTION, because using all the features at a time will blow up. </p>\n<p>Also I started with the <code>permutation_importance</code> as suggested by <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> in <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131\" target=\"_blank\">this</a> post. <br>\nI have done few experiments and indeed this improves cv as well as pblb score. But it is quite slow and wouldn't be possible to do it on all the features. Even picking one set of features finding perm_features then picking next and finding perm_features and repeating it will not be feasible. (Not atleast in my compute)</p>\n<p>So inspired by <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339071\" target=\"_blank\">this</a> post of <a href=\"https://www.kaggle.com/illidan7\" target=\"_blank\">@illidan7</a> , I tried the second best option i.e. <strong>split_importance</strong>. But in this both cv and LB score decreased for me. </p>\n<p>I think <strong>feature engineering</strong> and <strong>feature selection</strong> are the skills which will help win medal in this competition. <br>\nI will not ask about features people have created because that will not be fair ; ) </p>\n<hr>\n<p>1) What I want to know is the <strong>approach for feature selection</strong>. <br>\n2) Also is statistical and aggregation feature engineering enough for this competition or,<br>\n   Are people able to do some domain based feature engineering which improved the score. </p>\n<p>Apologies if some questions are naive. </p>",
      "rawMarkdown": "So far I have come to following conclusion.\n* Models like lgb, xgb, catboost perform well \n* Dart booster perform really well in this competition\n* Lag features of @thedevastator  works\n* Last - mean features of @ragnar123 works\n\nUsing these things we can get pblb score of `0.799` . \nBut to further increase score one of the important option is to do **FEATURE ENGINEERING**. \n\nI have created lots of features based on the statistical information/aggregation on different groups of features.  \nNow the question is how to do FEATURE SELECTION, because using all the features at a time will blow up. \n\nAlso I started with the `permutation_importance` as suggested by @ambrosm in [this](https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131) post. \nI have done few experiments and indeed this improves cv as well as pblb score. But it is quite slow and wouldn't be possible to do it on all the features. Even picking one set of features finding perm_features then picking next and finding perm_features and repeating it will not be feasible. (Not atleast in my compute)\n\nSo inspired by [this](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339071) post of @illidan7 , I tried the second best option i.e. **split_importance**. But in this both cv and LB score decreased for me. \n\nI think **feature engineering** and **feature selection** are the skills which will help win medal in this competition. \nI will not ask about features people have created because that will not be fair ; ) \n\n******************************************************************************************************************\n1) What I want to know is the **approach for feature selection**. \n2) Also is statistical and aggregation feature engineering enough for this competition or,\n   Are people able to do some domain based feature engineering which improved the score. \n\nApologies if some questions are naive. ",
      "votes": 13
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1870410": "So far I have come to following conclusion.\n* Models like lgb, xgb, catboost perform well \n* Dart booster perform really well in this competition\n* Lag features of @thedevastator  works\n* Last - mean features of @ragnar123 works\n\nUsing these things we can get pblb score of `0.799` . \nBut to further increase score one of the important option is to do **FEATURE ENGINEERING**. \n\nI have created lots of features based on the statistical information/aggregation on different groups of features.  \nNow the question is how to do FEATURE SELECTION, because using all the features at a time will blow up. \n\nAlso I started with the `permutation_importance` as suggested by @ambrosm in [this](https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131) post. \nI have done few experiments and indeed this improves cv as well as pblb score. But it is quite slow and wouldn't be possible to do it on all the features. Even picking one set of features finding perm_features then picking next and finding perm_features and repeating it will not be feasible. (Not atleast in my compute)\n\nSo inspired by [this](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339071) post of @illidan7 , I tried the second best option i.e. **split_importance**. But in this both cv and LB score decreased for me. \n\nI think **feature engineering** and **feature selection** are the skills which will help win medal in this competition. \nI will not ask about features people have created because that will not be fair ; ) \n\n******************************************************************************************************************\n1) What I want to know is the **approach for feature selection**. \n2) Also is statistical and aggregation feature engineering enough for this competition or,\n   Are people able to do some domain based feature engineering which improved the score. \n\nApologies if some questions are naive. "
  }
}