{
  "id": 330486,
  "title": "Some useful information I found about Features and Feature groups ",
  "url": "/competitions/amex-default-prediction/discussion/330486",
  "author_name": "",
  "post_date": "2022-06-12T16:47:00.374586Z",
  "votes": 56,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I was looking at various notebooks(most of them were based on boosting models) and I found that the different models are using different features(or feature groups) for splits, even the slightest changes in hyper-parameter were affecting the <code>feature_importance</code> list. Because of this, it is not only very hard to find the best features which we can carry on to our next models but it is also very hard to find and drop the irrelevant features. So I trained 100 XGBoost Models(each with slightly different hyper-parameters) on the same set of train and validation data to validate each feature's importance.</p>\n<ol>\n<li>First I made a <a href=\"https://www.kaggle.com/code/susnato/amex-data-preprocesing-feature-engineering/notebook?scriptVersionId=98169754\" target=\"_blank\">notebook</a> with a bunch of Feature Preprocessing techniques applied to the total data and then split the data into 5 folds.</li>\n<li>Then I used Weights and Biases Hyperparameter Sweep to train 100 models on different hyper-parameters(but used the same train and validation splits).</li>\n<li>After getting those feature_importance DataFrames from each of the models I aggregated the feature importance over all models.</li>\n</ol>\n<p>The results are here,<br>\n<strong>This plot shows the Top 100 features</strong> : <br>\n<img src=\"https://drive.google.com/file/d/18kaNHJZ0f8k_HKoB1emFkSJxKFXST2-Y/view?usp=sharing\" alt=\"\"></p>\n<p><strong>This plot shows feature importance of different feature groups</strong> : <br>\n<img src=\"https://drive.google.com/file/d/1f30qc3hX0_sxmWDfdSDaXPOoYNDXgnPl/view?usp=sharing\" alt=\"\"></p>\n<p>For the Feature Groups, the group <code>last</code> turned out to be most useful though <code>last2</code> or <code>last3</code> groups didn't affect that much.</p>\n<p>For anyone who is interested to look at the feature importance files click <a href=\"https://www.kaggle.com/datasets/susnato/amexxgboostfeature-importances/settings\" target=\"_blank\">here</a>.<br>\nFor anyone who wants to see the Notebook used to create these features and how these plots were created click <a href=\"https://www.kaggle.com/code/susnato/amex-data-preprocesing-feature-engineering/notebook?scriptVersionId=98169754\" target=\"_blank\">here</a>.</p>",
  "messages": [
    {
      "id": "1818409",
      "postDate": "06/12/2022 16:47:00",
      "content": "<p>I was looking at various notebooks(most of them were based on boosting models) and I found that the different models are using different features(or feature groups) for splits, even the slightest changes in hyper-parameter were affecting the <code>feature_importance</code> list. Because of this, it is not only very hard to find the best features which we can carry on to our next models but it is also very hard to find and drop the irrelevant features. So I trained 100 XGBoost Models(each with slightly different hyper-parameters) on the same set of train and validation data to validate each feature's importance.</p>\n<ol>\n<li>First I made a <a href=\"https://www.kaggle.com/code/susnato/amex-data-preprocesing-feature-engineering/notebook?scriptVersionId=98169754\" target=\"_blank\">notebook</a> with a bunch of Feature Preprocessing techniques applied to the total data and then split the data into 5 folds.</li>\n<li>Then I used Weights and Biases Hyperparameter Sweep to train 100 models on different hyper-parameters(but used the same train and validation splits).</li>\n<li>After getting those feature_importance DataFrames from each of the models I aggregated the feature importance over all models.</li>\n</ol>\n<p>The results are here,<br>\n<strong>This plot shows the Top 100 features</strong> : <br>\n<img src=\"https://drive.google.com/file/d/18kaNHJZ0f8k_HKoB1emFkSJxKFXST2-Y/view?usp=sharing\" alt=\"\"></p>\n<p><strong>This plot shows feature importance of different feature groups</strong> : <br>\n<img src=\"https://drive.google.com/file/d/1f30qc3hX0_sxmWDfdSDaXPOoYNDXgnPl/view?usp=sharing\" alt=\"\"></p>\n<p>For the Feature Groups, the group <code>last</code> turned out to be most useful though <code>last2</code> or <code>last3</code> groups didn't affect that much.</p>\n<p>For anyone who is interested to look at the feature importance files click <a href=\"https://www.kaggle.com/datasets/susnato/amexxgboostfeature-importances/settings\" target=\"_blank\">here</a>.<br>\nFor anyone who wants to see the Notebook used to create these features and how these plots were created click <a href=\"https://www.kaggle.com/code/susnato/amex-data-preprocesing-feature-engineering/notebook?scriptVersionId=98169754\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "I was looking at various notebooks(most of them were based on boosting models) and I found that the different models are using different features(or feature groups) for splits, even the slightest changes in hyper-parameter were affecting the `feature_importance` list. Because of this, it is not only very hard to find the best features which we can carry on to our next models but it is also very hard to find and drop the irrelevant features. So I trained 100 XGBoost Models(each with slightly different hyper-parameters) on the same set of train and validation data to validate each feature's importance.\n\n1. First I made a [notebook](https://www.kaggle.com/code/susnato/amex-data-preprocesing-feature-engineering/notebook?scriptVersionId=98169754) with a bunch of Feature Preprocessing techniques applied to the total data and then split the data into 5 folds.\n2. Then I used Weights and Biases Hyperparameter Sweep to train 100 models on different hyper-parameters(but used the same train and validation splits).\n3. After getting those feature_importance DataFrames from each of the models I aggregated the feature importance over all models.\n\nThe results are here,\n**This plot shows the Top 100 features** : \n![](https://drive.google.com/file/d/18kaNHJZ0f8k_HKoB1emFkSJxKFXST2-Y/view?usp=sharing)\n\n**This plot shows feature importance of different feature groups** : \n![](https://drive.google.com/file/d/1f30qc3hX0_sxmWDfdSDaXPOoYNDXgnPl/view?usp=sharing)\n\nFor the Feature Groups, the group `last` turned out to be most useful though `last2` or `last3` groups didn't affect that much.\n\nFor anyone who is interested to look at the feature importance files click [here](https://www.kaggle.com/datasets/susnato/amexxgboostfeature-importances/settings).\nFor anyone who wants to see the Notebook used to create these features and how these plots were created click [here](https://www.kaggle.com/code/susnato/amex-data-preprocesing-feature-engineering/notebook?scriptVersionId=98169754).",
      "votes": null
    },
    {
      "id": "1818436",
      "postDate": "06/12/2022 17:28:34",
      "content": "<p>Really cool experiment! I didn't think about training many models with different hyperparams on the same splits to understand the feature importance. this is creative! <br>\nI would suggest trying to do the same but with a different seed for the KFold so you can be sure you don't overfit the exact splits you try. </p>\n<p>Anyway, really cool! </p>",
      "rawMarkdown": "Really cool experiment! I didn't think about training many models with different hyperparams on the same splits to understand the feature importance. this is creative! \nI would suggest trying to do the same but with a different seed for the KFold so you can be sure you don't overfit the exact splits you try. \n\nAnyway, really cool!",
      "votes": null
    },
    {
      "id": "1818995",
      "postDate": "06/13/2022 10:47:56",
      "content": "<p>Really interesting, Great job! Have you had any insights about the noise in the data? <br>\nCheers</p>",
      "rawMarkdown": "Really interesting, Great job! Have you had any insights about the noise in the data? \nCheers",
      "votes": null
    },
    {
      "id": "1819138",
      "postDate": "06/13/2022 13:28:55",
      "content": "<p><a href=\"https://www.kaggle.com/susnato\" target=\"_blank\">@susnato</a> Thank you for sharing valuable insights. It helps</p>",
      "rawMarkdown": "susnato Thank you for sharing valuable insights. It helps",
      "votes": null
    },
    {
      "id": "1821009",
      "postDate": "06/15/2022 07:16:12",
      "content": "<p><a href=\"https://www.kaggle.com/susnato\" target=\"_blank\">@susnato</a> Thanks for sharing </p>",
      "rawMarkdown": "susnato Thanks for sharing",
      "votes": null
    },
    {
      "id": "1821109",
      "postDate": "06/15/2022 08:58:09",
      "content": "<p>Great idea! Thanks</p>",
      "rawMarkdown": "Great idea! Thanks",
      "votes": null
    },
    {
      "id": "1821362",
      "postDate": "06/15/2022 13:24:34",
      "content": "<p><a href=\"https://www.kaggle.com/susnato\" target=\"_blank\">@susnato</a> Good idea, this sound like some kind of Monte Carlo simulation, It is true that a small change in the hyperparameters list yield to changes in the features' importance.</p>",
      "rawMarkdown": "susnato Good idea, this sound like some kind of Monte Carlo simulation, It is true that a small change in the hyperparameters list yield to changes in the features' importance.",
      "votes": null
    },
    {
      "id": "1821408",
      "postDate": "06/15/2022 14:08:30",
      "content": "<p>Very Helpful 👍 Thanks for Sharing :)</p>",
      "rawMarkdown": "Very Helpful 👍 Thanks for Sharing :)",
      "votes": null
    },
    {
      "id": "1821576",
      "postDate": "06/15/2022 17:36:37",
      "content": "<p>Good idea. I also suggest digging into Autoencoders (even Denosing AE's) to find features/combinations easier.</p>",
      "rawMarkdown": "Good idea. I also suggest digging into Autoencoders (even Denosing AE's) to find features/combinations easier.",
      "votes": null
    },
    {
      "id": "1838960",
      "postDate": "07/01/2022 02:57:39",
      "content": "<p>Thanks for sharing your notebooks! </p>",
      "rawMarkdown": "Thanks for sharing your notebooks!",
      "votes": null
    },
    {
      "id": "1873163",
      "postDate": "07/27/2022 13:34:19",
      "content": "<p>Thanks for your great works!</p>",
      "rawMarkdown": "Thanks for your great works!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1818436,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "06/12/2022 17:28:34",
      "content": "<p>Really cool experiment! I didn't think about training many models with different hyperparams on the same splits to understand the feature importance. this is creative! <br>\nI would suggest trying to do the same but with a different seed for the KFold so you can be sure you don't overfit the exact splits you try. </p>\n<p>Anyway, really cool! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1818995,
      "author_name": "paulojunqueira",
      "author_url": "",
      "post_date": "06/13/2022 10:47:56",
      "content": "<p>Really interesting, Great job! Have you had any insights about the noise in the data? <br>\nCheers</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1819138,
      "author_name": "thamotharan",
      "author_url": "",
      "post_date": "06/13/2022 13:28:55",
      "content": "<p><a href=\"https://www.kaggle.com/susnato\" target=\"_blank\">@susnato</a> Thank you for sharing valuable insights. It helps</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1821009,
      "author_name": "samanemami",
      "author_url": "",
      "post_date": "06/15/2022 07:16:12",
      "content": "<p><a href=\"https://www.kaggle.com/susnato\" target=\"_blank\">@susnato</a> Thanks for sharing </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1821109,
      "author_name": "ashlockkerplord",
      "author_url": "",
      "post_date": "06/15/2022 08:58:09",
      "content": "<p>Great idea! Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1821362,
      "author_name": "omarjarir",
      "author_url": "",
      "post_date": "06/15/2022 13:24:34",
      "content": "<p><a href=\"https://www.kaggle.com/susnato\" target=\"_blank\">@susnato</a> Good idea, this sound like some kind of Monte Carlo simulation, It is true that a small change in the hyperparameters list yield to changes in the features' importance.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1821408,
      "author_name": "gaurav126",
      "author_url": "",
      "post_date": "06/15/2022 14:08:30",
      "content": "<p>Very Helpful 👍 Thanks for Sharing :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1821576,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "06/15/2022 17:36:37",
      "content": "<p>Good idea. I also suggest digging into Autoencoders (even Denosing AE's) to find features/combinations easier.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1838960,
      "author_name": "mohammadrahmati",
      "author_url": "",
      "post_date": "07/01/2022 02:57:39",
      "content": "<p>Thanks for sharing your notebooks! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1873163,
      "author_name": "ringonewton",
      "author_url": "",
      "post_date": "07/27/2022 13:34:19",
      "content": "<p>Thanks for your great works!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1818409": "I was looking at various notebooks(most of them were based on boosting models) and I found that the different models are using different features(or feature groups) for splits, even the slightest changes in hyper-parameter were affecting the `feature_importance` list. Because of this, it is not only very hard to find the best features which we can carry on to our next models but it is also very hard to find and drop the irrelevant features. So I trained 100 XGBoost Models(each with slightly different hyper-parameters) on the same set of train and validation data to validate each feature's importance.\n\n1. First I made a [notebook](https://www.kaggle.com/code/susnato/amex-data-preprocesing-feature-engineering/notebook?scriptVersionId=98169754) with a bunch of Feature Preprocessing techniques applied to the total data and then split the data into 5 folds.\n2. Then I used Weights and Biases Hyperparameter Sweep to train 100 models on different hyper-parameters(but used the same train and validation splits).\n3. After getting those feature_importance DataFrames from each of the models I aggregated the feature importance over all models.\n\nThe results are here,\n**This plot shows the Top 100 features** : \n![](https://drive.google.com/file/d/18kaNHJZ0f8k_HKoB1emFkSJxKFXST2-Y/view?usp=sharing)\n\n**This plot shows feature importance of different feature groups** : \n![](https://drive.google.com/file/d/1f30qc3hX0_sxmWDfdSDaXPOoYNDXgnPl/view?usp=sharing)\n\nFor the Feature Groups, the group `last` turned out to be most useful though `last2` or `last3` groups didn't affect that much.\n\nFor anyone who is interested to look at the feature importance files click [here](https://www.kaggle.com/datasets/susnato/amexxgboostfeature-importances/settings).\nFor anyone who wants to see the Notebook used to create these features and how these plots were created click [here](https://www.kaggle.com/code/susnato/amex-data-preprocesing-feature-engineering/notebook?scriptVersionId=98169754).",
    "1818436": "Really cool experiment! I didn't think about training many models with different hyperparams on the same splits to understand the feature importance. this is creative! \nI would suggest trying to do the same but with a different seed for the KFold so you can be sure you don't overfit the exact splits you try. \n\nAnyway, really cool!",
    "1818995": "Really interesting, Great job! Have you had any insights about the noise in the data? \nCheers",
    "1819138": "susnato Thank you for sharing valuable insights. It helps",
    "1821009": "susnato Thanks for sharing",
    "1821109": "Great idea! Thanks",
    "1821362": "susnato Good idea, this sound like some kind of Monte Carlo simulation, It is true that a small change in the hyperparameters list yield to changes in the features' importance.",
    "1821408": "Very Helpful 👍 Thanks for Sharing :)",
    "1821576": "Good idea. I also suggest digging into Autoencoders (even Denosing AE's) to find features/combinations easier.",
    "1838960": "Thanks for sharing your notebooks!",
    "1873163": "Thanks for your great works!"
  },
  "source": "meta"
}