{
  "id": 343191,
  "title": "🥢🪝Selected Features Dataset🪝🥢",
  "url": "/competitions/amex-default-prediction/discussion/343191",
  "author_name": "",
  "post_date": "2022-08-10T10:45:00.350999400Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello Everyone,<br>\nIn the <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/342798\" target=\"_blank\">previous post</a> I  share with you an aggregated feature engineered dataset (that you can find <a href=\"https://www.kaggle.com/datasets/schopenhacker75/feature-eng-agg?sort=recent-comments\" target=\"_blank\">here</a>)</p>\n<p>In this post I will share with you a dataset extracted from feature selection applied on the previous dataset :</p>\n<p>You can find it here : <a href=\"https://www.kaggle.com/datasets/schopenhacker75/amexfeatureselected\" target=\"_blank\">https://www.kaggle.com/datasets/schopenhacker75/amexfeatureselected</a></p>\n<h3>1 - Feature engineering</h3>\n<p>I applied on <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">raddars's dataset</a> some feature engineering functions described in <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/342798\" target=\"_blank\">this post</a></p>\n<h3>2 - Feature Selection</h3>\n<p>I selected TOP 200 features from the generated dataset using <strong>filter based techniques</strong> because this method is faster and less computationally expensive than other feature selection methods such as wrapper methods.</p>\n<p><strong>2.1. Drop Top missing values features</strong><br>\nAll features having more than 75% of missing values are ommited</p>\n<p><strong>2.2. Select top correlated features with the target</strong><br>\nWith this method we assume that high predictive features are highly correlated with the target</p>\n<p>🤩🤩<strong>[BONUS] : A fancy EDA notebook with plotly</strong>: <a href=\"https://www.kaggle.com/code/schopenhacker75/fancy-complete-eda🤩🤩\" target=\"_blank\">https://www.kaggle.com/code/schopenhacker75/fancy-complete-eda🤩🤩</a></p>",
  "messages": [
    {
      "id": "1892799",
      "postDate": "08/10/2022 10:45:00",
      "content": "<p>Hello Everyone,<br>\nIn the <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/342798\" target=\"_blank\">previous post</a> I  share with you an aggregated feature engineered dataset (that you can find <a href=\"https://www.kaggle.com/datasets/schopenhacker75/feature-eng-agg?sort=recent-comments\" target=\"_blank\">here</a>)</p>\n<p>In this post I will share with you a dataset extracted from feature selection applied on the previous dataset :</p>\n<p>You can find it here : <a href=\"https://www.kaggle.com/datasets/schopenhacker75/amexfeatureselected\" target=\"_blank\">https://www.kaggle.com/datasets/schopenhacker75/amexfeatureselected</a></p>\n<h3>1 - Feature engineering</h3>\n<p>I applied on <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">raddars's dataset</a> some feature engineering functions described in <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/342798\" target=\"_blank\">this post</a></p>\n<h3>2 - Feature Selection</h3>\n<p>I selected TOP 200 features from the generated dataset using <strong>filter based techniques</strong> because this method is faster and less computationally expensive than other feature selection methods such as wrapper methods.</p>\n<p><strong>2.1. Drop Top missing values features</strong><br>\nAll features having more than 75% of missing values are ommited</p>\n<p><strong>2.2. Select top correlated features with the target</strong><br>\nWith this method we assume that high predictive features are highly correlated with the target</p>\n<p>🤩🤩<strong>[BONUS] : A fancy EDA notebook with plotly</strong>: <a href=\"https://www.kaggle.com/code/schopenhacker75/fancy-complete-eda🤩🤩\" target=\"_blank\">https://www.kaggle.com/code/schopenhacker75/fancy-complete-eda🤩🤩</a></p>",
      "rawMarkdown": "Hello Everyone,\nIn the [previous post](https://www.kaggle.com/competitions/amex-default-prediction/discussion/342798) I  share with you an aggregated feature engineered dataset (that you can find [here](https://www.kaggle.com/datasets/schopenhacker75/feature-eng-agg?sort=recent-comments))\n\nIn this post I will share with you a dataset extracted from feature selection applied on the previous dataset :\n\nYou can find it here : https://www.kaggle.com/datasets/schopenhacker75/amexfeatureselected\n### 1 - Feature engineering\n I applied on [raddars's dataset](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format) some feature engineering functions described in [this post](https://www.kaggle.com/competitions/amex-default-prediction/discussion/342798)\n\n### 2 - Feature Selection \nI selected TOP 200 features from the generated dataset using **filter based techniques** because this method is faster and less computationally expensive than other feature selection methods such as wrapper methods.\n\n\n**2.1. Drop Top missing values features**\nAll features having more than 75% of missing values are ommited\n\n\n**2.2. Select top correlated features with the target**\nWith this method we assume that high predictive features are highly correlated with the target\n\n🤩🤩**[BONUS] : A fancy EDA notebook with plotly**: https://www.kaggle.com/code/schopenhacker75/fancy-complete-eda🤩🤩",
      "votes": null
    },
    {
      "id": "1896038",
      "postDate": "08/12/2022 14:38:53",
      "content": "<p>Thanks for sharing your work on feature selection!<br>\nI tried your dataset with my pipeline. It resulted in 0.783 on LB, while the same pipeline (and number of folds) with 1157 features (that I originally use) results in 0.800.</p>",
      "rawMarkdown": "Thanks for sharing your work on feature selection!\nI tried your dataset with my pipeline. It resulted in 0.783 on LB, while the same pipeline (and number of folds) with 1157 features (that I originally use) results in 0.800.",
      "votes": null
    },
    {
      "id": "1896175",
      "postDate": "08/12/2022 16:23:17",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/rasoulmojtahedzadeh\" target=\"_blank\">@rasoulmojtahedzadeh</a> thx for your feedback 😊!</p>\n<ol>\n<li>Actually I think that the comparison would be relevant if you compare with <strong>the same dataset</strong>(with/without feature selection). With lightgbm it resulted almost the same score for me (0.78)</li>\n<li>The correlation based feature selection is surely not the most efficient method however it is the fastest and the only method that worked for me through a kaggle notebook </li>\n<li>If you have the link to the dataset that produces 0.8 with LB,  thx to share it!! I am actually looking forward enhancing the actual dateset (and eventually share it)</li>\n</ol>",
      "rawMarkdown": "Hello @rasoulmojtahedzadeh thx for your feedback 😊!\n1. Actually I think that the comparison would be relevant if you compare with **the same dataset**(with/without feature selection). With lightgbm it resulted almost the same score for me (0.78)\n2. The correlation based feature selection is surely not the most efficient method however it is the fastest and the only method that worked for me through a kaggle notebook \n3. If you have the link to the dataset that produces 0.8 with LB,  thx to share it!! I am actually looking forward enhancing the actual dateset (and eventually share it)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1896038,
      "author_name": "rasoulmojtahedzadeh",
      "author_url": "",
      "post_date": "08/12/2022 14:38:53",
      "content": "<p>Thanks for sharing your work on feature selection!<br>\nI tried your dataset with my pipeline. It resulted in 0.783 on LB, while the same pipeline (and number of folds) with 1157 features (that I originally use) results in 0.800.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1896175,
          "author_name": "schopenhacker75",
          "author_url": "",
          "post_date": "08/12/2022 16:23:17",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/rasoulmojtahedzadeh\" target=\"_blank\">@rasoulmojtahedzadeh</a> thx for your feedback 😊!</p>\n<ol>\n<li>Actually I think that the comparison would be relevant if you compare with <strong>the same dataset</strong>(with/without feature selection). With lightgbm it resulted almost the same score for me (0.78)</li>\n<li>The correlation based feature selection is surely not the most efficient method however it is the fastest and the only method that worked for me through a kaggle notebook </li>\n<li>If you have the link to the dataset that produces 0.8 with LB,  thx to share it!! I am actually looking forward enhancing the actual dateset (and eventually share it)</li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1892799": "Hello Everyone,\nIn the [previous post](https://www.kaggle.com/competitions/amex-default-prediction/discussion/342798) I  share with you an aggregated feature engineered dataset (that you can find [here](https://www.kaggle.com/datasets/schopenhacker75/feature-eng-agg?sort=recent-comments))\n\nIn this post I will share with you a dataset extracted from feature selection applied on the previous dataset :\n\nYou can find it here : https://www.kaggle.com/datasets/schopenhacker75/amexfeatureselected\n### 1 - Feature engineering\n I applied on [raddars's dataset](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format) some feature engineering functions described in [this post](https://www.kaggle.com/competitions/amex-default-prediction/discussion/342798)\n\n### 2 - Feature Selection \nI selected TOP 200 features from the generated dataset using **filter based techniques** because this method is faster and less computationally expensive than other feature selection methods such as wrapper methods.\n\n\n**2.1. Drop Top missing values features**\nAll features having more than 75% of missing values are ommited\n\n\n**2.2. Select top correlated features with the target**\nWith this method we assume that high predictive features are highly correlated with the target\n\n🤩🤩**[BONUS] : A fancy EDA notebook with plotly**: https://www.kaggle.com/code/schopenhacker75/fancy-complete-eda🤩🤩",
    "1896038": "Thanks for sharing your work on feature selection!\nI tried your dataset with my pipeline. It resulted in 0.783 on LB, while the same pipeline (and number of folds) with 1157 features (that I originally use) results in 0.800.",
    "1896175": "Hello @rasoulmojtahedzadeh thx for your feedback 😊!\n1. Actually I think that the comparison would be relevant if you compare with **the same dataset**(with/without feature selection). With lightgbm it resulted almost the same score for me (0.78)\n2. The correlation based feature selection is surely not the most efficient method however it is the fastest and the only method that worked for me through a kaggle notebook \n3. If you have the link to the dataset that produces 0.8 with LB,  thx to share it!! I am actually looking forward enhancing the actual dateset (and eventually share it)"
  },
  "source": "meta"
}