{
  "id": 336557,
  "title": "Basic Feature Engineering - 1500 features",
  "url": "/competitions/amex-default-prediction/discussion/336557",
  "author_name": "",
  "post_date": "2022-07-11T20:34:44.001424600Z",
  "votes": 42,
  "comment_count": 20,
  "views": 0,
  "content": "<p>There have been a lot of good feature engineering notebooks published for this competition so far. I have put together some basic feature engineering in the notebook below</p>\n<p><a href=\"https://www.kaggle.com/code/illidan7/amex-basic-feature-engineering-1500-features\" target=\"_blank\">https://www.kaggle.com/code/illidan7/amex-basic-feature-engineering-1500-features</a></p>\n<p>If you would rather jump directly into the hell of feature selection, I have saved the result of this notebook as a dataset</p>\n<p><a href=\"https://www.kaggle.com/datasets/illidan7/amexfeatureeng\" target=\"_blank\">https://www.kaggle.com/datasets/illidan7/amexfeatureeng</a></p>\n<p>A lot of the features are based on what others have shared and I have tried to add in some of my own as well</p>\n<ul>\n<li><p>Date based features (Column S_2)</p></li>\n<li><p>\"After pay\" features (<a href=\"https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb\" target=\"_blank\">https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb</a>)</p></li>\n<li><p>Null columns handling </p>\n<ul>\n<li>&gt;30% null; Count num nulls</li>\n<li>&gt;90% null; keep only last</li></ul></li>\n<li><p>Categorical features</p>\n<ul>\n<li>cat1 features (The categorical features mentioned in the competition data page) <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/data\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/data</a></li>\n<li>cat2 features (Low cardinality features; &lt;=4 unique values)</li>\n<li>cat3 features (Low cardinality features; &gt;=8 and &lt;=21 unique values)</li></ul></li>\n<li><p>Last - First (<a href=\"https://www.kaggle.com/code/thedevastator/lag-features-are-all-you-need\" target=\"_blank\">https://www.kaggle.com/code/thedevastator/lag-features-are-all-you-need</a>)</p></li>\n<li><p>Last - mean features (<a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977</a>)</p></li>\n</ul>\n<p>Credits:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/333940\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/333940</a></li>\n<li><a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format</a></li>\n</ul>",
  "messages": [
    {
      "id": "1852151",
      "postDate": "07/11/2022 20:34:44",
      "content": "<p>There have been a lot of good feature engineering notebooks published for this competition so far. I have put together some basic feature engineering in the notebook below</p>\n<p><a href=\"https://www.kaggle.com/code/illidan7/amex-basic-feature-engineering-1500-features\" target=\"_blank\">https://www.kaggle.com/code/illidan7/amex-basic-feature-engineering-1500-features</a></p>\n<p>If you would rather jump directly into the hell of feature selection, I have saved the result of this notebook as a dataset</p>\n<p><a href=\"https://www.kaggle.com/datasets/illidan7/amexfeatureeng\" target=\"_blank\">https://www.kaggle.com/datasets/illidan7/amexfeatureeng</a></p>\n<p>A lot of the features are based on what others have shared and I have tried to add in some of my own as well</p>\n<ul>\n<li><p>Date based features (Column S_2)</p></li>\n<li><p>\"After pay\" features (<a href=\"https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb\" target=\"_blank\">https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb</a>)</p></li>\n<li><p>Null columns handling </p>\n<ul>\n<li>&gt;30% null; Count num nulls</li>\n<li>&gt;90% null; keep only last</li></ul></li>\n<li><p>Categorical features</p>\n<ul>\n<li>cat1 features (The categorical features mentioned in the competition data page) <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/data\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/data</a></li>\n<li>cat2 features (Low cardinality features; &lt;=4 unique values)</li>\n<li>cat3 features (Low cardinality features; &gt;=8 and &lt;=21 unique values)</li></ul></li>\n<li><p>Last - First (<a href=\"https://www.kaggle.com/code/thedevastator/lag-features-are-all-you-need\" target=\"_blank\">https://www.kaggle.com/code/thedevastator/lag-features-are-all-you-need</a>)</p></li>\n<li><p>Last - mean features (<a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977</a>)</p></li>\n</ul>\n<p>Credits:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/333940\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/333940</a></li>\n<li><a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format</a></li>\n</ul>",
      "rawMarkdown": "There have been a lot of good feature engineering notebooks published for this competition so far. I have put together some basic feature engineering in the notebook below\n\nhttps://www.kaggle.com/code/illidan7/amex-basic-feature-engineering-1500-features\n\nIf you would rather jump directly into the hell of feature selection, I have saved the result of this notebook as a dataset\n\nhttps://www.kaggle.com/datasets/illidan7/amexfeatureeng\n\nA lot of the features are based on what others have shared and I have tried to add in some of my own as well\n\n- Date based features (Column S_2)\n- \"After pay\" features (https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb)\n- Null columns handling \n    - \\>30% null; Count num nulls\n    - \\>90% null; keep only last\n\n- Categorical features\n    - cat1 features (The categorical features mentioned in the competition data page) https://www.kaggle.com/competitions/amex-default-prediction/data\n    - cat2 features (Low cardinality features; <=4 unique values)\n    - cat3 features (Low cardinality features; >=8 and <=21 unique values)\n\n\n- Last - First (https://www.kaggle.com/code/thedevastator/lag-features-are-all-you-need)\n- Last - mean features (https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977)\n\n\nCredits:\n- https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\n- https://www.kaggle.com/competitions/amex-default-prediction/discussion/333940\n- https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format",
      "votes": null
    },
    {
      "id": "1852170",
      "postDate": "07/11/2022 21:01:36",
      "content": "<p>What a coincidence, I just posted about isolating the best set of features. (<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/336546\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/336546</a>)<br>\nI am running your dataset in my pipeline, will update with the results.</p>\n<p>edit: 0.7950 CV, 41.2min training time (interestingly long)</p>",
      "rawMarkdown": "What a coincidence, I just posted about isolating the best set of features. (https://www.kaggle.com/competitions/amex-default-prediction/discussion/336546)\nI am running your dataset in my pipeline, will update with the results.\n\nedit: 0.7950 CV, 41.2min training time (interestingly long)",
      "votes": null
    },
    {
      "id": "1852371",
      "postDate": "07/12/2022 03:07:36",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/illidan7\" target=\"_blank\">@illidan7</a>; thanks for sharing it will help me once I jump back this weekend</p>",
      "rawMarkdown": "Hello @illidan7; thanks for sharing it will help me once I jump back this weekend",
      "votes": null
    },
    {
      "id": "1853131",
      "postDate": "07/12/2022 16:13:38",
      "content": "<p>Happy to hear it!</p>",
      "rawMarkdown": "Happy to hear it!",
      "votes": null
    },
    {
      "id": "1853137",
      "postDate": "07/12/2022 16:16:33",
      "content": "<p>Yeah I guess it is because it is a very wide feature set. I'm sure with some good feature selection, you can bring down that training time and even improve on the CV</p>\n<p>Thanks for sharing the results!</p>",
      "rawMarkdown": "Yeah I guess it is because it is a very wide feature set. I'm sure with some good feature selection, you can bring down that training time and even improve on the CV\n\nThanks for sharing the results!",
      "votes": null
    },
    {
      "id": "1853293",
      "postDate": "07/12/2022 19:26:36",
      "content": "<p>Null column handling sounds interesting to try , I haven't tried that one yet . Did it help your cv ? </p>",
      "rawMarkdown": "Null column handling sounds interesting to try , I haven't tried that one yet . Did it help your cv ?",
      "votes": null
    },
    {
      "id": "1853320",
      "postDate": "07/12/2022 20:03:21",
      "content": "<p>For now I have just separated these &gt;90% null columns so that I am not doing the groupby aggregations on them. Most likely those would not be useful and will just increase the column count</p>\n<p>For the &gt;30% null columns, I've just added an aggregation to count number of nulls in the series for the customer</p>\n<p>Have not experimented a whole lot with it yet. But in some initial experiments, they were not really that impactful</p>",
      "rawMarkdown": "For now I have just separated these >90% null columns so that I am not doing the groupby aggregations on them. Most likely those would not be useful and will just increase the column count\n\nFor the >30% null columns, I've just added an aggregation to count number of nulls in the series for the customer\n\nHave not experimented a whole lot with it yet. But in some initial experiments, they were not really that impactful",
      "votes": null
    },
    {
      "id": "1853371",
      "postDate": "07/12/2022 21:11:36",
      "content": "<p>Great thanks for you insights on this .. yup it will have specially the 90% population to reduce the size for training .. its effect we have to see . </p>",
      "rawMarkdown": "Great thanks for you insights on this .. yup it will have specially the 90% population to reduce the size for training .. its effect we have to see .",
      "votes": null
    },
    {
      "id": "1853484",
      "postDate": "07/12/2022 23:18:52",
      "content": "<p>great work! Did these FEs improve your CV?</p>",
      "rawMarkdown": "great work! Did these FEs improve your CV?",
      "votes": null
    },
    {
      "id": "1854283",
      "postDate": "07/13/2022 15:08:59",
      "content": "<p>Yes, combined with some good feature selection, it has improved CV for my XGB model</p>",
      "rawMarkdown": "Yes, combined with some good feature selection, it has improved CV for my XGB model",
      "votes": null
    },
    {
      "id": "1854343",
      "postDate": "07/13/2022 16:08:17",
      "content": "<p>great! good luck!</p>",
      "rawMarkdown": "great! good luck!",
      "votes": null
    },
    {
      "id": "1854553",
      "postDate": "07/13/2022 19:37:00",
      "content": "<p>Thanks for sharing . One place to get all (most of?) FE ideas shared for the competition.</p>",
      "rawMarkdown": "Thanks for sharing . One place to get all (most of?) FE ideas shared for the competition.",
      "votes": null
    },
    {
      "id": "1854691",
      "postDate": "07/13/2022 23:11:58",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "1855827",
      "postDate": "07/15/2022 00:55:42",
      "content": "<p>Thanks for sharing your work!</p>\n<p>I have a few comments:</p>\n<ul>\n<li>Have you tried any other feature engineering methods, such as target encoding or leave-one-out encoding?</li>\n<li>For the null columns, have you considered imputing the missing values using some sophisticated method?</li>\n<li>Have you tried any other model? maybe deep learning?<br>\nKeep up the good work!</li>\n</ul>",
      "rawMarkdown": "Thanks for sharing your work!\n\nI have a few comments:\n- Have you tried any other feature engineering methods, such as target encoding or leave-one-out encoding?\n- For the null columns, have you considered imputing the missing values using some sophisticated method?\n- Have you tried any other model? maybe deep learning?\nKeep up the good work!",
      "votes": null
    },
    {
      "id": "1856755",
      "postDate": "07/15/2022 15:14:46",
      "content": "<p>Unfortunately the answer to all 3 of your questions is no</p>\n<p>Would love to experiment with the things you have mentioned: encoding, imputation, models. Just haven't gotten around to it</p>\n<p>Here are some interesting links I saved to start exploring these </p>\n<ul>\n<li><a href=\"https://towardsdatascience.com/why-you-should-always-use-feature-embeddings-with-structured-datasets-7f280b40e716\" target=\"_blank\">https://towardsdatascience.com/why-you-should-always-use-feature-embeddings-with-structured-datasets-7f280b40e716</a></li>\n<li><a href=\"https://www.kaggle.com/code/azminetoushikwasi/imputation-different-techniques-with-codes\" target=\"_blank\">https://www.kaggle.com/code/azminetoushikwasi/imputation-different-techniques-with-codes</a></li>\n</ul>\n<p>Hope that helps!</p>",
      "rawMarkdown": "Unfortunately the answer to all 3 of your questions is no\n\nWould love to experiment with the things you have mentioned: encoding, imputation, models. Just haven't gotten around to it\n\nHere are some interesting links I saved to start exploring these \n\n- https://towardsdatascience.com/why-you-should-always-use-feature-embeddings-with-structured-datasets-7f280b40e716\n- https://www.kaggle.com/code/azminetoushikwasi/imputation-different-techniques-with-codes\n\nHope that helps!",
      "votes": null
    },
    {
      "id": "1857697",
      "postDate": "07/16/2022 09:54:22",
      "content": "<p>Thanks…this was quite helpful</p>",
      "rawMarkdown": "Thanks...this was quite helpful",
      "votes": null
    },
    {
      "id": "1858004",
      "postDate": "07/16/2022 15:20:17",
      "content": "<p>If I understand correctly, you also treat cat2 and cat3 as categorical features, how does this affect your cv scores? Do it improve your score? </p>",
      "rawMarkdown": "If I understand correctly, you also treat cat2 and cat3 as categorical features, how does this affect your cv scores? Do it improve your score?",
      "votes": null
    },
    {
      "id": "1858132",
      "postDate": "07/16/2022 17:30:12",
      "content": "<p>In this case, I have just identified them as features in the dataset which have low cardinality, apart from the categorical features mentioned in the competition data page. </p>\n<p>So that we can exclude them from some of the generic aggregations ('min', 'max', 'mean',  'std') which make more sense for float features with lots of unique values. CV score has not dropped a whole lot by excluding these aggregation features in my experiments</p>",
      "rawMarkdown": "In this case, I have just identified them as features in the dataset which have low cardinality, apart from the categorical features mentioned in the competition data page. \n\nSo that we can exclude them from some of the generic aggregations ('min', 'max', 'mean',  'std') which make more sense for float features with lots of unique values. CV score has not dropped a whole lot by excluding these aggregation features in my experiments",
      "votes": null
    },
    {
      "id": "1858178",
      "postDate": "07/16/2022 18:13:30",
      "content": "<p>This is amazing, Thank you for sharing it</p>",
      "rawMarkdown": "This is amazing, Thank you for sharing it",
      "votes": null
    },
    {
      "id": "1917638",
      "postDate": "08/28/2022 23:58:30",
      "content": "<p>This is brilliant, great share! </p>",
      "rawMarkdown": "This is brilliant, great share!",
      "votes": null
    },
    {
      "id": "1918239",
      "postDate": "08/29/2022 12:14:25",
      "content": "<p>Interesting notebook and thanks for sharing </p>",
      "rawMarkdown": "Interesting notebook and thanks for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1852170,
      "author_name": "mrandri19",
      "author_url": "",
      "post_date": "07/11/2022 21:01:36",
      "content": "<p>What a coincidence, I just posted about isolating the best set of features. (<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/336546\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/336546</a>)<br>\nI am running your dataset in my pipeline, will update with the results.</p>\n<p>edit: 0.7950 CV, 41.2min training time (interestingly long)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1853137,
          "author_name": "illidan7",
          "author_url": "",
          "post_date": "07/12/2022 16:16:33",
          "content": "<p>Yeah I guess it is because it is a very wide feature set. I'm sure with some good feature selection, you can bring down that training time and even improve on the CV</p>\n<p>Thanks for sharing the results!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1852371,
      "author_name": "cv13j0",
      "author_url": "",
      "post_date": "07/12/2022 03:07:36",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/illidan7\" target=\"_blank\">@illidan7</a>; thanks for sharing it will help me once I jump back this weekend</p>",
      "votes": null,
      "replies": [
        {
          "id": 1853131,
          "author_name": "illidan7",
          "author_url": "",
          "post_date": "07/12/2022 16:13:38",
          "content": "<p>Happy to hear it!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1853293,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "07/12/2022 19:26:36",
      "content": "<p>Null column handling sounds interesting to try , I haven't tried that one yet . Did it help your cv ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1853320,
          "author_name": "illidan7",
          "author_url": "",
          "post_date": "07/12/2022 20:03:21",
          "content": "<p>For now I have just separated these &gt;90% null columns so that I am not doing the groupby aggregations on them. Most likely those would not be useful and will just increase the column count</p>\n<p>For the &gt;30% null columns, I've just added an aggregation to count number of nulls in the series for the customer</p>\n<p>Have not experimented a whole lot with it yet. But in some initial experiments, they were not really that impactful</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1853371,
          "author_name": "gauravbrills",
          "author_url": "",
          "post_date": "07/12/2022 21:11:36",
          "content": "<p>Great thanks for you insights on this .. yup it will have specially the 90% population to reduce the size for training .. its effect we have to see . </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1853484,
      "author_name": "mohammadrahmati",
      "author_url": "",
      "post_date": "07/12/2022 23:18:52",
      "content": "<p>great work! Did these FEs improve your CV?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1854283,
          "author_name": "illidan7",
          "author_url": "",
          "post_date": "07/13/2022 15:08:59",
          "content": "<p>Yes, combined with some good feature selection, it has improved CV for my XGB model</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1854343,
          "author_name": "mohammadrahmati",
          "author_url": "",
          "post_date": "07/13/2022 16:08:17",
          "content": "<p>great! good luck!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1854553,
      "author_name": "nitishraj",
      "author_url": "",
      "post_date": "07/13/2022 19:37:00",
      "content": "<p>Thanks for sharing . One place to get all (most of?) FE ideas shared for the competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1854691,
      "author_name": "zexigong",
      "author_url": "",
      "post_date": "07/13/2022 23:11:58",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1855827,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "07/15/2022 00:55:42",
      "content": "<p>Thanks for sharing your work!</p>\n<p>I have a few comments:</p>\n<ul>\n<li>Have you tried any other feature engineering methods, such as target encoding or leave-one-out encoding?</li>\n<li>For the null columns, have you considered imputing the missing values using some sophisticated method?</li>\n<li>Have you tried any other model? maybe deep learning?<br>\nKeep up the good work!</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1856755,
          "author_name": "illidan7",
          "author_url": "",
          "post_date": "07/15/2022 15:14:46",
          "content": "<p>Unfortunately the answer to all 3 of your questions is no</p>\n<p>Would love to experiment with the things you have mentioned: encoding, imputation, models. Just haven't gotten around to it</p>\n<p>Here are some interesting links I saved to start exploring these </p>\n<ul>\n<li><a href=\"https://towardsdatascience.com/why-you-should-always-use-feature-embeddings-with-structured-datasets-7f280b40e716\" target=\"_blank\">https://towardsdatascience.com/why-you-should-always-use-feature-embeddings-with-structured-datasets-7f280b40e716</a></li>\n<li><a href=\"https://www.kaggle.com/code/azminetoushikwasi/imputation-different-techniques-with-codes\" target=\"_blank\">https://www.kaggle.com/code/azminetoushikwasi/imputation-different-techniques-with-codes</a></li>\n</ul>\n<p>Hope that helps!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1857697,
      "author_name": "jayeshgokhale",
      "author_url": "",
      "post_date": "07/16/2022 09:54:22",
      "content": "<p>Thanks…this was quite helpful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1858004,
      "author_name": "dzbsun",
      "author_url": "",
      "post_date": "07/16/2022 15:20:17",
      "content": "<p>If I understand correctly, you also treat cat2 and cat3 as categorical features, how does this affect your cv scores? Do it improve your score? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1858132,
          "author_name": "illidan7",
          "author_url": "",
          "post_date": "07/16/2022 17:30:12",
          "content": "<p>In this case, I have just identified them as features in the dataset which have low cardinality, apart from the categorical features mentioned in the competition data page. </p>\n<p>So that we can exclude them from some of the generic aggregations ('min', 'max', 'mean',  'std') which make more sense for float features with lots of unique values. CV score has not dropped a whole lot by excluding these aggregation features in my experiments</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1858178,
      "author_name": "tarundalal",
      "author_url": "",
      "post_date": "07/16/2022 18:13:30",
      "content": "<p>This is amazing, Thank you for sharing it</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1917638,
      "author_name": "drossa89",
      "author_url": "",
      "post_date": "08/28/2022 23:58:30",
      "content": "<p>This is brilliant, great share! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1918239,
      "author_name": "gazu468",
      "author_url": "",
      "post_date": "08/29/2022 12:14:25",
      "content": "<p>Interesting notebook and thanks for sharing </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1852151": "There have been a lot of good feature engineering notebooks published for this competition so far. I have put together some basic feature engineering in the notebook below\n\nhttps://www.kaggle.com/code/illidan7/amex-basic-feature-engineering-1500-features\n\nIf you would rather jump directly into the hell of feature selection, I have saved the result of this notebook as a dataset\n\nhttps://www.kaggle.com/datasets/illidan7/amexfeatureeng\n\nA lot of the features are based on what others have shared and I have tried to add in some of my own as well\n\n- Date based features (Column S_2)\n- \"After pay\" features (https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb)\n- Null columns handling \n    - \\>30% null; Count num nulls\n    - \\>90% null; keep only last\n\n- Categorical features\n    - cat1 features (The categorical features mentioned in the competition data page) https://www.kaggle.com/competitions/amex-default-prediction/data\n    - cat2 features (Low cardinality features; <=4 unique values)\n    - cat3 features (Low cardinality features; >=8 and <=21 unique values)\n\n\n- Last - First (https://www.kaggle.com/code/thedevastator/lag-features-are-all-you-need)\n- Last - mean features (https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977)\n\n\nCredits:\n- https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\n- https://www.kaggle.com/competitions/amex-default-prediction/discussion/333940\n- https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format",
    "1852170": "What a coincidence, I just posted about isolating the best set of features. (https://www.kaggle.com/competitions/amex-default-prediction/discussion/336546)\nI am running your dataset in my pipeline, will update with the results.\n\nedit: 0.7950 CV, 41.2min training time (interestingly long)",
    "1852371": "Hello @illidan7; thanks for sharing it will help me once I jump back this weekend",
    "1853131": "Happy to hear it!",
    "1853137": "Yeah I guess it is because it is a very wide feature set. I'm sure with some good feature selection, you can bring down that training time and even improve on the CV\n\nThanks for sharing the results!",
    "1853293": "Null column handling sounds interesting to try , I haven't tried that one yet . Did it help your cv ?",
    "1853320": "For now I have just separated these >90% null columns so that I am not doing the groupby aggregations on them. Most likely those would not be useful and will just increase the column count\n\nFor the >30% null columns, I've just added an aggregation to count number of nulls in the series for the customer\n\nHave not experimented a whole lot with it yet. But in some initial experiments, they were not really that impactful",
    "1853371": "Great thanks for you insights on this .. yup it will have specially the 90% population to reduce the size for training .. its effect we have to see .",
    "1853484": "great work! Did these FEs improve your CV?",
    "1854283": "Yes, combined with some good feature selection, it has improved CV for my XGB model",
    "1854343": "great! good luck!",
    "1854553": "Thanks for sharing . One place to get all (most of?) FE ideas shared for the competition.",
    "1854691": "Thanks for sharing!",
    "1855827": "Thanks for sharing your work!\n\nI have a few comments:\n- Have you tried any other feature engineering methods, such as target encoding or leave-one-out encoding?\n- For the null columns, have you considered imputing the missing values using some sophisticated method?\n- Have you tried any other model? maybe deep learning?\nKeep up the good work!",
    "1856755": "Unfortunately the answer to all 3 of your questions is no\n\nWould love to experiment with the things you have mentioned: encoding, imputation, models. Just haven't gotten around to it\n\nHere are some interesting links I saved to start exploring these \n\n- https://towardsdatascience.com/why-you-should-always-use-feature-embeddings-with-structured-datasets-7f280b40e716\n- https://www.kaggle.com/code/azminetoushikwasi/imputation-different-techniques-with-codes\n\nHope that helps!",
    "1857697": "Thanks...this was quite helpful",
    "1858004": "If I understand correctly, you also treat cat2 and cat3 as categorical features, how does this affect your cv scores? Do it improve your score?",
    "1858132": "In this case, I have just identified them as features in the dataset which have low cardinality, apart from the categorical features mentioned in the competition data page. \n\nSo that we can exclude them from some of the generic aggregations ('min', 'max', 'mean',  'std') which make more sense for float features with lots of unique values. CV score has not dropped a whole lot by excluding these aggregation features in my experiments",
    "1858178": "This is amazing, Thank you for sharing it",
    "1917638": "This is brilliant, great share!",
    "1918239": "Interesting notebook and thanks for sharing"
  },
  "source": "meta"
}