{
  "id": 342798,
  "title": "🤓 🤓 Feature Engineering Data 🚀🚀",
  "url": "/competitions/amex-default-prediction/discussion/342798",
  "author_name": "SchopenHacker75",
  "post_date": "2022-08-08T20:35:21.755000",
  "votes": 15,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hello Everyone,<br>\nWhile I was applying the feature engineering step, It was very annoying and difficult to finish it on both Train and test subsets. Because of the volume of data the noteboook was crashing on each cell😬 I had to devide  each of the train and the test datasets into 3 subsets to make it possible.<br>\nThe final dataset is here : <a href=\"https://www.kaggle.com/datasets/schopenhacker75/feature-eng-agg?sort=recent-comments\" target=\"_blank\">https://www.kaggle.com/datasets/schopenhacker75/feature-eng-agg?sort=recent-comments</a> feel free to use it for this competition</p>\n<p>The initial subset was taker from <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">here</a>, the associated discussion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">here</a> and the explanation notebook that I made around raddar's insights <a href=\"https://www.kaggle.com/code/schopenhacker75/data-deanonymization\" target=\"_blank\">here</a>. The Feature Engineering part is vey inspired from <a href=\"https://www.kaggle.com/code/huseyincot/amex-agg-data-how-it-created\" target=\"_blank\">this notebook</a>.</p>\n<p>🤩🤩<strong>[BONUS] : A fancy and complete EDA notebook with plotly</strong>: <a href=\"https://www.kaggle.com/code/schopenhacker75/fancy-complete-eda\" target=\"_blank\">https://www.kaggle.com/code/schopenhacker75/fancy-complete-eda</a>🤩🤩</p>\n<p>The data pre-processing script sample looks pretty much like this:</p>\n<pre><code>train = pd.read_parquet('/kaggle/input/amex-data-integer-dtypes-parquet-format/train.parquet')\n## Define some \ncolumns = train.columns\ncat_vars = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68', 'month', 'day_of_week' ]\nnum_vars = filter(lambda x : x not in cat_vars, columns)\nweek_days ={1: 'Mon', 2: 'Tue', 3: 'Wen', 4: 'Thu', 5: 'Fri', 6: 'Sat', 7: 'Sun'}\n\n#############\n# divide nums vars by AmEx Categories\ndelequincy_vars = filter(lambda x:x.startswith('D') and x not in cat_vars, columns)\nspend_vars = filter(lambda x:(x.startswith('S')) and (x not in cat_vars) and (x != 'S_2'),columns)\npayment_vars = filter(lambda x:x.startswith('P') and x not in cat_vars, columns)\nbalance_vars = filter(lambda x:x.startswith('B') and x not in cat_vars, columns)\nrisk_vars = filter(lambda x:x.startswith('R') and x not in cat_vars, columns)\n###########\ndef extract_date_vars(df, date_var='S_2', sort_by=['customer_ID','S_2'], week_days=week_days):\n    # change to datetime\n    df[date_var] = pd.to_datetime(df[date_var])\n    # sort by custoner ther by date \n    df = df.sort_values(by=sort_by)\n    # extract some date characteristics\n    # year has not a very\n    # month\n    df['month'] = df[date_var].dt.month\n    # day of week\n    df['day_of_week'] = df[date_var].apply(lambda x : x.isocalendar()[-1])\n    return df\n\ngroup_names = [\"delequincy_vars\", \"spend_vars\", \"payment_vars\", \"balance_vars\", \"risk_vars\"]\n# row rise aggregation\ndef row_rise_aggregation(df, group_vars=[delequincy_vars, spend_vars, payment_vars, balance_vars, risk_vars], group_names=group_names):\n    for group_name, group_var in zip(group_names, group_vars):\n        df[group_name+'_sum'] = df[group_var].sum(axis=1)\n        df[group_name+'_mean'] = df[group_var].mean(axis=1)\n    df['missing_values'] = df.isnull().sum(axis=1)\n    return df\n\ndef column_rise_aggregation(df):\n    #from https://www.kaggle.com/code/huseyincot/amex-agg-data-how-it-created\n    print('shape before engineering', df.shape )\n    group_names = filter(lambda x: '_vars' in x, df.columns)\n    num_agg = df.groupby(\"customer_ID\")[list(num_vars) + list(group_names)].agg(['mean', 'std', 'min', 'max', 'last'])\n    num_agg.columns = ['_'.join(x) for x in num_agg.columns]\n\n    cat_agg = df.groupby(\"customer_ID\")[list(cat_vars)].agg(['count', 'last', 'nunique', pd.Series.mode])\n    cat_agg.columns = ['_'.join(x) for x in cat_agg.columns]\n\n    mode_cols = filter(lambda x:x.endswith('_mode'), cat_agg.columns)\n    for col in mode_cols:\n        cat_agg[col] = cat_agg[col].apply(lambda x: random.choice(str(x).strip('[]').split()))\n    #concat the two dataframes\n    df = pd.concat([num_agg, cat_agg], axis=1)\n    del num_agg, cat_agg\n    print('shape after engineering', df.shape )\n    return df\n\n###############\n# apply on train\ntrain = extract_date_vars(train)\ntrain = row_rise_aggregation(train)\ntrain = column_rise_aggregation(train)\n\n## Left join with labels:\nlabels = pd.read_csv('/kaggle/input/amex-default-prediction/train_labels.csv')\nprint(labels.shape, labels['customer_ID'].nunique())\nlabels = labels.set_index('customer_ID')\ntrain['target'] = labels['target']\ndel labels\ngc.collect()\n# Save data\ntrain.reset_index().to_feather('feat_eng_agg_train.ftr')\n\n# same for test dataset\ntest = pd.read_parquet('/kaggle/input/amex-data-integer-dtypes-parquet-format/test.parquet')\ncustomers = test['customer_ID'].unique()\nprint(f\"Shape = {test.shape}, number of customers = {len(customers)}\")\ntest = extract_date_vars(test)\ntest = row_rise_aggregation(test)\ntest = column_rise_aggregation(test)\n\ntest.reset_index().to_feather('feat_eng_agg_test.ftr')\n</code></pre>",
  "messages": [
    {
      "id": 1890556,
      "postDate": "2022-08-08T20:35:21.757Z",
      "content": "<p>Hello Everyone,<br>\nWhile I was applying the feature engineering step, It was very annoying and difficult to finish it on both Train and test subsets. Because of the volume of data the noteboook was crashing on each cell😬 I had to devide  each of the train and the test datasets into 3 subsets to make it possible.<br>\nThe final dataset is here : <a href=\"https://www.kaggle.com/datasets/schopenhacker75/feature-eng-agg?sort=recent-comments\" target=\"_blank\">https://www.kaggle.com/datasets/schopenhacker75/feature-eng-agg?sort=recent-comments</a> feel free to use it for this competition</p>\n<p>The initial subset was taker from <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">here</a>, the associated discussion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">here</a> and the explanation notebook that I made around raddar's insights <a href=\"https://www.kaggle.com/code/schopenhacker75/data-deanonymization\" target=\"_blank\">here</a>. The Feature Engineering part is vey inspired from <a href=\"https://www.kaggle.com/code/huseyincot/amex-agg-data-how-it-created\" target=\"_blank\">this notebook</a>.</p>\n<p>🤩🤩<strong>[BONUS] : A fancy and complete EDA notebook with plotly</strong>: <a href=\"https://www.kaggle.com/code/schopenhacker75/fancy-complete-eda\" target=\"_blank\">https://www.kaggle.com/code/schopenhacker75/fancy-complete-eda</a>🤩🤩</p>\n<p>The data pre-processing script sample looks pretty much like this:</p>\n<pre><code>train = pd.read_parquet('/kaggle/input/amex-data-integer-dtypes-parquet-format/train.parquet')\n## Define some \ncolumns = train.columns\ncat_vars = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68', 'month', 'day_of_week' ]\nnum_vars = filter(lambda x : x not in cat_vars, columns)\nweek_days ={1: 'Mon', 2: 'Tue', 3: 'Wen', 4: 'Thu', 5: 'Fri', 6: 'Sat', 7: 'Sun'}\n\n#############\n# divide nums vars by AmEx Categories\ndelequincy_vars = filter(lambda x:x.startswith('D') and x not in cat_vars, columns)\nspend_vars = filter(lambda x:(x.startswith('S')) and (x not in cat_vars) and (x != 'S_2'),columns)\npayment_vars = filter(lambda x:x.startswith('P') and x not in cat_vars, columns)\nbalance_vars = filter(lambda x:x.startswith('B') and x not in cat_vars, columns)\nrisk_vars = filter(lambda x:x.startswith('R') and x not in cat_vars, columns)\n###########\ndef extract_date_vars(df, date_var='S_2', sort_by=['customer_ID','S_2'], week_days=week_days):\n    # change to datetime\n    df[date_var] = pd.to_datetime(df[date_var])\n    # sort by custoner ther by date \n    df = df.sort_values(by=sort_by)\n    # extract some date characteristics\n    # year has not a very\n    # month\n    df['month'] = df[date_var].dt.month\n    # day of week\n    df['day_of_week'] = df[date_var].apply(lambda x : x.isocalendar()[-1])\n    return df\n\ngroup_names = [\"delequincy_vars\", \"spend_vars\", \"payment_vars\", \"balance_vars\", \"risk_vars\"]\n# row rise aggregation\ndef row_rise_aggregation(df, group_vars=[delequincy_vars, spend_vars, payment_vars, balance_vars, risk_vars], group_names=group_names):\n    for group_name, group_var in zip(group_names, group_vars):\n        df[group_name+'_sum'] = df[group_var].sum(axis=1)\n        df[group_name+'_mean'] = df[group_var].mean(axis=1)\n    df['missing_values'] = df.isnull().sum(axis=1)\n    return df\n\ndef column_rise_aggregation(df):\n    #from https://www.kaggle.com/code/huseyincot/amex-agg-data-how-it-created\n    print('shape before engineering', df.shape )\n    group_names = filter(lambda x: '_vars' in x, df.columns)\n    num_agg = df.groupby(\"customer_ID\")[list(num_vars) + list(group_names)].agg(['mean', 'std', 'min', 'max', 'last'])\n    num_agg.columns = ['_'.join(x) for x in num_agg.columns]\n\n    cat_agg = df.groupby(\"customer_ID\")[list(cat_vars)].agg(['count', 'last', 'nunique', pd.Series.mode])\n    cat_agg.columns = ['_'.join(x) for x in cat_agg.columns]\n\n    mode_cols = filter(lambda x:x.endswith('_mode'), cat_agg.columns)\n    for col in mode_cols:\n        cat_agg[col] = cat_agg[col].apply(lambda x: random.choice(str(x).strip('[]').split()))\n    #concat the two dataframes\n    df = pd.concat([num_agg, cat_agg], axis=1)\n    del num_agg, cat_agg\n    print('shape after engineering', df.shape )\n    return df\n\n###############\n# apply on train\ntrain = extract_date_vars(train)\ntrain = row_rise_aggregation(train)\ntrain = column_rise_aggregation(train)\n\n## Left join with labels:\nlabels = pd.read_csv('/kaggle/input/amex-default-prediction/train_labels.csv')\nprint(labels.shape, labels['customer_ID'].nunique())\nlabels = labels.set_index('customer_ID')\ntrain['target'] = labels['target']\ndel labels\ngc.collect()\n# Save data\ntrain.reset_index().to_feather('feat_eng_agg_train.ftr')\n\n# same for test dataset\ntest = pd.read_parquet('/kaggle/input/amex-data-integer-dtypes-parquet-format/test.parquet')\ncustomers = test['customer_ID'].unique()\nprint(f\"Shape = {test.shape}, number of customers = {len(customers)}\")\ntest = extract_date_vars(test)\ntest = row_rise_aggregation(test)\ntest = column_rise_aggregation(test)\n\ntest.reset_index().to_feather('feat_eng_agg_test.ftr')\n</code></pre>",
      "rawMarkdown": "Hello Everyone,\nWhile I was applying the feature engineering step, It was very annoying and difficult to finish it on both Train and test subsets. Because of the volume of data the noteboook was crashing on each cell😬 I had to devide  each of the train and the test datasets into 3 subsets to make it possible.\nThe final dataset is here : [https://www.kaggle.com/datasets/schopenhacker75/feature-eng-agg?sort=recent-comments](https://www.kaggle.com/datasets/schopenhacker75/feature-eng-agg?sort=recent-comments) feel free to use it for this competition\n\nThe initial subset was taker from [here](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format), the associated discussion [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514) and the explanation notebook that I made around raddar's insights [here](https://www.kaggle.com/code/schopenhacker75/data-deanonymization). The Feature Engineering part is vey inspired from [this notebook](https://www.kaggle.com/code/huseyincot/amex-agg-data-how-it-created).\n\n🤩🤩**[BONUS] : A fancy and complete EDA notebook with plotly**: [https://www.kaggle.com/code/schopenhacker75/fancy-complete-eda](https://www.kaggle.com/code/schopenhacker75/fancy-complete-eda)🤩🤩\n\nThe data pre-processing script sample looks pretty much like this:\n```\ntrain = pd.read_parquet('/kaggle/input/amex-data-integer-dtypes-parquet-format/train.parquet')\n## Define some \ncolumns = train.columns\ncat_vars = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68', 'month', 'day_of_week' ]\nnum_vars = filter(lambda x : x not in cat_vars, columns)\nweek_days ={1: 'Mon', 2: 'Tue', 3: 'Wen', 4: 'Thu', 5: 'Fri', 6: 'Sat', 7: 'Sun'}\n\n#############\n# divide nums vars by AmEx Categories\ndelequincy_vars = filter(lambda x:x.startswith('D') and x not in cat_vars, columns)\nspend_vars = filter(lambda x:(x.startswith('S')) and (x not in cat_vars) and (x != 'S_2'),columns)\npayment_vars = filter(lambda x:x.startswith('P') and x not in cat_vars, columns)\nbalance_vars = filter(lambda x:x.startswith('B') and x not in cat_vars, columns)\nrisk_vars = filter(lambda x:x.startswith('R') and x not in cat_vars, columns)\n###########\ndef extract_date_vars(df, date_var='S_2', sort_by=['customer_ID','S_2'], week_days=week_days):\n    # change to datetime\n    df[date_var] = pd.to_datetime(df[date_var])\n    # sort by custoner ther by date \n    df = df.sort_values(by=sort_by)\n    # extract some date characteristics\n    # year has not a very\n    # month\n    df['month'] = df[date_var].dt.month\n    # day of week\n    df['day_of_week'] = df[date_var].apply(lambda x : x.isocalendar()[-1])\n    return df\n\ngroup_names = [\"delequincy_vars\", \"spend_vars\", \"payment_vars\", \"balance_vars\", \"risk_vars\"]\n# row rise aggregation\ndef row_rise_aggregation(df, group_vars=[delequincy_vars, spend_vars, payment_vars, balance_vars, risk_vars], group_names=group_names):\n    for group_name, group_var in zip(group_names, group_vars):\n        df[group_name+'_sum'] = df[group_var].sum(axis=1)\n        df[group_name+'_mean'] = df[group_var].mean(axis=1)\n    df['missing_values'] = df.isnull().sum(axis=1)\n    return df\n\ndef column_rise_aggregation(df):\n    #from https://www.kaggle.com/code/huseyincot/amex-agg-data-how-it-created\n    print('shape before engineering', df.shape )\n    group_names = filter(lambda x: '_vars' in x, df.columns)\n    num_agg = df.groupby(\"customer_ID\")[list(num_vars) + list(group_names)].agg(['mean', 'std', 'min', 'max', 'last'])\n    num_agg.columns = ['_'.join(x) for x in num_agg.columns]\n\n    cat_agg = df.groupby(\"customer_ID\")[list(cat_vars)].agg(['count', 'last', 'nunique', pd.Series.mode])\n    cat_agg.columns = ['_'.join(x) for x in cat_agg.columns]\n    \n    mode_cols = filter(lambda x:x.endswith('_mode'), cat_agg.columns)\n    for col in mode_cols:\n        cat_agg[col] = cat_agg[col].apply(lambda x: random.choice(str(x).strip('[]').split()))\n    #concat the two dataframes\n    df = pd.concat([num_agg, cat_agg], axis=1)\n    del num_agg, cat_agg\n    print('shape after engineering', df.shape )\n    return df\n\n###############\n# apply on train\ntrain = extract_date_vars(train)\ntrain = row_rise_aggregation(train)\ntrain = column_rise_aggregation(train)\n\n## Left join with labels:\nlabels = pd.read_csv('/kaggle/input/amex-default-prediction/train_labels.csv')\nprint(labels.shape, labels['customer_ID'].nunique())\nlabels = labels.set_index('customer_ID')\ntrain['target'] = labels['target']\ndel labels\ngc.collect()\n# Save data\ntrain.reset_index().to_feather('feat_eng_agg_train.ftr')\n\n# same for test dataset\ntest = pd.read_parquet('/kaggle/input/amex-data-integer-dtypes-parquet-format/test.parquet')\ncustomers = test['customer_ID'].unique()\nprint(f\"Shape = {test.shape}, number of customers = {len(customers)}\")\ntest = extract_date_vars(test)\ntest = row_rise_aggregation(test)\ntest = column_rise_aggregation(test)\n\ntest.reset_index().to_feather('feat_eng_agg_test.ftr')\n```",
      "votes": 15
    },
    {
      "id": 1904191,
      "postDate": "2022-08-18T03:20:09.027Z",
      "content": "<p>Impressive, thanks for sharing</p>",
      "rawMarkdown": "Impressive, thanks for sharing",
      "votes": 3,
      "replies": [
        {
          "id": 1904429,
          "postDate": "2022-08-18T07:21:05.297Z",
          "content": "<p>thx <a href=\"https://www.kaggle.com/tiger0\" target=\"_blank\">@tiger0</a> 😊</p>",
          "rawMarkdown": "thx @tiger0 😊"
        }
      ]
    },
    {
      "id": 1898063,
      "postDate": "2022-08-14T08:49:18.553Z",
      "content": "<p>great job,tkx</p>",
      "rawMarkdown": "great job,tkx",
      "votes": 1,
      "replies": [
        {
          "id": 1899110,
          "postDate": "2022-08-15T04:23:03.467Z",
          "content": "<p>thx 😊      </p>",
          "rawMarkdown": "thx 😊      "
        }
      ]
    },
    {
      "id": 1895484,
      "postDate": "2022-08-12T07:01:55.580Z",
      "content": "<p>Thanks for your sharing</p>",
      "rawMarkdown": "Thanks for your sharing",
      "votes": 1,
      "replies": [
        {
          "id": 1899112,
          "postDate": "2022-08-15T04:23:20.680Z",
          "content": "<p>Welcome 😊    </p>",
          "rawMarkdown": "Welcome 😊    "
        }
      ]
    },
    {
      "id": 1894239,
      "postDate": "2022-08-11T10:51:54.017Z",
      "content": "<p>Thanks for the opening topic.<br>\nI take the opportunity and share the <a href=\"https://www.kaggle.com/code/samanemami/feature-selection-various-approaches\" target=\"_blank\">feature engineering NoteBook (comprehensive approach)</a></p>",
      "rawMarkdown": "Thanks for the opening topic.\nI take the opportunity and share the [feature engineering NoteBook (comprehensive approach)](https://www.kaggle.com/code/samanemami/feature-selection-various-approaches)",
      "votes": 1,
      "replies": [
        {
          "id": 1894261,
          "postDate": "2022-08-11T11:05:23.833Z",
          "content": "<p>Thanks For sharing <a href=\"https://www.kaggle.com/samanemami\" target=\"_blank\">@samanemami</a> 👍</p>",
          "rawMarkdown": "Thanks For sharing @samanemami 👍",
          "votes": 1
        }
      ]
    },
    {
      "id": 1900816,
      "postDate": "2022-08-16T09:17:40.907Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1904191,
      "author_name": "Tiger0",
      "author_url": "",
      "post_date": "2022-08-18T03:20:09.027000",
      "content": "<p>Impressive, thanks for sharing</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1904429,
          "author_name": "SchopenHacker75",
          "author_url": "",
          "post_date": "2022-08-18T07:21:05.297000",
          "content": "<p>thx <a href=\"https://www.kaggle.com/tiger0\" target=\"_blank\">@tiger0</a> 😊</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1898063,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-14T08:49:18.553000",
      "content": "<p>great job,tkx</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1899110,
          "author_name": "SchopenHacker75",
          "author_url": "",
          "post_date": "2022-08-15T04:23:03.467000",
          "content": "<p>thx 😊      </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1895484,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-12T07:01:55.580000",
      "content": "<p>Thanks for your sharing</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1899112,
          "author_name": "SchopenHacker75",
          "author_url": "",
          "post_date": "2022-08-15T04:23:20.680000",
          "content": "<p>Welcome 😊    </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1894239,
      "author_name": "Saman",
      "author_url": "",
      "post_date": "2022-08-11T10:51:54.017000",
      "content": "<p>Thanks for the opening topic.<br>\nI take the opportunity and share the <a href=\"https://www.kaggle.com/code/samanemami/feature-selection-various-approaches\" target=\"_blank\">feature engineering NoteBook (comprehensive approach)</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1894261,
          "author_name": "SchopenHacker75",
          "author_url": "",
          "post_date": "2022-08-11T11:05:23.833000",
          "content": "<p>Thanks For sharing <a href=\"https://www.kaggle.com/samanemami\" target=\"_blank\">@samanemami</a> 👍</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1900816,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-16T09:17:40.907000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1890556": "Hello Everyone,\nWhile I was applying the feature engineering step, It was very annoying and difficult to finish it on both Train and test subsets. Because of the volume of data the noteboook was crashing on each cell😬 I had to devide  each of the train and the test datasets into 3 subsets to make it possible.\nThe final dataset is here : [https://www.kaggle.com/datasets/schopenhacker75/feature-eng-agg?sort=recent-comments](https://www.kaggle.com/datasets/schopenhacker75/feature-eng-agg?sort=recent-comments) feel free to use it for this competition\n\nThe initial subset was taker from [here](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format), the associated discussion [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514) and the explanation notebook that I made around raddar's insights [here](https://www.kaggle.com/code/schopenhacker75/data-deanonymization). The Feature Engineering part is vey inspired from [this notebook](https://www.kaggle.com/code/huseyincot/amex-agg-data-how-it-created).\n\n🤩🤩**[BONUS] : A fancy and complete EDA notebook with plotly**: [https://www.kaggle.com/code/schopenhacker75/fancy-complete-eda](https://www.kaggle.com/code/schopenhacker75/fancy-complete-eda)🤩🤩\n\nThe data pre-processing script sample looks pretty much like this:\n```\ntrain = pd.read_parquet('/kaggle/input/amex-data-integer-dtypes-parquet-format/train.parquet')\n## Define some \ncolumns = train.columns\ncat_vars = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68', 'month', 'day_of_week' ]\nnum_vars = filter(lambda x : x not in cat_vars, columns)\nweek_days ={1: 'Mon', 2: 'Tue', 3: 'Wen', 4: 'Thu', 5: 'Fri', 6: 'Sat', 7: 'Sun'}\n\n#############\n# divide nums vars by AmEx Categories\ndelequincy_vars = filter(lambda x:x.startswith('D') and x not in cat_vars, columns)\nspend_vars = filter(lambda x:(x.startswith('S')) and (x not in cat_vars) and (x != 'S_2'),columns)\npayment_vars = filter(lambda x:x.startswith('P') and x not in cat_vars, columns)\nbalance_vars = filter(lambda x:x.startswith('B') and x not in cat_vars, columns)\nrisk_vars = filter(lambda x:x.startswith('R') and x not in cat_vars, columns)\n###########\ndef extract_date_vars(df, date_var='S_2', sort_by=['customer_ID','S_2'], week_days=week_days):\n    # change to datetime\n    df[date_var] = pd.to_datetime(df[date_var])\n    # sort by custoner ther by date \n    df = df.sort_values(by=sort_by)\n    # extract some date characteristics\n    # year has not a very\n    # month\n    df['month'] = df[date_var].dt.month\n    # day of week\n    df['day_of_week'] = df[date_var].apply(lambda x : x.isocalendar()[-1])\n    return df\n\ngroup_names = [\"delequincy_vars\", \"spend_vars\", \"payment_vars\", \"balance_vars\", \"risk_vars\"]\n# row rise aggregation\ndef row_rise_aggregation(df, group_vars=[delequincy_vars, spend_vars, payment_vars, balance_vars, risk_vars], group_names=group_names):\n    for group_name, group_var in zip(group_names, group_vars):\n        df[group_name+'_sum'] = df[group_var].sum(axis=1)\n        df[group_name+'_mean'] = df[group_var].mean(axis=1)\n    df['missing_values'] = df.isnull().sum(axis=1)\n    return df\n\ndef column_rise_aggregation(df):\n    #from https://www.kaggle.com/code/huseyincot/amex-agg-data-how-it-created\n    print('shape before engineering', df.shape )\n    group_names = filter(lambda x: '_vars' in x, df.columns)\n    num_agg = df.groupby(\"customer_ID\")[list(num_vars) + list(group_names)].agg(['mean', 'std', 'min', 'max', 'last'])\n    num_agg.columns = ['_'.join(x) for x in num_agg.columns]\n\n    cat_agg = df.groupby(\"customer_ID\")[list(cat_vars)].agg(['count', 'last', 'nunique', pd.Series.mode])\n    cat_agg.columns = ['_'.join(x) for x in cat_agg.columns]\n    \n    mode_cols = filter(lambda x:x.endswith('_mode'), cat_agg.columns)\n    for col in mode_cols:\n        cat_agg[col] = cat_agg[col].apply(lambda x: random.choice(str(x).strip('[]').split()))\n    #concat the two dataframes\n    df = pd.concat([num_agg, cat_agg], axis=1)\n    del num_agg, cat_agg\n    print('shape after engineering', df.shape )\n    return df\n\n###############\n# apply on train\ntrain = extract_date_vars(train)\ntrain = row_rise_aggregation(train)\ntrain = column_rise_aggregation(train)\n\n## Left join with labels:\nlabels = pd.read_csv('/kaggle/input/amex-default-prediction/train_labels.csv')\nprint(labels.shape, labels['customer_ID'].nunique())\nlabels = labels.set_index('customer_ID')\ntrain['target'] = labels['target']\ndel labels\ngc.collect()\n# Save data\ntrain.reset_index().to_feather('feat_eng_agg_train.ftr')\n\n# same for test dataset\ntest = pd.read_parquet('/kaggle/input/amex-data-integer-dtypes-parquet-format/test.parquet')\ncustomers = test['customer_ID'].unique()\nprint(f\"Shape = {test.shape}, number of customers = {len(customers)}\")\ntest = extract_date_vars(test)\ntest = row_rise_aggregation(test)\ntest = column_rise_aggregation(test)\n\ntest.reset_index().to_feather('feat_eng_agg_test.ftr')\n```",
    "1904191": "Impressive, thanks for sharing",
    "1898063": "great job,tkx",
    "1895484": "Thanks for your sharing",
    "1894239": "Thanks for the opening topic.\nI take the opportunity and share the [feature engineering NoteBook (comprehensive approach)](https://www.kaggle.com/code/samanemami/feature-selection-various-approaches)",
    "1900816": ""
  }
}