{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"In this Notebook we will look at :\n1. The **Feature/Feature** correlations (helps detect redundant features)\n2. The **Feature/Target** correlations (helps determine feature importance)\nI am using Raddar's [denoised dataset](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format). Thank you Raddar for your great contribution!","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom colorama import Fore, Back, Style\nimport matplotlib.pyplot as plt\ntrain = pd.read_parquet('../input/amex-data-integer-dtypes-parquet-format/train.parquet')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-14T06:38:19.583661Z","iopub.execute_input":"2022-08-14T06:38:19.584172Z","iopub.status.idle":"2022-08-14T06:38:40.219446Z","shell.execute_reply.started":"2022-08-14T06:38:19.584065Z","shell.execute_reply":"2022-08-14T06:38:40.217494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:175%;text-align:center;display:fill;border-radius:5px;background-color:#016CC9;overflow:hidden;font-weight:500\">1. Feature/Feature Correlations</div>","metadata":{}},{"cell_type":"markdown","source":"We will look for **the most correlated pairs Feature/Feature** as an indicator of redundant features.\nWe first select the **last values** in the time series and drop the time dimension.","metadata":{"execution":{"iopub.status.busy":"2022-08-13T18:57:53.231979Z","iopub.execute_input":"2022-08-13T18:57:53.232437Z","iopub.status.idle":"2022-08-13T18:58:01.999833Z","shell.execute_reply.started":"2022-08-13T18:57:53.232399Z","shell.execute_reply":"2022-08-13T18:58:01.998610Z"}}},{"cell_type":"code","source":"features = [col for col in train.columns if col not in ['customer_ID', 'S_2','cid']]\ncid = pd.Categorical(train['customer_ID'], ordered=True)\nlast = (cid != np.roll(cid, -1)) # mask for last statement of every customer\ndf_last = (train.loc[last, features].set_index(np.asarray(cid[last])))\ndf_last=df_last.fillna(0)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T06:38:40.222268Z","iopub.execute_input":"2022-08-14T06:38:40.222652Z","iopub.status.idle":"2022-08-14T06:38:43.262579Z","shell.execute_reply.started":"2022-08-14T06:38:40.222619Z","shell.execute_reply":"2022-08-14T06:38:43.261479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The correlation matrix would be **too large to show in a heatmap**. We need a different way to look at it. First we transform the matrix into a list of feature1, feature2, correlation.","metadata":{}},{"cell_type":"code","source":"c=df_last.corr()\ns = c.unstack()\ncorr_ff=pd.DataFrame(s).reset_index()\ncorr_ff.columns=['feature1','feature2','corr']","metadata":{"execution":{"iopub.status.busy":"2022-08-14T06:38:43.264082Z","iopub.execute_input":"2022-08-14T06:38:43.264559Z","iopub.status.idle":"2022-08-14T06:39:27.069170Z","shell.execute_reply.started":"2022-08-14T06:38:43.264513Z","shell.execute_reply":"2022-08-14T06:39:27.067903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Then we remove the correlation between identical columns (= 1)","metadata":{}},{"cell_type":"code","source":"corr_ff.drop(corr_ff[corr_ff['feature1']==corr_ff['feature2']].index, inplace=True)\ncorr_ff.sort_values(by=['corr'],inplace=True,ascending=False)\ncorr_ff=corr_ff.reset_index(drop=True)\ncorr_ff.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T06:39:27.072394Z","iopub.execute_input":"2022-08-14T06:39:27.072899Z","iopub.status.idle":"2022-08-14T06:39:27.110898Z","shell.execute_reply.started":"2022-08-14T06:39:27.072855Z","shell.execute_reply":"2022-08-14T06:39:27.109801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"then we remove every second row as we have feature1,feature2 followed by feature2,feature1","metadata":{}},{"cell_type":"code","source":"corr_ff=corr_ff.loc[corr_ff.index % 2 == 0]","metadata":{"execution":{"iopub.status.busy":"2022-08-14T06:39:27.112609Z","iopub.execute_input":"2022-08-14T06:39:27.113532Z","iopub.status.idle":"2022-08-14T06:39:27.123287Z","shell.execute_reply.started":"2022-08-14T06:39:27.113487Z","shell.execute_reply":"2022-08-14T06:39:27.121819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15,10))\nplt.barh(corr_ff.index,corr_ff['corr'])\nplt.title('Feature/Feature correlations')\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-14T06:39:27.125504Z","iopub.execute_input":"2022-08-14T06:39:27.126170Z","iopub.status.idle":"2022-08-14T06:39:56.802627Z","shell.execute_reply.started":"2022-08-14T06:39:27.126125Z","shell.execute_reply":"2022-08-14T06:39:56.801243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We select the correlations **higher than 96% in absolute value**","metadata":{}},{"cell_type":"code","source":"highest_corr_ff=corr_ff[abs(corr_ff['corr'])>0.96]\nhighest_corr_ff","metadata":{"execution":{"iopub.status.busy":"2022-08-14T06:39:56.804301Z","iopub.execute_input":"2022-08-14T06:39:56.804724Z","iopub.status.idle":"2022-08-14T06:39:56.819893Z","shell.execute_reply.started":"2022-08-14T06:39:56.804690Z","shell.execute_reply":"2022-08-14T06:39:56.818684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have selected **24 strongly correlated feature pairs**. We can assume that **feature2 is redundant to feature1** and try to remove it. This may improve the model score. It will at least reduce the memory load and accelerate the model. That would be the following features:","metadata":{}},{"cell_type":"code","source":"features_highly_redundant=sorted(list(set(highest_corr_ff.feature2.values)))\nprint(f'{Fore.GREEN}{Style.BRIGHT}Potentially redundant Features:')\nprint(features_highly_redundant)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-14T06:39:56.821757Z","iopub.execute_input":"2022-08-14T06:39:56.822499Z","iopub.status.idle":"2022-08-14T06:39:56.831355Z","shell.execute_reply.started":"2022-08-14T06:39:56.822455Z","shell.execute_reply":"2022-08-14T06:39:56.830461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:175%;text-align:center;display:fill;border-radius:5px;background-color:#016CC9;overflow:hidden;font-weight:500\">2. Feature/Target Correlations</div>","metadata":{}},{"cell_type":"markdown","source":"We add the target column","metadata":{}},{"cell_type":"code","source":"df_last.index.name='customer_ID'\ndf_last = df_last.reset_index()\ntarget = pd.read_csv('../input/amex-default-prediction/train_labels.csv')\ndf_last=pd.merge(df_last,target,on='customer_ID',how='inner')","metadata":{"execution":{"iopub.status.busy":"2022-08-14T06:39:56.832715Z","iopub.execute_input":"2022-08-14T06:39:56.833159Z","iopub.status.idle":"2022-08-14T06:40:10.535058Z","shell.execute_reply.started":"2022-08-14T06:39:56.833113Z","shell.execute_reply":"2022-08-14T06:40:10.533804Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"c=df_last.corrwith(df_last['target'])","metadata":{"execution":{"iopub.status.busy":"2022-08-14T06:40:10.538216Z","iopub.execute_input":"2022-08-14T06:40:10.538584Z","iopub.status.idle":"2022-08-14T06:40:11.717701Z","shell.execute_reply.started":"2022-08-14T06:40:10.538552Z","shell.execute_reply":"2022-08-14T06:40:11.716479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We generate the **list of correlations feature/target.**","metadata":{}},{"cell_type":"code","source":"corr_ft=pd.DataFrame(c).reset_index()\ncorr_ft.columns=['Feature','corr']\ncorr_ft.sort_values(by=['corr'],inplace=True,ascending=False)\ncorr_ft=corr_ft.reset_index(drop=True)\ncorr_ft=corr_ft[corr_ft['Feature']!='target']\ncorr_ft","metadata":{"execution":{"iopub.status.busy":"2022-08-14T06:40:11.719346Z","iopub.execute_input":"2022-08-14T06:40:11.719737Z","iopub.status.idle":"2022-08-14T06:40:11.739662Z","shell.execute_reply.started":"2022-08-14T06:40:11.719704Z","shell.execute_reply":"2022-08-14T06:40:11.738155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15,10))\nplt.barh(corr_ft.index,corr_ft['corr'])\nplt.title('Feature/Target correlations')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T06:40:11.741173Z","iopub.execute_input":"2022-08-14T06:40:11.741549Z","iopub.status.idle":"2022-08-14T06:40:12.298560Z","shell.execute_reply.started":"2022-08-14T06:40:11.741516Z","shell.execute_reply":"2022-08-14T06:40:12.297298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We select the features with a correlation to the target **above 50%** in absolute value.","metadata":{}},{"cell_type":"code","source":"highest_corr_ft=corr_ft[abs(corr_ft['corr'])>0.5]\nhighest_corr_ft","metadata":{"execution":{"iopub.status.busy":"2022-08-14T06:40:12.299861Z","iopub.execute_input":"2022-08-14T06:40:12.300195Z","iopub.status.idle":"2022-08-14T06:40:12.313147Z","shell.execute_reply.started":"2022-08-14T06:40:12.300164Z","shell.execute_reply":"2022-08-14T06:40:12.312118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Based on the correlations the most important features are :","metadata":{}},{"cell_type":"code","source":"most_important_features=sorted(list(highest_corr_ft.Feature.values))\nprint(f'{Fore.GREEN}{Style.BRIGHT}Most important Features:')\nprint(most_important_features)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-14T06:40:12.314611Z","iopub.execute_input":"2022-08-14T06:40:12.315774Z","iopub.status.idle":"2022-08-14T06:40:12.324844Z","shell.execute_reply.started":"2022-08-14T06:40:12.315736Z","shell.execute_reply":"2022-08-14T06:40:12.323685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Similarly, we select the features with an absolute value correlation with the target **below 2%**. These are the least important features.","metadata":{}},{"cell_type":"code","source":"lowest_corr_ft=corr_ft[(abs(corr_ft['corr'])<0.02)]\nleast_important_features=sorted(list(lowest_corr_ft.Feature.values))\nprint(f'{Fore.GREEN}{Style.BRIGHT}Least important Features:')\nprint(least_important_features)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-14T06:40:12.326944Z","iopub.execute_input":"2022-08-14T06:40:12.327803Z","iopub.status.idle":"2022-08-14T06:40:12.342938Z","shell.execute_reply.started":"2022-08-14T06:40:12.327748Z","shell.execute_reply":"2022-08-14T06:40:12.341443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can also consider removing these features and see if the model perform better. It will at least perform faster and with less RAM usage.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:175%;text-align:center;display:fill;border-radius:5px;background-color:#016CC9;overflow:hidden;font-weight:500\">3. Conclusion</div>","metadata":{}},{"cell_type":"markdown","source":"We can try to remove these **potentially redundant features** :","metadata":{}},{"cell_type":"code","source":"print(f'{Fore.GREEN}{Style.BRIGHT}Potentially redundant Features:')\nprint(features_highly_redundant)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-14T06:40:12.345054Z","iopub.execute_input":"2022-08-14T06:40:12.345976Z","iopub.status.idle":"2022-08-14T06:40:12.354072Z","shell.execute_reply.started":"2022-08-14T06:40:12.345927Z","shell.execute_reply":"2022-08-14T06:40:12.352982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And we can try to remove these **least important features** :","metadata":{}},{"cell_type":"code","source":"print(f'{Fore.GREEN}{Style.BRIGHT}Least important Features:')\nprint(least_important_features)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-14T06:40:12.355770Z","iopub.execute_input":"2022-08-14T06:40:12.356180Z","iopub.status.idle":"2022-08-14T06:40:12.366749Z","shell.execute_reply.started":"2022-08-14T06:40:12.356094Z","shell.execute_reply":"2022-08-14T06:40:12.365473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally, These are the **most important Features**. We can try to create new Features based on them.","metadata":{}},{"cell_type":"code","source":"print(f'{Fore.GREEN}{Style.BRIGHT}Most important Features:')\nprint(most_important_features)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-14T06:42:10.988099Z","iopub.execute_input":"2022-08-14T06:42:10.990244Z","iopub.status.idle":"2022-08-14T06:42:10.998965Z","shell.execute_reply.started":"2022-08-14T06:42:10.990184Z","shell.execute_reply":"2022-08-14T06:42:10.997553Z"},"trusted":true},"execution_count":null,"outputs":[]}]}