{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:200%;text-align:center;display:fill;border-radius:5px;background-color:#da8b8f;overflow:hidden;font-weight:500\"><b>AmEx Default Prediction — EDA Baseline</b></div>\n","metadata":{}},{"cell_type":"code","source":"from IPython.core.display import HTML\nHTML(\"\"\"\n<style>\nfigcaption {\n  color: #4e181b;\n  font-style: italic;\n  font-size: 16px;\n  padding: 0px;\n  text-align: center;\n}\n</style>\n<img width=\"100%\" src=\"https://janbaca.net/wp-content/uploads/2015/06/Credit-Card-Design-Illustration.jpg?auto=compress&cs=tinysrgb&w=1260&h=750&dpr=1\">\n<figcaption>Published by: janbacanet</figcaption>\n\"\"\")","metadata":{"execution":{"iopub.status.busy":"2022-08-05T09:58:14.168998Z","iopub.status.idle":"2022-08-05T09:58:14.169903Z","shell.execute_reply.started":"2022-08-05T09:58:14.169563Z","shell.execute_reply":"2022-08-05T09:58:14.169612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h2 style='display:fill;color:#4e181b;background-color:#edc5c7;padding:20px'>   🗯️ Short history of Credit Cards </h2>\n\n\nThe concept of credit can be said to date back to at least 5,000 years ago in ancient **Mesopotamia**. Inscriptions on clay tablets from that time period show a record of transactions between Mesopotamian and neighboring merchants from Harappa, and are among the earliest known examples of an agreement to buy something in the moment but pay for it later. check out this interesting [BBC post](https://www.bbc.com/news/business-39870485) about Mesopotamian accounts.\n\nAmerican Express developed its first charge card in **1958**, allowing customers to pay their bill monthly in exchange for an annual fee. Merchants who accepted the card would pay American Express a percentage of the amount being charged, a precursor to the practice widely used today known as interchange fees.\n Later, in 1976, the BankAmericard changed its name to **“Visa”** a word that sounded the same in nearly every language.\n \n *These inforemations were extracted from [here](https://www.forbes.com/advisor/credit-cards/history-of-credit-cards/)*\n \n \n\n # <h2 style='display:fill;color:#4e181b;background-color:#edc5c7;padding:20px'>  🚫  Introduction on Payment Default </h2>\n \nEven though credit cards presents many advantages such us avoiding carrying a bulky wallet in your pocket, tracking the spending behaviour, fraud detection... On the other hand the major downside for credit card usage is the increasing tendency to default on their payments. Aggressive marketing strategies can encourage credit card use beyond payment capacity, thus increasing the bearer’s credit risk and resulting in defaults and losses that might have not been properly anticipated","metadata":{}},{"cell_type":"markdown","source":"\n\n## <h2 style='color:#4e181b;background-color:#edc5c7;padding:16px'>🙏 Credits</h2>\nThe code below is inspired from :\n* [@Pavithra Devi M](https://www.kaggle.com/code/ninjaac/amex-default-prediction-eda-lgbm)'s notebook : https://www.kaggle.com/code/ninjaac/amex-default-prediction-eda-lgbm\n* [@Kelli Belcher](https://www.kaggle.com/code/kellibelcher/amex-default-prediction-eda-lgbm-baseline)'s notebook https://www.kaggle.com/code/kellibelcher/amex-default-prediction-eda-lgbm-baseline\n\nThey merits many upvotes 🤗\n\n\n## <h2 style='color:#4e181b;background-color:#edc5c7;padding:16px'>🕵️‍ Data description</h2>\n[here](https://www.kaggle.com/competitions/amex-default-prediction/data) : The dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories:\n\n* `D_*` = Delinquency variables\n* `S_*` = Spend variables\n* `P_*` = Payment variables\n* `B_*` = Balance variables\n* `R_*` = Risk variables\n\n\nThe following features are categorical:\n\n`['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']`","metadata":{}},{"cell_type":"code","source":"## ESSENTIALS\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nimport gc\nimport pandas as pd\npd.set_option('display.max_columns', None)\n### warnings setting\nimport sys\nimport warnings\nif not sys.warnoptions:\n    warnings.simplefilter(\"ignore\")\nwarnings.filterwarnings(\"ignore\", category=DeprecationWarning)\n\n#### plots\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport matplotlib.colors\n\nsns.set(rc={'axes.facecolor':'#f9ecec', 'figure.facecolor':'#f9ecec'})\n\nimport plotly.express as px\nimport plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\nfrom plotly.offline import init_notebook_mode\n\n### Plotly settings\ntemp=dict(layout=go.Layout(font=dict(family=\"Ubuntu\", size=14), \n                           height=600, \n                         legend=dict(#traceorder='reversed',\n                            orientation=\"v\",\n                            y=1.15,\n                            x=0.9),\n#                      height=600,\n                           plot_bgcolor = '#f9ecec',\n                  paper_bgcolor = '#f9ecec'))\ntheme_palette={\n    'paid': '#da8b8f',\n    'default' : '#4e181b', \n    'test' : '#edc5c7' \n}\nSAVED=True\n\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-05T10:03:30.327690Z","iopub.execute_input":"2022-08-05T10:03:30.327990Z","iopub.status.idle":"2022-08-05T10:03:30.338143Z","shell.execute_reply.started":"2022-08-05T10:03:30.327967Z","shell.execute_reply":"2022-08-05T10:03:30.337123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n## <h2 style='color:#4e181b;background-color:#edc5c7;padding:16px'>🚚  Group data by Custom</h2>\n\nThe main objective of this notebook is to get a descriptive global picture of our date, instead of deeping down on every single transactions we will state out the global statement profile ","metadata":{}},{"cell_type":"code","source":"if not(SAVED):\n    train = pd.read_parquet('/kaggle/input/amex-data-integer-dtypes-parquet-format/train.parquet')\n    print('Shape', train.shape, train['customer_ID'].nunique())\n    train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T09:58:16.875268Z","iopub.execute_input":"2022-08-05T09:58:16.875947Z","iopub.status.idle":"2022-08-05T09:58:16.881535Z","shell.execute_reply.started":"2022-08-05T09:58:16.875884Z","shell.execute_reply":"2022-08-05T09:58:16.880514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"=> For each customer we have the spending history we will suppose that the default value corresponds **to the last spending date** (provided by `S_2` feature)\n\n\n(let's convert S_2 datetime feature)","metadata":{}},{"cell_type":"code","source":"if not(SAVED):\n    # change to datetime\n    train['S_2'] = pd.to_datetime(train['S_2'])\n    # sort by custoner ther by date \n    train = train.sort_values(by=['customer_ID','S_2'])\n    # keep the last paiment transaction \n    train = train.groupby('customer_ID').tail(1)\n    print(\"The training data begins on {} and ends on {}.\".format(train['S_2'].min().strftime('%d-%m-%Y'), train['S_2'].max().strftime('%d-%m-%Y')))\n\n    print(train.shape)\n    train = train.set_index('customer_ID')\n    gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T09:58:20.750116Z","iopub.execute_input":"2022-08-05T09:58:20.750651Z","iopub.status.idle":"2022-08-05T09:58:20.760340Z","shell.execute_reply.started":"2022-08-05T09:58:20.750605Z","shell.execute_reply":"2022-08-05T09:58:20.759429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if not(SAVED):\n    labels = pd.read_csv('/kaggle/input/amex-default-prediction/train_labels.csv')\n    print(labels.shape, labels['customer_ID'].nunique())\n    labels = labels.set_index('customer_ID')\n\n    train['target'] = labels['target']\n    del labels\n    gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T09:58:21.735852Z","iopub.execute_input":"2022-08-05T09:58:21.736559Z","iopub.status.idle":"2022-08-05T09:58:21.743379Z","shell.execute_reply.started":"2022-08-05T09:58:21.736522Z","shell.execute_reply":"2022-08-05T09:58:21.741562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if not(SAVED):\n    train.reset_index().to_feather('grouped_train.ftr')","metadata":{"execution":{"iopub.status.busy":"2022-08-05T09:58:24.404094Z","iopub.execute_input":"2022-08-05T09:58:24.404976Z","iopub.status.idle":"2022-08-05T09:58:24.410643Z","shell.execute_reply.started":"2022-08-05T09:58:24.404921Z","shell.execute_reply":"2022-08-05T09:58:24.409121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"we'll do the same for test data","metadata":{}},{"cell_type":"code","source":"if not(SAVED):\n    test =  pd.read_parquet('/kaggle/input/amex-data-integer-dtypes-parquet-format/test.parquet')\n    print(test.shape, test['customer_ID'].nunique())\n\n    test['S_2'] = pd.to_datetime(test['S_2'])\n    test = test.sort_values(by=['customer_ID','S_2'])\n    test = test.groupby('customer_ID').tail(1).set_index('customer_ID')\n    print(\"The test data begins on {} and ends on {}.\".format(test['S_2'].min().strftime('%d-%m-%Y'), test['S_2'].max().strftime('%d-%m-%Y')))\n    # lets get by user the last transaction history\n\n    gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T09:58:25.824956Z","iopub.execute_input":"2022-08-05T09:58:25.825380Z","iopub.status.idle":"2022-08-05T09:58:25.833286Z","shell.execute_reply.started":"2022-08-05T09:58:25.825330Z","shell.execute_reply":"2022-08-05T09:58:25.832116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# save data\nif not(SAVED):\n    test.reset_index().to_feather('grouped_test.ftr')\n    SAVED=True","metadata":{"execution":{"iopub.status.busy":"2022-08-05T09:58:26.499989Z","iopub.execute_input":"2022-08-05T09:58:26.500422Z","iopub.status.idle":"2022-08-05T09:58:26.507003Z","shell.execute_reply.started":"2022-08-05T09:58:26.500386Z","shell.execute_reply":"2022-08-05T09:58:26.505680Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Later\nif SAVED:\n    train = pd.read_feather('/kaggle/input/grouped-data/grouped_train.ftr').set_index('customer_ID')\n    test = pd.read_feather('/kaggle/input/grouped-data/grouped_test.ftr').set_index('customer_ID')","metadata":{"execution":{"iopub.status.busy":"2022-08-05T09:58:27.192830Z","iopub.execute_input":"2022-08-05T09:58:27.193750Z","iopub.status.idle":"2022-08-05T09:58:34.501600Z","shell.execute_reply.started":"2022-08-05T09:58:27.193710Z","shell.execute_reply":"2022-08-05T09:58:34.500021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <h2 style='color:#4e181b;background-color:#edc5c7;padding:16px'>🕳️ Missing Values</h2>","metadata":{}},{"cell_type":"code","source":"# Function to calculate missing values by column# Funct \n# from https://www.kaggle.com/parulpandey/starter-code-with-baseline\ndef missing_values_table(df):\n        # Total missing values by column\n        mis_val = df.isnull().sum()\n        \n        # Percentage of missing values by column\n        mis_val_percent = 100 * df.isnull().sum() / len(df)\n        \n        # build a table with the thw columns\n        mis_val_table = pd.concat([mis_val, mis_val_percent], axis=1)\n        \n        # Rename the columns\n        mis_val_table_ren_columns = mis_val_table.rename(\n        columns = {0 : 'Missing Values', 1 : '% of Total Values'})\n        \n        # Sort the table by percentage of missing descending\n        mis_val_table_ren_columns = mis_val_table_ren_columns[\n            mis_val_table_ren_columns.iloc[:,1] != 0].sort_values(\n        '% of Total Values', ascending=False).round(1)\n        \n        # Print some summary information\n        print (\"Your selected dataframe has \" + str(df.shape[1]) + \" columns.\\n\"      \n            \"There are \" + str(mis_val_table_ren_columns.shape[0]) +\n              \" columns that have missing values.\")\n        \n        # Return the dataframe with missing information\n        return mis_val_table_ren_columns\n\n# Missing values for training data\nmissing_values_train = missing_values_table(train)\n#cm = sns.color_palette('Set2', as_cmap=True)\nmissing_values_train[:20]#.style.background_gradient(cmap=cm)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T09:58:34.503729Z","iopub.execute_input":"2022-08-05T09:58:34.504119Z","iopub.status.idle":"2022-08-05T09:58:34.923338Z","shell.execute_reply.started":"2022-08-05T09:58:34.504084Z","shell.execute_reply":"2022-08-05T09:58:34.922139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = missing_values_train[:20]\ntemp=dict(layout=go.Layout(font=dict(family=\"Ubuntu\", size=14), \n                           height=600, \n                         legend=dict(#traceorder='reversed',\n                            orientation=\"v\",\n                            y=1.15,\n                            x=0.9),\n                           plot_bgcolor = '#f9ecec',\n                          paper_bgcolor = '#f9ecec'))\n\nfig = go.Figure()\nfig.add_trace(go.Bar(y=data.index, x= data['% of Total Values'],\n                     orientation='h',\n                     width=[0.6]*len(data),\n                     marker=dict(color=(data['% of Total Values'] < 80).astype('int'),\n                                 colorscale=[[0, theme_palette['default']], [1, theme_palette['paid']]], ),\n                     \n                     text = [\"<b>%.2f\"%(round(v ,2))+'%</b>' for v in data['% of Total Values']],\n                     textposition = 'inside',\n                     textfont_color = theme_palette['test']))\n\nfig.update_layout(template = temp,\n                  yaxis_automargin=False,\n                  height = 800,\n                  plot_bgcolor = '#f9ecec',\n                  paper_bgcolor = '#f9ecec',\n                  yaxis=dict(autorange=\"reversed\"),\n                  title={\n                      \"text\": \"<b>Features to Drop</b> — Threshold = 80%<BR />Missing values in pourcentage (%).<br> <br> \",\n                      \"x\":0.035,\n                      \"font_size\": 20,                    \n                  },\n                  margin={'pad':10},\n)\n\n","metadata":{"execution":{"iopub.status.busy":"2022-08-05T10:02:25.209186Z","iopub.execute_input":"2022-08-05T10:02:25.210702Z","iopub.status.idle":"2022-08-05T10:02:25.255044Z","shell.execute_reply.started":"2022-08-05T10:02:25.210619Z","shell.execute_reply":"2022-08-05T10:02:25.253698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"all Features with more than 80% will be dropped","metadata":{}},{"cell_type":"code","source":"# drop features with more than 80% of missing data\ndrop_cols = missing_values_train[missing_values_train[ '% of Total Values']>80].index.to_list()\ntrain = train.drop(drop_cols, axis=1)\ntest = test.drop(drop_cols, axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T09:58:37.051724Z","iopub.execute_input":"2022-08-05T09:58:37.052446Z","iopub.status.idle":"2022-08-05T09:58:37.789703Z","shell.execute_reply.started":"2022-08-05T09:58:37.052406Z","shell.execute_reply":"2022-08-05T09:58:37.788697Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T09:58:37.791613Z","iopub.execute_input":"2022-08-05T09:58:37.792047Z","iopub.status.idle":"2022-08-05T09:58:37.928039Z","shell.execute_reply.started":"2022-08-05T09:58:37.792014Z","shell.execute_reply":"2022-08-05T09:58:37.926758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <h2 style='color:#4e181b;background-color:#edc5c7;padding:16px'>⚒️ Regression Imputation</h2>\n\n\nThe idea of the Regression imputation is to fill the missing values with **a regressor** pre-trained on another feature (with less missing values rate) that is **highly correlated** to it. Check [this paper](https://stats-202.github.io/assets/files/9781119961154.ch8.pdf) to explore different methods of missing values imputation\n\nLet  go deep down the top remaining missing values features and see if they are correlated with some other features with low missing values rate  ","metadata":{}},{"cell_type":"code","source":"%time\ncorr = train.corr()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T09:58:39.284930Z","iopub.execute_input":"2022-08-05T09:58:39.286911Z","iopub.status.idle":"2022-08-05T09:59:19.863517Z","shell.execute_reply.started":"2022-08-05T09:58:39.286864Z","shell.execute_reply":"2022-08-05T09:59:19.862378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"remaining_missing = missing_values_table(train)\nremaining_missing[remaining_missing['% of Total Values']>50]\n","metadata":{"execution":{"iopub.status.busy":"2022-08-05T09:59:19.865342Z","iopub.execute_input":"2022-08-05T09:59:19.865763Z","iopub.status.idle":"2022-08-05T09:59:20.138867Z","shell.execute_reply.started":"2022-08-05T09:59:19.865720Z","shell.execute_reply":"2022-08-05T09:59:20.137655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets take the top remaining missing values feature **D_53** and find out the feature that is the most correlated to it","metadata":{}},{"cell_type":"code","source":"missing_var = 'D_53'\ncorr[missing_var].sort_values(key=abs, ascending=False)[:2].index[-1]","metadata":{"execution":{"iopub.status.busy":"2022-08-05T09:59:20.140171Z","iopub.execute_input":"2022-08-05T09:59:20.140992Z","iopub.status.idle":"2022-08-05T09:59:20.150095Z","shell.execute_reply.started":"2022-08-05T09:59:20.140956Z","shell.execute_reply":"2022-08-05T09:59:20.149057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**D_84** and **D_53** have **0.77** of Perason Correlation which means that these two features are **highly and positively** correlated.\n\nLet get a more local look: The best way to do it is a **scatter plot** ","metadata":{}},{"cell_type":"code","source":"\nsns.set(rc={'axes.facecolor':'#f9ecec', 'figure.facecolor':'#f9ecec'})\nfor missing_var in remaining_missing[remaining_missing['% of Total Values']>50].index:\n    corr_var = corr[missing_var].sort_values(key=abs, ascending=False)[:2].index[-1]\n    corr_val = corr[missing_var].sort_values(key=abs, ascending=False)[:2][-1]\n    plt.figure(figsize=(20,8))\n    sns.scatterplot(data=train, x=missing_var, y=corr_var,\n                    hue='target', \n                    palette=list(theme_palette.values())[:2]).set(title=f\"\\n\\n{corr_var} = f ( {missing_var} ) -- pearson_corr=\" + \"%.2f\"%(round(corr_val ,2)) )","metadata":{"execution":{"iopub.status.busy":"2022-08-05T09:59:20.152251Z","iopub.execute_input":"2022-08-05T09:59:20.152898Z","iopub.status.idle":"2022-08-05T09:59:48.910919Z","shell.execute_reply.started":"2022-08-05T09:59:20.152863Z","shell.execute_reply":"2022-08-05T09:59:48.909748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# with plotly it makes the html very heavy \nif False: \n    fig = go.Figure()\n    fig.add_trace(go.Scatter(x=train['D_84'], y=train['D_53'], mode='markers',\n                            marker_color=theme_palette['paid']))\n\n    fig.update_layout(template = temp,\n                      title={\n                          \"text\": \"<b>Scatter Plot between D_84 and D_53</b> <BR /> linear plot<br> <br> \",\n                          \"x\":0.48,\n                          \"font_size\": 18,\n\n                      })","metadata":{"execution":{"iopub.status.busy":"2022-08-05T09:59:48.913615Z","iopub.execute_input":"2022-08-05T09:59:48.914471Z","iopub.status.idle":"2022-08-05T09:59:48.921706Z","shell.execute_reply.started":"2022-08-05T09:59:48.914420Z","shell.execute_reply":"2022-08-05T09:59:48.920804Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will take **only top coorelations** to train the imputation regressor: correaltion >0.7","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import LinearRegression\n\nfor missing_var in remaining_missing[remaining_missing['% of Total Values']>50].index:\n    corr_val = corr[missing_var].sort_values(key=abs, ascending=False)[:2][-1]\n    if corr_val>0.7:\n        corr_var = corr[missing_var].sort_values(key=abs, ascending=False)[:2].index[-1]\n        print(f\"IMPUTE {missing_var} WITH {corr_var} <-> Pearson Corr = {corr_val}\")\n        # the target must not have missing values\n        tmp = train[train[missing_var].notna()]\n        # train on the top correlatef feature with the missing feature\n        X = tmp[[corr_var]].values\n        # target is the missing feature to impute\n        y = tmp[[missing_var]].values\n        # init regressor\n        reg = LinearRegression().fit(X, y)\n        train[f'{missing_var}_pred'] = np.round(reg.predict(train[[corr_var]].values))\n\n\n        train[missing_var] = np.where(train[missing_var].isna(), train[f'{missing_var}_pred'], train[missing_var])\n        train.drop(f'{missing_var}_pred', axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T09:59:48.923132Z","iopub.execute_input":"2022-08-05T09:59:48.924500Z","iopub.status.idle":"2022-08-05T09:59:49.803523Z","shell.execute_reply.started":"2022-08-05T09:59:48.924453Z","shell.execute_reply":"2022-08-05T09:59:49.802204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's check if **D_53** and **D_105** are still on the top missing values features","metadata":{}},{"cell_type":"code","source":"\nremaining_missing = missing_values_table(train)\nremaining_missing[remaining_missing['% of Total Values']>50]\n","metadata":{"execution":{"iopub.status.busy":"2022-08-05T09:59:49.806265Z","iopub.execute_input":"2022-08-05T09:59:49.806796Z","iopub.status.idle":"2022-08-05T09:59:50.082044Z","shell.execute_reply.started":"2022-08-05T09:59:49.806744Z","shell.execute_reply":"2022-08-05T09:59:50.081017Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Well not anymooore 🥳 🥳 🥳 ","metadata":{}},{"cell_type":"markdown","source":"## <h2 style='color:#4e181b;background-color:#edc5c7;padding:16px'> 🎯 Target Distribution</h2>","metadata":{}},{"cell_type":"code","source":"target=train.target.value_counts(normalize=True)\ntarget.rename(index={1:'Default',0:'Paid'},inplace=True)\ntarget","metadata":{"execution":{"iopub.status.busy":"2022-08-05T10:00:44.180895Z","iopub.execute_input":"2022-08-05T10:00:44.181262Z","iopub.status.idle":"2022-08-05T10:00:44.198705Z","shell.execute_reply.started":"2022-08-05T10:00:44.181231Z","shell.execute_reply":"2022-08-05T10:00:44.197443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n\ntarget=train.target.value_counts(normalize=True)\ntarget.rename(index={1:'Default',0:'Paid'},inplace=True)\n\n\n\nfig=go.Figure()\nfig.add_trace(go.Pie(labels=target.index, values=target*100,# hole=.45, \n                     showlegend=True,#sort=True, \n                     marker=dict(colors=list(theme_palette.values())),\n#                     marker=dict(colors=color,line=dict(color=pal,width=2.5)),\n                     hovertemplate = \"%{label} Amex Acoounts: <b>%{value:.2f}</b>%<extra></extra>\"))\nfig.update_layout(template=temp, \n                  title={\n                      \"text\":'<b>Default VS Paid Transaction</b><BR />Unballanced dataset',\n                      \"x\":0.035,\n                      \"font_size\": 20,\n                      \n                  },\n                  \n                  uniformtext_minsize=15,# width=700,\n#                  height=800,\n                  margin={'t':150, 'l':5})","metadata":{"execution":{"iopub.status.busy":"2022-08-05T09:59:52.186859Z","iopub.execute_input":"2022-08-05T09:59:52.187277Z","iopub.status.idle":"2022-08-05T09:59:52.369212Z","shell.execute_reply.started":"2022-08-05T09:59:52.187245Z","shell.execute_reply":"2022-08-05T09:59:52.367521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"=> we have an unballanced dataset here\n\nDifferent techniques can be applied to address this issue, for simplicity we will let the amazing **LightGBM** algorithm deal with this just by setting the `is_unbalance = True` (check the [doc](https://lightgbm.readthedocs.io/en/latest/Parameters.html#is_unbalance)) or `autoclass_weight` of **Catboost** model ([here](https://catboost.ai/en/docs/references/training-parameters/common#auto_class_weights) the related doc)\n\n\n## <h2 style='color:#4e181b;background-color:#edc5c7;padding:16px'> 📆 Statement Date Distribution</h2>\n\n<h3 style='color:#4e181b;background-color:#edc5c7;padding:10px'>  👉  -- Year --</h3>","metadata":{}},{"cell_type":"code","source":"train['year'] = train['S_2'].dt.year\ntrain['month'] = train['S_2'].dt.month\n\ntest['year'] = test['S_2'].dt.year\ntest['month'] = test['S_2'].dt.month","metadata":{"execution":{"iopub.status.busy":"2022-08-05T07:35:56.518121Z","iopub.execute_input":"2022-08-05T07:35:56.518570Z","iopub.status.idle":"2022-08-05T07:35:56.802659Z","shell.execute_reply.started":"2022-08-05T07:35:56.518535Z","shell.execute_reply":"2022-08-05T07:35:56.801607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = make_subplots(rows=1, cols=2)\n\n\n                  \ndata = train['year'].value_counts()\nfig.add_trace(go.Bar(x=data.index, y=data.values, name='train', marker_color=theme_palette['default']),\n             row=1, col=1)\n\ndata = test['year'].value_counts()\nfig.add_trace(go.Bar(x=data.index, y=data.values, name='test', marker_color=theme_palette['test']),\n             row=1, col=2)\n\nfig.update_layout(template=temp,#title=\"<b>Default VS Paid Transaction</b><BR />unballanced dataset\",   \n                  title={\n                      \"text\": \"<b>Train/Test Year Distrubution</b> <BR />Train-2018 / Test--2019<br> <br> \",\n                      \"x\":0.035,\n                      \"font_size\": 20,\n                      \n                  },\n\n) ","metadata":{"execution":{"iopub.status.busy":"2022-08-05T07:36:40.640415Z","iopub.execute_input":"2022-08-05T07:36:40.640818Z","iopub.status.idle":"2022-08-05T07:36:40.705660Z","shell.execute_reply.started":"2022-08-05T07:36:40.640787Z","shell.execute_reply":"2022-08-05T07:36:40.704591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style='color:#4e181b;background-color:#edc5c7;padding:10px'>  👉  -- Month --</h3>","metadata":{}},{"cell_type":"code","source":"fig = make_subplots(rows=1, cols=2)\ndata = train.reset_index().groupby(['year', 'month'])['customer_ID'].count()\ndata.index = data.index.to_series().apply(lambda x: f'{x[0]}-{x[1]}').values\nfig.add_trace(go.Bar(x=data.index, y=data.values, name='train', marker_color=theme_palette['default']),# marker_line_color=\"#EDD3B3\"),\n             row=1, col=1)\n\n\ndata = test.reset_index().groupby(['year', 'month'])['customer_ID'].count()\ndata.index = data.index.to_series().apply(lambda x: f'{x[0]}-{x[1]}').values\nfig.add_trace(go.Bar(x=data.index, y=data.values, name='test', marker_color=theme_palette['test']),\n\n              # marker_line_color=\"#c4464c\",),\n             row=1, col=2)\n\n\nfig.update_layout(template=temp,\n                    title={\n                      \"text\": \"<b>Train/Test Month Distrubution</b> <BR />Train-Mars / Test--April-Oct<br> <br> \",\n                      \"x\":0.035,\n                      \"font_size\": 20,\n                      \n                  }\n\n)         \nfig.update_xaxes(showticklabels=True, showtickprefix='none')\n","metadata":{"execution":{"iopub.status.busy":"2022-08-05T07:36:01.322885Z","iopub.execute_input":"2022-08-05T07:36:01.324122Z","iopub.status.idle":"2022-08-05T07:36:01.977689Z","shell.execute_reply.started":"2022-08-05T07:36:01.324077Z","shell.execute_reply":"2022-08-05T07:36:01.976381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style='color:#4e181b;background-color:#edc5c7;padding:10px'>  👉  -- Day of week --</h3>","metadata":{}},{"cell_type":"code","source":"week_days ={1: 'Mon', 2: 'Tue', 3: 'Wen', 4: 'Thu', 5: 'Fri', 6: 'Sat', 7: 'Sun'}\nfig = make_subplots(rows=1, cols=2)\ntrain['day_of_week'] = train['S_2'].apply(lambda x : x.isocalendar()[-1])\ntarget=pd.DataFrame(data={'Default': train.groupby(['day_of_week'])['target'].sum()})\ntarget['count'] = train['day_of_week'].value_counts()\ntarget.index = target.index.map(mapper=(lambda x: week_days[x]))\ntarget['Paid']= target['count'] - target['Default']\ntarget[\"perc_default\"] = (target['Default']*100)/target['count']\ntarget[\"perc_paid\"] = 100 - target[\"perc_default\"]\n\n#fig=go.Figure()\nfig.add_trace(go.Bar(x=target.index, y=target.Paid, name='Paid',\n                     text=target.perc_paid, texttemplate='%{text:.0f}%', \n                     textposition='inside',insidetextanchor=\"middle\",\n                     marker_color=theme_palette['paid'],\n#                     marker=dict(color=theme_palette['paid'],line=dict(color=theme_palette['paid'],width=1.5)),\n                     hovertemplate = \"<b>%{x}</b><br>Paid accounts: %{text:.2f}%\"))\n\nfig.add_trace(go.Bar(x=target.index, y=target.Default, name='Default',\n                     text=target.perc_default, texttemplate='%{text:.0f}%', \n                     textposition='inside',insidetextanchor=\"middle\",\n                     marker_color=theme_palette['default'],\n#                     marker=dict(color=color['default'],line=dict(color=pal[],width=1.5)),\n                     hovertemplate = \"<b>%{x}</b><br>Default accounts: %{text:.2f}%\"))\n\ntest['day_of_week'] = test['S_2'].apply(lambda x : x.isocalendar()[-1])\ndata = test.reset_index().groupby('day_of_week')['customer_ID'].count()\ndata.index = data.index.map(mapper=(lambda x: week_days[x]))\nfig.add_trace(go.Bar(x=data.index, y=data.values, name='Test', marker_color=theme_palette['test']), row=1, col=2)\n\n\nfig.update_layout(template=temp,#title='Distribution of Default by Day Of Week', \n                  title={\n                      \"text\": \"<b>Train/Test Week days Distrubution</b> <BR />Top1 = Saturday<br> <br> \",\n                     \"x\":0.035,\n                      \"font_size\": 18,\n                      \n                  },\n                  barmode='relative', yaxis_ticksuffix='%', #width=1400,\n\n                   margin={'pad':20},\n\n\n\n#                  legend=dict(orientation=\"h\", traceorder=\"reversed\", yanchor=\"bottom\",y=1.1,xanchor=\"right\", x=0)\n                 )\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T07:36:56.559419Z","iopub.execute_input":"2022-08-05T07:36:56.560631Z","iopub.status.idle":"2022-08-05T07:36:56.915973Z","shell.execute_reply.started":"2022-08-05T07:36:56.560559Z","shell.execute_reply":"2022-08-05T07:36:56.914781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style='color:#4e181b;background-color:#edc5c7;padding:10px'>  👉  -- Daily Statement Counts --</h3>","metadata":{}},{"cell_type":"markdown","source":"The following below is extracted from : [kellibelcher's notebook](https://www.kaggle.com/code/kellibelcher/amex-default-prediction-eda-lgbm-baseline)","metadata":{}},{"cell_type":"code","source":"## https://www.kaggle.com/code/kellibelcher/amex-default-prediction-eda-lgbm-baseline\nfig = make_subplots(rows=1, cols=2)\nplot_df=train.reset_index().groupby('S_2')['customer_ID'].nunique().reset_index()\n\nfig.add_trace(go.Scatter(x=plot_df['S_2'], \n                         y=plot_df['customer_ID'], mode='lines',\n                         name='train',\n                         line=dict(color=theme_palette['default'], width=3)),\n                        row=1, col=1)\nplot_df=test.reset_index().groupby('S_2')['customer_ID'].nunique().reset_index()\nfig.add_trace(go.Scatter(x=plot_df['S_2'], \n                         y=plot_df['customer_ID'], mode='lines',\n                         name='test',\n                         line=dict(color=theme_palette['paid'], width=3)),\n                        row=1, col=2)\nfig.update_layout(template=temp,  \n                  \n                  hovermode=\"x unified\", \n                height=550,\n#                yaxis_title='Number of Statements',\n                 margin={\"pad\":15},\n                 title={\n                      \"text\": \"<b>Train/Test Daily Customer Statements Count</b> <BR />Seasonal Trend on train data<br> <br> \",\n                      \"x\":0.035,\n                      \"font_size\": 18\n                      \n                  })\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T08:23:59.782334Z","iopub.execute_input":"2022-08-05T08:23:59.783205Z","iopub.status.idle":"2022-08-05T08:24:00.986563Z","shell.execute_reply.started":"2022-08-05T08:23:59.783165Z","shell.execute_reply":"2022-08-05T08:24:00.985221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <h2 style='color:#4e181b;background-color:#edc5c7;padding:16px'> 📈 Kernel Density Estimate Plot </h2>\n\nSeaborKernel Density Estimate (KDE) Plot allows to dress the **“shape”** of our features, as a kind of continuous replacement for the discrete histogram. ","metadata":{}},{"cell_type":"code","source":"cat_cols = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\n\ndef show_kde_plot(tag='S', title=\"Distribution of Spend Features\", fig_size=(16,20)): \n    kde_cols = filter(lambda x : (x.startswith(tag)) & (x not in cat_cols), train.columns)\n    plot_df = train[list(kde_cols) + ['target']]\n    n_cols = int(np.ceil(np.sqrt(plot_df.shape[1]-1)))\n    fig, ax = plt.subplots(n_cols, n_cols, figsize=fig_size)\n    fig.suptitle(title,fontsize=16)\n    row=0\n    col = list(range(n_cols))*n_cols\n    for i, column in enumerate(plot_df.columns[:-1]):\n        if (i!=0)&(i%n_cols==0):\n            row+=1\n        sns.kdeplot(x=column, hue='target', \n                    palette=list(theme_palette.values())[:2], hue_order=[1,0], \n                    label=['Default','Paid'], \n                    data=plot_df, \n                    fill=True, linewidth=2, legend=False, ax=ax[row,col[i]])\n\n        ax[row,col[i]].tick_params(left=False,bottom=False, labelrotation=15)\n        ax[row,col[i]].set(title='\\n\\n{}'.format(column), xlabel='', ylabel=('Density' if i%n_cols==0 else ''))\n    #    ax.set_xticklabels(ax.get_xticklabels(),rotation = 30)\n    for i in range(1,n_cols):\n        ax[n_cols-1, i].set_visible(False)\n    handles, _ = ax[0,0].get_legend_handles_labels() \n    fig.legend(labels=['Default','Paid'], handles=reversed(handles), ncol=2, bbox_to_anchor=(0.18, 0.985))\n    sns.despine(bottom=True, trim=True)\n    plt.tight_layout(rect=[0, 0.2, 1, 0.99])","metadata":{"execution":{"iopub.status.busy":"2022-08-05T07:38:07.892733Z","iopub.execute_input":"2022-08-05T07:38:07.893305Z","iopub.status.idle":"2022-08-05T07:38:07.910601Z","shell.execute_reply.started":"2022-08-05T07:38:07.893263Z","shell.execute_reply":"2022-08-05T07:38:07.909549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style='color:#4e181b;background-color:#edc5c7;padding:10px'>  1️⃣  Delinquency variables</h3>","metadata":{}},{"cell_type":"code","source":"show_kde_plot(tag='D', title=\"Distribution of Delinquency Features\", fig_size=(35, 27))","metadata":{"execution":{"iopub.status.busy":"2022-08-04T17:35:18.480609Z","iopub.execute_input":"2022-08-04T17:35:18.481407Z","iopub.status.idle":"2022-08-04T17:38:21.097060Z","shell.execute_reply.started":"2022-08-04T17:35:18.481364Z","shell.execute_reply":"2022-08-04T17:38:21.095768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style='color:#4e181b;background-color:#edc5c7;padding:10px'>  2️⃣  Spend Variables</h3>","metadata":{}},{"cell_type":"code","source":"show_kde_plot(tag='S', title=\"Distribution of Spend Features\", fig_size=(35,17))","metadata":{"execution":{"iopub.status.busy":"2022-08-04T17:38:21.099024Z","iopub.execute_input":"2022-08-04T17:38:21.099513Z","iopub.status.idle":"2022-08-04T17:39:19.109663Z","shell.execute_reply.started":"2022-08-04T17:38:21.099467Z","shell.execute_reply":"2022-08-04T17:39:19.108411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style='color:#4e181b;background-color:#edc5c7;padding:10px'>  3️⃣ Payment Variables</h3>","metadata":{}},{"cell_type":"code","source":"show_kde_plot(tag='P', title=\"Distribution of Payment Features\", fig_size=(35,6))","metadata":{"execution":{"iopub.status.busy":"2022-08-04T17:39:19.111160Z","iopub.execute_input":"2022-08-04T17:39:19.111865Z","iopub.status.idle":"2022-08-04T17:39:27.265001Z","shell.execute_reply.started":"2022-08-04T17:39:19.111827Z","shell.execute_reply":"2022-08-04T17:39:27.263660Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style='color:#4e181b;background-color:#edc5c7;padding:10px'>  4️⃣ Balance Variables</h3>","metadata":{}},{"cell_type":"code","source":"show_kde_plot(tag='B', title=\"Balance Features Distribution\", fig_size=(35,17))","metadata":{"execution":{"iopub.status.busy":"2022-08-04T17:39:27.267709Z","iopub.execute_input":"2022-08-04T17:39:27.268094Z","iopub.status.idle":"2022-08-04T17:40:50.672577Z","shell.execute_reply.started":"2022-08-04T17:39:27.268057Z","shell.execute_reply":"2022-08-04T17:40:50.671355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style='color:#4e181b;background-color:#edc5c7;padding:10px'>  5️⃣ Risk Variables</h3>","metadata":{}},{"cell_type":"code","source":"show_kde_plot(tag='R', title=\"Risk Features Distribution\", fig_size=(35,17))","metadata":{"execution":{"iopub.status.busy":"2022-08-04T17:40:50.674255Z","iopub.execute_input":"2022-08-04T17:40:50.674703Z","iopub.status.idle":"2022-08-04T17:41:58.670466Z","shell.execute_reply.started":"2022-08-04T17:40:50.674662Z","shell.execute_reply":"2022-08-04T17:41:58.669213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <h2 style='color:#4e181b;background-color:#edc5c7;padding:16px'> 🎹 Categorical Features</h2>\n\n\n\n#### Countplot","metadata":{"execution":{"iopub.status.busy":"2022-08-01T17:45:16.322702Z","iopub.execute_input":"2022-08-01T17:45:16.323412Z","iopub.status.idle":"2022-08-01T17:45:16.333625Z","shell.execute_reply.started":"2022-08-01T17:45:16.323373Z","shell.execute_reply":"2022-08-01T17:45:16.332348Z"}}},{"cell_type":"code","source":"#train['B_30'].value_counts()\n\ndef count_plot(col, title):\n    \n    fig = make_subplots(rows=1, cols=2)\n    target=pd.DataFrame(data={'Default': train.groupby([col])['target'].mean()*100})\n\n    target['Paid']= 100 - target['Default']\n\n    fig.add_trace(go.Bar(x=target.index, y=target.Paid, name='Paid',\n                         text=target.Paid, texttemplate='%{text:.0f}%', \n                         textposition='inside',insidetextanchor=\"middle\",\n                         marker=dict(color=theme_palette['paid'],line=dict(color=theme_palette['paid'],width=0)),\n                         hovertemplate = \"<b>%{x}</b> | Paid accounts: %{y:.2f}%\"),\n                        row=1, col=1)\n\n    fig.add_trace(go.Bar(x=target.index, y=target.Default, name='Default',\n                         text=target.Default, texttemplate='%{text:.0f}%', \n                         textposition='inside',insidetextanchor=\"middle\",\n                         marker=dict(color=theme_palette['default'],line=dict(color=theme_palette['default'],width=0)),\n                         hovertemplate = \"<b>%{x}</b> | Default accounts: %{y:.2f}%\"),\n                        row=1, col=1)\n    \n\n    \n    \n    data = test[col].value_counts()\n    fig.add_trace(go.Bar(x=data.index, y=data.values, name='Test', \n                         marker=dict(color=theme_palette['test'],line=dict(color=theme_palette['test'],width=0))), # marker_color=\"#c4464c\", marker_line_color=\"#673523\",),\n                         row=1, col=2)\n\n\n    fig.update_layout(template=temp,title={\n                    \"text\": f\"<b>{title}</b>\",\n                    \"x\":0.035,\n                      \"font_size\": 18\n                      },\n                      \n                      plot_bgcolor = '#f9ecec',\n                      paper_bgcolor = '#f9ecec',\n                      barmode='relative', yaxis_ticksuffix='%', #width=1400,\n                      height=550,\n                      legend=dict(orientation=\"h\", traceorder=\"reversed\", yanchor=\"bottom\",y=1.05,xanchor=\"left\", x=0.9))\n    fig.show()\n    \nfor col in cat_cols:\n    count_plot(col, f'Train/Test {col} Counts')","metadata":{"execution":{"iopub.status.busy":"2022-08-05T08:24:42.788456Z","iopub.execute_input":"2022-08-05T08:24:42.788933Z","iopub.status.idle":"2022-08-05T08:24:43.705318Z","shell.execute_reply.started":"2022-08-05T08:24:42.788881Z","shell.execute_reply":"2022-08-05T08:24:43.703936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n## <h2 style='color:#4e181b;background-color:#edc5c7;padding:16px'> 🤏 Correlation Analysis</h2>","metadata":{}},{"cell_type":"code","source":"# inspired from https://www.kaggle.com/code/ninjaac/amex-default-prediction-eda-lgbm\ncorr = train.corrwith(train['target'], axis=0)\ncorr = corr[corr.notna()].sort_values(ascending=False)\n\npos = corr[(corr>0) & (corr<1)]\nneg = corr[(corr<0)]\n\nfig = go.Figure()\nfig.add_trace(go.Bar(x=pos.index, y= pos.values,\n                     orientation='v',\n                     name='Positive Correlation',\n                     marker=dict(color=theme_palette['default'],line=dict(color=theme_palette['paid'],width=0)),\n                     text = [\"%.2f\" %(round(v ,2) *100) + '%' for v in pos.values],\n                     textposition = 'outside',\n                     textfont_color = '#4E1C1E'))\n\nfig.add_trace(go.Bar(x=neg.index, y= neg.values,\n                     orientation='v',\n                     name='Negative Correlation',\n                     marker=dict(color=theme_palette['paid'],line=dict(color=theme_palette['default'],width=0)),\n                     text = [\"%.2f\" %(round(v ,2) *100) + '%' for v in neg.values],\n                     textposition = 'outside',\n                     textfont_color = '#4E1C1E'))\n\nfig.update_layout(template = temp,\n                  \n                  title={\n                      \"text\": \"<b>Pearson Correaltion with The Payment Default Feature</b> <BR />Extreme elements are the top correlated features<br> <br> \",\n                      \"x\":0.035,\n                      \"font_size\": 18,\n                      \n                  },\n                 plot_bgcolor = '#f9ecec',\n                      paper_bgcolor = '#f9ecec',\n                 legend=dict(\n                            y=1.18,\n                            x=0.88))\n#fig.update_xaxes(range=[-2,2])","metadata":{"execution":{"iopub.status.busy":"2022-08-04T17:53:57.558193Z","iopub.execute_input":"2022-08-04T17:53:57.558668Z","iopub.status.idle":"2022-08-04T17:53:58.721876Z","shell.execute_reply.started":"2022-08-04T17:53:57.558627Z","shell.execute_reply":"2022-08-04T17:53:58.720543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:57:28.714283Z","iopub.execute_input":"2022-08-04T09:57:28.715227Z","iopub.status.idle":"2022-08-04T09:57:29.003748Z","shell.execute_reply.started":"2022-08-04T09:57:28.715183Z","shell.execute_reply":"2022-08-04T09:57:29.002367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"the most negatively correlated feature with the target is **P_2** with **-0.67** correlation. At the same time most positively correlated feature is **D_48** with **0.61** correlation","metadata":{}},{"cell_type":"markdown","source":"## <h2 style='color:#4e181b;background-color:#edc5c7;padding:16px'> 🧐 Zoom the TOP Correlations</h2>\n\nLets zoom for the TOP correlated features","metadata":{}},{"cell_type":"code","source":"sorted_corr = corr.sort_values(key=abs, ascending=False)[:11] # top but we have to  drop corr=1\n\n\npos = sorted_corr[(sorted_corr>0) & (sorted_corr<1)]\nneg = sorted_corr[(sorted_corr<0)].sort_values(ascending=False)\n\nfig = go.Figure()\nfig.add_trace(go.Bar(x=pos.index, y= pos.values,\n                     orientation='v',\n                     name='Positive',\n                     marker=dict(color=theme_palette['default'],line=dict(color=theme_palette['paid'],width=0)),\n                     text = [\"%.2f\" %(round(v ,2) *100) + '%' for v in pos.values],\n                     textposition = 'outside',\n                     textfont_color = '#4E1C1E'))\n\nfig.add_trace(go.Bar(x=neg.index, y= neg.values,\n                     orientation='v',\n                     name='Negative',\n                     marker=dict(color=theme_palette['paid'],line=dict(color=theme_palette['default'],width=0)),\n                     text = [\"%.2f\" %(round(v ,2) *100) + '%' for v in neg.values],\n                     textposition = 'outside',\n                     textfont_color = '#4E1C1E'))\n\nfig.update_layout(template = temp,\n                  title={\n                      \"text\": \"<b>Top-10 Correlated Features with the Payment Default Feature</b> <BR />Pearson Values > 0.5<br> <br> \",\n                      \"x\":0.035,\n                      \"font_size\": 18,\n                      \n                  },\n                 plot_bgcolor = '#f9ecec',\n                      paper_bgcolor = '#f9ecec',\n                 legend=dict(\n                            y=1.15,\n                            x=0.88))","metadata":{"execution":{"iopub.status.busy":"2022-08-04T17:54:35.413687Z","iopub.execute_input":"2022-08-04T17:54:35.415276Z","iopub.status.idle":"2022-08-04T17:54:35.460096Z","shell.execute_reply.started":"2022-08-04T17:54:35.415221Z","shell.execute_reply":"2022-08-04T17:54:35.458765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### ","metadata":{}},{"cell_type":"code","source":"","metadata":{"execution":{"iopub.status.busy":"2022-08-04T15:59:30.828889Z","iopub.execute_input":"2022-08-04T15:59:30.830087Z","iopub.status.idle":"2022-08-04T16:00:12.413771Z","shell.execute_reply.started":"2022-08-04T15:59:30.830029Z","shell.execute_reply":"2022-08-04T16:00:12.412525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}