{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <h1 style=\"font-family: Trebuchet MS; padding: 12px; font-size: 48px; color: #CD5C5C; text-align: center; line-height: 1.25;\"><b>🏬🔧 Data Pre-processing,<span style=\"color: #40E0D0\"> EDA & Feature Engineering 📉</span></b><br><span style=\"color: #DE3163; font-size: 24px\">American Express </span></h1>\n<hr>","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"font-family: Trebuchet MS; background-color: #CD5C5C; color: #FFFFFF; padding: 12px; line-height: 1.5;\">1. | Introduction 👋</div>\n<center>\n    <img src=\"https://www.americanexpress.com/content/dam/amex/be/nl/kaarten/gold-kaart/960-608-chg-gold-card.png\" alt=\"Mart\" width=\"80%\">\n</center>\n<br>","metadata":{}},{"cell_type":"markdown","source":"## <div style=\"font-family: Trebuchet MS; background-color: #CD5C5C; color: #FFFFFF; padding: 12px; line-height: 1.5;\">Data Set Problems 🤔</div>\n<div style=\"font-family: Segoe UI; line-height: 2; color: #000000; text-align: justify\">\n    👉 American Express is a globally integrated payments company. The largest payment card issuer in the world, they provide customers with access to products, insights, and experiences that enrich lives and build business success. <br>\n    👉 <mark>We’ll be apply our machine learning skills to predict credit default</mark> which allows lenders to optimize lending decisions. <br> \n    👉 <mark><b>Data pre-processing and feature engineering will be performed to prepare the dataset</b></mark> before it is used by the machine learning model.\n</div>\n\n## <div style=\"font-family: Trebuchet MS; background-color: #CD5C5C; color: #FFFFFF; padding: 12px; line-height: 1.5;\">Objectives of Notebook 📌</div>\n<div style=\"font-family: Segoe UI; line-height: 2; color: #000000; text-align: justify\">\n    👉 <b>This notebook aims to:</b>\n    <ul>\n        <li> Perform <mark><b>initial data exploration</b></mark>.</li>\n        <li> Perform <mark><b>data pre-processing</b></mark>.</li>\n        <li> Perform <mark><b>EDA</b></mark> and <mark><b>hypothesis testing (statistical and non-statistical)</b></mark> in cleaned data set.</li>\n        <li> Perform <mark><b>feature engineering (one-hot encoding, label encoding, and binning)</b></mark>.</li>\n    </ul>\n</div>","metadata":{}},{"cell_type":"markdown","source":"## <div style=\"font-family: Trebuchet MS; background-color: #CD5C5C; color: #FFFFFF; padding: 12px; line-height: 1.5;\">Data Set Description 🧾</div>\n<div style=\"font-family: Segoe UI; line-height: 2; color: #000000; text-align: justify\">\n    👉 The dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories::\n    <ul>\n        <li> <mark><b>D_* </b></mark> = Delinquency variables,</li>\n        <li> <mark><b>S_* </b></mark> = Spend variables,</li>\n        <li> <mark><b>P_* </b></mark> = Payment variables </li>\n        <li> <mark><b>B_* </b></mark> = Balance variables </li>\n        <li> <mark><b>R_* </b></mark> = = Risk variables  </li>        \n    </ul><br>\n    \nwith the following features being categorical:\n\n['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']   \n\nOur task is to predict, for each customer_ID, the probability of a future payment default (target = 1).\n**Note that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.**\n    ","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"font-family: Trebuchet MS; background-color: #CD5C5C; color: #FFFFFF; padding: 12px; line-height: 1.5;\">2. | Importing Libraries 📚</div>\n<div style=\"font-family: Segoe UI; line-height: 2; color: #000000; text-align: justify\">\n    👉 <b>Importing libraries</b> that will be used in this notebook.\n</div>","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport seaborn as sns\nimport matplotlib.pyplot as plt \nimport missingno as mso\nimport plotly.graph_objects as go\nimport plotly.offline as po\nfrom plotly.offline import download_plotlyjs, init_notebook_mode, plot, iplot\nimport plotly.express as px\nimport random\nimport plotly.figure_factory as ff\nfrom plotly.subplots import make_subplots\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-31T09:38:45.280979Z","iopub.execute_input":"2022-07-31T09:38:45.281837Z","iopub.status.idle":"2022-07-31T09:38:50.176927Z","shell.execute_reply.started":"2022-07-31T09:38:45.281731Z","shell.execute_reply":"2022-07-31T09:38:50.175952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"font-family: Trebuchet MS; background-color: #CD5C5C; color: #FFFFFF; padding: 12px; line-height: 1.5;\">3. | Reading Dataset 👓</div>\n<div style=\"font-family: Segoe UI; line-height: 2; color: #000000; text-align: justify\">\n    👉 After importing libraries, <b>the dataset that will be used will be imported</b>.\n</div>","metadata":{}},{"cell_type":"code","source":"train_data=pd.read_feather(\"/kaggle/input/amexfeather/train_data.ftr\")\ntrain_labels=pd.read_csv(\"/kaggle/input/amex-default-prediction/train_labels.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:38:56.027829Z","iopub.execute_input":"2022-07-31T09:38:56.028171Z","iopub.status.idle":"2022-07-31T09:39:16.761388Z","shell.execute_reply.started":"2022-07-31T09:38:56.028138Z","shell.execute_reply":"2022-07-31T09:39:16.760356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# --- Reading Dataset ---\ntrain_data.head().style.background_gradient(cmap='Greens').set_properties(**{'font-family': 'Segoe UI'}).hide_index()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:39:27.950187Z","iopub.execute_input":"2022-07-31T09:39:27.950556Z","iopub.status.idle":"2022-07-31T09:39:28.462299Z","shell.execute_reply.started":"2022-07-31T09:39:27.950523Z","shell.execute_reply":"2022-07-31T09:39:28.461404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# --- Reading Dataset ---\ntrain_labels.head().style.background_gradient(cmap='Greens').set_properties(**{'font-family': 'Segoe UI'}).hide_index()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:39:33.142019Z","iopub.execute_input":"2022-07-31T09:39:33.142853Z","iopub.status.idle":"2022-07-31T09:39:33.159724Z","shell.execute_reply.started":"2022-07-31T09:39:33.142814Z","shell.execute_reply":"2022-07-31T09:39:33.158735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"font-family: Trebuchet MS; background-color: #CD5C5C; color: #FFFFFF; padding: 12px; line-height: 1.5;\">4. | Initial Data Exploration 🔍</div>\n<div style=\"font-family: Segoe UI; line-height: 2; color: #000000; text-align: justify\">\n    👉 This section will focused on <b>initial data exploration</b> before pre-process the data.\n</div>","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"font-family: Trebuchet MS; background-color: #db8a8a; color: #FFFFFF; padding: 12px; line-height: 1.5;\">4.1 | Train Labels Data Exploration 🔍</div>\n<div style=\"font-family: Segoe UI; line-height: 2; color: #000000; text-align: justify\">\n    👉 This section will focused on <b>Train Labels data set exploration</b> before pre-process the data.\n</div>","metadata":{}},{"cell_type":"code","source":"# Shape of the train label data set\n\nprint('\\033[1m'\"Shape of the train label data file\\n\"'\\033[0m',train_labels.shape)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:39:39.113094Z","iopub.execute_input":"2022-07-31T09:39:39.113505Z","iopub.status.idle":"2022-07-31T09:39:39.119324Z","shell.execute_reply.started":"2022-07-31T09:39:39.113469Z","shell.execute_reply":"2022-07-31T09:39:39.118228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Data Type\nprint('\\033[1m'\"Data types of each column in train label data file\\n\"'\\033[0m',train_labels.dtypes)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:39:43.258231Z","iopub.execute_input":"2022-07-31T09:39:43.258623Z","iopub.status.idle":"2022-07-31T09:39:43.265066Z","shell.execute_reply.started":"2022-07-31T09:39:43.258589Z","shell.execute_reply":"2022-07-31T09:39:43.264153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Missing value in the table \n\nprint('\\033[1m'\"Missing value present in each column of train label data file\\n\"'\\033[0m',train_labels.isna().any())\n","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:39:48.492115Z","iopub.execute_input":"2022-07-31T09:39:48.492482Z","iopub.status.idle":"2022-07-31T09:39:48.544515Z","shell.execute_reply.started":"2022-07-31T09:39:48.492445Z","shell.execute_reply":"2022-07-31T09:39:48.543536Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check for the duplicated values \n\nprint('\\033[1m'\"Duplicate value present in each column of train label data file\\n\"'\\033[0m',train_labels.customer_ID.duplicated().any())\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:39:52.003866Z","iopub.execute_input":"2022-07-31T09:39:52.004213Z","iopub.status.idle":"2022-07-31T09:39:52.083771Z","shell.execute_reply.started":"2022-07-31T09:39:52.004182Z","shell.execute_reply":"2022-07-31T09:39:52.082698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# No of unique customers \n\nprint('\\033[1m'\"No of Unique customers in the  train label data file\\n\"'\\033[0m',train_labels.customer_ID.nunique())\n","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:39:55.261062Z","iopub.execute_input":"2022-07-31T09:39:55.261432Z","iopub.status.idle":"2022-07-31T09:39:55.452931Z","shell.execute_reply.started":"2022-07-31T09:39:55.261382Z","shell.execute_reply":"2022-07-31T09:39:55.451864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Count of the Target\n\nprint('\\033[1m'\"Value count of the target column in the train label data file\\n\"'\\033[0m',train_labels.target.value_counts())\n","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:39:58.989604Z","iopub.execute_input":"2022-07-31T09:39:58.990096Z","iopub.status.idle":"2022-07-31T09:39:59.002747Z","shell.execute_reply.started":"2022-07-31T09:39:58.990051Z","shell.execute_reply":"2022-07-31T09:39:59.001496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target=train_labels.target.value_counts(normalize=True)\ntarget.rename(index={1:'Default',0:'Paid'},inplace=True)\ncolors = ['#17becf', '#E1396C']\ndata = go.Pie(\nvalues= target,\nlabels= target.index,\nmarker=dict(colors=colors),\ntextinfo='label+percent'\n)\nlayout = go.Layout(\ntitle=dict(text = \"Target Distribution\",x=0.46,y=0.95,font_size=20)\n)\nfig = go.Figure(data=data,layout=layout)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:40:02.055632Z","iopub.execute_input":"2022-07-31T09:40:02.056050Z","iopub.status.idle":"2022-07-31T09:40:02.225956Z","shell.execute_reply.started":"2022-07-31T09:40:02.056013Z","shell.execute_reply":"2022-07-31T09:40:02.221499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del train_labels","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:40:08.513727Z","iopub.execute_input":"2022-07-31T09:40:08.514063Z","iopub.status.idle":"2022-07-31T09:40:08.519047Z","shell.execute_reply.started":"2022-07-31T09:40:08.514032Z","shell.execute_reply":"2022-07-31T09:40:08.518035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"font-family: Trebuchet MS; background-color: #db8a8a; color: #FFFFFF; padding: 12px; line-height: 1.5;\">4.1 | Train Data Exploration 🔍</div>\n<div style=\"font-family: Segoe UI; line-height: 2; color: #000000; text-align: justify\">\n    👉 This section will focused on <b>Train  data set exploration</b> before pre-process the data.\n</div>","metadata":{}},{"cell_type":"code","source":"# Shape of the train data set\n\nprint('\\033[1m'\"Shape of the train  data file\\n\"'\\033[0m',train_data.shape)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:40:20.569403Z","iopub.execute_input":"2022-07-31T09:40:20.570292Z","iopub.status.idle":"2022-07-31T09:40:20.576487Z","shell.execute_reply.started":"2022-07-31T09:40:20.570240Z","shell.execute_reply":"2022-07-31T09:40:20.575164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Data Type\nprint('\\033[1m'\"Data types of each column in train label data file\\n\"'\\033[0m',train_data.dtypes)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:40:23.755973Z","iopub.execute_input":"2022-07-31T09:40:23.756326Z","iopub.status.idle":"2022-07-31T09:40:23.766542Z","shell.execute_reply.started":"2022-07-31T09:40:23.756293Z","shell.execute_reply":"2022-07-31T09:40:23.765562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.info(max_cols=200, show_counts=True)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:40:27.450067Z","iopub.execute_input":"2022-07-31T09:40:27.450432Z","iopub.status.idle":"2022-07-31T09:40:32.829598Z","shell.execute_reply.started":"2022-07-31T09:40:27.450383Z","shell.execute_reply":"2022-07-31T09:40:32.828643Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Following are the categorical variables 'B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68'. We are going to analyze it w.r.t to Target","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"font-family: Trebuchet MS; background-color: #db8a8a; color: #FFFFFF; padding: 12px; line-height: 1.5;\">4.1.1 | Analyzing the Balance categorical variables 🔍</div>\n<div style=\"font-family: Segoe UI; line-height: 2; color: #000000; text-align: justify\">\n    👉 This section will focused on <b>Analyzing the Balance categorical variables</b> before pre-process the data.\n</div>\n \n ","metadata":{}},{"cell_type":"code","source":"train_data[\"B_30\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:40:41.784361Z","iopub.execute_input":"2022-07-31T09:40:41.785078Z","iopub.status.idle":"2022-07-31T09:40:41.831295Z","shell.execute_reply.started":"2022-07-31T09:40:41.785039Z","shell.execute_reply":"2022-07-31T09:40:41.830378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# --- Setting Colors, Labels, Order ---\nblack_grad = ['#100C07', '#3E3B39', '#6D6A6A', '#9B9A9C', '#CAC9CD']\ncyan_grad = ['#142459', '#176BA0', '#19AADE', '#1AC9E6', '#87EAFA']\ncolors=cyan_grad\nlabels=train_data['B_30'].dropna().unique()\norder=train_data['B_30'].value_counts().index\n\n# --- Size for Both Figures ---\nplt.figure(figsize=(16, 8))\nplt.suptitle('B_30 Distribution', fontweight='heavy', fontsize='16', fontfamily='sans-serif', \n             color=black_grad[0])\n\n# --- Histogram ---\ncountplt = plt.subplot(1, 2, 1)\nplt.title('Histogram', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nax = sns.countplot(x='B_30', data=train_data, palette=colors, order=order, edgecolor=black_grad[2], alpha=0.85)\nfor rect in ax.patches:\n    ax.text (rect.get_x()+rect.get_width()/2, rect.get_height()+100,rect.get_height(), horizontalalignment='center',\n             fontsize=12, bbox=dict(facecolor='none', edgecolor=black_grad[0], linewidth=0.15, boxstyle='round'))\n\nplt.xlabel('B_30 Distribution', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.ylabel('Total', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.grid(axis='y', alpha=0.4)\ncountplt\n\n# --- Pie Chart ---\nplt.subplot(1, 2, 2)\nplt.title('Pie Chart', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nplt.pie(train_data['B_30'].value_counts(), colors=colors, labels=order, pctdistance=0.67, autopct='%.2f%%',\n        wedgeprops=dict(alpha=0.8, edgecolor=black_grad[1]), textprops={'fontsize':12})\ncentre=plt.Circle((0, 0), 0.45, fc='white', edgecolor=black_grad[1])\nplt.gcf().gca().add_artist(centre)\n\n# --- Count Categorical Labels w/out Dropping Null Walues ---\nprint('\\033[36m*' * 29)\nprint('\\033[1m'+'.: B_30 Content Total :.'+'\\033[0m')\nprint('\\033[36m*' * 29+'\\033[0m')\ntrain_data.B_38.value_counts(dropna=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:40:46.032744Z","iopub.execute_input":"2022-07-31T09:40:46.033090Z","iopub.status.idle":"2022-07-31T09:40:46.678828Z","shell.execute_reply.started":"2022-07-31T09:40:46.033059Z","shell.execute_reply":"2022-07-31T09:40:46.677775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# --- Setting Colors, Labels, Order ---\nblack_grad = ['#100C07', '#3E3B39', '#6D6A6A', '#9B9A9C', '#CAC9CD']\ncyan_grad = ['#142459', '#176BA0', '#19AADE', '#1AC9E6', '#87EAFA']\ncolors=cyan_grad\nlabels=train_data['B_38'].dropna().unique()\norder=train_data['B_38'].value_counts().index\n\n# --- Size for Both Figures ---\nplt.figure(figsize=(16, 8))\nplt.suptitle('B_38 Distribution', fontweight='heavy', fontsize='16', fontfamily='sans-serif', \n             color=black_grad[0])\n\n# --- Histogram ---\ncountplt = plt.subplot(1, 2, 1)\nplt.title('Histogram', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nax = sns.countplot(x='B_38', data=train_data, palette=colors, order=order, edgecolor=black_grad[2], alpha=0.85)\nfor rect in ax.patches:\n    ax.text (rect.get_x()+rect.get_width()/2, rect.get_height()+100,rect.get_height(), horizontalalignment='center',\n             fontsize=12, bbox=dict(facecolor='none', edgecolor=black_grad[0], linewidth=0.15, boxstyle='round'))\n\nplt.xlabel('B_38 Distribution', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.ylabel('Total', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.grid(axis='y', alpha=0.4)\ncountplt\n\n# --- Pie Chart ---\nplt.subplot(1, 2, 2)\nplt.title('Pie Chart', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nplt.pie(train_data['B_38'].value_counts(), colors=colors, labels=order, pctdistance=0.67, autopct='%.2f%%',\n        wedgeprops=dict(alpha=0.8, edgecolor=black_grad[1]), textprops={'fontsize':12})\ncentre=plt.Circle((0, 0), 0.45, fc='white', edgecolor=black_grad[1])\nplt.gcf().gca().add_artist(centre)\n\n# --- Count Categorical Labels w/out Dropping Null Walues ---\nprint('\\033[36m*' * 29)\nprint('\\033[1m'+'.: B_38 Content Total :.'+'\\033[0m')\nprint('\\033[36m*' * 29+'\\033[0m')\ntrain_data.B_38.value_counts(dropna=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:40:52.428224Z","iopub.execute_input":"2022-07-31T09:40:52.428904Z","iopub.status.idle":"2022-07-31T09:40:53.092221Z","shell.execute_reply.started":"2022-07-31T09:40:52.428863Z","shell.execute_reply":"2022-07-31T09:40:53.091340Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"font-family: Trebuchet MS; background-color: #db8a8a; color: #FFFFFF; padding: 12px; line-height: 1.5;\">4.1.2| Analyzing the Delinquency categorical variables 🔍</div>\n<div style=\"font-family: Segoe UI; line-height: 2; color: #000000; text-align: justify\">\n    👉 This section will focused on <b>Analyzing the Delinquency categorical variables</b> before pre-process the data.\n</div>\n\n\n\n","metadata":{}},{"cell_type":"code","source":"# --- Setting Colors, Labels, Order ---\npurple_grad = ['#491D8B', '#6929C4', '#8A3FFC', '#A56EFF', '#BE95FF']\ncolors=purple_grad\nlabels=train_data['D_114'].dropna().unique()\norder=train_data['D_114'].value_counts().index\n\n# --- Size for Both Figures ---\nplt.figure(figsize=(18, 8))\nplt.suptitle('D_114 Distribution', fontweight='heavy', fontsize='16', fontfamily='sans-serif', \n             color=black_grad[0])\n\n# --- Histogram ---\ncountplt = plt.subplot(1, 2, 1)\nplt.title('Histogram', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nax = sns.countplot(x='D_114', data=train_data, palette=colors, order=order, edgecolor=black_grad[2], alpha=0.85)\nfor rect in ax.patches:\n    ax.text (rect.get_x()+rect.get_width()/2, rect.get_height()+20,rect.get_height(), horizontalalignment='center', \n             fontsize=12, bbox=dict(facecolor='none', edgecolor=black_grad[0], linewidth=0.15, boxstyle='round'))\nplt.tight_layout(rect=[0, 0.04, 1, 0.965])\nplt.xlabel('D_114 Distribution', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.ylabel('Total', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.grid(axis='y', alpha=0.4)\ncountplt\n\n# --- Pie Chart ---\nplt.subplot(1, 2, 2)\nplt.title('Pie Chart', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nplt.pie(train_data['D_114'].value_counts(), colors=colors, labels=order, pctdistance=0.67, autopct='%.2f%%', \n        wedgeprops=dict(alpha=0.8, edgecolor=black_grad[1]), textprops={'fontsize':12})\ncentre=plt.Circle((0, 0), 0.45, fc='white', edgecolor=black_grad[1])\nplt.gcf().gca().add_artist(centre);\n\n# --- Count Categorical Labels w/out Dropping Null Walues ---\nprint('\\033[36m*' * 30)\nprint('\\033[1m'+'.: D_114 Total :.'+'\\033[0m')\nprint('\\033[36m*' * 30+'\\033[0m')\ntrain_data.D_114.value_counts(dropna=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:41:05.460181Z","iopub.execute_input":"2022-07-31T09:41:05.460552Z","iopub.status.idle":"2022-07-31T09:41:06.066691Z","shell.execute_reply.started":"2022-07-31T09:41:05.460519Z","shell.execute_reply":"2022-07-31T09:41:06.065471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# --- Setting Colors, Labels, Order ---\npurple_grad = ['#491D8B', '#6929C4', '#8A3FFC', '#A56EFF', '#BE95FF']\ncolors=purple_grad\nlabels=train_data['D_116'].dropna().unique()\norder=train_data['D_116'].value_counts().index\n\n# --- Size for Both Figures ---\nplt.figure(figsize=(18, 8))\nplt.suptitle('D_116 Distribution', fontweight='heavy', fontsize='16', fontfamily='sans-serif', \n             color=black_grad[0])\n\n# --- Histogram ---\ncountplt = plt.subplot(1, 2, 1)\nplt.title('Histogram', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nax = sns.countplot(x='D_116', data=train_data, palette=colors, order=order, edgecolor=black_grad[2], alpha=0.85)\nfor rect in ax.patches:\n    ax.text (rect.get_x()+rect.get_width()/2, rect.get_height()+20,rect.get_height(), horizontalalignment='center', \n             fontsize=12, bbox=dict(facecolor='none', edgecolor=black_grad[0], linewidth=0.15, boxstyle='round'))\nplt.tight_layout(rect=[0, 0.04, 1, 0.965])\nplt.xlabel('D_116 Distribution', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.ylabel('Total', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.grid(axis='y', alpha=0.4)\ncountplt\n\n# --- Pie Chart ---\nplt.subplot(1, 2, 2)\nplt.title('Pie Chart', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nplt.pie(train_data['D_116'].value_counts(), colors=colors, labels=order, pctdistance=0.67, autopct='%.2f%%', \n        wedgeprops=dict(alpha=0.8, edgecolor=black_grad[1]), textprops={'fontsize':12})\ncentre=plt.Circle((0, 0), 0.45, fc='white', edgecolor=black_grad[1])\nplt.gcf().gca().add_artist(centre);\n\n# --- Count Categorical Labels w/out Dropping Null Walues ---\nprint('\\033[36m*' * 30)\nprint('\\033[1m'+'.: D_116 Total :.'+'\\033[0m')\nprint('\\033[36m*' * 30+'\\033[0m')\ntrain_data.D_116.value_counts(dropna=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:41:13.138930Z","iopub.execute_input":"2022-07-31T09:41:13.139753Z","iopub.status.idle":"2022-07-31T09:41:14.569160Z","shell.execute_reply.started":"2022-07-31T09:41:13.139711Z","shell.execute_reply":"2022-07-31T09:41:14.568250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# --- Setting Colors, Labels, Order ---\npurple_grad = ['#491D8B', '#6929C4', '#8A3FFC', '#A56EFF', '#BE95FF']\ncolors=purple_grad\nlabels=train_data['D_117'].dropna().unique()\norder=train_data['D_117'].value_counts().index\n\n# --- Size for Both Figures ---\nplt.figure(figsize=(18, 8))\nplt.suptitle('D_117 Distribution', fontweight='heavy', fontsize='16', fontfamily='sans-serif', \n             color=black_grad[0])\n\n# --- Histogram ---\ncountplt = plt.subplot(1, 2, 1)\nplt.title('Histogram', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nax = sns.countplot(x='D_117', data=train_data, palette=colors, order=order, edgecolor=black_grad[2], alpha=0.85)\nfor rect in ax.patches:\n    ax.text (rect.get_x()+rect.get_width()/2, rect.get_height()+20,rect.get_height(), horizontalalignment='center', \n             fontsize=12, bbox=dict(facecolor='none', edgecolor=black_grad[0], linewidth=0.15, boxstyle='round'))\nplt.tight_layout(rect=[0, 0.04, 1, 0.965])\nplt.xlabel('D_117 Distribution', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.ylabel('Total', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.grid(axis='y', alpha=0.4)\ncountplt\n\n# --- Pie Chart ---\nplt.subplot(1, 2, 2)\nplt.title('Pie Chart', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nplt.pie(train_data['D_117'].value_counts(), colors=colors, labels=order, pctdistance=0.67, autopct='%.2f%%', \n        wedgeprops=dict(alpha=0.8, edgecolor=black_grad[1]), textprops={'fontsize':12})\ncentre=plt.Circle((0, 0), 0.45, fc='white', edgecolor=black_grad[1])\nplt.gcf().gca().add_artist(centre);\n\n# --- Count Categorical Labels w/out Dropping Null Walues ---\nprint('\\033[36m*' * 30)\nprint('\\033[1m'+'.: D_117 Total :.'+'\\033[0m')\nprint('\\033[36m*' * 30+'\\033[0m')\ntrain_data.D_117.value_counts(dropna=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:41:21.257334Z","iopub.execute_input":"2022-07-31T09:41:21.258424Z","iopub.status.idle":"2022-07-31T09:41:21.937718Z","shell.execute_reply.started":"2022-07-31T09:41:21.258375Z","shell.execute_reply":"2022-07-31T09:41:21.936594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# --- Setting Colors, Labels, Order ---\npurple_grad = ['#491D8B', '#6929C4', '#8A3FFC', '#A56EFF', '#BE95FF']\ncolors=purple_grad\nlabels=train_data['D_120'].dropna().unique()\norder=train_data['D_120'].value_counts().index\n\n# --- Size for Both Figures ---\nplt.figure(figsize=(18, 8))\nplt.suptitle('D_120 Distribution', fontweight='heavy', fontsize='16', fontfamily='sans-serif', \n             color=black_grad[0])\n\n# --- Histogram ---\ncountplt = plt.subplot(1, 2, 1)\nplt.title('Histogram', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nax = sns.countplot(x='D_120', data=train_data, palette=colors, order=order, edgecolor=black_grad[2], alpha=0.85)\nfor rect in ax.patches:\n    ax.text (rect.get_x()+rect.get_width()/2, rect.get_height()+20,rect.get_height(), horizontalalignment='center', \n             fontsize=12, bbox=dict(facecolor='none', edgecolor=black_grad[0], linewidth=0.15, boxstyle='round'))\nplt.tight_layout(rect=[0, 0.04, 1, 0.965])\nplt.xlabel('D_120 Distribution', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.ylabel('Total', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.grid(axis='y', alpha=0.4)\ncountplt\n\n# --- Pie Chart ---\nplt.subplot(1, 2, 2)\nplt.title('Pie Chart', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nplt.pie(train_data['D_120'].value_counts(), colors=colors, labels=order, pctdistance=0.67, autopct='%.2f%%', \n        wedgeprops=dict(alpha=0.8, edgecolor=black_grad[1]), textprops={'fontsize':12})\ncentre=plt.Circle((0, 0), 0.45, fc='white', edgecolor=black_grad[1])\nplt.gcf().gca().add_artist(centre);\n\n# --- Count Categorical Labels w/out Dropping Null Walues ---\nprint('\\033[36m*' * 30)\nprint('\\033[1m'+'.: D_120 Total :.'+'\\033[0m')\nprint('\\033[36m*' * 30+'\\033[0m')\ntrain_data.D_120.value_counts(dropna=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:41:30.677951Z","iopub.execute_input":"2022-07-31T09:41:30.678348Z","iopub.status.idle":"2022-07-31T09:41:31.273308Z","shell.execute_reply.started":"2022-07-31T09:41:30.678316Z","shell.execute_reply":"2022-07-31T09:41:31.272338Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# --- Setting Colors, Labels, Order ---\npurple_grad = ['#491D8B', '#6929C4', '#8A3FFC', '#A56EFF', '#BE95FF']\ncolors=purple_grad\nlabels=train_data['D_126'].dropna().unique()\norder=train_data['D_126'].value_counts().index\n\n# --- Size for Both Figures ---\nplt.figure(figsize=(18, 8))\nplt.suptitle('D_126 Distribution', fontweight='heavy', fontsize='16', fontfamily='sans-serif', \n             color=black_grad[0])\n\n# --- Histogram ---\ncountplt = plt.subplot(1, 2, 1)\nplt.title('Histogram', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nax = sns.countplot(x='D_126', data=train_data, palette=colors, order=order, edgecolor=black_grad[2], alpha=0.85)\nfor rect in ax.patches:\n    ax.text (rect.get_x()+rect.get_width()/2, rect.get_height()+20,rect.get_height(), horizontalalignment='center', \n             fontsize=12, bbox=dict(facecolor='none', edgecolor=black_grad[0], linewidth=0.15, boxstyle='round'))\nplt.tight_layout(rect=[0, 0.04, 1, 0.965])\nplt.xlabel('D_126 Distribution', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.ylabel('Total', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.grid(axis='y', alpha=0.4)\ncountplt\n\n# --- Pie Chart ---\nplt.subplot(1, 2, 2)\nplt.title('Pie Chart', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nplt.pie(train_data['D_126'].value_counts(), colors=colors, labels=order, pctdistance=0.67, autopct='%.2f%%', \n        wedgeprops=dict(alpha=0.8, edgecolor=black_grad[1]), textprops={'fontsize':12})\ncentre=plt.Circle((0, 0), 0.45, fc='white', edgecolor=black_grad[1])\nplt.gcf().gca().add_artist(centre);\n\n# --- Count Categorical Labels w/out Dropping Null Walues ---\nprint('\\033[36m*' * 30)\nprint('\\033[1m'+'.: D_126 Total :.'+'\\033[0m')\nprint('\\033[36m*' * 30+'\\033[0m')\ntrain_data.D_126.value_counts(dropna=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:41:38.300204Z","iopub.execute_input":"2022-07-31T09:41:38.300574Z","iopub.status.idle":"2022-07-31T09:41:38.942549Z","shell.execute_reply.started":"2022-07-31T09:41:38.300541Z","shell.execute_reply":"2022-07-31T09:41:38.941468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# --- Setting Colors, Labels, Order ---\npurple_grad = ['#491D8B', '#6929C4', '#8A3FFC', '#A56EFF', '#BE95FF']\ncolors=purple_grad\nlabels=train_data['D_63'].dropna().unique()\norder=train_data['D_63'].value_counts().index\n\n# --- Size for Both Figures ---\nplt.figure(figsize=(18, 8))\nplt.suptitle('D_63 Distribution', fontweight='heavy', fontsize='16', fontfamily='sans-serif', \n             color=black_grad[0])\n\n# --- Histogram ---\ncountplt = plt.subplot(1, 2, 1)\nplt.title('Histogram', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nax = sns.countplot(x='D_63', data=train_data, palette=colors, order=order, edgecolor=black_grad[2], alpha=0.85)\nfor rect in ax.patches:\n    ax.text (rect.get_x()+rect.get_width()/2, rect.get_height()+20,rect.get_height(), horizontalalignment='center', \n             fontsize=12, bbox=dict(facecolor='none', edgecolor=black_grad[0], linewidth=0.15, boxstyle='round'))\nplt.tight_layout(rect=[0, 0.04, 1, 0.965])\nplt.xlabel('D_63 Distribution', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.ylabel('Total', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.grid(axis='y', alpha=0.4)\ncountplt\n\n# --- Pie Chart ---\nplt.subplot(1, 2, 2)\nplt.title('Pie Chart', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nplt.pie(train_data['D_63'].value_counts(), colors=colors, labels=order, pctdistance=0.67, autopct='%.2f%%', \n        wedgeprops=dict(alpha=0.8, edgecolor=black_grad[1]), textprops={'fontsize':12})\ncentre=plt.Circle((0, 0), 0.45, fc='white', edgecolor=black_grad[1])\nplt.gcf().gca().add_artist(centre);\n\n# --- Count Categorical Labels w/out Dropping Null Walues ---\nprint('\\033[36m*' * 30)\nprint('\\033[1m'+'.: D_63 Total :.'+'\\033[0m')\nprint('\\033[36m*' * 30+'\\033[0m')\ntrain_data.D_63.value_counts(dropna=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:41:45.053936Z","iopub.execute_input":"2022-07-31T09:41:45.054291Z","iopub.status.idle":"2022-07-31T09:41:45.740331Z","shell.execute_reply.started":"2022-07-31T09:41:45.054258Z","shell.execute_reply":"2022-07-31T09:41:45.739340Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# --- Setting Colors, Labels, Order ---\npurple_grad = ['#491D8B', '#6929C4', '#8A3FFC', '#A56EFF', '#BE95FF']\ncolors=purple_grad\nlabels=train_data['D_64'].dropna().unique()\norder=train_data['D_64'].value_counts().index\n\n# --- Size for Both Figures ---\nplt.figure(figsize=(18, 8))\nplt.suptitle('D_64 Distribution', fontweight='heavy', fontsize='16', fontfamily='sans-serif', \n             color=black_grad[0])\n\n# --- Histogram ---\ncountplt = plt.subplot(1, 2, 1)\nplt.title('Histogram', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nax = sns.countplot(x='D_64', data=train_data, palette=colors, order=order, edgecolor=black_grad[2], alpha=0.85)\nfor rect in ax.patches:\n    ax.text (rect.get_x()+rect.get_width()/2, rect.get_height()+20,rect.get_height(), horizontalalignment='center', \n             fontsize=12, bbox=dict(facecolor='none', edgecolor=black_grad[0], linewidth=0.15, boxstyle='round'))\nplt.tight_layout(rect=[0, 0.04, 1, 0.965])\nplt.xlabel('D_64 Distribution', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.ylabel('Total', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.grid(axis='y', alpha=0.4)\ncountplt\n\n# --- Pie Chart ---\nplt.subplot(1, 2, 2)\nplt.title('Pie Chart', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nplt.pie(train_data['D_64'].value_counts(), colors=colors, labels=order, pctdistance=0.67, autopct='%.2f%%', \n        wedgeprops=dict(alpha=0.8, edgecolor=black_grad[1]), textprops={'fontsize':12})\ncentre=plt.Circle((0, 0), 0.45, fc='white', edgecolor=black_grad[1])\nplt.gcf().gca().add_artist(centre);\n\n# --- Count Categorical Labels w/out Dropping Null Walues ---\nprint('\\033[36m*' * 30)\nprint('\\033[1m'+'.: D_64 Total :.'+'\\033[0m')\nprint('\\033[36m*' * 30+'\\033[0m')\ntrain_data.D_64.value_counts(dropna=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:41:51.182341Z","iopub.execute_input":"2022-07-31T09:41:51.183028Z","iopub.status.idle":"2022-07-31T09:41:51.837719Z","shell.execute_reply.started":"2022-07-31T09:41:51.182989Z","shell.execute_reply":"2022-07-31T09:41:51.836709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# --- Setting Colors, Labels, Order ---\npurple_grad = ['#491D8B', '#6929C4', '#8A3FFC', '#A56EFF', '#BE95FF']\ncolors=purple_grad\nlabels=train_data['D_66'].dropna().unique()\norder=train_data['D_66'].value_counts().index\n\n# --- Size for Both Figures ---\nplt.figure(figsize=(18, 8))\nplt.suptitle('D_66 Distribution', fontweight='heavy', fontsize='16', fontfamily='sans-serif', \n             color=black_grad[0])\n\n# --- Histogram ---\ncountplt = plt.subplot(1, 2, 1)\nplt.title('Histogram', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nax = sns.countplot(x='D_66', data=train_data, palette=colors, order=order, edgecolor=black_grad[2], alpha=0.85)\nfor rect in ax.patches:\n    ax.text (rect.get_x()+rect.get_width()/2, rect.get_height()+20,rect.get_height(), horizontalalignment='center', \n             fontsize=12, bbox=dict(facecolor='none', edgecolor=black_grad[0], linewidth=0.15, boxstyle='round'))\nplt.tight_layout(rect=[0, 0.04, 1, 0.965])\nplt.xlabel('D_66 Distribution', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.ylabel('Total', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.grid(axis='y', alpha=0.4)\ncountplt\n\n# --- Pie Chart ---\nplt.subplot(1, 2, 2)\nplt.title('Pie Chart', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nplt.pie(train_data['D_66'].value_counts(), colors=colors, labels=order, pctdistance=0.67, autopct='%.2f%%', \n        wedgeprops=dict(alpha=0.8, edgecolor=black_grad[1]), textprops={'fontsize':12})\ncentre=plt.Circle((0, 0), 0.45, fc='white', edgecolor=black_grad[1])\nplt.gcf().gca().add_artist(centre);\n\n# --- Count Categorical Labels w/out Dropping Null Walues ---\nprint('\\033[36m*' * 30)\nprint('\\033[1m'+'.: D_66 Total :.'+'\\033[0m')\nprint('\\033[36m*' * 30+'\\033[0m')\ntrain_data.D_66.value_counts(dropna=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:41:58.348369Z","iopub.execute_input":"2022-07-31T09:41:58.349062Z","iopub.status.idle":"2022-07-31T09:41:58.814292Z","shell.execute_reply.started":"2022-07-31T09:41:58.349021Z","shell.execute_reply":"2022-07-31T09:41:58.813379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# --- Setting Colors, Labels, Order ---\npurple_grad = ['#491D8B', '#6929C4', '#8A3FFC', '#A56EFF', '#BE95FF']\ncolors=purple_grad\nlabels=train_data['D_68'].dropna().unique()\norder=train_data['D_68'].value_counts().index\n\n# --- Size for Both Figures ---\nplt.figure(figsize=(18, 8))\nplt.suptitle('D_68 Distribution', fontweight='heavy', fontsize='16', fontfamily='sans-serif', \n             color=black_grad[0])\n\n# --- Histogram ---\ncountplt = plt.subplot(1, 2, 1)\nplt.title('Histogram', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nax = sns.countplot(x='D_68', data=train_data, palette=colors, order=order, edgecolor=black_grad[2], alpha=0.85)\nfor rect in ax.patches:\n    ax.text (rect.get_x()+rect.get_width()/2, rect.get_height()+20,rect.get_height(), horizontalalignment='center', \n             fontsize=12, bbox=dict(facecolor='none', edgecolor=black_grad[0], linewidth=0.15, boxstyle='round'))\nplt.tight_layout(rect=[0, 0.04, 1, 0.965])\nplt.xlabel('D_68 Distribution', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.ylabel('Total', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.grid(axis='y', alpha=0.4)\ncountplt\n\n# --- Pie Chart ---\nplt.subplot(1, 2, 2)\nplt.title('Pie Chart', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nplt.pie(train_data['D_68'].value_counts(), colors=colors, labels=order, pctdistance=0.67, autopct='%.2f%%', \n        wedgeprops=dict(alpha=0.8, edgecolor=black_grad[1]), textprops={'fontsize':12})\ncentre=plt.Circle((0, 0), 0.45, fc='white', edgecolor=black_grad[1])\nplt.gcf().gca().add_artist(centre);\n\n# --- Count Categorical Labels w/out Dropping Null Walues ---\nprint('\\033[36m*' * 30)\nprint('\\033[1m'+'.: D_68 Total :.'+'\\033[0m')\nprint('\\033[36m*' * 30+'\\033[0m')\ntrain_data.D_68.value_counts(dropna=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:42:05.079935Z","iopub.execute_input":"2022-07-31T09:42:05.080273Z","iopub.status.idle":"2022-07-31T09:42:05.773296Z","shell.execute_reply.started":"2022-07-31T09:42:05.080243Z","shell.execute_reply":"2022-07-31T09:42:05.772392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"font-family: Trebuchet MS; background-color: #db8a8a; color: #FFFFFF; padding: 12px; line-height: 1.5;\">4.1.3| Analyzing the categorical variables w.r.t Target 🔍</div>\n<div style=\"font-family: Segoe UI; line-height: 2; color: #000000; text-align: justify\">\n    👉 This section will focused on <b>Analyzing the categorical variables w.r.t Target</b> before pre-process the data.\n</div>\n","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12,10))\nplt.title('Histogram', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nax = sns.countplot(x='B_30', data=train_data,hue='target', palette=colors, order=order, edgecolor=black_grad[2], alpha=0.85)\nplt.tight_layout(rect=[0, 0.04, 1, 0.965])\nplt.xlabel('B_30 Distribution w.r.t target', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.ylabel('Total', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.grid(axis='y', alpha=0.4)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:42:12.888329Z","iopub.execute_input":"2022-07-31T09:42:12.889365Z","iopub.status.idle":"2022-07-31T09:42:13.599290Z","shell.execute_reply.started":"2022-07-31T09:42:12.889320Z","shell.execute_reply":"2022-07-31T09:42:13.598390Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12,10))\nplt.title('Histogram', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nax = sns.countplot(x='B_38', data=train_data,hue='target', palette=colors, order=order, edgecolor=black_grad[2], alpha=0.85)\nplt.tight_layout(rect=[0, 0.04, 1, 0.965])\nplt.xlabel('B_38 Distribution w.r.t target', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.ylabel('Total', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.grid(axis='y', alpha=0.4)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:42:19.248249Z","iopub.execute_input":"2022-07-31T09:42:19.248751Z","iopub.status.idle":"2022-07-31T09:42:20.332246Z","shell.execute_reply.started":"2022-07-31T09:42:19.248706Z","shell.execute_reply":"2022-07-31T09:42:20.331325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12,10))\nplt.title('Histogram', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nax = sns.countplot(x='D_114', data=train_data,hue='target', palette=colors, order=order, edgecolor=black_grad[2], alpha=0.85)\nplt.tight_layout(rect=[0, 0.04, 1, 0.965])\nplt.xlabel('D_114 Distribution w.r.t target', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.ylabel('Total', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.grid(axis='y', alpha=0.4)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:42:24.875087Z","iopub.execute_input":"2022-07-31T09:42:24.875451Z","iopub.status.idle":"2022-07-31T09:42:25.539350Z","shell.execute_reply.started":"2022-07-31T09:42:24.875398Z","shell.execute_reply":"2022-07-31T09:42:25.538334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12,10))\nplt.title('Histogram', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nax = sns.countplot(x='D_116', data=train_data,hue='target', palette=colors, order=order, edgecolor=black_grad[2], alpha=0.85)\nplt.tight_layout(rect=[0, 0.04, 1, 0.965])\nplt.xlabel('D_116 Distribution w.r.t target', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.ylabel('Total', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.grid(axis='y', alpha=0.4)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:42:30.116800Z","iopub.execute_input":"2022-07-31T09:42:30.117148Z","iopub.status.idle":"2022-07-31T09:42:30.910512Z","shell.execute_reply.started":"2022-07-31T09:42:30.117115Z","shell.execute_reply":"2022-07-31T09:42:30.909460Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12,10))\nplt.title('Histogram', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nax = sns.countplot(x='D_117', data=train_data,hue='target', palette=colors, order=order, edgecolor=black_grad[2], alpha=0.85)\nplt.tight_layout(rect=[0, 0.04, 1, 0.965])\nplt.xlabel('D_117 Distribution w.r.t target', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.ylabel('Total', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.grid(axis='y', alpha=0.4)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:42:35.905654Z","iopub.execute_input":"2022-07-31T09:42:35.906557Z","iopub.status.idle":"2022-07-31T09:42:36.553395Z","shell.execute_reply.started":"2022-07-31T09:42:35.906517Z","shell.execute_reply":"2022-07-31T09:42:36.552371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12,10))\nplt.title('Histogram', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nax = sns.countplot(x='D_120', data=train_data,hue='target', palette=colors, order=order, edgecolor=black_grad[2], alpha=0.85)\nplt.tight_layout(rect=[0, 0.04, 1, 0.965])\nplt.xlabel('D_120 Distribution w.r.t target', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.ylabel('Total', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.grid(axis='y', alpha=0.4)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:42:42.099499Z","iopub.execute_input":"2022-07-31T09:42:42.099977Z","iopub.status.idle":"2022-07-31T09:42:42.838308Z","shell.execute_reply.started":"2022-07-31T09:42:42.099935Z","shell.execute_reply":"2022-07-31T09:42:42.836493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12,10))\nplt.title('Histogram', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nax = sns.countplot(x='D_126', data=train_data,hue='target', palette=colors, order=order, edgecolor=black_grad[2], alpha=0.85)\nplt.tight_layout(rect=[0, 0.04, 1, 0.965])\nplt.xlabel('D_126 Distribution w.r.t target', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.ylabel('Total', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.grid(axis='y', alpha=0.4)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:42:49.528980Z","iopub.execute_input":"2022-07-31T09:42:49.529338Z","iopub.status.idle":"2022-07-31T09:42:50.201638Z","shell.execute_reply.started":"2022-07-31T09:42:49.529304Z","shell.execute_reply":"2022-07-31T09:42:50.200718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12,10))\nplt.title('Histogram', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nax = sns.countplot(x='D_66', data=train_data,hue='target', palette=colors, order=order, edgecolor=black_grad[2], alpha=0.85)\nplt.tight_layout(rect=[0, 0.04, 1, 0.965])\nplt.xlabel('D_66 Distribution w.r.t target', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.ylabel('Total', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.grid(axis='y', alpha=0.4)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:42:55.045661Z","iopub.execute_input":"2022-07-31T09:42:55.046040Z","iopub.status.idle":"2022-07-31T09:42:55.516290Z","shell.execute_reply.started":"2022-07-31T09:42:55.046003Z","shell.execute_reply":"2022-07-31T09:42:55.515294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12,10))\nplt.title('Histogram', fontweight='bold', fontsize=14, fontfamily='sans-serif', color=black_grad[0])\nax = sns.countplot(x='D_68', data=train_data,hue='target', palette=colors, order=order, edgecolor=black_grad[2], alpha=0.85)\nplt.tight_layout(rect=[0, 0.04, 1, 0.965])\nplt.xlabel('D_68 Distribution w.r.t target', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.ylabel('Total', fontweight='bold', fontsize=11, fontfamily='sans-serif', color=black_grad[1])\nplt.grid(axis='y', alpha=0.4)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:42:59.980991Z","iopub.execute_input":"2022-07-31T09:42:59.981332Z","iopub.status.idle":"2022-07-31T09:43:00.693324Z","shell.execute_reply.started":"2022-07-31T09:42:59.981301Z","shell.execute_reply":"2022-07-31T09:43:00.692334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"font-family: Trebuchet MS; background-color: #db8a8a; color: #FFFFFF; padding: 12px; line-height: 1.5;\">4.1.4| Analyzing the numerical attributes 🔍</div>\n<div style=\"font-family: Segoe UI; line-height: 2; color: #000000; text-align: justify\">\n    👉 This section will focused on <b>Analyzing the numerical attributes</b> before pre-process the data.\n</div>\n","metadata":{}},{"cell_type":"markdown","source":"## Data analysis on the Balance attributes","metadata":{}},{"cell_type":"code","source":"# Categorcal columns \n\ncategorical_cols=['B_30', 'B_38', 'D_63', 'D_64', 'D_66', 'D_68',\n          'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'target']","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:43:14.422199Z","iopub.execute_input":"2022-07-31T09:43:14.422567Z","iopub.status.idle":"2022-07-31T09:43:14.427432Z","shell.execute_reply.started":"2022-07-31T09:43:14.422535Z","shell.execute_reply":"2022-07-31T09:43:14.426471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\ncols=[col for col in train_data.columns if (col.startswith(('B','T'))) & (col not in categorical_cols[:-1])]\n\ntrain_balance_cols =train_data[cols]","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:43:24.600495Z","iopub.execute_input":"2022-07-31T09:43:24.600847Z","iopub.status.idle":"2022-07-31T09:43:25.266796Z","shell.execute_reply.started":"2022-07-31T09:43:24.600816Z","shell.execute_reply":"2022-07-31T09:43:25.265797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_balance_cols.select_dtypes(exclude='object').describe().T.style.background_gradient(cmap='RdPu').set_properties(**{'font-family': 'Segoe UI'})","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:43:30.524518Z","iopub.execute_input":"2022-07-31T09:43:30.525088Z","iopub.status.idle":"2022-07-31T09:44:01.515850Z","shell.execute_reply.started":"2022-07-31T09:43:30.525049Z","shell.execute_reply":"2022-07-31T09:44:01.514940Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# --- Plot Missing Values ---\nmso.bar(train_balance_cols, fontsize=9, color=[purple_grad[0], purple_grad[0], purple_grad[0], purple_grad[0], purple_grad[0], purple_grad[0],\n                               purple_grad[0], purple_grad[0], purple_grad[0], purple_grad[0], purple_grad[1], purple_grad[1]], \n        figsize=(15, 8), sort='descending', labels=True)\n\n# --- Title & Subtitle Settings ---\nplt.suptitle('Missing Values in each Columns', fontweight='heavy', x=0.124, y=1.22, ha='left',fontsize='16', \n             fontfamily='sans-serif', color=black_grad[0])\nplt.title('Almost all columns have  missing value.\\n\\nThe total of missing values in each column is less than 25%, which means that imputation can still be done to fill in the missing values in the\\ntwo columns.', \n          fontsize='8', fontfamily='sans-serif', loc='left', color=black_grad[1], pad=5)\nplt.grid(axis='both', alpha=0);\n\n# --- Total Missing Values in each Columns ---\nprint('\\033[36m*' * 43)\nprint('\\033[1m'+'.: Total Missing Values in each Columns :.'+'\\033[0m')\nprint('\\033[36m*' * 43+'\\033[0m')\ntrain_balance_cols.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:44:15.387818Z","iopub.execute_input":"2022-07-31T09:44:15.388166Z","iopub.status.idle":"2022-07-31T09:44:21.794383Z","shell.execute_reply.started":"2022-07-31T09:44:15.388134Z","shell.execute_reply":"2022-07-31T09:44:21.793361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#plot distribution for selected B variables with respect to target label\nnrows = 8\nncols = 5\nfig, axes = plt.subplots(figsize=(24,22)) \nb_columns=['B_1', 'B_2', 'B_3', 'B_4', 'B_5', 'B_6', 'B_7', 'B_8', 'B_9', 'B_10',\n       'B_11', 'B_12', 'B_13', 'B_14', 'B_15', 'B_16', 'B_17', 'B_18', 'B_19',\n       'B_20', 'B_21', 'B_22', 'B_23', 'B_24', 'B_25', 'B_26', 'B_27', 'B_28',\n       'B_29', 'B_31', 'B_32', 'B_33', 'B_36', 'B_37', 'B_39', 'B_40', 'B_41',\n       'B_42']\nfor i, col in enumerate(b_columns):\n    ax=fig.add_subplot(nrows, ncols, i+1)\n    sns.kdeplot(x=train_data[col],hue=train_data['target'], multiple=\"stack\",palette=[\"#FF3333\" ,\"#00CC00\"],ax=ax)\n    \nfig.tight_layout()  \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:45:32.202917Z","iopub.execute_input":"2022-07-31T09:45:32.203260Z","iopub.status.idle":"2022-07-31T09:57:01.261070Z","shell.execute_reply.started":"2022-07-31T09:45:32.203229Z","shell.execute_reply":"2022-07-31T09:57:01.260238Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del train_data[col]","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:57:09.661201Z","iopub.execute_input":"2022-07-31T09:57:09.662162Z","iopub.status.idle":"2022-07-31T09:57:09.668574Z","shell.execute_reply.started":"2022-07-31T09:57:09.662110Z","shell.execute_reply":"2022-07-31T09:57:09.667543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr=train_balance_cols.corr()\nplt.figure(figsize=(24,22))\nax = sns.heatmap(corr,cmap=\"YlGnBu\", linewidths=.5, vmin=-1, vmax=1, center=0, annot=True, annot_kws={'fontsize':12,'fontweight':'bold'}, cbar=False,fmt='.2f' )\nplt.yticks(rotation=0)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T09:44:39.347562Z","iopub.execute_input":"2022-07-31T09:44:39.347903Z","iopub.status.idle":"2022-07-31T09:45:03.630565Z","shell.execute_reply.started":"2022-07-31T09:44:39.347873Z","shell.execute_reply":"2022-07-31T09:45:03.629510Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del train_balance_cols","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data analysis on the Spend attributes","metadata":{}},{"cell_type":"code","source":"cols=[col for col in train_data.columns if (col.startswith(('S','T'))) & (col not in categorical_cols[:-1])]\n\ntrain_spend_cols =train_data[cols]","metadata":{"execution":{"iopub.status.busy":"2022-07-30T07:05:20.223494Z","iopub.execute_input":"2022-07-30T07:05:20.223859Z","iopub.status.idle":"2022-07-30T07:05:20.627219Z","shell.execute_reply.started":"2022-07-30T07:05:20.223824Z","shell.execute_reply":"2022-07-30T07:05:20.626198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_spend_cols.select_dtypes(exclude='object').describe().T.style.background_gradient(cmap='RdPu').set_properties(**{'font-family': 'Segoe UI'})","metadata":{"execution":{"iopub.status.busy":"2022-07-30T07:05:20.628862Z","iopub.execute_input":"2022-07-30T07:05:20.629249Z","iopub.status.idle":"2022-07-30T07:05:38.428201Z","shell.execute_reply.started":"2022-07-30T07:05:20.629211Z","shell.execute_reply":"2022-07-30T07:05:38.426867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# --- Plot Missing Values ---\nmso.bar(train_spend_cols, fontsize=9, color=[purple_grad[0], purple_grad[0], purple_grad[0], purple_grad[0], purple_grad[0], purple_grad[0],\n                               purple_grad[0], purple_grad[0], purple_grad[0], purple_grad[0], purple_grad[1], purple_grad[1]], \n        figsize=(15, 8), sort='descending', labels=True)\n\n# --- Title & Subtitle Settings ---\nplt.suptitle('Missing Values in each Columns', fontweight='heavy', x=0.124, y=1.22, ha='left',fontsize='16', \n             fontfamily='sans-serif', color=black_grad[0])\nplt.title('Almost all columns have no missing value except \\n\\nThe total of missing values in each column is less than 25%, which means that imputation can still be done to fill in the missing values in the\\nthe columns.', \n          fontsize='8', fontfamily='sans-serif', loc='left', color=black_grad[1], pad=5)\nplt.grid(axis='both', alpha=0);\n\n# --- Total Missing Values in each Columns ---\nprint('\\033[36m*' * 43)\nprint('\\033[1m'+'.: Total Missing Values in each Columns :.'+'\\033[0m')\nprint('\\033[36m*' * 43+'\\033[0m')\ntrain_spend_cols.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-30T07:05:38.429990Z","iopub.execute_input":"2022-07-30T07:05:38.430407Z","iopub.status.idle":"2022-07-30T07:05:42.383793Z","shell.execute_reply.started":"2022-07-30T07:05:38.430368Z","shell.execute_reply":"2022-07-30T07:05:42.382670Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#plot distribution for selected B variables with respect to target label\nnrows = 5\nncols = 5\nfig, axes = plt.subplots(figsize=(24,22)) \ns_columns=['S_2', 'S_3', 'S_5', 'S_6', 'S_7', 'S_8', 'S_9', 'S_11', 'S_12', 'S_13',\n       'S_15', 'S_16', 'S_17', 'S_18', 'S_19', 'S_20', 'S_22', 'S_23', 'S_24',\n       'S_25', 'S_26', 'S_27']\nfor i, col in enumerate(s_columns):\n    ax=fig.add_subplot(nrows, ncols, i+1)\n    sns.kdeplot(x=train_data[col],hue=train_data['target'], multiple=\"stack\",palette=[\"#FF3333\" ,\"#00CC00\"],ax=ax)\n    \nfig.tight_layout()  \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-30T07:05:42.394794Z","iopub.execute_input":"2022-07-30T07:05:42.395225Z","iopub.status.idle":"2022-07-30T07:12:50.443086Z","shell.execute_reply.started":"2022-07-30T07:05:42.395187Z","shell.execute_reply":"2022-07-30T07:12:50.441858Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del train_data[col]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr=train_spend_cols.corr()\nplt.figure(figsize=(24,22))\nax = sns.heatmap(corr,cmap=\"YlGnBu\", linewidths=.5, vmin=-1, vmax=1, center=0, annot=True, annot_kws={'fontsize':12,'fontweight':'bold'}, cbar=False,fmt='.2f' )\nplt.yticks(rotation=0)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-30T07:12:50.445053Z","iopub.execute_input":"2022-07-30T07:12:50.445418Z","iopub.status.idle":"2022-07-30T07:12:59.745560Z","shell.execute_reply.started":"2022-07-30T07:12:50.445359Z","shell.execute_reply":"2022-07-30T07:12:59.744627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del train_spend_cols","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data analysis on the Deliquency attributes","metadata":{}},{"cell_type":"code","source":"cols=[col for col in train_data.columns if (col.startswith(('D','T'))) & (col not in categorical_cols[:-1])]\n\ntrain_deliquency_cols =train_data[cols]","metadata":{"execution":{"iopub.status.busy":"2022-07-30T07:12:59.746961Z","iopub.execute_input":"2022-07-30T07:12:59.747476Z","iopub.status.idle":"2022-07-30T07:13:01.315813Z","shell.execute_reply.started":"2022-07-30T07:12:59.747442Z","shell.execute_reply":"2022-07-30T07:13:01.314793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_deliquency_cols.select_dtypes(exclude='object').describe().T.style.background_gradient(cmap='RdPu').set_properties(**{'font-family': 'Segoe UI'})","metadata":{"execution":{"iopub.status.busy":"2022-07-30T07:13:01.317279Z","iopub.execute_input":"2022-07-30T07:13:01.317637Z","iopub.status.idle":"2022-07-30T07:14:14.280534Z","shell.execute_reply.started":"2022-07-30T07:13:01.317600Z","shell.execute_reply":"2022-07-30T07:14:14.279276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# --- Plot Missing Values ---\nmso.bar(train_deliquency_cols, fontsize=9, color=[purple_grad[0], purple_grad[0], purple_grad[0], purple_grad[0], purple_grad[0], purple_grad[0],\n                               purple_grad[0], purple_grad[0], purple_grad[0], purple_grad[0], purple_grad[1], purple_grad[1]], \n        figsize=(24, 22), sort='descending', labels=True)\n\n# --- Title & Subtitle Settings ---\nplt.suptitle('Missing Values in each Columns', fontweight='heavy', x=0.124, y=1.22, ha='left',fontsize='16', \n             fontfamily='sans-serif', color=black_grad[0])\nplt.title('Almost all columns have no missing value except \\n\\nThe total of missing values in each column is less than 25%, which means that imputation can still be done to fill in the missing values in the\\nthe columns.', \n          fontsize='8', fontfamily='sans-serif', loc='left', color=black_grad[1], pad=5)\nplt.grid(axis='both', alpha=0);\n\n# --- Total Missing Values in each Columns ---\nprint('\\033[36m*' * 43)\nprint('\\033[1m'+'.: Total Missing Values in each Columns :.'+'\\033[0m')\nprint('\\033[36m*' * 43+'\\033[0m')\ntrain_deliquency_cols.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-30T07:14:14.281956Z","iopub.execute_input":"2022-07-30T07:14:14.282408Z","iopub.status.idle":"2022-07-30T07:14:31.995502Z","shell.execute_reply.started":"2022-07-30T07:14:14.282372Z","shell.execute_reply":"2022-07-30T07:14:31.994605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#plot distribution for selected B variables with respect to target label\nnrows = 18\nncols = 5\nfig, axes = plt.subplots(figsize=(24,22)) \nd_columns=['D_39', 'D_41', 'D_42', 'D_43', 'D_44', 'D_45', 'D_46', 'D_47', 'D_48',\n       'D_49', 'D_50', 'D_51', 'D_52', 'D_53', 'D_54', 'D_55', 'D_56', 'D_58',\n       'D_59', 'D_60', 'D_61', 'D_62', 'D_65', 'D_69', 'D_70', 'D_71', 'D_72',\n       'D_73', 'D_74', 'D_75', 'D_76', 'D_77', 'D_78', 'D_79', 'D_80', 'D_81',\n       'D_82', 'D_83', 'D_84', 'D_86', 'D_88', 'D_89', 'D_91', 'D_92',\n       'D_93', 'D_94', 'D_96', 'D_102', 'D_103', 'D_104', 'D_105',\n       'D_107', 'D_108', 'D_109', 'D_110', 'D_111', 'D_112', 'D_113', 'D_115',\n       'D_118', 'D_119', 'D_121', 'D_122', 'D_123', 'D_124', 'D_125', 'D_127',\n       'D_128', 'D_129', 'D_130', 'D_131', 'D_132', 'D_133', 'D_134', 'D_135',\n       'D_136', 'D_137', 'D_138', 'D_139', 'D_140', 'D_141', 'D_142', 'D_143',\n       'D_144', 'D_145']\nfor i, col in enumerate(d_columns):\n    ax=fig.add_subplot(nrows, ncols, i+1)\n    sns.kdeplot(x=train_data[col],hue=train_data['target'], multiple=\"stack\",palette=[\"#FF3333\" ,\"#00CC00\"],ax=ax)\n    \nfig.tight_layout()  \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-30T07:14:32.007849Z","iopub.execute_input":"2022-07-30T07:14:32.008483Z","iopub.status.idle":"2022-07-30T07:37:39.247282Z","shell.execute_reply.started":"2022-07-30T07:14:32.008448Z","shell.execute_reply":"2022-07-30T07:37:39.246425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del train_data[col]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del train_deliquency_cols","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Work In progress","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}