{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**Created by Sanskar Hasija**\n\n**🚀AMEX-Default Prediction- Detailed EDA📊📈**\n\n**26 May 2022**\n","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","papermill":{"duration":0.072306,"end_time":"2022-04-01T23:16:53.76377","exception":false,"start_time":"2022-04-01T23:16:53.691464","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# <center> AMEX-DEFAULT PREDICTION- DETAILED EDA📊 </center>\n## <center>If you find this notebook useful, support with an upvote👍</center>","metadata":{"papermill":{"duration":0.067511,"end_time":"2022-04-01T23:16:53.903176","exception":false,"start_time":"2022-04-01T23:16:53.835665","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# Table of Contents\n<a id=\"toc\"></a>\n- [1. Introduction](#1)\n- [2. Imports](#2)\n- [3. Data Loading and Preperation](#3)\n    - [3.1 Exploring Train Data](#3.1)\n    - [3.2 Exploring Test Data](#3.2)\n    - [3.3 Submission File](#3.3)\n- [4. EDA](#4)\n    - [4.1 Null Value Distribution](#4.1)\n    - [4.2 Continuos and Categorical Data Distribution](#4.2)\n    - [4.3 Target Distribution ](#4.3)\n    - [4.4 Continuos Features Distribution  ](#4.4)\n    - [4.5 Categorical Features Distribution ](#4.5)","metadata":{"papermill":{"duration":0.06763,"end_time":"2022-04-01T23:16:54.040531","exception":false,"start_time":"2022-04-01T23:16:53.972901","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<a id=\"1\"></a>\n# **<center><span style=\"color:#00BFC4;\">Introduction  </span></center>**","metadata":{"papermill":{"duration":0.067786,"end_time":"2022-04-01T23:16:54.175996","exception":false,"start_time":"2022-04-01T23:16:54.10821","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"![](https://raw.githubusercontent.com/sanskar-hasija/kaggle/main/images/amex-header.png)","metadata":{"papermill":{"duration":0.067962,"end_time":"2022-04-01T23:16:54.311814","exception":false,"start_time":"2022-04-01T23:16:54.243852","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"**The competition is organised by `American Express` and for `Credit default prediction`**\n\n**In this competition, you are supposed to predict predict Credit default prediction.Submissions are evaluated on a custom evaluation metric which is described as follows :**\n\n<center><b>M = 0.5*(G+D)</b></center><br>\n\n\n\n<b>Here G is the Normalized Gini Coefficient,and D is the default rate captured at 4%</b>","metadata":{"papermill":{"duration":0.068232,"end_time":"2022-04-01T23:16:54.448761","exception":false,"start_time":"2022-04-01T23:16:54.380529","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<a href=\"#toc\" role=\"button\" aria-pressed=\"true\" >⬆️Back to Table of Contents ⬆️</a>","metadata":{"papermill":{"duration":0.067049,"end_time":"2022-04-01T23:16:54.583724","exception":false,"start_time":"2022-04-01T23:16:54.516675","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<a id=\"2\"></a>\n# **<center><span style=\"color:#00BFC4;\">Imports  </span></center>**","metadata":{"papermill":{"duration":0.068173,"end_time":"2022-04-01T23:16:54.720368","exception":false,"start_time":"2022-04-01T23:16:54.652195","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport plotly.express as px\nimport matplotlib.pyplot as plt\nimport plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\n\n\nimport time\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"_kg_hide-input":false,"papermill":{"duration":2.518164,"end_time":"2022-04-01T23:18:14.667553","exception":false,"start_time":"2022-04-01T23:18:12.149389","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-05-25T23:47:34.493253Z","iopub.execute_input":"2022-05-25T23:47:34.494475Z","iopub.status.idle":"2022-05-25T23:47:36.671521Z","shell.execute_reply.started":"2022-05-25T23:47:34.494351Z","shell.execute_reply":"2022-05-25T23:47:36.670008Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a href=\"#toc\" role=\"button\" aria-pressed=\"true\" >⬆️Back to Table of Contents ⬆️</a>","metadata":{"papermill":{"duration":0.069467,"end_time":"2022-04-01T23:18:14.807198","exception":false,"start_time":"2022-04-01T23:18:14.737731","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<a id=\"3\"></a>\n# **<center><span style=\"color:#00BFC4;\">Data Loading and Preparation </span></center>**","metadata":{"papermill":{"duration":0.068777,"end_time":"2022-04-01T23:18:14.945385","exception":false,"start_time":"2022-04-01T23:18:14.876608","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"##### I have created a Parquet version of dataset files for faster loading under the constrain of low memory in Kaggle Kernels.\n\n#### Dataset link - www.kaggle.com/datasets/odins0n/amex-parquet\n#### Example Notebook to open Parquet files from the dataset - www.kaggle.com/code/odins0n/load-parquet-files-with-low-memory","metadata":{}},{"cell_type":"code","source":"train = pd.read_parquet('../input/amex-parquet/train_data.parquet')\nsubmission = pd.read_csv(\"../input/amex-default-prediction/sample_submission.csv\")\nRANDOM_STATE = 12 ","metadata":{"_kg_hide-input":true,"papermill":{"duration":0.190055,"end_time":"2022-04-01T23:18:15.212934","exception":false,"start_time":"2022-04-01T23:18:15.022879","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-05-25T23:47:36.678609Z","iopub.execute_input":"2022-05-25T23:47:36.679007Z","iopub.status.idle":"2022-05-25T23:48:08.744753Z","shell.execute_reply.started":"2022-05-25T23:47:36.678973Z","shell.execute_reply":"2022-05-25T23:48:08.743698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-05-25T23:48:08.746371Z","iopub.execute_input":"2022-05-25T23:48:08.746833Z","iopub.status.idle":"2022-05-25T23:48:08.892554Z","shell.execute_reply.started":"2022-05-25T23:48:08.746786Z","shell.execute_reply":"2022-05-25T23:48:08.891711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:#e76f51;\"> Column Descriptions  : </span>\n\n`customer_ID` = Unique Customer ID<br>\n`- D_*` = Delinquency variables<br>\n`S_*` = Spend variables<br>\n`P_*` = Payment variables<br>\n`B_*` = Balance variables<br>\n`R_*` = Risk variables<br>\n\n","metadata":{"papermill":{"duration":0.068193,"end_time":"2022-04-01T23:18:15.350875","exception":false,"start_time":"2022-04-01T23:18:15.282682","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<a id=\"3.1\"></a>\n## <span style=\"color:#e76f51;\"> Exploring Train Data : </span>","metadata":{"papermill":{"duration":0.068668,"end_time":"2022-04-01T23:18:15.488781","exception":false,"start_time":"2022-04-01T23:18:15.420113","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:14px; font-family:verdana; line-height: 1.7em;\">\n    📌 &nbsp;<b><u>Observations in Train Data:</u></b><br>\n \n* <i> There are total of <b><u>190</u></b> columns and <b><u>5531451</u></b> rows in <b><u>train</u></b> data.</i><br>\n* <i> Train data contains <b><u>890116722</u></b> observation with <b><u>160858968</u></b>  missing values.</i><br>\n</div>","metadata":{"papermill":{"duration":0.069187,"end_time":"2022-04-01T23:18:15.626805","exception":false,"start_time":"2022-04-01T23:18:15.557618","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"### <span style=\"color:#e76f51;\"> Quick view of Train Data : </span>","metadata":{"papermill":{"duration":0.068616,"end_time":"2022-04-01T23:18:15.764223","exception":false,"start_time":"2022-04-01T23:18:15.695607","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"Below are the first 5 rows of train dataset:","metadata":{"papermill":{"duration":0.067464,"end_time":"2022-04-01T23:18:15.904188","exception":false,"start_time":"2022-04-01T23:18:15.836724","status":"completed"},"tags":[]}},{"cell_type":"code","source":"train.head()","metadata":{"_kg_hide-input":true,"papermill":{"duration":0.098549,"end_time":"2022-04-01T23:18:16.072851","exception":false,"start_time":"2022-04-01T23:18:15.974302","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-05-25T23:48:08.894540Z","iopub.execute_input":"2022-05-25T23:48:08.895104Z","iopub.status.idle":"2022-05-25T23:48:08.935542Z","shell.execute_reply.started":"2022-05-25T23:48:08.895063Z","shell.execute_reply":"2022-05-25T23:48:08.934277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'\\033[94mNumber of rows in train data: {train.shape[0]}')\nprint(f'\\033[94mNumber of columns in train data: {train.shape[1]}')\nprint(f'\\033[94mNumber of values in train data: {train.count().sum()}')\nprint(f'\\033[94mNumber missing values in train data: {sum(train.isna().sum())}')","metadata":{"_kg_hide-input":true,"papermill":{"duration":0.100334,"end_time":"2022-04-01T23:18:16.244456","exception":false,"start_time":"2022-04-01T23:18:16.144122","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-05-25T23:48:08.937132Z","iopub.execute_input":"2022-05-25T23:48:08.937641Z","iopub.status.idle":"2022-05-25T23:48:14.761874Z","shell.execute_reply.started":"2022-05-25T23:48:08.937596Z","shell.execute_reply":"2022-05-25T23:48:14.760963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:#e76f51;\"> Column Wise missing values : </span>","metadata":{"papermill":{"duration":0.069632,"end_time":"2022-04-01T23:18:16.385832","exception":false,"start_time":"2022-04-01T23:18:16.3162","status":"completed"},"tags":[]}},{"cell_type":"code","source":"print(f'\\033[94m')\nprint(train.isna().sum().sort_values(ascending = False))","metadata":{"_kg_hide-input":true,"papermill":{"duration":0.089076,"end_time":"2022-04-01T23:18:16.544989","exception":false,"start_time":"2022-04-01T23:18:16.455913","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-05-25T23:48:14.763034Z","iopub.execute_input":"2022-05-25T23:48:14.764252Z","iopub.status.idle":"2022-05-25T23:48:17.483588Z","shell.execute_reply.started":"2022-05-25T23:48:14.764194Z","shell.execute_reply":"2022-05-25T23:48:17.482777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:#e76f51;\"> Basic statistics of training data : </span>","metadata":{"papermill":{"duration":0.070649,"end_time":"2022-04-01T23:18:16.687369","exception":false,"start_time":"2022-04-01T23:18:16.61672","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"Below is the basic statistics for each variables which contain information on `count`, `mean`, `standard deviation`, `minimum`, `1st quartile`, `median`, `3rd quartile` and `maximum`.","metadata":{"papermill":{"duration":0.070386,"end_time":"2022-04-01T23:18:16.82821","exception":false,"start_time":"2022-04-01T23:18:16.757824","status":"completed"},"tags":[]}},{"cell_type":"code","source":"train.iloc[:, :-1].describe().T.sort_values(by='std' , ascending = False)\\\n                     .style.background_gradient(cmap='GnBu')\\\n                     .bar(subset=[\"max\"], color='#F8766D')\\\n                     .bar(subset=[\"mean\",], color='#00BFC4')","metadata":{"_kg_hide-input":true,"papermill":{"duration":0.1139,"end_time":"2022-04-01T23:18:17.012174","exception":false,"start_time":"2022-04-01T23:18:16.898274","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-05-25T23:48:17.485060Z","iopub.execute_input":"2022-05-25T23:48:17.486271Z","iopub.status.idle":"2022-05-25T23:49:06.181076Z","shell.execute_reply.started":"2022-05-25T23:48:17.486220Z","shell.execute_reply":"2022-05-25T23:49:06.180242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"3.3\"></a>\n## <span style=\"color:#e76f51;\"> Submission File </span>","metadata":{"papermill":{"duration":0.07332,"end_time":"2022-04-01T23:18:18.687214","exception":false,"start_time":"2022-04-01T23:18:18.613894","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"### <span style=\"color:#e76f51;\"> Quick view of Submission File </span>","metadata":{"papermill":{"duration":0.072978,"end_time":"2022-04-01T23:18:18.833771","exception":false,"start_time":"2022-04-01T23:18:18.760793","status":"completed"},"tags":[]}},{"cell_type":"code","source":"submission.head()","metadata":{"_kg_hide-input":true,"papermill":{"duration":0.086129,"end_time":"2022-04-01T23:18:18.992298","exception":false,"start_time":"2022-04-01T23:18:18.906169","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-05-25T23:49:06.182433Z","iopub.execute_input":"2022-05-25T23:49:06.182934Z","iopub.status.idle":"2022-05-25T23:49:06.191795Z","shell.execute_reply.started":"2022-05-25T23:49:06.182901Z","shell.execute_reply":"2022-05-25T23:49:06.191006Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a href=\"#toc\" role=\"button\" aria-pressed=\"true\" >⬆️Back to Table of Contents ⬆️</a>","metadata":{"papermill":{"duration":0.075137,"end_time":"2022-04-01T23:18:19.141093","exception":false,"start_time":"2022-04-01T23:18:19.065956","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<a id=\"4\"></a>\n# **<center><span style=\"color:#00BFC4;\"> EDA </span></center>**","metadata":{"papermill":{"duration":0.073449,"end_time":"2022-04-01T23:18:19.288709","exception":false,"start_time":"2022-04-01T23:18:19.21526","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<a id=\"4.1\"></a>\n## <span style=\"color:#e76f51;\"> Null Value Distribution  </span>","metadata":{"papermill":{"duration":0.075799,"end_time":"2022-04-01T23:18:20.015208","exception":false,"start_time":"2022-04-01T23:18:19.939409","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:14px; font-family:verdana; line-height: 1.7em;\">\n    📌 &nbsp;<b><u>Observations in Null Value Distribution :</u></b><br>\n \n* <i> The maximum of missing value in an row is <b><u>102</u></b> and the lowest is <b><u>9</u></b> missing values.</i><br>\n* <i> All the rows have atleast <b><u>9</u></b> missing values.</i><br>\n* <i> <b><u>D88 </u></b>feature hax maximum number of missing values with a total of <b><u>5527586</u></b> missing values.</i><br>\n* <i> <b><u>68 </u></b> features have no missing values whereas <b><u>122</u></b> features have atleast 1 missing values </i><br>\n</div>\n","metadata":{"papermill":{"duration":0.075603,"end_time":"2022-04-01T23:18:20.167656","exception":false,"start_time":"2022-04-01T23:18:20.092053","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<a id=\"4.2.1\"></a>\n### <span style=\"color:#e76f51;\">Column wise Null Value Distribution   </span>","metadata":{"papermill":{"duration":0.07502,"end_time":"2022-04-01T23:18:20.318781","exception":false,"start_time":"2022-04-01T23:18:20.243761","status":"completed"},"tags":[]}},{"cell_type":"code","source":"train_null = pd.DataFrame(train.isna().sum())\ntrain_null = train_null[train_null[0]>0]\ntrain_null = train_null.sort_values(by = 0 ,ascending = True)\nfig = px.bar(x=train_null[0],y=train_null.index,color_discrete_sequence = [\"#DE3163\"])\nfig.update_layout(showlegend=False, \n                  title_text=\"Column Wise Null Value Distribution\", \n                  title_x=0.5,\n                  xaxis_title=\"Missing Value Count\",\n                  yaxis_title=\"Feature Name\")\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-05-25T23:49:19.845946Z","iopub.execute_input":"2022-05-25T23:49:19.846839Z","iopub.status.idle":"2022-05-25T23:49:23.666136Z","shell.execute_reply.started":"2022-05-25T23:49:19.846785Z","shell.execute_reply":"2022-05-25T23:49:23.663756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"4.7.2\"></a>\n### <span style=\"color:#e76f51;\">Row wise Null Value Distribution   </span>","metadata":{"papermill":{"duration":0.075576,"end_time":"2022-04-01T23:18:20.929072","exception":false,"start_time":"2022-04-01T23:18:20.853496","status":"completed"},"tags":[]}},{"cell_type":"code","source":"missing_train_row = train.isna().sum(axis=1)\nmissing_train_row = pd.DataFrame(missing_train_row.value_counts()/train.shape[0]).reset_index()\nmissing_train_row.columns = ['no', 'count']\nmissing_train_row[\"count\"] = missing_train_row[\"count\"]*100\n\nfig = px.bar(x=missing_train_row[\"no\"], \n                     y=missing_train_row[\"count\"] ,\n             color_discrete_sequence = [\"#DE3163\"])\nfig.update_layout(showlegend=False, \n                  title_text=\"Row wise Null Value Distribution\", \n                  title_x=0.5,\n                  xaxis_title=\"Number of Rows\",\n                  yaxis_title=\"Percentage of Missing Values\")\nfig.show()","metadata":{"_kg_hide-input":true,"papermill":{"duration":0.145743,"end_time":"2022-04-01T23:18:21.152699","exception":false,"start_time":"2022-04-01T23:18:21.006956","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-05-25T23:49:23.667731Z","iopub.execute_input":"2022-05-25T23:49:23.668071Z","iopub.status.idle":"2022-05-25T23:49:27.504709Z","shell.execute_reply.started":"2022-05-25T23:49:23.668043Z","shell.execute_reply":"2022-05-25T23:49:27.503757Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:#e76f51;\">Dealing with missing value (reference)  </span>\nSome references on how to deal with missing value:\n- [Missing Values](https://www.kaggle.com/alexisbcook/missing-values) by [Alexis Cook](https://www.kaggle.com/alexisbcook)\n- [Data Cleaning Challenge: Handling missing values](https://www.kaggle.com/rtatman/data-cleaning-challenge-handling-missing-values) by [Rachael Tatman](https://www.kaggle.com/rtatman)\n- [A Guide to Handling Missing values in Python ](https://www.kaggle.com/parulpandey/a-guide-to-handling-missing-values-in-python) by [Parul Pandey](https://www.kaggle.com/parulpandey)\n\nSome models that have capability to handle missing value by default are:\n- XGBoost: https://xgboost.readthedocs.io/en/latest/faq.html\n- LightGBM: https://lightgbm.readthedocs.io/en/latest/Advanced-Topics.html\n- Catboost: https://catboost.ai/docs/concepts/algorithm-missing-values-processing.html","metadata":{"papermill":{"duration":0.102464,"end_time":"2022-04-01T23:18:21.332081","exception":false,"start_time":"2022-04-01T23:18:21.229617","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<a href=\"#toc\" role=\"button\" aria-pressed=\"true\" >⬆️Back to Table of Contents ⬆️</a>","metadata":{"papermill":{"duration":0.076931,"end_time":"2022-04-01T23:18:21.496212","exception":false,"start_time":"2022-04-01T23:18:21.419281","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<a id=\"4.2\"></a>\n## <span style=\"color:#e76f51;\">Continuos and Categorical Data Distribution </span>","metadata":{"papermill":{"duration":0.078675,"end_time":"2022-04-01T23:18:21.653153","exception":false,"start_time":"2022-04-01T23:18:21.574478","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:14px; font-family:verdana; line-height: 1.7em;\">\n    📌 &nbsp;<b><u>Observations in Null Value Distribution :</u></b><br>\n \n* <i> There are a total of  <b><u>190</u></b> features, out of which  <b><u>177</u></b> features are continous, <b><u>1</u></b> feature represents date and <b><u>11</u></b> features are categorical.</i><br>\n</div>","metadata":{"papermill":{"duration":0.077294,"end_time":"2022-04-01T23:18:21.807746","exception":false,"start_time":"2022-04-01T23:18:21.730452","status":"completed"},"tags":[]}},{"cell_type":"code","source":"FEATURES = list(train.columns[2:190])\nTARGET = \"target\"\ncat_features = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\ncont_features = [col for col in FEATURES if col not in cat_features and TARGET]\nlabels=['Categorical', 'Continuos']\nvalues= [len(cat_features), len(cont_features)]\ncolors = ['#DE3163', '#58D68D']\n\nprint(f'\\033[94mTotal number of features: {len(FEATURES) + 2   }')\nprint(f'\\033[94mNumber of categorical features: {len(cat_features)}')\nprint(f'\\033[94mNumber of continuos features: {len(cont_features)}')\n\nfig = go.Figure(data=[go.Pie(\n    labels=labels, \n    values=values, pull=[0.1, 0 ],\n    marker=dict(colors=colors, \n                line=dict(color='#000000', \n                          width=2))\n)])\nfig.show()\n","metadata":{"_kg_hide-input":true,"papermill":{"duration":0.153081,"end_time":"2022-04-01T23:18:22.039531","exception":false,"start_time":"2022-04-01T23:18:21.88645","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-05-25T23:49:54.640251Z","iopub.execute_input":"2022-05-25T23:49:54.640684Z","iopub.status.idle":"2022-05-25T23:49:54.658655Z","shell.execute_reply.started":"2022-05-25T23:49:54.640649Z","shell.execute_reply":"2022-05-25T23:49:54.657280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a href=\"#toc\" role=\"button\" aria-pressed=\"true\" >⬆️Back to Table of Contents ⬆️</a>","metadata":{"papermill":{"duration":0.07837,"end_time":"2022-04-01T23:18:22.197661","exception":false,"start_time":"2022-04-01T23:18:22.119291","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<a id=\"4.3\"></a>\n## <span style=\"color:#e76f51;\">  Target Distribution </span>","metadata":{"papermill":{"duration":0.084501,"end_time":"2022-04-01T23:18:25.007428","exception":false,"start_time":"2022-04-01T23:18:24.922927","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:14px; font-family:verdana; line-height: 1.7em;\">\n    📌 &nbsp;<b><u>Observations in Null Value Distribution :</u></b><br>\n \n* <i>There are two target values - <b><u>0</u></b> and <b><u>1</u></b>.</i><br>\n* <i>Percentage of Target <b><u>0</u></b> and Target <b><u>1</u></b> are <b><u>74.11%</u></b> and <b><u>25.89%</u></b> respectively. </i><br>\n</div>","metadata":{"execution":{"iopub.execute_input":"2022-02-24T00:39:27.972624Z","iopub.status.busy":"2022-02-24T00:39:27.972052Z","iopub.status.idle":"2022-02-24T00:39:27.984571Z","shell.execute_reply":"2022-02-24T00:39:27.983081Z","shell.execute_reply.started":"2022-02-24T00:39:27.972572Z"},"papermill":{"duration":0.083605,"end_time":"2022-04-01T23:18:25.175269","exception":false,"start_time":"2022-04-01T23:18:25.091664","status":"completed"},"tags":[]}},{"cell_type":"code","source":"target_df = pd.DataFrame(train['target'].value_counts()).reset_index()\ntarget_df.columns = ['target', 'count']\nfig = px.bar(data_frame =target_df, \n             x = 'target',\n             y = 'count'\n            ) \nfig.update_traces(marker_color =['#58D68D','#DE3163'], \n                  marker_line_color='rgb(0,0,0)',\n                  marker_line_width=2,)\nfig.update_layout(title = \"Target Distribution\",\n                  template = \"plotly_white\",\n                  title_x = 0.5)\nprint(\"\\033[94mPercentage of Target = 0: {:.2f} %\".format(target_df[\"count\"][0]*100 / (target_df[\"count\"][0]+ target_df[\"count\"][1])))\nprint(\"\\033[94mPercentage of Target = 1: {:.2f} %\".format(target_df[\"count\"][1]* 100 / (target_df[\"count\"][0]+ target_df[\"count\"][1])))\nfig.show()","metadata":{"_kg_hide-input":true,"papermill":{"duration":0.345777,"end_time":"2022-04-01T23:18:25.60636","exception":false,"start_time":"2022-04-01T23:18:25.260583","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-05-25T23:50:05.464102Z","iopub.execute_input":"2022-05-25T23:50:05.464968Z","iopub.status.idle":"2022-05-25T23:50:05.584341Z","shell.execute_reply.started":"2022-05-25T23:50:05.464926Z","shell.execute_reply":"2022-05-25T23:50:05.583138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a href=\"#toc\" role=\"button\" aria-pressed=\"true\" >⬆️Back to Table of Contents ⬆️</a>","metadata":{"papermill":{"duration":0.085541,"end_time":"2022-04-01T23:18:25.77698","exception":false,"start_time":"2022-04-01T23:18:25.691439","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<a id=\"4.4\"></a>\n## <span style=\"color:#e76f51;\"> Continuos Features Distribution  </span>","metadata":{}},{"cell_type":"code","source":"RANDOM_SPLIT = 100000\nncols = 5\nnrows = 36\nn_features = cont_features\nfig, axes = plt.subplots(nrows, ncols, figsize=(25, 15*8))\n\nfor r in range(nrows):\n    for c in range(ncols):\n        if r*ncols+c == len(cont_features):\n            break\n        col = n_features[r*ncols+c]\n        sns.histplot(data= train.iloc[:RANDOM_SPLIT],  x=col, ax=axes[r, c], hue= \"target\", bins = 20, palette =['#DE3163','#58D68D'])\n        axes[r,c].legend()\n        axes[r, c].set_ylabel('')\n        axes[r, c].set_xlabel(col, fontsize=8)\n        axes[r, c].tick_params(labelsize=5, width=0.5)\n        axes[r, c].xaxis.offsetText.set_fontsize(6)\n        axes[r, c].yaxis.offsetText.set_fontsize(4)\nfig.delaxes(axes[35][2])\nfig.delaxes(axes[35][3])   \nfig.delaxes(axes[35][4])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-05-25T23:50:21.571110Z","iopub.execute_input":"2022-05-25T23:50:21.571557Z","iopub.status.idle":"2022-05-25T23:51:07.269575Z","shell.execute_reply.started":"2022-05-25T23:50:21.571524Z","shell.execute_reply":"2022-05-25T23:51:07.268587Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"4.5\"></a>\n## <span style=\"color:#e76f51;\"> Categorical Features Distribution  </span>","metadata":{}},{"cell_type":"code","source":"sns.set_style(style='white')\nncols = 5\nnrows = int(len(cat_features) / ncols + (len(FEATURES) % ncols > 0)) \n\nfig, axes = plt.subplots(nrows, ncols, figsize=(18, 15), facecolor='#EAEAF2')\n\nfor r in range(nrows):\n    for c in range(ncols):\n        if r*ncols+c >= len(cat_features):\n            break\n        col = cat_features[r*ncols+c]\n        sns.countplot(data=train.iloc[:RANDOM_SPLIT] , x = col, ax=axes[r, c], hue = \"target\", palette =['#DE3163','#58D68D'])\n        axes[r, c].set_ylabel('')\n        axes[r, c].set_xlabel(col, fontsize=8, fontweight='bold')\n        axes[r, c].tick_params(labelsize=5, width=0.5)\n        axes[r, c].xaxis.offsetText.set_fontsize(4)\n        axes[r, c].yaxis.offsetText.set_fontsize(4)\nfig.delaxes(axes[2][1])     \nfig.delaxes(axes[2][2]) \nfig.delaxes(axes[2][3]) \nfig.delaxes(axes[2][4]) \nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:14px; font-family:verdana; line-height: 1.7em;\">\n    \n    \n### <center>Work in Progress 🙂</center>\n### <center>If you have any feedback or find anything wrong, please let me know!</center>\n","metadata":{"papermill":{"duration":0.094536,"end_time":"2022-04-01T23:18:44.24344","exception":false,"start_time":"2022-04-01T23:18:44.148904","status":"completed"},"tags":[]}}]}