{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":50160,"databundleVersionId":7602123,"sourceType":"competition"}],"dockerImageVersionId":30646,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Welcome to the starter notebook for Home Credit - Credit Risk Model Stability. In this competition, you will be predicting default of clients based on internal and external information that are available for each client. Scoring is performed using custom metric that not only evaluates the AUC of predictions but also considers the stability of predictions model across the data range of the test set. \n\nCurrent Work done:\n1. Data Loading\n2. Basic EDA\n\nI will be updating notebook as soon as I make some progress best of luck to everyone :)","metadata":{"_kg_hide-input":false,"_kg_hide-output":false}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport polars as pl\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-02-06T08:04:17.793604Z","iopub.execute_input":"2024-02-06T08:04:17.796298Z","iopub.status.idle":"2024-02-06T08:04:19.104326Z","shell.execute_reply.started":"2024-02-06T08:04:17.796258Z","shell.execute_reply":"2024-02-06T08:04:19.103224Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rc = {\n    \"axes.facecolor\": \"#F8F8F8\",\n    \"figure.facecolor\": \"#F8F8F8\",\n    \"axes.edgecolor\": \"#000000\",\n    \"grid.color\": \"#EBEBE7\" + \"30\",\n    \"font.family\": \"serif\",\n    \"axes.labelcolor\": \"#000000\",\n    \"xtick.color\": \"#000000\",\n    \"ytick.color\": \"#000000\",\n    \"grid.alpha\": 0.4,\n}\n\nsns.set(rc=rc)\npalette = ['#302c36', '#037d97', '#E4591E', '#C09741',\n           '#EC5B6D', '#90A6B1', '#6ca957', '#D8E3E2']\n\nfrom colorama import Style, Fore\nblk = Style.BRIGHT + Fore.BLACK\nmgt = Style.BRIGHT + Fore.MAGENTA\nred = Style.BRIGHT + Fore.RED\nblu = Style.BRIGHT + Fore.BLUE\nres = Style.RESET_ALL\n\nplt.style.use('fivethirtyeight')","metadata":{"execution":{"iopub.status.busy":"2024-02-06T05:46:15.272045Z","iopub.execute_input":"2024-02-06T05:46:15.273126Z","iopub.status.idle":"2024-02-06T05:46:15.283407Z","shell.execute_reply.started":"2024-02-06T05:46:15.273084Z","shell.execute_reply":"2024-02-06T05:46:15.282266Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1. Data Exploration\n\nThis dataset contains a large number of tables as a result of utilizing diverse data sources and the varying levels of data aggregation used while preparing the dataset. \n\n- Note: All files listed below are found in both `.csv` and `.parquet` formats.\n\n## 1.1 Base Tables:\nBase tables store the basic information about the observation and `case_id`. This is a unique identification of every observation and you need to use it to join the other tables to base tables.\n\n### Feature Overview of Base Tables:\n- **case_id** - This is the unique identifier for each credit case. You'll need this ID to join relevant tables to the base table.\n- **date_decision** - This refers to the date when a decision was made regarding the approval of the loan.\n- **WEEK_NUM** - This is the week number used for aggregation. In the test sample, WEEK_NUM continues sequentially from the last training value of WEEK_NUM.\n- **MONTH** - This column represents the month and is intended for aggregation purposes.\n- **target** - This is the target value, determined after a certain period based on whether or not the client defaulted on the specific credit case\n\n> All the information is available in overview and data section too.\n\nLet's Explore the train data.","metadata":{}},{"cell_type":"code","source":"TRAIN = '/kaggle/input/home-credit-credit-risk-model-stability/csv_files/train'\nTEST = '/kaggle/input/home-credit-credit-risk-model-stability/csv_files/test'","metadata":{"execution":{"iopub.status.busy":"2024-02-06T08:04:10.985294Z","iopub.execute_input":"2024-02-06T08:04:10.985657Z","iopub.status.idle":"2024-02-06T08:04:11.014715Z","shell.execute_reply.started":"2024-02-06T08:04:10.985624Z","shell.execute_reply":"2024-02-06T08:04:11.013779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_base = pd.read_csv(f'{TRAIN}/train_base.csv')\ntest_base = pd.read_csv(f'{TEST}/test_base.csv')\ndisplay(train_base.head())","metadata":{"execution":{"iopub.status.busy":"2024-02-06T08:04:23.729834Z","iopub.execute_input":"2024-02-06T08:04:23.732122Z","iopub.status.idle":"2024-02-06T08:04:24.912297Z","shell.execute_reply.started":"2024-02-06T08:04:23.732085Z","shell.execute_reply":"2024-02-06T08:04:24.911214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train_base.info())","metadata":{"execution":{"iopub.status.busy":"2024-02-06T08:04:28.189735Z","iopub.execute_input":"2024-02-06T08:04:28.190311Z","iopub.status.idle":"2024-02-06T08:04:28.296748Z","shell.execute_reply.started":"2024-02-06T08:04:28.190283Z","shell.execute_reply":"2024-02-06T08:04:28.295675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- As we can see we have to `NULL` values in base file.` \n- We can also convert `date_decision` into date_time formate. \n\nSeperating year and month from `MONTH` column and converting `date_decision` into date_time formate.","metadata":{}},{"cell_type":"code","source":"train_base['year'] = train_base['MONTH'] // 100  \ntrain_base['month'] = train_base['MONTH'] % 100   \n\ntrain_base['date_decision'] = pd.to_datetime(train_base['date_decision'], format='%Y-%m-%d')\ndisplay(train_base.head())","metadata":{"execution":{"iopub.status.busy":"2024-02-06T10:16:36.628158Z","iopub.execute_input":"2024-02-06T10:16:36.628563Z","iopub.status.idle":"2024-02-06T10:16:36.788265Z","shell.execute_reply.started":"2024-02-06T10:16:36.628535Z","shell.execute_reply":"2024-02-06T10:16:36.787042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train_base.info())","metadata":{"execution":{"iopub.status.busy":"2024-02-06T08:05:54.339798Z","iopub.execute_input":"2024-02-06T08:05:54.340305Z","iopub.status.idle":"2024-02-06T08:05:54.368077Z","shell.execute_reply.started":"2024-02-06T08:05:54.340267Z","shell.execute_reply":"2024-02-06T08:05:54.366999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Target:\nTarget is defined as whether or not the client defaulted on the specific credit case.","metadata":{}},{"cell_type":"code","source":"f,ax=plt.subplots(1,2,figsize=(19,7))\ntrain_base['target'].value_counts().plot.pie(autopct='%1.1f%%',ax=ax[0],shadow=True)\n# ax[0].set_title('Pie-Plot')\nax[0].set_ylabel('')\nsns.countplot(x='target',data=train_base,ax=ax[1])\n# ax[1].set_title('Count-Plot')\nplt.suptitle('Target Value Anaysis')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-06T05:49:31.143140Z","iopub.execute_input":"2024-02-06T05:49:31.143573Z","iopub.status.idle":"2024-02-06T05:49:31.669271Z","shell.execute_reply.started":"2024-02-06T05:49:31.143531Z","shell.execute_reply":"2024-02-06T05:49:31.668093Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- As we can see data is highly imbalanced.\n\n### Month:\nRepresents month.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(18, 4))\nfig = sns.histplot(data=train_base, x='month', hue=\"target\", bins=50,)\n# plt.ylim(0,10000)\nplt.title('Month Distribution')\nplt.show()\n\n# ---------------------------\n\nfig, ax = plt.subplots(figsize=(20, 4))\nfig = sns.histplot(data=train_base, x='month', hue=\"target\", bins=50,)\nplt.ylim(0,10000)\n# plt.title('Month Distribution')\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-02-06T06:26:07.903935Z","iopub.execute_input":"2024-02-06T06:26:07.904571Z","iopub.status.idle":"2024-02-06T06:26:11.034084Z","shell.execute_reply.started":"2024-02-06T06:26:07.904509Z","shell.execute_reply":"2024-02-06T06:26:11.032748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Week:\nRepresents week number In test data `WEEK_NUM` continues sequentially from the last training value of `WEEK_NUM`.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(18, 4))\nfig = sns.histplot(data=train_base, x='WEEK_NUM', hue=\"target\", bins=50)\n# plt.ylim(0,15000)\nplt.ticklabel_format(style = 'plain')\nplt.title('Week Distribution')\nplt.show()\n\n# ----------------------\n\nfig, ax = plt.subplots(figsize=(18, 4))\nfig = sns.histplot(data=train_base, x='WEEK_NUM', hue=\"target\", bins=50)\nplt.ylim(0,10000)\nplt.ticklabel_format(style = 'plain')\n# plt.title('Week Distribution')\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-02-06T06:25:39.853814Z","iopub.execute_input":"2024-02-06T06:25:39.854203Z","iopub.status.idle":"2024-02-06T06:25:43.875266Z","shell.execute_reply.started":"2024-02-06T06:25:39.854173Z","shell.execute_reply":"2024-02-06T06:25:43.873792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Date Decision:\nThis refers to the date when a decision was made regarding the approval of the loan.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(18, 4))\nfig = sns.histplot(data=train_base, x='date_decision', hue=\"target\", bins=50)\n# plt.ylim(0,15000)\n\nplt.title('Date Distribution')\nplt.show()\n\n# ----------------------\n\nfig, ax = plt.subplots(figsize=(18, 4))\nfig = sns.histplot(data=train_base, x='date_decision', hue=\"target\", bins=50)\nplt.ylim(0,10000)\n\n# plt.title('Week Distribution')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-06T06:25:50.609697Z","iopub.execute_input":"2024-02-06T06:25:50.610173Z","iopub.status.idle":"2024-02-06T06:25:54.516090Z","shell.execute_reply.started":"2024-02-06T06:25:50.610135Z","shell.execute_reply":"2024-02-06T06:25:54.514544Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that we have explored the base tables let's explore more tables.\n\n### Depth values:\n\n- depth=0 - These are static features directly tied to a specific `case_id`.\n- depth=1 - Each case_id has an associated historical record, indexed by `num_group1`.\n- depth=2 - Each case_id has an associated historical record, indexed by both `num_group1` and `num_group2`.\n\n## 1.2 Depth = 0\nAs mentioned earlier all this features are directly tied with base data also for depth=0 tables, predictors can be directly used as features. All the depth=0 files are as below:\n\n**1. static_0** - This is internal data source.\n\n- Train Files:\n    - train_static_0_0.csv\n    - train_static_0_1.csv\n    \n- Test Files:\n    - test_static_0_0.csv\n    - test_static_0_1.csv\n    - test_static_0_2.csv\n    \n**2. static_cb_0** - This is external data source.\n\n- Train Files:\n    - train_static_cb_0.csv\n    \n- Test Files:\n    - test_static_cb_0.csv\n    \nLet's load all the files.","metadata":{}},{"cell_type":"code","source":"train_static_0 = pd.concat(\n    [\n        pd.read_csv(f'{TRAIN}/train_static_0_0.csv'), \n        pd.read_csv(f'{TRAIN}/train_static_0_1.csv')\n    ], axis=0)\n\ntrain_static_cb_0 = pd.read_csv(f'{TRAIN}/train_static_cb_0.csv')\n\n# ----------------\n\ntest_static_0 = pd.concat(\n    [\n        pd.read_csv(f'{TEST}/test_static_0_0.csv'), \n        pd.read_csv(f'{TEST}/test_static_0_1.csv'),\n        pd.read_csv(f'{TEST}/test_static_0_2.csv'),\n    ], axis=0)\n\ntest_static_cb_0 = pd.read_csv(f'{TEST}/test_static_cb_0.csv')","metadata":{"execution":{"iopub.status.busy":"2024-02-06T09:48:49.485879Z","iopub.execute_input":"2024-02-06T09:48:49.486445Z","iopub.status.idle":"2024-02-06T09:49:41.561084Z","shell.execute_reply.started":"2024-02-06T09:48:49.486379Z","shell.execute_reply":"2024-02-06T09:49:41.559916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_static_0.head(3)","metadata":{"execution":{"iopub.status.busy":"2024-02-06T10:00:20.366025Z","iopub.execute_input":"2024-02-06T10:00:20.366408Z","iopub.status.idle":"2024-02-06T10:00:20.394213Z","shell.execute_reply.started":"2024-02-06T10:00:20.366381Z","shell.execute_reply":"2024-02-06T10:00:20.392674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_static_cb_0.head(3)","metadata":{"execution":{"iopub.status.busy":"2024-02-06T10:00:19.370990Z","iopub.execute_input":"2024-02-06T10:00:19.371408Z","iopub.status.idle":"2024-02-06T10:00:19.398053Z","shell.execute_reply.started":"2024-02-06T10:00:19.371375Z","shell.execute_reply":"2024-02-06T10:00:19.396447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### To be Continue....\n\n![](https://media.giphy.com/media/v1.Y2lkPTc5MGI3NjExdnd4cGJ1aWk2dGFpaW1xYTg4MDRyaHg3dm5rNWZ2OTBtamVkNHFndSZlcD12MV9pbnRlcm5hbF9naWZfYnlfaWQmY3Q9Zw/BsrrUcNzNstabBXwCJ/giphy.gif)","metadata":{}}]}