{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":50160,"databundleVersionId":7921029,"sourceType":"competition"}],"dockerImageVersionId":30698,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Project Introduction\n****","metadata":{}},{"cell_type":"markdown","source":"The **goal** of this competition is to determine how likely a customer is going to default on an issued loan.\n\nThe submission will be scored with a custom metric that will take into account how well the model performs in future. A decline in performance will be penalized. The goal is to create a model that is stable and performs well in the future.","metadata":{}},{"cell_type":"markdown","source":"* This is a **binary classification challenge** to predict home loan credit defaulters (target variable). GINI is the metric for the base classifier in this challenge\n\n* We also have to additionally assess the stability of GINI measure across time in the evaluation period. We need to score the classifier on a weekly basis and then assess the stability of the weekly GINI score using a regression model against time. Stability measure penalizes models that wane off in prediction capabilities across time.\n\nDepth values:\n* depth=0 - These are static features directly tied to a specific case_id.\n* depth=1 - Each case_id has an associated historical record, indexed by num_group1.\n* depth=2 - Each case_id has an associated historical record, indexed by both num_group1 and num_group2.","metadata":{}},{"cell_type":"markdown","source":"# Import Libraries\n****","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport datetime\nimport os\nimport gc\nimport polars as pl\nfrom sklearn.base import BaseEstimator, ClassifierMixin\nfrom pathlib import Path\nfrom glob import glob\nimport warnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)\nfrom sklearn.model_selection import StratifiedGroupKFold\nfrom sklearn.base import BaseEstimator, ClassifierMixin\nimport lightgbm as lgb","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:13.531754Z","iopub.execute_input":"2024-05-24T12:07:13.532386Z","iopub.status.idle":"2024-05-24T12:07:18.887782Z","shell.execute_reply.started":"2024-05-24T12:07:13.532346Z","shell.execute_reply":"2024-05-24T12:07:18.886945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We'll use Polars here - the goal of Polars is to provide a lightning fast DataFrame library that:\n\n* Utilizes all available cores on the machine.\n* Optimizes queries to reduce unneeded work/memory allocations.\n* Handles datasets much larger than available RAM.","metadata":{}},{"cell_type":"markdown","source":"# Import Data\n****","metadata":{}},{"cell_type":"code","source":"data_path = \"/kaggle/input/home-credit-credit-risk-model-stability/\"","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:18.889353Z","iopub.execute_input":"2024-05-24T12:07:18.889636Z","iopub.status.idle":"2024-05-24T12:07:18.893694Z","shell.execute_reply.started":"2024-05-24T12:07:18.889611Z","shell.execute_reply":"2024-05-24T12:07:18.892657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.read_csv(data_path + \"csv_files/train/train_base.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:18.894697Z","iopub.execute_input":"2024-05-24T12:07:18.894981Z","iopub.status.idle":"2024-05-24T12:07:20.112721Z","shell.execute_reply.started":"2024-05-24T12:07:18.894959Z","shell.execute_reply":"2024-05-24T12:07:20.111939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Initial Data Exploration\n****","metadata":{}},{"cell_type":"code","source":"train_df.head(2)","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:20.114634Z","iopub.execute_input":"2024-05-24T12:07:20.114950Z","iopub.status.idle":"2024-05-24T12:07:20.133334Z","shell.execute_reply.started":"2024-05-24T12:07:20.114917Z","shell.execute_reply":"2024-05-24T12:07:20.132455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.describe()","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:20.134607Z","iopub.execute_input":"2024-05-24T12:07:20.135152Z","iopub.status.idle":"2024-05-24T12:07:20.265735Z","shell.execute_reply.started":"2024-05-24T12:07:20.135120Z","shell.execute_reply":"2024-05-24T12:07:20.264804Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.isna().any()","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:20.267705Z","iopub.execute_input":"2024-05-24T12:07:20.268085Z","iopub.status.idle":"2024-05-24T12:07:20.410675Z","shell.execute_reply.started":"2024-05-24T12:07:20.268053Z","shell.execute_reply":"2024-05-24T12:07:20.409784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* There aro no null values, good for us!","metadata":{}},{"cell_type":"code","source":"train_df.dtypes","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:20.411946Z","iopub.execute_input":"2024-05-24T12:07:20.412338Z","iopub.status.idle":"2024-05-24T12:07:20.420124Z","shell.execute_reply.started":"2024-05-24T12:07:20.412306Z","shell.execute_reply":"2024-05-24T12:07:20.419373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* From the first glance, we can see that date_decision column is of the object type, which we will need to convert to a date.","metadata":{}},{"cell_type":"code","source":"train_df[\"date_decision\"] = pd.to_datetime(train_df[\"date_decision\"])","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:20.421186Z","iopub.execute_input":"2024-05-24T12:07:20.421470Z","iopub.status.idle":"2024-05-24T12:07:20.712946Z","shell.execute_reply.started":"2024-05-24T12:07:20.421449Z","shell.execute_reply":"2024-05-24T12:07:20.712187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Number of rows in the dataframe:","metadata":{}},{"cell_type":"code","source":"len(train_df.index)","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:20.713934Z","iopub.execute_input":"2024-05-24T12:07:20.714182Z","iopub.status.idle":"2024-05-24T12:07:20.720782Z","shell.execute_reply.started":"2024-05-24T12:07:20.714161Z","shell.execute_reply":"2024-05-24T12:07:20.719925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Number of distinct case ids:","metadata":{}},{"cell_type":"code","source":"train_df[\"case_id\"].nunique()","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:20.724252Z","iopub.execute_input":"2024-05-24T12:07:20.724549Z","iopub.status.idle":"2024-05-24T12:07:20.780202Z","shell.execute_reply.started":"2024-05-24T12:07:20.724526Z","shell.execute_reply":"2024-05-24T12:07:20.779353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Explorative Data Analysis\n****","metadata":{}},{"cell_type":"markdown","source":"To better understand the data, the nuances, the rules in it, and the connections between the variables, we'll convey univariate and bivariate data analysis.","metadata":{}},{"cell_type":"markdown","source":"## Univariate Data Analysis\n****","metadata":{"execution":{"iopub.status.busy":"2024-04-25T17:16:40.077758Z","iopub.execute_input":"2024-04-25T17:16:40.078162Z","iopub.status.idle":"2024-04-25T17:16:40.084623Z","shell.execute_reply.started":"2024-04-25T17:16:40.078133Z","shell.execute_reply":"2024-04-25T17:16:40.083246Z"}}},{"cell_type":"code","source":"for column in train_df.columns:\n    print(column)","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:20.781581Z","iopub.execute_input":"2024-05-24T12:07:20.781957Z","iopub.status.idle":"2024-05-24T12:07:20.787143Z","shell.execute_reply.started":"2024-05-24T12:07:20.781925Z","shell.execute_reply":"2024-05-24T12:07:20.786286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Month\n****","metadata":{}},{"cell_type":"code","source":"def countplot_variable(var):\n    sns.countplot(train_df, x = train_df[var])\n    plt.xticks(rotation = 90)\n    plt.xlabel(var)\n    plt.ylabel(\"Case Count\")\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:20.788260Z","iopub.execute_input":"2024-05-24T12:07:20.788565Z","iopub.status.idle":"2024-05-24T12:07:20.797391Z","shell.execute_reply.started":"2024-05-24T12:07:20.788537Z","shell.execute_reply":"2024-05-24T12:07:20.796581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"countplot_variable('MONTH')","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:20.798513Z","iopub.execute_input":"2024-05-24T12:07:20.799203Z","iopub.status.idle":"2024-05-24T12:07:22.122280Z","shell.execute_reply.started":"2024-05-24T12:07:20.799173Z","shell.execute_reply":"2024-05-24T12:07:22.121375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* It's clearly visible that there was a big drop in the number of cases since the start of 2020. This could be due to **Covid19**. Since it was a pandemic, and (we hope) once in a lifetime situation, we should bring that variable later as the flag. The numbers after that were definitely affected, and it probably took longet to recover. (*check literature)\n\n* What we find in the literature: *\"The observed drop in inquiries could be due to a drop in underlying credit demand, a drop credit supply affecting inquiries either directly or indirectly by heightening consumers’ in expectation of being turned down for credit, or a lack of opportunity for car and home sales to take place due to physical restrictions on movement and economic activity. **Regardless of the reason, the data indicate all types of inquiries dropped significantly and that consumers with higher credit scores have more flexibility in substituting away from applying for credit than consumers with lower credit scores.**\"\n\n\n* Probably, the number of cases maybe grew back to normal, but that the number of cases approved decreased. Let's find some data/articles on this topic, and think about how to put this in the data.","metadata":{}},{"cell_type":"code","source":"def countplot_with_target_hue(df, var):\n    sns.countplot(train_df, x = df[var], hue = df['target'])\n    plt.xticks(rotation = 90)\n    plt.xlabel(\"Decision Month\")\n    plt.ylabel(\"Case Count\")\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:22.123414Z","iopub.execute_input":"2024-05-24T12:07:22.123700Z","iopub.status.idle":"2024-05-24T12:07:22.129291Z","shell.execute_reply.started":"2024-05-24T12:07:22.123676Z","shell.execute_reply":"2024-05-24T12:07:22.128288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"countplot_with_target_hue(train_df, 'MONTH')","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:22.130394Z","iopub.execute_input":"2024-05-24T12:07:22.130650Z","iopub.status.idle":"2024-05-24T12:07:23.890667Z","shell.execute_reply.started":"2024-05-24T12:07:22.130627Z","shell.execute_reply":"2024-05-24T12:07:23.889740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* As expected, most of the cases get rejected. Let's explore what's the ratio of rejected vs accepted loans:","metadata":{}},{"cell_type":"code","source":"def ratio_per_period(df, var):\n    count_per_target = df[[var,'case_id','target']].groupby([var,'target']).count()\n    count_per_period = df[[var,'case_id']].groupby([var]).count()\n    ratio = round(count_per_target/count_per_period, 4)*100\n    ratio = ratio.reset_index()\n    return ratio","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:23.891739Z","iopub.execute_input":"2024-05-24T12:07:23.892027Z","iopub.status.idle":"2024-05-24T12:07:23.897538Z","shell.execute_reply.started":"2024-05-24T12:07:23.892003Z","shell.execute_reply":"2024-05-24T12:07:23.896640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ratio_monthly = ratio_per_period(train_df, 'MONTH')\nratio_monthly[ratio_monthly['target'] == 1].head(2)","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:23.898533Z","iopub.execute_input":"2024-05-24T12:07:23.898784Z","iopub.status.idle":"2024-05-24T12:07:24.037142Z","shell.execute_reply.started":"2024-05-24T12:07:23.898762Z","shell.execute_reply":"2024-05-24T12:07:24.036013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ratio_monthly[ratio_monthly['target'] == 1].groupby(['target']).mean()","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:24.038442Z","iopub.execute_input":"2024-05-24T12:07:24.038731Z","iopub.status.idle":"2024-05-24T12:07:24.050862Z","shell.execute_reply.started":"2024-05-24T12:07:24.038708Z","shell.execute_reply":"2024-05-24T12:07:24.049947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The different way of getting the same information (just for fun):\n\n*please don't comment on this definition of fun","metadata":{}},{"cell_type":"code","source":"ratio_monthly[ratio_monthly['target'] == 1].mean()","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:24.052032Z","iopub.execute_input":"2024-05-24T12:07:24.052309Z","iopub.status.idle":"2024-05-24T12:07:24.062875Z","shell.execute_reply.started":"2024-05-24T12:07:24.052278Z","shell.execute_reply":"2024-05-24T12:07:24.062010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that the average across all months for the approved credits was 2.967%.\nLet's dive into the average before Covid19 and during (and after?) Covid19. To do that, we need to define Covid19 exact dates/months.","metadata":{}},{"cell_type":"markdown","source":"Covid19 began mostly in March 2020, which is after when we see a bigger drop (outside of the seasonality). Credit applications mostly recovered to pre-pandemic levels for all credit types considered by May 2021 (citing literature - available on the link below).\n\nSince that data is not available for training, we might need to extrapolate it.\n\nFirst, let's do the average before March 2020 and after.","metadata":{}},{"cell_type":"code","source":"ratio_monthly[(ratio_monthly['target'] == 1) & (ratio_monthly['MONTH'] < 202003)].mean()","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:24.064020Z","iopub.execute_input":"2024-05-24T12:07:24.064282Z","iopub.status.idle":"2024-05-24T12:07:24.077491Z","shell.execute_reply.started":"2024-05-24T12:07:24.064261Z","shell.execute_reply":"2024-05-24T12:07:24.076665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ratio_monthly[(ratio_monthly['target'] == 1) & (ratio_monthly['MONTH'] >= 202003)].mean()","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:24.078583Z","iopub.execute_input":"2024-05-24T12:07:24.078932Z","iopub.status.idle":"2024-05-24T12:07:24.090349Z","shell.execute_reply.started":"2024-05-24T12:07:24.078884Z","shell.execute_reply":"2024-05-24T12:07:24.089500Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"difference_between_credit_approved_percentage = ratio_monthly[(ratio_monthly['target'] == 1) & (ratio_monthly['MONTH'] < 202003)].mean()[2] - ratio_monthly[(ratio_monthly['target'] == 1) & (ratio_monthly['MONTH'] >= 202003)].mean()[2]\ndifference_between_credit_approved_percentage","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:24.091253Z","iopub.execute_input":"2024-05-24T12:07:24.091503Z","iopub.status.idle":"2024-05-24T12:07:24.104061Z","shell.execute_reply.started":"2024-05-24T12:07:24.091481Z","shell.execute_reply":"2024-05-24T12:07:24.103219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The difference of almost 1% (0.74%) is clear, for the average rate of approved credits before and after covid. \nThe effect of Covid19 was:\n* there was a smaller number of applications\n* the average rate of approval decreased","metadata":{}},{"cell_type":"markdown","source":"Since the data during pandemic was an exception (again, hopefully), we'll look into the further analysis on the data before and after March 2020.\n\nSo, let's peek into the seasonality of the months in terms of credit approval and credit case count, for period before and after.\n\nWe'll feature engineer the name of the month:","metadata":{}},{"cell_type":"code","source":"train_df['Month_Name'] = train_df['date_decision'].dt.strftime(\"%B\")","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:24.105166Z","iopub.execute_input":"2024-05-24T12:07:24.105528Z","iopub.status.idle":"2024-05-24T12:07:34.514437Z","shell.execute_reply.started":"2024-05-24T12:07:24.105499Z","shell.execute_reply":"2024-05-24T12:07:34.513466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since I don't believe most of the things tech-related until I see them with my own eyes, let's double-check month names:","metadata":{}},{"cell_type":"code","source":"train_df.sample(10)","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:34.515495Z","iopub.execute_input":"2024-05-24T12:07:34.515819Z","iopub.status.idle":"2024-05-24T12:07:34.573079Z","shell.execute_reply.started":"2024-05-24T12:07:34.515794Z","shell.execute_reply":"2024-05-24T12:07:34.572003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Good to go! Now let's see the trends per month (case count, ratio of approved cases)","metadata":{}},{"cell_type":"code","source":"countplot_with_target_hue(train_df, 'Month_Name')","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:34.574357Z","iopub.execute_input":"2024-05-24T12:07:34.574697Z","iopub.status.idle":"2024-05-24T12:07:37.777058Z","shell.execute_reply.started":"2024-05-24T12:07:34.574672Z","shell.execute_reply":"2024-05-24T12:07:37.776186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"When looking at all the data mixed up, it's visible that January and September have most credit application cases. We don't want this data poluted with Covid data, so let's split before Covid and during Covid:","metadata":{}},{"cell_type":"code","source":"countplot_with_target_hue(train_df[train_df['MONTH'] < 202003], 'Month_Name')","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:37.778144Z","iopub.execute_input":"2024-05-24T12:07:37.778429Z","iopub.status.idle":"2024-05-24T12:07:40.448068Z","shell.execute_reply.started":"2024-05-24T12:07:37.778399Z","shell.execute_reply":"2024-05-24T12:07:40.447213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In pre-pandemic data, most of the cases occured in January, February, and December. It naturally dropped in warmer months (from March to October).","metadata":{}},{"cell_type":"code","source":"countplot_with_target_hue(train_df[train_df['MONTH'] >= 202003], 'Month_Name')","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:40.449096Z","iopub.execute_input":"2024-05-24T12:07:40.449352Z","iopub.status.idle":"2024-05-24T12:07:41.261699Z","shell.execute_reply.started":"2024-05-24T12:07:40.449329Z","shell.execute_reply":"2024-05-24T12:07:41.260741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In post-pandemic data, most of the cases occured in September and March. There are some drops different from the ones in the pre-pandemic data (i.e. July and October).","metadata":{}},{"cell_type":"markdown","source":"Now let's see the ratio of approved credit before and after pandemic, by month name.","metadata":{}},{"cell_type":"code","source":"ratio_per_month_name = ratio_per_period(train_df, 'Month_Name')","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:41.268877Z","iopub.execute_input":"2024-05-24T12:07:41.269179Z","iopub.status.idle":"2024-05-24T12:07:41.693874Z","shell.execute_reply.started":"2024-05-24T12:07:41.269156Z","shell.execute_reply":"2024-05-24T12:07:41.693046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ratio_per_month_name[ratio_per_month_name['target'] == 1].sort_values(by = 'case_id', ascending = False)","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:41.694964Z","iopub.execute_input":"2024-05-24T12:07:41.695244Z","iopub.status.idle":"2024-05-24T12:07:41.709146Z","shell.execute_reply.started":"2024-05-24T12:07:41.695220Z","shell.execute_reply":"2024-05-24T12:07:41.708200Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It's visible that in February, October, and November, there is the highest ratio of credit being approved.\nThis is different than what we saw, January and September having the highest number of credit applications (for mixed data as well). It's safe to say that here the higher number of applications doesn't affect the increase/decrease in the ratio of approved credit.\n\nNow let's split the data into pre-pandemic and post-pandemic:","metadata":{}},{"cell_type":"code","source":"ratio_per_month_name_pre_pandemic = ratio_per_period(train_df[train_df['MONTH'] < 202003], 'Month_Name')","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:41.710288Z","iopub.execute_input":"2024-05-24T12:07:41.710557Z","iopub.status.idle":"2024-05-24T12:07:42.118160Z","shell.execute_reply.started":"2024-05-24T12:07:41.710535Z","shell.execute_reply":"2024-05-24T12:07:42.117154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ratio_per_month_name_pre_pandemic[ratio_per_month_name_pre_pandemic['target'] == 1].sort_values(by = 'case_id', ascending = False)","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:42.119347Z","iopub.execute_input":"2024-05-24T12:07:42.119650Z","iopub.status.idle":"2024-05-24T12:07:42.131806Z","shell.execute_reply.started":"2024-05-24T12:07:42.119627Z","shell.execute_reply":"2024-05-24T12:07:42.130938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Pre-pandemic, October, February, and November had the highest ratio of credit applications approved. Except for February which is over-indexed, the rest of the months aren't the months with the highest number of credit application. Again, the number of applications doesn't impact the number of approvals.","metadata":{}},{"cell_type":"code","source":"ratio_per_month_name_post_pandemic = ratio_per_period(train_df[train_df['MONTH'] >= 202003], 'Month_Name')\nratio_per_month_name_post_pandemic[ratio_per_month_name_post_pandemic['target'] == 1].sort_values(by = 'case_id', ascending = False)","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:42.132946Z","iopub.execute_input":"2024-05-24T12:07:42.133224Z","iopub.status.idle":"2024-05-24T12:07:42.245189Z","shell.execute_reply.started":"2024-05-24T12:07:42.133201Z","shell.execute_reply":"2024-05-24T12:07:42.244108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In Covid, March had most credit applications, and it had the highest ratio of approvals. It probably took some time for the system to change and to take on the \"new normal\". After that, the drop is visible, and it doesn't go even close to the March-levels of approval. Other months aren't that affected by the number of applications.","metadata":{}},{"cell_type":"markdown","source":"To conclude on this part, it's clear that both pre-pandemic and post-pandemic, the number of applications didn't impact the number of approvals. However, the ratio of credit approved dropped significantly (0.73% on average) during pandemic.","metadata":{}},{"cell_type":"markdown","source":"## Week Num\n****","metadata":{}},{"cell_type":"markdown","source":"Let's add another feature; day of the week.","metadata":{}},{"cell_type":"code","source":"train_df['Day_of_the_Week'] = train_df['date_decision'].dt.day_name()","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:42.246384Z","iopub.execute_input":"2024-05-24T12:07:42.246697Z","iopub.status.idle":"2024-05-24T12:07:42.817327Z","shell.execute_reply.started":"2024-05-24T12:07:42.246671Z","shell.execute_reply":"2024-05-24T12:07:42.816299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.sample(10)","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:42.818748Z","iopub.execute_input":"2024-05-24T12:07:42.819118Z","iopub.status.idle":"2024-05-24T12:07:42.873983Z","shell.execute_reply.started":"2024-05-24T12:07:42.819086Z","shell.execute_reply":"2024-05-24T12:07:42.873025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"countplot_with_target_hue(train_df[train_df['MONTH'] < 202003], 'Day_of_the_Week')","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:42.875272Z","iopub.execute_input":"2024-05-24T12:07:42.875633Z","iopub.status.idle":"2024-05-24T12:07:45.196182Z","shell.execute_reply.started":"2024-05-24T12:07:42.875601Z","shell.execute_reply":"2024-05-24T12:07:45.195243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Pre-pandemic, Saturdays and Fridays were when there were most decisions made for applications.","metadata":{}},{"cell_type":"markdown","source":"There's a lot more datasets to read and join onto this base dataset, so let's continue!","metadata":{}},{"cell_type":"code","source":"ROOT            = Path(\"/kaggle/input/home-credit-credit-risk-model-stability\")\nTRAIN_DIR       = ROOT / \"parquet_files\" / \"train\"\nTEST_DIR        = ROOT / \"parquet_files\" / \"test\"","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:45.197353Z","iopub.execute_input":"2024-05-24T12:07:45.197633Z","iopub.status.idle":"2024-05-24T12:07:45.202450Z","shell.execute_reply.started":"2024-05-24T12:07:45.197610Z","shell.execute_reply":"2024-05-24T12:07:45.201532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's read all the files, and make the necessary changes on each:","metadata":{}},{"cell_type":"code","source":"class Pipeline:\n    \n    @staticmethod\n    def set_table_dtypes(df):\n        for col in df.columns:\n            if col in [\"case_id\", \"WEEK_NUM\", \"num_group1\", \"num_group2\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Int32))\n            elif col in [\"date_decision\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Date))\n            elif col[-1] in (\"P\", \"A\"):\n                df = df.with_columns(pl.col(col).cast(pl.Float64))\n            elif col[-1] in (\"M\"):\n                df = df.with_columns(pl.col(col).cast(pl.String))\n            elif col[-1] in (\"D\"):\n                df = df.with_columns(pl.col(col).cast(pl.Date))\n                \n        return df\n    \n    @staticmethod\n    def handle_dates(df):\n        for col in df.columns:\n            if col[-1] in (\"D\",):\n                df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))\n                df = df.with_columns(pl.col(col).dt.total_days())\n                df = df.with_columns(pl.col(col).cast(pl.Float32))\n                \n        df = df.drop(\"date_decision\", \"MONTH\")\n        \n        return df\n    \n    @staticmethod\n    def filter_cols(df):\n        for col in df.columns:\n            if col not in [\"target\", \"case_id\", \"WEEK_NUM\"]:\n                isnull = df[col].is_null().mean()\n\n                if isnull > 0.95:\n                    df = df.drop(col)\n                    \n        for col in df.columns:\n            if (col not in [\"target\", \"case_id\", \"WEEK_NUM\"]) & (df[col].dtype == pl.String):\n                freq = df[col].n_unique()\n\n                if (freq == 1) | (freq > 200):\n                    df = df.drop(col)\n\n        return df\n    ","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:45.203495Z","iopub.execute_input":"2024-05-24T12:07:45.204064Z","iopub.status.idle":"2024-05-24T12:07:45.217602Z","shell.execute_reply.started":"2024-05-24T12:07:45.204039Z","shell.execute_reply":"2024-05-24T12:07:45.216701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we'll create an Aggregator, to aggregate level 1 and level 2 data:","metadata":{}},{"cell_type":"code","source":"class Aggregator:\n    \n    @staticmethod\n    def num_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"P\", \"A\")]\n        \n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        \n        return expr_max\n    \n    @staticmethod\n    def date_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"D\")]\n        \n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        \n        return expr_max\n    \n    @staticmethod\n    def string_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"M\")]\n        \n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        \n        return expr_max\n    \n    @staticmethod\n    def other_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"T\", \"L\")]\n        \n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        return expr_max\n    \n    @staticmethod\n    def count_expr(df):\n        cols = [col for col in df.columns if \"num_group\" in col]\n        \n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        return expr_max\n    \n    @staticmethod\n    def get_exprs(df):\n        exprs = Aggregator.num_expr(df) + \\\n                Aggregator.num_expr(df) + \\\n                Aggregator.date_expr(df) + \\\n                Aggregator.string_expr(df) + \\\n                Aggregator.other_expr(df) + \\\n                Aggregator.count_expr(df)\n        \n        return exprs\n    ","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:45.218808Z","iopub.execute_input":"2024-05-24T12:07:45.219177Z","iopub.status.idle":"2024-05-24T12:07:45.232543Z","shell.execute_reply.started":"2024-05-24T12:07:45.219147Z","shell.execute_reply":"2024-05-24T12:07:45.231812Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's create the same thing, but shorter:","metadata":{}},{"cell_type":"code","source":"class Aggregator_two:\n    \n    @staticmethod\n    def level_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"P\", \"A\", \"D\", \"M\", \"T\", \"L\")]\n        \n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        \n        return expr_max\n    \n    @staticmethod\n    def count_expr(df):\n        cols = [col for col in df.columns if \"num_group\" in col]\n        \n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        return expr_max\n    \n    @staticmethod\n    def get_exprs(df):\n        exprs = Aggregator_two.level_expr(df) + \\\n                Aggregator.count_expr(df)\n        \n        return exprs","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:45.233715Z","iopub.execute_input":"2024-05-24T12:07:45.234262Z","iopub.status.idle":"2024-05-24T12:07:45.247175Z","shell.execute_reply.started":"2024-05-24T12:07:45.234237Z","shell.execute_reply":"2024-05-24T12:07:45.246404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Reading the files in a path, handling levels 1 and 2, and appending datasets","metadata":{}},{"cell_type":"code","source":"def read_file(path, depth=None):\n    df = pl.read_parquet(path)\n    df = df.pipe(Pipeline.set_table_dtypes)\n    \n    if depth in [1, 2]:\n        df = df.group_by(\"case_id\").agg(Aggregator_two.get_exprs(df))\n    \n    return df\n\ndef read_files(regex_path, depth = None):\n    chunks = []\n    for path in glob(str(regex_path)):\n        df = pl.read_parquet(path)\n        df = df.pipe(Pipeline.set_table_dtypes)\n        \n        if depth in [1, 2]:\n            df = df.group_by(\"case_id\").agg(Aggregator_two.get_exprs(df))\n        \n        chunks.append(df)\n        \n    df = pl.concat(chunks, how=\"vertical_relaxed\")\n    df = df.unique(subset=[\"case_id\"])\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:45.248198Z","iopub.execute_input":"2024-05-24T12:07:45.248449Z","iopub.status.idle":"2024-05-24T12:07:45.260580Z","shell.execute_reply.started":"2024-05-24T12:07:45.248427Z","shell.execute_reply":"2024-05-24T12:07:45.259843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Feature engineering, similar to what we did in the base file, to get week days and months","metadata":{}},{"cell_type":"code","source":"def feature_eng(df_base, depth_0, depth_1, depth_2):\n    df_base = (df_base.with_columns(\n                                weekday_decision = pl.col(\"date_decision\").dt.weekday(),\n                                month_decision = pl.col(\"date_decision\").dt.month(),\n                                )\n            )\n    \n    for i, df in enumerate(depth_0 + depth_1 + depth_2):\n        df_base = df_base.join(df, how = \"left\", on = \"case_id\", suffix = f\"_{i}\")\n        \n    df_base = df_base.pipe(Pipeline.handle_dates)\n        \n    return df_base","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:45.261609Z","iopub.execute_input":"2024-05-24T12:07:45.261881Z","iopub.status.idle":"2024-05-24T12:07:45.271389Z","shell.execute_reply.started":"2024-05-24T12:07:45.261860Z","shell.execute_reply":"2024-05-24T12:07:45.270542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def to_pandas(df, cat_cols = None):\n    df = df.to_pandas()\n    \n    if cat_cols is None:\n        cat_cols = list(df.select_dtypes(\"object\").columns)\n    \n    df[cat_cols] = df[cat_cols].astype(\"category\")\n        \n    return df, cat_cols","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:45.272433Z","iopub.execute_input":"2024-05-24T12:07:45.273123Z","iopub.status.idle":"2024-05-24T12:07:45.281944Z","shell.execute_reply.started":"2024-05-24T12:07:45.273099Z","shell.execute_reply":"2024-05-24T12:07:45.281145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Reading train files:","metadata":{}},{"cell_type":"code","source":"data_store_train = {\n    \"df_base\": read_file(TRAIN_DIR / \"train_base.parquet\"),\n    \"depth_0\": [\n        read_file(TRAIN_DIR / \"train_static_cb_0.parquet\"),\n        read_files(TRAIN_DIR / \"train_static_0_*.parquet\"),\n    ],\n    \"depth_1\": [\n        read_files(TRAIN_DIR / \"train_applprev_1_*.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_a_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_b_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_c_1.parquet\", 1),\n        read_files(TRAIN_DIR / \"train_credit_bureau_a_1_*.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_credit_bureau_b_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_other_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_person_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_deposit_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_debitcard_1.parquet\", 1),\n    ],\n    \"depth_2\": [\n        read_file(TRAIN_DIR / \"train_credit_bureau_b_2.parquet\", 2),\n        read_files(TRAIN_DIR / \"train_credit_bureau_a_2_*.parquet\", 2),\n    ]\n}","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:07:45.282883Z","iopub.execute_input":"2024-05-24T12:07:45.283163Z","iopub.status.idle":"2024-05-24T12:09:49.987987Z","shell.execute_reply.started":"2024-05-24T12:07:45.283142Z","shell.execute_reply":"2024-05-24T12:09:49.986956Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Feature engineering, removing columns","metadata":{}},{"cell_type":"code","source":"df_train = feature_eng(**data_store_train)\nprint(\"train data shape:\\t\", df_train.shape)","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:09:49.989336Z","iopub.execute_input":"2024-05-24T12:09:49.989597Z","iopub.status.idle":"2024-05-24T12:09:59.325601Z","shell.execute_reply.started":"2024-05-24T12:09:49.989576Z","shell.execute_reply":"2024-05-24T12:09:59.324709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Reading test files:","metadata":{}},{"cell_type":"code","source":"data_store_test = {\n    \"df_base\": read_file(TEST_DIR / \"test_base.parquet\"),\n    \"depth_0\": [\n        read_file(TEST_DIR / \"test_static_cb_0.parquet\"),\n        read_files(TEST_DIR / \"test_static_0_*.parquet\"),\n    ],\n    \"depth_1\": [\n        read_files(TEST_DIR / \"test_applprev_1_*.parquet\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_a_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_b_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_c_1.parquet\", 1),\n        read_files(TEST_DIR / \"test_credit_bureau_a_1_*.parquet\", 1),\n        read_file(TEST_DIR / \"test_credit_bureau_b_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_other_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_person_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_deposit_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_debitcard_1.parquet\", 1),\n    ],\n    \"depth_2\": [\n        read_file(TEST_DIR / \"test_credit_bureau_b_2.parquet\", 2),\n        read_files(TEST_DIR / \"test_credit_bureau_a_2_*.parquet\", 2),\n    ]\n}","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:09:59.326685Z","iopub.execute_input":"2024-05-24T12:09:59.326953Z","iopub.status.idle":"2024-05-24T12:09:59.636181Z","shell.execute_reply.started":"2024-05-24T12:09:59.326931Z","shell.execute_reply":"2024-05-24T12:09:59.635420Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test = feature_eng(**data_store_test)\nprint(\"train data shape:\\t\", df_test.shape)","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:09:59.637264Z","iopub.execute_input":"2024-05-24T12:09:59.637577Z","iopub.status.idle":"2024-05-24T12:09:59.679137Z","shell.execute_reply.started":"2024-05-24T12:09:59.637554Z","shell.execute_reply":"2024-05-24T12:09:59.678242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = df_train.pipe(Pipeline.filter_cols)\nprint(\"train data shape:\\t\", df_train.shape)","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:09:59.680361Z","iopub.execute_input":"2024-05-24T12:09:59.681054Z","iopub.status.idle":"2024-05-24T12:10:02.089694Z","shell.execute_reply.started":"2024-05-24T12:09:59.681021Z","shell.execute_reply":"2024-05-24T12:10:02.088697Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_test = df_test.pipe(Pipeline.filter_cols)\ndf_test = df_test.select([col for col in df_train.columns if col != \"target\"])\nprint(\"test data shape:\\t\", df_test.shape)","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:10:02.090962Z","iopub.execute_input":"2024-05-24T12:10:02.091260Z","iopub.status.idle":"2024-05-24T12:10:02.099291Z","shell.execute_reply.started":"2024-05-24T12:10:02.091235Z","shell.execute_reply":"2024-05-24T12:10:02.098237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train, cat_cols = to_pandas(df_train)\ndf_test, cat_cols = to_pandas(df_test, cat_cols)","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:10:02.100421Z","iopub.execute_input":"2024-05-24T12:10:02.100659Z","iopub.status.idle":"2024-05-24T12:10:22.727243Z","shell.execute_reply.started":"2024-05-24T12:10:02.100640Z","shell.execute_reply":"2024-05-24T12:10:22.726469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del data_store_train, data_store_test\n\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:10:22.728359Z","iopub.execute_input":"2024-05-24T12:10:22.728673Z","iopub.status.idle":"2024-05-24T12:10:23.186278Z","shell.execute_reply.started":"2024-05-24T12:10:22.728649Z","shell.execute_reply":"2024-05-24T12:10:23.185324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Explorative Data Analysis\n****","metadata":{}},{"cell_type":"markdown","source":"Doing EDA on all train data, not only base data. Let's check out the trends in overall data!","metadata":{}},{"cell_type":"code","source":"# placeholder","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:10:23.187465Z","iopub.execute_input":"2024-05-24T12:10:23.187773Z","iopub.status.idle":"2024-05-24T12:10:23.195368Z","shell.execute_reply.started":"2024-05-24T12:10:23.187748Z","shell.execute_reply":"2024-05-24T12:10:23.194544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Checking correlation between the variables to exclude the high-correlation ones:","metadata":{}},{"cell_type":"code","source":"df_train.select_dtypes(['number'])","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:10:23.196320Z","iopub.execute_input":"2024-05-24T12:10:23.196594Z","iopub.status.idle":"2024-05-24T12:10:24.604254Z","shell.execute_reply.started":"2024-05-24T12:10:23.196573Z","shell.execute_reply":"2024-05-24T12:10:24.603366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''\n# Create correlation matrix\ncorr_matrix = df_train.select_dtypes(['number']).corr().abs()\n\n# Select upper triangle of correlation matrix\nupper = corr_matrix.where(np.triu(np.ones(corr_matrix.shape), k = 1).astype(bool))\n\n# Find features with correlation greater than 0.95\nto_drop = [column for column in upper.columns if any(upper[column] > 0.95)]\n\n# Drop features \ndf_train.drop(to_drop, axis=1, inplace = True)\n'''","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:10:24.605521Z","iopub.execute_input":"2024-05-24T12:10:24.605815Z","iopub.status.idle":"2024-05-24T12:10:24.612244Z","shell.execute_reply.started":"2024-05-24T12:10:24.605790Z","shell.execute_reply":"2024-05-24T12:10:24.611276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#print(\"Number of numerical columns dropped: \", len(to_drop))","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:10:24.613415Z","iopub.execute_input":"2024-05-24T12:10:24.613666Z","iopub.status.idle":"2024-05-24T12:10:24.621661Z","shell.execute_reply.started":"2024-05-24T12:10:24.613644Z","shell.execute_reply":"2024-05-24T12:10:24.620743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training the Model\n****","metadata":{}},{"cell_type":"markdown","source":"Long awaited step, let's train the model! This is just the first version from which we'll hopefully progress. Let's use something simple for the first run.","metadata":{}},{"cell_type":"code","source":"class VotingModel(BaseEstimator, ClassifierMixin):\n    def __init__(self, estimators):\n        super().__init__()\n        self.estimators = estimators\n        \n    def fit(self, X, y = None):\n        return self\n    \n    def predict(self, X):\n        y_preds = [estimator.predict(X) for estimator in self.estimators]\n        return np.mean(y_preds, axis = 0)\n    \n    def predict_proba(self, X):\n        y_preds = [estimator.predict_proba(X) for estimator in self.estimators]\n        return np.mean(y_preds, axis = 0)","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:10:24.622801Z","iopub.execute_input":"2024-05-24T12:10:24.623162Z","iopub.status.idle":"2024-05-24T12:10:24.635631Z","shell.execute_reply.started":"2024-05-24T12:10:24.623133Z","shell.execute_reply":"2024-05-24T12:10:24.634769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = df_train.drop(columns = [\"target\", \"case_id\", \"WEEK_NUM\"])\ny = df_train[\"target\"]\nweeks = df_train[\"WEEK_NUM\"]\n\ncv = StratifiedGroupKFold(n_splits = 5, shuffle = False)\n\nparams = {\n    \"boosting_type\": \"gbdt\",\n    \"objective\": \"binary\",\n    \"metric\": \"auc\",\n    \"max_depth\": 8,\n    \"learning_rate\": 0.05,\n    \"n_estimators\": 1000,\n    \"colsample_bytree\": 0.8, \n    \"colsample_bynode\": 0.8,\n    \"verbose\": -1,\n    \"random_state\": 42,\n    \"device\": \"gpu\",\n}\n\nfitted_models = []\n\nfor idx_train, idx_valid in cv.split(X, y, groups = weeks):\n    X_train, y_train = X.iloc[idx_train], y.iloc[idx_train]\n    X_valid, y_valid = X.iloc[idx_valid], y.iloc[idx_valid]\n\n    model = lgb.LGBMClassifier(**params)\n    model.fit(\n        X_train, y_train,\n        eval_set = [(X_valid, y_valid)],\n        callbacks = [lgb.log_evaluation(100), lgb.early_stopping(100)]\n    )\n\n    fitted_models.append(model)\n\nmodel = VotingModel(fitted_models)","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:10:24.636606Z","iopub.execute_input":"2024-05-24T12:10:24.636867Z","iopub.status.idle":"2024-05-24T12:27:27.183567Z","shell.execute_reply.started":"2024-05-24T12:10:24.636838Z","shell.execute_reply":"2024-05-24T12:27:27.182588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prediction\n****","metadata":{}},{"cell_type":"code","source":"#df_test.drop(to_drop, axis=1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:27:27.185064Z","iopub.execute_input":"2024-05-24T12:27:27.185424Z","iopub.status.idle":"2024-05-24T12:27:27.189373Z","shell.execute_reply.started":"2024-05-24T12:27:27.185392Z","shell.execute_reply":"2024-05-24T12:27:27.188523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test = df_test.drop(columns=[\"WEEK_NUM\"])\nX_test = X_test.set_index(\"case_id\")","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:27:27.190481Z","iopub.execute_input":"2024-05-24T12:27:27.190722Z","iopub.status.idle":"2024-05-24T12:27:27.208348Z","shell.execute_reply.started":"2024-05-24T12:27:27.190701Z","shell.execute_reply":"2024-05-24T12:27:27.207365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred = pd.Series(model.predict_proba(X_test)[:, 1], index = X_test.index)","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:27:27.209548Z","iopub.execute_input":"2024-05-24T12:27:27.210117Z","iopub.status.idle":"2024-05-24T12:27:27.488561Z","shell.execute_reply.started":"2024-05-24T12:27:27.210082Z","shell.execute_reply":"2024-05-24T12:27:27.487528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission\n****","metadata":{}},{"cell_type":"code","source":"df_subm = pd.read_csv(ROOT / \"sample_submission.csv\")\ndf_subm = df_subm.set_index(\"case_id\")\n\ndf_subm[\"score\"] = y_pred\ndf_subm.to_csv(\"submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:27:27.489866Z","iopub.execute_input":"2024-05-24T12:27:27.490225Z","iopub.status.idle":"2024-05-24T12:27:27.515570Z","shell.execute_reply.started":"2024-05-24T12:27:27.490197Z","shell.execute_reply":"2024-05-24T12:27:27.514852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Feature importance\n****","metadata":{}},{"cell_type":"code","source":"#feat_imp = pd.Series(model.feature_importances_, index=X.columns)\n#feat_imp.nlargest(30).plot(kind='barh', figsize=(8,10))","metadata":{"execution":{"iopub.status.busy":"2024-05-24T12:27:27.516801Z","iopub.execute_input":"2024-05-24T12:27:27.517160Z","iopub.status.idle":"2024-05-24T12:27:27.521525Z","shell.execute_reply.started":"2024-05-24T12:27:27.517129Z","shell.execute_reply":"2024-05-24T12:27:27.520504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Evaluation\n****","metadata":{}},{"cell_type":"markdown","source":"Submissions are evaluated using a gini stability metric. A gini score is calculated for predictions corresponding to each WEEK_NUM.\n\n> **gini = 2∗AUC−1**\n\nA linear regression, a⋅x+b\n, is fit through the weekly gini scores, and a falling_rate is calculated as min(0,a)\n. This is used to penalize models that drop off in predictive ability.\n\nFinally, the variability of the predictions are calculated by taking the standard deviation of the residuals from the above linear regression, applying a penalty to model variablity.\n\nThe final metric is calculated as:\n\n> **stability metric = mean(gini) + 88.0⋅min(0,a) − 0.5⋅std(residuals)**","metadata":{}},{"cell_type":"markdown","source":"This means we are going for:\n* stability metric as high as possible\n* gini as high as possible\n* AUC as high as possible\n* falling rate as low as possible (linear factor a as low as possible)\n* resudials residuals as low as possible","metadata":{}},{"cell_type":"markdown","source":"# Literature\n****","metadata":{}},{"cell_type":"markdown","source":"* [How Covid19 Affected Credit Applications](https://files.consumerfinance.gov/f/documents/cfpb_issue-brief_early-effects-covid-19-credit-applications_2020-04.pdf)\n\n* [How Credit Applications Recovered to Pre-Pandemic Levels](https://files.consumerfinance.gov/f/documents/cfpb_recovery-of-credit-applications-pre-pandemic-levels_report_2021-07.pdf)","metadata":{}},{"cell_type":"markdown","source":"# Credits\n****","metadata":{}},{"cell_type":"markdown","source":"[* https://www.kaggle.com/code/greysky/home-credit-baseline/notebook](*https://www.kaggle.com/code/greysky/home-credit-baseline/notebook)","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}