{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":50160,"databundleVersionId":7921029,"sourceType":"competition"}],"dockerImageVersionId":30664,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## <div style= \"font-family: Cambria; font-weight:bold; letter-spacing: 0px; color:white; font-size:150%; text-align:left;padding:3.0px; background: #607BB0; border-bottom: 8px solid #53599A; border-radius: 5px;\">Differences between PIT and TTC credit scoring models<br><div>\n\nThe main idea of this notebook is:\n    \n- to present the theoretical background behind risk rating philosophy, i.e. point in time (PIT) and through the cycle (TTC) models;\n- provide practical examples how the model's stability could be evaluated. \n    \n## <div style= \"font-family: Cambria; font-weight:bold; letter-spacing: 0px; color:white; font-size:100%; text-align:left;padding:3.0px; background: #607BB0; border-bottom: 8px solid #53599A; border-radius: 5px;\">Table of contents<br><div>\n<a id=\"toc\"></a>\n- [1 Theory](#1)\n    - [1.1 TTC vs PIT models](#1.1)\n    - [1.2 Data limitations](#1.2)\n    - [1.3 Estimating model's time horizon](#1.3)\n- [2 Prepare project](#2)\n    - [2.1 Load data](#2.1)\n    - [2.2 Feature engineering](#2.2)\n- [3 Feature analysis](#3)\n    - [3.1 Historical default rates](#3.1)\n    - [3.2 Train/test/valid split](#3.2)\n        - [3.2.1 Out-of-time split](#3.2.1)\n        - [3.2.2 Random split](#3.2.2)\n        - [3.2.3 Compare splits](#3.2.3)\n- [4 Modeling](#4)\n    - [4.1 Ensembling](#4.1)\n- [5 Model validation](#5)\n- [6 Submit predictions](#6)\n- [7 Changelog](#7)\n- [Refencences](#8)    ","metadata":{}},{"cell_type":"code","source":"# data processing libraries\nimport polars as pl\nimport numpy as np\nimport pandas as pd\n\n# LightGBM modeling\nimport lightgbm as lgb\n\n# sklearn functions\nfrom sklearn.model_selection import train_test_split, StratifiedGroupKFold\n# for validating models\nfrom sklearn.metrics import roc_auc_score, roc_curve\n\n# for hypertuning\nfrom hyperopt import hp, fmin, tpe, Trials\n\n# for visualizing data\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n# define default colors for plots in notebook\nfrom matplotlib import cycler\nfrom matplotlib.colors import LinearSegmentedColormap\nCOLORS = [\"#068D9D\", \"#53599A\", \"#607BB0\", \"#6D9DC5\", \"#77BECF\", \"#80DED9\", \"#AEECEF\"]\nplt.rc('axes', facecolor='#E6E6E6', edgecolor='none', axisbelow=True, grid=True, prop_cycle=cycler('color', COLORS))\n\n# to save all trained models\nALL_MODELS = dict()\n\n# constants\nPATH = \"/kaggle/input/home-credit-credit-risk-model-stability/\"\nSEED = 123","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:49:47.977423Z","iopub.execute_input":"2024-03-24T05:49:47.977808Z","iopub.status.idle":"2024-03-24T05:49:51.674503Z","shell.execute_reply.started":"2024-03-24T05:49:47.977778Z","shell.execute_reply":"2024-03-24T05:49:51.673420Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"1\"></a>\n# <b>1 <span style='color:#53599A'>Theory</span></b>\n\n<a id=\"1.1\"></a>\n## <b>1.1 <span style='color:#53599A'>TTC vs PIT models</span></b>\n\nOriginally point in time (**PIT**) and through the cycle (**TTC**) model credit scoring model definitions were introduced by the Basel committee in the early 2000s [**[1]**](#ref1). In the recent European Banking Authority report [**[2]**](#ref2), much more emphasis was placed on the rating philosophy,i.e. how the business cycle interacts with the\nrelated systematic variability in default experience when evaluating models' stability. This means that before to any validation it should be identified whether the model is PIT, TTC, or in between, i.e. hybrid. \n\nLet's investigate those two edge cases. During the economic cycle, one can observe periods/quarters of increased observed default rates (**ODF**), i.e. economic downturn. There are also periods considered \"good times\", e.g. e.g. low-interest rates, when low default rates are observed.\n\n![TTC vs PIT](https://www.dropbox.com/scl/fi/6qa7ezn0q43lpmye28h3d/ttc_pit_1.png?rlkey=8ipxfnncd09ywsenobeoxvkxs&dl=1)\n\nPIT models are sensitive to the economic cycle and reflect the borrower’s current ability to repay debt. On the other hand, TTC models are less sensitive to economic cycles and reflect the borrower’s ability to repay debt over an entire economic cycle. Hybrid models fall somewhere in between. These models are used for different risk management purposes.\n\n<a id=\"1.2\"></a>\n## <b>1.2 <span style='color:#53599A'>Data litimations</span></b>\n\nBefore starting to build a new credit scoring model one should always take into account challenges due to data limitations:\n\n* Credit risk models rely on past loan performance to predict future defaults. However, if the available historical data is sparse or covers only a short period, it becomes difficult to capture the full range of credit events, e.g. years with a large number of defaults and years/quarters with low ODFs;\n* To build a robust model it requires a substantial dataset spanning various economic cycles, i.e. both good and bad times. Limited historical data may lead to biased estimates and inadequate risk assessments;\n* Take into account data quality especially if third-party data sources are used. For example, due to changed data governance rules/regulations the quality of external risk rating vendors could change as well;\n* It is also advised to incorporate macro variables into modeling which could be used as a proxy to estimate missing economical cycles from the available data about customers.\n\nBelow is a theoretical example of what could happen if the model is built without taking data limitations into account. [A practical example](#example.1) using competition data can be found in here [3.2.1 Out-of-time split](#3.2.1) sub-section.\n\n![Example #1](https://www.dropbox.com/scl/fi/75xybcuq0mzlsjxsrlri0/ttc_pit_example_1.png?rlkey=fkvhtkwycpdubdo21prd81vz5&dl=1)\n\n<a id=\"1.3\"></a>\n## <b>1.3 <span style='color:#53599A'>Estimating model's time horizon</span></b>\n\nIn one of the most recent EBA report [**[2]**](#ref2) for the *Assessment of the core model performance validation function needs to*:\n\n> Risk quantification: The estimates should meet all regulatory requirements. To this end, the validation of the risk parameter estimates should include a comparison of realised DR with estimated PDs for each grade or pool, and analogous analysis for LGDs and CFs where institutions received permission to use own estimates for those risk parameters, taking into account **the rating philosophy**.\n\nIn other words, before the model is put into production, model developers and model validators need to evaluate whatever new model is TTC, PIT, or hybrid as explained above. Below is an example using synthetic data that illustrates error depending on time, i.e. whatever model is more TTC or PIT.\n\n<div>\n<br>\n<a href=\"#toc\" style=\"background-color: #607BB0; color: #ffffff; padding: 7px 10px; text-decoration: none; border-radius: 50px;\">Back to top</a><a id=\"toc\"></a>\n</div>","metadata":{}},{"cell_type":"code","source":"# syntetic quaterly data for pricing model\nquarters = np.arange(0, 60)\n\n# generate random downturn quarters (make all predictions positive values)\ndownturn = abs(np.random.normal(5, 2, 20))\ndownturn.sort()\ndownturn[10:] = downturn[9::-1]\n\n# generate random good quarters (make all predictions positive values)\ngoodtimes = abs(np.random.normal(1.3, 0.5, 40))\n\n# observed defaults frequncies\nodf = np.array(downturn.tolist() + goodtimes.tolist())\n# add more noise\nfor i in range(0, 60, 3):\n    np.random.shuffle(odf[i:i+3])\n    \n# fake PIT model\npit = odf * np.random.uniform(0.7, 1.3, 60)\nttc = odf.mean() * np.random.uniform(0.7, 1.3, 60)\n\nfig, ax = plt.subplots(1, 1, figsize=(16, 4))\nax.plot(quarters[:20], odf[:20], \"o\", color = \"r\", label = \"Historical data (downturn)\", markersize = 10)\nax.plot(quarters[20:], odf[20:], \"o\", color = \"g\", label = \"Historical data (good times)\", markersize = 10)\n\nax.plot(quarters, pit, \"-\", label = \"PIT model\", zorder = -1, linewidth=4)\nax.plot(quarters, ttc, \"-\", label = \"TTC model\", zorder = -1, linewidth=4)\n\nax.set_xlabel(\"Quarter\", fontsize = 14)\nax.set_ylabel(\"Avg. ODF/PD, %\", fontsize = 14)\n\n# add white background to legend\nlegend = ax.legend(frameon=1)\nframe = legend.get_frame()\nframe.set_facecolor('w')\n\nplt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:49:51.676844Z","iopub.execute_input":"2024-03-24T05:49:51.677665Z","iopub.status.idle":"2024-03-24T05:49:52.262928Z","shell.execute_reply.started":"2024-03-24T05:49:51.677603Z","shell.execute_reply":"2024-03-24T05:49:52.261628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 2, figsize=(16, 4))\n\nsns.regplot(x=pit, y=odf, ax=ax[0], label=f\"Corr. {np.corrcoef(pit, odf)[0, 1]:.3f}\")\nsns.regplot(x=ttc, y=odf, ax=ax[1], label=f\"Corr. {np.corrcoef(ttc, odf)[0, 1]:.3f}\", color = COLORS[1])\n\nax[0].set_xlabel(\"PIT PD, %\", fontsize = 14)\nax[1].set_xlabel(\"TTC PD, %\", fontsize = 14)\n\nfor i in range(2):\n    ax[i].set_ylabel(\"ODF\", fontsize = 14)\n    # add white background to legend\n    legend = ax[i].legend(frameon=1)\n    frame = legend.get_frame()\n    frame.set_facecolor('w')\n\nplt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:49:52.264912Z","iopub.execute_input":"2024-03-24T05:49:52.265713Z","iopub.status.idle":"2024-03-24T05:49:53.199826Z","shell.execute_reply.started":"2024-03-24T05:49:52.265673Z","shell.execute_reply":"2024-03-24T05:49:53.198686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"2\"></a>\n# <b>2 <span style='color:#53599A'>Prepare project</span></b>\n\n<a id=\"2.1\"></a>\n## <b>2.1 <span style='color:#53599A'>Load data</span></b>\n\n<div>\n<br>\n<a href=\"#toc\" style=\"background-color: #607BB0; color: #ffffff; padding: 7px 10px; text-decoration: none; border-radius: 50px;\">Back to top</a><a id=\"toc\"></a>\n</div>","metadata":{}},{"cell_type":"code","source":"def set_table_dtypes(df: pl.DataFrame) -> pl.DataFrame:\n    \"\"\"\n    Helper function taken from notebook:\n    https://www.kaggle.com/code/jetakow/home-credit-2024-starter-notebook\n    \"\"\"\n    # implement here all desired dtypes for tables\n    # the following is just an example\n    for col in df.columns:\n        # last letter of column name will help you determine the type\n        if col[-1] in (\"P\", \"A\"):\n            df = df.with_columns(pl.col(col).cast(pl.Float64).alias(col))\n\n    return df\n\ndef convert_strings(df: pd.DataFrame) -> pd.DataFrame:\n    \"\"\"\n    Helper function taken from notebook:\n    https://www.kaggle.com/code/jetakow/home-credit-2024-starter-notebook\n    \"\"\"\n    for col in df.columns:  \n        if df[col].dtype.name in ['object', 'string']:\n            df[col] = df[col].astype(\"string\").astype('category')\n            current_categories = df[col].cat.categories\n            new_categories = current_categories.to_list() + [\"Unknown\"]\n            new_dtype = pd.CategoricalDtype(categories=new_categories, ordered=True)\n            df[col] = df[col].astype(new_dtype)\n    return df","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:49:53.203195Z","iopub.execute_input":"2024-03-24T05:49:53.204104Z","iopub.status.idle":"2024-03-24T05:49:53.217055Z","shell.execute_reply.started":"2024-03-24T05:49:53.204058Z","shell.execute_reply":"2024-03-24T05:49:53.215671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# training dataset\ntrain_basetable = pl.read_csv(PATH + \"csv_files/train/train_base.csv\")\ntrain_static = pl.concat(\n    [\n        pl.read_csv(PATH + \"csv_files/train/train_static_0_0.csv\").pipe(set_table_dtypes),\n        pl.read_csv(PATH + \"csv_files/train/train_static_0_1.csv\").pipe(set_table_dtypes),\n    ],\n    how=\"vertical_relaxed\",\n)\ntrain_static_cb = pl.read_csv(PATH + \"csv_files/train/train_static_cb_0.csv\").pipe(set_table_dtypes)\ntrain_person_1 = pl.read_csv(PATH + \"csv_files/train/train_person_1.csv\").pipe(set_table_dtypes) \ntrain_credit_bureau_b_2 = pl.read_csv(PATH + \"csv_files/train/train_credit_bureau_b_2.csv\").pipe(set_table_dtypes) ","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:49:53.219515Z","iopub.execute_input":"2024-03-24T05:49:53.219916Z","iopub.status.idle":"2024-03-24T05:50:11.413418Z","shell.execute_reply.started":"2024-03-24T05:49:53.219886Z","shell.execute_reply":"2024-03-24T05:50:11.411553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# test dataset\ntest_basetable = pl.read_csv(PATH + \"csv_files/test/test_base.csv\")\ntest_static = pl.concat(\n    [\n        pl.read_csv(PATH + \"csv_files/test/test_static_0_0.csv\").pipe(set_table_dtypes),\n        pl.read_csv(PATH + \"csv_files/test/test_static_0_1.csv\").pipe(set_table_dtypes),\n        pl.read_csv(PATH + \"csv_files/test/test_static_0_2.csv\").pipe(set_table_dtypes),\n    ],\n    how=\"vertical_relaxed\",\n)\ntest_static_cb = pl.read_csv(PATH + \"csv_files/test/test_static_cb_0.csv\").pipe(set_table_dtypes)\ntest_person_1 = pl.read_csv(PATH + \"csv_files/test/test_person_1.csv\").pipe(set_table_dtypes) \ntest_credit_bureau_b_2 = pl.read_csv(PATH + \"csv_files/test/test_credit_bureau_b_2.csv\").pipe(set_table_dtypes) ","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:50:11.415521Z","iopub.execute_input":"2024-03-24T05:50:11.415922Z","iopub.status.idle":"2024-03-24T05:50:11.483463Z","shell.execute_reply.started":"2024-03-24T05:50:11.415891Z","shell.execute_reply":"2024-03-24T05:50:11.482090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"2.2\"></a>\n## <b>2.2 <span style='color:#53599A'>Feature engineering</span></b>\n\n<div>\n<br>\n<a href=\"#toc\" style=\"background-color: #607BB0; color: #ffffff; padding: 7px 10px; text-decoration: none; border-radius: 50px;\">Back to top</a><a id=\"toc\"></a>\n</div>","metadata":{}},{"cell_type":"code","source":"# We need to use aggregation functions in tables with depth > 1, so tables that contain num_group1 column or \n# also num_group2 column.\ntrain_person_1_feats_1 = train_person_1.group_by(\"case_id\").agg(\n    pl.col(\"mainoccupationinc_384A\").max().alias(\"mainoccupationinc_384A_max\"),\n    (pl.col(\"incometype_1044T\") == \"SELFEMPLOYED\").max().alias(\"mainoccupationinc_384A_any_selfemployed\")\n)\n\n# Here num_group1=0 has special meaning, it is the person who applied for the loan.\ntrain_person_1_feats_2 = train_person_1.select([\"case_id\", \"num_group1\", \"housetype_905L\"]).filter(\n    pl.col(\"num_group1\") == 0\n).drop(\"num_group1\").rename({\"housetype_905L\": \"person_housetype\"})\n\n# Here we have num_goup1 and num_group2, so we need to aggregate again.\ntrain_credit_bureau_b_2_feats = train_credit_bureau_b_2.group_by(\"case_id\").agg(\n    pl.col(\"pmts_pmtsoverdue_635A\").max().alias(\"pmts_pmtsoverdue_635A_max\"),\n    (pl.col(\"pmts_dpdvalue_108P\") > 31).max().alias(\"pmts_dpdvalue_108P_over31\")\n)\n\n# We will process in this examples only A-type and M-type columns, so we need to select them.\nselected_static_cols = []\nfor col in train_static.columns:\n    if col[-1] in (\"A\", \"M\"):\n        selected_static_cols.append(col)\n\nselected_static_cb_cols = []\nfor col in train_static_cb.columns:\n    if col[-1] in (\"A\", \"M\"):\n        selected_static_cb_cols.append(col)\n\n# Join all tables together.\ndata = train_basetable.join(\n    train_static.select([\"case_id\"]+selected_static_cols), how=\"left\", on=\"case_id\"\n).join(\n    train_static_cb.select([\"case_id\"]+selected_static_cb_cols), how=\"left\", on=\"case_id\"\n).join(\n    train_person_1_feats_1, how=\"left\", on=\"case_id\"\n).join(\n    train_person_1_feats_2, how=\"left\", on=\"case_id\"\n).join(\n    train_credit_bureau_b_2_feats, how=\"left\", on=\"case_id\"\n)","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:50:11.485247Z","iopub.execute_input":"2024-03-24T05:50:11.485701Z","iopub.status.idle":"2024-03-24T05:50:14.143295Z","shell.execute_reply.started":"2024-03-24T05:50:11.485654Z","shell.execute_reply":"2024-03-24T05:50:14.142386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_person_1_feats_1 = test_person_1.group_by(\"case_id\").agg(\n    pl.col(\"mainoccupationinc_384A\").max().alias(\"mainoccupationinc_384A_max\"),\n    (pl.col(\"incometype_1044T\") == \"SELFEMPLOYED\").max().alias(\"mainoccupationinc_384A_any_selfemployed\")\n)\n\ntest_person_1_feats_2 = test_person_1.select([\"case_id\", \"num_group1\", \"housetype_905L\"]).filter(\n    pl.col(\"num_group1\") == 0\n).drop(\"num_group1\").rename({\"housetype_905L\": \"person_housetype\"})\n\ntest_credit_bureau_b_2_feats = test_credit_bureau_b_2.group_by(\"case_id\").agg(\n    pl.col(\"pmts_pmtsoverdue_635A\").max().alias(\"pmts_pmtsoverdue_635A_max\"),\n    (pl.col(\"pmts_dpdvalue_108P\") > 31).max().alias(\"pmts_dpdvalue_108P_over31\")\n)\n\ndata_submission = test_basetable.join(\n    test_static.select([\"case_id\"]+selected_static_cols), how=\"left\", on=\"case_id\"\n).join(\n    test_static_cb.select([\"case_id\"]+selected_static_cb_cols), how=\"left\", on=\"case_id\"\n).join(\n    test_person_1_feats_1, how=\"left\", on=\"case_id\"\n).join(\n    test_person_1_feats_2, how=\"left\", on=\"case_id\"\n).join(\n    test_credit_bureau_b_2_feats, how=\"left\", on=\"case_id\"\n)","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:50:14.144776Z","iopub.execute_input":"2024-03-24T05:50:14.145168Z","iopub.status.idle":"2024-03-24T05:50:14.161834Z","shell.execute_reply.started":"2024-03-24T05:50:14.145132Z","shell.execute_reply":"2024-03-24T05:50:14.160667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"3\"></a>\n# <b>3 <span style='color:#53599A'>Feature analysis</span></b>\n\n<a id=\"3.1\"></a>\n## <b>3.1 <span style='color:#53599A'>Historical default rates</span></b>\n\nThere are 22 months or 5.5 quarters of data provided. In 2020-04 one can observe a sharp decrease of ODF. In addition, the number of observations also is smaller as compared to periods before 2020-04 (see right y-axis label).\n\n<div>\n<br>\n<a href=\"#toc\" style=\"background-color: #607BB0; color: #ffffff; padding: 7px 10px; text-decoration: none; border-radius: 50px;\">Back to top</a><a id=\"toc\"></a>\n</div>","metadata":{}},{"cell_type":"code","source":"# select data for plotting\ndf_plot = data[['MONTH', 'target']].to_pandas().groupby('MONTH').agg({'target': ['mean', 'count']}).reset_index()\ndf_plot.columns = [\"\".join(_) for _ in df_plot.columns]\n# convert to date\ndf_plot['Date'] = pd.to_datetime(df_plot['MONTH'], format='%Y%m', errors='coerce')\n\nfig = plt.figure(constrained_layout=True, figsize=(16, 6))\ngs = fig.add_gridspec(2, 3)\nax1 = fig.add_subplot(gs[0, :-1])\nax2 = fig.add_subplot(gs[1, :-1])\nax3 = fig.add_subplot(gs[:, -1:])\n\nax1.plot(df_plot[\"Date\"], df_plot['targetmean'] * 100, \"o-\", color = COLORS[1])\nax2.plot(df_plot[\"Date\"], df_plot['targetcount'] / 1_000, \"o-\", color = COLORS[2])\nsns.regplot(x=df_plot['targetmean'] * 100, y=df_plot['targetcount'] / 1_000, ax=ax3)\n# calculate correlation Pearson coefficient\n_corr = df_plot[['targetmean', 'targetcount']].corr().iloc[0, 1]\n\nax1.set_xlabel(\"Date\", fontsize=12)\nax1.set_ylabel(\"Monthly ODF, %\", fontsize=12)\nax2.set_xlabel(\"Date\", fontsize=12)\nax2.set_ylabel(\"Number of observations, k\", fontsize=12)\nax3.set_xlabel(\"Monthly ODF, %\", fontsize=12)\nax3.set_ylabel(\"Number of observations, k\", fontsize=12)\nax3.set_title(f\"Pearson correlation coefficient {_corr:.3f}\");","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:50:14.163773Z","iopub.execute_input":"2024-03-24T05:50:14.164670Z","iopub.status.idle":"2024-03-24T05:50:15.950775Z","shell.execute_reply.started":"2024-03-24T05:50:14.164625Z","shell.execute_reply":"2024-03-24T05:50:15.949813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"80% of observations are up to `2020-FEB`.  while 84.7% up to structural change.","metadata":{}},{"cell_type":"code","source":"df_plot[\"targetcount_cumsum\"] = df_plot['targetcount'].cumsum()\ndf_plot[\"Share, %\"] = (df_plot[\"targetcount_cumsum\"] / df_plot['targetcount'].sum() * 100).round(3)\n_df = df_plot[[\"Date\", \"targetcount\", \"targetcount_cumsum\", \"Share, %\"]]\n_df.columns = [\"Date\", \"Observations, N\", \"Observations cumsum, N\", \"Share, %\"]\n_df","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:50:15.956204Z","iopub.execute_input":"2024-03-24T05:50:15.956957Z","iopub.status.idle":"2024-03-24T05:50:15.980288Z","shell.execute_reply.started":"2024-03-24T05:50:15.956920Z","shell.execute_reply":"2024-03-24T05:50:15.979444Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"3.2\"></a>\n## <b>3.2 <span style='color:#53599A'>Train/test/valid split</span></b>\n\nLet's investigate the differences of model's performance which depends on how train/test/validation samples were split:\n\n* random splitting;\n* out-of-time splitting.\n\n<a id=\"3.2.1\"></a>\n## <b>3.2.1 <span style='color:#53599A'>Out-of-time splitting</span></b>\n\nInstead of a random splitting, out of time splitting is should used in order to access historical economical conditions.\n\n<div>\n<br>\n<a href=\"#toc\" style=\"background-color: #607BB0; color: #ffffff; padding: 7px 10px; text-decoration: none; border-radius: 50px;\">Back to top</a><a id=\"toc\"></a>\n</div>","metadata":{}},{"cell_type":"code","source":"# select data for plotting\ndf_plot = data[['MONTH', 'target']].to_pandas().groupby('MONTH').agg({'target': ['mean', 'count']}).reset_index()\ndf_plot.columns = [\"\".join(_) for _ in df_plot.columns]\n# convert to date\ndf_plot['Date'] = pd.to_datetime(df_plot['MONTH'], format='%Y%m', errors='coerce')\n\nfig, ax = plt.subplots(1, 1, figsize=(16, 4))\n\n_oot_date = 202004\n\n_size = df_plot.loc[df_plot.MONTH < _oot_date, \"targetcount\"].sum() / df_plot[\"targetcount\"].sum() * 100\nax.plot(df_plot.loc[df_plot.MONTH < _oot_date, \"Date\"], df_plot.loc[df_plot.MONTH < _oot_date, \"targetmean\"] * 100, \"o-\", label = f\"Train + valid sample ({_size:.2f}%)\")\nax.plot(df_plot.loc[df_plot.MONTH >= _oot_date, \"Date\"], df_plot.loc[df_plot.MONTH >= _oot_date, \"targetmean\"] * 100, \"o-\", label = f\"Test sample ({100-_size:.2f}%)\")\n\nax.set_xlabel(\"Date\", fontsize = 14)\nax.set_ylabel(\"ODF, %\", fontsize = 14)\n\nax.set_title(\"Out of time split\", fontsize = 16)\n\n# add white background to legend\nlegend = ax.legend(frameon=1)\nframe = legend.get_frame()\nframe.set_facecolor('w')\n\nplt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:50:15.981383Z","iopub.execute_input":"2024-03-24T05:50:15.982265Z","iopub.status.idle":"2024-03-24T05:50:16.686964Z","shell.execute_reply.started":"2024-03-24T05:50:15.982232Z","shell.execute_reply":"2024-03-24T05:50:16.686018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# select IDs for training, validation and testing\ncase_ids_train = data.filter(pl.col(\"MONTH\") < 202004)[\"case_id\"].unique().shuffle(seed=SEED)\ncase_ids_test = data.filter(pl.col(\"MONTH\") >= 202004)[\"case_id\"].unique().shuffle(seed=SEED)\ncase_ids_train, case_ids_valid = train_test_split(case_ids_train, train_size=0.8, random_state=SEED)\n\ncols_pred = []\nfor col in data.columns:\n    if col[-1].isupper() and col[:-1].islower():\n        cols_pred.append(col)\n        \ndef from_polars_to_pandas(case_ids: pl.DataFrame) -> pl.DataFrame:\n    \"\"\"\n    Function taken from notebook:\n    https://www.kaggle.com/code/jetakow/home-credit-2024-starter-notebook\n    \"\"\"\n    return (\n        data.filter(pl.col(\"case_id\").is_in(case_ids))[[\"case_id\", \"WEEK_NUM\", \"MONTH\", \"target\"]].to_pandas(),\n        data.filter(pl.col(\"case_id\").is_in(case_ids))[cols_pred].to_pandas(),\n        data.filter(pl.col(\"case_id\").is_in(case_ids))[\"target\"].to_pandas()\n    )\n\n# select data for different samples\nbase_train, X_train, y_train = from_polars_to_pandas(case_ids_train)\nbase_valid, X_valid, y_valid = from_polars_to_pandas(case_ids_valid)\nbase_test, X_test, y_test = from_polars_to_pandas(case_ids_test)\n\n# convert columns to strings\nfor df in [X_train, X_valid, X_test]:\n    df = convert_strings(df)\n    \n# save results to dict\nout_of_time_data = {\"base_train\": base_train, \"X_train\": X_train, \"y_train\": y_train,\n                    \"base_valid\": base_valid, \"X_valid\": X_valid, \"y_valid\": y_valid,\n                    \"base_test\": base_test, \"X_test\": X_test, \"y_test\": y_test}","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:50:16.688100Z","iopub.execute_input":"2024-03-24T05:50:16.689182Z","iopub.status.idle":"2024-03-24T05:50:25.630818Z","shell.execute_reply.started":"2024-03-24T05:50:16.689146Z","shell.execute_reply":"2024-03-24T05:50:25.629550Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Train basic `LightGBM` model (same params as in the example notebook).","metadata":{}},{"cell_type":"code","source":"%%time\n# oot - out of time\noot_lgb_train = lgb.Dataset(out_of_time_data[\"X_train\"], label=out_of_time_data[\"y_train\"])\noot_lgb_valid = lgb.Dataset(out_of_time_data[\"X_valid\"], label=out_of_time_data[\"y_valid\"], reference=oot_lgb_train)\n\n# use the same params for random spliting\nPARAMS = {\n    \"boosting_type\": \"gbdt\",\n    \"objective\": \"binary\",\n    \"metric\": \"auc\",\n    \"max_depth\": 3,\n    \"num_leaves\": 31,\n    \"learning_rate\": 0.05,\n    \"feature_fraction\": 0.9,\n    \"bagging_fraction\": 0.8,\n    \"bagging_freq\": 5,\n    \"n_estimators\": 1000,\n    \"verbose\": -1,\n}\n\ngbm_oot = lgb.train(\n    PARAMS,\n    oot_lgb_train,\n    valid_sets = oot_lgb_valid,\n    callbacks = [lgb.log_evaluation(50), lgb.early_stopping(10)]\n)","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:50:25.632043Z","iopub.execute_input":"2024-03-24T05:50:25.632365Z","iopub.status.idle":"2024-03-24T05:52:13.568875Z","shell.execute_reply.started":"2024-03-24T05:50:25.632338Z","shell.execute_reply":"2024-03-24T05:52:13.567663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# add model and data which was trained on out of time sample\nALL_MODELS['OOT_basic'] = {\"model\": gbm_oot, \"data\": out_of_time_data, \"lgb_train\": oot_lgb_train, \"lgb_valid\": oot_lgb_valid}","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:52:13.570405Z","iopub.execute_input":"2024-03-24T05:52:13.571237Z","iopub.status.idle":"2024-03-24T05:52:13.576674Z","shell.execute_reply.started":"2024-03-24T05:52:13.571205Z","shell.execute_reply":"2024-03-24T05:52:13.575396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def make_predictions(gbm_model, gbm_data):\n    \"\"\"\n    Custom function to make predictions inplace using gbm\n    \"\"\"\n    # create data list\n    data_lst = [(gbm_data['base_train'], gbm_data['X_train']),\n                (gbm_data['base_valid'], gbm_data['X_valid']),\n                (gbm_data['base_test'], gbm_data['X_test'])]\n    for base, X in data_lst:\n        y_pred = gbm_model.predict(X, num_iteration=gbm_model.best_iteration)\n        base[\"pred\"] = y_pred","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:52:13.578575Z","iopub.execute_input":"2024-03-24T05:52:13.579041Z","iopub.status.idle":"2024-03-24T05:52:13.590913Z","shell.execute_reply.started":"2024-03-24T05:52:13.579000Z","shell.execute_reply":"2024-03-24T05:52:13.589759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nmake_predictions(gbm_model = ALL_MODELS[\"OOT_basic\"][\"model\"], gbm_data = ALL_MODELS[\"OOT_basic\"][\"data\"])","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:52:13.592346Z","iopub.execute_input":"2024-03-24T05:52:13.592690Z","iopub.status.idle":"2024-03-24T05:52:36.697276Z","shell.execute_reply.started":"2024-03-24T05:52:13.592656Z","shell.execute_reply":"2024-03-24T05:52:36.696077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"example.1\"></a>\nPractical example when model is build on limited historical data.","metadata":{}},{"cell_type":"code","source":"# select data for plotting\ndf_plot = pd.concat([ALL_MODELS[\"OOT_basic\"][\"data\"][\"base_train\"],\n                     ALL_MODELS[\"OOT_basic\"][\"data\"][\"base_valid\"],\n                     ALL_MODELS[\"OOT_basic\"][\"data\"][\"base_test\"]])\ndf_plot = df_plot.groupby('MONTH').agg({'target': ['mean', 'count'], \"pred\": ['mean']}).reset_index()\ndf_plot.columns = [\"\".join(_) for _ in df_plot.columns]\n# convert to date\ndf_plot['Date'] = pd.to_datetime(df_plot['MONTH'], format='%Y%m', errors='coerce')\n\nfig, ax = plt.subplots(1, 1, figsize=(16, 4))\n\n_size = df_plot.loc[df_plot.MONTH < 202004, \"targetcount\"].sum() / df_plot[\"targetcount\"].sum() * 100\nax.scatter(df_plot.loc[df_plot.MONTH < 202004, \"Date\"], df_plot.loc[df_plot.MONTH < 202004, \"targetmean\"] * 100,  label = f\"Train + valid sample ({_size:.2f}%)\")\nax.scatter(df_plot.loc[df_plot.MONTH >= 202004, \"Date\"], df_plot.loc[df_plot.MONTH >= 202004, \"targetmean\"] * 100, label = f\"Test sample ({100-_size:.2f}%)\")\n\n# plot average predictions\nax.plot(df_plot[\"Date\"], df_plot[\"predmean\"] * 100, \"o-\", label = \"Average predicted scores each month\", color = \"red\", zorder = 10)\n\nax.set_xlabel(\"Date\", fontsize = 14)\nax.set_ylabel(\"ODF, %\", fontsize = 14)\n\nax.set_title(\"Predictions using model trained on out of time split\", fontsize = 16)\n\n# add white background to legend\nlegend = ax.legend(frameon=1)\nframe = legend.get_frame()\nframe.set_facecolor('w')\n\nplt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:52:36.698476Z","iopub.execute_input":"2024-03-24T05:52:36.698796Z","iopub.status.idle":"2024-03-24T05:52:37.506335Z","shell.execute_reply.started":"2024-03-24T05:52:36.698769Z","shell.execute_reply":"2024-03-24T05:52:37.504952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"3.2.2\"></a>\n## <b>3.2.2 <span style='color:#53599A'>Random splitting</span></b>\n\n<div>\n<br>\n<a href=\"#toc\" style=\"background-color: #607BB0; color: #ffffff; padding: 7px 10px; text-decoration: none; border-radius: 50px;\">Back to top</a><a id=\"toc\"></a>\n</div>","metadata":{}},{"cell_type":"code","source":"%%time\n# select IDs for training\ncase_ids = data[\"case_id\"].unique().shuffle(seed=SEED)\ncase_ids_train, case_ids_test = train_test_split(case_ids, train_size=0.6, random_state=SEED)\ncase_ids_valid, case_ids_test = train_test_split(case_ids_test, train_size=0.5, random_state=SEED)\n\n# select data for different samples\nbase_train, X_train, y_train = from_polars_to_pandas(case_ids_train)\nbase_valid, X_valid, y_valid = from_polars_to_pandas(case_ids_valid)\nbase_test, X_test, y_test = from_polars_to_pandas(case_ids_test)\n\n# convert columns to strings\nfor df in [X_train, X_valid, X_test]:\n    df = convert_strings(df)\n    \n# save results to dict\nrandom_split_data = {\"base_train\": base_train, \"X_train\": X_train, \"y_train\": y_train,\n                     \"base_valid\": base_valid, \"X_valid\": X_valid, \"y_valid\": y_valid,\n                     \"base_test\": base_test, \"X_test\": X_test, \"y_test\": y_test}","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:52:37.507990Z","iopub.execute_input":"2024-03-24T05:52:37.508462Z","iopub.status.idle":"2024-03-24T05:52:46.062539Z","shell.execute_reply.started":"2024-03-24T05:52:37.508421Z","shell.execute_reply":"2024-03-24T05:52:46.061185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# rs - random split\nrs_lgb_train = lgb.Dataset(random_split_data[\"X_train\"], label=random_split_data[\"y_train\"])\nrs_lgb_valid = lgb.Dataset(random_split_data[\"X_valid\"], label=random_split_data[\"y_valid\"], reference=rs_lgb_train)\n\n# use the same params for random spliting\nPARAMS = {\n    \"boosting_type\": \"gbdt\",\n    \"objective\": \"binary\",\n    \"metric\": \"auc\",\n    \"max_depth\": 3,\n    \"num_leaves\": 31,\n    \"learning_rate\": 0.05,\n    \"feature_fraction\": 0.9,\n    \"bagging_fraction\": 0.8,\n    \"bagging_freq\": 5,\n    \"n_estimators\": 1000,\n    \"verbose\": -1,\n}\n\ngbm_rs = lgb.train(\n    PARAMS,\n    rs_lgb_train,\n    valid_sets = rs_lgb_valid,\n    callbacks = [lgb.log_evaluation(50), lgb.early_stopping(10)]\n)","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:52:46.064135Z","iopub.execute_input":"2024-03-24T05:52:46.064584Z","iopub.status.idle":"2024-03-24T05:54:26.398670Z","shell.execute_reply.started":"2024-03-24T05:52:46.064544Z","shell.execute_reply":"2024-03-24T05:54:26.397743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# add model and data which was trained on out of time sample\nALL_MODELS['RANDOM_basic'] = {\"model\": gbm_rs, \"data\": random_split_data,\n                              \"lgb_train\": rs_lgb_train, \"lgb_valid\": rs_lgb_valid}","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:54:26.399730Z","iopub.execute_input":"2024-03-24T05:54:26.400234Z","iopub.status.idle":"2024-03-24T05:54:26.405967Z","shell.execute_reply.started":"2024-03-24T05:54:26.400205Z","shell.execute_reply":"2024-03-24T05:54:26.404686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nmake_predictions(gbm_model = ALL_MODELS[\"RANDOM_basic\"][\"model\"], gbm_data = ALL_MODELS[\"RANDOM_basic\"][\"data\"])","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:54:26.407699Z","iopub.execute_input":"2024-03-24T05:54:26.408032Z","iopub.status.idle":"2024-03-24T05:54:50.351969Z","shell.execute_reply.started":"2024-03-24T05:54:26.408004Z","shell.execute_reply":"2024-03-24T05:54:50.349084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# select data for plotting\ndf_plot = pd.concat([ALL_MODELS[\"RANDOM_basic\"][\"data\"][\"base_train\"],\n                     ALL_MODELS[\"RANDOM_basic\"][\"data\"][\"base_valid\"],\n                     ALL_MODELS[\"RANDOM_basic\"][\"data\"][\"base_test\"]])\ndf_plot = df_plot.groupby('MONTH').agg({'target': ['mean', 'count'], \"pred\": ['mean']}).reset_index()\ndf_plot.columns = [\"\".join(_) for _ in df_plot.columns]\n# convert to date\ndf_plot['Date'] = pd.to_datetime(df_plot['MONTH'], format='%Y%m', errors='coerce')\n\nfig, ax = plt.subplots(1, 1, figsize=(16, 4))\n\nax.scatter(df_plot[\"Date\" ], df_plot[\"targetmean\"] * 100, label = \"All known data sample\")\n\n# plot average predictions\nax.plot(df_plot[\"Date\"], df_plot[\"predmean\"] * 100, \"o-\", label = f\"Average predicted scores each month\", color = \"red\", zorder = 10)\n\nax.set_xlabel(\"Date\", fontsize = 14)\nax.set_ylabel(\"ODF, %\", fontsize = 14)\n\nax.set_title(\"Predictions using model trained on random sample split\", fontsize = 16)\n\n# add white background to legend\nlegend = ax.legend(frameon=1)\nframe = legend.get_frame()\nframe.set_facecolor('w')\n\nplt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:54:50.353660Z","iopub.execute_input":"2024-03-24T05:54:50.354301Z","iopub.status.idle":"2024-03-24T05:54:51.095728Z","shell.execute_reply.started":"2024-03-24T05:54:50.354270Z","shell.execute_reply":"2024-03-24T05:54:51.094866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"3.2.3\"></a>\n## <b>3.2.3 <span style='color:#53599A'>Compare splits</span></b>\n\n<div>\n<br>\n<a href=\"#toc\" style=\"background-color: #607BB0; color: #ffffff; padding: 7px 10px; text-decoration: none; border-radius: 50px;\">Back to top</a><a id=\"toc\"></a>\n</div>","metadata":{}},{"cell_type":"code","source":"def gini_stability(base, w_fallingrate=88.0, w_resstd=-0.5):\n    \"\"\"\n    Sligtly modified function from notebook:\n    https://www.kaggle.com/code/jetakow/home-credit-2024-starter-notebook\n    \"\"\"\n    # calculate weekly gini\n    gini_in_time = base.loc[:, [\"WEEK_NUM\", \"target\", \"pred\"]]\\\n        .sort_values(\"WEEK_NUM\")\\\n        .groupby(\"WEEK_NUM\")[[\"target\", \"pred\"]]\\\n        .apply(lambda x: 2*roc_auc_score(x[\"target\"], x[\"pred\"])-1)\n    \n    # calculate monthly gini\n    gini_in_time_m = base.loc[:, [\"MONTH\", \"target\", \"pred\"]]\\\n        .sort_values(\"MONTH\")\\\n        .groupby(\"MONTH\")[[\"target\", \"pred\"]]\\\n        .apply(lambda x: 2*roc_auc_score(x[\"target\"], x[\"pred\"])-1)\n    \n    # TODO: refractor this code chunk\n    x = gini_in_time.index.values\n    y = gini_in_time.tolist()\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    \n    # calculate competition score\n    score = avg_gini + w_fallingrate * min(0, a) + w_resstd * res_std\n    \n    # TODO: refractor this code chunk\n    x_m = np.arange(gini_in_time_m.shape[0])\n    y_m = gini_in_time_m.tolist()\n    a_m, b_m = np.polyfit(x_m, y_m, 1)\n    y_hat_m = a_m*x_m + b_m\n    residuals_m = y_m - y_hat_m\n    res_std_m = np.std(residuals_m)\n    avg_gini_m = np.mean(gini_in_time_m)\n    \n    # calculate competition score\n    score_m = avg_gini_m + w_fallingrate * min(0, a_m) + w_resstd * res_std_m\n    \n    # return average GINI scores DataFrame object\n    gini_in_time = gini_in_time.reset_index().rename(columns = {0: \"gini\"})\n    gini_in_time_m = gini_in_time_m.reset_index().rename(columns = {0: \"gini\"})\n    gini_in_time_m['MONTH'] = pd.to_datetime(gini_in_time_m['MONTH'], format='%Y%m', errors='coerce')\n    \n    # return results as dictonary\n    results = {\n        # weekly\n        \"gini_in_time\": gini_in_time,\n        \"avg_gini\": avg_gini,\n        \"w_fallingrate\": w_fallingrate,\n        \"a\": a,\n        \"b\": b,\n        \"w_resstd\": w_resstd,\n        \"res_std\": res_std,\n        \"score\": score,\n        # monthly\n        \"gini_in_time_m\": gini_in_time_m,\n        \"avg_gini_m\": avg_gini_m,\n        \"a_m\": a_m,\n        \"b_m\": b_m,\n        \"res_std_m\": res_std_m,\n        \"score_m\": score_m,\n    }\n    \n    return results","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:54:51.097242Z","iopub.execute_input":"2024-03-24T05:54:51.097882Z","iopub.status.idle":"2024-03-24T05:54:51.112781Z","shell.execute_reply.started":"2024-03-24T05:54:51.097851Z","shell.execute_reply":"2024-03-24T05:54:51.111709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# calculate stabitlity scores (ss) for different models and datasets\nss_oot_train = gini_stability(ALL_MODELS[\"OOT_basic\"][\"data\"][\"base_train\"])\nss_oot_valid = gini_stability(ALL_MODELS[\"OOT_basic\"][\"data\"][\"base_valid\"])\nss_oot_test = gini_stability(ALL_MODELS[\"OOT_basic\"][\"data\"][\"base_test\"])\n\nss_rs_train = gini_stability(ALL_MODELS[\"RANDOM_basic\"][\"data\"][\"base_train\"])\nss_rs_valid = gini_stability(ALL_MODELS[\"RANDOM_basic\"][\"data\"][\"base_valid\"])\nss_rs_test = gini_stability(ALL_MODELS[\"RANDOM_basic\"][\"data\"][\"base_test\"])","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:54:51.114230Z","iopub.execute_input":"2024-03-24T05:54:51.115181Z","iopub.status.idle":"2024-03-24T05:54:54.522137Z","shell.execute_reply.started":"2024-03-24T05:54:51.115141Z","shell.execute_reply":"2024-03-24T05:54:54.520814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Out-off-time split\nprint(f'The stability score on the train set using out-of-time split is: {ss_oot_train[\"score\"]:.3f}')\nprint(f'The stability score on the valid set using out-of-time split is: {ss_oot_valid[\"score\"]:.3f}')\nprint(f'The stability score on the test set using out-of-time split is: {ss_oot_test[\"score\"]:.3f}\\n')\n\n# Random split\nprint(f'The stability score on the train set using random split is: {ss_rs_train[\"score\"]:.3f}')\nprint(f'The stability score on the valid set using random split is: {ss_rs_valid[\"score\"]:.3f}')\nprint(f'The stability score on the test set using random split is: {ss_rs_test[\"score\"]:.3f}') ","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:54:54.523962Z","iopub.execute_input":"2024-03-24T05:54:54.525036Z","iopub.status.idle":"2024-03-24T05:54:54.533041Z","shell.execute_reply.started":"2024-03-24T05:54:54.524985Z","shell.execute_reply":"2024-03-24T05:54:54.531803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def add_gini(ax, data, row_idx, col_idx, label=\"unknown\", time_col_nm=\"WEEK_NUM\", a=\"\", b=\"\"):\n    \"\"\"\n    Helper function for visualizing temporal gini score changes\n    \"\"\"\n    # add temporal Gini score distributions\n    ax[row_idx, col_idx].scatter(data[time_col_nm], data[\"gini\"], label = label, color=COLORS[col_idx])\n    # add fitted lines if value are provided\n    if a != \"\":\n        # dashed line y values\n        if time_col_nm == \"WEEK_NUM\":\n            _y = a * data[time_col_nm] + b\n        else:\n            _y = a * data[time_col_nm].index + b\n        ax[row_idx, col_idx].plot(data[time_col_nm], _y,\n                                  color=\"r\", linestyle=\"--\",\n                                  label= \"$G_{\\mathrm{fit}}(i) =$\" + f\" {a:.3f} i + {b:.3f}\")","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:54:54.534830Z","iopub.execute_input":"2024-03-24T05:54:54.535774Z","iopub.status.idle":"2024-03-24T05:54:54.551320Z","shell.execute_reply.started":"2024-03-24T05:54:54.535731Z","shell.execute_reply":"2024-03-24T05:54:54.549805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(2, 3, figsize=(16, 6), sharex=True, sharey=True)\n\n# out-of-time\nadd_gini(ax=ax, data=ss_oot_train[\"gini_in_time\"], row_idx=0, col_idx=0, label=f\"train = {ss_oot_train['score']:.3f} score\", a=ss_oot_train['a'], b=ss_oot_train['b'])\nadd_gini(ax=ax, data=ss_oot_valid[\"gini_in_time\"], row_idx=0, col_idx=1, label=f\"valid = {ss_oot_valid['score']:.3f} score\", a=ss_oot_valid['a'], b=ss_oot_valid['b'])\nadd_gini(ax=ax, data=ss_oot_test[\"gini_in_time\"], row_idx=0, col_idx=2, label=f\"test = {ss_oot_test['score']:.3f} score\", a=ss_oot_test['a'], b=ss_oot_test['b'])\n\n# random split\nadd_gini(ax=ax, data=ss_rs_train[\"gini_in_time\"], row_idx=1, col_idx=0, label=f\"train = {ss_rs_train['score']:.3f} score\", a=ss_rs_train['a'], b=ss_rs_train['b'])\nadd_gini(ax=ax, data=ss_rs_valid[\"gini_in_time\"], row_idx=1, col_idx=1, label=f\"valid = {ss_rs_valid['score']:.3f} score\", a=ss_rs_valid['a'], b=ss_rs_valid['b'])\nadd_gini(ax=ax, data=ss_rs_test[\"gini_in_time\"], row_idx=1, col_idx=2, label=f\"test = {ss_rs_test['score']:.3f} score\", a=ss_rs_test['a'], b=ss_rs_test['b'])\n\n# add titles\nax[0, 1].set_title(\"Predictions using model trained on out-of-time sample split\", fontsize = 16)\nax[1, 1].set_title(\"Predictions using model trained on random split\", fontsize = 16)\n\nfor i in range(2):\n    ax[i, 0].set_ylabel(\"Gini\", fontsize = 14)\nfor i in range(3):\n    ax[1, i].set_xlabel(\"Week number\", fontsize = 14)\n\n# add white background to legend\nfor i in range(2):\n    for ii in range(3):\n        legend = ax[i, ii].legend(frameon=1)\n        frame = legend.get_frame()\n        frame.set_facecolor('w')\n\nplt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:54:54.552604Z","iopub.execute_input":"2024-03-24T05:54:54.552945Z","iopub.status.idle":"2024-03-24T05:54:56.830461Z","shell.execute_reply.started":"2024-03-24T05:54:54.552917Z","shell.execute_reply":"2024-03-24T05:54:56.829605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(2, 3, figsize=(16, 6), sharex=True, sharey=True)\n\n# out-of-time\nadd_gini(ax=ax, data=ss_oot_train[\"gini_in_time_m\"], row_idx=0, col_idx=0, label=f\"train = {ss_oot_train['score_m']:.3f} score\", time_col_nm=\"MONTH\", a=ss_oot_train['a_m'], b=ss_oot_train['b_m'])\nadd_gini(ax=ax, data=ss_oot_valid[\"gini_in_time_m\"], row_idx=0, col_idx=1, label=f\"valid = {ss_oot_valid['score_m']:.3f} score\", time_col_nm=\"MONTH\", a=ss_oot_valid['a_m'], b=ss_oot_valid['b_m'])\nadd_gini(ax=ax, data=ss_oot_test[\"gini_in_time_m\"], row_idx=0, col_idx=2, label=f\"test = {ss_oot_test['score_m']:.3f} score\", time_col_nm=\"MONTH\", a=ss_oot_test['a_m'], b=ss_oot_test['b_m'])\n\n# random split\nadd_gini(ax=ax, data=ss_rs_train[\"gini_in_time_m\"], row_idx=1, col_idx=0, label=f\"train = {ss_rs_train['score_m']:.3f} score\", time_col_nm=\"MONTH\", a=ss_rs_train['a_m'], b=ss_rs_train['b_m'])\nadd_gini(ax=ax, data=ss_rs_valid[\"gini_in_time_m\"], row_idx=1, col_idx=1, label=f\"valid = {ss_rs_valid['score_m']:.3f} score\", time_col_nm=\"MONTH\", a=ss_rs_valid['a_m'], b=ss_rs_valid['b_m'])\nadd_gini(ax=ax, data=ss_rs_test[\"gini_in_time_m\"], row_idx=1, col_idx=2, label=f\"test = {ss_rs_test['score_m']:.3f} score\", time_col_nm=\"MONTH\", a=ss_rs_test['a_m'], b=ss_rs_test['b_m'])\n\n# add titles\nax[0, 1].set_title(\"Predictions using model trained on out-of-time sample split\", fontsize = 16)\nax[1, 1].set_title(\"Predictions using model trained on random split\", fontsize = 16)\n\nfor i in range(2):\n    ax[i, 0].set_ylabel(\"Gini\", fontsize = 14)\nfor i in range(3):\n    ax[1, i].set_xlabel(\"Date\", fontsize = 14)\n\n# add white background to legend\nfor i in range(2):\n    for ii in range(3):\n        legend = ax[i, ii].legend(frameon=1)\n        frame = legend.get_frame()\n        frame.set_facecolor('w')\n\nplt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:54:56.835204Z","iopub.execute_input":"2024-03-24T05:54:56.836140Z","iopub.status.idle":"2024-03-24T05:55:00.256830Z","shell.execute_reply.started":"2024-03-24T05:54:56.836106Z","shell.execute_reply":"2024-03-24T05:55:00.255678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 2, figsize=(16, 4), sharex=True, sharey=True)\n\nfor __ in ['base_train', 'base_valid', 'base_test']:\n    _df = ALL_MODELS[\"OOT_basic\"][\"data\"][__]\n    fpr, tpr, _ = roc_curve(y_true=_df[\"target\"], y_score=_df[\"pred\"])\n    _auc = roc_auc_score(_df[\"target\"], _df[\"pred\"])\n    ax[0].plot(fpr, tpr, label = f'{__.split(\"_\")[1]} (AUC = {_auc:.3f})')\n    \n    _df = ALL_MODELS[\"RANDOM_basic\"][\"data\"][__]\n    fpr, tpr, _ =roc_curve(y_true=_df[\"target\"], y_score=_df[\"pred\"])\n    _auc = roc_auc_score(_df[\"target\"], _df[\"pred\"])\n    ax[1].plot(fpr, tpr, label = f'{__.split(\"_\")[1]} (AUC = {_auc:.3f})')\n    \n    \n# add titles\nax[0].set_title(\"Model trained on out-of-time sample split\", fontsize = 16)\nax[1].set_title(\"Model trained on random split\", fontsize = 16)\n\nfor i in range(2):\n    # add white background to legend\n    legend = ax[i].legend(frameon=1)\n    frame = legend.get_frame()\n    frame.set_facecolor('w')\n    # add labels\n    ax[i].set_xlabel(\"False positive rate\", fontsize = 14)\n    ax[i].set_ylabel(\"True positive rate\", fontsize = 14)","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:55:00.258357Z","iopub.execute_input":"2024-03-24T05:55:00.258798Z","iopub.status.idle":"2024-03-24T05:55:03.146799Z","shell.execute_reply.started":"2024-03-24T05:55:00.258760Z","shell.execute_reply":"2024-03-24T05:55:03.145577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"4\"></a>\n# <b>4 <span style='color:#53599A'>Modeling</span></b>\n\n<a id=\"4.1\"></a>\n## <b>4.1 <span style='color:#53599A'>Ensembling</span></b>\n\n<div>\n<br>\n<a href=\"#toc\" style=\"background-color: #607BB0; color: #ffffff; padding: 7px 10px; text-decoration: none; border-radius: 50px;\">Back to top</a><a id=\"toc\"></a>\n</div>","metadata":{}},{"cell_type":"code","source":"%%time\n# select IDs for training\ncase_ids = data[\"case_id\"].unique().shuffle(seed=SEED*2)\ncase_ids_train, case_ids_test = train_test_split(case_ids, train_size=0.8, random_state=SEED*2)\n\n# select data for different samples\nbase_train, X_train, y_train = from_polars_to_pandas(case_ids_train)\nbase_test, X_test, y_test = from_polars_to_pandas(case_ids_test)\n\n# convert columns to strings\nfor df in [X_train, X_test]:\n    df = convert_strings(df)\n    \n# save results to dict\nrandom_split_best_data = {\"base_train\": base_train, \"X_train\": X_train, \"y_train\": y_train,\n                          \"base_test\": base_test, \"X_test\": X_test, \"y_test\": y_test}","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:55:03.148272Z","iopub.execute_input":"2024-03-24T05:55:03.149523Z","iopub.status.idle":"2024-03-24T05:55:11.486691Z","shell.execute_reply.started":"2024-03-24T05:55:03.149481Z","shell.execute_reply":"2024-03-24T05:55:11.485419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nNO_MODELS = 5\n\n# use week numbers for group spliting\nweeks = base_train['WEEK_NUM']\n\ncv = StratifiedGroupKFold(n_splits = NO_MODELS, shuffle=False)\n\nensemble_models = {}\n\n# COMPUTE CV SCORE WITH 5 GROUP K FOLD\nfor i, (train_index, test_index) in enumerate(cv.split(X_train, y_train, groups=weeks)):\n    print('#'*25)\n    print('### Fold',i+1)\n    print('#'*25)\n    \n    # parameters were taken from this notebook:\n    # https://www.kaggle.com/code/greysky/home-credit-baseline\n    lgb_params = {\n        \"boosting_type\": \"gbdt\",\n        'objective' : 'binary',\n        'metric' : 'auc',\n        \"learning_rate\": 0.05,\n        \"n_estimators\": 1500,\n        \"colsample_bytree\": 0.8, \n        \"colsample_bynode\": 0.8,\n        \"verbose\": -1,\n        \"reg_alpha\": 0.1,\n        \"reg_lambda\": 10,\n        \"extra_trees\":True,\n        'num_leaves':64,\n        \"random_state\": SEED + i\n    }\n    \n            \n    # TRAIN DATA\n    train_x = X_train.loc[train_index]\n    train_y = y_train.loc[train_index]\n\n    # VALID DATA\n    valid_x = X_train.iloc[test_index]\n    valid_y = y_train.iloc[test_index]\n    \n    # Convert to DataSet\n    train_dataset = lgb.Dataset(train_x, label=train_y)\n    valid_dataset = lgb.Dataset(valid_x, label=valid_y, reference=train_dataset)\n\n    # TRAIN MODEL\n    clf = lgb.train(\n        lgb_params,\n        train_dataset,\n        valid_sets = valid_dataset,\n        callbacks = [lgb.log_evaluation(200), lgb.early_stopping(100)]\n    )\n\n    # SAVE MODEL\n    ensemble_models[f'model_{i}'] = clf\n        \n    print()","metadata":{"execution":{"iopub.status.busy":"2024-03-24T05:56:55.927792Z","iopub.execute_input":"2024-03-24T05:56:55.928233Z","iopub.status.idle":"2024-03-24T06:11:44.442490Z","shell.execute_reply.started":"2024-03-24T05:56:55.928204Z","shell.execute_reply":"2024-03-24T06:11:44.441208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf_preds = pd.DataFrame()\nfor i in range(NO_MODELS):\n    clf = ensemble_models[f'model_{i}'] \n    df_preds[f'preds_{i}'] = clf.predict(X_test, num_iteration=clf.best_iteration)\n    \n# add average predictions\nrandom_split_best_data['base_test']['pred'] = df_preds.mean(axis = 1)","metadata":{"execution":{"iopub.status.busy":"2024-03-24T06:11:44.444236Z","iopub.execute_input":"2024-03-24T06:11:44.444586Z","iopub.status.idle":"2024-03-24T06:13:00.363392Z","shell.execute_reply.started":"2024-03-24T06:11:44.444557Z","shell.execute_reply":"2024-03-24T06:13:00.361730Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"5\"></a>\n# <b>5 <span style='color:#53599A'>Model validation</span></b>\n\n<div>\n<br>\n<a href=\"#toc\" style=\"background-color: #607BB0; color: #ffffff; padding: 7px 10px; text-decoration: none; border-radius: 50px;\">Back to top</a><a id=\"toc\"></a>\n</div>","metadata":{}},{"cell_type":"code","source":"def validate_model(data, chart = True):\n    \"\"\"\n    Input:\n        preds, array like object\n        target, array like object\n        chart, boolean, if true return matplolib else table of summary statistics\n    Output:\n        pandas DataFrame or matplolib object\n        \n    Needs equal length prediction and taget column vectors\n    \"\"\"\n    # create DataFrame for plotting and summary stats\n    _df = data.copy()\n    _roc = data.loc[:, [\"WEEK_NUM\", \"target\", \"pred\"]]\\\n                   .sort_values(\"WEEK_NUM\").groupby(\"WEEK_NUM\")[[\"target\", \"pred\"]]\\\n                   .apply(lambda x: roc_auc_score(x[\"target\"], x[\"pred\"]))\n    _df['ROC_AUC'] = _df['WEEK_NUM'].map(_roc)\n    _df['GINI'] = 2 * _df['ROC_AUC'] - 1\n    \n    # stability score\n    x = _df[['WEEK_NUM', 'GINI']].drop_duplicates()['WEEK_NUM']\n    y = _df[['WEEK_NUM', 'GINI']].drop_duplicates()['GINI']\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = _df[['WEEK_NUM', 'GINI']].drop_duplicates()['GINI'].mean()\n    w_fallingrate = 88.0\n    w_resstd = -0.5 \n    # calculate competition score\n    score = avg_gini + w_fallingrate * min(0, a) + w_resstd * res_std\n    \n    if chart:\n        fig, ax = plt.subplots(2, 2, figsize=(16, 10))\n        \n        # temporal changes of gini scores\n        _df_plot = _df.groupby('WEEK_NUM')['GINI'].mean().reset_index()\n        ax[0, 0].plot(_df_plot['WEEK_NUM'], _df_plot['GINI'], \"o-\", label = 'Weekly GINI scores')\n        # add linear fit\n        ax[0, 0].plot(_df_plot['WEEK_NUM'], _df_plot['WEEK_NUM'] * a + b, \"r--\", label = 'linear fit')\n        ax[0, 0].set_ylabel(\"Gini score\")\n        ax[0, 0].set_xlabel(\"Week number\")\n        \n        # GINI score boxplot\n        sns.boxplot(data=_df_plot, x=\"GINI\", ax = ax[0, 1], color = COLORS[2])\n        ax[0, 1].set_xlabel(\"Gini score\")\n        \n        # GINI score histogram\n        ax[1, 1].hist(_df_plot['GINI'], bins=20, label=\"Weekly GINI scores\", color = COLORS[2])\n        ax[1, 1].set_ylabel(\"Count\")\n        ax[1, 1].set_xlabel(\"Gini score\")\n        \n        # GINI vs number of observations\n        _df_plot = _df.groupby('WEEK_NUM').agg({\"GINI\": ['mean', \"count\"]})\n        sns.regplot(x=_df_plot.iloc[:, 1], y=_df_plot.iloc[:, 0], ax=ax[1, 0], label=f'Corr. {_df_plot.corr().values[0, 1]:.3f}', color = COLORS[1])\n        ax[1, 0].set_ylabel(\"Weekly Gini score\")\n        ax[1, 0].set_xlabel(\"Number of applications\")\n        \n        # add white backgrounds to legends\n        for i in range(2):\n            for ii in range(2):\n                legend = ax[i, ii].legend(frameon=1)\n                frame = legend.get_frame()\n                frame.set_facecolor('w')\n        \n    # return DataFrame with summary stats\n    else:\n        _df_stats = pd.DataFrame({\"dataset\": [\"test\"]})\n        _df_stats['Avg. weekly ROC AUC'] = _df[['WEEK_NUM', 'ROC_AUC']].drop_duplicates()['ROC_AUC'].mean().round(3)\n        _df_stats['Std. of weekly ROC AUC'] = _df[['WEEK_NUM', 'ROC_AUC']].drop_duplicates()['ROC_AUC'].std().round(3)\n        _df_stats['Avg. weekly GINI'] = _df[['WEEK_NUM', 'GINI']].drop_duplicates()['GINI'].mean().round(3)\n        _df_stats['Std. of weekly GINI'] = _df[['WEEK_NUM', 'GINI']].drop_duplicates()['GINI'].std().round(3)\n        _df_stats['Stability score'] = round(score, 3)\n        \n        return _df_stats","metadata":{"execution":{"iopub.status.busy":"2024-03-24T06:13:00.365190Z","iopub.execute_input":"2024-03-24T06:13:00.365537Z","iopub.status.idle":"2024-03-24T06:13:00.394027Z","shell.execute_reply.started":"2024-03-24T06:13:00.365510Z","shell.execute_reply":"2024-03-24T06:13:00.392786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"_data = random_split_best_data['base_test'].copy()\nvalidate_model(_data)","metadata":{"execution":{"iopub.status.busy":"2024-03-24T06:13:00.396747Z","iopub.execute_input":"2024-03-24T06:13:00.397190Z","iopub.status.idle":"2024-03-24T06:13:01.934497Z","shell.execute_reply.started":"2024-03-24T06:13:00.397151Z","shell.execute_reply":"2024-03-24T06:13:01.933209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"validate_model(_data, False)","metadata":{"execution":{"iopub.status.busy":"2024-03-24T06:13:01.936035Z","iopub.execute_input":"2024-03-24T06:13:01.936443Z","iopub.status.idle":"2024-03-24T06:13:02.382669Z","shell.execute_reply.started":"2024-03-24T06:13:01.936409Z","shell.execute_reply":"2024-03-24T06:13:02.381457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"6\"></a>\n# <b>6 <span style='color:#53599A'>Submit predictions</span></b>\n\n<div>\n<br>\n<a href=\"#toc\" style=\"background-color: #607BB0; color: #ffffff; padding: 7px 10px; text-decoration: none; border-radius: 50px;\">Back to top</a><a id=\"toc\"></a>\n</div>","metadata":{}},{"cell_type":"code","source":"X_submission = data_submission[cols_pred].to_pandas()\nX_submission = convert_strings(X_submission)\ncategorical_cols = X_train.select_dtypes(include=['category']).columns\n\nfor col in categorical_cols:\n    train_categories = set(X_train[col].cat.categories)\n    submission_categories = set(X_submission[col].cat.categories)\n    new_categories = submission_categories - train_categories\n    X_submission.loc[X_submission[col].isin(new_categories), col] = \"Unknown\"\n    new_dtype = pd.CategoricalDtype(categories=train_categories, ordered=True)\n    X_train[col] = X_train[col].astype(new_dtype)\n    X_submission[col] = X_submission[col].astype(new_dtype)","metadata":{"execution":{"iopub.status.busy":"2024-03-23T20:02:42.871701Z","iopub.execute_input":"2024-03-23T20:02:42.872092Z","iopub.status.idle":"2024-03-23T20:02:42.993516Z","shell.execute_reply.started":"2024-03-23T20:02:42.872062Z","shell.execute_reply":"2024-03-23T20:02:42.992535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf_preds = pd.DataFrame()\nfor i in range(NO_MODELS):\n    clf = ensemble_models[f'model_{i}'] \n    df_preds[f'preds_{i}'] = clf.predict(X_submission, num_iteration=clf.best_iteration)\n    \n# average predictions\ny_submission_pred = df_preds.mean(axis = 1)","metadata":{"execution":{"iopub.status.busy":"2024-03-23T20:02:42.994952Z","iopub.execute_input":"2024-03-23T20:02:42.995382Z","iopub.status.idle":"2024-03-23T20:02:43.086369Z","shell.execute_reply.started":"2024-03-23T20:02:42.995343Z","shell.execute_reply":"2024-03-23T20:02:43.085262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame({\n    \"case_id\": data_submission[\"case_id\"].to_numpy(),\n    \"score\": y_submission_pred\n}).set_index('case_id')\nsubmission.to_csv(\"./submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-03-23T20:02:43.088010Z","iopub.execute_input":"2024-03-23T20:02:43.088529Z","iopub.status.idle":"2024-03-23T20:02:43.100596Z","shell.execute_reply.started":"2024-03-23T20:02:43.088489Z","shell.execute_reply":"2024-03-23T20:02:43.099369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"7\"></a>\n# <b>7 <span style='color:#53599A'>Changelog</span></b>\n\n* **v1** added basic data loading and processing code from https://www.kaggle.com/code/jetakow/home-credit-2024-starter-notebook notebook by @jetakow. Created general structure for the upcomming notebook.\n* **v2** added basic descriptions from TTC and PIT PD models. In addition, created visualizations for the provided historical default data. An example section was created that visually differeces between models trained on out-of-time and random samples.\n* **v3** remade a chart in [3.1 Historical default rate](#3.1) section. Sligthly re-fractored code to have one dictornary with all train/valid/test samples and trained models so it is easier to compare multiple approaches (work still in progress).\n* **v4** finished re-fractoring model code - now all data and trained models are saved into dictonary for easy comparison. Added new [3.2.3 Compare splits](#3.2.3) sub-section for comparing 2 models trained on different sample spliting using competition stability equation. As expected, model trained on random sample scored better on stability score.\n* **v5** compared two trained models using the gini stability score. These scores were also visualized using weekly and monthly time scales. I have also added references for previous TODO placeholders.\n* **v6** calculated stability scores on monthly data. Added fitted score lines on stability score (inspired by @eliocordeiropereira notebook: https://www.kaggle.com/code/eliocordeiropereira/the-starter-notebook-thoroughly-explained) and ROC charts in [3.2.3 Compare splits](#3.2.3) sub-section.\n* **v7** fixed small code typos with the last 2 charts in notebook version 6.\n* **v8** added submission using model which was trained on out-of-time sample.\n* **v9** added diagnostic charts and tables in the validation section and turned-off the internet so one can submit predictions from the notebook. On 2023-03-17 this notebook version scored 0.384 on public LB and placed 475/621 (top 76.49%).\n* **v10** added theoretical background to [1.2 Data limitations](#1.2) and [1.3 Estimating model's time horizon](#1.3) sub-sections.\n* **v11** tuned hyperparameters for `lgbm` model using `hyperopt` method as described in https://www.kaggle.com/code/pshikk/lgbm-hyperopt#Training-LightGBM notebook by @pshikk. Unfortunetly this did not result in good **LB** score.\n* **v12** I have decided to try ensembling method as described by @kimtaehun in notebook: https://www.kaggle.com/code/kimtaehun/lightgbm-baseline-with-aggregated-log-data using 3 models. On 2023-03-23 this notebook version scored 0.388 on public LB and placed 680/954 (top 71.28%).\n* **v13** fixed minor typos in the table of contents and increased the number of models in ensembling to 5. Public **LB** score was smaller for this notebook as compared to to previous versions.\n* **v14** for training ensemble model data was split by month, i.e. added `groups=months` to `StratifiedKFold()`.  Public **LB** score was similiar for this notebook **version 9**.\n* **v15** changed cross-validation method from `StratifiedKFold()` to `GroupKFold` and increased the number of cross-validation splits from 3 to 5. No significant improvement on LB score was observed.\n* **v16** used ` StratifiedGroupKFold` for cross-validation with weeks as group input. This idea was taken from notebook https://www.kaggle.com/code/daviddirethucus/home-credit-risk-lightgbm by @daviddirethucus.\n\n<a id=\"8\"></a>\n# <b><span style='color:#53599A'>Refencences</span></b>\n\n<a id=\"ref1\"><b>[1]</b></a> Basel Committee on Bank Supervision. (2000). Range of Practice in Banks? Internal Ratings Systems. Discussion paper January, Basel, Switzerland.\n<br>\n<a id=\"ref2\"><b>[2]</b></a> Supervisory handbook on the validation of rating systems under the internal ratings based approach, last update 10 August 2023.\n\n<div>\n<br>\n<a href=\"#toc\" style=\"background-color: #607BB0; color: #ffffff; padding: 7px 10px; text-decoration: none; border-radius: 50px;\">Back to top</a><a id=\"toc\"></a>\n</div>","metadata":{}}]}