{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":50160,"databundleVersionId":7921029,"sourceType":"competition"}],"dockerImageVersionId":30699,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Credit Risk","metadata":{}},{"cell_type":"markdown","source":"With the outbreak of the credit crisis in 2008, it became clear that risk, and in particular credit risk, had\nbeen largely underestimated. In response to this, new  requirements have been introduced that\nsignificantly strengthened bank solvency [1].\n\n\nThis competition aims to create a model to predict which clients are more likely to default on their loans, considering the temporal stability of the model. \n\n\nThe dataset used is made up of parquet files from different sources. Some files belong to level 0, others to levels 1 and 2. The files in the last 2 levels need aggregation techniques to feed the models. \n\n\nThe files contain social demographic data such as the client's birth, education, salary, and marital status. In addition, data regarding its credit history, such as the value of credits granted, their annuity, the interest rate, the number of payments to be made, the date when the credit was created, and the financial institution that issues the loan, among others. \n\nOther files include data regarding tax deductions, debit card transactions, and previous applications in Home Credit.  ","metadata":{}},{"cell_type":"code","source":"import os\nimport gc\nfrom glob import glob\nfrom pathlib import Path\nfrom datetime import datetime\n\nimport numpy as np\nimport pandas as pd\nimport polars as pl\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport xgboost as xgb\nimport lightgbm as lgb\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import roc_auc_score\nimport warnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)\npath='/kaggle/input/home-credit-credit-risk-model-stability/parquet_files/train/'\npath2='/kaggle/input/home-credit-credit-risk-model-stability/parquet_files/test/'","metadata":{"execution":{"iopub.status.busy":"2024-05-27T14:18:39.639351Z","iopub.execute_input":"2024-05-27T14:18:39.639995Z","iopub.status.idle":"2024-05-27T14:18:39.646201Z","shell.execute_reply.started":"2024-05-27T14:18:39.639966Z","shell.execute_reply":"2024-05-27T14:18:39.645270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Data Preprocessing**\n\nPart of the code used for reading files and preprocessing data was based on Notebooks [2] and [3]. In the data preprocessing step, a datatype was assigned to certain feature columns as some data was of different types. The datatypes assigned were integer, float, character, date, and categorical. Columns with more than 30% null data were removed, including categorical data columns with more than 200 types or that by mistake only had 1 single value recorded.\n\nThe code below is responsible for carrying out part of that procedure. The set_table_dtypes() method assigns the appropriate data types to each column of data. handle_dates() transforms date data into a format that can be used to train predictive models. filter_cols is the method used to filter the columns. The read_files() function is used to read and process the files of Credit Burea Provider A.\n\nto_pandas() function converts Polars DataFrames to Pandas,  one of the data formats that XGBoost and LightGBM models accept for training. Additionally, it transforms the corresponding columns into categorical data.\n","metadata":{}},{"cell_type":"code","source":"class Pipeline:\n    @staticmethod\n    def set_table_dtypes(df):\n        for col in df.columns:\n            if col in [\"case_id\", \"WEEK_NUM\", \"num_group1\", \"num_group2\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Int32))\n            elif col in [\"date_decision\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Date))\n            elif col[-1] in (\"P\", \"A\"):\n                df = df.with_columns(pl.col(col).cast(pl.Float64))\n            elif col[-1] in (\"M\",):\n                df = df.with_columns(pl.col(col).cast(pl.String))\n            elif col[-1] in (\"D\",):\n                df = df.with_columns(pl.col(col).cast(pl.Date))            \n\n        return df\n    \n    @staticmethod\n    def handle_dates(df):\n        for col in df.columns:\n            if col[-1] in (\"D\",):\n                df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))\n                df = df.with_columns(pl.col(col).dt.total_days())\n                df = df.with_columns(pl.col(col).cast(pl.Float32))\n                \n        df = df.drop(\"date_decision\", \"MONTH\")\n\n        return df\n    \n    @staticmethod\n    def filter_cols(df):\n        for col in df.columns:\n            if col not in [\"target\", \"case_id\", \"WEEK_NUM\"]:\n                isnull = df[col].is_null().mean()\n\n                if isnull > 0.3:\n                    df = df.drop(col)\n\n        for col in df.columns:\n            if (col not in [\"target\", \"case_id\", \"WEEK_NUM\"]) & (df[col].dtype == pl.String):\n                freq = df[col].n_unique()\n\n                if (freq == 1) | (freq > 200):\n                    df = df.drop(col)\n\n        return df\n    \n\ndef read_files(regex_path, depth=None):\n    chunks = []\n    for path in glob(str(regex_path)):\n        df = pl.read_parquet(path).filter(pl.col('num_group1')==0)\n        df = df.pipe(Pipeline.set_table_dtypes)    \n        chunks.append(df)\n        \n    df = pl.concat(chunks, how=\"vertical_relaxed\")\n    return df\n\ndef to_pandas(df_data, cat_cols=None):\n    df_data = df_data.to_pandas()\n    \n    if cat_cols is None:\n        cat_cols = list(df_data.select_dtypes(\"object\").columns)\n    \n    df_data[cat_cols] = df_data[cat_cols].astype(\"category\")\n    \n    return df_data, cat_cols\n","metadata":{"execution":{"iopub.status.busy":"2024-05-27T14:18:39.654757Z","iopub.execute_input":"2024-05-27T14:18:39.655049Z","iopub.status.idle":"2024-05-27T14:18:39.671350Z","shell.execute_reply.started":"2024-05-27T14:18:39.655024Z","shell.execute_reply":"2024-05-27T14:18:39.670500Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Feature Engineering**\n\nThis is one of the key procedures in this competition and with greater sensitivity in the metric used to evaluate the stability of the model.  If numerous features are used, it was observed that the model could increase its AUC score, but the stability metric decreased. Therefore, the techniques used to add features should be studied in detail.\n\nAside from the data features corresponding to the level 0 files, one of the features that boost the stability of my model was created from the tax deduction files. It was constructed by adding the total tax deduction of each client in the last year.   Another of the most important characteristics was the financial institution that issued the active credits to a specific client. This was obtained from two Credit Bureau providers (Credit bureau providers A, and B).\n\nThere was similar data recorded with different labels in different files,  when adding the base data that would feed the predictive model, redundancy was generated, which caused the stability metric to decrease. Therefore, some files were discarded, and features that generated an acceptable AUC score (greater than 80%) were selected, keeping the stability metric as high as possible.\n\nBelow is the code for reading different data file sources that feed the models, with their respective aggregation functions. The bureau_taxes variable is responsible for joining all the data, after being grouped according to the client whose credit was approved ('case_id').","metadata":{}},{"cell_type":"code","source":"train_basetable=pl.read_parquet(f'{path}train_base.parquet').pipe(Pipeline.set_table_dtypes).with_columns(\n            month_decision = pl.col(\"date_decision\").dt.month(),\n            weekday_decision = pl.col(\"date_decision\").dt.weekday(),\n        )\n\ntrain_static_0 = pl.concat([pl.read_parquet(f'{path}train_static_0_0.parquet').pipe(Pipeline.set_table_dtypes),\n                              pl.read_parquet(f'{path}train_static_0_1.parquet').pipe(Pipeline.set_table_dtypes),\n                             ],\n                             how=\"vertical_relaxed\").pipe(Pipeline.filter_cols)\ntrain_static_cb=pl.read_parquet(f'{path}train_static_cb_0.parquet').pipe(Pipeline.set_table_dtypes).pipe(Pipeline.filter_cols)\n\n\ntrain_person_1=pl.read_parquet(f'{path}train_person_1.parquet').filter(pl.col('num_group1')==0).pipe(Pipeline.set_table_dtypes).pipe(Pipeline.filter_cols)\ntrain_person1_feats=train_person_1.select('case_id','birth_259D','mainoccupationinc_384A').pipe(Pipeline.set_table_dtypes)\n\n\nbureau_a_1=read_files(f'{path}train_credit_bureau_a_1_*.parquet').select('case_id', 'financialinstitution_591M')\n\nbureau= pl.concat([bureau_a_1,\npl.read_parquet(f'{path}train_credit_bureau_b_1.parquet').pipe(Pipeline.set_table_dtypes)\\\n.rename({\"credor_3940957M\": \"financialinstitution_591M\"}).select('case_id', 'financialinstitution_591M'),\n], how=\"vertical_relaxed\")\n        \n\n\ntaxes=pl.concat([pl.read_parquet(f'{path}train_tax_registry_a_1.parquet').pipe(Pipeline.set_table_dtypes).group_by('case_id').agg(pl.col('amount_4527230A').sum().alias('amount_taxA'),pl.col('recorddate_4527225D').sort().last().alias('recorddate_taxD')),\npl.read_parquet(f'{path}train_tax_registry_b_1.parquet').pipe(Pipeline.set_table_dtypes).group_by('case_id').agg(pl.col('amount_4917619A').sum().alias('amount_taxA'),pl.col('deductiondate_4917603D').sort().last().alias('recorddate_taxD')),\npl.read_parquet(f'{path}train_tax_registry_c_1.parquet').pipe(Pipeline.set_table_dtypes).group_by('case_id').agg(pl.col('pmtamount_36A').sum().alias('amount_taxA'),pl.col('processingdate_168D').sort().last().alias('recorddate_taxD')),],\n                     how=\"vertical_relaxed\").pipe(Pipeline.filter_cols)\n\n\nbureau_taxes= train_basetable.join(train_static_0, how=\"left\", on=\"case_id\"\n).join(\n    train_static_cb, how=\"left\", on=\"case_id\"\n).join(\n    train_person1_feats, how=\"left\", on=\"case_id\"\n).join(\n   bureau, how=\"left\", on=\"case_id\"\n).join(\n   taxes, how=\"left\", on=\"case_id\"\n).pipe(Pipeline.handle_dates).drop('num_group1')\nbureau_taxes = bureau_taxes.unique(subset=[\"case_id\"])\ndf_train=bureau_taxes.drop('dateofbirth_337D')\n","metadata":{"execution":{"iopub.status.busy":"2024-05-27T14:18:39.684634Z","iopub.execute_input":"2024-05-27T14:18:39.685063Z","iopub.status.idle":"2024-05-27T14:19:16.960416Z","shell.execute_reply.started":"2024-05-27T14:18:39.685041Z","shell.execute_reply":"2024-05-27T14:19:16.959616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.shape[0]","metadata":{"execution":{"iopub.status.busy":"2024-05-27T14:19:16.961847Z","iopub.execute_input":"2024-05-27T14:19:16.962133Z","iopub.status.idle":"2024-05-27T14:19:16.968776Z","shell.execute_reply.started":"2024-05-27T14:19:16.962109Z","shell.execute_reply":"2024-05-27T14:19:16.967753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(df_train['target']==1).mean()*100","metadata":{"execution":{"iopub.status.busy":"2024-05-27T14:19:16.970131Z","iopub.execute_input":"2024-05-27T14:19:16.970469Z","iopub.status.idle":"2024-05-27T14:19:16.983002Z","shell.execute_reply.started":"2024-05-27T14:19:16.970440Z","shell.execute_reply":"2024-05-27T14:19:16.982037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are **1,526,659** clients with loans in Home Credit, of which **3.14 %** defaulted on their debt. It is a highly imbalanced dataset, the AUC score can be used to evaluate the model's performance in this kind of dataset.","metadata":{}},{"cell_type":"markdown","source":"The following code randomizes and splits the client's data into training, validation, and test data. 80% of the data is used as training data, in the remaining 20%, 10% is used as validation data, and the other 10% as test data.\n\nEach data set is separated into three matrices with the prefixes base_, X_, and y_. Taking the training data (base_train, X_train, y_train) as an example, the base_train matrix contains the 'WEEKS' column that will later be used to calculate the stability metric, X_train is the features matrix, and y_train the labels, it is  1 if the client defaulted on his debt and 0 otherwise.\n\n118 features will be used for training the model.","metadata":{}},{"cell_type":"code","source":"data=df_train\ncase_ids = data[\"case_id\"].unique().shuffle(seed=1)\ncase_ids_train, case_ids_test = train_test_split(case_ids, train_size=0.8, random_state=1)\ncase_ids_valid, case_ids_test = train_test_split(case_ids_test, train_size=0.5, random_state=1)\n\n\ncols_pred = [\"case_id\"]\nfor col in data.columns:\n    if col[-1].isupper() and col[:-1].islower():\n        cols_pred.append(col)\n\n\n\ndef from_polars_to_pandas(case_ids: pl.DataFrame) -> pl.DataFrame:\n    return (\n        data.filter(pl.col(\"case_id\").is_in(case_ids))[[\"case_id\", \"WEEK_NUM\", \"target\"]].to_pandas(),\n        data.filter(pl.col(\"case_id\").is_in(case_ids))[cols_pred],\n        data.filter(pl.col(\"case_id\").is_in(case_ids))[\"target\"].to_pandas()\n    )\n\nbase_train, X_train, y_train = from_polars_to_pandas(case_ids_train)\n\nX_train, cat_cols=to_pandas(X_train, cat_cols=None)\n\nbase_valid, X_valid, y_valid = from_polars_to_pandas(case_ids_valid)\n\nX_valid, cat_cols=to_pandas(X_valid, cat_cols=cat_cols)\nbase_test, X_test, y_test = from_polars_to_pandas(case_ids_test)\n\nX_test, cat_cols=to_pandas(X_test, cat_cols=cat_cols)\n\n\nX_train=X_train.set_index(\"case_id\")\nX_valid=X_valid.set_index(\"case_id\")   \nX_test=X_test.set_index(\"case_id\")","metadata":{"execution":{"iopub.status.busy":"2024-05-27T14:19:16.985560Z","iopub.execute_input":"2024-05-27T14:19:16.985921Z","iopub.status.idle":"2024-05-27T14:19:25.964580Z","shell.execute_reply.started":"2024-05-27T14:19:16.985892Z","shell.execute_reply":"2024-05-27T14:19:25.963741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(X_train.shape)\nprint(X_valid.shape)\nprint(X_test.shape)\n","metadata":{"execution":{"iopub.status.busy":"2024-05-27T14:19:25.965683Z","iopub.execute_input":"2024-05-27T14:19:25.965997Z","iopub.status.idle":"2024-05-27T14:19:25.971149Z","shell.execute_reply.started":"2024-05-27T14:19:25.965971Z","shell.execute_reply":"2024-05-27T14:19:25.970172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model\n\nIn the research process, it was observed that models designed with few characteristics could achieve high values in the AUC score. Furthermore, considering that this type of model is used by workers in the credit underwriting area, it should be simple and explainable. Based on this, only 118 features were used, the model employed was the LGBMClassifier with a max_depth=4 and 500 estimators. It was observed that by increasing the max_depth parameter, the AUC score increased but the stability metric decreased, at least, for the selected feature matrix.","metadata":{}},{"cell_type":"code","source":"params = {\n    \"boosting_type\": \"gbdt\",\n    \"objective\": \"binary\",\n    \"metric\": \"auc\",\n    \"max_depth\": 4,\n    \"num_leaves\": 31,\n    \"learning_rate\": 0.05,\n    \"feature_fraction\": 0.9,\n    \"bagging_fraction\": 0.8,\n    \"bagging_freq\": 5,\n    \"n_estimators\": 500,\n    \"verbose\": -1,\n    \"device\": \"gpu\"\n}","metadata":{"execution":{"iopub.status.busy":"2024-05-27T14:40:39.800716Z","iopub.execute_input":"2024-05-27T14:40:39.801396Z","iopub.status.idle":"2024-05-27T14:40:39.806390Z","shell.execute_reply.started":"2024-05-27T14:40:39.801364Z","shell.execute_reply":"2024-05-27T14:40:39.805446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lgb_model=lgb.LGBMClassifier(**params)\nlgb_model.fit(\n        X_train, y_train,\n        eval_set=[(X_valid, y_valid)],\n        callbacks=[lgb.log_evaluation(50), lgb.early_stopping(20)]\n    )","metadata":{"execution":{"iopub.status.busy":"2024-05-27T14:40:43.633130Z","iopub.execute_input":"2024-05-27T14:40:43.633966Z","iopub.status.idle":"2024-05-27T14:42:13.696902Z","shell.execute_reply.started":"2024-05-27T14:40:43.633933Z","shell.execute_reply":"2024-05-27T14:42:13.695963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Feature Importance**","metadata":{}},{"cell_type":"code","source":"fig1, ax1 = plt.subplots(figsize=(10, 15))\n#plt.title('Feature Importance', fontsize=20)\nplt.xlabel('xlabel', fontsize=15)\nplt.ylabel('ylabel', fontsize=15) \nlgb.plot_importance(lgb_model,max_num_features =20,ax=ax1,grid=False,title=None,xlabel='Importance')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-05-27T14:19:25.984633Z","iopub.status.idle":"2024-05-27T14:19:25.984991Z","shell.execute_reply.started":"2024-05-27T14:19:25.984808Z","shell.execute_reply":"2024-05-27T14:19:25.984828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Among the 10 most important features to predict the probability of default of a client, are the financial institutions that have approved a loan, the year of birth, the total amount of tax deduction, the value of the credit, the total number of loan payments made by the client, the annuity, the number of payments made with more than 10 days past due, the last tax deduction date, the number of income payments in last 9 months, and the number of people sharing the customer's phone.","metadata":{}},{"cell_type":"markdown","source":"**Stability Metric**\n\nThe following code defines the gini_stability() function which is used as the stability metric of the model. In addition, this stability value and the AUC score are calculated in the 10% of the data assigned as test data.","metadata":{}},{"cell_type":"code","source":"def gini_stability(base, w_fallingrate=88.0, w_resstd=-0.5):\n    gini_in_time = base.loc[:, [\"WEEK_NUM\", \"target\", \"score\"]]\\\n        .sort_values(\"WEEK_NUM\")\\\n        .groupby(\"WEEK_NUM\")[[\"target\", \"score\"]]\\\n        .apply(lambda x: 2*roc_auc_score(x[\"target\"], x[\"score\"])-1).tolist()\n    \n    x = np.arange(len(gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    return avg_gini + w_fallingrate * min(0, a) + w_resstd * res_std\n\n\ny_pred =lgb_model.predict_proba(X_test)[:,1]\nbase_test[\"score\"]=y_pred\n\nprint(f'The AUC score on the test set is: {roc_auc_score(base_test[\"target\"], base_test[\"score\"])}') \nstability_score_test = gini_stability(base_test)\n\nprint(f'The stability score on the test set is: {stability_score_test}')  \n","metadata":{"execution":{"iopub.status.busy":"2024-05-27T14:19:25.986089Z","iopub.status.idle":"2024-05-27T14:19:25.986443Z","shell.execute_reply.started":"2024-05-27T14:19:25.986283Z","shell.execute_reply":"2024-05-27T14:19:25.986297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gini_in_time = base_test.loc[:, [\"WEEK_NUM\", \"target\", \"score\"]]\\\n        .sort_values(\"WEEK_NUM\")\\\n        .groupby(\"WEEK_NUM\")[[\"target\", \"score\"]]\\\n        .apply(lambda x: 2*roc_auc_score(x[\"target\"], x[\"score\"])-1).tolist()\n    \nx = np.arange(len(gini_in_time))\ny = gini_in_time\n\n\na,b = np.polyfit(x, y, deg=1)\ny_hat = a*x + b\nresiduals = y - y_hat\nres_std = np.std(residuals)\navg_gini = np.mean(gini_in_time)\n\nfig2, ax2 = plt.subplots(figsize=(6, 4))\nplt.title('Gini Score in Time', fontsize=12)\nplt.xlabel('Week Num', fontsize=10)\nplt.ylabel('Gini Score', fontsize=10) \nax2.scatter(x, y, s=60, alpha=0.7, edgecolors=\"k\")\n\nxseq = np.linspace(0, 100, num=100)\nax2.plot(xseq,a * xseq+b, color=\"k\", lw=2.5)","metadata":{"execution":{"iopub.status.busy":"2024-05-27T14:19:25.987404Z","iopub.status.idle":"2024-05-27T14:19:25.987715Z","shell.execute_reply.started":"2024-05-27T14:19:25.987562Z","shell.execute_reply":"2024-05-27T14:19:25.987575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the graph above, the Y axis represents the Gini score during weekly time intervals. Although in week 66 an abrupt drop is observed in the Gini score, being less than 0.3, this drop cannot be attributed to a decrease in the model's predictive capacity over time, since Gini values score greater than 0.8 are obtained in week 80.","metadata":{}},{"cell_type":"markdown","source":" ","metadata":{}},{"cell_type":"markdown","source":"**Reading submission data files**","metadata":{}},{"cell_type":"code","source":"test_bureau_a_1=read_files(f'{path2}test_credit_bureau_a_1_*.parquet').select('case_id', 'financialinstitution_591M')\n\ntest_bureau= pl.concat([test_bureau_a_1,\npl.read_parquet(f'{path2}test_credit_bureau_b_1.parquet').pipe(Pipeline.set_table_dtypes)\\\n.rename({\"credor_3940957M\": \"financialinstitution_591M\"}).select('case_id', 'financialinstitution_591M'),\n], how=\"vertical_relaxed\")\n\n","metadata":{"execution":{"iopub.status.busy":"2024-05-27T14:19:25.988890Z","iopub.status.idle":"2024-05-27T14:19:25.989266Z","shell.execute_reply.started":"2024-05-27T14:19:25.989087Z","shell.execute_reply":"2024-05-27T14:19:25.989101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_basetable=pl.read_parquet(f'{path2}test_base.parquet').pipe(Pipeline.set_table_dtypes).with_columns(\n            month_decision = pl.col(\"date_decision\").dt.month(),\n            weekday_decision = pl.col(\"date_decision\").dt.weekday(),\n        )\n\ntest_static_0 = pl.concat([pl.read_parquet(f'{path2}test_static_0_0.parquet').pipe(Pipeline.set_table_dtypes),\n                              pl.read_parquet(f'{path2}test_static_0_1.parquet').pipe(Pipeline.set_table_dtypes),\n                             pl.read_parquet(f'{path2}test_static_0_2.parquet').pipe(Pipeline.set_table_dtypes),],\n                             how=\"vertical_relaxed\")\n\n\ntest_static_cb=pl.read_parquet(f'{path2}test_static_cb_0.parquet').pipe(Pipeline.set_table_dtypes)\n\n\n\ntest_person_1=pl.read_parquet(f'{path2}test_person_1.parquet').filter(pl.col('num_group1')==0).pipe(Pipeline.set_table_dtypes)\ntest_person1_feats=test_person_1.select('case_id','birth_259D','mainoccupationinc_384A').pipe(Pipeline.set_table_dtypes)\n\n\ntest_taxes=pl.concat([pl.read_parquet(f'{path2}test_tax_registry_a_1.parquet').pipe(Pipeline.set_table_dtypes).group_by('case_id').agg(pl.col('amount_4527230A').sum().alias('amount_taxA'),pl.col('recorddate_4527225D').sort().last().alias('recorddate_taxD')),\npl.read_parquet(f'{path2}test_tax_registry_b_1.parquet').pipe(Pipeline.set_table_dtypes).group_by('case_id').agg(pl.col('amount_4917619A').sum().alias('amount_taxA'),pl.col('deductiondate_4917603D').sort().last().alias('recorddate_taxD')),\npl.read_parquet(f'{path2}test_tax_registry_c_1.parquet').pipe(Pipeline.set_table_dtypes).group_by('case_id').agg(pl.col('pmtamount_36A').sum().alias('amount_taxA'),pl.col('processingdate_168D').sort().last().alias('recorddate_taxD')),],\n                     how=\"vertical_relaxed\")\n","metadata":{"execution":{"iopub.status.busy":"2024-05-27T14:19:25.990731Z","iopub.status.idle":"2024-05-27T14:19:25.991168Z","shell.execute_reply.started":"2024-05-27T14:19:25.990933Z","shell.execute_reply":"2024-05-27T14:19:25.990951Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_bureau_taxes= test_basetable.join(test_static_0, how=\"left\", on=\"case_id\"\n).join(\n    test_static_cb, how=\"left\", on=\"case_id\"\n).join(\n    test_person1_feats, how=\"left\", on=\"case_id\"\n).join(\n   test_bureau, how=\"left\", on=\"case_id\"\n).join(\n   test_taxes, how=\"left\", on=\"case_id\"\n).pipe(Pipeline.handle_dates).drop('num_group1')\ntest_bureau_taxes = test_bureau_taxes.unique(subset=[\"case_id\"])\ndf_test=test_bureau_taxes.drop('dateofbirth_337D')","metadata":{"execution":{"iopub.status.busy":"2024-05-27T14:19:25.992791Z","iopub.status.idle":"2024-05-27T14:19:25.993165Z","shell.execute_reply.started":"2024-05-27T14:19:25.992983Z","shell.execute_reply":"2024-05-27T14:19:25.992999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_base,test_X = df_test.select('case_id','WEEK_NUM'), df_test[cols_pred]\n\ntest_X, cat_cols=to_pandas(test_X, cat_cols=cat_cols)\ntest_base=test_base.to_pandas()\ntest_X=test_X.set_index(\"case_id\")\n","metadata":{"execution":{"iopub.status.busy":"2024-05-27T14:19:25.994280Z","iopub.status.idle":"2024-05-27T14:19:25.994578Z","shell.execute_reply.started":"2024-05-27T14:19:25.994429Z","shell.execute_reply":"2024-05-27T14:19:25.994442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred = pd.Series(lgb_model.predict_proba(test_X)[:,1], index=test_X.index)\ndf_subm = pd.read_csv('/kaggle/input/home-credit-credit-risk-model-stability/sample_submission.csv')\ndf_subm = df_subm.set_index(\"case_id\")\n\ndf_subm[\"score\"] = y_pred\ndf_subm.to_csv(\"submission.csv\")\ndf_subm","metadata":{"execution":{"iopub.status.busy":"2024-05-27T14:19:25.995946Z","iopub.status.idle":"2024-05-27T14:19:25.996263Z","shell.execute_reply.started":"2024-05-27T14:19:25.996092Z","shell.execute_reply":"2024-05-27T14:19:25.996104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# References\n\n[1]. B. Baesens, K. Smedts,\"Boosting Credit Risk Models\", Baesens,British Accounting Review,Forthcoming, 2023.\n\n**Kaggle Notebooks**\n\n[2]. https://www.kaggle.com/code/jetakow/home-credit-2024-starter-notebook\n\n[3]. https://www.kaggle.com/code/greysky/home-credit-baseline\n","metadata":{}}]}