{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":50160,"databundleVersionId":7921029,"sourceType":"competition"}],"dockerImageVersionId":30646,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Developing a Stable and Reliable Loan Default Prediction Model for Enhanced Financial Inclusion","metadata":{}},{"cell_type":"markdown","source":"# Rational\n\nConsumer finance providers face significant challenges in accurately predicting loan default risk, especially for clients with little or no credit history. Traditional methods often fail to address the dynamic nature of client behavior, leading to unstable and frequently updated scorecards. A stable and reliable model is essential for making informed lending decisions that balance risk and accessibility.\n\n# Introduction\n\nThis notebook aims to develop a predictive model to determine which clients are more likely to default on their loans. The goal is to provide a stable and reliable solution that remains effective over time, thus aiding consumer finance providers in making better lending decisions. This project is part of a competition hosted by Home Credit, an international consumer finance provider known for responsible lending practices and financial inclusion efforts.\n\nThe competition runs from February 5, 2024, to May 28, 2024, and focuses on using data science to improve the prediction of loan default risk. \n\nA key evaluation criterion is the gini stability metric, which measures the model's predictive performance and stability over time.\n\n# Objective\n\nThe primary objective of this project is to build a predictive model that accurately assesses the likelihood of loan default while maintaining stability over time. The model should:\n\n1. Utilize a robust data preprocessing and feature engineering pipeline.\n2. Employ advanced machine learning techniques to maximize predictive accuracy.\n3. Ensure stability in predictions across different time periods to minimize the need for frequent model updates.\n\n# Deliverables\n\n## Data Preprocessing and Feature Engineering:\n\n**1. Read and concatenate multiple datasets.**\n\n    * Set appropriate data types for columns.\n    * Handle date features and filter out irrelevant or low-quality columns.\n    * Engineer additional features to enrich the dataset.\n\n**2. Model Training and Evaluation:**\n\n    * Implement a cross-validation strategy using StratifiedGroupKFold to ensure stable performance evaluation.\n    * Train a LightGBM classifier and utilize early stopping and logging callbacks.\n    * Aggregate predictions from multiple cross-validation folds using a custom VotingModel.\n\n**3. Prediction and Submission:**\n\n    * Prepare the test dataset by setting the correct data types and indices.\n    * Use the trained model to predict probabilities for the test set.\n    * Create a submission file with case_id and predicted scores.\n\n**4. Model Stability Assessment:**\n\n    * Evaluate the model using the gini stability metric.\n    * Ensure the model's predictions are stable over different weeks to avoid performance drop-offs.\n\n\nBy achieving these deliverables, the notebook will contribute to developing a reliable and stable predictive model that can help consumer finance providers make better lending decisions and potentially improve financial inclusion for individuals with limited credit history.","metadata":{}},{"cell_type":"markdown","source":"# Import Libraries","metadata":{}},{"cell_type":"code","source":"import warnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)\nimport os\nimport gc\nimport numpy as np\nimport pandas as pd\nimport polars as pl\nprint(pl.__version__)\nfrom glob import glob\nfrom pathlib import Path\nfrom datetime import datetime\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn.model_selection import TimeSeriesSplit, GroupKFold, StratifiedGroupKFold\nfrom sklearn.base import BaseEstimator, RegressorMixin\nfrom sklearn.metrics import roc_auc_score\nimport lightgbm as lgb\n\n","metadata":{"execution":{"iopub.status.busy":"2024-07-15T16:56:06.347285Z","iopub.execute_input":"2024-07-15T16:56:06.347749Z","iopub.status.idle":"2024-07-15T16:56:11.393353Z","shell.execute_reply.started":"2024-07-15T16:56:06.347715Z","shell.execute_reply":"2024-07-15T16:56:11.392138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ROOT            = Path(\"/kaggle/input/home-credit-credit-risk-model-stability\")\n# ROOT            = Path(\"./input\")\n\nTRAIN_DIR       = ROOT / \"csv_files\" / \"train\"\nTEST_DIR        = ROOT / \"csv_files\" / \"test\"","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.679609Z","iopub.execute_input":"2024-06-13T08:38:44.680037Z","iopub.status.idle":"2024-06-13T08:38:44.687629Z","shell.execute_reply.started":"2024-06-13T08:38:44.679998Z","shell.execute_reply":"2024-06-13T08:38:44.686663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Pre-Processing","metadata":{}},{"cell_type":"markdown","source":"# Pipeline Class for Changing the Data Types","metadata":{}},{"cell_type":"code","source":"import polars as pl\nfrom glob import glob\n\nclass Pipeline:\n    \n    '''This decorator indicates that the method does not depend on the instance of the class and can be \n    called on the class itself.\n    The set_table_dtypes method is a static method designed to standardize the data types of columns \n    in a Polars DataFrame (df). This method ensures that each column is cast to an appropriate data type \n    based on specific rules and conventions.'''\n    \n    @staticmethod \n    def set_table_dtypes(df):\n        for col in df.columns:\n            if col in [\"case_id\", \"WEEK_NUM\", \"num_group1\", \"num_group2\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Int64))\n            elif col in [\"date_decision\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Date))\n            elif col[-1] in (\"P\", \"A\"):\n                df = df.with_columns(pl.col(col).cast(pl.Float64))\n            elif col[-1] in (\"M\",):\n                df = df.with_columns(pl.col(col).cast(pl.String))\n            elif col[-1] in (\"D\",):\n                df = df.with_columns(pl.col(col).cast(pl.Date))\n        return df\n    \n    '''The handle_dates function is designed to process date columns in a Polars DataFrame (df). It \n    modifies columns whose names end with \"D\" by calculating the number of days between the dates in \n    these columns and a reference date column (date_decision).'''\n       \n    @staticmethod\n    def handle_dates(df):\n        for col in df.columns:\n            if col[-1] in (\"D\",):\n                df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))\n                df = df.with_columns(pl.col(col).dt.total_days())\n        df = df.drop(\"date_decision\", \"MONTH\")\n        return df\n\n    '''The filter_cols function is designed to clean and filter the columns of a Polars DataFrame (df). \n    It performs two main tasks: removing columns with a high proportion of missing values and removing \n    categorical columns with low or excessively high cardinality'''\n    \n    @staticmethod\n    def filter_cols(df):\n        for col in df.columns:\n            if col not in [\"target\", \"case_id\", \"WEEK_NUM\"]:\n                isnull = df[col].is_null().mean()\n\n                if isnull > 0.95:\n                    df = df.drop(col)\n\n        for col in df.columns:\n            if (col not in [\"target\", \"case_id\", \"WEEK_NUM\"]) & (df[col].dtype == pl.String):\n                freq = df[col].n_unique()\n\n                if (freq == 1) | (freq > 200):\n                    df = df.drop(col)\n\n        return df\n","metadata":{"execution":{"iopub.status.busy":"2024-07-16T10:17:07.004073Z","iopub.execute_input":"2024-07-16T10:17:07.004611Z","iopub.status.idle":"2024-07-16T10:17:07.350680Z","shell.execute_reply.started":"2024-07-16T10:17:07.004578Z","shell.execute_reply":"2024-07-16T10:17:07.349320Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Aggregator Class to summerise/transform the data by grouping and aggregating the values","metadata":{}},{"cell_type":"code","source":"class Aggregator:\n    \n    '''Generates aggregation expressions for numeric columns (identified by suffixes \"P\" or \"A\"). \n    It creates expressions to find the maximum value of these columns.'''\n    \n    @staticmethod\n    def num_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"P\", \"A\")]\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        return expr_max\n\n    '''Generates aggregation expressions for date columns (identified by suffix \"D\"). \n    It creates expressions to find the maximum value of these columns.'''\n    \n    @staticmethod\n    def date_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"D\",)]\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        return expr_max\n\n    '''Generates aggregation expressions for string columns (identified by suffix \"M\"). \n    It creates expressions to find the maximum value of these columns.'''\n    \n    @staticmethod\n    def str_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"M\",)]\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        return expr_max\n\n    '''Generates aggregation expressions for other types of columns (identified by suffixes \"T\" or \"L\"). \n    It creates expressions to find the maximum value of these columns.'''\n    \n    @staticmethod\n    def other_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"T\", \"L\")]\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        return expr_max\n\n    '''Generates aggregation expressions for columns related to grouping \n    (columns containing \"num_group\"). It creates expressions to find the maximum value of these columns.'''\n    \n    @staticmethod\n    def count_expr(df):\n        cols = [col for col in df.columns if \"num_group\" in col]\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        return expr_max\n\n    '''Combines all the aggregation expressions from the other methods into a single list of expressions.'''\n    \n    @staticmethod\n    def get_exprs(df):\n        exprs = Aggregator.num_expr(df) + \\\n                Aggregator.date_expr(df) + \\\n                Aggregator.str_expr(df) + \\\n                Aggregator.other_expr(df) + \\\n                Aggregator.count_expr(df)\n        return exprs\n","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.710225Z","iopub.status.idle":"2024-06-13T08:38:44.710707Z","shell.execute_reply.started":"2024-06-13T08:38:44.710434Z","shell.execute_reply":"2024-06-13T08:38:44.710451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# read_file method to read single .csv file","metadata":{}},{"cell_type":"code","source":"def read_file(path, depth=None):\n    df = pl.read_csv(path)\n    df = df.pipe(Pipeline.set_table_dtypes)\n\n    if depth in [1, 2]:\n        df = df.group_by(\"case_id\").agg(Aggregator.get_exprs(df))\n    return df","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.712624Z","iopub.status.idle":"2024-06-13T08:38:44.713146Z","shell.execute_reply.started":"2024-06-13T08:38:44.712885Z","shell.execute_reply":"2024-06-13T08:38:44.712907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# read_files method to read multiple .csv files","metadata":{}},{"cell_type":"code","source":"def read_files(regex_path, depth=None):\n    chunks = []\n    for path in glob(str(regex_path)):\n        chunks.append(pl.read_csv(path).pipe(Pipeline.set_table_dtypes))\n    df = pl.concat(chunks, how=\"vertical_relaxed\")\n    if depth in [1, 2]:\n        df = df.group_by(\"case_id\").agg(Aggregator.get_exprs(df))\n    return df","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.714657Z","iopub.status.idle":"2024-06-13T08:38:44.715183Z","shell.execute_reply.started":"2024-06-13T08:38:44.714923Z","shell.execute_reply":"2024-06-13T08:38:44.714946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature Engineering","metadata":{}},{"cell_type":"markdown","source":"### Feature Enggineering methods for creating new features and merging dataframes","metadata":{}},{"cell_type":"code","source":"'''This feature_eng function is designed to perform feature engineering on a base DataFrame (df_base) along with additional DataFrames (depth_0, depth_1, and depth_2). Here's a breakdown of what it does:\n\nDate Features Addition:\n\nIt adds two new columns to df_base: month_decision and weekday_decision.\nThese columns are derived from the date_decision column, capturing the month and weekday information.\n\nJoining Additional DataFrames:\n\nIt iterates over the lists depth_0, depth_1, and depth_2, which contain additional DataFrames to be joined with df_base.\nFor each DataFrame in these lists, it performs a left join with df_base on the case_id column.\nIt adds a suffix to the column names of the joined DataFrames to differentiate them from existing columns in df_base.\n\nDate Handling:\n\nAfter all joins are performed, it applies the Pipeline.handle_dates method static method defined above\nto handle date-related columns in the resulting DataFrame.\n\nReturn Resulting DataFrame:\n\nThe resulting DataFrame, after all feature engineering steps, is returned.'''\n\n\ndef feature_eng(df_base, depth_0, depth_1, depth_2):\n    df_base = (\n        df_base\n        .with_columns(\n            month_decision = pl.col(\"date_decision\").dt.month(),\n            weekday_decision = pl.col(\"date_decision\").dt.weekday(),\n        )\n    )\n    for i, df in enumerate(depth_0 + depth_1 + depth_2):\n        df_base = df_base.join(df, how=\"left\", on=\"case_id\", suffix=f\"_{i}\")\n    df_base = df_base.pipe(Pipeline.handle_dates)\n    return df_base\n\n\n''' to_pandas method converts Polars Dataframe to Pandas Dataframe.It identifies columns of type object as \ncategorical if cat_cols is not provided. It converts the identified categorical columns to the \ncategory data type in Pandas.Finally, it returns the converted DataFrame and the list of categorical columns.'''\n\ndef to_pandas(df_data, cat_cols=None):\n    df_data = df_data.to_pandas()\n    if cat_cols is None:\n        cat_cols = list(df_data.select_dtypes(\"object\").columns)\n    df_data[cat_cols] = df_data[cat_cols].astype(\"category\")\n    return df_data, cat_cols","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.716545Z","iopub.status.idle":"2024-06-13T08:38:44.717027Z","shell.execute_reply.started":"2024-06-13T08:38:44.716805Z","shell.execute_reply":"2024-06-13T08:38:44.716824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Organizing and Loading Multiple Train and Test Datasets into data_store","metadata":{}},{"cell_type":"code","source":"data_store = {\n    \"df_base\": read_file(TRAIN_DIR / \"train_base.csv\"),\n    \"depth_0\": [\n        read_file(TRAIN_DIR / \"train_static_cb_0.csv\"),\n        read_files(TRAIN_DIR / \"train_static_0_*.csv\"),\n    ],\n    \"depth_1\": [\n        read_files(TRAIN_DIR / \"train_applprev_1_*.csv\", 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_a_1.csv\", 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_b_1.csv\", 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_c_1.csv\", 1),\n        read_file(TRAIN_DIR / \"train_credit_bureau_b_1.csv\", 1),\n        read_file(TRAIN_DIR / \"train_other_1.csv\", 1),\n        read_file(TRAIN_DIR / \"train_person_1.csv\", 1),\n        read_file(TRAIN_DIR / \"train_deposit_1.csv\", 1),\n        read_file(TRAIN_DIR / \"train_debitcard_1.csv\", 1),\n    ],\n    \"depth_2\": [\n        read_file(TRAIN_DIR / \"train_credit_bureau_b_2.csv\", 2),\n    ]\n}","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.718201Z","iopub.status.idle":"2024-06-13T08:38:44.718569Z","shell.execute_reply.started":"2024-06-13T08:38:44.718383Z","shell.execute_reply":"2024-06-13T08:38:44.718398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = feature_eng(**data_store) # Unpacking the contents of the dictionary using ** Operator\nprint(\"train data shape:\\t\", df_train.shape)","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.720235Z","iopub.status.idle":"2024-06-13T08:38:44.720620Z","shell.execute_reply.started":"2024-06-13T08:38:44.720418Z","shell.execute_reply":"2024-06-13T08:38:44.720433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_store = {\n    \"df_base\": read_file(TEST_DIR / \"test_base.csv\"),\n    \"depth_0\": [\n        read_file(TEST_DIR / \"test_static_cb_0.csv\"),\n        read_files(TEST_DIR / \"test_static_0_*.csv\"),\n    ],\n    \"depth_1\": [\n        read_files(TEST_DIR / \"test_applprev_1_*.csv\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_a_1.csv\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_b_1.csv\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_c_1.csv\", 1),\n        read_file(TEST_DIR / \"test_credit_bureau_b_1.csv\", 1),\n        read_file(TEST_DIR / \"test_other_1.csv\", 1),\n        read_file(TEST_DIR / \"test_person_1.csv\", 1),\n        read_file(TEST_DIR / \"test_deposit_1.csv\", 1),\n        read_file(TEST_DIR / \"test_debitcard_1.csv\", 1),\n    ],\n    \"depth_2\": [\n        read_file(TEST_DIR / \"test_credit_bureau_b_2.csv\", 2),\n    ]\n}","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.721823Z","iopub.status.idle":"2024-06-13T08:38:44.722214Z","shell.execute_reply.started":"2024-06-13T08:38:44.722028Z","shell.execute_reply":"2024-06-13T08:38:44.722045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test = feature_eng(**data_store) # Unpacking the contents of the dictionary using ** Operator\nprint(\"test data shape:\\t\", df_test.shape)","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.723267Z","iopub.status.idle":"2024-06-13T08:38:44.723708Z","shell.execute_reply.started":"2024-06-13T08:38:44.723472Z","shell.execute_reply":"2024-06-13T08:38:44.723489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train and Test Dataset","metadata":{}},{"cell_type":"code","source":"df_train = df_train.pipe(Pipeline.filter_cols)\ndf_test = df_test.select([col for col in df_train.columns if col != \"target\"])\n\nprint(\"train data shape:\\t\", df_train.shape)\nprint(\"test data shape:\\t\", df_test.shape)","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.725398Z","iopub.status.idle":"2024-06-13T08:38:44.725954Z","shell.execute_reply.started":"2024-06-13T08:38:44.725680Z","shell.execute_reply":"2024-06-13T08:38:44.725703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train, cat_cols = to_pandas(df_train)\ndf_test, cat_cols = to_pandas(df_test, cat_cols)","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.727767Z","iopub.status.idle":"2024-06-13T08:38:44.728181Z","shell.execute_reply.started":"2024-06-13T08:38:44.727994Z","shell.execute_reply":"2024-06-13T08:38:44.728012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## del data_store\n\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2024-05-20T20:47:09.358517Z","iopub.execute_input":"2024-05-20T20:47:09.358919Z","iopub.status.idle":"2024-05-20T20:47:09.569362Z","shell.execute_reply.started":"2024-05-20T20:47:09.358889Z","shell.execute_reply":"2024-05-20T20:47:09.568057Z"}}},{"cell_type":"code","source":"class VotingModel(BaseEstimator, RegressorMixin):\n    def __init__(self, estimators):\n        super().__init__()\n        self.estimators = estimators\n\n    def fit(self, X, y=None):\n        return self\n\n    def predict(self, X):\n        y_preds = [estimator.predict(X) for estimator in self.estimators]\n        return np.mean(y_preds, axis=0)\n\n    def predict_proba(self, X):\n        y_preds = [estimator.predict_proba(X) for estimator in self.estimators]\n        return np.mean(y_preds, axis=0)","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.729569Z","iopub.status.idle":"2024-06-13T08:38:44.729997Z","shell.execute_reply.started":"2024-06-13T08:38:44.729802Z","shell.execute_reply":"2024-06-13T08:38:44.729818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = df_train.drop(columns=[\"target\", \"case_id\",\"WEEK_NUM\"])\ny = df_train[\"target\"]\nweeks = df_train[\"WEEK_NUM\"]","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.732766Z","iopub.status.idle":"2024-06-13T08:38:44.733197Z","shell.execute_reply.started":"2024-06-13T08:38:44.732971Z","shell.execute_reply":"2024-06-13T08:38:44.732993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''import optuna\nfrom sklearn.model_selection import cross_validate\nfrom lightgbm import LGBMClassifier\n\ndef objective(trial):\n    max_depth = trial.suggest_int('max_depth', 3, 30)\n    n_estimators = trial.suggest_int('n_estimators', 1, 1000)\n    gamma = trial.suggest_float('gamma', 0, 1)\n    reg_alpha = trial.suggest_float('reg_alpha', 0, 1)\n    reg_lambda = trial.suggest_float('reg_lambda', 0, 1)\n    min_child_weight = trial.suggest_int('min_child_weight', 0, 10)\n    subsample = trial.suggest_float('subsample', 0, 1)\n    colsample_bytree = trial.suggest_float('colsample_bytree', 0, 1)\n    learning_rate = trial.suggest_float('learning_rate', 0, 1)\n    \n#     print('Training the model with', X.shape[1], 'features')\n    \n#       LightGBM\n    params = {'learning_rate': learning_rate,\n              'n_estimators': n_estimators,\n              'max_depth': max_depth,\n              'lambda_l1': reg_alpha,\n              'lambda_l2': reg_lambda,\n              'colsample_bytree': colsample_bytree, \n              'subsample': subsample,    \n              'min_child_samples': min_child_weight,\n              'class_weight': 'balanced'}\n    \n    clf = LGBMClassifier(**params, verbose = -1, verbosity = -1)\n    \n    cv_results = cross_validate(clf,X,y, cv=5, scoring='accuracy')\n    \n    validation_score = np.mean(cv_results['test_score'])\n    \n    return validation_score'''","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.734300Z","iopub.status.idle":"2024-06-13T08:38:44.734715Z","shell.execute_reply.started":"2024-06-13T08:38:44.734490Z","shell.execute_reply":"2024-06-13T08:38:44.734506Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''study = optuna.create_study(direction=\"maximize\")\nstudy.optimize(objective, n_trials= 2)'''","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.735886Z","iopub.status.idle":"2024-06-13T08:38:44.736256Z","shell.execute_reply.started":"2024-06-13T08:38:44.736074Z","shell.execute_reply":"2024-06-13T08:38:44.736090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Cross Validation to assess the performance of an LightGBM classifier\n\nThis following code snippet is performing cross-validation using the StratifiedGroupKFold technique. Here's a breakdown of what it's doing:\n\n1. **StratifiedGroupKFold(n_splits=5, shuffle=False):** This initializes a StratifiedGroupKFold object with 5 splits and shuffle set to False, ensuring that each fold maintains the same class distribution and group membership.\n\n2. **fitted_models = [] and cv_scores = []:** These lists will store the trained models and the corresponding cross-validation scores, respectively.\n\n3. The loop iterates over each fold generated by the cross-validator **(cv.split(X, y, groups=weeks)). Within each iteration:**\n\n    * It splits the data into training and validation sets **(X_train, y_train, X_valid, y_valid)**     based on the current fold indices.\n    * Trains an LightGBM classifier **(lgb.LGBMClassifier())** on the training data.\n    * Evaluates the model on the validation set using the AUC score.\n    * Appends the trained model to *fitted_models* and the AUC score to **cv_scores.**\n\n\n4. Finally, it creates a VotingModel instance with the fitted models and prints the cross-validation AUC scores.","metadata":{}},{"cell_type":"code","source":"cv = StratifiedGroupKFold(n_splits=5, shuffle=False)\n\n\nfitted_models = []\ncv_scores = []\n\n\nfor idx_train, idx_valid in cv.split(X, y, groups=weeks):\n    X_train, y_train = X.iloc[idx_train], y.iloc[idx_train]\n    X_valid, y_valid = X.iloc[idx_valid], y.iloc[idx_valid]\n\n    print(\"Valid week range: \", (weeks.iloc[idx_valid].min(), weeks.iloc[idx_valid].max()))\n\n    model = lgb.LGBMClassifier()\n    model.fit(\n        X_train, y_train,\n        eval_set=[(X_valid, y_valid)],\n        callbacks=[lgb.log_evaluation(50), lgb.early_stopping(50)]\n    )\n\n    fitted_models.append(model)\n\n\n    y_pred_valid = model.predict_proba(X_valid)[:, 1]\n    auc_score = roc_auc_score(y_valid, y_pred_valid)\n    cv_scores.append(auc_score)\n\nmodel = VotingModel(fitted_models)\nprint(\"CV AUC scores: \", cv_scores)","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.737741Z","iopub.status.idle":"2024-06-13T08:38:44.738105Z","shell.execute_reply.started":"2024-06-13T08:38:44.737927Z","shell.execute_reply":"2024-06-13T08:38:44.737942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test = df_test.drop(columns=[\"WEEK_NUM\"])\nX_test = X_test.set_index(\"case_id\")\n\nX_test[['pmtcount_693L', 'pmtscount_423L', 'deferredmnthsnum_166L', 'max_credacc_transactions_402L']] = X_test[['pmtcount_693L', 'pmtscount_423L', 'deferredmnthsnum_166L', 'max_credacc_transactions_402L']].astype(float)\n\nlgb_pred = pd.Series(model.predict_proba(X_test)[:, 1], index=X_test.index)","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.739333Z","iopub.status.idle":"2024-06-13T08:38:44.739729Z","shell.execute_reply.started":"2024-06-13T08:38:44.739507Z","shell.execute_reply":"2024-06-13T08:38:44.739521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_subm = pd.read_csv(ROOT / \"sample_submission.csv\")\ndf_subm = df_subm.set_index(\"case_id\")\n\ndf_subm[\"score\"] = lgb_pred","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.741945Z","iopub.status.idle":"2024-06-13T08:38:44.742351Z","shell.execute_reply.started":"2024-06-13T08:38:44.742147Z","shell.execute_reply":"2024-06-13T08:38:44.742163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Check null: \", df_subm[\"score\"].isnull().any())\n","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.744343Z","iopub.status.idle":"2024-06-13T08:38:44.744769Z","shell.execute_reply.started":"2024-06-13T08:38:44.744540Z","shell.execute_reply":"2024-06-13T08:38:44.744557Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_subm.head()","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.745812Z","iopub.status.idle":"2024-06-13T08:38:44.746168Z","shell.execute_reply.started":"2024-06-13T08:38:44.745997Z","shell.execute_reply":"2024-06-13T08:38:44.746012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_subm.to_csv(\"submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-06-13T08:38:44.747833Z","iopub.status.idle":"2024-06-13T08:38:44.748205Z","shell.execute_reply.started":"2024-06-13T08:38:44.748021Z","shell.execute_reply":"2024-06-13T08:38:44.748037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}