{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":50160,"databundleVersionId":7921029,"sourceType":"competition"}],"dockerImageVersionId":30673,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"\n**1. Introduction:**\n   - Provide background information on the project's objective, which is to develop a credit risk model using LightGBM.\n   - Explain the significance of credit risk modeling in financial institutions for assessing the likelihood of default by borrowers.\n   - Introduce the dataset used for training the model, highlighting its relevance to the problem statement.\n   - Briefly outline the main steps involved in the project, such as data preprocessing, feature engineering, model selection, and evaluation.\n\n**2. Methodology:**\n   - **Data Collection and Preprocessing:**\n     - Describe the source of the dataset and any data cleaning steps performed to handle missing values or outliers.\n   - **Feature Engineering Techniques:**\n     - Explain the process of creating new features or transforming existing ones to improve model performance.\n   - **Model Selection: LightGBM Classifier:**\n     - Justify the choice of LightGBM as the primary modeling algorithm, highlighting its advantages for handling large datasets and categorical features.\n   - **Cross-Validation Strategy: Stratified Group K-Fold:**\n     - Provide rationale for using Stratified Group K-Fold cross-validation to ensure balanced splits while preserving temporal or group information in the data.\n   - **Training Process and Hyperparameter Tuning:**\n     - Detail the training process of the LightGBM model, including hyperparameter tuning techniques such as grid search or random search.\n   - **Evaluation Metrics: AUC Score:**\n     - Define the Area Under the ROC Curve (AUC) as the primary evaluation metric for assessing model performance on binary classification tasks.\n   - **Ensemble Modeling Approach: Voting Model:**\n     - Explain the concept of ensemble learning and the implementation of a Voting Model to combine predictions from multiple LightGBM models.\n\n**3. Results:**\n   - **Performance Evaluation on Validation Set:**\n     - Present the AUC scores achieved by each fold of the cross-validation process and discuss any variations observed.\n   - **Feature Importance Analysis:**\n     - Visualize and discuss the importance of features derived from the LightGBM models, highlighting key predictors of credit risk.\n   - **Model Comparison and Selection of the Best Model:**\n     - Compare the performance of individual LightGBM models and the ensemble Voting Model to select the best-performing model.\n   - **Discussion on the Significance of Key Features:**\n     - Interpret the impact of key features identified by the model on predicting credit risk and their potential implications for decision-making.\n\n**4. Discussion:**\n   - **Interpretation of Results:**\n     - Provide insights into the findings from the modeling process and their alignment with domain knowledge.\n   - **Potential Business Implications:**\n     - Discuss how the developed credit risk model can be utilized by financial institutions to optimize lending decisions and mitigate risks.\n   - **Limitations of the Study:**\n     - Acknowledge any limitations or constraints encountered during the project, such as data quality issues or assumptions made.\n   - **Future Work and Possible Improvements:**\n     - Suggest areas for future research or enhancements to the credit risk model, such as incorporating additional data sources or exploring alternative modeling techniques.\n\n**5. Conclusion:**\n   - **Summary of Findings:**\n     - Recap the key findings and outcomes of the project, emphasizing the effectiveness of the developed credit risk model.\n   - **Overall Conclusions Drawn from the Project:**\n     - Summarize the project's contributions to addressing the problem of credit risk assessment and its implications for stakeholders.\n   - **Final Thoughts and Recommendations:**\n     - Provide concluding remarks on the significance of the project and offer recommendations for further action or decision-making based on the results obtained.\n\n","metadata":{}},{"cell_type":"markdown","source":"Here's a step-by-step guide on how to proceed:\n\n1. **Install Required Libraries:**\n   - Make sure you have Python installed on your system. You can download and install Python from the official website: [python.org](https://www.python.org/).\n   - Use pip, Python's package installer, to install the required libraries. You can do this by running the following command in your terminal or command prompt:\n     ```\n     pip install polars lightgbm pandas seaborn matplotlib scikit-learn imbalanced-learn\n     ```\n\n2. **Download the Dataset:**\n   - If the dataset is not provided in the code snippet, you'll need to obtain it separately. Ensure that you have access to the dataset or replace the paths with the correct locations of your dataset files.\n\n3. **Run the Code:**\n   - Copy the provided code into a Python script file (with a .py extension) or a Jupyter Notebook.\n   - Update any file paths or configurations as needed to match your directory structure or dataset location.\n   - Execute the script or notebook cells to run the code.\n\n4. **Review Output:**\n   - After running the code, review the output, which may include visualizations, performance metrics, or saved files such as CSVs containing model predictions.\n\n5. **Troubleshooting:**\n   - If you encounter any errors during execution, carefully read the error messages to identify the issue. Common issues may include missing dependencies, incorrect file paths, or incompatible data types.\n\n","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-03-29T14:16:48.282688Z","iopub.execute_input":"2024-03-29T14:16:48.283359Z","iopub.status.idle":"2024-03-29T14:16:49.495997Z","shell.execute_reply.started":"2024-03-29T14:16:48.283321Z","shell.execute_reply":"2024-03-29T14:16:49.494792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os  # Importing the os module for operating system functionalities\nimport gc  # Importing the gc module for garbage collection\nfrom glob import glob  # Importing the glob function from the glob module for file path expansion\nfrom pathlib import Path  # Importing the Path class from the pathlib module for working with file paths\nfrom datetime import datetime  # Importing the datetime class from the datetime module for working with dates and times\n\nimport numpy as np  # Importing numpy library as np for numerical computing\nimport pandas as pd  # Importing pandas library as pd for data manipulation and analysis\nimport polars as pl  # Importing polars library as pl for data manipulation and analysis\n\nimport matplotlib.pyplot as plt  # Importing pyplot module from matplotlib library for creating visualizations\nimport seaborn as sns  # Importing seaborn library for statistical data visualization\n\nfrom sklearn.model_selection import TimeSeriesSplit, GroupKFold, StratifiedGroupKFold  # Importing classes for cross-validation strategies\nfrom sklearn.base import BaseEstimator, RegressorMixin  # Importing base classes for creating custom estimators\n\nimport joblib  # Importing joblib library for saving and loading models\n\nimport lightgbm as lgb  # Importing LightGBM library for gradient boosting framework\n\nimport warnings  # Importing warnings module for handling warnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)  # Ignoring FutureWarning to avoid cluttering output","metadata":{"execution":{"iopub.status.busy":"2024-03-29T14:16:55.320844Z","iopub.execute_input":"2024-03-29T14:16:55.321487Z","iopub.status.idle":"2024-03-29T14:16:59.062498Z","shell.execute_reply.started":"2024-03-29T14:16:55.321438Z","shell.execute_reply":"2024-03-29T14:16:59.061326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import sys  # Importing the sys module for system-specific parameters and functions\nfrom pathlib import Path  # Importing the Path class from the pathlib module for working with file paths\nimport subprocess  # Importing the subprocess module for spawning new processes\nimport os  # Importing the os module for operating system functionalities\nimport gc  # Importing the gc module for garbage collection\nfrom glob import glob  # Importing the glob function from the glob module for file path expansion\n\nimport numpy as np  # Importing numpy library as np for numerical computing\nimport pandas as pd  # Importing pandas library as pd for data manipulation and analysis\nimport polars as pl  # Importing polars library as pl for data manipulation and analysis\nfrom datetime import datetime  # Importing the datetime class from the datetime module for working with dates and times\nimport seaborn as sns  # Importing seaborn library for statistical data visualization\nimport matplotlib.pyplot as plt  # Importing pyplot module from matplotlib library for creating visualizations\n\nimport warnings  # Importing warnings module for handling warnings\nwarnings.filterwarnings('ignore')  # Ignoring warnings to avoid cluttering output\n\nROOT = '/kaggle/input/home-credit-credit-risk-model-stability'  # Defining ROOT as the root directory path","metadata":{"execution":{"iopub.status.busy":"2024-03-29T14:17:16.418928Z","iopub.execute_input":"2024-03-29T14:17:16.419610Z","iopub.status.idle":"2024-03-29T14:17:16.428983Z","shell.execute_reply.started":"2024-03-29T14:17:16.419574Z","shell.execute_reply":"2024-03-29T14:17:16.427308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import TimeSeriesSplit, GroupKFold, StratifiedGroupKFold  # Importing classes for cross-validation strategies\nfrom sklearn.base import BaseEstimator, RegressorMixin  # Importing base classes for creating custom estimators\nfrom sklearn.metrics import roc_auc_score  # Importing the roc_auc_score function for evaluating model performance\nimport lightgbm as lgb  # Importing LightGBM library for gradient boosting framework\n\nfrom imblearn.over_sampling import SMOTE  # Importing SMOTE from imbalanced-learn library for oversampling\nfrom sklearn.preprocessing import OrdinalEncoder  # Importing OrdinalEncoder from scikit-learn for encoding categorical features\nfrom sklearn.impute import KNNImputer  # Importing KNNImputer from scikit-learn for imputing missing values using K-nearest neighbors","metadata":{"execution":{"iopub.status.busy":"2024-03-29T14:17:49.536862Z","iopub.execute_input":"2024-03-29T14:17:49.537706Z","iopub.status.idle":"2024-03-29T14:17:50.015000Z","shell.execute_reply.started":"2024-03-29T14:17:49.537668Z","shell.execute_reply":"2024-03-29T14:17:50.013617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class Pipeline:\n    @staticmethod\n    def set_table_dtypes(df):  # Set table data types\n        for col in df.columns:\n            if col in [\"case_id\", \"WEEK_NUM\", \"num_group1\", \"num_group2\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Int64))  # Cast columns to Int64\n            elif col in [\"date_decision\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Date))  # Cast columns to Date\n            elif col[-1] in (\"P\", \"A\"):\n                df = df.with_columns(pl.col(col).cast(pl.Float64))  # Cast columns to Float64\n            elif col[-1] in (\"M\",):\n                df = df.with_columns(pl.col(col).cast(pl.String))  # Cast columns to String\n            elif col[-1] in (\"D\",):\n                df = df.with_columns(pl.col(col).cast(pl.Date))  # Cast columns to Date\n        return df\n\n    @staticmethod\n    def handle_dates(df):  # Handle dates\n        for col in df.columns:\n            if col[-1] in (\"D\",):\n                df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))  # Calculate time difference\n                df = df.with_columns(pl.col(col).dt.total_days())  # Extract total days\n        df = df.drop(\"date_decision\", \"MONTH\")  # Drop unnecessary columns\n        return df\n\n    @staticmethod\n    def filter_cols(df):  # Filter columns\n        for col in df.columns:\n            if col not in [\"target\", \"case_id\", \"WEEK_NUM\"]:\n                isnull = df[col].is_null().mean()\n                if isnull > 0.7:\n                    df = df.drop(col)  # Drop columns with high null percentage\n        \n        for col in df.columns:\n            if (col not in [\"target\", \"case_id\", \"WEEK_NUM\"]) & (df[col].dtype == pl.String):\n                freq = df[col].n_unique()\n                if (freq == 1) | (freq > 200):\n                    df = df.drop(col)  # Drop columns with low variability or too many unique values\n        return df\n\nclass Aggregator:\n    \n    @staticmethod\n    def num_expr(df):  # Numerical expressions\n        cols = [col for col in df.columns if col[-1] in (\"P\", \"A\")]\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]  # Aggregate maximum values\n        return expr_max\n    \n    @staticmethod\n    def date_expr(df):  # Date expressions\n        cols = [col for col in df.columns if col[-1] in (\"D\")]\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]  # Aggregate maximum values\n        return expr_max\n    \n    @staticmethod\n    def str_expr(df):  # String expressions\n        cols = [col for col in df.columns if col[-1] in (\"M\",)]\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]  # Aggregate maximum values\n        return expr_max\n    \n    @staticmethod\n    def other_expr(df):  # Other expressions\n        cols = [col for col in df.columns if col[-1] in (\"T\", \"L\")]\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]  # Aggregate maximum values\n        return expr_max\n    \n    @staticmethod\n    def count_expr(df):  # Count expressions\n        cols = [col for col in df.columns if \"num_group\" in col]\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]  # Aggregate maximum values\n        return expr_max\n    \n    @staticmethod\n    def get_exprs(df):  # Get expressions\n        exprs = Aggregator.num_expr(df) + \\\n                Aggregator.date_expr(df) + \\\n                Aggregator.str_expr(df) + \\\n                Aggregator.other_expr(df) + \\\n                Aggregator.count_expr(df)  # Concatenate expressions\n        return exprs","metadata":{"execution":{"iopub.status.busy":"2024-03-29T14:19:24.352934Z","iopub.execute_input":"2024-03-29T14:19:24.353708Z","iopub.status.idle":"2024-03-29T14:19:24.380964Z","shell.execute_reply.started":"2024-03-29T14:19:24.353667Z","shell.execute_reply":"2024-03-29T14:19:24.379404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def read_file(path, depth=None):\n    \"\"\"\n    Read a Parquet file and perform data preprocessing.\n\n    Args:\n        path (str): Path to the Parquet file.\n        depth (int, optional): Depth of aggregation. Defaults to None.\n\n    Returns:\n        pl.DataFrame: Preprocessed DataFrame.\n    \"\"\"\n    df = pl.read_parquet(path)  # Read Parquet file\n    df = df.pipe(Pipeline.set_table_dtypes)  # Set table data types\n    if depth in [1, 2]:\n        df = df.group_by(\"case_id\").agg(Aggregator.get_exprs(df))  # Aggregate data\n    return df\n\ndef read_files(regex_path, depth=None):\n    \"\"\"\n    Read multiple Parquet files matching a regex pattern and perform data preprocessing.\n\n    Args:\n        regex_path (str): Regular expression pattern for file paths.\n        depth (int, optional): Depth of aggregation. Defaults to None.\n\n    Returns:\n        pl.DataFrame: Concatenated and preprocessed DataFrame.\n    \"\"\"\n    chunks = []\n    \n    for path in glob(str(regex_path)):\n        df = pl.read_parquet(path)  # Read Parquet file\n        df = df.pipe(Pipeline.set_table_dtypes)  # Set table data types\n        if depth in [1, 2]:\n            df = df.group_by(\"case_id\").agg(Aggregator.get_exprs(df))  # Aggregate data\n        chunks.append(df)\n    \n    df = pl.concat(chunks, how=\"vertical_relaxed\")  # Concatenate DataFrames\n    df = df.unique(subset=[\"case_id\"])  # Remove duplicate case_ids\n    return df\n\ndef feature_eng(df_base, depth_0, depth_1, depth_2):\n    \"\"\"\n    Perform feature engineering on a base DataFrame.\n\n    Args:\n        df_base (pl.DataFrame): Base DataFrame.\n        depth_0 (list): List of DataFrames for depth 0 aggregation.\n        depth_1 (list): List of DataFrames for depth 1 aggregation.\n        depth_2 (list): List of DataFrames for depth 2 aggregation.\n\n    Returns:\n        pl.DataFrame: DataFrame with engineered features.\n    \"\"\"\n    df_base = (\n        df_base\n        .with_columns(\n            month_decision=pl.col(\"date_decision\").dt.month(),\n            weekday_decision=pl.col(\"date_decision\").dt.weekday(),\n        )\n    )\n    for i, df in enumerate(depth_0 + depth_1 + depth_2):\n        df_base = df_base.join(df, how=\"left\", on=\"case_id\", suffix=f\"_{i}\")  # Join DataFrames\n    df_base = df_base.pipe(Pipeline.handle_dates)  # Handle dates\n    return df_base\n\ndef to_pandas(df_data, cat_cols=None):\n    \"\"\"\n    Convert DataFrame to Pandas DataFrame and perform type conversion.\n\n    Args:\n        df_data (pl.DataFrame): DataFrame to convert.\n        cat_cols (list, optional): List of categorical columns. Defaults to None.\n\n    Returns:\n        pd.DataFrame: Pandas DataFrame.\n        list: List of categorical columns.\n    \"\"\"\n    df_data = df_data.to_pandas()  # Convert DataFrame to Pandas DataFrame\n    if cat_cols is None:\n        cat_cols = list(df_data.select_dtypes(\"object\").columns)  # Select categorical columns\n    df_data[cat_cols] = df_data[cat_cols].astype(\"category\")  # Convert categorical columns to category type\n    return df_data, cat_cols","metadata":{"execution":{"iopub.status.busy":"2024-03-29T14:20:24.213635Z","iopub.execute_input":"2024-03-29T14:20:24.214209Z","iopub.status.idle":"2024-03-29T14:20:24.231555Z","shell.execute_reply.started":"2024-03-29T14:20:24.214167Z","shell.execute_reply":"2024-03-29T14:20:24.229956Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ROOT = Path(\"/kaggle/input/home-credit-credit-risk-model-stability\")  # Define root directory path\nTRAIN_DIR = ROOT / \"parquet_files\" / \"train\"  # Define train directory path\nTEST_DIR = ROOT / \"parquet_files\" / \"test\"  # Define test directory path","metadata":{"execution":{"iopub.status.busy":"2024-03-29T14:20:51.987152Z","iopub.execute_input":"2024-03-29T14:20:51.987615Z","iopub.status.idle":"2024-03-29T14:20:51.994521Z","shell.execute_reply.started":"2024-03-29T14:20:51.987581Z","shell.execute_reply":"2024-03-29T14:20:51.993059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_store = {\n    \"df_base\": read_file(TRAIN_DIR / \"train_base.parquet\"),  # Read base DataFrame\n    \"depth_0\": [\n        read_file(TRAIN_DIR / \"train_static_cb_0.parquet\"),  # Read depth 0 DataFrame\n        read_files(TRAIN_DIR / \"train_static_0_*.parquet\"),  # Read multiple depth 0 DataFrames\n    ],\n    \"depth_1\": [\n        read_files(TRAIN_DIR / \"train_applprev_1_*.parquet\", 1),  # Read multiple depth 1 DataFrames\n        read_file(TRAIN_DIR / \"train_tax_registry_a_1.parquet\", 1),  # Read depth 1 DataFrame\n        read_file(TRAIN_DIR / \"train_tax_registry_b_1.parquet\", 1),  # Read depth 1 DataFrame\n        read_file(TRAIN_DIR / \"train_tax_registry_c_1.parquet\", 1),  # Read depth 1 DataFrame\n        read_files(TRAIN_DIR / \"train_credit_bureau_a_1_*.parquet\", 1),  # Read multiple depth 1 DataFrames\n        read_file(TRAIN_DIR / \"train_credit_bureau_b_1.parquet\", 1),  # Read depth 1 DataFrame\n        read_file(TRAIN_DIR / \"train_other_1.parquet\", 1),  # Read depth 1 DataFrame\n        read_file(TRAIN_DIR / \"train_person_1.parquet\", 1),  # Read depth 1 DataFrame\n        read_file(TRAIN_DIR / \"train_deposit_1.parquet\", 1),  # Read depth 1 DataFrame\n        read_file(TRAIN_DIR / \"train_debitcard_1.parquet\", 1),  # Read depth 1 DataFrame\n    ],\n    \"depth_2\": [\n        read_file(TRAIN_DIR / \"train_credit_bureau_b_2.parquet\", 2),  # Read depth 2 DataFrame\n        read_files(TRAIN_DIR / \"train_credit_bureau_a_2_*.parquet\", 2),  # Read multiple depth 2 DataFrames\n        read_file(TRAIN_DIR / \"train_applprev_2.parquet\", 2),  # Read depth 2 DataFrame\n        read_file(TRAIN_DIR / \"train_person_2.parquet\", 2)  # Read depth 2 DataFrame\n    ]\n}\n\ndf_train = feature_eng(**data_store)  # Perform feature engineering\nprint(\"train data shape:\\t\", df_train.shape)  # Print shape of train data","metadata":{"execution":{"iopub.status.busy":"2024-03-29T14:23:04.235552Z","iopub.execute_input":"2024-03-29T14:23:04.236022Z","iopub.status.idle":"2024-03-29T14:25:46.816124Z","shell.execute_reply.started":"2024-03-29T14:23:04.235989Z","shell.execute_reply":"2024-03-29T14:25:46.814640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_store = {\n    \"df_base\": read_file(TEST_DIR / \"test_base.parquet\"),  # Read base DataFrame\n    \"depth_0\": [\n        read_file(TEST_DIR / \"test_static_cb_0.parquet\"),  # Read depth 0 DataFrame\n        read_files(TEST_DIR / \"test_static_0_*.parquet\"),  # Read multiple depth 0 DataFrames\n    ],\n    \"depth_1\": [\n        read_files(TEST_DIR / \"test_applprev_1_*.parquet\", 1),  # Read multiple depth 1 DataFrames\n        read_file(TEST_DIR / \"test_tax_registry_a_1.parquet\", 1),  # Read depth 1 DataFrame\n        read_file(TEST_DIR / \"test_tax_registry_b_1.parquet\", 1),  # Read depth 1 DataFrame\n        read_file(TEST_DIR / \"test_tax_registry_c_1.parquet\", 1),  # Read depth 1 DataFrame\n        read_files(TEST_DIR / \"test_credit_bureau_a_1_*.parquet\", 1),  # Read multiple depth 1 DataFrames\n        read_file(TEST_DIR / \"test_credit_bureau_b_1.parquet\", 1),  # Read depth 1 DataFrame\n        read_file(TEST_DIR / \"test_other_1.parquet\", 1),  # Read depth 1 DataFrame\n        read_file(TEST_DIR / \"test_person_1.parquet\", 1),  # Read depth 1 DataFrame\n        read_file(TEST_DIR / \"test_deposit_1.parquet\", 1),  # Read depth 1 DataFrame\n        read_file(TEST_DIR / \"test_debitcard_1.parquet\", 1),  # Read depth 1 DataFrame\n    ],\n    \"depth_2\": [\n        read_file(TEST_DIR / \"test_credit_bureau_b_2.parquet\", 2),  # Read depth 2 DataFrame\n        read_files(TEST_DIR / \"test_credit_bureau_a_2_*.parquet\", 2),  # Read multiple depth 2 DataFrames\n        read_file(TEST_DIR / \"test_applprev_2.parquet\", 2),  # Read depth 2 DataFrame\n        read_file(TEST_DIR / \"test_person_2.parquet\", 2)  # Read depth 2 DataFrame\n    ]\n}","metadata":{"execution":{"iopub.status.busy":"2024-03-29T14:26:28.117749Z","iopub.execute_input":"2024-03-29T14:26:28.118228Z","iopub.status.idle":"2024-03-29T14:26:29.095925Z","shell.execute_reply.started":"2024-03-29T14:26:28.118178Z","shell.execute_reply":"2024-03-29T14:26:29.094414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test = feature_eng(**data_store)  # Perform feature engineering on test data\nprint(\"test data shape:\\t\", df_test.shape)  # Print shape of test data","metadata":{"execution":{"iopub.status.busy":"2024-03-29T14:27:15.474303Z","iopub.execute_input":"2024-03-29T14:27:15.474693Z","iopub.status.idle":"2024-03-29T14:27:15.528639Z","shell.execute_reply.started":"2024-03-29T14:27:15.474664Z","shell.execute_reply":"2024-03-29T14:27:15.527293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Drop the insignificant features from train data\ndf_train = df_train.pipe(Pipeline.filter_cols)\n\n# Select columns from test data excluding the target column\ndf_test = df_test.select([col for col in df_train.columns if col != \"target\"])\n\nprint(\"train data shape:\\t\", df_train.shape)  # Print shape of train data\nprint(\"test data shape:\\t\", df_test.shape)  # Print shape of test data","metadata":{"execution":{"iopub.status.busy":"2024-03-29T14:27:20.651889Z","iopub.execute_input":"2024-03-29T14:27:20.652941Z","iopub.status.idle":"2024-03-29T14:27:24.300207Z","shell.execute_reply.started":"2024-03-29T14:27:20.652886Z","shell.execute_reply":"2024-03-29T14:27:24.298982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Convert train data to Pandas DataFrame and get categorical columns\ndf_train, cat_cols = to_pandas(df_train)\n\n# Convert test data to Pandas DataFrame and get categorical columns\ndf_test, cat_cols = to_pandas(df_test, cat_cols)\n\n# Delete data_store dictionary and perform garbage collection\ndel data_store\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2024-03-29T14:27:34.738789Z","iopub.execute_input":"2024-03-29T14:27:34.739942Z","iopub.status.idle":"2024-03-29T14:27:59.420668Z","shell.execute_reply.started":"2024-03-29T14:27:34.739900Z","shell.execute_reply":"2024-03-29T14:27:59.419479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the first few rows of the train data\ndisplay(df_train.head())\n\n# Display the first few rows of the test data\ndisplay(df_test.head())","metadata":{"execution":{"iopub.status.busy":"2024-03-29T14:28:20.639566Z","iopub.execute_input":"2024-03-29T14:28:20.639976Z","iopub.status.idle":"2024-03-29T14:28:20.708782Z","shell.execute_reply.started":"2024-03-29T14:28:20.639947Z","shell.execute_reply":"2024-03-29T14:28:20.707719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = df_train.drop(columns=[\"target\", \"case_id\", \"WEEK_NUM\"])\ny = df_train[\"target\"]\nweeks = df_train[\"WEEK_NUM\"]\n\ncv = StratifiedGroupKFold(n_splits=5, shuffle=False)\n\nparams = {\n    \"boosting_type\": \"gbdt\",\n    \"objective\": \"binary\",\n    \"metric\": \"auc\",\n    \"max_depth\": 10,  \n    \"learning_rate\": 0.05,\n    \"n_estimators\": 2000,  \n    \"colsample_bytree\": 0.8,\n    \"colsample_bynode\": 0.8,\n    \"verbose\": -1,\n    \"random_state\": 42,\n    \"reg_alpha\": 0.1,\n    \"reg_lambda\": 10,\n    \"extra_trees\":True,\n    'num_leaves':64,\n    \"device\": \"cpu\",  # Change device to CPU\n    \"verbose\": -1,\n}\n\nfitted_models = []\ncv_scores = []\n\nfor idx_train, idx_valid in cv.split(X, y, groups=weeks):\n    X_train, y_train = X.iloc[idx_train], y.iloc[idx_train]\n    X_valid, y_valid = X.iloc[idx_valid], y.iloc[idx_valid]\n    \n    model = lgb.LGBMClassifier(**params)\n    model.fit(\n        X_train, y_train,\n        eval_set = [(X_valid, y_valid)],\n        callbacks = [lgb.log_evaluation(400), lgb.early_stopping(5)] )\n    fitted_models.append(model)\n    \n    y_pred_valid = model.predict_proba(X_valid)[:,1]\n    auc_score = roc_auc_score(y_valid, y_pred_valid)\n    cv_scores.append(auc_score)\n    \nprint(\"CV AUC scores: \", cv_scores)\nprint(\"Maximum CV AUC score: \", max(cv_scores))\n","metadata":{"execution":{"iopub.status.busy":"2024-03-29T14:51:52.775579Z","iopub.execute_input":"2024-03-29T14:51:52.776011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class VotingModel(BaseEstimator, RegressorMixin):\n    \"\"\"\n    A custom ensemble model that averages predictions from multiple estimators.\n\n    Args:\n        estimators (list): List of estimators to be included in the ensemble.\n\n    Attributes:\n        estimators (list): List of estimators.\n    \"\"\"\n\n    def __init__(self, estimators):\n        super().__init__()\n        self.estimators = estimators\n        \n    def fit(self, X, y=None):\n        \"\"\"\n        Fit the ensemble model.\n\n        Args:\n            X (array-like): Input features.\n            y (array-like, optional): Target values. Defaults to None.\n\n        Returns:\n            self: Returns an instance of self.\n        \"\"\"\n        return self\n    \n    def predict(self, X):\n        \"\"\"\n        Make predictions using the ensemble model.\n\n        Args:\n            X (array-like): Input features.\n\n        Returns:\n            array-like: Predicted target values.\n        \"\"\"\n        y_preds = [estimator.predict(X) for estimator in self.estimators]\n        return np.mean(y_preds, axis=0)\n    \n    def predict_proba(self, X):\n        \"\"\"\n        Make probability predictions using the ensemble model.\n\n        Args:\n            X (array-like): Input features.\n\n        Returns:\n            array-like: Predicted probabilities.\n        \"\"\"\n        y_preds = [estimator.predict_proba(X) for estimator in self.estimators]\n        return np.mean(y_preds, axis=0)\n\n# Instantiate the VotingModel with the fitted models\nmodel = VotingModel(fitted_models)\n\n# Plot feature importance of one of the fitted models\nlgb.plot_importance(fitted_models[2], importance_type=\"split\", figsize=(10,50))\nplt.show()\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the column names (features)\nfeatures = X_train.columns\n\n# Get the feature importances from the third fitted LightGBM model (index 2)\nimportances = fitted_models[2].feature_importances_\n\n# Create a DataFrame to store feature importances and sort them\nfeature_importance = pd.DataFrame({'importance': importances, 'features': features}).sort_values('importance', ascending=False).reset_index(drop=True)\n\n# Display the DataFrame\nprint(feature_importance)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"drop_list = []\n\n# Iterate over the rows of the feature importance DataFrame\nfor i, f in feature_importance.iterrows():\n    # If the importance score is less than 80, add the feature to the drop list\n    if f['importance'] < 80:\n        drop_list.append(f['features'])\n\n# Print the number of features to be dropped\nprint(f\"Number of features which are not important: {len(drop_list)}\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(drop_list)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Drop the WEEK_NUM column from the test dataset and set case_id as the index\nX_test = df_test.drop(columns=[\"WEEK_NUM\"])\nX_test = X_test.set_index(\"case_id\")\n\n# Make predictions using the trained LightGBM model on the test dataset\nlgb_pred = pd.Series(model.predict_proba(X_test)[:, 1], index=X_test.index)\n\n# Read the sample submission file\ndf_subm = pd.read_csv(ROOT / \"sample_submission.csv\")\ndf_subm = df_subm.set_index(\"case_id\")  # Set case_id as the index\n\n# Assign the predicted scores to the corresponding case_id in the submission DataFrame\ndf_subm[\"score\"] = lgb_pred\n\n# Display the submission DataFrame (if needed)\nprint(df_subm)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_subm.head())","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_subm.to_csv(\"submission.csv\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}