{"metadata":{"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":50160,"databundleVersionId":7921029,"sourceType":"competition"},{"sourceId":33095,"sourceType":"modelInstanceVersion","modelInstanceId":27710},{"sourceId":33096,"sourceType":"modelInstanceVersion","modelInstanceId":27711}],"dockerImageVersionId":30699,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.10.13"},"papermill":{"default_parameters":{},"duration":17.915251,"end_time":"2024-04-18T01:06:04.765569","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2024-04-18T01:05:46.850318","version":"2.5.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"code step by step:\n\n1. **Import Libraries**: In the first part, we import all the necessary libraries and modules that we'll be using throughout the code. These include libraries for data manipulation (`numpy`, `pandas`, `polars`), machine learning models (`lightgbm`, `catboost`), file handling (`joblib`, `Path`), and other utilities (`gc`, `glob`).\n\n2. **Define Classes for Data Processing**: The code defines two classes: `Pipeline` and `Aggregator`. These classes encapsulate methods for data preprocessing and feature engineering, respectively.\n\n3. **Define Functions for Data Reading and Feature Engineering**: Several functions are defined to read data files, perform feature engineering, and convert data to a more memory-efficient format. These functions include `read_file`, `read_files`, `feature_eng`, `to_pandas`, and `reduce_mem_usage`.\n\n4. **Load Trained Models and Model Metadata**: The code loads trained models (`lgb_models`, `cat_models`) and their associated metadata (`lgb_notebook_info`, `cat_notebook_info`) from disk. These models are later used for making predictions on the test data.\n\n5. **Define Test Data Paths and Load Test Data**: Test data paths are defined, and the test data is loaded using the previously defined functions for reading data files. The loaded test data is stored in a dictionary called `data_store`.\n\n6. **Perform Feature Engineering on Test Data**: The loaded test data is passed through the feature engineering pipeline (`feature_eng`) to generate features required for making predictions.\n\n7. **Generate Predictions**: The `VotingModel` class is used to generate predictions on the test data. This class averages the predictions from multiple individual models to obtain the final prediction probabilities.\n\n8. **Save Predictions to Submission File**: The predicted probabilities are saved to a CSV file (`submission.csv`) in the format required for submission. The submission file is based on a sample submission file provided earlier (`sample_submission.csv`).\n\n9. **Display Submission DataFrame**: Finally, the submission DataFrame (`df_subm`) is displayed, showing the case IDs and corresponding predicted scores.\n\n","metadata":{}},{"cell_type":"code","source":"import joblib  # Import joblib for saving and loading models\nfrom pathlib import Path  # Import Path for working with file paths\nimport gc  # Import gc for garbage collection\nfrom glob import glob  # Import glob for file matching\nimport numpy as np  # Import numpy for numerical computing\nimport pandas as pd  # Import pandas for data manipulation\nimport polars as pl  # Import polars for fast data manipulation\nfrom sklearn.base import BaseEstimator, RegressorMixin  # Import BaseEstimator and RegressorMixin from sklearn.base\nfrom sklearn.metrics import roc_auc_score  # Import roc_auc_score from sklearn.metrics\nimport lightgbm as lgb  # Import lightgbm for gradient boosting\n\nimport warnings  # Import warnings to ignore warnings\nwarnings.filterwarnings('ignore')  # Ignore warnings\n\nROOT = '/kaggle/input/home-credit-credit-risk-model-stability'  # Define ROOT path","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","papermill":{"duration":6.044727,"end_time":"2024-04-18T01:05:56.081208","exception":false,"start_time":"2024-04-18T01:05:50.036481","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-04-20T16:49:38.511999Z","iopub.execute_input":"2024-04-20T16:49:38.512375Z","iopub.status.idle":"2024-04-20T16:49:43.683354Z","shell.execute_reply.started":"2024-04-20T16:49:38.512324Z","shell.execute_reply":"2024-04-20T16:49:43.682352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class Pipeline:\n    # Method to set data types for specific columns in a DataFrame\n    def set_table_dtypes(df):\n        for col in df.columns:\n            if col in [\"case_id\", \"WEEK_NUM\", \"num_group1\", \"num_group2\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Int64))\n            elif col in [\"date_decision\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Date))\n            elif col[-1] in (\"P\", \"A\"):\n                df = df.with_columns(pl.col(col).cast(pl.Float64))\n            elif col[-1] in (\"M\",):\n                df = df.with_columns(pl.col(col).cast(pl.String))\n            elif col[-1] in (\"D\",):\n                df = df.with_columns(pl.col(col).cast(pl.Date))\n        return df\n\n    # Method to handle date columns and calculate time differences\n    def handle_dates(df):\n        for col in df.columns:\n            if col[-1] in (\"D\",):\n                df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))  # Calculate time differences\n                df = df.with_columns(pl.col(col).dt.total_days())  # Convert time differences to total days\n        df = df.drop(\"date_decision\", \"MONTH\")  # Drop unnecessary columns\n        return df\n\n    # Method to filter out columns based on missing values and frequency\n    def filter_cols(df):\n        for col in df.columns:\n            if col not in [\"target\", \"case_id\", \"WEEK_NUM\"]:\n                isnull = df[col].is_null().mean()\n                if isnull > 0.7:\n                    df = df.drop(col)  # Drop columns with more than 70% missing values\n        \n        for col in df.columns:\n            if (col not in [\"target\", \"case_id\", \"WEEK_NUM\"]) & (df[col].dtype == pl.String):\n                freq = df[col].n_unique()\n                if (freq == 1) | (freq > 200):\n                    df = df.drop(col)  # Drop columns with only one unique value or more than 200 unique values\n        \n        return df","metadata":{"papermill":{"duration":0.034893,"end_time":"2024-04-18T01:05:56.121806","exception":false,"start_time":"2024-04-18T01:05:56.086913","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-04-20T16:49:43.685157Z","iopub.execute_input":"2024-04-20T16:49:43.685481Z","iopub.status.idle":"2024-04-20T16:49:43.697671Z","shell.execute_reply.started":"2024-04-20T16:49:43.685444Z","shell.execute_reply":"2024-04-20T16:49:43.696751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class Aggregator:\n    # Method to aggregate numerical features\n    def num_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"P\", \"A\")]  # Select numerical columns\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]  # Calculate max\n        expr_last = [pl.last(col).alias(f\"last_{col}\") for col in cols]  # Calculate last\n        expr_mean = [pl.mean(col).alias(f\"mean_{col}\") for col in cols]  # Calculate mean\n        expr_median = [pl.median(col).alias(f\"median_{col}\") for col in cols]  # Calculate median\n        expr_var = [pl.var(col).alias(f\"var_{col}\") for col in cols]  # Calculate variance\n        return expr_max + expr_last + expr_mean \n\n    # Method to aggregate date features\n    def date_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"D\")]  # Select date columns\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]  # Calculate max\n        expr_last = [pl.last(col).alias(f\"last_{col}\") for col in cols]  # Calculate last\n        expr_mean = [pl.mean(col).alias(f\"mean_{col}\") for col in cols]  # Calculate mean\n        expr_median = [pl.median(col).alias(f\"median_{col}\") for col in cols]  # Calculate median\n        return expr_max + expr_last + expr_mean \n\n    # Method to aggregate string features\n    def str_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"M\",)]  # Select string columns\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]  # Calculate max\n        expr_last = [pl.last(col).alias(f\"last_{col}\") for col in cols]  # Calculate last\n        return expr_max + expr_last\n\n    # Method to aggregate other features\n    def other_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"T\", \"L\")]  # Select other columns\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]  # Calculate max\n        expr_last = [pl.last(col).alias(f\"last_{col}\") for col in cols]  # Calculate last\n        return expr_max + expr_last\n\n    # Method to aggregate count features\n    def count_expr(df):\n        cols = [col for col in df.columns if \"num_group\" in col]  # Select count columns\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]  # Calculate max\n        expr_last = [pl.last(col).alias(f\"last_{col}\") for col in cols]  # Calculate last\n        return expr_max + expr_last\n\n    # Method to get all aggregation expressions\n    def get_exprs(df):\n        exprs = Aggregator.num_expr(df) + \\\n                Aggregator.date_expr(df) + \\\n                Aggregator.str_expr(df) + \\\n                Aggregator.other_expr(df) + \\\n                Aggregator.count_expr(df)\n        return exprs","metadata":{"execution":{"iopub.status.busy":"2024-04-20T16:49:43.698983Z","iopub.execute_input":"2024-04-20T16:49:43.69932Z","iopub.status.idle":"2024-04-20T16:49:43.717216Z","shell.execute_reply.started":"2024-04-20T16:49:43.699297Z","shell.execute_reply":"2024-04-20T16:49:43.716319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def read_file(path, depth=None):\n    # Read parquet file into a Polars DataFrame\n    df = pl.read_parquet(path)\n    # Set table data types using Pipeline method\n    df = df.pipe(Pipeline.set_table_dtypes)\n    # Aggregate features if depth is specified\n    if depth in [1, 2]:\n        df = df.group_by(\"case_id\").agg(Aggregator.get_exprs(df)) \n    return df\n","metadata":{"papermill":{"duration":0.032096,"end_time":"2024-04-18T01:05:56.159326","exception":false,"start_time":"2024-04-18T01:05:56.12723","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-04-20T16:49:43.719222Z","iopub.execute_input":"2024-04-20T16:49:43.719527Z","iopub.status.idle":"2024-04-20T16:49:43.730492Z","shell.execute_reply.started":"2024-04-20T16:49:43.719505Z","shell.execute_reply":"2024-04-20T16:49:43.729674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def read_files(regex_path, depth=None):\n    chunks = []\n    # Iterate over files matching the regex pattern\n    for path in glob(str(regex_path)):\n        # Read parquet file into a Polars DataFrame\n        df = pl.read_parquet(path)\n        # Set table data types using Pipeline method\n        df = df.pipe(Pipeline.set_table_dtypes)\n        # Aggregate features if depth is specified\n        if depth in [1, 2]:\n            df = df.group_by(\"case_id\").agg(Aggregator.get_exprs(df))\n        chunks.append(df)\n    # Concatenate DataFrames and drop duplicate rows based on \"case_id\"\n    df = pl.concat(chunks, how=\"vertical_relaxed\").unique(subset=[\"case_id\"])\n    return df","metadata":{"execution":{"iopub.status.busy":"2024-04-20T16:49:43.731533Z","iopub.execute_input":"2024-04-20T16:49:43.731815Z","iopub.status.idle":"2024-04-20T16:49:43.74083Z","shell.execute_reply.started":"2024-04-20T16:49:43.731792Z","shell.execute_reply":"2024-04-20T16:49:43.740124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def feature_eng(df_base, depth_0, depth_1, depth_2):\n    # Add month and weekday features based on \"date_decision\"\n    df_base = df_base.with_columns(\n        month_decision = pl.col(\"date_decision\").dt.month(),\n        weekday_decision = pl.col(\"date_decision\").dt.weekday(),\n    )\n    # Join additional depth DataFrames\n    for i, df in enumerate(depth_0 + depth_1 + depth_2):\n        df_base = df_base.join(df, how=\"left\", on=\"case_id\", suffix=f\"_{i}\")\n    # Handle dates using Pipeline method\n    df_base = df_base.pipe(Pipeline.handle_dates)\n    return df_base","metadata":{"execution":{"iopub.status.busy":"2024-04-20T16:49:43.741783Z","iopub.execute_input":"2024-04-20T16:49:43.742044Z","iopub.status.idle":"2024-04-20T16:49:43.751152Z","shell.execute_reply.started":"2024-04-20T16:49:43.742012Z","shell.execute_reply":"2024-04-20T16:49:43.750207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def to_pandas(df_data, cat_cols=None):\n    # Convert Polars DataFrame to pandas DataFrame\n    df_data = df_data.to_pandas()\n    # Convert categorical columns to category data type\n    if cat_cols is None:\n        cat_cols = list(df_data.select_dtypes(\"object\").columns)\n    df_data[cat_cols] = df_data[cat_cols].astype(\"category\")\n    return df_data, cat_cols","metadata":{"execution":{"iopub.status.busy":"2024-04-20T16:49:43.752272Z","iopub.execute_input":"2024-04-20T16:49:43.752612Z","iopub.status.idle":"2024-04-20T16:49:43.764564Z","shell.execute_reply.started":"2024-04-20T16:49:43.752588Z","shell.execute_reply":"2024-04-20T16:49:43.763682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def reduce_mem_usage(df):\n    \"\"\" \n    Iterate through all the columns of a dataframe and modify the data type\n    to reduce memory usage.\n    \"\"\"\n    start_mem = df.memory_usage().sum() / 1024**2  # Memory usage before optimization\n    print('Memory usage of dataframe is {:.2f} MB'.format(start_mem))\n    \n    for col in df.columns:\n        col_type = df[col].dtype\n        if str(col_type)==\"category\":\n            continue\n        \n        if col_type != object:\n            c_min = df[col].min()\n            c_max = df[col].max()\n            if str(col_type)[:3] == 'int':\n                if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                    df[col] = df[col].astype(np.int8)\n                elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                    df[col] = df[col].astype(np.int16)\n                elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                    df[col] = df[col].astype(np.int32)\n                elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                    df[col] = df[col].astype(np.int64)  \n            else:\n                if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                    df[col] = df[col].astype(np.float16)\n                elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                    df[col] = df[col].astype(np.float32)\n                else:\n                    df[col] = df[col].astype(np.float64)\n        else:\n            continue\n    end_mem = df.memory_usage().sum() / 1024**2  # Memory usage after optimization\n    print('Memory usage after optimization is: {:.2f} MB'.format(end_mem))\n    print('Decreased by {:.1f}%'.format(100 * (start_mem - end_mem) / start_mem))\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2024-04-20T16:49:43.765726Z","iopub.execute_input":"2024-04-20T16:49:43.766024Z","iopub.status.idle":"2024-04-20T16:49:43.779552Z","shell.execute_reply.started":"2024-04-20T16:49:43.766001Z","shell.execute_reply":"2024-04-20T16:49:43.778709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lgb_notebook_info = joblib.load('/kaggle/input/homecredit-models-public/other/lgb/1/notebook_info.joblib')\n\n# Print notebook information\nprint(f\"- [lgb] notebook_start_time: {lgb_notebook_info['notebook_start_time']}\")\nprint(f\"- [lgb] description: {lgb_notebook_info['description']}\")\n\n# Load columns and categorical columns\ncols = lgb_notebook_info['cols']\ncat_cols = lgb_notebook_info['cat_cols']\nprint(f\"- [lgb] len(cols): {len(cols)}\")\nprint(f\"- [lgb] len(cat_cols): {len(cat_cols)}\")\n\n# Load LightGBM models\nlgb_models = joblib.load('/kaggle/input/homecredit-models-public/other/lgb/1/lgb_models.joblib')\nlgb_models","metadata":{"papermill":{"duration":0.722458,"end_time":"2024-04-18T01:05:56.898184","exception":false,"start_time":"2024-04-18T01:05:56.175726","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-04-20T16:49:43.780587Z","iopub.execute_input":"2024-04-20T16:49:43.780841Z","iopub.status.idle":"2024-04-20T16:49:44.792315Z","shell.execute_reply.started":"2024-04-20T16:49:43.780819Z","shell.execute_reply":"2024-04-20T16:49:44.791459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load categorical model notebook information\ncat_notebook_info = joblib.load('/kaggle/input/homecredit-models-public/other/cat/1/notebook_info.joblib')\n\n# Print notebook information\nprint(f\"- [cat] notebook_start_time: {cat_notebook_info['notebook_start_time']}\")\nprint(f\"- [cat] description: {cat_notebook_info['description']}\")\n\n# Load categorical models\ncat_models = joblib.load('/kaggle/input/homecredit-models-public/other/cat/1/cat_models.joblib')\ncat_models","metadata":{"papermill":{"duration":4.878543,"end_time":"2024-04-18T01:06:01.784082","exception":false,"start_time":"2024-04-18T01:05:56.905539","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-04-20T16:49:44.795724Z","iopub.execute_input":"2024-04-20T16:49:44.796015Z","iopub.status.idle":"2024-04-20T16:49:49.486001Z","shell.execute_reply.started":"2024-04-20T16:49:44.795991Z","shell.execute_reply":"2024-04-20T16:49:49.485113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the root directory path\nROOT = Path(\"/kaggle/input/home-credit-credit-risk-model-stability\")\n\n# Define the directory path for the test data\nTEST_DIR = ROOT / \"parquet_files\" / \"test\"\n\n# Create a dictionary to store different dataframes generated from reading parquet files\ndata_store = {\n    # Read the base test data and store it with the key 'df_base'\n    \"df_base\": read_file(TEST_DIR / \"test_base.parquet\"),\n    \n    # Read depth 0 data, which includes static data and additional files matching a pattern\n    \"depth_0\": [\n        read_file(TEST_DIR / \"test_static_cb_0.parquet\"),\n        read_files(TEST_DIR / \"test_static_0_*.parquet\"),\n    ],\n    \n    # Read depth 1 data, including various files related to applicant previous applications, tax registries,\n    # credit bureau data, and other information\n    \"depth_1\": [\n        read_files(TEST_DIR / \"test_applprev_1_*.parquet\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_a_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_b_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_c_1.parquet\", 1),\n        read_files(TEST_DIR / \"test_credit_bureau_a_1_*.parquet\", 1),\n        read_file(TEST_DIR / \"test_credit_bureau_b_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_other_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_person_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_deposit_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_debitcard_1.parquet\", 1),\n    ],\n    \n    # Read depth 2 data, which includes additional credit bureau data, applicant previous applications,\n    # and personal information\n    \"depth_2\": [\n        read_file(TEST_DIR / \"test_credit_bureau_b_2.parquet\", 2),\n        read_files(TEST_DIR / \"test_credit_bureau_a_2_*.parquet\", 2),\n        read_file(TEST_DIR / \"test_applprev_2.parquet\", 2),\n        read_file(TEST_DIR / \"test_person_2.parquet\", 2)\n    ]\n}","metadata":{"papermill":{"duration":0.545928,"end_time":"2024-04-18T01:06:02.349483","exception":false,"start_time":"2024-04-18T01:06:01.803555","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-04-20T16:49:49.487303Z","iopub.execute_input":"2024-04-20T16:49:49.487608Z","iopub.status.idle":"2024-04-20T16:49:49.918279Z","shell.execute_reply.started":"2024-04-20T16:49:49.487583Z","shell.execute_reply":"2024-04-20T16:49:49.917458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Perform feature engineering on the test data using the provided data store\ndf_test = feature_eng(**data_store)\n\n# Print the shape of the test data before further processing\nprint(\"test data shape:\\t\", df_test.shape)\n\n# Clean up memory by deleting the data store and running garbage collection\ndel data_store\ngc.collect()\n\n# Select columns of interest from the test data\ndf_test = df_test.select(['case_id'] + cols)\n\n# Convert the test data to a pandas DataFrame and optimize memory usage\ndf_test, cat_cols = to_pandas(df_test, cat_cols)\ndf_test = reduce_mem_usage(df_test)\n\n# Set the case_id column as the index of the DataFrame\ndf_test = df_test.set_index('case_id')\n\n# Print the shape of the test data after processing\nprint(\"test data shape:\\t\", df_test.shape)\n\n# Run garbage collection to clean up memory\ngc.collect()","metadata":{"papermill":{"duration":0.654075,"end_time":"2024-04-18T01:06:03.010344","exception":false,"start_time":"2024-04-18T01:06:02.356269","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-04-20T16:49:49.919488Z","iopub.execute_input":"2024-04-20T16:49:49.919808Z","iopub.status.idle":"2024-04-20T16:49:50.487384Z","shell.execute_reply.started":"2024-04-20T16:49:49.919782Z","shell.execute_reply":"2024-04-20T16:49:50.486432Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test","metadata":{"papermill":{"duration":0.05776,"end_time":"2024-04-18T01:06:03.074529","exception":false,"start_time":"2024-04-18T01:06:03.016769","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-04-20T16:49:50.488561Z","iopub.execute_input":"2024-04-20T16:49:50.488842Z","iopub.status.idle":"2024-04-20T16:49:50.535121Z","shell.execute_reply.started":"2024-04-20T16:49:50.488819Z","shell.execute_reply":"2024-04-20T16:49:50.53418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class VotingModel(BaseEstimator, RegressorMixin):\n    def __init__(self, estimators):\n        super().__init__()\n        self.estimators = estimators\n        \n    def fit(self, X, y=None):\n        \"\"\"\n        Fit the VotingModel.\n        \n        Parameters:\n        - X: array-like or sparse matrix of shape (n_samples, n_features)\n            The input samples.\n        - y: array-like of shape (n_samples,), default=None\n            The target values.\n            \n        Returns:\n        - self: object\n            Returns self.\n        \"\"\"\n        return self\n    \n    def predict(self, X):\n        \"\"\"\n        Predict regression target for X.\n        \n        Parameters:\n        - X: array-like or sparse matrix of shape (n_samples, n_features)\n            The input samples.\n            \n        Returns:\n        - y_preds: array-like of shape (n_samples,)\n            The predicted target values.\n        \"\"\"\n        y_preds = [estimator.predict(X) for estimator in self.estimators]\n        return np.mean(y_preds, axis=0)\n     \n    def predict_proba(self, X):      \n        \"\"\"\n        Predict class probabilities for X.\n        \n        Parameters:\n        - X: array-like or sparse matrix of shape (n_samples, n_features)\n            The input samples.\n            \n        Returns:\n        - proba: array-like of shape (n_samples, n_classes)\n            Class probabilities of the input samples.\n        \"\"\"\n        # lgb\n        y_preds = [estimator.predict_proba(X) for estimator in self.estimators[:5]]\n        \n        # cat        \n        X[cat_cols] = X[cat_cols].astype(str)\n        y_preds += [estimator.predict_proba(X) for estimator in self.estimators[-5:]]\n        \n        return np.mean(y_preds, axis=0)","metadata":{"papermill":{"duration":0.019084,"end_time":"2024-04-18T01:06:03.118007","exception":false,"start_time":"2024-04-18T01:06:03.098923","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-04-20T16:49:50.536635Z","iopub.execute_input":"2024-04-20T16:49:50.536945Z","iopub.status.idle":"2024-04-20T16:49:50.545881Z","shell.execute_reply.started":"2024-04-20T16:49:50.536919Z","shell.execute_reply":"2024-04-20T16:49:50.544997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = VotingModel(lgb_models + cat_models)\nlen(model.estimators)","metadata":{"papermill":{"duration":0.015923,"end_time":"2024-04-18T01:06:03.142213","exception":false,"start_time":"2024-04-18T01:06:03.12629","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-04-20T16:49:50.547043Z","iopub.execute_input":"2024-04-20T16:49:50.547322Z","iopub.status.idle":"2024-04-20T16:49:50.560941Z","shell.execute_reply.started":"2024-04-20T16:49:50.547297Z","shell.execute_reply":"2024-04-20T16:49:50.560135Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Predict probabilities for the test data\ny_pred = pd.Series(model.predict_proba(df_test)[:, 1], index=df_test.index)\n\n# Read the sample submission file\ndf_subm = pd.read_csv(ROOT / \"sample_submission.csv\")\ndf_subm = df_subm.set_index(\"case_id\")\n\n# Assign predicted probabilities to the submission dataframe\ndf_subm[\"score\"] = y_pred\n\n# Save the submission dataframe to a CSV file\ndf_subm.to_csv(\"submission.csv\")\n\n# Display the submission dataframe\ndf_subm","metadata":{"papermill":{"duration":0.671929,"end_time":"2024-04-18T01:06:03.821453","exception":false,"start_time":"2024-04-18T01:06:03.149524","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-04-20T16:49:50.562233Z","iopub.execute_input":"2024-04-20T16:49:50.562667Z","iopub.status.idle":"2024-04-20T16:49:51.152563Z","shell.execute_reply.started":"2024-04-20T16:49:50.562636Z","shell.execute_reply":"2024-04-20T16:49:51.151631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"papermill":{"duration":0.007392,"end_time":"2024-04-18T01:06:03.836532","exception":false,"start_time":"2024-04-18T01:06:03.82914","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]}]}