{"metadata":{"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":50160,"databundleVersionId":7921029,"sourceType":"competition"},{"sourceId":179641090,"sourceType":"kernelVersion"},{"sourceId":179741017,"sourceType":"kernelVersion"}],"dockerImageVersionId":30699,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.10.13"},"papermill":{"default_parameters":{},"duration":450.088551,"end_time":"2024-05-26T05:22:41.609561","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2024-05-26T05:15:11.521010","version":"2.5.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"\n```python\n# Handle warning messages\nimport warnings\nwarnings.filterwarnings('ignore')\n```\n- **Purpose**: Suppress warnings to keep the output clean. This is particularly useful during development to avoid clutter from non-critical warnings.\n\n```python\n# System management\nimport os\nimport gc\nimport sys\n```\n- **Purpose**: Import system management modules:\n  - `os`: Provides functions to interact with the operating system.\n  - `gc`: Provides an interface to the garbage collector, useful for managing memory.\n  - `sys`: Provides access to some variables used or maintained by the Python interpreter and to functions that interact strongly with the interpreter.\n\n```python\n# Data preprocessing\nimport numpy as np\nimport polars as pl\nimport pandas as pd\n```\n- **Purpose**: Import data preprocessing libraries:\n  - `numpy`: A fundamental package for numerical computations in Python.\n  - `polars`: A fast DataFrame library for Rust and Python.\n  - `pandas`: A powerful data manipulation and analysis library for Python.\n\n```python\n# File management\nimport joblib\nfrom glob import glob\nfrom pathlib import Path\n```\n- **Purpose**: Import file management libraries:\n  - `joblib`: Used for saving and loading Python objects, including models.\n  - `glob`: Provides a function to search for files matching a specified pattern.\n  - `Path` from `pathlib`: An object-oriented filesystem paths library.\n\n```python\n# Feature reduction\nfrom itertools import combinations, permutations\n```\n- **Purpose**: Import functions for combinatorial operations:\n  - `combinations` and `permutations`: Functions from the `itertools` module to generate combinations and permutations, useful for feature engineering and selection.\n\n```python\n# Model development\nimport lightgbm as lgb\nfrom catboost import CatBoostClassifier, Pool\nfrom sklearn.base import BaseEstimator, RegressorMixin\n```\n- **Purpose**: Import libraries and classes for model development:\n  - `lightgbm`: A gradient boosting framework that uses tree-based learning algorithms.\n  - `CatBoostClassifier` and `Pool` from `catboost`: A gradient boosting library optimized for categorical features.\n  - `BaseEstimator` and `RegressorMixin` from `sklearn.base`: Base classes from scikit-learn to create custom estimators and regressor models.\n\n### Detailed Explanations\n\n1. **Handling Warnings**:\n   ```python\n   import warnings\n   warnings.filterwarnings('ignore')\n   ```\n   - This block suppresses all warning messages. It's often used to avoid non-critical warnings cluttering the console output during development.\n\n2. **System Management Imports**:\n   ```python\n   import os\n   import gc\n   import sys\n   ```\n   - `os`: Provides functionality to interact with the operating system, such as file and directory operations.\n   - `gc`: Controls the garbage collector, which is responsible for automatic memory management in Python. Useful for managing memory, especially when dealing with large datasets.\n   - `sys`: Provides access to system-specific parameters and functions. It can be used to manipulate the Python runtime environment.\n\n3. **Data Preprocessing Imports**:\n   ```python\n   import numpy as np\n   import polars as pl\n   import pandas as pd\n   ```\n   - `numpy`: A fundamental library for numerical computations, supporting large, multi-dimensional arrays and matrices.\n   - `polars`: A fast DataFrame library for data manipulation and analysis. It is known for its performance, especially with large datasets.\n   - `pandas`: A versatile data analysis and manipulation library. It is widely used for data cleaning, preparation, and analysis.\n\n4. **File Management Imports**:\n   ```python\n   import joblib\n   from glob import glob\n   from pathlib import Path\n   ```\n   - `joblib`: Used for serializing Python objects, which is useful for saving and loading machine learning models.\n   - `glob`: Provides a function for pathname pattern matching, allowing for flexible file searching.\n   - `Path` from `pathlib`: Offers an object-oriented approach to filesystem paths, making file operations more intuitive and readable.\n\n5. **Feature Reduction Imports**:\n   ```python\n   from itertools import combinations, permutations\n   ```\n   - These functions generate combinations and permutations of input iterables. They are often used in feature engineering and selection processes to explore different feature combinations.\n\n6. **Model Development Imports**:\n   ```python\n   import lightgbm as lgb\n   from catboost import CatBoostClassifier, Pool\n   from sklearn.base import BaseEstimator, RegressorMixin\n   ```\n   - `lightgbm`: A high-performance, distributed, and efficient gradient boosting framework that uses tree-based learning algorithms.\n   - `CatBoostClassifier` and `Pool`: From the `catboost` library, designed for handling categorical features efficiently.\n   - `BaseEstimator` and `RegressorMixin`: Base classes from scikit-learn, allowing for the creation of custom estimator and regressor models that can integrate seamlessly with scikit-learn's ecosystem.","metadata":{}},{"cell_type":"code","source":"# System management\nimport os  # Import the os module for system-related operations\nimport gc  # Import the gc module for garbage collection operations\nimport sys  # Import the sys module for system-specific functions","metadata":{"papermill":{"duration":5.290874,"end_time":"2024-05-26T05:15:19.592866","exception":false,"start_time":"2024-05-26T05:15:14.301992","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-29T14:53:02.614072Z","iopub.execute_input":"2024-05-29T14:53:02.614476Z","iopub.status.idle":"2024-05-29T14:53:02.619489Z","shell.execute_reply.started":"2024-05-29T14:53:02.614445Z","shell.execute_reply":"2024-05-29T14:53:02.618426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Data preprocessing\nimport numpy as np  # Import the numpy library for numerical operations\nimport polars as pl  # Import the polars library for data manipulation\nimport pandas as pd  # Import the pandas library for data manipulation","metadata":{"execution":{"iopub.status.busy":"2024-05-29T14:53:02.621075Z","iopub.execute_input":"2024-05-29T14:53:02.621369Z","iopub.status.idle":"2024-05-29T14:53:02.630842Z","shell.execute_reply.started":"2024-05-29T14:53:02.621346Z","shell.execute_reply":"2024-05-29T14:53:02.629969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# File management\nimport joblib  # Import the joblib library for saving and loading models\nfrom glob import glob  # Import the glob function for file path pattern matching\nfrom pathlib import Path  # Import the Path class for representing file paths\n\n# Feature reduction\nfrom itertools import combinations, permutations  # Import itertools for generating combinations and permutations of elements","metadata":{"execution":{"iopub.status.busy":"2024-05-29T14:53:02.631763Z","iopub.execute_input":"2024-05-29T14:53:02.632036Z","iopub.status.idle":"2024-05-29T14:53:02.643046Z","shell.execute_reply.started":"2024-05-29T14:53:02.632014Z","shell.execute_reply":"2024-05-29T14:53:02.642191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Model development\nimport lightgbm as lgb  # Import lightgbm for gradient boosting framework\nfrom catboost import CatBoostClassifier, Pool  # Import CatBoost for gradient boosting model\nfrom sklearn.base import BaseEstimator, RegressorMixin  # Import BaseEstimator and RegressorMixin for custom model development","metadata":{"execution":{"iopub.status.busy":"2024-05-29T14:53:02.644962Z","iopub.execute_input":"2024-05-29T14:53:02.645263Z","iopub.status.idle":"2024-05-29T14:53:02.656884Z","shell.execute_reply.started":"2024-05-29T14:53:02.645218Z","shell.execute_reply":"2024-05-29T14:53:02.656062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Handle warning messages\nimport warnings  # Import the warnings module to manage warnings\nwarnings.filterwarnings('ignore')  # Ignore warning messages","metadata":{"execution":{"iopub.status.busy":"2024-05-29T14:53:02.657968Z","iopub.execute_input":"2024-05-29T14:53:02.658311Z","iopub.status.idle":"2024-05-29T14:53:02.667269Z","shell.execute_reply.started":"2024-05-29T14:53:02.658279Z","shell.execute_reply":"2024-05-29T14:53:02.666440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n1. **Paths to Data:**\n   - `train_path`: This variable holds the path to the training data in the parquet format.\n   - `test_path`: This variable holds the path to the test data in the parquet format.\n   - `subm_data`: This variable holds the path to the submission data, which typically contains a sample submission format for competitions or evaluation purposes.\n\n2. **Paths to Pretrained Models:**\n   - `lgb`: This variable holds the path to a pretrained LightGBM classifier model file (`LGB.pkl`).\n   - `ctb`: This variable holds the path to a pretrained CatBoost classifier model file (`CTB.pkl`).\n\n3. **Thresholds for Dropping Columns:**\n   - `corr_threshold`: This variable holds the threshold value for correlation between columns. If the correlation between two columns exceeds this threshold, one of them might be dropped.\n   - `drop_threshold`: This variable holds the threshold for dropping columns based on their importance.\n   - `drop_frequency`: This variable holds the frequency threshold for dropping columns.","metadata":{}},{"cell_type":"code","source":"class CFG:\n    # Paths to (.parquet) train/test data\n    train_path = Path(\"/kaggle/input/home-credit-credit-risk-model-stability/parquet_files/train\")  # Path to training data in parquet format\n    test_path = Path(\"/kaggle/input/home-credit-credit-risk-model-stability/parquet_files/test\")  # Path to test data in parquet format\n    \n    # Path to submission data\n    subm_data = \"/kaggle/input/home-credit-credit-risk-model-stability/sample_submission.csv\"  # Path to submission data\n    \n    # Paths to pretrained LightGBM and CatBoost classifiers\n    lgb = \"/kaggle/input/crms-train-lgb/LGB.pkl\"  # Path to pretrained LightGBM classifier\n    ctb = \"/kaggle/input/crms-train-ctb/CTB.pkl\"  # Path to pretrained CatBoost classifier\n    \n    # Thresholds for dropping columns\n    corr_threshold = 0.9  # Threshold for correlation, above which columns will be dropped\n    drop_threshold = 0.7  # Threshold for dropping columns based on importance\n    drop_frequency = 100  # Frequency threshold for dropping columns","metadata":{"papermill":{"duration":0.015549,"end_time":"2024-05-26T05:15:19.616388","exception":false,"start_time":"2024-05-26T05:15:19.600839","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-29T14:53:02.668302Z","iopub.execute_input":"2024-05-29T14:53:02.668646Z","iopub.status.idle":"2024-05-29T14:53:02.677360Z","shell.execute_reply.started":"2024-05-29T14:53:02.668622Z","shell.execute_reply":"2024-05-29T14:53:02.676485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n1. **`set_table_dtypes(df)`**:\n   - This method sets the data types for columns in a DataFrame (`df`) based on their names or suffixes.\n   - It iterates over each column in the DataFrame and checks if the column name is in predefined lists or if it ends with specific suffixes.\n   - Depending on the column name or suffix, it casts the column to a specific data type using Polars library's `cast()` method.\n   - After iterating over all columns, it returns the modified DataFrame.\n\n2. **`handle_dates(df)`**:\n   - This method handles date-related columns in the DataFrame.\n   - It iterates over each column in the DataFrame and checks if the column name ends with a specific suffix (\"D\" indicating date-related columns).\n   - For such columns, it calculates the difference between the values in that column and the 'date_decision' column, and then converts the result to total days.\n   - It drops unnecessary date-related columns ('date_decision' and 'MONTH') from the DataFrame.\n   - Finally, it returns the modified DataFrame.\n\n3. **`filter_cols(df)`**:\n   - This method filters columns based on specified conditions.\n   - It first iterates over each column and checks if it's not in a predefined list of essential columns (\"target\", \"case_id\", \"WEEK_NUM\").\n   - For such columns, it calculates the proportion of null values (`isnull`) and drops the column if the proportion exceeds the threshold specified in `CFG.drop_threshold`.\n   - Then, it iterates over columns again and checks if the column is not in the essential list and is of type String. For such columns, it counts the unique values (`freq`) and drops the column if the count is 1 or exceeds the frequency threshold specified in `CFG.drop_frequency`.\n   - Finally, it returns the modified DataFrame.\n","metadata":{}},{"cell_type":"code","source":"class Pipeline:\n    @staticmethod\n    def set_table_dtypes(df):\n        # Set data types for columns based on their names\n        for col in df.columns:\n            if col in [\"case_id\", \"WEEK_NUM\", \"num_group1\", \"num_group2\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Int32))  # Convert specified columns to Int32 type\n            elif col in [\"date_decision\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Date))  # Convert 'date_decision' column to Date type\n            elif col.endswith((\"P\", \"A\")):\n                df = df.with_columns(pl.col(col).cast(pl.Float32))  # Convert columns ending with 'P' or 'A' to Float32 type\n            elif col.endswith((\"M\")):\n                df = df.with_columns(pl.col(col).cast(pl.String))  # Convert columns ending with 'M' to String type\n            elif col.endswith((\"D\")):\n                df = df.with_columns(pl.col(col).cast(pl.Date))  # Convert columns ending with 'D' to Date type\n                \n        return df  # Return the modified DataFrame\n\n    @staticmethod\n    def handle_dates(df):\n        # Handle date-related columns\n        for col in df.columns:\n            if col.endswith((\"D\")):\n                df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))  # Calculate the difference between column and 'date_decision'\n                df = df.with_columns(pl.col(col).dt.total_days())  # Convert the column to total days\n                \n        # Drop unnecessary date-related columns\n        df = df.drop(\"date_decision\", \"MONTH\")  # Drop 'date_decision' and 'MONTH' columns\n        \n        return df  # Return the modified DataFrame\n\n    @staticmethod\n    def filter_cols(df):\n        # Filter columns based on specified conditions\n        for col in df.columns:\n            if col not in [\"target\", \"case_id\", \"WEEK_NUM\"]:\n                isnull = df[col].is_null().mean()  # Calculate the proportion of null values\n                if isnull > CFG.drop_threshold:  # If proportion exceeds the threshold\n                    df = df.drop(col)  # Drop the column\n        \n        for col in df.columns:\n            if (col not in [\"target\", \"case_id\", \"WEEK_NUM\"]) & (df[col].dtype == pl.String):\n                freq = df[col].n_unique()  # Count the unique values\n                if (freq == 1) | (freq > CFG.drop_frequency):  # If frequency is 1 or exceeds frequency threshold\n                    df = df.drop(col)  # Drop the column\n        \n        return df  # Return the modified DataFrame","metadata":{"papermill":{"duration":0.022214,"end_time":"2024-05-26T05:15:19.646123","exception":false,"start_time":"2024-05-26T05:15:19.623909","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-29T14:53:02.679000Z","iopub.execute_input":"2024-05-29T14:53:02.679421Z","iopub.status.idle":"2024-05-29T14:53:02.692564Z","shell.execute_reply.started":"2024-05-29T14:53:02.679391Z","shell.execute_reply":"2024-05-29T14:53:02.691579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\n1. **`numeric_group_expr(df, N=200)`**:\n   - This method aggregates numeric transformation columns.\n   - It first identifies numeric columns by selecting those ending with 'P' or 'A'.\n   - For each selected numeric column, it calculates various statistics such as differences, means, mins, maxs, and sums.\n   - Specifically, it calculates:\n     - The difference between the max and min values of the last `N` non-null values (`diff1_{col}`).\n     - The difference between the max and min values of the first `N` non-null values (`diff2_{col}`).\n     - The difference between the mean of the last `N` non-null values and the mean of the first `N` non-null values (`diff3_{col}`).\n     - The difference between the sum of the last `N` non-null values and the sum of the first `N` non-null values (`diff4_{col}`).\n     - The difference between the last and first non-null values (`diff5_{col}`).\n     - The last non-null value (`last_{col}`).\n     - The difference between the last non-null value and the mean of all values (`diff6_{col}`).\n     - The mean, min, max, and sum of all values (`mean_{col}`, `min_{col}`, `max_{col}`, `sum_{col}`).\n   - It returns a list of these aggregation expressions.\n\n2. **`temporal_group_expr(df)`**:\n   - This method aggregates temporal transformation columns.\n   - It identifies temporal columns by selecting those ending with 'D'.\n   - For each selected temporal column, it calculates the minimum and maximum values of all non-null values.\n   - It returns a list of these aggregation expressions.\n\n3. **`other_group_expr(df)`**:\n   - This method aggregates other (categorical and unspecified) transformation columns.\n   - It identifies such columns by selecting those ending with 'M', 'T', or 'L'.\n   - For each selected column, it extracts the first and last non-null values.\n   - It returns a list of these aggregation expressions.\n\n4. **`num_group_expr(df)`**:\n   - This method aggregates num_group expressions.\n   - It selects columns containing 'num_group'.\n   - For each selected column, it calculates the maximum value.\n   - It returns a list of these aggregation expressions.\n\n5. **`get_exprs(df)`**:\n   - This method concatenates all the aggregation expressions obtained from the previous methods.\n   - It returns a single list containing all aggregation expressions for the DataFrame `df`.\n","metadata":{}},{"cell_type":"code","source":"class Aggregator:\n    \n    # Aggregate numeric transformation columns\n    @staticmethod\n    def numeric_group_expr(df, N=200):\n        cols = [col for col in df.columns if col.endswith((\"P\", \"A\"))]   # Select numeric columns ending with 'P' or 'A'\n        exprs = []\n        for col in cols:\n            exprs.extend([\n                # Extract the difference between the max of the last N non-null values and the min of the last N non-null values\n                (pl.col(col).filter(pl.col(col).is_not_null()).slice(-N, None).max() - \n                 pl.col(col).filter(pl.col(col).is_not_null()).slice(-N, None).min()).alias(f\"diff1_{col}\"),\n                \n                # Extract the difference between the max of the first N non-null values and the min of the first N non-null values\n                (pl.col(col).filter(pl.col(col).is_not_null()).slice(0, N).max() -\n                 pl.col(col).filter(pl.col(col).is_not_null()).slice(0, N).min()).alias(f\"diff2_{col}\"),\n                \n                # Extract the difference between the mean of the last N non-null values and the mean of the first N non-null values\n                (pl.col(col).filter(pl.col(col).is_not_null()).slice(-N, None).mean() -\n                 pl.col(col).filter(pl.col(col).is_not_null()).slice(0, N).mean()).alias(f\"diff3_{col}\"),\n                \n                # Extract the difference between the sum of the last N non-null values and the sum of the first N non-null values\n                (pl.col(col).filter(pl.col(col).is_not_null()).slice(-N, None).sum() -\n                 pl.col(col).filter(pl.col(col).is_not_null()).slice(0, N).sum()).alias(f\"diff4_{col}\"),\n                \n                # Extract the difference between the last and first non-null values\n                (pl.col(col).filter(pl.col(col).is_not_null()).last() -\n                 pl.col(col).filter(pl.col(col).is_not_null()).first()).alias(f\"diff5_{col}\"),\n                \n                # Extract the last non-null value\n                pl.col(col).filter(pl.col(col).is_not_null()).last().alias(f\"last_{col}\"),\n                \n                # Extract the difference between the last non-null value and the mean of all values\n                (pl.col(col).filter(pl.col(col).is_not_null()).last() -\n                 pl.col(col).mean()).alias(f\"diff6_{col}\"),\n                \n                # Extract the mean, min, max and sum of all values\n                pl.col(col).mean().alias(f\"mean_{col}\"),\n                pl.col(col).min().alias(f\"min_{col}\"),\n                pl.col(col).max().alias(f\"max_{col}\"),\n                pl.col(col).sum().alias(f\"sum_{col}\")\n            ])\n            \n        return exprs\n    \n    # Aggregate temporal transformation columns\n    @staticmethod\n    def temporal_group_expr(df):\n        cols = [col for col in df.columns if col.endswith(\"D\")]  # Select temporal columns ending with 'D'\n        exprs = []\n        for col in cols:\n            exprs.extend([\n                # Extract the min and max of all non-null values\n                pl.col(col).filter(pl.col(col).is_not_null()).min().alias(f\"min_{col}\"),\n                pl.col(col).filter(pl.col(col).is_not_null()).max().alias(f\"max_{col}\")\n            ])\n        \n        return exprs\n    \n    # Aggregate other (categorical and unspecified) transformation columns\n    @staticmethod\n    def other_group_expr(df):\n        cols = [col for col in df.columns if col.endswith((\"M\", \"T\", \"L\"))]  # Select other columns ending with 'M', 'T', or 'L'\n        exprs = []\n        for col in cols:\n            exprs.extend([     \n                # Extract the first and last non-null values\n                pl.col(col).filter(pl.col(col).is_not_null()).first().alias(f\"first_{col}\"),\n                pl.col(col).filter(pl.col(col).is_not_null()).last().alias(f\"last_{col}\")\n            ])\n        \n        return exprs\n    \n    # Aggregate num_group expressions\n    @staticmethod\n    def num_group_expr(df):\n        cols = [col for col in df.columns if \"num_group\" in col]  # Select columns containing 'num_group'\n        exprs = []\n        for col in cols:\n            exprs.extend([      \n                # Extract the max of all values\n                pl.col(col).max().alias(f\"max_{col}\")\n            ])\n        \n        return exprs\n    \n    # Concatenate all aggregation expressions\n    @staticmethod\n    def get_exprs(df):\n        exprs = []\n        exprs.extend(Aggregator.numeric_group_expr(df))  # Numeric group expressions\n        exprs.extend(Aggregator.temporal_group_expr(df))  # Temporal group expressions\n        exprs.extend(Aggregator.other_group_expr(df))  # Other group expressions\n        exprs.extend(Aggregator.num_group_expr(df))  # Num group expressions\n\n        return exprs\n","metadata":{"papermill":{"duration":0.03145,"end_time":"2024-05-26T05:15:19.685411","exception":false,"start_time":"2024-05-26T05:15:19.653961","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-29T14:53:02.846103Z","iopub.execute_input":"2024-05-29T14:53:02.846459Z","iopub.status.idle":"2024-05-29T14:53:02.870043Z","shell.execute_reply.started":"2024-05-29T14:53:02.846432Z","shell.execute_reply":"2024-05-29T14:53:02.868973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\n1. **`read_file(self, path, depth=None)`**:\n   - This method reads data from a single file specified by the `path` parameter.\n   - It reads the data into a Polars DataFrame using the `pl.read_parquet()` function.\n   - It applies data type setting using the `Pipeline.set_table_dtypes` method.\n   - If `depth` is specified (either 1 or 2), it aggregates the data based on the \"case_id\" column using the `Aggregator.get_exprs()` method.\n   - It returns the processed DataFrame.\n\n2. **`read_files(self, regex_path, depth=None)`**:\n   - This method reads data from multiple files matching the regex pattern specified by `regex_path`.\n   - It iterates through each file, reads it into a Polars DataFrame, sets data types using `Pipeline.set_table_dtypes`, and applies depth-based aggregation if specified.\n   - It concatenates all the chunks of data vertically using `pl.concat()` and ensures uniqueness based on the \"case_id\" column.\n   - It returns the concatenated and unique DataFrame.\n\n3. **`feature_eng(self, df_base, depth_0, depth_1, depth_2)`**:\n   - This method performs feature engineering on a base DataFrame (`df_base`) by adding additional columns such as `month_decision` and `weekday_decision`.\n   - It then joins additional DataFrames based on \"case_id\" from the `depth_0`, `depth_1`, and `depth_2` lists.\n   - It handles date columns using the `Pipeline.handle_dates` method.\n   - It returns the DataFrame with the engineered features.\n\n4. **`to_pandas(self, df_data, cat_cols=None)`**:\n   - This method converts a Polars DataFrame `df_data` to a Pandas DataFrame.\n   - Optionally, it converts specified columns to categorical type (if `cat_cols` is provided).\n   - It returns the converted DataFrame and the list of categorical columns.\n\n5. **`reduce_mem_usage(self, df)`**:\n   - This method optimizes the memory usage of a DataFrame by downcasting numeric data types to the smallest possible types.\n   - It iterates over each column, checks its data type, and downcasts it accordingly.\n   - It prints the memory usage before and after optimization.\n   - It returns the DataFrame with optimized memory usage.\n","metadata":{}},{"cell_type":"code","source":"class DataUtilities:\n    def read_file(self, path, depth=None):\n        # Read data from a single file\n        df = pl.read_parquet(path, low_memory=True)  # Read data into a Polars DataFrame\n        \n        # Set data types using custom pipeline\n        df = df.pipe(Pipeline.set_table_dtypes)  # Apply data type setting\n        \n        # Apply depth-based aggregation if specified\n        if depth in [1, 2]:\n            df = df.group_by(\"case_id\").agg(Aggregator.get_exprs(df))  # Aggregate based on \"case_id\"\n        \n        return df  # Return the processed DataFrame\n\n    def read_files(self, regex_path, depth=None):\n        # Read data from multiple files matching the regex pattern\n        chunks = []\n\n        for path in glob(str(regex_path)):\n            df = pl.read_parquet(path, low_memory=True)  # Read data into a Polars DataFrame\n            \n            # Set data types using custom pipeline\n            df = df.pipe(Pipeline.set_table_dtypes)  # Apply data type setting\n            \n            # Apply depth-based aggregation if specified\n            if depth in [1, 2]:\n                df = df.group_by(\"case_id\").agg(Aggregator.get_exprs(df))  # Aggregate based on \"case_id\"\n                \n            chunks.append(df)\n\n        # Concatenate data chunks and ensure uniqueness based on \"case_id\"\n        df = pl.concat(chunks, how=\"vertical_relaxed\")  # Concatenate chunks vertically\n        df = df.unique(subset=[\"case_id\"])  # Ensure uniqueness based on \"case_id\"\n        \n        return df  # Return the concatenated and unique DataFrame\n\n    def feature_eng(self, df_base, depth_0, depth_1, depth_2):\n        # Perform feature engineering on the base DataFrame\n        df_base = (\n            df_base\n            .with_columns(\n                month_decision = pl.col(\"date_decision\").dt.month(),  # Extract month from \"date_decision\"\n                weekday_decision = pl.col(\"date_decision\").dt.weekday(),  # Extract weekday from \"date_decision\"\n            )\n        )\n        \n        # Join additional dataframes based on \"case_id\"\n        for i, df in enumerate(depth_0 + depth_1 + depth_2):\n            df_base = df_base.join(df, how=\"left\", on=\"case_id\", suffix=f\"_{i}\")  # Join based on \"case_id\"\n        \n        # Handle date columns using custom pipeline\n        df_base = df_base.pipe(Pipeline.handle_dates)  # Handle date columns\n        \n        return df_base  # Return the DataFrame with engineered features\n\n    def to_pandas(self, df_data, cat_cols=None):\n        # Convert DataFrame to Pandas format\n        df_data = df_data.to_pandas()  # Convert to Pandas DataFrame\n        \n        # Convert specified columns to categorical type\n        if cat_cols is None:\n            cat_cols = list(df_data.select_dtypes(\"object\").columns)  # Select categorical columns if not specified\n        df_data[cat_cols] = df_data[cat_cols].astype(\"category\")  # Convert to categorical type\n        \n        return df_data, cat_cols  # Return the converted DataFrame and the list of categorical columns\n    \n    def reduce_mem_usage(self, df):       \n        start_mem = df.memory_usage().sum() / 1024**2  # Initial memory usage\n        \n        # Iterate over each column\n        for col in df.columns:\n            col_type = df[col].dtype\n\n            # Skip categorical columns\n            if str(col_type) == \"category\":\n                continue\n\n            # Optimize integer columns\n            if col_type != object and str(col_type)[:3] == 'int':\n                c_min = df[col].min()\n                c_max = df[col].max()\n                if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                    df[col] = df[col].astype(np.int8)\n                elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                    df[col] = df[col].astype(np.int16)\n                elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                    df[col] = df[col].astype(np.int32)\n\n            # Optimize float columns\n            elif col_type != object:\n                c_min = df[col].min()\n                c_max = df[col].max()\n                if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                    df[col] = df[col].astype(np.float16)\n                elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                    df[col] = df[col].astype(np.float32)\n\n        end_mem = df.memory_usage().sum() / 1024**2  # Final memory usage\n        print('Memory usage after optimization is: {:.2f} MB'.format(end_mem))\n        print('Decreased by {:.1f}%'.format(100 * (start_mem - end_mem) / start_mem))\n\n        return df  # Return the DataFrame with optimized memory usage\n","metadata":{"papermill":{"duration":0.029915,"end_time":"2024-05-26T05:15:19.722948","exception":false,"start_time":"2024-05-26T05:15:19.693033","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-29T14:53:02.871806Z","iopub.execute_input":"2024-05-29T14:53:02.872089Z","iopub.status.idle":"2024-05-29T14:53:02.892971Z","shell.execute_reply.started":"2024-05-29T14:53:02.872064Z","shell.execute_reply":"2024-05-29T14:53:02.892102Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"du = DataUtilities()","metadata":{"papermill":{"duration":0.022401,"end_time":"2024-05-26T05:15:19.754105","exception":false,"start_time":"2024-05-26T05:15:19.731704","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-29T14:53:02.893962Z","iopub.execute_input":"2024-05-29T14:53:02.894205Z","iopub.status.idle":"2024-05-29T14:53:02.905311Z","shell.execute_reply.started":"2024-05-29T14:53:02.894184Z","shell.execute_reply":"2024-05-29T14:53:02.904446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n```python\nclass FeatureReduction:\n    # Initialize the FeatureReduction object with a correlation threshold\n    def __init__(self, threshold):\n        self.threshold = threshold\n```\n- This line defines the `FeatureReduction` class.\n- The `__init__` method is a constructor method that initializes objects of this class.\n- It takes a parameter `threshold` to set the correlation threshold for feature reduction.\n- `self.threshold` stores the correlation threshold value for the object.\n\n```python\n    # Group columns based on the count of NaN values in each column\n    def _group_columns_by_nan_count(self, df):\n        # Select numeric columns\n        nums = df.select_dtypes(exclude='category').columns\n```\n- This method `_group_columns_by_nan_count` groups columns based on the count of NaN (missing) values in each column.\n- It takes a DataFrame `df` as input.\n- It selects numeric columns by excluding columns of type 'category'.\n\n```python\n        # Create a DataFrame indicating NaN values\n        nans_df = df[nums].isna()\n        nans_groups = {}\n```\n- It creates a new DataFrame `nans_df` indicating NaN values in the selected numeric columns.\n- `nans_groups` is initialized as an empty dictionary to store the groups of columns based on NaN counts.\n\n```python\n        # Iterate over numeric columns\n        for col in nums:\n            # Count the number of NaN values in each column\n            cur_group = nans_df[col].sum()\n            try:\n                # Append the column to the corresponding group based on the count of NaN values\n                nans_groups[cur_group].append(col)\n            except:\n                nans_groups[cur_group] = [col]\n```\n- It iterates over the numeric columns.\n- For each column, it calculates the count of NaN values (`cur_group`).\n- It appends the column to the corresponding group in `nans_groups` based on the count of NaN values.\n\n```python\n        # Clean up and release memory\n        del nans_df\n        gc.collect()\n```\n- This code block cleans up the temporary DataFrame `nans_df` and releases memory.\n\n```python\n        return nans_groups\n```\n- The method returns `nans_groups`, which contains groups of columns based on NaN counts.\n\n```python\n    # Reduce each group of columns to a single column based on the one with the highest number of unique values\n    def _reduce_group(self, grps, df):\n        # Initialize the list to store the selected columns\n        use = []\n```\n- This method `_reduce_group` reduces each group of columns to a single column based on the one with the highest number of unique values.\n- It takes `grps` (groups of columns) and `df` (DataFrame) as inputs.\n- It initializes an empty list `use` to store the selected columns.\n\n```python\n        # Iterate over each group of columns\n        for g in grps:\n            mx = 0\n            vx = g[0]\n```\n- It iterates over each group of columns (`g`) within `grps`.\n- `mx` is initialized to store the maximum number of unique values, and `vx` is initialized to store the column with the highest number of unique values.\n\n```python\n            # Find the column with the highest number of unique values in the group\n            for gg in g:\n                n = df[gg].nunique()\n                if n > mx:\n                    mx = n\n                    vx = gg\n```\n- It iterates over each column (`gg`) within the group (`g`).\n- For each column, it calculates the number of unique values (`n`).\n- If the number of unique values is greater than the current maximum (`mx`), it updates `mx` and assigns the column to `vx`.\n\n```python\n            # Append the selected column to the list\n            use.append(vx)\n```\n- After finding the column with the highest number of unique values in the group, it appends the column (`vx`) to the list `use`.\n\n```python\n        return use\n```\n- The method returns the list `use`, which contains the selected columns from each group.\n\n```python\n    # Group columns based on their correlation with a given threshold\n    def _group_columns_by_correlation(self, matrix):\n        # Calculate the correlation matrix\n        correlation_matrix = matrix.corr()\n```\n- This method `_group_columns_by_correlation` groups columns based on their correlation with a given threshold.\n- It takes a DataFrame `matrix` as input.\n- It calculates the correlation matrix using the `corr()` method.\n\n```python\n        # Initialize the list to store groups of correlated columns\n        groups = []\n        \n        # Initialize the list of remaining columns\n        remaining_cols = list(matrix.columns)\n```\n- It initializes an empty list `groups` to store groups of correlated columns.\n- `remaining_cols` is initialized to store the list of column names from the correlation matrix.\n\n```python\n        # Iterate until all columns are assigned to a group\n        while remaining_cols:\n            # Pop the first column from the list\n            col = remaining_cols.pop(0)\n```\n- This loop iterates until all columns are assigned to a group.\n- It removes and returns the first column name from `remaining_cols` and assigns it to `col`.\n\n","metadata":{}},{"cell_type":"code","source":"class FeatureReduction:\n    # Initialize the FeatureReduction object with a correlation threshold\n    def __init__(self, threshold):\n        self.threshold = threshold\n\n    # Group columns based on the count of NaN values in each column\n    def _group_columns_by_nan_count(self, df):\n        # Select numeric columns\n        nums = df.select_dtypes(exclude='category').columns\n        \n        # Create a DataFrame indicating NaN values\n        nans_df = df[nums].isna()\n        nans_groups = {}\n        \n        # Iterate over numeric columns\n        for col in nums:\n            # Count the number of NaN values in each column\n            cur_group = nans_df[col].sum()\n            try:\n                # Append the column to the corresponding group based on the count of NaN values\n                nans_groups[cur_group].append(col)\n            except:\n                nans_groups[cur_group] = [col]\n        \n        # Clean up and release memory\n        del nans_df\n        gc.collect()\n        \n        return nans_groups\n\n    # Reduce each group of columns to a single column based on the one with the highest number of unique values\n    def _reduce_group(self, grps, df):\n        # Initialize the list to store the selected columns\n        use = []\n        \n        # Iterate over each group of columns\n        for g in grps:\n            mx = 0\n            vx = g[0]\n            \n            # Find the column with the highest number of unique values in the group\n            for gg in g:\n                n = df[gg].nunique()\n                if n > mx:\n                    mx = n\n                    vx = gg\n                    \n            # Append the selected column to the list\n            use.append(vx)\n            \n        return use\n\n    # Group columns based on their correlation with a given threshold\n    def _group_columns_by_correlation(self, matrix):\n        # Calculate the correlation matrix\n        correlation_matrix = matrix.corr()\n        \n        # Initialize the list to store groups of correlated columns\n        groups = []\n        \n        # Initialize the list of remaining columns\n        remaining_cols = list(matrix.columns)\n        \n        # Iterate until all columns are assigned to a group\n        while remaining_cols:\n            # Pop the first column from the list\n            col = remaining_cols.pop(0)\n            \n            # Create a new group with the popped column\n            group = [col]\n            \n            # Initialize the list of correlated columns\n            correlated_cols = [col]\n            \n            # Iterate over the remaining columns to find correlated ones\n            for c in remaining_cols:\n                # Check if the correlation between the current column and the popped column is above the threshold\n                if correlation_matrix.loc[col, c] >= self.threshold:\n                    # Add the correlated column to the group and the list of correlated columns\n                    group.append(c)\n                    correlated_cols.append(c)\n                    \n            # Append the group to the list of groups\n            groups.append(group)\n            \n            # Update the list of remaining columns by removing those that are already in a group\n            remaining_cols = [c for c in remaining_cols if c not in correlated_cols]\n            \n        return groups\n\n    # Reduce features based on groups of columns with similar NaN counts and correlated columns\n    def reduce_features(self, df):\n        # Group columns based on NaN counts\n        nans_groups = self._group_columns_by_nan_count(df)\n        \n        # Initialize a list to store selected features\n        uses = []\n        \n        # Iterate over the groups of columns with similar NaN counts\n        for k, v in nans_groups.items():\n            # Check if there are multiple columns in the group\n            if len(v) > 1:\n                # Extract the group of columns\n                Vs = nans_groups[k]\n                \n                # Group columns within the group based on correlation\n                grps = self._group_columns_by_correlation(df[Vs])\n                \n                # Reduce the group to a single representative column\n                use = self._reduce_group(grps, df)\n                \n                # Append the selected column to the list of features\n                uses = uses + use\n            else:\n                # If there is only one column in the group, directly append it to the list of features\n                uses = uses + v\n                \n        # Include categorical columns in the selected features\n        uses = uses + list(df.select_dtypes(include='category').columns)\n        \n        # Return the DataFrame with reduced features\n        return df[uses]\n","metadata":{"papermill":{"duration":0.035536,"end_time":"2024-05-26T05:15:19.800493","exception":false,"start_time":"2024-05-26T05:15:19.764957","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-29T14:53:02.908135Z","iopub.execute_input":"2024-05-29T14:53:02.908698Z","iopub.status.idle":"2024-05-29T14:53:02.924565Z","shell.execute_reply.started":"2024-05-29T14:53:02.908672Z","shell.execute_reply":"2024-05-29T14:53:02.923682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fr = FeatureReduction(threshold=CFG.corr_threshold)","metadata":{"papermill":{"duration":0.01701,"end_time":"2024-05-26T05:15:19.826952","exception":false,"start_time":"2024-05-26T05:15:19.809942","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-29T14:53:02.925634Z","iopub.execute_input":"2024-05-29T14:53:02.925898Z","iopub.status.idle":"2024-05-29T14:53:02.936779Z","shell.execute_reply.started":"2024-05-29T14:53:02.925875Z","shell.execute_reply":"2024-05-29T14:53:02.935883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\n```python\ntrain_data_store = {\n    \"df_base\": du.read_file(CFG.train_path / \"train_base.parquet\"),\n```\n- `df_base`: Represents the base dataframe for training, read from the \"train_base.parquet\" file.\n\n```python\n    \"depth_0\": [\n        du.read_file(CFG.train_path / \"train_static_cb_0.parquet\"),\n        du.read_files(CFG.train_path / \"train_static_0_*.parquet\"),\n    ],\n```\n- `depth_0`: Represents the first level of depth. It includes:\n  - Data from \"train_static_cb_0.parquet\".\n  - Data from multiple files matching the pattern \"train_static_0_*.parquet\".\n\n```python\n    \"depth_1\": [\n        du.read_files(CFG.train_path / \"train_applprev_1_*.parquet\", 1),\n        du.read_file(CFG.train_path / \"train_tax_registry_a_1.parquet\", 1),\n        du.read_file(CFG.train_path / \"train_tax_registry_b_1.parquet\", 1),\n        du.read_file(CFG.train_path / \"train_tax_registry_c_1.parquet\", 1),\n        du.read_files(CFG.train_path / \"train_credit_bureau_a_1_*.parquet\", 1),\n        du.read_file(CFG.train_path / \"train_credit_bureau_b_1.parquet\", 1),\n        du.read_file(CFG.train_path / \"train_other_1.parquet\", 1),\n        du.read_file(CFG.train_path / \"train_person_1.parquet\", 1),\n        du.read_file(CFG.train_path / \"train_deposit_1.parquet\", 1),\n        du.read_file(CFG.train_path / \"train_debitcard_1.parquet\", 1),\n    ],\n```\n- `depth_1`: Represents the second level of depth. It includes various types of data:\n  - Data from multiple files matching the pattern \"train_applprev_1_*.parquet\".\n  - Data from individual files like \"train_tax_registry_a_1.parquet\", \"train_tax_registry_b_1.parquet\", etc., each representing a specific category.\n  - Data from multiple files matching the pattern \"train_credit_bureau_a_1_*.parquet\".\n  - Other specific data files like \"train_other_1.parquet\", \"train_person_1.parquet\", etc.\n\n```python\n    \"depth_2\": [\n        du.read_file(CFG.train_path / \"train_credit_bureau_b_2.parquet\", 2),\n        du.read_files(CFG.train_path / \"train_credit_bureau_a_2_*.parquet\", 2),\n        du.read_file(CFG.train_path / \"train_applprev_2.parquet\", 2),\n        du.read_file(CFG.train_path / \"train_person_2.parquet\", 2)\n    ]\n}\n```\n- `depth_2`: Represents the third level of depth. It includes:\n  - Data from specific files like \"train_credit_bureau_b_2.parquet\", \"train_applprev_2.parquet\", etc.\n  - Data from multiple files matching the pattern \"train_credit_bureau_a_2_*.parquet\".\n\nOverall, `train_data_store` stores various levels of training data, each fetched from different files or sets of files, organized based on their depth in the feature hierarchy.","metadata":{}},{"cell_type":"code","source":"train_data_store = {\n    # Base dataframe for training\n    \"df_base\": du.read_file(CFG.train_path / \"train_base.parquet\"),\n\n    # First level of depth\n    \"depth_0\": [\n        # Data from \"train_static_cb_0.parquet\"\n        du.read_file(CFG.train_path / \"train_static_cb_0.parquet\"),\n        # Data from multiple files matching the pattern \"train_static_0_*.parquet\"\n        du.read_files(CFG.train_path / \"train_static_0_*.parquet\"),\n    ],\n\n    # Second level of depth\n    \"depth_1\": [\n        # Data from multiple files matching the pattern \"train_applprev_1_*.parquet\"\n        du.read_files(CFG.train_path / \"train_applprev_1_*.parquet\", 1),\n        # Data from \"train_tax_registry_a_1.parquet\"\n        du.read_file(CFG.train_path / \"train_tax_registry_a_1.parquet\", 1),\n        # Data from \"train_tax_registry_b_1.parquet\"\n        du.read_file(CFG.train_path / \"train_tax_registry_b_1.parquet\", 1),\n        # Data from \"train_tax_registry_c_1.parquet\"\n        du.read_file(CFG.train_path / \"train_tax_registry_c_1.parquet\", 1),\n        # Data from multiple files matching the pattern \"train_credit_bureau_a_1_*.parquet\"\n        du.read_files(CFG.train_path / \"train_credit_bureau_a_1_*.parquet\", 1),\n        # Data from \"train_credit_bureau_b_1.parquet\"\n        du.read_file(CFG.train_path / \"train_credit_bureau_b_1.parquet\", 1),\n        # Data from \"train_other_1.parquet\"\n        du.read_file(CFG.train_path / \"train_other_1.parquet\", 1),\n        # Data from \"train_person_1.parquet\"\n        du.read_file(CFG.train_path / \"train_person_1.parquet\", 1),\n        # Data from \"train_deposit_1.parquet\"\n        du.read_file(CFG.train_path / \"train_deposit_1.parquet\", 1),\n        # Data from \"train_debitcard_1.parquet\"\n        du.read_file(CFG.train_path / \"train_debitcard_1.parquet\", 1),\n    ],\n\n    # Third level of depth\n    \"depth_2\": [\n        # Data from \"train_credit_bureau_b_2.parquet\"\n        du.read_file(CFG.train_path / \"train_credit_bureau_b_2.parquet\", 2),\n        # Data from multiple files matching the pattern \"train_credit_bureau_a_2_*.parquet\"\n        du.read_files(CFG.train_path / \"train_credit_bureau_a_2_*.parquet\", 2),\n        # Data from \"train_applprev_2.parquet\"\n        du.read_file(CFG.train_path / \"train_applprev_2.parquet\", 2),\n        # Data from \"train_person_2.parquet\"\n        du.read_file(CFG.train_path / \"train_person_2.parquet\", 2)\n    ]\n}","metadata":{"papermill":{"duration":256.873114,"end_time":"2024-05-26T05:19:36.712612","exception":false,"start_time":"2024-05-26T05:15:19.839498","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-29T14:53:02.937934Z","iopub.execute_input":"2024-05-29T14:53:02.938274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = du.feature_eng(**train_data_store)\nprint(f\"The shape of the train data is: {train_data.shape}\")","metadata":{"papermill":{"duration":17.325503,"end_time":"2024-05-26T05:19:54.046770","exception":false,"start_time":"2024-05-26T05:19:36.721267","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del train_data_store\ngc.collect()","metadata":{"papermill":{"duration":0.709221,"end_time":"2024-05-26T05:19:54.764898","exception":false,"start_time":"2024-05-26T05:19:54.055677","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = train_data.pipe(Pipeline.filter_cols)","metadata":{"papermill":{"duration":5.879436,"end_time":"2024-05-26T05:20:00.652859","exception":false,"start_time":"2024-05-26T05:19:54.773423","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data, cat_cols = du.to_pandas(train_data)","metadata":{"papermill":{"duration":33.32419,"end_time":"2024-05-26T05:20:33.985772","exception":false,"start_time":"2024-05-26T05:20:00.661582","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = du.reduce_mem_usage(train_data)","metadata":{"papermill":{"duration":13.783018,"end_time":"2024-05-26T05:20:47.777666","exception":false,"start_time":"2024-05-26T05:20:33.994648","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"The shape of the train data after filtering is: {train_data.shape}\")","metadata":{"papermill":{"duration":0.016212,"end_time":"2024-05-26T05:20:47.802831","exception":false,"start_time":"2024-05-26T05:20:47.786619","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = fr.reduce_features(train_data)","metadata":{"papermill":{"duration":104.59587,"end_time":"2024-05-26T05:22:32.407347","exception":false,"start_time":"2024-05-26T05:20:47.811477","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\n```python\ntest_data_store = {\n```\nThis line initializes the `test_data_store` dictionary.\n\n```python\n    \"df_base\": du.read_file(CFG.test_path / \"test_base.parquet\"),\n```\nThis line reads the base testing data from the \"test_base.parquet\" file and stores it in a dataframe under the key \"df_base\".\n\n```python\n    \"depth_0\": [\n```\nThis line begins a list for the first level of depth in the feature hierarchy.\n\n```python\n        du.read_file(CFG.test_path / \"test_static_cb_0.parquet\"),\n```\nThis line reads data from the \"test_static_cb_0.parquet\" file and adds it to the list under the first level of depth.\n\n```python\n        du.read_files(CFG.test_path / \"test_static_0_*.parquet\"),\n```\nThis line reads data from multiple files matching the pattern \"test_static_0_*.parquet\" and adds them to the list under the first level of depth.\n\n```python\n    ],\n```\nThis line closes the list for the first level of depth.\n\n```python\n    \"depth_1\": [\n```\nThis line begins a list for the second level of depth in the feature hierarchy.\n\n```python\n        du.read_files(CFG.test_path / \"test_applprev_1_*.parquet\", 1),\n```\nThis line reads data from multiple files matching the pattern \"test_applprev_1_*.parquet\" with depth 1 and adds them to the list under the second level of depth.\n\n```python\n        du.read_file(CFG.test_path / \"test_tax_registry_a_1.parquet\", 1),\n```\nThis line reads data from the \"test_tax_registry_a_1.parquet\" file with depth 1 and adds it to the list under the second level of depth.\n\n```python\n        ...\n```\nSimilar lines follow, reading data from different files with depth 1 and adding them to the list under the second level of depth.\n\n```python\n    ],\n```\nThis line closes the list for the second level of depth.\n\n```python\n    \"depth_2\": [\n```\nThis line begins a list for the third level of depth in the feature hierarchy.\n\n```python\n        du.read_file(CFG.test_path / \"test_credit_bureau_b_2.parquet\", 2),\n```\nThis line reads data from the \"test_credit_bureau_b_2.parquet\" file with depth 2 and adds it to the list under the third level of depth.\n\n```python\n        ...\n```\nSimilar lines follow, reading data from different files with depth 2 and adding them to the list under the third level of depth.\n\n```python\n    ]\n}\n```\nThis line closes the dictionary for `test_data_store`. \n\nOverall, `test_data_store` organizes testing data into different levels of depth, with each level containing data from specific files or patterns of files.","metadata":{}},{"cell_type":"code","source":"test_data_store = {\n    # Base dataframe for testing\n    \"df_base\": du.read_file(CFG.test_path / \"test_base.parquet\"),\n\n    # First level of depth\n    \"depth_0\": [\n        # Data from \"test_static_cb_0.parquet\"\n        du.read_file(CFG.test_path / \"test_static_cb_0.parquet\"),\n        # Data from multiple files matching the pattern \"test_static_0_*.parquet\"\n        du.read_files(CFG.test_path / \"test_static_0_*.parquet\"),\n    ],\n\n    # Second level of depth\n    \"depth_1\": [\n        # Data from multiple files matching the pattern \"test_applprev_1_*.parquet\"\n        du.read_files(CFG.test_path / \"test_applprev_1_*.parquet\", 1),\n        # Data from \"test_tax_registry_a_1.parquet\"\n        du.read_file(CFG.test_path / \"test_tax_registry_a_1.parquet\", 1),\n        # Data from \"test_tax_registry_b_1.parquet\"\n        du.read_file(CFG.test_path / \"test_tax_registry_b_1.parquet\", 1),\n        # Data from \"test_tax_registry_c_1.parquet\"\n        du.read_file(CFG.test_path / \"test_tax_registry_c_1.parquet\", 1),\n        # Data from multiple files matching the pattern \"test_credit_bureau_a_1_*.parquet\"\n        du.read_files(CFG.test_path / \"test_credit_bureau_a_1_*.parquet\", 1),\n        # Data from \"test_credit_bureau_b_1.parquet\"\n        du.read_file(CFG.test_path / \"test_credit_bureau_b_1.parquet\", 1),\n        # Data from \"test_other_1.parquet\"\n        du.read_file(CFG.test_path / \"test_other_1.parquet\", 1),\n        # Data from \"test_person_1.parquet\"\n        du.read_file(CFG.test_path / \"test_person_1.parquet\", 1),\n        # Data from \"test_deposit_1.parquet\"\n        du.read_file(CFG.test_path / \"test_deposit_1.parquet\", 1),\n        # Data from \"test_debitcard_1.parquet\"\n        du.read_file(CFG.test_path / \"test_debitcard_1.parquet\", 1),\n    ],\n\n    # Third level of depth\n    \"depth_2\": [\n        # Data from \"test_credit_bureau_b_2.parquet\"\n        du.read_file(CFG.test_path / \"test_credit_bureau_b_2.parquet\", 2),\n        # Data from multiple files matching the pattern \"test_credit_bureau_a_2_*.parquet\"\n        du.read_files(CFG.test_path / \"test_credit_bureau_a_2_*.parquet\", 2),\n        # Data from \"test_applprev_2.parquet\"\n        du.read_file(CFG.test_path / \"test_applprev_2.parquet\", 2),\n        # Data from \"test_person_2.parquet\"\n        du.read_file(CFG.test_path / \"test_person_2.parquet\", 2)\n    ]\n}\n","metadata":{"papermill":{"duration":0.52701,"end_time":"2024-05-26T05:22:32.945165","exception":false,"start_time":"2024-05-26T05:22:32.418155","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data = du.feature_eng(**test_data_store)\nprint(f\"The shape of the test data is: {test_data.shape}\")","metadata":{"papermill":{"duration":0.09381,"end_time":"2024-05-26T05:22:33.048627","exception":false,"start_time":"2024-05-26T05:22:32.954817","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del test_data_store\ngc.collect()","metadata":{"papermill":{"duration":0.120147,"end_time":"2024-05-26T05:22:33.177723","exception":false,"start_time":"2024-05-26T05:22:33.057576","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data = test_data.select([col for col in train_data.columns if col != \"target\"])","metadata":{"papermill":{"duration":0.019933,"end_time":"2024-05-26T05:22:33.206955","exception":false,"start_time":"2024-05-26T05:22:33.187022","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data, cat_cols = du.to_pandas(test_data, cat_cols)","metadata":{"papermill":{"duration":0.084223,"end_time":"2024-05-26T05:22:33.300288","exception":false,"start_time":"2024-05-26T05:22:33.216065","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data = du.reduce_mem_usage(test_data)","metadata":{"papermill":{"duration":0.258995,"end_time":"2024-05-26T05:22:33.568410","exception":false,"start_time":"2024-05-26T05:22:33.309415","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"The shape of the train data after feature reduction is: {train_data.shape}\")\nprint(f\"The shape of the test data after feature reduction is: {test_data.shape}\")","metadata":{"papermill":{"duration":0.016981,"end_time":"2024-05-26T05:22:33.594842","exception":false,"start_time":"2024-05-26T05:22:33.577861","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_mem = train_data.memory_usage().sum() / 1024**2\nprint(\"Memory usage of train data is {:.2f} MB\".format(train_mem))\ntest_mem = test_data.memory_usage().sum() / 1024**2\nprint(\"Memory usage of test data is {:.2f} MB\".format(test_mem))","metadata":{"papermill":{"duration":0.049503,"end_time":"2024-05-26T05:22:33.653820","exception":false,"start_time":"2024-05-26T05:22:33.604317","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del train_data, train_mem, test_mem\ngc.collect()","metadata":{"papermill":{"duration":0.354365,"end_time":"2024-05-26T05:22:34.017857","exception":false,"start_time":"2024-05-26T05:22:33.663492","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here's an explanation of the `ModelEnsemble` class:\n\n```python\nclass ModelEnsemble(BaseEstimator, RegressorMixin):\n```\nThis line defines a class called `ModelEnsemble` which inherits from both `BaseEstimator` and `RegressorMixin`, implying that instances of this class can be used as scikit-learn estimators.\n\n```python\n    # Constructor to initialize the model with a list of estimators\n    def __init__(self, estimators):\n        super().__init__()\n        self.estimators = estimators\n```\nThis `__init__` method initializes a `ModelEnsemble` object with a list of estimators. It calls the constructor of the superclass (`BaseEstimator`) and assigns the list of estimators to the `estimators` attribute of the object.\n\n```python\n    # Average predictions from all estimators\n    def predict(self, X):\n        y_preds = [estimator.predict(X) for estimator in self.estimators]\n        return np.mean(y_preds, axis=0)\n```\nThe `predict` method takes input data `X` and returns the average prediction made by each estimator in the ensemble. It iterates over each estimator in `self.estimators`, predicts using each estimator for input `X`, and then calculates the mean prediction across all estimators.\n\n```python\n    # Average predicted probabilities from all estimators\n    def predict_proba(self, X):\n        y_preds = [estimator.predict_proba(X) for estimator in self.estimators]\n        return np.mean(y_preds, axis=0)\n```\nThe `predict_proba` method behaves similarly to `predict`, but instead of returning predictions, it returns the average predicted probabilities for each class from all estimators in the ensemble.\n\n```python\n    # Save the ensemble model\n    def save_model(self, filepath):\n        joblib.dump(self, filepath)\n```\nThe `save_model` method saves the ensemble model to a file using the `joblib.dump` function. It takes a file path `filepath` as input.\n\n```python\n    # Load the ensemble model\n    @classmethod\n    def load_model(cls, filepath):\n        return joblib.load(filepath)\n```\nThe `load_model` method is a class method (decorated with `@classmethod`) used to load a saved ensemble model from a file. It takes a file path `filepath` as input and returns the loaded model using `joblib.load`.","metadata":{}},{"cell_type":"code","source":"class ModelEnsemble(BaseEstimator, RegressorMixin):\n    # Constructor to initialize the model with a list of estimators\n    def __init__(self, estimators):\n        super().__init__()  # Call the constructor of the base class\n        self.estimators = estimators  # Store the list of estimators\n    \n    # Average predictions from all estimators\n    def predict(self, X):\n        # Collect predictions from each estimator\n        y_preds = [estimator.predict(X) for estimator in self.estimators]\n        # Return the mean of predictions from all estimators\n        return np.mean(y_preds, axis=0)\n    \n    # Average predicted probabilities from all estimators\n    def predict_proba(self, X):\n        # Collect predicted probabilities from each estimator\n        y_preds = [estimator.predict_proba(X) for estimator in self.estimators]\n        # Return the mean of predicted probabilities from all estimators\n        return np.mean(y_preds, axis=0)\n    \n    # Save the ensemble model to a file\n    def save_model(self, filepath):\n        joblib.dump(self, filepath)  # Use joblib to save the model to the specified file path\n\n    # Load the ensemble model from a file\n    @classmethod\n    def load_model(cls, filepath):\n        # Use joblib to load the model from the specified file path and return it\n        return joblib.load(filepath)\n","metadata":{"papermill":{"duration":0.019511,"end_time":"2024-05-26T05:22:34.047361","exception":false,"start_time":"2024-05-26T05:22:34.027850","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\n1. **Class Definition:**\n   ```python\n   class ModelUtilities:\n   ```\n   - Defines a class named `ModelUtilities`.\n\n2. **Method `infer_lgb`:**\n   ```python\n   def infer_lgb(self, model, test_data, subm_data):\n   ```\n   - This method takes a LightGBM model (`model`), test data (`test_data`), and submission data file path (`subm_data`) as input.\n   \n   ```python\n   X_test = test_data.drop(columns=[\"WEEK_NUM\"]).set_index(\"case_id\")\n   ```\n   - Drops the \"WEEK_NUM\" column from the test data and sets \"case_id\" as the index.\n   \n   ```python\n   preds = model.predict_proba(X_test)[:, 1]\n   ```\n   - Uses the LightGBM model to predict probabilities, and selects the probabilities of the positive class.\n   \n   ```python\n   subm_df = pd.read_csv(subm_data)\n   subm_df = subm_df.set_index(\"case_id\")\n   ```\n   - Reads the submission data from a CSV file and sets \"case_id\" as the index.\n   \n   ```python\n   subm_df[\"score\"] = preds\n   display(subm_df)\n   ```\n   - Assigns the predicted scores to the \"score\" column of the submission DataFrame and displays it.\n   \n   ```python\n   return subm_df\n   ```\n   - Returns the updated submission DataFrame.\n\n3. **Method `infer_ctb`:**\n   ```python\n   def infer_ctb(self, model, test_data, cat_cols, subm_data):\n   ```\n   - This method takes a CatBoost model (`model`), test data (`test_data`), a list of categorical columns (`cat_cols`), and submission data file path (`subm_data`) as input.\n   \n   ```python\n   test_data[cat_cols] = test_data[cat_cols].astype(str)\n   ```\n   - Converts the specified categorical columns to string type.\n   \n   ```python\n   X_test = test_data.drop(columns=[\"WEEK_NUM\"]).set_index(\"case_id\")\n   preds = model.predict_proba(X_test)[:, 1]\n   subm_df = pd.read_csv(subm_data)\n   subm_df = subm_df.set_index(\"case_id\")\n   subm_df[\"score\"] = preds\n   display(subm_df)\n   ```\n   - Prepares the test data, makes predictions, reads and updates the submission DataFrame similarly to `infer_lgb`.\n   \n   ```python\n   return subm_df\n   ```\n   - Returns the updated submission DataFrame.\n\n4. **Method `inference`:**\n   ```python\n   def inference(self, pred_dfs):\n   ```\n   - This method takes a list of prediction DataFrames (`pred_dfs`) as input.\n   \n   ```python\n   combined_preds = np.mean([df[\"score\"].values for df in pred_dfs], axis=0)\n   ```\n   - Averages the 'score' column from each DataFrame in `pred_dfs` to combine predictions.\n   \n   ```python\n   final_preds = pred_dfs[0].copy()\n   final_preds[\"score\"] = combined_preds\n   display(final_preds)\n   ```\n   - Copies the first DataFrame in `pred_dfs`, assigns the combined predictions to its 'score' column, and displays it.\n   \n   ```python\n   return final_preds\n   ```\n   - Returns the final DataFrame with combined predictions.","metadata":{}},{"cell_type":"code","source":"class ModelUtilities:\n    # Method to make predictions using a LightGBM model and prepare submission data\n    def infer_lgb(self, model, test_data, subm_data):\n        # Prepare test data by dropping \"WEEK_NUM\" and setting \"case_id\" as the index\n        X_test = test_data.drop(columns=[\"WEEK_NUM\"]).set_index(\"case_id\")\n\n        # Make predictions using the LightGBM model, taking the probability of the positive class\n        preds = model.predict_proba(X_test)[:, 1]\n\n        # Read the submission data from the CSV file and set \"case_id\" as the index\n        subm_df = pd.read_csv(subm_data)\n        subm_df = subm_df.set_index(\"case_id\")\n\n        # Assign the predicted scores to the submission DataFrame\n        subm_df[\"score\"] = preds\n        display(subm_df)  # Display the submission DataFrame for verification\n\n        return subm_df  # Return the updated submission DataFrame\n    \n    # Method to make predictions using a CatBoost model and prepare submission data\n    def infer_ctb(self, model, test_data, cat_cols, subm_data):\n        # Convert categorical columns to string type\n        test_data[cat_cols] = test_data[cat_cols].astype(str)\n\n        # Prepare test data by dropping \"WEEK_NUM\" and setting \"case_id\" as the index\n        X_test = test_data.drop(columns=[\"WEEK_NUM\"]).set_index(\"case_id\")\n\n        # Make predictions using the CatBoost model, taking the probability of the positive class\n        preds = model.predict_proba(X_test)[:, 1]\n\n        # Read the submission data from the CSV file and set \"case_id\" as the index\n        subm_df = pd.read_csv(subm_data)\n        subm_df = subm_df.set_index(\"case_id\")\n\n        # Assign the predicted scores to the submission DataFrame\n        subm_df[\"score\"] = preds\n        display(subm_df)  # Display the submission DataFrame for verification\n\n        return subm_df  # Return the updated submission DataFrame\n\n    # Method to combine predictions from multiple DataFrames by averaging their 'score' columns\n    def inference(self, pred_dfs):\n        # Combine predictions by averaging the 'score' column from each DataFrame in pred_dfs\n        combined_preds = np.mean([df[\"score\"].values for df in pred_dfs], axis=0)\n        \n        # Create a final DataFrame by copying the first DataFrame in pred_dfs\n        final_preds = pred_dfs[0].copy()\n        # Assign the combined predictions to the 'score' column of the final DataFrame\n        final_preds[\"score\"] = combined_preds\n        \n        display(final_preds)  # Display the final DataFrame for verification\n        return final_preds  # Return the final DataFrame with combined predictions\n","metadata":{"papermill":{"duration":0.020986,"end_time":"2024-05-26T05:22:34.078166","exception":false,"start_time":"2024-05-26T05:22:34.057180","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mu = ModelUtilities()","metadata":{"papermill":{"duration":0.01599,"end_time":"2024-05-26T05:22:34.103945","exception":false,"start_time":"2024-05-26T05:22:34.087955","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load pretrained LightGBM and CatBoost classifiers\nlgb = ModelEnsemble.load_model(CFG.lgb)\nctb = ModelEnsemble.load_model(CFG.ctb)","metadata":{"papermill":{"duration":6.020404,"end_time":"2024-05-26T05:22:40.134410","exception":false,"start_time":"2024-05-26T05:22:34.114006","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lgb_preds = mu.infer_lgb(lgb, test_data, CFG.subm_data)","metadata":{"papermill":{"duration":0.448707,"end_time":"2024-05-26T05:22:40.593677","exception":false,"start_time":"2024-05-26T05:22:40.144970","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ctb_preds = mu.infer_ctb(ctb, test_data, cat_cols, CFG.subm_data) ","metadata":{"papermill":{"duration":0.211287,"end_time":"2024-05-26T05:22:40.815666","exception":false,"start_time":"2024-05-26T05:22:40.604379","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Combine predictions\npreds = [lgb_preds, ctb_preds]\nfinal_preds = mu.inference(preds)","metadata":{"papermill":{"duration":0.021737,"end_time":"2024-05-26T05:22:40.848445","exception":false,"start_time":"2024-05-26T05:22:40.826708","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_preds.to_csv(\"submission.csv\")","metadata":{"papermill":{"duration":0.019087,"end_time":"2024-05-26T05:22:40.878357","exception":false,"start_time":"2024-05-26T05:22:40.859270","status":"completed"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]}]}