{"metadata":{"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":50160,"databundleVersionId":7921029,"sourceType":"competition"},{"sourceId":7964074,"sourceType":"datasetVersion","datasetId":4685368}],"dockerImageVersionId":30674,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<p align=\"center\" >\n<img src=\"https://upload.wikimedia.org/wikipedia/commons/thumb/2/2b/Home_Credit_logo.svg/1920px-Home_Credit_logo.svg.png\" width=\"200\"/>\n</p>\n\n# Credit Risk Model Stability 🏦\n\n# 📑 Introduction\n​\n> Home Credit B.V. is an international non-bank financial institution founded in 1997 in the Czech Republic and headquartered in Netherlands. The company operates in 9 countries and focuses on installment lending primarily to people with little or no credit history. As of 30 June 2020 the Group has cumulatively served over 135.4 million customers. Major shareholder of the Group is PPF, a privately held international investment group whose founder and major beneficiary was Petr Kellner with 88.62% stake.\n​\n# 📝 Evaluation\n​\n ><p>Submissions are evaluated using a gini stability metric. A gini score is calculated for predictions corresponding to each <code>WEEK_NUM</code>.  </p>\n\n$$\n\\text{Gini} = 2 \\times \\text{Area under ROC curve} - 1 \n$$\n\n\n>A `linear regression model`, $$y = ax + b $$, is fit through the weekly Gini scores, and a <code>falling rate</code> is calculated as: $$min(0, a)$$\nThis is used to penalize models that drop off in predictive ability.\n\n>Finally, the variability of the predictions are calculated by taking the standard deviation of the residuals from the above linear regression, applying a penalty to model variablity.\n\n>The final metric is calculated as:\n\n$$\n\\text{stability metric} = mean(gini) +88.0 \\cdot min(0, a) -0.5 \\cdot std(\\text{residuals})\n$$\n​\n# 🎯 The Goal\n​\n> The goal of this competition is to predict which clients are more likely to default on their loans. The evaluation will favor solutions that are stable over time.\nYour participation may offer consumer finance providers a more reliable and longer-lasting way to assess a potential client’s default risk.\n​\n\n\n","metadata":{}},{"cell_type":"markdown","source":"# 💾 Data Description --> (first 50)\n\n----\n-----\nHere is the information on this particular data set:\n\n<table border=\"1\" class=\"dataframe\">\n  <thead>\n    <tr style=\"text-align: right;\">\n      <th></th>\n      <th>Featurelist</th>\n      <th>Description</th>\n    </tr>\n  </thead>\n  <tbody>\n    <tr>\n      <th>0</th>\n      <td>actualdpd_943P</td>\n      <td>Days Past Due (DPD) of previous contract (actual).</td>\n    </tr>\n    <tr>\n      <th>1</th>\n      <td>actualdpdtolerance_344P</td>\n      <td>DPD of client with tolerance.</td>\n    </tr>\n    <tr>\n      <th>2</th>\n      <td>addres_district_368M</td>\n      <td>District of the person's address.</td>\n    </tr>\n    <tr>\n      <th>3</th>\n      <td>addres_role_871L</td>\n      <td>Role of person's address.</td>\n    </tr>\n    <tr>\n      <th>4</th>\n      <td>addres_zip_823M</td>\n      <td>Zip code of the address.</td>\n    </tr>\n    <tr>\n      <th>5</th>\n      <td>amount_1115A</td>\n      <td>Credit amount of the active contract provided by the credit bureau.</td>\n    </tr>\n    <tr>\n      <th>6</th>\n      <td>amount_416A</td>\n      <td>Deposit amount.</td>\n    </tr>\n    <tr>\n      <th>7</th>\n      <td>amount_4527230A</td>\n      <td>Tax deductions amount tracked by the government registry.</td>\n    </tr>\n    <tr>\n      <th>8</th>\n      <td>amount_4917619A</td>\n      <td>Tax deductions amount tracked by the government registry.</td>\n    </tr>\n    <tr>\n      <th>9</th>\n      <td>amtdebitincoming_4809443A</td>\n      <td>Incoming debit card transactions amount.</td>\n    </tr>\n    <tr>\n      <th>10</th>\n      <td>amtdebitoutgoing_4809440A</td>\n      <td>Outgoing debit card transactions amount.</td>\n    </tr>\n    <tr>\n      <th>11</th>\n      <td>amtdepositbalance_4809441A</td>\n      <td>Deposit balance of client.</td>\n    </tr>\n    <tr>\n      <th>12</th>\n      <td>amtdepositincoming_4809444A</td>\n      <td>Amount of incoming deposits to client's account.</td>\n    </tr>\n    <tr>\n      <th>13</th>\n      <td>amtdepositoutgoing_4809442A</td>\n      <td>Amount of outgoing deposits from client's account.</td>\n    </tr>\n    <tr>\n      <th>14</th>\n      <td>amtinstpaidbefduel24m_4187115A</td>\n      <td>Number of instalments paid before due date in the last 24 months.</td>\n    </tr>\n    <tr>\n      <th>15</th>\n      <td>annualeffectiverate_199L</td>\n      <td>Interest rate of the closed contracts.</td>\n    </tr>\n    <tr>\n      <th>16</th>\n      <td>annualeffectiverate_63L</td>\n      <td>Interest rate for the active contracts.</td>\n    </tr>\n    <tr>\n      <th>17</th>\n      <td>annuity_780A</td>\n      <td>Monthly annuity amount.</td>\n    </tr>\n    <tr>\n      <th>18</th>\n      <td>annuity_853A</td>\n      <td>Monthly annuity for previous applications.</td>\n    </tr>\n    <tr>\n      <th>19</th>\n      <td>annuitynextmonth_57A</td>\n      <td>Next month's amount of annuity.</td>\n    </tr>\n    <tr>\n      <th>20</th>\n      <td>applicationcnt_361L</td>\n      <td>Number of applications associated with the same email address as the client.</td>\n    </tr>\n    <tr>\n      <th>21</th>\n      <td>applications30d_658L</td>\n      <td>Number of applications made by the client in the last 30 days.</td>\n    </tr>\n    <tr>\n      <th>22</th>\n      <td>applicationscnt_1086L</td>\n      <td>Number of applications associated with the same phone number.</td>\n    </tr>\n    <tr>\n      <th>23</th>\n      <td>applicationscnt_464L</td>\n      <td>Number of applications made in the last 30 days by other clients with the same employer as the applicant.</td>\n    </tr>\n    <tr>\n      <th>24</th>\n      <td>applicationscnt_629L</td>\n      <td>Number of applications with the same employer in the last 7 days.</td>\n    </tr>\n    <tr>\n      <th>25</th>\n      <td>applicationscnt_867L</td>\n      <td>Number of applications associated with the same mobile phone.</td>\n    </tr>\n    <tr>\n      <th>26</th>\n      <td>approvaldate_319D</td>\n      <td>Approval Date of Previous Application</td>\n    </tr>\n    <tr>\n      <th>27</th>\n      <td>assignmentdate_238D</td>\n      <td>Tax authority data - date of assignment.</td>\n    </tr>\n    <tr>\n      <th>28</th>\n      <td>assignmentdate_4527235D</td>\n      <td>Tax authority data - Date of assignment.</td>\n    </tr>\n    <tr>\n      <th>29</th>\n      <td>assignmentdate_4955616D</td>\n      <td>Tax authority assignment date.</td>\n    </tr>\n    <tr>\n      <th>30</th>\n      <td>avgdbddpdlast24m_3658932P</td>\n      <td>Average days past or before due of payment during the last 24 months.</td>\n    </tr>\n</tbody>\n</table>","metadata":{}},{"cell_type":"markdown","source":"# Let's see the code... 👨‍💻","metadata":{}},{"cell_type":"markdown","source":"# <p style=\" font-weight:bold; letter-spacing: 2px; color:black; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid black;\">1. Import Libraries</p>","metadata":{}},{"cell_type":"code","source":"import os\nimport gc\nfrom glob import glob\nfrom pathlib import Path\nfrom datetime import datetime\n\nimport numpy as np\nimport pandas as pd\nimport polars as pl\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom sklearn.model_selection import StratifiedGroupKFold\nfrom sklearn.base import BaseEstimator, ClassifierMixin\n\nimport lightgbm as lgb\n\nimport warnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.status.busy":"2024-03-28T10:47:44.582912Z","iopub.execute_input":"2024-03-28T10:47:44.583728Z","iopub.status.idle":"2024-03-28T10:47:49.336390Z","shell.execute_reply.started":"2024-03-28T10:47:44.583701Z","shell.execute_reply":"2024-03-28T10:47:49.335426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\" font-weight:bold; letter-spacing: 2px; color:black; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid black;\">2. Definition of the class: Pipeline</p>","metadata":{}},{"cell_type":"code","source":"class Pipeline:\n    \n    @staticmethod\n    def set_table_dtypes(df):\n        \"\"\"\n        Method to set the data types of the DataFrame columns.\n\n        Parameters:\n        - df: DataFrame\n            DataFrame containing the data to be transformed\n\n        Returns:\n        - df: DataFrame\n            Transformed DataFrame with correct data types\n        \"\"\"\n        for col in df.columns:\n            if col in [\"case_id\", \"WEEK_NUM\", \"num_group1\", \"num_group2\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Int32)) # if the columns are among those listed, cast them to Int32\n            elif col in [\"date_decision\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Date)) # if the column is date_decision, cast it to Date\n            elif col[-1] in (\"P\", \"A\"):\n                df = df.with_columns(pl.col(col).cast(pl.Float64)) # if the column ends with P or A, cast it to Float64\n            elif col[-1] in (\"M\",):\n                df = df.with_columns(pl.col(col).cast(pl.String)) # if the column ends with M, cast it to String\n            elif col[-1] in (\"D\",):\n                df = df.with_columns(pl.col(col).cast(pl.Date)) # if the column ends with D, cast it to Date\n\n        return df\n\n    @staticmethod\n    def handle_dates(df):\n        \"\"\"\n        Method to handle dates in the DataFrame.\n\n        Parameters:\n        - df: DataFrame\n            DataFrame containing the data to be transformed\n\n        Returns:\n        - df: DataFrame\n            Transformed DataFrame with dates handled correctly\n        \"\"\"\n        for col in df.columns:\n            if col[-1] in (\"D\",):\n                # calculate the difference between the values of the column ending with D and date_decision\n                df = df.with_columns(pl.col(col) - pl.col(\"date_decision\")) \n                df = df.with_columns(pl.col(col).dt.total_days()) # transform the difference to days\n                df = df.with_columns(pl.col(col).cast(pl.Float32)) # transform the difference to Float32\n\n        df = df.drop(\"date_decision\", \"MONTH\") # remove the date_decision and MONTH columns\n\n        return df\n\n    @staticmethod\n    def filter_cols(df):\n        \"\"\"\n        Method to filter the columns of the DataFrame.\n\n        Removes columns with a high rate of null values and columns with a single frequency or high frequency of unique values.\n\n        Parameters:\n        - df: DataFrame\n            DataFrame containing the data to be filtered\n\n        Returns:\n        - df: DataFrame\n            Filtered DataFrame\n        \"\"\"\n        # remove columns with a high rate of null values (95%)\n        for col in df.columns:\n            if col not in [\"target\", \"case_id\", \"WEEK_NUM\"]:\n                isnull = df[col].is_null().mean()\n\n                if isnull > 0.95:\n                    df = df.drop(col)\n                    \n        # remove columns with a single frequency or high frequency of unique values\n        for col in df.columns:\n            if (col not in [\"target\", \"case_id\", \"WEEK_NUM\"]) & (df[col].dtype == pl.String):\n                freq = df[col].n_unique()\n\n                if (freq == 1) | (freq > 200):\n                    df = df.drop(col)\n\n        return df\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\" font-weight:bold; letter-spacing: 2px; color:black; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid black;\">3. Definition of the class: Aggregator</p>","metadata":{}},{"cell_type":"code","source":"class Aggregator:\n\n    @staticmethod\n    def num_expr(df):\n        \"\"\"\n        Method to generate aggregate expressions for numeric columns.\n        \n        Generates an expression to calculate the maximum value (max) for each numeric column.\n\n        Parameters:\n        - df: DataFrame\n            DataFrame containing the data\n\n        Returns:\n        - expr_max: list of aggregate expressions\n            List of expressions to calculate the maximum value for each numeric column\n        \"\"\"\n        cols = [col for col in df.columns if col[-1] in (\"P\", \"A\")] # list of columns ending with P or A\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols] # expression to calculate the maximum value for each column ending with P or A\n        return expr_max\n\n    @staticmethod\n    def date_expr(df):\n        \"\"\"\n        Method to generate aggregate expressions for date columns.\n\n        Generates an expression to calculate the maximum value (max) for each date column.\n\n        Parameters:\n        - df: DataFrame\n            DataFrame containing the data\n\n        Returns:\n        - expr_max: list of aggregate expressions\n            List of expressions to calculate the maximum value for each date column\n        \"\"\"\n        cols = [col for col in df.columns if col[-1] in (\"D\",)] # list of columns ending with D\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols] # expression to calculate the maximum value for each column ending with D\n        return expr_max\n\n    @staticmethod\n    def str_expr(df):\n        \"\"\"\n        Method to generate aggregate expressions for string columns.\n\n        Generates an expression to calculate the maximum value (max) for each string column.\n\n        Parameters:\n        - df: DataFrame\n            DataFrame containing the data\n\n        Returns:\n        - expr_max: list of aggregate expressions\n            List of expressions to calculate the maximum value for each string column\n        \"\"\"\n        cols = [col for col in df.columns if col[-1] in (\"M\",)] # list of columns ending with M\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols] # expression to calculate the maximum value for each column ending with M\n        return expr_max\n\n    @staticmethod\n    def other_expr(df):\n        \"\"\"\n        Method to generate aggregate expressions for other columns.\n\n        Generates an expression to calculate the maximum value (max) for each column with suffix \"T\" or \"L\".\n\n        Parameters:\n        - df: DataFrame\n            DataFrame containing the data\n\n        Returns:\n        - expr_max: list of aggregate expressions\n            List of expressions to calculate the maximum value for each column with suffix \"T\" or \"L\"\n        \"\"\"\n        cols = [col for col in df.columns if col[-1] in (\"T\", \"L\")] # list of columns ending with T or L\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols] # expression to calculate the maximum value for each column ending with T or L\n        return expr_max\n\n    @staticmethod\n    def count_expr(df):\n        \"\"\"\n        Method to generate aggregate expressions for columns containing \"num_group\".\n\n        Generates an expression to calculate the maximum value (max) for each column containing \"num_group\".\n\n        Parameters:\n        - df: DataFrame\n            DataFrame containing the data\n\n        Returns:\n        - expr_max: list of aggregate expressions\n            List of expressions to calculate the maximum value for each column containing \"num_group\"\n        \"\"\"\n        cols = [col for col in df.columns if \"num_group\" in col] # list of columns containing \"num_group\"\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols] # expression to calculate the maximum value for each column containing \"num_group\"\n        return expr_max\n\n    @staticmethod\n    def get_exprs(df):\n        \"\"\"\n        Method to get all aggregate expressions.\n\n        Concatenates the aggregate expressions generated by the previous methods.\n\n        Parameters:\n        - df: DataFrame\n            DataFrame containing the data\n\n        Returns:\n        - exprs: list of aggregate expressions\n            Complete list of aggregate expressions\n        \"\"\"\n        \n        exprs = Aggregator.num_expr(df) + \\\n                Aggregator.date_expr(df) + \\\n                Aggregator.str_expr(df) + \\\n                Aggregator.other_expr(df) + \\\n                Aggregator.count_expr(df) # concatenation of aggregate expressions\n                \n        return exprs\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\" font-weight:bold; letter-spacing: 2px; color:black; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid black;\">4. Definition of: \"read_file\" & \"read_files\" functions </p>\n\n> - `read_file` reads only one file (.parquet)\n> - `read_files` reads all files with a specific prefix (.parquet)","metadata":{}},{"cell_type":"code","source":"def read_file(path, depth=None):\n    \"\"\"\n    Reads a Parquet file, applies the transformations defined in Pipeline and Aggregator,\n    and returns the transformed DataFrame.\n\n    Parameters:\n    - path: string\n        Path of the Parquet file to read.\n    - depth: int, optional\n        Level of data aggregation. If depth is 1 or 2, the data is aggregated by \"case_id\".\n\n    Returns:\n    - df: DataFrame\n        Transformed DataFrame.\n    \"\"\"\n    df = pl.read_parquet(path)  # Reads the Parquet file\n    df = df.pipe(Pipeline.set_table_dtypes)  # Applies Pipeline transformations\n    \n    if depth in [1, 2]: # If depth is 1 or 2, aggregates the data by \"case_id\"\n        df = df.group_by(\"case_id\").agg(Aggregator.get_exprs(df))  # Groups by case_id and applies Aggregator transformations (get_exprs)\n    return df\n\ndef read_files(regex_path, depth=None):\n    \"\"\"\n    Reads a series of Parquet files matching a regular expression, applies the transformations\n    defined in Pipeline and Aggregator, and returns the concatenated and unique DataFrame.\n\n    Parameters:\n    - regex_path: string\n        Regular expression for the path of Parquet files to read.\n    - depth: int, optional\n        Level of data aggregation. If depth is 1 or 2, the data is aggregated by \"case_id\".\n\n    Returns:\n    - df: DataFrame\n        Concatenated and unique DataFrame.\n    \"\"\"\n    chunks = []\n    for path in glob(str(regex_path)):  # Iterates over files matching the specified path\n        # (with glob--> includes all patterns *)\n\n        df = pl.read_parquet(path)  # Reads the Parquet file(s)\n        df = df.pipe(Pipeline.set_table_dtypes)  # Applies Pipeline transformations (set_table_dtypes)\n        \n        if depth in [1, 2]:\n            df = df.group_by(\"case_id\").agg(Aggregator.get_exprs(df))  # Applies Aggregator transformations\n        \n        # Adds the DataFrame to the list (since multiple files are read)\n        chunks.append(df)\n        \n    df = pl.concat(chunks, how=\"vertical_relaxed\")  # Concatenates the DataFrames\n    df = df.unique(subset=[\"case_id\"])  # Removes duplicate rows based on \"case_id\"\n    \n    return df\n\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\" font-weight:bold; letter-spacing: 2px; color:black; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid black;\">5. Definition of: \"feature_eng\" function </p>","metadata":{}},{"cell_type":"code","source":"def feature_eng(df_base, depth_0, depth_1, depth_2):\n    \"\"\"\n    Performs feature engineering on the base DataFrame and integrates data from different aggregation levels.\n\n    Parameters:\n    - df_base: DataFrame\n        Base DataFrame on which to perform feature engineering.\n    - depth_0: list of DataFrames\n        List of DataFrames at the first level of aggregation.\n    - depth_1: list of DataFrames\n        List of DataFrames at the second level of aggregation.\n    - depth_2: list of DataFrames\n        List of DataFrames at the third level of aggregation.\n\n    Returns:\n    - df_base: DataFrame\n        DataFrame with feature engineering performed and integrated data.\n    \"\"\"\n    \n    # Adds new columns for month and weekday from \"date_decision\"\n    df_base = (\n        df_base\n        .with_columns(\n            month_decision = pl.col(\"date_decision\").dt.month(),\n            weekday_decision = pl.col(\"date_decision\").dt.weekday(),\n        )\n    )\n    \n    # Joins data from different aggregation levels to the base DataFrame\n    for i, df in enumerate(depth_0 + depth_1 + depth_2):\n        df_base = df_base.join(df, how=\"left\", on=\"case_id\", suffix=f\"_{i}\")\n        \n    # Applies date handling transformations\n    df_base = df_base.pipe(Pipeline.handle_dates)\n    \n    return df_base\n\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\" font-weight:bold; letter-spacing: 2px; color:black; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid black;\">6. Definition of: \"to_pandas\" function --> from generic DF to pandas DF </p>","metadata":{}},{"cell_type":"code","source":"def to_pandas(df_data, cat_cols=None):\n    \"\"\"\n    Converts a DataFrame from a specific format to a Pandas DataFrame and performs type operations for categorical columns.\n\n    Parameters:\n    - df_data: DataFrame\n        DataFrame to convert.\n    - cat_cols: list of strings, optional\n        List of categorical columns.\n\n    Returns:\n    - df_data: Pandas DataFrame\n        Converted DataFrame.\n    - cat_cols: list of strings\n        List of categorical columns.\n    \"\"\"\n    \n    # Converts the DataFrame to a Pandas DataFrame\n    df_data = df_data.to_pandas()\n    \n    # If categorical columns are not specified, identifies columns of type \"object\" as categorical columns\n    if cat_cols is None:\n        cat_cols = list(df_data.select_dtypes(\"object\").columns)\n    \n    # Converts identified categorical columns to \"category\" type\n    df_data[cat_cols] = df_data[cat_cols].astype(\"category\")\n    \n    return df_data, cat_cols\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\" font-weight:bold; letter-spacing: 2px; color:black; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid black;\">7. Save the paths </p>","metadata":{}},{"cell_type":"code","source":"ROOT            = Path(\"/kaggle/input/home-credit-credit-risk-model-stability\")\nTRAIN_DIR       = ROOT / \"parquet_files\" / \"train\"\nTEST_DIR        = ROOT / \"parquet_files\" / \"test\"","metadata":{"execution":{"iopub.execute_input":"2024-03-11T19:56:46.658359Z","iopub.status.busy":"2024-03-11T19:56:46.657949Z","iopub.status.idle":"2024-03-11T19:56:46.668308Z","shell.execute_reply":"2024-03-11T19:56:46.667301Z","shell.execute_reply.started":"2024-03-11T19:56:46.658316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\" font-weight:bold; letter-spacing: 2px; color:black; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid black;\">8. Read Train & Test files</p>","metadata":{}},{"cell_type":"markdown","source":" **Train files**","metadata":{}},{"cell_type":"code","source":"\"\"\" - read_file = returns a DataFrame (reads a single file)\n- read_files = returns a concatenated and unique DataFrame (reads multiple files using a pattern *) \"\"\"\n\n\n# CREATE a dictionary with train data (comprising a series of lists containing the DataFrames for different levels)\ndata_store = {\n    \"df_base\": read_file(TRAIN_DIR / \"train_base.parquet\"),\n    \"depth_0\": [\n        read_file(TRAIN_DIR / \"train_static_cb_0.parquet\"),\n        read_files(TRAIN_DIR / \"train_static_0_*.parquet\"),  # read using a pattern\n    ],\n    \"depth_1\": [\n        read_files(TRAIN_DIR / \"train_applprev_1_*.parquet\", 1),  # read using a pattern\n        read_file(TRAIN_DIR / \"train_tax_registry_a_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_b_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_c_1.parquet\", 1),\n        read_files(TRAIN_DIR / \"train_credit_bureau_a_1_*.parquet\", 1),  # read using a pattern\n        read_file(TRAIN_DIR / \"train_credit_bureau_b_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_other_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_person_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_deposit_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_debitcard_1.parquet\", 1),\n    ],\n    \"depth_2\": [\n        read_file(TRAIN_DIR / \"train_credit_bureau_b_2.parquet\", 2),\n        read_files(TRAIN_DIR / \"train_credit_bureau_a_2_*.parquet\", 2),  # read using a pattern\n    ]\n}\n\n","metadata":{"execution":{"iopub.execute_input":"2024-03-11T19:56:46.670129Z","iopub.status.busy":"2024-03-11T19:56:46.669639Z","iopub.status.idle":"2024-03-11T19:59:06.750594Z","shell.execute_reply":"2024-03-11T19:59:06.749785Z","shell.execute_reply.started":"2024-03-11T19:56:46.670094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\" The data_store dictionary is passed as an argument to the feature_eng function using the ** operator, \nwhich expands the dictionary into keyword arguments:\nthus data_store is expanded into the dictionary keys that are passed to the function (df_base, depth_0, depth_1, depth_2).\n\"\"\"\ndf_train = feature_eng(**data_store)\n\nprint(\"train data shape:\\t\", df_train.shape)\n\n","metadata":{"execution":{"iopub.execute_input":"2024-03-11T19:59:06.754345Z","iopub.status.busy":"2024-03-11T19:59:06.754038Z","iopub.status.idle":"2024-03-11T19:59:20.573791Z","shell.execute_reply":"2024-03-11T19:59:20.572796Z","shell.execute_reply.started":"2024-03-11T19:59:06.754321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Test files**\n","metadata":{}},{"cell_type":"code","source":"data_store = {\n    \"df_base\": read_file(TEST_DIR / \"test_base.parquet\"),\n    \"depth_0\": [\n        read_file(TEST_DIR / \"test_static_cb_0.parquet\"),\n        read_files(TEST_DIR / \"test_static_0_*.parquet\"),\n    ],\n    \"depth_1\": [\n        read_files(TEST_DIR / \"test_applprev_1_*.parquet\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_a_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_b_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_c_1.parquet\", 1),\n        read_files(TEST_DIR / \"test_credit_bureau_a_1_*.parquet\", 1),\n        read_file(TEST_DIR / \"test_credit_bureau_b_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_other_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_person_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_deposit_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_debitcard_1.parquet\", 1),\n    ],\n    \"depth_2\": [\n        read_file(TEST_DIR / \"test_credit_bureau_b_2.parquet\", 2),\n        read_files(TEST_DIR / \"test_credit_bureau_a_2_*.parquet\", 2),\n    ]\n}","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test = feature_eng(**data_store)\n\nprint(\"test data shape:\\t\", df_test.shape)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\" font-weight:bold; letter-spacing: 2px; color:black; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid black;\">9. Perform Feature Elimination</p>","metadata":{}},{"cell_type":"code","source":"df_train = df_train.pipe(Pipeline.filter_cols)\n# select all columns present in df_train except the target column\ndf_test = df_test.select([col for col in df_train.columns if col != \"target\"]) \n\nprint(\"train data shape:\\t\", df_train.shape)\nprint(\"test data shape:\\t\", df_test.shape)\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\" font-weight:bold; letter-spacing: 2px; color:black; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid black;\">10. Converts the df to pandas</p>","metadata":{}},{"cell_type":"code","source":"# 'cat_cols' will contain the list of identified categorical columns\ndf_train, cat_cols = to_pandas(df_train)\ndf_test, cat_cols = to_pandas(df_test, cat_cols)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\" font-weight:bold; letter-spacing: 2px; color:black; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid black;\">11. Clear memory</p>","metadata":{}},{"cell_type":"code","source":"\"\"\" The declaration of data_store removes the data_store dictionary from memory, making the\npreviously occupied memory space available for other purposes. \"\"\"\n\ndel data_store\n\n# Garbage Collection: gc.collect() is a call to Python's gc (garbage collector) module\n# to perform garbage collection, i.e., free up memory from unreferenced and unused objects.\ngc.collect()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\" font-weight:bold; letter-spacing: 2px; color:black; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid black;\">12.Exploratory Data Analysis 🔍</p>","metadata":{}},{"cell_type":"code","source":"print(\"Train is duplicated:\\t\", df_train[\"case_id\"].duplicated().any())\n# print week range for the train\nprint(\"Train Week Range:\\t\", (df_train[\"WEEK_NUM\"].min(), df_train[\"WEEK_NUM\"].max()))\n\nprint()\n\nprint(\"Test is duplicated:\\t\", df_test[\"case_id\"].duplicated().any())\n# print week range for the test\nprint(\"Test Week Range:\\t\", (df_test[\"WEEK_NUM\"].min(), df_test[\"WEEK_NUM\"].max()))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# TREND OF THE TARGET VARIABLE OVER TIME.\n\nsns.lineplot(\n    data=df_train,\n    x=\"WEEK_NUM\",\n    y=\"target\",\n)\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\" font-weight:bold; letter-spacing: 2px; color:black; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid black;\">13. Train the model 🏋️</p>","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import roc_auc_score\n\nX = df_train.drop(columns=[\"target\", \"case_id\", \"WEEK_NUM\"])\ny = df_train[\"target\"]\nweeks = df_train[\"WEEK_NUM\"]\n\ncv = StratifiedGroupKFold(n_splits=5, shuffle=False)\n\n# Definition of LightGBM model parameters:\n\"\"\" params = {\n    \"boosting_type\": \"gbdt\",\n    \"objective\": \"binary\",\n    \"metric\": \"auc\",\n    \"max_depth\": 8,\n    \"learning_rate\": 0.05,\n    \"n_estimators\": 1000,\n    \"colsample_bytree\": 0.8, \n    \"colsample_bynode\": 0.8,\n    \"verbose\": -1,\n    \"random_state\": 42,\n    \"device\": \"gpu\",\n} \"\"\"\n\n# 1) also try these parameters\nparams = {\n    \"boosting_type\": \"gbdt\",\n    \"objective\": \"binary\",\n    \"metric\": \"auc\",\n    \"max_depth\": 10,  \n    \"learning_rate\": 0.05,\n    \"n_estimators\": 2000,  \n    \"colsample_bytree\": 0.8,\n    \"colsample_bynode\": 0.8,\n    \"verbose\": -1,\n    \"random_state\": 42,\n    \"reg_alpha\": 0.1,\n    \"reg_lambda\": 10,\n    \"extra_trees\":True,\n    'num_leaves':64,\n    \"device\": \"gpu\", \n    \"verbose\": -1\n} \n\"\"\"\n# 2) also try these parameters\nparams = {\n    \"boosting_type\": \"gbdt\",\n    \"objective\": \"binary\",\n    \"metric\": \"auc\",\n    \"max_depth\": 8,\n    \"learning_rate\": 0.03,\n    \"n_estimators\": 1000,\n    \"colsample_bytree\": 0.8, \n    \"colsample_bynode\": 0.8,\n    \"verbose\": -1,\n    \"random_state\": 42,\n    \"device\": \"gpu\",\n}\n\"\"\"\n\nfitted_models = []\ncv_scores = []\n\n\"\"\"Using a for loop, the LGBMClassifier model is trained\non each fold of cross-validation. During training,\nthe features and label of the training set (X_train and y_train) are used,\nand the model is validated on the validation set (X_valid and y_valid).\nThe trained models are saved in the fitted_models list.\"\"\"\n\n# Training the model with StratifiedGroupKFold cross-validation\nfor idx_train, idx_valid in cv.split(X, y, groups=weeks):\n    X_train, y_train = X.iloc[idx_train], y.iloc[idx_train]\n    X_valid, y_valid = X.iloc[idx_valid], y.iloc[idx_valid]\n\n    model = lgb.LGBMClassifier(**params)\n    model.fit(\n        X_train, y_train,\n        eval_set=[(X_valid, y_valid)],\n        # initial and 2): \n        # callbacks=[lgb.log_evaluation(100), lgb.early_stopping(100)]\n        callbacks = [lgb.log_evaluation(200), lgb.early_stopping(60)] \n    )\n\n    fitted_models.append(model)\n\n    y_pred_valid = model.predict_proba(X_valid)[:, 1]\n    auc_score = roc_auc_score(y_valid, y_pred_valid)\n    cv_scores.append(auc_score)\n    \nprint(\"CV AUC scores: \", cv_scores)\nprint(\"Maximum CV AUC score: \", max(cv_scores))\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\" font-weight:bold; letter-spacing: 2px; color:black; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid black;\">14. Voting Model ↕️ </p>","metadata":{}},{"cell_type":"code","source":"from sklearn.base import BaseEstimator, ClassifierMixin\n\nclass VotingModel(BaseEstimator, ClassifierMixin):\n    \"\"\"\n    Implementation of a voting model for classification.\n    Combines predictions from a list of estimators to produce a final prediction.\n    \"\"\"\n\n    def __init__(self, estimators):\n        \"\"\"\n        Class constructor.\n\n        Parameters:\n        - estimators: list of estimators to use for predictions\n        \"\"\"\n        super().__init__()\n        self.estimators = estimators  # Stores the list of estimators\n\n    def fit(self, X, y=None):\n        \"\"\"\n        Placeholder for the training method.\n        \n        Does not perform training, but simply returns the current instance.\n\n        Parameters:\n        - X: array-like, shape (n_samples, n_features)\n            Training data\n        - y: array-like, shape (n_samples,), optional (default=None)\n            Training labels\n\n        Returns:\n        - self: current instance of the VotingModel object\n        \"\"\"\n        return self\n\n    def predict(self, X):\n        \"\"\"\n        Method to make predictions based on the estimators.\n\n        Uses the predict function of each estimator to get the predictions\n        and then calculates the arithmetic mean of these predictions.\n\n        Parameters:\n        - X: array-like, shape (n_samples, n_features)\n            Data on which to make predictions\n\n        Returns:\n        - y_pred: array-like, shape (n_samples,)\n            Average predictions of the estimators\n        \"\"\"\n        # Gets predictions from each estimator and stores them in a list\n        y_preds = [estimator.predict(X) for estimator in self.estimators]\n        # Calculates the arithmetic mean of the predictions\n        return np.mean(y_preds, axis=0)\n\n    def predict_proba(self, X):\n        \"\"\"\n        Method to get predicted probabilities based on the estimators.\n\n        Uses the predict_proba function of each estimator to get the\n        predicted probabilities and then calculates the arithmetic mean of these probabilities.\n\n        Parameters:\n        - X: array-like, shape (n_samples, n_features)\n            Data on which to get predicted probabilities\n\n        Returns:\n        - y_prob: array-like, shape (n_samples, n_classes)\n            Average predicted probabilities of the estimators\n        \"\"\"\n        # Gets predicted probabilities from each estimator and stores them in a list\n        y_preds = [estimator.predict_proba(X) for estimator in self.estimators]\n        # Calculates the arithmetic mean of the predicted probabilities\n        return np.mean(y_preds, axis=0)\n\n# Creating the VotingModel model\nmodel = VotingModel(fitted_models)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\" font-weight:bold; letter-spacing: 2px; color:black; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid black;\">15. Plot Feature Importance 📊</p>","metadata":{}},{"cell_type":"code","source":"lgb.plot_importance(fitted_models[2], figsize=(10,45))\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\" font-weight:bold; letter-spacing: 2px; color:black; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid black;\">16. Predictions 🔮</p>","metadata":{}},{"cell_type":"code","source":"# Prepare test data\nX_test = df_test.drop(columns=[\"WEEK_NUM\"])\nX_test = X_test.set_index(\"case_id\")\n\n# Predictions on the test set\ny_pred = pd.Series(model.predict_proba(X_test)[:, 1], index=X_test.index)\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\" font-weight:bold; letter-spacing: 2px; color:black; font-size:100%; text-align:left;padding: 0px; border-bottom: 3px solid black;\">17. Submission ✅</p> ","metadata":{}},{"cell_type":"code","source":"df_subm = pd.read_csv(ROOT / \"sample_submission.csv\")\n# Reads the submission file and sets \"case_id\" as the index\ndf_subm = df_subm.set_index(\"case_id\")\n\ndf_subm[\"score\"] = y_pred","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Check null: \", df_subm[\"score\"].isnull().any())\ndf_subm.head()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_subm.to_csv(\"submission.csv\")","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:19:29.485101Z","iopub.status.busy":"2024-03-11T20:19:29.484376Z","iopub.status.idle":"2024-03-11T20:19:29.493778Z","shell.execute_reply":"2024-03-11T20:19:29.492359Z","shell.execute_reply.started":"2024-03-11T20:19:29.485058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Reference: [notebook](https://www.kaggle.com/code/greysky/home-credit-baseline)\n","metadata":{}}]}