{"metadata":{"kaggle":{"accelerator":"none","dataSources":[{"sourceId":50160,"databundleVersionId":7921029,"sourceType":"competition"}],"dockerImageVersionId":30664,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false},"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"papermill":{"default_parameters":{},"duration":1221.035913,"end_time":"2024-03-11T20:40:31.474791","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2024-03-11T20:20:10.438878","version":"2.5.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Гимранов Артур","metadata":{}},{"cell_type":"markdown","source":"В данном ноутбуке я проводил исследование и готовил данные к финальной посылки, здесь можно найти все комментарии","metadata":{}},{"cell_type":"markdown","source":"TLDR: Дополнил агрегацию данных по группам функциями pl.median, pl.mean и pl.min в дополнение к pl.max. Это потребовало дополнительных усилий для обработки дат и частого удаления пустых столбцов, чтобы избежать переполнения памяти. Выделил несколько наборов признаков для оптимизации с использованием Optuna и провёл настройку параметров через этот инструмент. Также упомянута неудачная попытка более эффективного отбора категориальных признаков.","metadata":{}},{"cell_type":"markdown","source":"Ссылка на ноутбук из посылки: https://www.kaggle.com/code/toobrainless/credit/notebook","metadata":{}},{"cell_type":"code","source":"# Note: I'm looking for a job in Europe, if you like my work don't hesitate to reach =)\n\nimport os\nimport gc\nfrom glob import glob\nfrom pathlib import Path\nfrom datetime import datetime\n\nimport numpy as np\nimport pandas as pd\nimport polars as pl\n\nfrom datetime import datetime\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom sklearn.model_selection import StratifiedGroupKFold\nfrom sklearn.base import BaseEstimator, ClassifierMixin\n\nimport lightgbm as lgb\n\nimport warnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.execute_input":"2024-03-11T20:20:13.177936Z","iopub.status.busy":"2024-03-11T20:20:13.177674Z","iopub.status.idle":"2024-03-11T20:20:19.340830Z","shell.execute_reply":"2024-03-11T20:20:19.340062Z"},"papermill":{"duration":6.175783,"end_time":"2024-03-11T20:20:19.343026","exception":false,"start_time":"2024-03-11T20:20:13.167243","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Pre-Fitted Voting Model","metadata":{"papermill":{"duration":0.008777,"end_time":"2024-03-11T20:20:19.361579","exception":false,"start_time":"2024-03-11T20:20:19.352802","status":"completed"},"tags":[]}},{"cell_type":"code","source":"class VotingModel(BaseEstimator, ClassifierMixin):\n    def __init__(self, estimators):\n        super().__init__()\n        self.estimators = estimators\n        \n    def fit(self, X, y=None):\n        return self\n    \n    def predict(self, X):\n        y_preds = [estimator.predict(X) for estimator in self.estimators]\n        return np.mean(y_preds, axis=0)\n    \n    def predict_proba(self, X):\n        y_preds = [estimator.predict_proba(X) for estimator in self.estimators]\n        return np.mean(y_preds, axis=0)","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:20:19.381759Z","iopub.status.busy":"2024-03-11T20:20:19.381081Z","iopub.status.idle":"2024-03-11T20:20:19.387550Z","shell.execute_reply":"2024-03-11T20:20:19.386761Z"},"papermill":{"duration":0.018278,"end_time":"2024-03-11T20:20:19.389382","exception":false,"start_time":"2024-03-11T20:20:19.371104","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Pipeline","metadata":{"papermill":{"duration":0.008615,"end_time":"2024-03-11T20:20:19.406942","exception":false,"start_time":"2024-03-11T20:20:19.398327","status":"completed"},"tags":[]}},{"cell_type":"code","source":"some_date = datetime(2024, 3, 16)\n\nclass Pipeline:\n    @staticmethod\n    def set_table_dtypes(df):\n        for col in df.columns:\n            if col in [\"case_id\", \"WEEK_NUM\", \"num_group1\", \"num_group2\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Int32))\n            elif col in [\"date_decision\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Date))\n            elif col[-1] in (\"P\", \"A\"):\n                df = df.with_columns(pl.col(col).cast(pl.Float64))\n            elif col[-1] in (\"M\",):\n                df = df.with_columns(pl.col(col).cast(pl.String))\n            elif col[-1] in (\"D\",):\n                df = df.with_columns(pl.col(col).cast(pl.Date))            \n\n        return df\n    \n    @staticmethod\n    def handle_dates1(df):\n        for col in df.columns:\n            if col[-1] in (\"D\",):\n                df = df.with_columns((pl.col(col) - some_date).alias(col))\n                \n        df = df.drop(\"date_decision\", \"MONTH\")\n\n        return df\n    \n    @staticmethod\n    def handle_dates2(df):\n        for col in df.columns:\n            if col[-1] in (\"D\",) or col[-2:] in (\"D#\",):\n                df = df.with_columns(pl.col(col) + (some_date - pl.col(\"date_decision\")))\n                df = df.with_columns(pl.col(col).dt.total_days())\n                df = df.with_columns(pl.col(col).cast(pl.Float32))\n                \n        df = df.drop(\"date_decision\", \"MONTH\")\n\n        return df\n    \n    @staticmethod\n    def filter_cols(df):\n        for col in df.columns:\n            if col not in [\"target\", \"case_id\", \"WEEK_NUM\"]:\n                isnull = df[col].is_null().mean()\n\n                if isnull > 0.95:\n                    df = df.drop(col)\n\n        for col in df.columns:\n            if (col not in [\"target\", \"case_id\", \"WEEK_NUM\"]) & (df[col].dtype == pl.String):\n                freq = df[col].n_unique()\n\n                if (freq == 1) | (freq > 200):\n                    df = df.drop(col)\n\n        return df","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:20:19.425745Z","iopub.status.busy":"2024-03-11T20:20:19.425481Z","iopub.status.idle":"2024-03-11T20:20:19.437672Z","shell.execute_reply":"2024-03-11T20:20:19.436838Z"},"papermill":{"duration":0.023479,"end_time":"2024-03-11T20:20:19.439407","exception":false,"start_time":"2024-03-11T20:20:19.415928","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Automatic Aggregation","metadata":{"papermill":{"duration":0.008664,"end_time":"2024-03-11T20:20:19.457100","exception":false,"start_time":"2024-03-11T20:20:19.448436","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"Здесь я добавил несколько аггрегаций.","metadata":{}},{"cell_type":"code","source":"class Aggregator:\n    @staticmethod\n    def num_expr(df, agg):\n        cols = [col for col in df.columns if col[-1] in (\"P\", \"A\")]\n\n        expr_max = [agg(col).alias(f\"{agg.__name__}_{col}#\") for col in cols]\n\n        return expr_max\n\n    @staticmethod\n    def date_expr(df, agg):\n        cols = [col for col in df.columns if col[-1] in (\"D\",)]\n\n        expr_max = [agg(col).alias(f\"{agg.__name__}_{col}#\") for col in cols]\n\n        return expr_max\n\n    @staticmethod\n    def str_expr(df, agg):\n        cols = [col for col in df.columns if col[-1] in (\"M\",)]\n        \n        expr_max = [agg(col).alias(f\"{agg.__name__}_{col}#\") for col in cols]\n\n        return expr_max\n\n    @staticmethod\n    def other_expr(df, agg):\n        cols = [col for col in df.columns if col[-1] in (\"T\", \"L\")]\n        \n        expr_max = [agg(col).alias(f\"{agg.__name__}_{col}#\") for col in cols]\n\n        return expr_max\n    \n    @staticmethod\n    def count_expr(df, agg):\n        cols = [col for col in df.columns if \"num_group\" in col]\n\n        expr_max = [agg(col).alias(f\"{agg.__name__}_{col}#\") for col in cols]\n\n        return expr_max\n\n    @staticmethod\n    def get_exprs(df, agg):\n        exprs = Aggregator.num_expr(df, agg) + \\\n                Aggregator.date_expr(df, agg) + \\\n                Aggregator.other_expr(df, agg) + \\\n                Aggregator.count_expr(df, agg)\n        #                 Aggregator.str_expr(df, agg) + \\\n\n        return exprs","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:20:19.476571Z","iopub.status.busy":"2024-03-11T20:20:19.476292Z","iopub.status.idle":"2024-03-11T20:20:19.486434Z","shell.execute_reply":"2024-03-11T20:20:19.485637Z"},"papermill":{"duration":0.021569,"end_time":"2024-03-11T20:20:19.488325","exception":false,"start_time":"2024-03-11T20:20:19.466756","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### File I/O","metadata":{"papermill":{"duration":0.009719,"end_time":"2024-03-11T20:20:19.506836","exception":false,"start_time":"2024-03-11T20:20:19.497117","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def read_file(path, is_train, depth=None):\n    df = pl.read_parquet(path)\n    df = df.pipe(Pipeline.set_table_dtypes)\n    if depth is not None:\n        df = df.pipe(Pipeline.handle_dates1)\n    \n    if depth in [1, 2]:\n        df = df.group_by(\"case_id\").agg(Aggregator.get_exprs(df, pl.max) + Aggregator.get_exprs(df, pl.min) + \\\n                               Aggregator.get_exprs(df, pl.mean) + Aggregator.get_exprs(df, pl.median))\n    if is_train:\n        df = df.pipe(Pipeline.filter_cols)\n    return df\n\ndef read_files(regex_path, is_train, depth=None):\n    chunks = []\n    for path in glob(str(regex_path)):\n        df = read_file(path, False, depth)\n#         print(df.shape)\n        chunks.append(df)\n    \n    df = pl.concat(chunks, how=\"vertical_relaxed\")\n    df = df.unique(subset=[\"case_id\"])\n    if is_train:\n        df = df.pipe(Pipeline.filter_cols)\n\n    return df","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:20:19.525391Z","iopub.status.busy":"2024-03-11T20:20:19.525145Z","iopub.status.idle":"2024-03-11T20:20:19.531829Z","shell.execute_reply":"2024-03-11T20:20:19.531102Z"},"papermill":{"duration":0.018108,"end_time":"2024-03-11T20:20:19.533703","exception":false,"start_time":"2024-03-11T20:20:19.515595","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Feature Engineering","metadata":{"papermill":{"duration":0.00854,"end_time":"2024-03-11T20:20:19.550995","exception":false,"start_time":"2024-03-11T20:20:19.542455","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def feature_eng(df_base, depth_0, depth_1, depth_2):\n    df_base = (\n        df_base\n        .with_columns(\n            month_decision = pl.col(\"date_decision\").dt.month(),\n            weekday_decision = pl.col(\"date_decision\").dt.weekday(),\n        )\n    )\n        \n    for i, df in enumerate(depth_0 + depth_1 + depth_2):\n        df_base = df_base.join(df, how=\"left\", on=\"case_id\", suffix=f\"_{i}\")\n        \n    df_base = df_base.pipe(Pipeline.handle_dates2)\n    \n    return df_base","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:20:19.570143Z","iopub.status.busy":"2024-03-11T20:20:19.569547Z","iopub.status.idle":"2024-03-11T20:20:19.574976Z","shell.execute_reply":"2024-03-11T20:20:19.574273Z"},"papermill":{"duration":0.01688,"end_time":"2024-03-11T20:20:19.576796","exception":false,"start_time":"2024-03-11T20:20:19.559916","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def to_pandas(df_data, cat_cols=None):\n    df_data = df_data.to_pandas()\n    \n    if cat_cols is None:\n        cat_cols = list(df_data.select_dtypes(\"object\").columns)\n    \n    df_data[cat_cols] = df_data[cat_cols].astype(\"category\")\n    \n    return df_data, cat_cols","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:20:19.595389Z","iopub.status.busy":"2024-03-11T20:20:19.595137Z","iopub.status.idle":"2024-03-11T20:20:19.599640Z","shell.execute_reply":"2024-03-11T20:20:19.598867Z"},"papermill":{"duration":0.015752,"end_time":"2024-03-11T20:20:19.601370","exception":false,"start_time":"2024-03-11T20:20:19.585618","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Configuration","metadata":{"papermill":{"duration":0.00861,"end_time":"2024-03-11T20:20:19.619092","exception":false,"start_time":"2024-03-11T20:20:19.610482","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# ROOT            = Path(\"/kaggle/input/home-credit-credit-risk-model-stability\")\nROOT = Path(\".\")\nTRAIN_DIR       = ROOT / \"parquet_files\" / \"train\"\nTEST_DIR        = ROOT / \"parquet_files\" / \"test\"","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:20:19.637746Z","iopub.status.busy":"2024-03-11T20:20:19.637467Z","iopub.status.idle":"2024-03-11T20:20:19.641419Z","shell.execute_reply":"2024-03-11T20:20:19.640728Z"},"papermill":{"duration":0.015255,"end_time":"2024-03-11T20:20:19.643247","exception":false,"start_time":"2024-03-11T20:20:19.627992","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Train Files Read & Feature Engineering","metadata":{"papermill":{"duration":0.008715,"end_time":"2024-03-11T20:20:19.660605","exception":false,"start_time":"2024-03-11T20:20:19.651890","status":"completed"},"tags":[]}},{"cell_type":"code","source":"data_store = {\n    \"df_base\": read_file(TRAIN_DIR / \"train_base.parquet\", True),\n    \"depth_0\": [\n        read_file(TRAIN_DIR / \"train_static_cb_0.parquet\", True, 0),\n        read_files(TRAIN_DIR / \"train_static_0_*.parquet\", True, 0),\n    ],\n    \"depth_1\": [\n        read_files(TRAIN_DIR / \"train_applprev_1_*.parquet\", True, 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_a_1.parquet\", True, 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_b_1.parquet\", True, 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_c_1.parquet\", True, 1),\n        read_files(TRAIN_DIR / \"train_credit_bureau_a_1_*.parquet\", True, 1),\n        read_file(TRAIN_DIR / \"train_credit_bureau_b_1.parquet\", True, 1),\n        read_file(TRAIN_DIR / \"train_other_1.parquet\", True, 1),\n        read_file(TRAIN_DIR / \"train_person_1.parquet\", True, 1),\n        read_file(TRAIN_DIR / \"train_deposit_1.parquet\", True, 1),\n        read_file(TRAIN_DIR / \"train_debitcard_1.parquet\", True, 1),\n    ],\n    \"depth_2\": [\n        read_file(TRAIN_DIR / \"train_credit_bureau_b_2.parquet\", True, 2),\n        read_files(TRAIN_DIR / \"train_credit_bureau_a_2_*.parquet\", True, 2),\n    ]\n}","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:20:19.678920Z","iopub.status.busy":"2024-03-11T20:20:19.678686Z","iopub.status.idle":"2024-03-11T20:22:26.542632Z","shell.execute_reply":"2024-03-11T20:22:26.541531Z"},"papermill":{"duration":126.876005,"end_time":"2024-03-11T20:22:26.545249","exception":false,"start_time":"2024-03-11T20:20:19.669244","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = feature_eng(**data_store)\n\nprint(\"train data shape:\\t\", df_train.shape)","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:22:26.565337Z","iopub.status.busy":"2024-03-11T20:22:26.564999Z","iopub.status.idle":"2024-03-11T20:22:38.552398Z","shell.execute_reply":"2024-03-11T20:22:38.551267Z"},"papermill":{"duration":11.999532,"end_time":"2024-03-11T20:22:38.554478","exception":false,"start_time":"2024-03-11T20:22:26.554946","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Test Files Read & Feature Engineering","metadata":{"papermill":{"duration":0.009104,"end_time":"2024-03-11T20:22:38.572989","exception":false,"start_time":"2024-03-11T20:22:38.563885","status":"completed"},"tags":[]}},{"cell_type":"code","source":"data_store = {\n    \"df_base\": read_file(TEST_DIR / \"test_base.parquet\", False,),\n    \"depth_0\": [\n        read_file(TEST_DIR / \"test_static_cb_0.parquet\", False, 0),\n        read_files(TEST_DIR / \"test_static_0_*.parquet\", False, 0),\n    ],\n    \"depth_1\": [\n        read_files(TEST_DIR / \"test_applprev_1_*.parquet\", False, 1),\n        read_file(TEST_DIR / \"test_tax_registry_a_1.parquet\", False, 1),\n        read_file(TEST_DIR / \"test_tax_registry_b_1.parquet\", False, 1),\n        read_file(TEST_DIR / \"test_tax_registry_c_1.parquet\", False, 1),\n        read_files(TEST_DIR / \"test_credit_bureau_a_1_*.parquet\", False, 1),\n        read_file(TEST_DIR / \"test_credit_bureau_b_1.parquet\", False, 1),\n        read_file(TEST_DIR / \"test_other_1.parquet\", False, 1),\n        read_file(TEST_DIR / \"test_person_1.parquet\", False, 1),\n        read_file(TEST_DIR / \"test_deposit_1.parquet\", False, 1),\n        read_file(TEST_DIR / \"test_debitcard_1.parquet\", False, 1),\n    ],\n    \"depth_2\": [\n        read_file(TEST_DIR / \"test_credit_bureau_b_2.parquet\", False, 2),\n        read_files(TEST_DIR / \"test_credit_bureau_a_2_*.parquet\", False, 2),\n    ]\n}","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:22:38.592481Z","iopub.status.busy":"2024-03-11T20:22:38.591784Z","iopub.status.idle":"2024-03-11T20:22:39.168839Z","shell.execute_reply":"2024-03-11T20:22:39.167993Z"},"papermill":{"duration":0.589237,"end_time":"2024-03-11T20:22:39.171145","exception":false,"start_time":"2024-03-11T20:22:38.581908","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test = feature_eng(**data_store)\n\nprint(\"test data shape:\\t\", df_test.shape)","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:22:39.190830Z","iopub.status.busy":"2024-03-11T20:22:39.190540Z","iopub.status.idle":"2024-03-11T20:22:39.230327Z","shell.execute_reply":"2024-03-11T20:22:39.229525Z"},"papermill":{"duration":0.051709,"end_time":"2024-03-11T20:22:39.232155","exception":false,"start_time":"2024-03-11T20:22:39.180446","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Feature Elimination","metadata":{"papermill":{"duration":0.008776,"end_time":"2024-03-11T20:22:39.250093","exception":false,"start_time":"2024-03-11T20:22:39.241317","status":"completed"},"tags":[]}},{"cell_type":"code","source":"df_train = df_train.pipe(Pipeline.filter_cols)\ndf_test = df_test.select([col for col in df_train.columns if col != \"target\"])\n\nprint(\"train data shape:\\t\", df_train.shape)\nprint(\"test data shape:\\t\", df_test.shape)","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:22:39.269947Z","iopub.status.busy":"2024-03-11T20:22:39.269394Z","iopub.status.idle":"2024-03-11T20:22:42.015546Z","shell.execute_reply":"2024-03-11T20:22:42.014587Z"},"papermill":{"duration":2.758307,"end_time":"2024-03-11T20:22:42.017667","exception":false,"start_time":"2024-03-11T20:22:39.259360","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Pandas Conversion","metadata":{"papermill":{"duration":0.009039,"end_time":"2024-03-11T20:22:42.036115","exception":false,"start_time":"2024-03-11T20:22:42.027076","status":"completed"},"tags":[]}},{"cell_type":"code","source":"df_train, cat_cols = to_pandas(df_train)\ndf_test, cat_cols = to_pandas(df_test, cat_cols)","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:22:42.057274Z","iopub.status.busy":"2024-03-11T20:22:42.056714Z","iopub.status.idle":"2024-03-11T20:23:00.747089Z","shell.execute_reply":"2024-03-11T20:23:00.746301Z"},"papermill":{"duration":18.703961,"end_time":"2024-03-11T20:23:00.749238","exception":false,"start_time":"2024-03-11T20:22:42.045277","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Garbage Collection","metadata":{"papermill":{"duration":0.00925,"end_time":"2024-03-11T20:23:00.767926","exception":false,"start_time":"2024-03-11T20:23:00.758676","status":"completed"},"tags":[]}},{"cell_type":"code","source":"del data_store\n\ngc.collect()","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:23:00.787299Z","iopub.status.busy":"2024-03-11T20:23:00.786797Z","iopub.status.idle":"2024-03-11T20:23:00.911801Z","shell.execute_reply":"2024-03-11T20:23:00.911010Z"},"papermill":{"duration":0.137016,"end_time":"2024-03-11T20:23:00.913914","exception":false,"start_time":"2024-03-11T20:23:00.776898","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### EDA","metadata":{"papermill":{"duration":0.009084,"end_time":"2024-03-11T20:23:00.932380","exception":false,"start_time":"2024-03-11T20:23:00.923296","status":"completed"},"tags":[]}},{"cell_type":"code","source":"print(\"Train is duplicated:\\t\", df_train[\"case_id\"].duplicated().any())\nprint(\"Train Week Range:\\t\", (df_train[\"WEEK_NUM\"].min(), df_train[\"WEEK_NUM\"].max()))\n\nprint()\n\nprint(\"Test is duplicated:\\t\", df_test[\"case_id\"].duplicated().any())\nprint(\"Test Week Range:\\t\", (df_test[\"WEEK_NUM\"].min(), df_test[\"WEEK_NUM\"].max()))","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:23:00.952356Z","iopub.status.busy":"2024-03-11T20:23:00.951810Z","iopub.status.idle":"2024-03-11T20:23:00.978038Z","shell.execute_reply":"2024-03-11T20:23:00.977167Z"},"papermill":{"duration":0.038139,"end_time":"2024-03-11T20:23:00.979829","exception":false,"start_time":"2024-03-11T20:23:00.941690","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Training","metadata":{"papermill":{"duration":0.01291,"end_time":"2024-03-11T20:23:17.567408","exception":false,"start_time":"2024-03-11T20:23:17.554498","status":"completed"},"tags":[]}},{"cell_type":"code","source":"X = df_train.drop(columns=[\"target\", \"case_id\", \"WEEK_NUM\"])\ny = df_train[\"target\"]\nweeks = df_train[\"WEEK_NUM\"]\n\ncv = StratifiedGroupKFold(n_splits=5, shuffle=False)\n\nparams = {\n    \"boosting_type\": \"gbdt\",\n    \"objective\": \"binary\",\n    \"metric\": \"auc\",\n    \"max_depth\": 8,\n    \"learning_rate\": 0.05,\n    \"n_estimators\": 1000,\n    \"colsample_bytree\": 0.8, \n    \"colsample_bynode\": 0.8,\n    \"verbose\": -1,\n    \"random_state\": 42,\n#     \"device\": \"gpu\",\n}\n\nfitted_models = []\n\nfor idx_train, idx_valid in cv.split(X, y, groups=weeks):\n    X_train, y_train = X.iloc[idx_train], y.iloc[idx_train]\n    X_valid, y_valid = X.iloc[idx_valid], y.iloc[idx_valid]\n\n    model = lgb.LGBMClassifier(**params)\n    model.fit(\n        X_train, y_train,\n        eval_set=[(X_valid, y_valid)],\n        callbacks=[lgb.log_evaluation(100), lgb.early_stopping(100)]\n    )\n\n    fitted_models.append(model)\n\nmodel = VotingModel(fitted_models)","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:23:17.595206Z","iopub.status.busy":"2024-03-11T20:23:17.594787Z","iopub.status.idle":"2024-03-11T20:40:29.888472Z","shell.execute_reply":"2024-03-11T20:40:29.887312Z"},"papermill":{"duration":1032.311042,"end_time":"2024-03-11T20:40:29.890585","exception":false,"start_time":"2024-03-11T20:23:17.579543","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"После агреггаций у меня вышло слишком много фичей, и чтобы все поместилось в оперативу на каггл я решил выбрать лучшие с помощью feature_importances_","metadata":{}},{"cell_type":"code","source":"import lightgbm\n\nlightgbm.plot_importance(fitted_models[0])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\nfeature_importances = np.zeros((len(fitted_models[0].feature_importances_), len(fitted_models)))\n\nfor i, model in enumerate(fitted_models):\n    feature_importances[:, i] = model.feature_importances_\n\navg_feature_importances = np.mean(feature_importances, axis=1)\n\nfeatures_df = pd.DataFrame({\n    'Feature': fitted_models[0].feature_name_,\n    'Importance': avg_feature_importances\n})\n\nfeatures_df = features_df.sort_values(by='Importance', ascending=False)\n\nN = 300\ntop_features = features_df.head(N)\n\nprint(top_features)\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Посмотрел на этот график и решил, что возьмем топ 50 которые вносят основной импат, это может иметь смысл если мы переобучаемся из-за остальных фичей. Также решил взять топ 300, вообще можно брать и все, но это протсо не влазит на каггл","metadata":{}},{"cell_type":"code","source":"plt.plot(features_df.Importance.to_numpy())","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### HP tuning","metadata":{}},{"cell_type":"markdown","source":"Здесь я перебираю параметры по гайду с каггл","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import roc_auc_score\n\ndef gini_stability(base, w_fallingrate=88.0, w_resstd=-0.5):\n    gini_in_time = (\n        base.loc[:, [\"WEEK_NUM\", \"target\", \"score\"]]\n        .sort_values(\"WEEK_NUM\")\n        .groupby(\"WEEK_NUM\")[[\"target\", \"score\"]]\n        .apply(lambda x: 2 * roc_auc_score(x[\"target\"], x[\"score\"]) - 1)\n        .tolist()\n    )\n\n    x = np.arange(len(gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a * x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    return avg_gini + w_fallingrate * min(0, a) + w_resstd * res_std, gini_in_time\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X300 = df_train[top_features.head(300).Feature.to_list()]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from optuna.integration import LightGBMPruningCallback\nimport optuna\nimport json\n\ndef objective(trial, X, y, weeks, trial_name):\n    param_grid = {\n        # \"device_type\": trial.suggest_categorical(\"device_type\", ['gpu']),\n        \"n_estimators\": trial.suggest_categorical(\"n_estimators\", [5000]),\n        \"learning_rate\": trial.suggest_float(\"learning_rate\", 0.01, 0.3),\n        \"num_leaves\": trial.suggest_int(\"num_leaves\", 20, 3000, step=20),\n        \"max_depth\": trial.suggest_int(\"max_depth\", 3, 12),\n        \"min_data_in_leaf\": trial.suggest_int(\"min_data_in_leaf\", 200, 10000, step=100),\n        \"lambda_l1\": trial.suggest_int(\"lambda_l1\", 0, 100, step=5),\n        \"lambda_l2\": trial.suggest_int(\"lambda_l2\", 0, 100, step=5),\n        \"min_gain_to_split\": trial.suggest_float(\"min_gain_to_split\", 0, 15),\n        \"bagging_fraction\": trial.suggest_float(\n            \"bagging_fraction\", 0.2, 0.95, step=0.1\n        ),\n        \"bagging_freq\": trial.suggest_categorical(\"bagging_freq\", [1]),\n        \"feature_fraction\": trial.suggest_float(\n            \"feature_fraction\", 0.2, 0.95, step=0.1\n        ),\n    }\n\n    cv = StratifiedGroupKFold(n_splits=5, shuffle=False)\n\n    gini_scores = np.empty(5)\n    for idx, (train_idx, test_idx) in enumerate(cv.split(X, y, groups=weeks)):\n        X_train, X_test = X.iloc[train_idx], X.iloc[test_idx]\n        y_train, y_test = y[train_idx], y[test_idx]\n\n        model = lgb.LGBMClassifier(objective=\"binary\", boosting_type=\"gbdt\", verbose=-1, **param_grid)\n        model.fit(\n            X_train,\n            y_train,\n            eval_set=[(X_test, y_test)],\n            eval_metric=\"auc\",\n            callbacks=[\n                lgb.log_evaluation(100), lgb.early_stopping(100)\n            ],\n        )\n        \n        preds = model.predict_proba(X_test)[:, 1]\n        base = df_train.iloc[test_idx][[\"WEEK_NUM\", \"target\"]]\n        base[\"score\"] = preds\n        gini_scores[idx] = gini_stability(base)[0]\n        \n    with open(f\"{trial_name}.txt\", \"w\") as f:\n        print(np.mean(gini_scores), file=f)\n        print(json.dumps(param_grid, indent=4), file=f)\n        print(file=f)\n\n    return np.mean(gini_scores)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X300 = df_train[top_features.head(300).Feature.to_list()]\nstudy300 = optuna.create_study(direction=\"maximize\", study_name=\"LGBM Classifier 300 feature set\")\nfunc = lambda trial: objective(trial, X300, y, weeks, \"final_300_feats\")\nstudy300.optimize(func, n_trials=100)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X50 = df_train[top_features.head(50).Feature.to_list()]\nstudy50 = optuna.create_study(direction=\"maximize\", study_name=\"LGBM Classifier 50 feature set\")\nfunc = lambda trial: objective(trial, X50, y, weeks, \"final_50_feats\")\nstudy50.optimize(func, n_trials=100)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"study300.best_trial.params","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"study50.best_trial.params","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Для 50 фичей результаты были плохие, поэтому я забросил эту идею","metadata":{}},{"cell_type":"markdown","source":"### Cat columns processing","metadata":{}},{"cell_type":"markdown","source":"Я планировал применить трансформацию к категориальным признакам, чтобы для каждой категории и пользователя подсчитывать количество её появлений в каждом столбце. К сожалению, в финальной попытке мне не удалось это использовать из-за того, это сильно увеличивало количество признаков.","metadata":{}},{"cell_type":"code","source":"def flat_cat(df, cat_col, num_users=1526659, id_col=\"case_id\"):\n    df = df[[id_col, cat_col]]\n    cat_count_df = df.group_by([id_col, cat_col]).len().group_by(cat_col).len()\n    cat_count_df = cat_count_df.filter((cat_count_df[\"len\"] / num_users) > 0.05)\n    most_importance = cat_count_df[cat_col]\n    \n    df = df.filter(df[cat_col].is_in(most_importance))\n    df = df.with_columns((pl.col(cat_col).cast(pl.String) + f\"_{cat_col}\").alias(cat_col))\n    \n    return df.pivot(values=cat_col, index=id_col, columns=cat_col, aggregate_function=\"len\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def process_df_flat_cat(df):\n    chunks = []\n    for col in df.columns:\n        if col[-1] in (\"M\", ):\n#             print(col)\n            chunks.append(flat_cat(df, col))     \n    df_base = chunks[0]\n    for i, df_tmp in enumerate(chunks[1:]):\n        df_base = df_base.join(df_tmp, how=\"outer\", on=\"case_id\").drop(\"case_id_right\")\n    return df_base","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Пример работы\n\nprocess_df_flat_cat(pl.read_parquet(\"parquet_files/train/train_credit_bureau_a_1_*.parquet\"))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Prediction","metadata":{"papermill":{"duration":0.015335,"end_time":"2024-03-11T20:40:29.921663","exception":false,"start_time":"2024-03-11T20:40:29.906328","status":"completed"},"tags":[]}},{"cell_type":"code","source":"X_test = df_test.drop(columns=[\"WEEK_NUM\"])\nX_test = X_test.set_index(\"case_id\")\n\ny_pred = pd.Series(model.predict_proba(X_test)[:, 1], index=X_test.index)","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:40:29.953970Z","iopub.status.busy":"2024-03-11T20:40:29.953671Z","iopub.status.idle":"2024-03-11T20:40:30.231819Z","shell.execute_reply":"2024-03-11T20:40:30.230927Z"},"papermill":{"duration":0.29704,"end_time":"2024-03-11T20:40:30.234187","exception":false,"start_time":"2024-03-11T20:40:29.937147","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Submission","metadata":{"papermill":{"duration":0.015506,"end_time":"2024-03-11T20:40:30.266313","exception":false,"start_time":"2024-03-11T20:40:30.250807","status":"completed"},"tags":[]}},{"cell_type":"code","source":"df_subm = pd.read_csv(ROOT / \"sample_submission.csv\")\ndf_subm = df_subm.set_index(\"case_id\")\n\ndf_subm[\"score\"] = y_pred","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:40:30.299204Z","iopub.status.busy":"2024-03-11T20:40:30.298476Z","iopub.status.idle":"2024-03-11T20:40:30.315227Z","shell.execute_reply":"2024-03-11T20:40:30.314249Z"},"papermill":{"duration":0.035461,"end_time":"2024-03-11T20:40:30.317145","exception":false,"start_time":"2024-03-11T20:40:30.281684","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Check null: \", df_subm[\"score\"].isnull().any())\n\ndf_subm.head()","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:40:30.349432Z","iopub.status.busy":"2024-03-11T20:40:30.349153Z","iopub.status.idle":"2024-03-11T20:40:30.362280Z","shell.execute_reply":"2024-03-11T20:40:30.361363Z"},"papermill":{"duration":0.031645,"end_time":"2024-03-11T20:40:30.364269","exception":false,"start_time":"2024-03-11T20:40:30.332624","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_subm.to_csv(\"submission.csv\")","metadata":{"execution":{"iopub.execute_input":"2024-03-11T20:40:30.397489Z","iopub.status.busy":"2024-03-11T20:40:30.397241Z","iopub.status.idle":"2024-03-11T20:40:30.403177Z","shell.execute_reply":"2024-03-11T20:40:30.402519Z"},"papermill":{"duration":0.024175,"end_time":"2024-03-11T20:40:30.404950","exception":false,"start_time":"2024-03-11T20:40:30.380775","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"papermill":{"duration":0.015608,"end_time":"2024-03-11T20:40:30.436155","exception":false,"start_time":"2024-03-11T20:40:30.420547","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]}]}