{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":50160,"databundleVersionId":7921029,"sourceType":"competition"}],"dockerImageVersionId":30648,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"I just wanted to understand [@greysky](https://www.kaggle.com/greysky)'s amazing [Home Credit Baseline](https://www.kaggle.com/code/greysky/home-credit-baseline) notebook.\n\n- 메모리 문제로 확인 코드와 함께하면 training이 돌지 못함 -> 노트북 2개로 분리\n    - Home Credit Baseline (1/2, Explained, KR) : 이 노트북. training 전까지의 코드 설명\n    - [Home Credit Baseline (2/2, Explained, KR)](https://www.kaggle.com/code/sunghoshim/home-credit-baseline-2-2-explained-kr): 앞부분 코드는 스킵하고 training","metadata":{}},{"cell_type":"markdown","source":"### Reference\n- {Notebook} [Home Credit Baseline](https://www.kaggle.com/code/greysky/home-credit-baseline)\n- {Notebook} [Home Credit Risk (LightGBM)](https://www.kaggle.com/code/daviddirethucus/home-credit-risk-lightgbm/notebook)","metadata":{}},{"cell_type":"code","source":"import os\nimport gc\nimport joblib\nfrom glob import glob\nfrom pathlib import Path\nfrom datetime import datetime\n\nimport numpy as np\nimport pandas as pd\nimport polars as pl\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom sklearn.model_selection import StratifiedGroupKFold\nfrom sklearn.base import BaseEstimator, ClassifierMixin\nfrom sklearn.metrics import roc_auc_score\n\nimport lightgbm as lgb\n\nimport warnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-03-31T03:19:43.403542Z","iopub.execute_input":"2024-03-31T03:19:43.403988Z","iopub.status.idle":"2024-03-31T03:19:43.411324Z","shell.execute_reply.started":"2024-03-31T03:19:43.403959Z","shell.execute_reply":"2024-03-31T03:19:43.410358Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 1. Prepare Utility Classes","metadata":{}},{"cell_type":"markdown","source":"## . . 1.1. Pre-Fitted Voting Model\n- 요렇게 모델을 구현 sklearn의 다른 클래스들에서 바로 사용할 수 있나 봄\n- 요 노트북에서는 cv의 모델들을 평균 내려고 사용 (`predict_proba()`)\n- Estimator : Regressor, Classifier 등을 포함하는 상위 개념\n- {sklearn docs} [BaseEstimator](https://scikit-learn.org/stable/modules/generated/sklearn.base.BaseEstimator.html)\n- {sklearn docs} [ClassifierMixin](https://scikit-learn.org/stable/modules/generated/sklearn.base.ClassifierMixin.html)\n- {sklearn docs} [Developing scikit-learn estimators > Rolling your own estimator](https://scikit-learn.org/stable/developers/develop.html#rolling-your-own-estimator)","metadata":{}},{"cell_type":"code","source":"class VotingModel(BaseEstimator, ClassifierMixin):\n    def __init__(self, estimators):\n        super().__init__()\n        self.estimators = estimators\n        \n    def fit(self, X, y=None):\n        return self\n    \n    def predict(self, X):\n        y_preds = [estimator.predict(X) for estimator in self.estimators]\n        return np.mean(y_preds, axis=0)\n    \n    def predict_proba(self, X):\n        y_preds = [estimator.predict_proba(X) for estimator in self.estimators]\n        return np.mean(y_preds, axis=0)","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:14.434716Z","iopub.execute_input":"2024-03-31T03:00:14.435467Z","iopub.status.idle":"2024-03-31T03:00:14.442497Z","shell.execute_reply.started":"2024-03-31T03:00:14.435432Z","shell.execute_reply":"2024-03-31T03:00:14.441517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## . . 1.2. Pipeline\n- 전처리 관련 함수들","metadata":{}},{"cell_type":"code","source":"class Pipeline:\n    @staticmethod\n    def set_table_dtypes(df):\n        \"\"\"\n        column 명에 따라 해당 column의 데이터타입 지정\n        \"\"\"\n        for col in df.columns:\n            if col in [\"case_id\", \"WEEK_NUM\", \"num_group1\", \"num_group2\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Int32))\n            elif col in [\"date_decision\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Date))\n            elif col[-1] in (\"P\", \"A\"):\n                df = df.with_columns(pl.col(col).cast(pl.Float64))\n            elif col[-1] in (\"M\",):\n                df = df.with_columns(pl.col(col).cast(pl.String))\n            elif col[-1] in (\"D\",):\n                df = df.with_columns(pl.col(col).cast(pl.Date))            \n\n        return df\n    \n    @staticmethod\n    def handle_dates(df):\n        \"\"\"\n        `date_decision` 기준으로 지난 날짜로 표현\n        \"\"\"\n        for col in df.columns:\n            if col[-1] in (\"D\",):\n                df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))\n                df = df.with_columns(pl.col(col).dt.total_days().cast(pl.Int32))\n                \n        df = df.drop(\"date_decision\", \"MONTH\")\n\n        return df\n    \n    @staticmethod\n    def filter_cols(df):\n        \"\"\"\n        null이 너무 많거나, 다 똑같거나, 다른 게 너무 많으면 그 column은 무시함\n        \"\"\"\n        for col in df.columns:\n            if col not in [\"target\", \"case_id\", \"WEEK_NUM\"]:\n                isnull = df[col].is_null().mean()\n\n                if isnull > 0.7:  # 0.7로 바꿨음\n                    print(f'Pipeline.filter_cols(). drop {col} column. (isnull ratio: {isnull})')\n                    df = df.drop(col)\n\n        for col in df.columns:\n            if (col not in [\"target\", \"case_id\", \"WEEK_NUM\"]) & (df[col].dtype == pl.String):\n                freq = df[col].n_unique()\n\n                if (freq == 1) | (freq > 200):\n                    print(f'Pipeline.filter_cols(). drop {col} column. (n_unique: {freq})')\n                    df = df.drop(col)\n\n        return df","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:14.444081Z","iopub.execute_input":"2024-03-31T03:00:14.444398Z","iopub.status.idle":"2024-03-31T03:00:14.464459Z","shell.execute_reply.started":"2024-03-31T03:00:14.444373Z","shell.execute_reply":"2024-03-31T03:00:14.463534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Pipeline 확인","metadata":{}},{"cell_type":"code","source":"# Show head(2) and tail(2) only\n# https://docs.pola.rs/py-polars/html/reference/config.html\n# https://docs.pola.rs/py-polars/html/reference/api/polars.Config.set_tbl_rows.html\npl.Config.set_tbl_rows(4)","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:14.466652Z","iopub.execute_input":"2024-03-31T03:00:14.466946Z","iopub.status.idle":"2024-03-31T03:00:14.484296Z","shell.execute_reply.started":"2024-03-31T03:00:14.466923Z","shell.execute_reply":"2024-03-31T03:00:14.483296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_temp_base = pl.read_parquet('/kaggle/input/home-credit-credit-risk-model-stability/parquet_files/train/train_base.parquet')\ndf_temp_base","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:14.485535Z","iopub.execute_input":"2024-03-31T03:00:14.485855Z","iopub.status.idle":"2024-03-31T03:00:14.740062Z","shell.execute_reply.started":"2024-03-31T03:00:14.485813Z","shell.execute_reply":"2024-03-31T03:00:14.739147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Pipeline.set_table_dtypes() 적용 => 타입 바뀌었음\ndf_temp_base = df_temp_base.pipe(Pipeline.set_table_dtypes)\ndf_temp_base","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:14.741346Z","iopub.execute_input":"2024-03-31T03:00:14.741648Z","iopub.status.idle":"2024-03-31T03:00:15.148755Z","shell.execute_reply.started":"2024-03-31T03:00:14.741624Z","shell.execute_reply":"2024-03-31T03:00:15.147792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# train_static_0 테이블\ndf_temp_static_0_0 = pl.read_parquet(\n    '/kaggle/input/home-credit-credit-risk-model-stability/parquet_files/train/train_static_0_0.parquet'\n).pipe(Pipeline.set_table_dtypes)\nprint('df_temp_static_0_0.shape:', df_temp_static_0_0.shape)\n\ndf_temp_static_0_1 = pl.read_parquet(\n    '/kaggle/input/home-credit-credit-risk-model-stability/parquet_files/train/train_static_0_1.parquet'\n).pipe(Pipeline.set_table_dtypes)\nprint('df_temp_static_0_1.shape:', df_temp_static_0_1.shape)\n\ndf_temp_static_0 = pl.concat(\n    [df_temp_static_0_0, df_temp_static_0_1],\n    how = 'vertical'\n)\ndf_temp_static_0","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:15.150187Z","iopub.execute_input":"2024-03-31T03:00:15.150580Z","iopub.status.idle":"2024-03-31T03:00:21.449040Z","shell.execute_reply.started":"2024-03-31T03:00:15.150546Z","shell.execute_reply":"2024-03-31T03:00:21.448045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_temp_joined = df_temp_base.join(\n    df_temp_static_0,\n    on = 'case_id',\n    how = 'left',\n)\ndf_temp_joined","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:21.450174Z","iopub.execute_input":"2024-03-31T03:00:21.450518Z","iopub.status.idle":"2024-03-31T03:00:22.577624Z","shell.execute_reply.started":"2024-03-31T03:00:21.450492Z","shell.execute_reply":"2024-03-31T03:00:22.576593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 'D'로 끝나는 column들 확인\ndf_temp_joined.select(\n    pl.col('^.*D$')\n)","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:22.578818Z","iopub.execute_input":"2024-03-31T03:00:22.579104Z","iopub.status.idle":"2024-03-31T03:00:22.588693Z","shell.execute_reply.started":"2024-03-31T03:00:22.579080Z","shell.execute_reply":"2024-03-31T03:00:22.587682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# hadle_dates() 적용. \ndf_temp_handle_dates = df_temp_joined.pipe(Pipeline.handle_dates)\nprint('df_temp_handle_dates.shape:', df_temp_handle_dates.shape)\ndf_temp_handle_dates.select(\n    pl.col('^.*D$')\n)","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:22.592648Z","iopub.execute_input":"2024-03-31T03:00:22.592979Z","iopub.status.idle":"2024-03-31T03:00:23.786372Z","shell.execute_reply.started":"2024-03-31T03:00:22.592956Z","shell.execute_reply":"2024-03-31T03:00:23.785425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# filter_cols() 적용.\ndf_temp_handle_dates = df_temp_handle_dates.pipe(Pipeline.filter_cols)\ndf_temp_handle_dates","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:23.787767Z","iopub.execute_input":"2024-03-31T03:00:23.788064Z","iopub.status.idle":"2024-03-31T03:00:24.578521Z","shell.execute_reply.started":"2024-03-31T03:00:23.788041Z","shell.execute_reply":"2024-03-31T03:00:24.577563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## . . 1.3. Automatic Aggregation\n- depth1, depth2 테이블에서 groupby 해서 뽑아낼 column 들을 명시하기 위한 expressions 만듬\n- max() 값을 취하고 컬럼명에 'max_' 를 prefix 로 붙임","metadata":{}},{"cell_type":"code","source":"class Aggregator:\n    @staticmethod\n    def num_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"P\", \"A\")]\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        return expr_max\n\n    @staticmethod\n    def date_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"D\",)]\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        return expr_max\n\n    @staticmethod\n    def str_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"M\",)]\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        return expr_max\n\n    @staticmethod\n    def other_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"T\", \"L\")]\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        return expr_max\n    \n    @staticmethod\n    def count_expr(df):\n        cols = [col for col in df.columns if \"num_group\" in col]\n        expr_max = [pl.max(col).alias(f\"max_{col}\") for col in cols]\n        return expr_max\n\n    @staticmethod\n    def get_exprs(df):\n        exprs = Aggregator.num_expr(df) + \\\n                Aggregator.date_expr(df) + \\\n                Aggregator.str_expr(df) + \\\n                Aggregator.other_expr(df) + \\\n                Aggregator.count_expr(df)\n\n        return exprs","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:24.579544Z","iopub.execute_input":"2024-03-31T03:00:24.579811Z","iopub.status.idle":"2024-03-31T03:00:24.591163Z","shell.execute_reply.started":"2024-03-31T03:00:24.579784Z","shell.execute_reply":"2024-03-31T03:00:24.590138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Aggregator 확인","metadata":{}},{"cell_type":"code","source":"df_temp_person_1 = pl.read_parquet(\n    '/kaggle/input/home-credit-credit-risk-model-stability/parquet_files/train/train_person_1.parquet'\n).pipe(Pipeline.set_table_dtypes)\ndf_temp_person_1","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:24.592517Z","iopub.execute_input":"2024-03-31T03:00:24.592851Z","iopub.status.idle":"2024-03-31T03:00:26.630802Z","shell.execute_reply.started":"2024-03-31T03:00:24.592825Z","shell.execute_reply":"2024-03-31T03:00:26.629821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"exprs = Aggregator.num_expr(df_temp_person_1)\nfor expr in exprs:\n    print(expr)","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:26.632047Z","iopub.execute_input":"2024-03-31T03:00:26.632366Z","iopub.status.idle":"2024-03-31T03:00:26.637422Z","shell.execute_reply.started":"2024-03-31T03:00:26.632340Z","shell.execute_reply":"2024-03-31T03:00:26.636423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_temp_person_1.group_by('case_id').agg(\n    exprs\n)","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:26.638676Z","iopub.execute_input":"2024-03-31T03:00:26.639026Z","iopub.status.idle":"2024-03-31T03:00:26.853215Z","shell.execute_reply.started":"2024-03-31T03:00:26.639002Z","shell.execute_reply":"2024-03-31T03:00:26.852015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"exprs = Aggregator.date_expr(df_temp_person_1)\nfor expr in exprs:\n    print(expr)","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:26.854951Z","iopub.execute_input":"2024-03-31T03:00:26.855331Z","iopub.status.idle":"2024-03-31T03:00:26.860580Z","shell.execute_reply.started":"2024-03-31T03:00:26.855293Z","shell.execute_reply":"2024-03-31T03:00:26.859472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"exprs = Aggregator.str_expr(df_temp_person_1)\nfor expr in exprs:\n    print(expr)","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:26.862018Z","iopub.execute_input":"2024-03-31T03:00:26.862674Z","iopub.status.idle":"2024-03-31T03:00:26.873349Z","shell.execute_reply.started":"2024-03-31T03:00:26.862641Z","shell.execute_reply":"2024-03-31T03:00:26.872437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"exprs = Aggregator.other_expr(df_temp_person_1)\nfor expr in exprs:\n    print(expr)","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:26.874581Z","iopub.execute_input":"2024-03-31T03:00:26.874937Z","iopub.status.idle":"2024-03-31T03:00:26.887478Z","shell.execute_reply.started":"2024-03-31T03:00:26.874909Z","shell.execute_reply":"2024-03-31T03:00:26.886415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"exprs = Aggregator.count_expr(df_temp_person_1)\nfor expr in exprs:\n    print(expr)","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:26.889063Z","iopub.execute_input":"2024-03-31T03:00:26.889422Z","iopub.status.idle":"2024-03-31T03:00:26.899751Z","shell.execute_reply.started":"2024-03-31T03:00:26.889392Z","shell.execute_reply":"2024-03-31T03:00:26.898799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"exprs = Aggregator.get_exprs(df_temp_person_1)\nprint('len(exprs):', len(exprs))\nexprs[:5]","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:26.901401Z","iopub.execute_input":"2024-03-31T03:00:26.901755Z","iopub.status.idle":"2024-03-31T03:00:26.913350Z","shell.execute_reply.started":"2024-03-31T03:00:26.901725Z","shell.execute_reply":"2024-03-31T03:00:26.912428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_temp_person_1_agg = df_temp_person_1.group_by('case_id').agg(\n    exprs\n)\ndf_temp_person_1_agg","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:26.914725Z","iopub.execute_input":"2024-03-31T03:00:26.915106Z","iopub.status.idle":"2024-03-31T03:00:29.565967Z","shell.execute_reply.started":"2024-03-31T03:00:26.915074Z","shell.execute_reply":"2024-03-31T03:00:29.564976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## . . 1.4. File I/O\n- 테이블 읽고 dtypes 변경\n- depth 있는 테이블이면, `group_by()` 적용해서, `case_id` 별로 한 row만 있도록 함\n- `read_files()` 는 여러 테이블들 concat 까지","metadata":{}},{"cell_type":"code","source":"def read_file(path, depth=None):\n    print(f\"- read_file(). {str(path).split('/')[-1]}\")\n    df = pl.read_parquet(path)\n    df = df.pipe(Pipeline.set_table_dtypes)\n    \n    if depth in [1, 2]:\n        df = df.group_by(\"case_id\").agg(Aggregator.get_exprs(df))\n    \n    return df\n\ndef read_files(regex_path, depth=None):\n    print(f\"- read_files(). {str(regex_path).split('/')[-1]}\")\n    chunks = []\n    for path in glob(str(regex_path)):\n        print(f\"  . read_parquet(). {path.split('/')[-1]}\")\n        df = pl.read_parquet(path)\n        df = df.pipe(Pipeline.set_table_dtypes)\n        \n        if depth in [1, 2]:\n            df = df.group_by(\"case_id\").agg(Aggregator.get_exprs(df))\n        \n        chunks.append(df)\n        \n    df = pl.concat(chunks, how=\"vertical_relaxed\")\n    df = df.unique(subset=[\"case_id\"])\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:29.567338Z","iopub.execute_input":"2024-03-31T03:00:29.567623Z","iopub.status.idle":"2024-03-31T03:00:29.578516Z","shell.execute_reply.started":"2024-03-31T03:00:29.567599Z","shell.execute_reply":"2024-03-31T03:00:29.577570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### File I/O 확인","metadata":{}},{"cell_type":"code","source":"df_temp_person_1 = read_file(\n    '/kaggle/input/home-credit-credit-risk-model-stability/parquet_files/train/train_person_1.parquet',\n    depth=1,\n)\ndf_temp_person_1","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:29.579693Z","iopub.execute_input":"2024-03-31T03:00:29.580645Z","iopub.status.idle":"2024-03-31T03:00:34.468676Z","shell.execute_reply.started":"2024-03-31T03:00:29.580618Z","shell.execute_reply":"2024-03-31T03:00:34.467630Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"regex_path = (\n    '/kaggle/input/home-credit-credit-risk-model-stability/parquet_files/train/' +\n    'train_static_0_*.parquet'\n)\nregex_path","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:34.469998Z","iopub.execute_input":"2024-03-31T03:00:34.470337Z","iopub.status.idle":"2024-03-31T03:00:34.476767Z","shell.execute_reply.started":"2024-03-31T03:00:34.470309Z","shell.execute_reply":"2024-03-31T03:00:34.475763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"matching_files = glob(regex_path)\nmatching_files","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:34.477860Z","iopub.execute_input":"2024-03-31T03:00:34.478193Z","iopub.status.idle":"2024-03-31T03:00:34.503560Z","shell.execute_reply.started":"2024-03-31T03:00:34.478167Z","shell.execute_reply":"2024-03-31T03:00:34.502520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_temp_static_0 = read_files(regex_path)\ndf_temp_static_0","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:34.504886Z","iopub.execute_input":"2024-03-31T03:00:34.505279Z","iopub.status.idle":"2024-03-31T03:00:42.432401Z","shell.execute_reply.started":"2024-03-31T03:00:34.505218Z","shell.execute_reply":"2024-03-31T03:00:42.431365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del df_temp_person_1_agg\ndel df_temp_person_1\ndel df_temp_handle_dates\ndel df_temp_joined\ndel df_temp_static_0\ndel df_temp_static_0_0\ndel df_temp_static_0_1\ndel df_temp_base\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:42.433659Z","iopub.execute_input":"2024-03-31T03:00:42.433967Z","iopub.status.idle":"2024-03-31T03:00:42.719111Z","shell.execute_reply.started":"2024-03-31T03:00:42.433940Z","shell.execute_reply":"2024-03-31T03:00:42.718169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## . . 1.5. Feature Engineering","metadata":{}},{"cell_type":"code","source":"def feature_eng(df_base, depth_0, depth_1, depth_2):\n    '''\n    테이블 들 다 전처리, join 해서 하나의 table 로 만듬\n    - df_base 는 DataFrame\n    - 나머지는 DataFrame의 list\n    '''\n    df_base = (\n        df_base\n        .with_columns(\n            month_decision = pl.col(\"date_decision\").dt.month(),\n            weekday_decision = pl.col(\"date_decision\").dt.weekday(),\n        )\n    )\n    \n    print('len(depth_0 + depth_1 + depth_2):', len(depth_0 + depth_1 + depth_2))\n\n    # DataFrame 하나씩 df_base 에 join. depth 알 수 있게 suffix 붙임\n    for i, df in enumerate(depth_0 + depth_1 + depth_2):\n        df_base = df_base.join(df, how=\"left\", on=\"case_id\", suffix=f\"_{i}\")\n\n    df_base = df_base.pipe(Pipeline.handle_dates)\n    \n    return df_base","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:42.724845Z","iopub.execute_input":"2024-03-31T03:00:42.725197Z","iopub.status.idle":"2024-03-31T03:00:42.733463Z","shell.execute_reply.started":"2024-03-31T03:00:42.725168Z","shell.execute_reply":"2024-03-31T03:00:42.732484Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def to_pandas(df_data, cat_cols=None):\n    '''\n    다른 패키지에서 쓸 수 있다록 pandas Dataframe 으로 바꿈.\n    - cat_cols : 요걸로 columns 받아서 category 타입으로 바꿈 (test dataset)\n    - train dataset 에서는 (cat_cols=None) object 타입인 애들은 category 타입으로 바꾸고, 해당 cat_cols 도 return 해줌\n    '''\n    df_data = df_data.to_pandas()\n    \n    if cat_cols is None:\n        cat_cols = list(df_data.select_dtypes(\"object\").columns)\n    \n    df_data[cat_cols] = df_data[cat_cols].astype(\"category\")\n    \n    return df_data, cat_cols","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:42.734625Z","iopub.execute_input":"2024-03-31T03:00:42.734937Z","iopub.status.idle":"2024-03-31T03:00:42.751086Z","shell.execute_reply.started":"2024-03-31T03:00:42.734912Z","shell.execute_reply":"2024-03-31T03:00:42.750279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2. Prepare Datasets","metadata":{}},{"cell_type":"markdown","source":"## . . 2.1. Configuration","metadata":{}},{"cell_type":"code","source":"ROOT            = Path(\"/kaggle/input/home-credit-credit-risk-model-stability\")\nTRAIN_DIR       = ROOT / \"parquet_files\" / \"train\"\nTEST_DIR        = ROOT / \"parquet_files\" / \"test\"","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:42.752225Z","iopub.execute_input":"2024-03-31T03:00:42.752867Z","iopub.status.idle":"2024-03-31T03:00:42.763809Z","shell.execute_reply.started":"2024-03-31T03:00:42.752833Z","shell.execute_reply":"2024-03-31T03:00:42.762828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## . . 2.2. Train Files Read & Feature Engineering\n- https://www.kaggle.com/code/daviddirethucus/home-credit-risk-lightgbm?scriptVersionId=169160112&cellId=10","metadata":{}},{"cell_type":"code","source":"%%time\n\ndata_store = {\n    \"df_base\": read_file(TRAIN_DIR / \"train_base.parquet\"),\n    \"depth_0\": [\n        read_file(TRAIN_DIR / \"train_static_cb_0.parquet\"),\n        read_files(TRAIN_DIR / \"train_static_0_*.parquet\"),\n    ],\n    \"depth_1\": [\n        read_files(TRAIN_DIR / \"train_applprev_1_*.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_a_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_b_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_c_1.parquet\", 1),\n        read_files(TRAIN_DIR / \"train_credit_bureau_a_1_*.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_credit_bureau_b_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_other_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_person_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_deposit_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_debitcard_1.parquet\", 1),\n    ],\n    \"depth_2\": [\n        read_file(TRAIN_DIR / \"train_credit_bureau_b_2.parquet\", 2),\n        read_files(TRAIN_DIR / \"train_credit_bureau_a_2_*.parquet\", 2),\n        read_file(TRAIN_DIR / \"train_applprev_2.parquet\", 2),  # 추가\n        read_file(TRAIN_DIR / \"train_person_2.parquet\", 2),  # 추가\n    ]\n}","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:00:42.765202Z","iopub.execute_input":"2024-03-31T03:00:42.765860Z","iopub.status.idle":"2024-03-31T03:02:58.019429Z","shell.execute_reply.started":"2024-03-31T03:00:42.765826Z","shell.execute_reply":"2024-03-31T03:02:58.018327Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n# keyword argument unpacking (**-operator)\n# https://docs.python.org/3/tutorial/controlflow.html#unpacking-argument-lists\ndf_train = feature_eng(**data_store)\ndf_train","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:03:08.909776Z","iopub.execute_input":"2024-03-31T03:03:08.910217Z","iopub.status.idle":"2024-03-31T03:03:24.353788Z","shell.execute_reply.started":"2024-03-31T03:03:08.910181Z","shell.execute_reply":"2024-03-31T03:03:24.352587Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## . . 2.3. Test Files Read & Feature Engineering","metadata":{}},{"cell_type":"code","source":"%%time\n\ndata_store = {\n    \"df_base\": read_file(TEST_DIR / \"test_base.parquet\"),\n    \"depth_0\": [\n        read_file(TEST_DIR / \"test_static_cb_0.parquet\"),\n        read_files(TEST_DIR / \"test_static_0_*.parquet\"),\n    ],\n    \"depth_1\": [\n        read_files(TEST_DIR / \"test_applprev_1_*.parquet\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_a_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_b_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_c_1.parquet\", 1),\n        read_files(TEST_DIR / \"test_credit_bureau_a_1_*.parquet\", 1),\n        read_file(TEST_DIR / \"test_credit_bureau_b_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_other_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_person_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_deposit_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_debitcard_1.parquet\", 1),\n    ],\n    \"depth_2\": [\n        read_file(TEST_DIR / \"test_credit_bureau_b_2.parquet\", 2),\n        read_files(TEST_DIR / \"test_credit_bureau_a_2_*.parquet\", 2),\n        read_file(TEST_DIR / \"test_applprev_2.parquet\", 2),\n        read_file(TEST_DIR / \"test_person_2.parquet\", 2)\n    ]\n}","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:04:06.028520Z","iopub.execute_input":"2024-03-31T03:04:06.029440Z","iopub.status.idle":"2024-03-31T03:04:06.693373Z","shell.execute_reply.started":"2024-03-31T03:04:06.029406Z","shell.execute_reply":"2024-03-31T03:04:06.692423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test = feature_eng(**data_store)\ndf_test","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:04:16.220287Z","iopub.execute_input":"2024-03-31T03:04:16.220735Z","iopub.status.idle":"2024-03-31T03:04:16.278018Z","shell.execute_reply.started":"2024-03-31T03:04:16.220706Z","shell.execute_reply":"2024-03-31T03:04:16.277050Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## . . 2.4. Feature Elimination","metadata":{}},{"cell_type":"code","source":"df_train = df_train.pipe(Pipeline.filter_cols)\ndf_test = df_test.select([col for col in df_train.columns if col != \"target\"])\n\nprint(\"train data shape:\\t\", df_train.shape)\nprint(\"test data shape:\\t\", df_test.shape)","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:04:27.614729Z","iopub.execute_input":"2024-03-31T03:04:27.615139Z","iopub.status.idle":"2024-03-31T03:04:30.646769Z","shell.execute_reply.started":"2024-03-31T03:04:27.615107Z","shell.execute_reply":"2024-03-31T03:04:30.645735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:04:33.520407Z","iopub.execute_input":"2024-03-31T03:04:33.521151Z","iopub.status.idle":"2024-03-31T03:04:33.536696Z","shell.execute_reply.started":"2024-03-31T03:04:33.521122Z","shell.execute_reply":"2024-03-31T03:04:33.535734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## . . 2.5. Pandas Conversion","metadata":{}},{"cell_type":"code","source":"%%time\n\ndf_train, cat_cols = to_pandas(df_train)\nprint('len(cat_cols):', len(cat_cols))\nprint('df_train.shape:', df_train.shape)\ndf_train","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:04:39.597660Z","iopub.execute_input":"2024-03-31T03:04:39.598508Z","iopub.status.idle":"2024-03-31T03:04:58.943140Z","shell.execute_reply.started":"2024-03-31T03:04:39.598471Z","shell.execute_reply":"2024-03-31T03:04:58.942150Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test, cat_cols = to_pandas(df_test, cat_cols)\nprint('len(cat_cols):', len(cat_cols))\nprint('df_test.shape:', df_test.shape)\ndf_test","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:05:03.081095Z","iopub.execute_input":"2024-03-31T03:05:03.082159Z","iopub.status.idle":"2024-03-31T03:05:03.169019Z","shell.execute_reply.started":"2024-03-31T03:05:03.082113Z","shell.execute_reply":"2024-03-31T03:05:03.168068Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## . . 2.6. Garbage Collection","metadata":{}},{"cell_type":"code","source":"del data_store\n\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:05:06.783106Z","iopub.execute_input":"2024-03-31T03:05:06.784064Z","iopub.status.idle":"2024-03-31T03:05:06.941796Z","shell.execute_reply.started":"2024-03-31T03:05:06.784030Z","shell.execute_reply":"2024-03-31T03:05:06.940602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3. EDA\n- https://seaborn.pydata.org/generated/seaborn.lineplot.html\n\n> By default, the plot aggregates over multiple `y` values at each value of `x` and shows an estimate of the central tendency and a confidence interval for that estimate.","metadata":{}},{"cell_type":"code","source":"print(\"Train is duplicated:\\t\", df_train[\"case_id\"].duplicated().any())\nprint(\"Train Week Range:\\t\", (df_train[\"WEEK_NUM\"].min(), df_train[\"WEEK_NUM\"].max()))\n\nprint()\n\nprint(\"Test is duplicated:\\t\", df_test[\"case_id\"].duplicated().any())\nprint(\"Test Week Range:\\t\", (df_test[\"WEEK_NUM\"].min(), df_test[\"WEEK_NUM\"].max()))","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:05:09.814289Z","iopub.execute_input":"2024-03-31T03:05:09.815050Z","iopub.status.idle":"2024-03-31T03:05:09.843533Z","shell.execute_reply.started":"2024-03-31T03:05:09.815019Z","shell.execute_reply":"2024-03-31T03:05:09.842494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.lineplot(\n    data=df_train,\n    x=\"WEEK_NUM\",\n    y=\"target\",\n)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-31T03:05:10.215337Z","iopub.execute_input":"2024-03-31T03:05:10.215752Z","iopub.status.idle":"2024-03-31T03:05:27.600974Z","shell.execute_reply.started":"2024-03-31T03:05:10.215721Z","shell.execute_reply":"2024-03-31T03:05:27.599929Z"},"trusted":true},"execution_count":null,"outputs":[]}]}