{"metadata":{"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":50160,"databundleVersionId":7921029,"sourceType":"competition"}],"dockerImageVersionId":30648,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.10.13"},"papermill":{"default_parameters":{},"duration":888.861714,"end_time":"2024-02-08T13:46:18.589602","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2024-02-08T13:31:29.727888","version":"2.5.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction","metadata":{}},{"cell_type":"markdown","source":"**Based on:**\n\nhttps://www.kaggle.com/code/greysky/home-credit-baseline\n\nChanges:\n1. Removed model training code (contains only data collection and preproccessing; utility script)\n1. Added function prepare_df\n1. Added new aggregations: min, mean, mode, first, last, n_unique\n\n\n**Related notebooks**\n\nTraining model-1 notebook:\n\nhttps://www.kaggle.com/andreynesterov/home-credit-baseline-training\n\nInference notebook:\n\nhttps://www.kaggle.com/andreynesterov/home-credit-baseline-inference","metadata":{}},{"cell_type":"markdown","source":"# Dependencies","metadata":{}},{"cell_type":"code","source":"import os\nimport gc\nfrom glob import glob\nfrom pathlib import Path\nfrom datetime import datetime\n\nimport numpy as np\nimport pandas as pd\nimport polars as pl\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport warnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)\npd.set_option('display.max_columns', None)\npd.set_option('display.max_rows', 500)","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","papermill":{"duration":6.169561,"end_time":"2024-02-08T13:31:38.616695","exception":false,"start_time":"2024-02-08T13:31:32.447134","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-02-12T05:54:45.849931Z","iopub.execute_input":"2024-02-12T05:54:45.850536Z","iopub.status.idle":"2024-02-12T05:54:48.524594Z","shell.execute_reply.started":"2024-02-12T05:54:45.850503Z","shell.execute_reply":"2024-02-12T05:54:48.523419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Configuration","metadata":{"papermill":{"duration":0.008354,"end_time":"2024-02-08T13:31:38.872593","exception":false,"start_time":"2024-02-08T13:31:38.864239","status":"completed"},"tags":[]}},{"cell_type":"code","source":"class CFG:\n    root_dir = Path(\"/kaggle/input/home-credit-credit-risk-model-stability/\")\n    train_dir = Path(\"/kaggle/input/home-credit-credit-risk-model-stability/parquet_files/train\")\n    test_dir = Path(\"/kaggle/input/home-credit-credit-risk-model-stability/parquet_files/test\")","metadata":{"papermill":{"duration":0.01558,"end_time":"2024-02-08T13:31:38.89682","exception":false,"start_time":"2024-02-08T13:31:38.88124","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-02-12T05:54:48.526762Z","iopub.execute_input":"2024-02-12T05:54:48.527213Z","iopub.status.idle":"2024-02-12T05:54:48.532388Z","shell.execute_reply.started":"2024-02-12T05:54:48.527184Z","shell.execute_reply":"2024-02-12T05:54:48.531166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Feature definitions","metadata":{}},{"cell_type":"code","source":"if __name__ == '__main__':\n    feature_definitions_df = pd.read_csv(CFG.root_dir / \"feature_definitions.csv\")\n    display(feature_definitions_df)\n    pd.reset_option(\"display.max_rows\", 0)","metadata":{"_kg_hide-output":true,"scrolled":true,"execution":{"iopub.status.busy":"2024-02-12T05:54:48.533986Z","iopub.execute_input":"2024-02-12T05:54:48.534419Z","iopub.status.idle":"2024-02-12T05:54:48.597126Z","shell.execute_reply.started":"2024-02-12T05:54:48.534392Z","shell.execute_reply":"2024-02-12T05:54:48.596077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Collection and Preprocessing","metadata":{}},{"cell_type":"markdown","source":"### Pipeline","metadata":{"papermill":{"duration":0.007632,"end_time":"2024-02-08T13:31:38.673586","exception":false,"start_time":"2024-02-08T13:31:38.665954","status":"completed"},"tags":[]}},{"cell_type":"code","source":"class Pipeline:\n    @staticmethod\n    def set_table_dtypes(df):\n        for col in df.columns:\n            if col in [\"case_id\", \"WEEK_NUM\", \"num_group1\", \"num_group2\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Int64))\n            elif col in [\"date_decision\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Date))\n            elif col[-1] in (\"P\", \"A\"):\n                df = df.with_columns(pl.col(col).cast(pl.Float64))\n            elif col[-1] in (\"M\",):\n                df = df.with_columns(pl.col(col).cast(pl.String))\n            elif col[-1] in (\"D\",):\n                df = df.with_columns(pl.col(col).cast(pl.Date))            \n\n        return df\n    \n    @staticmethod\n    def handle_dates(df):\n        for col in df.columns:\n            if col[-1] in (\"D\",):\n                df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))\n                df = df.with_columns(pl.col(col).dt.total_days())\n                \n        df = df.drop(\"date_decision\", \"MONTH\")\n\n        return df\n    \n    @staticmethod\n    def filter_cols(df):\n        for col in df.columns:\n            if col not in [\"target\", \"case_id\", \"WEEK_NUM\"]:\n                isnull = df[col].is_null().mean()\n\n                if isnull > 0.95:\n                    df = df.drop(col)\n\n        for col in df.columns:\n            if (col not in [\"target\", \"case_id\", \"WEEK_NUM\"]) & (df[col].dtype == pl.String):\n                freq = df[col].n_unique()\n\n                if (freq == 1) | (freq > 200):\n                    df = df.drop(col)\n\n        return df","metadata":{"papermill":{"duration":0.022599,"end_time":"2024-02-08T13:31:38.704022","exception":false,"start_time":"2024-02-08T13:31:38.681423","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-02-12T05:54:48.598236Z","iopub.execute_input":"2024-02-12T05:54:48.598867Z","iopub.status.idle":"2024-02-12T05:54:48.611444Z","shell.execute_reply.started":"2024-02-12T05:54:48.598836Z","shell.execute_reply":"2024-02-12T05:54:48.61041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Automatic Aggregation","metadata":{"papermill":{"duration":0.00774,"end_time":"2024-02-08T13:31:38.719798","exception":false,"start_time":"2024-02-08T13:31:38.712058","status":"completed"},"tags":[]}},{"cell_type":"code","source":"class Aggregator:\n    num_aggregators = [pl.max, pl.min, pl.first, pl.last, pl.mean]\n    str_aggregators = [pl.max, pl.min, pl.first, pl.last] # n_unique\n    group_aggregators = [pl.max, pl.min, pl.first, pl.last]\n    \n    @staticmethod\n    def num_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"P\", \"A\")]\n        expr_all = []\n        for method in Aggregator.num_aggregators:\n            expr = [method(col).alias(f\"{method.__name__}_{col}\") for col in cols]\n            expr_all += expr\n\n        return expr_all\n\n    @staticmethod\n    def date_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"D\",)]\n        expr_all = []\n        for method in Aggregator.num_aggregators:\n            expr = [method(col).alias(f\"{method.__name__}_{col}\") for col in cols]  \n            expr_all += expr\n\n        return expr_all\n\n    @staticmethod\n    def str_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"M\",)]\n        \n        expr_all = []\n        for method in Aggregator.str_aggregators:\n            expr = [method(col).alias(f\"{method.__name__}_{col}\") for col in cols]  \n            expr_all += expr\n            \n        expr_mode = [\n            pl.col(col)\n            .drop_nulls()\n            .mode()\n            .first()\n            .alias(f\"mode_{col}\")\n            for col in cols\n        ]\n\n        return expr_all + expr_mode\n\n    @staticmethod\n    def other_expr(df):\n        cols = [col for col in df.columns if col[-1] in (\"T\", \"L\")]\n        \n        expr_all = []\n        for method in Aggregator.str_aggregators:\n            expr = [method(col).alias(f\"{method.__name__}_{col}\") for col in cols]  \n            expr_all += expr\n\n        return expr_all\n    \n    @staticmethod\n    def count_expr(df):\n        cols = [col for col in df.columns if \"num_group\" in col]\n\n        expr_all = []\n        for method in Aggregator.group_aggregators:\n            expr = [method(col).alias(f\"{method.__name__}_{col}\") for col in cols]  \n            expr_all += expr\n            \n#         if len(cols) > 0:\n#             method = pl.count\n#             expr = [method(col).alias(f\"{method.__name__}_{col}\") for col in [cols[0]]]\n#             expr_all += expr\n\n        return expr_all\n\n    @staticmethod\n    def get_exprs(df):\n        exprs = Aggregator.num_expr(df) + \\\n                Aggregator.date_expr(df) + \\\n                Aggregator.str_expr(df) + \\\n                Aggregator.other_expr(df) + \\\n                Aggregator.count_expr(df)\n\n        return exprs","metadata":{"papermill":{"duration":0.021321,"end_time":"2024-02-08T13:31:38.750003","exception":false,"start_time":"2024-02-08T13:31:38.728682","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-02-12T05:54:48.615075Z","iopub.execute_input":"2024-02-12T05:54:48.615486Z","iopub.status.idle":"2024-02-12T05:54:48.631187Z","shell.execute_reply.started":"2024-02-12T05:54:48.615458Z","shell.execute_reply":"2024-02-12T05:54:48.629796Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### File I/O","metadata":{"papermill":{"duration":0.007841,"end_time":"2024-02-08T13:31:38.76613","exception":false,"start_time":"2024-02-08T13:31:38.758289","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def read_file(path, depth=None):\n    df = pl.read_parquet(path)\n    df = df.pipe(Pipeline.set_table_dtypes)\n    \n    if depth in [1, 2]:\n        df = df.group_by(\"case_id\").agg(Aggregator.get_exprs(df))\n    \n    return df\n\ndef read_files(regex_path, depth=None):\n    chunks = []\n    for path in glob(str(regex_path)):\n        chunks.append(pl.read_parquet(path).pipe(Pipeline.set_table_dtypes))\n        \n    df = pl.concat(chunks, how=\"vertical_relaxed\")\n    if depth in [1, 2]:\n        df = df.group_by(\"case_id\").agg(Aggregator.get_exprs(df))\n    \n    return df","metadata":{"papermill":{"duration":0.017149,"end_time":"2024-02-08T13:31:38.791428","exception":false,"start_time":"2024-02-08T13:31:38.774279","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-02-12T05:54:48.632673Z","iopub.execute_input":"2024-02-12T05:54:48.633092Z","iopub.status.idle":"2024-02-12T05:54:48.64202Z","shell.execute_reply.started":"2024-02-12T05:54:48.633063Z","shell.execute_reply":"2024-02-12T05:54:48.640814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Feature Engineering","metadata":{"papermill":{"duration":0.007901,"end_time":"2024-02-08T13:31:38.807502","exception":false,"start_time":"2024-02-08T13:31:38.799601","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def feature_eng(df_base, depth_0, depth_1, depth_2):\n    df_base = (\n        df_base\n        .with_columns(\n            month_decision = pl.col(\"date_decision\").dt.month(),\n            weekday_decision = pl.col(\"date_decision\").dt.weekday(),\n        )\n    )\n        \n    for i, df in enumerate(depth_0 + depth_1 + depth_2):\n        df_base = df_base.join(df, how=\"left\", on=\"case_id\", suffix=f\"_{i}\")\n        \n    df_base = df_base.pipe(Pipeline.handle_dates)\n    \n    return df_base","metadata":{"papermill":{"duration":0.016482,"end_time":"2024-02-08T13:31:38.832632","exception":false,"start_time":"2024-02-08T13:31:38.81615","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-02-12T05:54:48.643488Z","iopub.execute_input":"2024-02-12T05:54:48.643854Z","iopub.status.idle":"2024-02-12T05:54:48.65153Z","shell.execute_reply.started":"2024-02-12T05:54:48.643826Z","shell.execute_reply":"2024-02-12T05:54:48.650288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def to_pandas(df_data, cat_cols=None):\n    df_data = df_data.to_pandas()\n    \n    if cat_cols is None:\n        cat_cols = list(df_data.select_dtypes(\"object\").columns)\n    \n    df_data[cat_cols] = df_data[cat_cols].astype(\"category\")\n    \n    return df_data","metadata":{"papermill":{"duration":0.015339,"end_time":"2024-02-08T13:31:38.85625","exception":false,"start_time":"2024-02-08T13:31:38.840911","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-02-12T05:54:48.652935Z","iopub.execute_input":"2024-02-12T05:54:48.653227Z","iopub.status.idle":"2024-02-12T05:54:48.661862Z","shell.execute_reply.started":"2024-02-12T05:54:48.653203Z","shell.execute_reply":"2024-02-12T05:54:48.660886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### from https://www.kaggle.com/code/batprem/home-credit-risk-mode-utility-scripts\n\ndef reduce_mem_usage(df, float16_as32=True):\n    \"\"\" iterate through all the columns of a dataframe and modify the data type\n        to reduce memory usage.        \n    \"\"\"\n    start_mem = df.memory_usage().sum() / 1024**2\n    print('Memory usage of dataframe is {:.2f} MB'.format(start_mem))\n    \n    for col in df.columns:\n        col_type = df[col].dtype\n        if str(col_type)==\"category\":\n            continue\n        \n        if col_type != object:\n            c_min = df[col].min()\n            c_max = df[col].max()\n            if str(col_type)[:3] == 'int':\n                if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                    df[col] = df[col].astype(np.int8)\n                elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                    df[col] = df[col].astype(np.int16)\n                elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                    df[col] = df[col].astype(np.int32)\n                elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                    df[col] = df[col].astype(np.int64)  \n            else:\n                if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                    if float16_as32:\n                        df[col] = df[col].astype(np.float32)\n                    else:\n                        df[col] = df[col].astype(np.float16)                    \n                elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                    df[col] = df[col].astype(np.float32)\n                else:\n                    df[col] = df[col].astype(np.float64)\n        else:\n            df[col] = df[col].astype('category')\n    end_mem = df.memory_usage().sum() / 1024**2\n    print('Memory usage after optimization is: {:.2f} MB'.format(end_mem))\n    print('Decreased by {:.1f}%'.format(100 * (start_mem - end_mem) / start_mem))\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2024-02-12T05:54:48.663563Z","iopub.execute_input":"2024-02-12T05:54:48.663954Z","iopub.status.idle":"2024-02-12T05:54:48.68008Z","shell.execute_reply.started":"2024-02-12T05:54:48.663918Z","shell.execute_reply":"2024-02-12T05:54:48.678958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prepare df","metadata":{"papermill":{"duration":0.007961,"end_time":"2024-02-08T13:31:38.913224","exception":false,"start_time":"2024-02-08T13:31:38.905263","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def prepare_df(data_dir, cat_cols=None, mode=\"train\", display_store=False, train_cols=[]):\n    print(\"Collecting data...\")\n    data_store = {\n        \"df_base\": read_file(data_dir / f\"{mode}_base.parquet\"),\n        \"depth_0\": [\n            read_file(data_dir / f\"{mode}_static_cb_0.parquet\"),\n            read_files(data_dir / f\"{mode}_static_0_*.parquet\"),\n        ],\n        \"depth_1\": [\n            read_files(data_dir / f\"{mode}_applprev_1_*.parquet\", 1),\n            read_file(data_dir / f\"{mode}_tax_registry_a_1.parquet\", 1),\n            read_file(data_dir / f\"{mode}_tax_registry_b_1.parquet\", 1),\n            read_file(data_dir / f\"{mode}_tax_registry_c_1.parquet\", 1),\n            read_file(data_dir / f\"{mode}_credit_bureau_b_1.parquet\", 1),\n            read_file(data_dir / f\"{mode}_other_1.parquet\", 1),\n            read_file(data_dir / f\"{mode}_person_1.parquet\", 1),\n            read_file(data_dir / f\"{mode}_deposit_1.parquet\", 1),\n            read_file(data_dir / f\"{mode}_debitcard_1.parquet\", 1),\n        ],\n        \"depth_2\": [\n            read_file(data_dir / f\"{mode}_credit_bureau_b_2.parquet\", 2),\n        ]\n    }\n    if display_store:\n        display(data_store)\n    \n    print(\"Feature engeneering...\")\n    feats_df = feature_eng(**data_store)\n    print(\"  feats_df shape:\\t\", feats_df.shape)\n    \n    del data_store\n    gc.collect()\n    \n    print(\"Filter cols...\")\n    if mode == \"train\":\n        feats_df = feats_df.pipe(Pipeline.filter_cols)\n    else:\n        train_cols = feats_df.columns if len(train_cols) == 0 else train_cols\n        feats_df = feats_df.select([col for col in train_cols if col != \"target\"])\n    print(\"  feats_df shape:\\t\", feats_df.shape)\n    \n    print(\"Convert to pandas...\")\n    feats_df = to_pandas(feats_df, cat_cols)\n    return feats_df","metadata":{"execution":{"iopub.status.busy":"2024-02-12T05:54:48.681676Z","iopub.execute_input":"2024-02-12T05:54:48.682089Z","iopub.status.idle":"2024-02-12T05:54:48.698301Z","shell.execute_reply.started":"2024-02-12T05:54:48.682053Z","shell.execute_reply":"2024-02-12T05:54:48.697046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if __name__ == '__main__':\n    train_df = prepare_df(CFG.train_dir)\n    cat_cols = list(train_df.select_dtypes(\"category\").columns)","metadata":{"_kg_hide-output":true,"scrolled":true,"execution":{"iopub.status.busy":"2024-02-12T05:54:48.699834Z","iopub.execute_input":"2024-02-12T05:54:48.700213Z","iopub.status.idle":"2024-02-12T05:57:46.735341Z","shell.execute_reply.started":"2024-02-12T05:54:48.700181Z","shell.execute_reply":"2024-02-12T05:57:46.734187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if __name__ == '__main__':\n    display(train_df)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-02-12T05:57:46.736662Z","iopub.execute_input":"2024-02-12T05:57:46.736981Z","iopub.status.idle":"2024-02-12T05:57:47.316972Z","shell.execute_reply.started":"2024-02-12T05:57:46.736955Z","shell.execute_reply":"2024-02-12T05:57:47.315612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if __name__ == '__main__':\n    display(cat_cols)","metadata":{"_kg_hide-output":true,"scrolled":true,"execution":{"iopub.status.busy":"2024-02-12T05:57:47.31851Z","iopub.execute_input":"2024-02-12T05:57:47.318888Z","iopub.status.idle":"2024-02-12T05:57:47.329952Z","shell.execute_reply.started":"2024-02-12T05:57:47.318854Z","shell.execute_reply":"2024-02-12T05:57:47.328491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if __name__ == '__main__':\n    test_df = prepare_df(CFG.test_dir, cat_cols=cat_cols, mode=\"test\", train_cols=train_df.columns)","metadata":{"_kg_hide-output":true,"scrolled":true,"execution":{"iopub.status.busy":"2024-02-12T05:57:47.333713Z","iopub.execute_input":"2024-02-12T05:57:47.334013Z","iopub.status.idle":"2024-02-12T05:57:47.808895Z","shell.execute_reply.started":"2024-02-12T05:57:47.333987Z","shell.execute_reply":"2024-02-12T05:57:47.807609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if __name__ == '__main__':\n    display(test_df)","metadata":{"execution":{"iopub.status.busy":"2024-02-12T05:57:47.810369Z","iopub.execute_input":"2024-02-12T05:57:47.810768Z","iopub.status.idle":"2024-02-12T05:57:48.325459Z","shell.execute_reply.started":"2024-02-12T05:57:47.810731Z","shell.execute_reply":"2024-02-12T05:57:48.3244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Reduce memory usage and save","metadata":{}},{"cell_type":"code","source":"if __name__ == '__main__':\n    train_df = reduce_mem_usage(train_df)\n    test_df = reduce_mem_usage(test_df)","metadata":{"execution":{"iopub.status.busy":"2024-02-12T05:57:48.326748Z","iopub.execute_input":"2024-02-12T05:57:48.327491Z","iopub.status.idle":"2024-02-12T05:57:57.638882Z","shell.execute_reply.started":"2024-02-12T05:57:48.327458Z","shell.execute_reply":"2024-02-12T05:57:57.637838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if __name__ == '__main__':\n    train_df.to_parquet(\"train_full.parquet\")","metadata":{"execution":{"iopub.status.busy":"2024-02-12T05:57:57.64005Z","iopub.execute_input":"2024-02-12T05:57:57.64036Z","iopub.status.idle":"2024-02-12T05:58:30.141901Z","shell.execute_reply.started":"2024-02-12T05:57:57.640312Z","shell.execute_reply":"2024-02-12T05:58:30.141094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### EDA","metadata":{"papermill":{"duration":0.008701,"end_time":"2024-02-08T13:32:31.919166","exception":false,"start_time":"2024-02-08T13:32:31.910465","status":"completed"},"tags":[]}},{"cell_type":"code","source":"if __name__ == '__main__':\n    print(\"Train is duplicated:\\t\", train_df[\"case_id\"].duplicated().any())\n    print(\"Train Week Range:\\t\", (train_df[\"WEEK_NUM\"].min(), train_df[\"WEEK_NUM\"].max()))\n\n    print()\n\n    print(\"Test is duplicated:\\t\", test_df[\"case_id\"].duplicated().any())\n    print(\"Test Week Range:\\t\", (test_df[\"WEEK_NUM\"].min(), test_df[\"WEEK_NUM\"].max()))","metadata":{"papermill":{"duration":0.050383,"end_time":"2024-02-08T13:32:31.9784","exception":false,"start_time":"2024-02-08T13:32:31.928017","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-02-12T05:58:30.143228Z","iopub.execute_input":"2024-02-12T05:58:30.144418Z","iopub.status.idle":"2024-02-12T05:58:30.181173Z","shell.execute_reply.started":"2024-02-12T05:58:30.144377Z","shell.execute_reply":"2024-02-12T05:58:30.180254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if __name__ == '__main__':\n    sns.lineplot(\n        data=train_df,\n        x=\"WEEK_NUM\",\n        y=\"target\",\n    )\n    plt.show()","metadata":{"papermill":{"duration":16.920109,"end_time":"2024-02-08T13:32:48.907503","exception":false,"start_time":"2024-02-08T13:32:31.987394","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-02-12T05:58:30.184876Z","iopub.execute_input":"2024-02-12T05:58:30.185258Z","iopub.status.idle":"2024-02-12T05:58:47.946045Z","shell.execute_reply.started":"2024-02-12T05:58:30.185231Z","shell.execute_reply":"2024-02-12T05:58:47.945008Z"},"trusted":true},"execution_count":null,"outputs":[]}]}