{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":50160,"databundleVersionId":7921029,"sourceType":"competition"},{"sourceId":7487967,"sourceType":"datasetVersion","datasetId":3933894}],"dockerImageVersionId":30648,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Home Credit - Credit Risk Model Stability\n\nby Косса Николай Евгеньевич\n\nОригинальный код: https://www.kaggle.com/code/jirkaborovec/credit-risk-eda-xgboost-depth-0-gpu\n\n- Было: 0.370\n- Стало: 0.527 (версия 20)\n- Вторая хорошая попытка: 0.514 (версия 46)\n\n## Улучшения\n\n- Добавлена аггрегация для числовых признаков с глубиной один (средняя, максимум, минимум)\n- Использованы новые источники данных (большая часть таблиц с глубиной 1, которые представлены в тестовых данных)\n- Добавлен подбор гиперпараметров (перебирал я больше, чем есть в коде, но поскольку это долго, оставил только наиболее влияющие на результат)\n- Некоторые стандартные параметры модели изменены для лучшего качества\n- Использована линейная комбинация двух моделей для оптимизации целевого скора\n\n## Отличие от прошлой версии\n\n - Починил линейную комбинацию моделей (раньше objective был одинаковый, теперь свой для каждой из валидационных метрик)\n - Добавил нормализацию для численных признаков, в том числе сгенерированных\n - Подобрал веса для позитивных и негативных примеров (раньше брал просто отношение 0 и 1, которое было слишком большим для нормальной работы, около 30)\n - Бросил затею с генерацией новых кат. признаков, так как слишком сложный формат данных и не получилось раздебагать мои функции\n - Добавил удаление кат. признаков с 1 значением, чтобы не забивать модель бесполезными факторами\n\n## Итоги\nПеребить свой прошлый скор не получилось до того, как кончится квота на GPU, но сделал задел на будущее улучшение. Разбираться с глубиной 2 было довольно сложно, времени на это не хватало, и это и является основной точкой для будущих улучшений в работе.\n\n","metadata":{}},{"cell_type":"code","source":"import warnings\n# warnings.simplefilter(\"ignore\")\nwarnings.simplefilter(\"ignore\", UserWarning)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-03-15T12:29:50.692957Z","iopub.execute_input":"2024-03-15T12:29:50.693741Z","iopub.status.idle":"2024-03-15T12:29:50.698146Z","shell.execute_reply.started":"2024-03-15T12:29:50.693705Z","shell.execute_reply":"2024-03-15T12:29:50.697223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"USE_HYPERTUNING_AUC = False\nUSE_HYPERTUNING_RMSE = False\nbest_hyperparams_auc = {'max_depth': 6, 'eta': 0.025, 'scale_pos_weight': 0.175}\nbest_hyperparams_rmse = {'max_depth': 6, 'eta': 0.025, 'scale_pos_weight': 0.75}","metadata":{"execution":{"iopub.status.busy":"2024-03-15T12:29:50.938060Z","iopub.execute_input":"2024-03-15T12:29:50.938680Z","iopub.status.idle":"2024-03-15T12:29:50.942941Z","shell.execute_reply.started":"2024-03-15T12:29:50.938644Z","shell.execute_reply":"2024-03-15T12:29:50.941968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os, glob\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom pathlib import Path\n\nPATH_DATASET = Path(\"/kaggle/input/home-credit-credit-risk-model-stability\")\nPATH_PARQUETS = PATH_DATASET / \"parquet_files\"\nPARQUETS_TRAIN = PATH_PARQUETS / \"train\"\nPARQUETS_TEST = PATH_PARQUETS / \"test\"\npd.set_option('display.max_columns', None)\npd.set_option('display.max_rows', None)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-03-15T12:29:51.347114Z","iopub.execute_input":"2024-03-15T12:29:51.347391Z","iopub.status.idle":"2024-03-15T12:29:51.353038Z","shell.execute_reply.started":"2024-03-15T12:29:51.347367Z","shell.execute_reply":"2024-03-15T12:29:51.352091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Explore the training data\n\n**Borrowed from the competition describtion:**\n\nTable Description\nThis dataset contains a large number of tables as a result of utilizing diverse data sources and the varying levels of data aggregation used while preparing the dataset. Note: All files listed below are found in both .csv and .parquet formats.\n\n### Depth values\n\n- **depth=0** - These are static features directly tied to a specific case_id.\n- **depth=1** - Each case_id has an associated historical record, indexed by num_group1.\n- **depth=2** - Each case_id has an associated historical record, indexed by both num_group1 and num_group2.\n\nYou can read more about Credit bureau (CB) here https://en.wikipedia.org/wiki/Credit_bureau.","metadata":{}},{"cell_type":"markdown","source":"Various predictors were transformed, therefore we have the following notation for similar groups of transformations\n\n- **P** - Transform DPD (Days past due)\n- **M** - Masking categories\n- **A** - Transform amount\n- **D** - Transform date\n- **T** - Unspecified Transform\n- **L** - Unspecified Transform\n\nPlease note that transformations within a group are denoted by a capital letter at the end of the predictor name (e.g., maxdbddpdtollast6m_4187119P). We hope that this will simplify the manipulation with predictors.","metadata":{}},{"cell_type":"code","source":"def _short_array(arr):\n    if len(arr) <= 5:\n        return repr(arr)\n    return f\"[{', '.join(map(str, arr[:2]))}, ..., {', '.join(map(str, arr[-2:]))}]\"\n\n# taking inpiration with column names from https://www.kaggle.com/code/greysky/home-credit-baseline\n# изменил так, чтобы возвращались колонки, сгруппированные по типам\ndef convert_dtypes(df, columns=None):\n    cols_cat = []\n    cols_num = []\n    cols_oth = []\n    for col, dt in dict(df.dtypes).items():\n        if col.startswith(\"for\"):\n            df[col] = df[col].fillna(0).astype(\"int16\")\n            cols_num.append(col)\n        elif \"num\" in col or \"cnt\" in col:\n            df[col] = df[col].fillna(0).astype(\"int32\")\n            cols_num.append(col)\n        elif col.startswith(\"pct\"):\n            df[col] = df[col].astype(\"float16\")\n            cols_num.append(col)\n        elif col[-1] in (\"A\", \"P\"):\n            df[col] = df[col].astype(\"float32\")\n            cols_num.append(col)\n        elif col[-1] in (\"D\", ):\n            df[col] = pd.to_datetime(df[col])\n            cols_oth.append(col)\n        elif col[-1] in (\"M\", \"L\"):\n            if col[-1] == \"L\" and dt.name.startswith(\"int\"):\n                df[col] = df[col].astype(\"int32\")\n                cols_num.append(col)\n            elif col[-1] == \"L\" and dt.name.startswith(\"float\"):\n                df[col] = df[col].astype(\"float32\")\n                cols_num.append(col)\n            else:\n                if columns is None:\n                    uq = list(df[col].unique())\n                    if len(uq) < 2:\n                        print(f'deleted {col}')\n                        del df[col]\n                    else:\n                        print(f'{col} -> #{len(uq)} -> {_short_array(uq)}')\n                        df[col] = df[col].astype(\"category\")\n                        cols_cat.append(col)\n                elif col in columns:\n                    uq = list(df[col].unique())\n                    print(f'{col} -> #{len(uq)} -> {_short_array(uq)}')\n                    df[col] = df[col].astype(\"category\")\n                    cols_cat.append(col)\n                else:\n                    print(f'deleted {col}')\n                    del df[col]\n#         if col[-1] in (\"A\", \"P\", \"M\", \"L\"):\n#             cols.append(col)\n    return cols_cat, cols_num, cols_oth","metadata":{"execution":{"iopub.status.busy":"2024-03-15T11:48:57.366774Z","iopub.execute_input":"2024-03-15T11:48:57.367249Z","iopub.status.idle":"2024-03-15T11:48:57.383353Z","shell.execute_reply.started":"2024-03-15T11:48:57.367222Z","shell.execute_reply":"2024-03-15T11:48:57.382454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Base tables\n\nBase tables store the basic information about the observation and case_id. This is a unique identification of every observation and you need to use it to join the other tables to base tables.\n\ntrain_base.csv","metadata":{}},{"cell_type":"code","source":"df_train = pd.read_parquet(PARQUETS_TRAIN / \"train_base.parquet\")\nprint(f\"size: {len(df_train)}\")\ndisplay(df_train.head())","metadata":{"execution":{"iopub.status.busy":"2024-03-15T11:48:58.350372Z","iopub.execute_input":"2024-03-15T11:48:58.350791Z","iopub.status.idle":"2024-03-15T11:48:58.577583Z","shell.execute_reply.started":"2024-03-15T11:48:58.350762Z","shell.execute_reply":"2024-03-15T11:48:58.576409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train[\"date_decision\"] = pd.to_datetime(df_train[\"date_decision\"]).dt.date\n# delete redundat cols\ndel df_train[\"MONTH\"], df_train[\"WEEK_NUM\"]","metadata":{"execution":{"iopub.status.busy":"2024-03-15T11:48:58.767538Z","iopub.execute_input":"2024-03-15T11:48:58.767912Z","iopub.status.idle":"2024-03-15T11:48:59.525320Z","shell.execute_reply.started":"2024-03-15T11:48:58.767883Z","shell.execute_reply":"2024-03-15T11:48:59.524115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def cnt_lag_depth1_num(df, cols):\n    stats = ['mean', 'max', 'min']\n    if 'num_group1' in cols:\n        cols.remove('num_group1')\n    if 'num_group2' in cols:\n        cols.remove('num_group2')\n    res = df[['case_id'] + cols].groupby('case_id', as_index=False).agg({col: stats for col in cols})\n    for col in cols:\n        for stat in stats:\n                res[f'{col}_{stat}'] = res[col][stat]\n        del res[col]\n    return res.droplevel(1, axis=1)\n    \ndef merge_parquets(df, name, prefix=\"train\", suffixes=(\"_x\", \"_y\"),  folder=PATH_PARQUETS, columns=None):\n    df_ = pd.concat(\n        [pd.read_parquet(p) for p in glob.glob(str(folder / prefix / f\"{prefix}_{name}*.parquet\"))],\n    )\n    df_.drop_duplicates(inplace=True)\n    cols_cat, cols_num, cols_oth = convert_dtypes(df_, columns)\n    cols_all = cols_cat + cols_num + cols_oth + [\"case_id\"]\n    df_ = df_[cols_all]\n\n    if \"num_group1\" in df_.columns:\n        lag_depth1_num = cnt_lag_depth1_num(df_, cols_num)\n        del df_[\"num_group1\"]\n\n        df_ = df_.groupby(['case_id'], as_index=False).first().reset_index()\n        df_ = df_.merge(lag_depth1_num, how=\"left\", on=\"case_id\")\n\n    df = df.merge(df_, how=\"left\", on=\"case_id\", suffixes=suffixes)\n    print(f\"fused size: {len(df)}\")\n    return df","metadata":{"execution":{"iopub.status.busy":"2024-03-15T11:48:59.527481Z","iopub.execute_input":"2024-03-15T11:48:59.528154Z","iopub.status.idle":"2024-03-15T11:48:59.539058Z","shell.execute_reply.started":"2024-03-15T11:48:59.528118Z","shell.execute_reply":"2024-03-15T11:48:59.537953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class Scaler:\n    def __init__(self):\n        self.scales = {}\n        self.fitted = False\n\n    def fit(self, df, ignore=[]):\n        self.scales = {}\n        for col in df.columns:\n            if col in ignore:\n                continue\n            try:\n                self.scales[col] = {'mean': df[col].mean(), 'std': df[col].std()}\n            except:\n                continue\n        self.fitted = True\n        return self\n    \n    def scale(self, df):\n        assert self.fitted, 'Fit the Scaler first'\n        for col in self.scales:\n            df[col] = (df[col] - self.scales[col]['mean']) / self.scales[col]['std']\n        return df","metadata":{"execution":{"iopub.status.busy":"2024-03-15T11:49:00.241973Z","iopub.execute_input":"2024-03-15T11:49:00.242980Z","iopub.status.idle":"2024-03-15T11:49:00.251920Z","shell.execute_reply.started":"2024-03-15T11:49:00.242936Z","shell.execute_reply":"2024-03-15T11:49:00.250921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Depth 0\n\n### static_0 | Properties: depth=0, internal data source\n\n- train_static_0_0.csv\n- train_static_0_1.csv\n\n### static_cb_0 | Properties: depth=0, external data source\n\n- train_static_cb_0.csv","metadata":{}},{"cell_type":"code","source":"for name in [\"static_0\", \"static_cb_0\"]:\n    df_train = merge_parquets(df_train, name, suffixes=('', f'_{name}'))","metadata":{"execution":{"iopub.status.busy":"2024-03-15T11:49:01.635259Z","iopub.execute_input":"2024-03-15T11:49:01.635656Z","iopub.status.idle":"2024-03-15T11:49:41.063132Z","shell.execute_reply.started":"2024-03-15T11:49:01.635603Z","shell.execute_reply":"2024-03-15T11:49:41.062084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Depth 1\n\n### applprev_1 | Properties: depth=1, internal data source\n\n- train_applprev_1_0.csv\n- train_applprev_1_1.csv\n\n### other_1 | Properties: depth=1, internal data source\n\n- train_other_1.csv\n\n### tax_registry_a_1 | Properties: depth=1, external data source, Tax registry provider A\n\n- train_tax_registry_a_1.csv\n\n### tax_registry_b_1 | Properties: depth=1, external data source, Tax registry provider B\n\n- train_tax_registry_b_1.csv\n\n### tax_registry_c_1 | Properties: depth=1, external data source, Tax registry provider C\n\n- train_tax_registry_c_1.csv\n\n### credit_bureau_a_1 | Properties: depth=1, external data source, Credit bureau provider A\n\n- train_credit_bureau_a_1_0.csv\n- train_credit_bureau_a_1_1.csv\n- train_credit_bureau_a_1_2.csv\n- train_credit_bureau_a_1_3.csv\n\n### credit_bureau_b_1 | Properties: depth=1, external data source, Credit bureau provider B\n\n- train_credit_bureau_b_1.csv\n\n### deposit_1 | Properties: depth=1, internal data source\n\n- train_deposit_1.csv\n\n### person_1 | Properties: depth=1, internal data source\n\n- train_person_1.csv\n\n### debitcard_1 | Properties: depth=1, internal data source\n\n- train_debitcard_1.csv","metadata":{}},{"cell_type":"code","source":"for name in [\n    \"other_1\",\n    \"deposit_1\",\n    \"debitcard_1\",\n    \"person_1\",\n    \"tax_registry_a_1\",\n    \"tax_registry_b_1\"\n    # \"tax_registry_c_1\"\n]:\n    df_train = merge_parquets(df_train, name, suffixes=('', f'_{name}'))","metadata":{"execution":{"iopub.status.busy":"2024-03-15T11:49:41.064762Z","iopub.execute_input":"2024-03-15T11:49:41.065067Z","iopub.status.idle":"2024-03-15T11:50:16.842319Z","shell.execute_reply.started":"2024-03-15T11:49:41.065041Z","shell.execute_reply":"2024-03-15T11:50:16.841330Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_columns = []\nthreshold = 0.5\n\nfor col, dt in dict(df_train.dtypes).items():\n    if dt.name == 'category' or dt.name.startswith('datetime64'):\n        if (((df_train[col] == np.NaN) | (df_train[col] == np.datetime64(\"NaT\"))).sum() / df_train.shape[0]) < threshold:\n            train_columns.append(col)\n    else:\n        if (df_train[col].isna().sum() / df_train.shape[0]) < threshold:\n            train_columns.append(col)\n            \nfor col in df_train.columns:\n    if col not in train_columns:\n        del df_train[col]","metadata":{"execution":{"iopub.status.busy":"2024-03-15T11:50:16.844135Z","iopub.execute_input":"2024-03-15T11:50:16.844469Z","iopub.status.idle":"2024-03-15T11:50:17.736175Z","shell.execute_reply.started":"2024-03-15T11:50:16.844440Z","shell.execute_reply":"2024-03-15T11:50:17.735186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Depth 2\n\n### applprev_2 | Properties: depth=2, internal data source\n\n- train_applprev_2.csv\n\n### person_2 | Properties: depth=2, internal data source\n\n- train_person_2.csv\n\n### redit_bureau_a_2 | Properties: depth=2, external data source, Credit bureau provider A\n\n- train_credit_bureau_a_2_0.csv\n- train_credit_bureau_a_2_1.csv\n- train_credit_bureau_a_2_2.csv\n- train_credit_bureau_a_2_3.csv\n- train_credit_bureau_a_2_4.csv\n- train_credit_bureau_a_2_5.csv\n- train_credit_bureau_a_2_6.csv\n- train_credit_bureau_a_2_7.csv\n- train_credit_bureau_a_2_8.csv\n- train_credit_bureau_a_2_9.csv\n- train_credit_bureau_a_2_10.csv\n\n### credit_bureau_b_2 | Properties: depth=2, external data source, Credit bureau provider B\n\n- train_credit_bureau_b_2.csv","metadata":{}},{"cell_type":"code","source":"# TODO","metadata":{"execution":{"iopub.status.busy":"2024-03-15T11:50:17.737511Z","iopub.execute_input":"2024-03-15T11:50:17.738345Z","iopub.status.idle":"2024-03-15T11:50:17.743480Z","shell.execute_reply.started":"2024-03-15T11:50:17.738311Z","shell.execute_reply":"2024-03-15T11:50:17.742637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Brows the data","metadata":{}},{"cell_type":"code","source":"print(f\"data size: {len(df_train)}\")\nprint(f\"unique: {len(df_train['case_id'].unique())}\")","metadata":{"execution":{"iopub.status.busy":"2024-03-15T11:50:17.745953Z","iopub.execute_input":"2024-03-15T11:50:17.746311Z","iopub.status.idle":"2024-03-15T11:50:17.800404Z","shell.execute_reply.started":"2024-03-15T11:50:17.746272Z","shell.execute_reply":"2024-03-15T11:50:17.799227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(df_train.head().T)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-03-15T11:50:17.802070Z","iopub.execute_input":"2024-03-15T11:50:17.802456Z","iopub.status.idle":"2024-03-15T11:50:18.009414Z","shell.execute_reply.started":"2024-03-15T11:50:17.802422Z","shell.execute_reply":"2024-03-15T11:50:18.008513Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"t = df_train.groupby('date_decision')['target'].value_counts(normalize=True).mul(100)\nt.unstack().plot.bar(stacked=True, figsize=(14, 2), legend=True)","metadata":{"execution":{"iopub.status.busy":"2024-03-15T11:50:18.010483Z","iopub.execute_input":"2024-03-15T11:50:18.010777Z","iopub.status.idle":"2024-03-15T11:50:26.136821Z","shell.execute_reply.started":"2024-03-15T11:50:18.010752Z","shell.execute_reply":"2024-03-15T11:50:26.135947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.groupby('date_decision')['target'].mean().plot(\n    figsize=(14, 2), grid=True,\n    xlabel=\"date of decision\",\n    ylabel=\"day mean / proxi ratio\",\n)","metadata":{"execution":{"iopub.status.busy":"2024-03-15T11:50:26.137912Z","iopub.execute_input":"2024-03-15T11:50:26.138185Z","iopub.status.idle":"2024-03-15T11:50:26.727215Z","shell.execute_reply.started":"2024-03-15T11:50:26.138160Z","shell.execute_reply":"2024-03-15T11:50:26.726360Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(df_train.dtypes)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-03-15T11:50:26.728432Z","iopub.execute_input":"2024-03-15T11:50:26.728758Z","iopub.status.idle":"2024-03-15T11:50:26.743273Z","shell.execute_reply.started":"2024-03-15T11:50:26.728732Z","shell.execute_reply":"2024-03-15T11:50:26.742220Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols_cat, cols_num, cols_oth = convert_dtypes(df_train)\ncols_all = cols_cat + cols_num + [\"case_id\", \"target\"]\ndf_train = df_train[cols_all]\ndisplay(df_train.dtypes)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-03-15T11:50:26.744466Z","iopub.execute_input":"2024-03-15T11:50:26.744827Z","iopub.status.idle":"2024-03-15T11:50:29.189271Z","shell.execute_reply.started":"2024-03-15T11:50:26.744793Z","shell.execute_reply":"2024-03-15T11:50:29.188379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train simple XGBoost","metadata":{}},{"cell_type":"code","source":"scaler = Scaler().fit(df_train, ignore=['case_id', 'target'])\ndf_train = scaler.scale(df_train)","metadata":{"execution":{"iopub.status.busy":"2024-03-15T11:50:29.192281Z","iopub.execute_input":"2024-03-15T11:50:29.192564Z","iopub.status.idle":"2024-03-15T11:50:31.646211Z","shell.execute_reply.started":"2024-03-15T11:50:29.192539Z","shell.execute_reply":"2024-03-15T11:50:31.645385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\ndf_train.replace([np.inf, -np.inf], np.nan, inplace=True)\ntrain_cols = [c for c in df_train.columns if c not in (\"case_id\", \"date_decision\", \"target\")]\nX_train, X_valid, y_train, y_valid = train_test_split(\n    df_train[train_cols], df_train['target'], test_size=0.2)","metadata":{"execution":{"iopub.status.busy":"2024-03-15T11:50:31.647294Z","iopub.execute_input":"2024-03-15T11:50:31.647555Z","iopub.status.idle":"2024-03-15T11:50:36.068043Z","shell.execute_reply.started":"2024-03-15T11:50:31.647533Z","shell.execute_reply":"2024-03-15T11:50:36.067002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install -U xgboost -f /kaggle/input/xgboost-python-package/ --no-index","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-03-15T11:50:36.069460Z","iopub.execute_input":"2024-03-15T11:50:36.070309Z","iopub.status.idle":"2024-03-15T11:50:48.269738Z","shell.execute_reply.started":"2024-03-15T11:50:36.070273Z","shell.execute_reply":"2024-03-15T11:50:48.268683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import xgboost as xgb\n\nprint(xgb.__version__)","metadata":{"execution":{"iopub.status.busy":"2024-03-15T11:50:48.271671Z","iopub.execute_input":"2024-03-15T11:50:48.272072Z","iopub.status.idle":"2024-03-15T11:50:48.277592Z","shell.execute_reply.started":"2024-03-15T11:50:48.272032Z","shell.execute_reply":"2024-03-15T11:50:48.276680Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"small_size = int(X_train.shape[0] * 0.2)\nX_train_small, y_train_small = X_train[:small_size], y_train[:small_size]","metadata":{"execution":{"iopub.status.busy":"2024-03-15T11:50:48.278972Z","iopub.execute_input":"2024-03-15T11:50:48.279286Z","iopub.status.idle":"2024-03-15T11:50:48.288013Z","shell.execute_reply.started":"2024-03-15T11:50:48.279260Z","shell.execute_reply":"2024-03-15T11:50:48.287103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"default_params = {\n    'scale_pos_weight':0.2,\n    'device':'cuda',\n    'enable_categorical':True,\n    'n_estimators': 1000,\n    'subsample':0.6,\n    'colsample_bytree':0.6,\n    'min_child_weight':1,\n    'random_state':42,\n}\n\nhyperparams = {\n    'max_depth': [4, 5, 6],\n    'eta': [0.0175, 0.025, 0.05, 0.075],\n    'scale_pos_weight': [0.175, 0.25, 0.5, 0.75]\n}","metadata":{"execution":{"iopub.status.busy":"2024-03-15T11:50:48.289106Z","iopub.execute_input":"2024-03-15T11:50:48.289399Z","iopub.status.idle":"2024-03-15T11:50:48.297214Z","shell.execute_reply.started":"2024-03-15T11:50:48.289375Z","shell.execute_reply":"2024-03-15T11:50:48.296405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from itertools import product\n\ndef hyperparam_tuning(hyperparams, default_params, objective='binary:logistic', eval_metric='auc'):\n    def print_kwargs(**kwargs):\n        print(kwargs)\n        \n    # было лень самому придумать как построить грид:\n    # https://stackoverflow.com/questions/64878301/iterate-over-all-possible-combinations-of-params-inside-a-python-dictionary-pas\n    param_names = list(hyperparams.keys())                                       \n    param_values = (zip(param_names, x) for x in product(*hyperparams.values()))\n\n    best_params = {}\n    best_score = 0\n\n    for params in param_values:\n        params = dict(params)\n        print_kwargs(**params)\n\n        model_params = {**default_params, **params}\n\n        model = xgb.XGBClassifier(\n            objective=objective,\n            eval_metric=eval_metric,\n            **model_params)\n        model.fit(\n            X_train_small, y_train_small,\n            eval_set=[(X_valid, y_valid)],\n            early_stopping_rounds=8,\n            verbose=False,\n        )\n        pred = np.clip(model.predict_proba(X_valid)[:, 1], 0, 1)\n        print(f'score:{model.evals_result()[\"validation_0\"][\"auc\"][-1]}')\n        print('=======================================')\n        if model.evals_result()['validation_0']['auc'][-1] > best_score:\n            best_score = model.best_score\n            best_params = params\n    return best_params","metadata":{"execution":{"iopub.status.busy":"2024-03-15T12:30:03.331264Z","iopub.execute_input":"2024-03-15T12:30:03.331888Z","iopub.status.idle":"2024-03-15T12:30:03.342734Z","shell.execute_reply.started":"2024-03-15T12:30:03.331855Z","shell.execute_reply":"2024-03-15T12:30:03.341610Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if USE_HYPERTUNING_AUC:\n    best_params = hyperparam_tuning(hyperparams, default_params, objective='binary:logistic', eval_metric='auc')\nelse:\n    best_params = best_hyperparams_auc","metadata":{"execution":{"iopub.status.busy":"2024-03-15T12:30:03.793437Z","iopub.execute_input":"2024-03-15T12:30:03.793735Z","iopub.status.idle":"2024-03-15T12:30:03.798373Z","shell.execute_reply.started":"2024-03-15T12:30:03.793710Z","shell.execute_reply":"2024-03-15T12:30:03.797290Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_auc = xgb.XGBClassifier(\n    objective='binary:logistic', eval_metric='auc', **{**best_params,**default_params}\n)\n\n# Training the model on the training data\nmodel_auc.fit(\n    X_train, y_train,\n    eval_set=[(X_valid, y_valid)],\n    early_stopping_rounds=25,\n    verbose=True,\n)","metadata":{"execution":{"iopub.status.busy":"2024-03-15T12:30:05.368130Z","iopub.execute_input":"2024-03-15T12:30:05.368841Z","iopub.status.idle":"2024-03-15T12:31:22.080997Z","shell.execute_reply.started":"2024-03-15T12:30:05.368808Z","shell.execute_reply":"2024-03-15T12:31:22.080054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def hyperparam_tuning(hyperparams, default_params, objective='reg:squaredlogerror', eval_metric='rmse'):\n    def print_kwargs(**kwargs):\n        print(kwargs)\n        \n    # было лень самому придумать как построить грид:\n    # https://stackoverflow.com/questions/64878301/iterate-over-all-possible-combinations-of-params-inside-a-python-dictionary-pas\n    param_names = list(hyperparams.keys())                                       \n    param_values = (zip(param_names, x) for x in product(*hyperparams.values()))\n\n    best_params = {}\n    best_score = 0\n\n    for params in param_values:\n        params = dict(params)\n        print_kwargs(**params)\n\n        model_params = {**default_params, **params}\n\n        model = xgb.XGBClassifier(\n            objective=objective,\n            eval_metric=eval_metric,\n            **model_params)\n        model.fit(\n            X_train_small, y_train_small,\n            eval_set=[(X_valid, y_valid)],\n            early_stopping_rounds=8,\n            verbose=False,\n        )\n        pred = np.clip(model.predict_proba(X_valid)[:, 1], 0, 1)\n        print(f'score:{model.evals_result()[\"validation_0\"][eval_metric][-1]}')\n        print('=======================================')\n        if model.evals_result()['validation_0'][eval_metric][-1] < best_score:\n            best_score = model.best_score\n            best_params = params\n    return best_params","metadata":{"execution":{"iopub.status.busy":"2024-03-15T12:31:22.091095Z","iopub.execute_input":"2024-03-15T12:31:22.091470Z","iopub.status.idle":"2024-03-15T12:31:22.101544Z","shell.execute_reply.started":"2024-03-15T12:31:22.091438Z","shell.execute_reply":"2024-03-15T12:31:22.100607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if USE_HYPERTUNING_RMSE:\n    best_params = hyperparam_tuning(hyperparams, default_params, objective='reg:squarederror', eval_metric='rmse')\nelse:\n    best_params = best_hyperparams_rmse","metadata":{"execution":{"iopub.status.busy":"2024-03-15T12:31:22.103418Z","iopub.execute_input":"2024-03-15T12:31:22.103771Z","iopub.status.idle":"2024-03-15T12:31:22.112288Z","shell.execute_reply.started":"2024-03-15T12:31:22.103746Z","shell.execute_reply":"2024-03-15T12:31:22.111369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_rmse = xgb.XGBClassifier(\n    objective='reg:squarederror', eval_metric='rmse', **{**best_params,**default_params}\n)\n\n# Training the model on the training data\nmodel_rmse.fit(\n    X_train, y_train,\n    eval_set=[(X_valid, y_valid)],\n    early_stopping_rounds=25,\n    verbose=True,\n)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-03-15T12:31:22.123975Z","iopub.execute_input":"2024-03-15T12:31:22.124305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_columns_dtypes = dict(df_train.dtypes)\ndel df_train","metadata":{"execution":{"iopub.status.busy":"2024-03-14T20:47:48.538892Z","iopub.status.idle":"2024-03-14T20:47:48.539335Z","shell.execute_reply.started":"2024-03-14T20:47:48.539102Z","shell.execute_reply":"2024-03-14T20:47:48.539121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loading test data","metadata":{}},{"cell_type":"markdown","source":"## base data","metadata":{"execution":{"iopub.status.busy":"2024-02-07T11:36:19.157007Z","iopub.execute_input":"2024-02-07T11:36:19.157411Z","iopub.status.idle":"2024-02-07T11:36:19.166433Z","shell.execute_reply.started":"2024-02-07T11:36:19.157376Z","shell.execute_reply":"2024-02-07T11:36:19.165721Z"}}},{"cell_type":"code","source":"df_test = pd.read_parquet(PARQUETS_TEST / \"test_base.parquet\")\nprint(f\"size: {len(df_test)}\")\ndisplay(df_test.head())","metadata":{"execution":{"iopub.status.busy":"2024-03-14T20:47:48.540751Z","iopub.status.idle":"2024-03-14T20:47:48.541208Z","shell.execute_reply.started":"2024-03-14T20:47:48.540957Z","shell.execute_reply":"2024-03-14T20:47:48.540977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test[\"date_decision\"] = pd.to_datetime(df_test[\"date_decision\"]).dt.date\n# delete redundat cols\ndel df_test[\"MONTH\"], df_test[\"WEEK_NUM\"]","metadata":{"execution":{"iopub.status.busy":"2024-03-14T20:47:48.542211Z","iopub.status.idle":"2024-03-14T20:47:48.542654Z","shell.execute_reply.started":"2024-03-14T20:47:48.542430Z","shell.execute_reply":"2024-03-14T20:47:48.542449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Depth 0","metadata":{}},{"cell_type":"code","source":"for name in [\"static_0\", \"static_cb_0\"]:\n    df_test = merge_parquets(df_test, name, \"test\", suffixes=('', f'_{name}'), columns=train_cols)","metadata":{"execution":{"iopub.status.busy":"2024-03-14T20:47:48.544021Z","iopub.status.idle":"2024-03-14T20:47:48.544489Z","shell.execute_reply.started":"2024-03-14T20:47:48.544239Z","shell.execute_reply":"2024-03-14T20:47:48.544258Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Depth 1","metadata":{}},{"cell_type":"code","source":"# df_test = merge_parquets(df_test, \"applprev_1\", \"test\")","metadata":{"execution":{"iopub.status.busy":"2024-03-14T20:47:48.545743Z","iopub.status.idle":"2024-03-14T20:47:48.546220Z","shell.execute_reply.started":"2024-03-14T20:47:48.545966Z","shell.execute_reply":"2024-03-14T20:47:48.545986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for name in [\n    \"other_1\",\n    \"deposit_1\",\n    \"debitcard_1\",\n    \"person_1\",\n    \"tax_registry_a_1\",\n    \"tax_registry_b_1\"\n    # \"tax_registry_c_1\"\n]:\n    df_test = merge_parquets(df_test, name, \"test\", suffixes=('', f'_{name}'), columns=train_cols)","metadata":{"execution":{"iopub.status.busy":"2024-03-14T20:47:48.547738Z","iopub.status.idle":"2024-03-14T20:47:48.548182Z","shell.execute_reply.started":"2024-03-14T20:47:48.547949Z","shell.execute_reply":"2024-03-14T20:47:48.547969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Overview","metadata":{}},{"cell_type":"code","source":"for col, dt in train_columns_dtypes.items():\n    if col not in train_cols:\n        continue\n    try:\n        df_test[col] = df_test[col].astype(dt)\n    except:\n        print(f\"failed converting {col} to {dt}\")","metadata":{"execution":{"iopub.status.busy":"2024-03-14T20:47:48.549640Z","iopub.status.idle":"2024-03-14T20:47:48.550073Z","shell.execute_reply.started":"2024-03-14T20:47:48.549849Z","shell.execute_reply":"2024-03-14T20:47:48.549868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(df_test.head().T)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-03-14T20:47:48.551485Z","iopub.status.idle":"2024-03-14T20:47:48.551930Z","shell.execute_reply.started":"2024-03-14T20:47:48.551694Z","shell.execute_reply":"2024-03-14T20:47:48.551714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols_cat, cols_num, cols_oth = convert_dtypes(df_test, train_cols)\ncols_all = cols_cat + cols_num + [\"case_id\"]\ndf_test = df_test[cols_all]\ndisplay(df_test.dtypes)\n\nfor col in df_test.columns:\n    if col not in train_columns:\n        del df_test[col]","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-03-14T20:47:48.553283Z","iopub.status.idle":"2024-03-14T20:47:48.553766Z","shell.execute_reply.started":"2024-03-14T20:47:48.553530Z","shell.execute_reply":"2024-03-14T20:47:48.553550Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test = scaler.scale(df_test)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-03-14T20:00:46.517290Z","iopub.execute_input":"2024-03-14T20:00:46.517665Z","iopub.status.idle":"2024-03-14T20:00:46.521522Z","shell.execute_reply.started":"2024-03-14T20:00:46.517631Z","shell.execute_reply":"2024-03-14T20:00:46.520562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Predict and submit","metadata":{}},{"cell_type":"code","source":"!head /kaggle/input/home-credit-credit-risk-model-stability/sample_submission.csv","metadata":{"execution":{"iopub.status.busy":"2024-03-14T20:00:50.539525Z","iopub.execute_input":"2024-03-14T20:00:50.540178Z","iopub.status.idle":"2024-03-14T20:00:51.539101Z","shell.execute_reply.started":"2024-03-14T20:00:50.540146Z","shell.execute_reply":"2024-03-14T20:00:51.537887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test.replace([np.inf, -np.inf], np.nan, inplace=True)\npreds_proba_auc = model_auc.predict_proba(df_test[train_cols])\npreds_proba_rmse = model_rmse.predict_proba(df_test[train_cols])\n\ndf_test[\"score\"] = (np.clip(preds_proba_auc[:, 1], 0, 1) * 4 + np.clip(preds_proba_rmse[:, 1], 0, 1)) / 5\ndisplay(df_test[[\"case_id\", \"score\"]].head(10).T)","metadata":{"execution":{"iopub.status.busy":"2024-03-14T20:03:03.374638Z","iopub.execute_input":"2024-03-14T20:03:03.375014Z","iopub.status.idle":"2024-03-14T20:03:03.901241Z","shell.execute_reply.started":"2024-03-14T20:03:03.374986Z","shell.execute_reply":"2024-03-14T20:03:03.900193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# df_test[\"case_id\"] = df_test[\"case_id\"].astype(\"int32\")\n# df_test[\"score\"] = df_test[\"score\"].astype(\"float32\")","metadata":{"execution":{"iopub.status.busy":"2024-03-14T20:03:06.155906Z","iopub.execute_input":"2024-03-14T20:03:06.156533Z","iopub.status.idle":"2024-03-14T20:03:06.160651Z","shell.execute_reply.started":"2024-03-14T20:03:06.156503Z","shell.execute_reply":"2024-03-14T20:03:06.159701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test[[\"case_id\", \"score\"]].to_csv(\"submission.csv\", float_format='%.3f', index=False)\n\n!head submission.csv","metadata":{"execution":{"iopub.status.busy":"2024-03-14T20:03:06.609663Z","iopub.execute_input":"2024-03-14T20:03:06.610613Z","iopub.status.idle":"2024-03-14T20:03:07.604599Z","shell.execute_reply.started":"2024-03-14T20:03:06.610569Z","shell.execute_reply":"2024-03-14T20:03:07.603582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}