{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":50160,"databundleVersionId":7921029,"sourceType":"competition"}],"dockerImageVersionId":30664,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-03-17T08:43:30.748703Z","iopub.execute_input":"2024-03-17T08:43:30.748972Z","iopub.status.idle":"2024-03-17T08:43:31.102053Z","shell.execute_reply.started":"2024-03-17T08:43:30.748951Z","shell.execute_reply":"2024-03-17T08:43:31.101203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Черновик для дз 2 Савченко Михаил\n\nОригинальный код: https://www.kaggle.com/code/jirkaborovec/credit-risk-eda-xgboost-depth-0-gpu\nбудем улучшать\n\nчто было сделано относительно изначального решения:\n\nдобавлены новые признаки, проведена фильтраия малозначимых признаков, перебраны гиперпараметры модели, получена новая модель с лучшей точностью на валидационной выборке.  На тесте","metadata":{}},{"cell_type":"code","source":"import warnings\n# warnings.simplefilter(\"ignore\")\nwarnings.simplefilter(\"ignore\", UserWarning)\nimport os, glob\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom pathlib import Path\n\nPATH_DATASET = Path(\"/kaggle/input/home-credit-credit-risk-model-stability\")\nPATH_PARQUETS = PATH_DATASET / \"parquet_files\"\nPARQUETS_TRAIN = PATH_PARQUETS / \"train\"\nPARQUETS_TEST = PATH_PARQUETS / \"test\"\npd.set_option('display.max_columns', None)\npd.set_option('display.max_rows', None)","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:43:34.610990Z","iopub.execute_input":"2024-03-17T08:43:34.611383Z","iopub.status.idle":"2024-03-17T08:43:35.060571Z","shell.execute_reply.started":"2024-03-17T08:43:34.611359Z","shell.execute_reply":"2024-03-17T08:43:35.059628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = pd.read_parquet(PARQUETS_TRAIN / \"train_base.parquet\")\nprint(f\"size: {len(df_train)}\")\ndisplay(df_train.head())","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:43:37.381552Z","iopub.execute_input":"2024-03-17T08:43:37.381908Z","iopub.status.idle":"2024-03-17T08:43:37.642458Z","shell.execute_reply.started":"2024-03-17T08:43:37.381886Z","shell.execute_reply":"2024-03-17T08:43:37.641638Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"ура-ура мы подключились к данным","metadata":{}},{"cell_type":"code","source":"!cat /kaggle/input/home-credit-credit-risk-model-stability/feature_definitions.csv","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:43:40.541519Z","iopub.execute_input":"2024-03-17T08:43:40.542267Z","iopub.status.idle":"2024-03-17T08:43:40.858836Z","shell.execute_reply.started":"2024-03-17T08:43:40.542236Z","shell.execute_reply":"2024-03-17T08:43:40.857704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Дэээээ, немало, надо как-то привести в порядок","metadata":{}},{"cell_type":"markdown","source":"**Explore the training data**\nBorrowed from the competition describtion:\n\nTable Description This dataset contains a large number of tables as a result of utilizing diverse data sources and the varying levels of data aggregation used while preparing the dataset. Note: All files listed below are found in both .csv and .parquet formats.\n\nDepth values\n\ndepth=0 - These are static features directly tied to a specific case_id.\n\ndepth=1 - Each case_id has an associated historical record, indexed by num_group1.\n\ndepth=2 - Each case_id has an associated historical record, indexed by both num_group1 and num_group2.\n\nYou can read more about Credit bureau (CB) here https://en.wikipedia.org/wiki/Credit_bureau.\n\nVarious predictors were transformed, therefore we have the following notation for similar groups of transformations\n\nP - Transform DPD (Days past due)\n\nM - Masking categories\n\nA - Transform amount\n\nD - Transform date\n\nT - Unspecified Transform\n\nL - Unspecified Transform\n\nPlease note that transformations within a group are denoted by a capital letter at the end of the predictor name (e.g., maxdbddpdtollast6m_4187119P). We hope that this will simplify the manipulation with predictors.","metadata":{}},{"cell_type":"code","source":"def _short_array(arr):\n    if len(arr) <= 5:\n        return repr(arr)\n    return f\"[{', '.join(map(str, arr[:2]))}, ..., {', '.join(map(str, arr[-2:]))}]\"\n\n# taking inpiration with column names from https://www.kaggle.com/code/greysky/home-credit-baseline\ndef convert_dtypes(df):\n    cols = []\n    for col, dt in dict(df.dtypes).items():\n        if col.startswith(\"for\"):\n            df[col] = df[col].fillna(0).astype(\"int16\")\n        elif \"num\" in col or \"cnt\" in col:\n            df[col] = df[col].fillna(0).astype(\"int32\")\n        elif col.startswith(\"pct\"):\n            df[col] = df[col].astype(\"float16\")\n        elif col[-1] in (\"A\", \"P\"):\n            df[col] = df[col].astype(\"float32\")\n        elif col[-1] in (\"D\", ):\n            df[col] = pd.to_datetime(df[col])\n        elif col[-1] in (\"M\", \"L\"):\n            if col[-1] == \"L\" and dt.name.startswith(\"int\"):\n                df[col] = df[col].astype(\"int32\")\n            elif col[-1] == \"L\" and dt.name.startswith(\"float\"):\n                df[col] = df[col].astype(\"float32\")\n            else:\n                uq = list(df[col].unique())\n                print(f'{col} -> #{len(uq)} -> {_short_array(uq)}')\n                df[col] = df[col].astype(\"category\")\n        if col[-1] in (\"A\", \"P\", \"M\", \"L\"):\n            cols.append(col)\n    return cols","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:43:52.210811Z","iopub.execute_input":"2024-03-17T08:43:52.211115Z","iopub.status.idle":"2024-03-17T08:43:52.220691Z","shell.execute_reply.started":"2024-03-17T08:43:52.211092Z","shell.execute_reply":"2024-03-17T08:43:52.219929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"это суперудобно. исходный кернел крутой, респект мужику\n\nтеперь будем собирать трейновый датасет","metadata":{}},{"cell_type":"code","source":"df_train[\"date_decision\"] = pd.to_datetime(df_train[\"date_decision\"]).dt.date\n# delete redundat cols\ndel df_train[\"MONTH\"], df_train[\"WEEK_NUM\"]","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:44:10.011847Z","iopub.execute_input":"2024-03-17T08:44:10.012161Z","iopub.status.idle":"2024-03-17T08:44:10.479346Z","shell.execute_reply.started":"2024-03-17T08:44:10.012139Z","shell.execute_reply":"2024-03-17T08:44:10.478491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def merge_parquets(df, name, prefix=\"train\", folder=PATH_PARQUETS):\n    df_ = pd.concat(\n        [pd.read_parquet(p) for p in glob.glob(str(folder / prefix / f\"{prefix}_{name}*.parquet\"))],\n    )\n    if \"num_group1\" in df_.columns:\n        del df_[\"num_group1\"]\n    df_.drop_duplicates(inplace=True)\n    print(f\"{name} size: {len(df_)} with features: {len(df_.columns)}\")\n    display(df_.head())\n    cols = convert_dtypes(df_) + [\"case_id\"]\n    df_ = df_[cols]\n    if len(df_) > len(df_[\"case_id\"].unique()):\n        df_ = df_.groupby(['case_id'], as_index=False).first()\n    df = df.merge(df_, how=\"left\", on=\"case_id\")\n    print(f\"fused size: {len(df)}\")\n    return df","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:44:13.843113Z","iopub.execute_input":"2024-03-17T08:44:13.843406Z","iopub.status.idle":"2024-03-17T08:44:13.849589Z","shell.execute_reply.started":"2024-03-17T08:44:13.843385Z","shell.execute_reply":"2024-03-17T08:44:13.848854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = merge_parquets(df_train, \"static_0\")","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:44:20.161819Z","iopub.execute_input":"2024-03-17T08:44:20.162140Z","iopub.status.idle":"2024-03-17T08:44:41.127367Z","shell.execute_reply.started":"2024-03-17T08:44:20.162114Z","shell.execute_reply":"2024-03-17T08:44:41.126452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = merge_parquets(df_train, \"static_cb_0\")\n","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:44:48.919003Z","iopub.execute_input":"2024-03-17T08:44:48.919302Z","iopub.status.idle":"2024-03-17T08:44:54.012515Z","shell.execute_reply.started":"2024-03-17T08:44:48.919275Z","shell.execute_reply":"2024-03-17T08:44:54.011723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"data size: {len(df_train)}\")\nprint(f\"unique: {len(df_train['case_id'].unique())}\")","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:44:57.002433Z","iopub.execute_input":"2024-03-17T08:44:57.002757Z","iopub.status.idle":"2024-03-17T08:44:57.040256Z","shell.execute_reply.started":"2024-03-17T08:44:57.002735Z","shell.execute_reply":"2024-03-17T08:44:57.039566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Мы хотим добавить новых признаков:\n\nдавайте долбвим\n\nИ тут оказывается что памяти не хватает и пупупу, так что методом перебора(добавляли по файлу и перезапусали, смотрели на результат) оставили файл с самыми ценными фичами\n\nА почему так? а потому что РАМы не хватает даже два файла прицепить((((((","metadata":{}},{"cell_type":"code","source":"df_train = merge_parquets(df_train, \"deposit_1\")\n","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:45:02.241150Z","iopub.execute_input":"2024-03-17T08:45:02.241425Z","iopub.status.idle":"2024-03-17T08:45:02.547705Z","shell.execute_reply.started":"2024-03-17T08:45:02.241405Z","shell.execute_reply":"2024-03-17T08:45:02.546779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:45:16.703269Z","iopub.execute_input":"2024-03-17T08:45:16.703573Z","iopub.status.idle":"2024-03-17T08:45:16.808768Z","shell.execute_reply.started":"2024-03-17T08:45:16.703552Z","shell.execute_reply":"2024-03-17T08:45:16.807684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.groupby('date_decision')['target'].mean().plot(\n    figsize=(14, 2), grid=True,\n    xlabel=\"date of decision\",\n    ylabel=\"day mean / proxi ratio\",\n)","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:45:26.473267Z","iopub.execute_input":"2024-03-17T08:45:26.473582Z","iopub.status.idle":"2024-03-17T08:45:26.896985Z","shell.execute_reply.started":"2024-03-17T08:45:26.473561Z","shell.execute_reply":"2024-03-17T08:45:26.896371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Данные подготовили, начинаем обучение**","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\ndf_train.replace([np.inf, -np.inf], np.nan, inplace=True)\ntrain_cols = [c for c in df_train.columns if c not in (\"case_id\", \"date_decision\", \"target\")]\nX_train, X_valid, y_train, y_valid = train_test_split(\n    df_train[train_cols], df_train['target'], test_size=0.2)","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:45:30.983383Z","iopub.execute_input":"2024-03-17T08:45:30.983856Z","iopub.status.idle":"2024-03-17T08:45:32.837306Z","shell.execute_reply.started":"2024-03-17T08:45:30.983832Z","shell.execute_reply":"2024-03-17T08:45:32.836691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import xgboost as xgb\n\nprint(xgb.__version__)","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:45:35.391256Z","iopub.execute_input":"2024-03-17T08:45:35.391785Z","iopub.status.idle":"2024-03-17T08:45:35.545962Z","shell.execute_reply.started":"2024-03-17T08:45:35.391761Z","shell.execute_reply":"2024-03-17T08:45:35.545203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = xgb.XGBClassifier(\n    device=\"cuda\",\n    objective='binary:logistic',\n    enable_categorical=True,\n    eval_metric='auc',\n    #learning_rate=0.05,\n    subsample=1,\n    colsample_bytree=1,\n    min_child_weight=1,\n    #gamma=0.7,\n    #reg_alpha=0.7,\n    max_depth=20,\n    n_estimators=800,\n    random_state=42,\n)\n\n# Training the model on the training data\nmodel.fit(\n    X_train, y_train,\n    eval_set=[(X_valid, y_valid)],\n    early_stopping_rounds=5,\n    verbose=True,\n)\n\nprint(model)\n","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:46:52.131451Z","iopub.execute_input":"2024-03-17T08:46:52.131780Z","iopub.status.idle":"2024-03-17T08:50:44.442553Z","shell.execute_reply.started":"2024-03-17T08:46:52.131758Z","shell.execute_reply":"2024-03-17T08:50:44.441745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Нам хочется как-то улучшиться для этого посмотрим какие фичи имеют наибольшие веса","metadata":{}},{"cell_type":"code","source":"feature_importances = model.feature_importances_\n\n# Сопоставление важности с именами колонок и сортировка\nfeature_names = X_train.columns\nfeature_importances_df = pd.DataFrame({'Feature': feature_names, 'Importance': feature_importances})\nfeature_importances_df = feature_importances_df.sort_values(by='Importance', ascending=False)\n\n# Вывод важности фич\nprint(feature_importances_df)\n\n","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:52:06.095066Z","iopub.execute_input":"2024-03-17T08:52:06.097338Z","iopub.status.idle":"2024-03-17T08:52:06.132780Z","shell.execute_reply.started":"2024-03-17T08:52:06.097304Z","shell.execute_reply":"2024-03-17T08:52:06.131772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Это уже наши улучшения**\n\nМы видим что одна фича имеет критический вес(кстати по сравнению с оригиналом её вес уменьшился на 5%. Это означает, что наша модель стала умнее), и многие имеют вес примерно никакой, давайте избавимся от тех клонок,\nгде вес никакой и попробуем обучиться заново","metadata":{}},{"cell_type":"code","source":"zero_importance_features = feature_importances_df[feature_importances_df['Importance'] == 0]['Feature'].tolist()\nzero_importance_features.append(\"case_id\")\nzero_importance_features.append(\"date_decision\")\nzero_importance_features.append(\"target\")\n\ntrain_cols_2 = [c for c in df_train.columns if c not in zero_importance_features]\nX_train_2, X_valid_2, y_train_2, y_valid_2 = train_test_split(\n    df_train[train_cols_2], df_train['target'], test_size=0.2)","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:52:33.442208Z","iopub.execute_input":"2024-03-17T08:52:33.442555Z","iopub.status.idle":"2024-03-17T08:52:35.065031Z","shell.execute_reply.started":"2024-03-17T08:52:33.442533Z","shell.execute_reply":"2024-03-17T08:52:35.064142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_2 = xgb.XGBClassifier(\n    device=\"cuda\",\n    objective='binary:logistic',\n    enable_categorical=True,\n    eval_metric='auc',\n    #learning_rate=0.05,\n    subsample=1,\n    colsample_bytree=1,\n    min_child_weight=1,\n    #gamma=0.7,\n    #reg_alpha=0.7,\n    max_depth=20,\n    n_estimators=800,\n    random_state=42,\n)\n\nmodel_2.fit(\n    X_train_2, y_train_2,\n    eval_set=[(X_valid_2, y_valid_2)],\n    early_stopping_rounds=5,\n    verbose=True,\n)\n\nprint(model)","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:53:27.331143Z","iopub.execute_input":"2024-03-17T08:53:27.331474Z","iopub.status.idle":"2024-03-17T08:57:26.853807Z","shell.execute_reply.started":"2024-03-17T08:53:27.331432Z","shell.execute_reply":"2024-03-17T08:57:26.853179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\nага, не очень полегчало, но данные мы почистили и добавили признаков\n\nТеперь хочется найти наиболее подходящие гиперпараметры. Для этого запустим автоматический переьор -- с богом!","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import RandomizedSearchCV\n\nparam_dist = {\n    'learning_rate': np.linspace(0.01, 0.2, 10),\n    'subsample': [0.5, 0.7, 1.0],\n    'colsample_bytree': [0.5, 0.7, 1.0],\n    'min_child_weight': [1, 3, 5],\n    'gamma': [0, 0.1, 0.3],\n    'reg_alpha': [0, 0.01, 0.1],\n    'max_depth': [6, 10, 15, 20],\n    'n_estimators': [100, 200, 500, 800]\n}\n\nmodel_3 = xgb.XGBClassifier(\n    device=\"cuda\",\n    objective='binary:logistic',\n    enable_categorical=True,\n    eval_metric='auc',\n    random_state=42,\n)\n\n# Настройка RandomizedSearchCV\nrandom_search = RandomizedSearchCV(\n    model_3, \n    param_distributions=param_dist, \n    n_iter=10,  \n    scoring='roc_auc', \n    error_score=0, \n    verbose=3, \n    n_jobs=-1, \n    cv=3,\n    random_state=42\n)\n\n# Запуск поискане будем делать т к долго и не нужно (уже нашли лучшие гиперпараметры)\n#random_search.fit(\n#    X_train_2, \n#    y_train_2, \n#    eval_set=[(X_valid_2, y_valid_2)], \n#    early_stopping_rounds=5, \n#    verbose=True,\n#)\n\n#print(f\"Лучшие параметры: {random_search.best_params_}\")\n\n#best_model = random_search.best_estimator_","metadata":{"execution":{"iopub.status.busy":"2024-03-16T15:40:41.190263Z","iopub.execute_input":"2024-03-16T15:40:41.190744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"я не хочу перезапусать подбор ещё раз, потому что долго бежало. Но переборами были достигнуты результаты в 0.8, что на 4% лучше чем без перебора в оригинальном ноутбуке\n\nЗначит можно с гордостью говорить, что мы молодцы\n\nдалее запущено обучение с лучшими гиперпараметрами","metadata":{}},{"cell_type":"code","source":"model_fin = xgb.XGBClassifier(\n    device=\"cuda\",\n    objective='binary:logistic',\n    enable_categorical=True,\n    eval_metric='auc',\n    learning_rate=0.03111111111111111,\n    subsample=1,\n    colsample_bytree=0.5,\n    min_child_weight=5,\n    gamma=0,\n    reg_alpha=0.01,\n    max_depth=20,\n    n_estimators=200, #можно и 800 но надоело переобучать\n    random_state=42,\n)\n\nmodel_fin.fit(\n    X_train_2, y_train_2,\n    eval_set=[(X_valid_2, y_valid_2)],\n    early_stopping_rounds=5,\n    verbose=True,\n)\n\nprint(model)","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:58:45.391431Z","iopub.execute_input":"2024-03-17T08:58:45.391758Z","iopub.status.idle":"2024-03-17T09:23:12.214915Z","shell.execute_reply.started":"2024-03-17T08:58:45.391737Z","shell.execute_reply":"2024-03-17T09:23:12.213662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"ну как бы улучшение есть -- я большой молодец\n\nОсталось выгрузить результаты","metadata":{}},{"cell_type":"code","source":"df_test = pd.read_parquet(PARQUETS_TEST / \"test_base.parquet\")\nprint(f\"size: {len(df_test)}\")\ndisplay(df_test.head())","metadata":{"execution":{"iopub.status.busy":"2024-03-17T09:25:13.202832Z","iopub.execute_input":"2024-03-17T09:25:13.203213Z","iopub.status.idle":"2024-03-17T09:25:13.229865Z","shell.execute_reply.started":"2024-03-17T09:25:13.203189Z","shell.execute_reply":"2024-03-17T09:25:13.229032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test[\"date_decision\"] = pd.to_datetime(df_test[\"date_decision\"]).dt.date\n# delete redundat cols\ndel df_test[\"MONTH\"], df_test[\"WEEK_NUM\"]","metadata":{"execution":{"iopub.status.busy":"2024-03-17T09:25:17.723554Z","iopub.execute_input":"2024-03-17T09:25:17.723864Z","iopub.status.idle":"2024-03-17T09:25:17.731046Z","shell.execute_reply.started":"2024-03-17T09:25:17.723843Z","shell.execute_reply":"2024-03-17T09:25:17.730265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for name in [\"static_0\", \"static_cb_0\", \"deposit_1\"]:\n    df_test = merge_parquets(df_test, name, \"test\")","metadata":{"execution":{"iopub.status.busy":"2024-03-17T09:25:22.562346Z","iopub.execute_input":"2024-03-17T09:25:22.562698Z","iopub.status.idle":"2024-03-17T09:25:22.924671Z","shell.execute_reply.started":"2024-03-17T09:25:22.562674Z","shell.execute_reply":"2024-03-17T09:25:22.923782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for name in [\n    # \"other_1\",\n    # \"deposit_1\",\n    # \"debitcard_1\",\n    \"person_1\",\n    # \"tax_registry_a_1\",\n    # \"tax_registry_b_1\",\n    # \"tax_registry_c_1\"\n]:\n    df_test = merge_parquets(df_test, name, \"test\")","metadata":{"execution":{"iopub.status.busy":"2024-03-17T09:25:28.171047Z","iopub.execute_input":"2024-03-17T09:25:28.171836Z","iopub.status.idle":"2024-03-17T09:25:28.237245Z","shell.execute_reply.started":"2024-03-17T09:25:28.171805Z","shell.execute_reply":"2024-03-17T09:25:28.236539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = df_train.apply(lambda x: x.astype(np.float32) if x.dtype == np.float16 else x)\ndf_train.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:37:46.083418Z","iopub.execute_input":"2024-03-17T08:37:46.083793Z","iopub.status.idle":"2024-03-17T08:37:46.653067Z","shell.execute_reply.started":"2024-03-17T08:37:46.083764Z","shell.execute_reply":"2024-03-17T08:37:46.651446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_columns_dtypes = dict(df_train.dtypes)\n#del df_train\nfor col, dt in train_columns_dtypes.items():\n    if col not in train_cols_2:\n        continue\n    try:\n        df_test[col] = df_test[col].astype(dt)\n    except:\n        print(f\"failed converting {col} to {dt}\")","metadata":{"execution":{"iopub.status.busy":"2024-03-17T09:25:35.863190Z","iopub.execute_input":"2024-03-17T09:25:35.863546Z","iopub.status.idle":"2024-03-17T09:25:35.905740Z","shell.execute_reply.started":"2024-03-17T09:25:35.863513Z","shell.execute_reply":"2024-03-17T09:25:35.904934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-17T08:32:48.102516Z","iopub.execute_input":"2024-03-17T08:32:48.102944Z","iopub.status.idle":"2024-03-17T08:32:48.222244Z","shell.execute_reply.started":"2024-03-17T08:32:48.102912Z","shell.execute_reply":"2024-03-17T08:32:48.221185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test.replace([np.inf, -np.inf], np.nan, inplace=True)\npreds_proba = model_fin.predict_proba(df_test[train_cols_2])\n\ndf_test[\"score\"] = np.clip(preds_proba[:, 1], 0, 1)\ndisplay(df_test[[\"case_id\", \"date_decision\", \"score\"]].head(10).T)","metadata":{"execution":{"iopub.status.busy":"2024-03-17T09:25:40.540003Z","iopub.execute_input":"2024-03-17T09:25:40.540322Z","iopub.status.idle":"2024-03-17T09:25:40.650986Z","shell.execute_reply.started":"2024-03-17T09:25:40.540300Z","shell.execute_reply":"2024-03-17T09:25:40.650307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test[[\"case_id\", \"score\"]].to_csv(\"submission.csv\", float_format='%.3f', index=False)\n\n!head submission.csv","metadata":{"execution":{"iopub.status.busy":"2024-03-17T09:25:45.600177Z","iopub.execute_input":"2024-03-17T09:25:45.601036Z","iopub.status.idle":"2024-03-17T09:25:46.029760Z","shell.execute_reply.started":"2024-03-17T09:25:45.600995Z","shell.execute_reply":"2024-03-17T09:25:46.028740Z"},"trusted":true},"execution_count":null,"outputs":[]}]}