{"metadata":{"kaggle":{"accelerator":"none","dataSources":[{"sourceId":50160,"databundleVersionId":7921029,"sourceType":"competition"}],"dockerImageVersionId":30664,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.10.13"},"papermill":{"default_parameters":{},"duration":771.875757,"end_time":"2024-03-16T22:46:54.059247","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2024-03-16T22:34:02.183490","version":"2.5.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"В этом ноутбуке осуществляется обучение модели с использованием заранее выбранных признаков и гиперпараметров, а также предсказание для тестовой выборки.\nКод, с помощью которого выбирались признаки и гиперпараметры, представлен в ноутбуке, но запускался отдельно (чтобы сэкономить время при посылке в соревнование).\n- [Отбор признаков](https://www.kaggle.com/code/trickmanoff/credit-risk-analysis-submit-2?scriptVersionId=167428534)\n- [Подбор гиперпараметров](https://www.kaggle.com/code/trickmanoff/credit-risk-analysis-submit/log?scriptVersionId=167368350)","metadata":{}},{"cell_type":"code","source":"# all imports are here\nimport json\nimport re\nimport time\nfrom copy import copy\nfrom dataclasses import dataclass\nfrom pathlib import Path\nfrom typing import Dict, List, Sequence, Tuple\n\nimport catboost\nimport numpy as np\nimport optuna\nimport pandas as pd\nimport polars as pl\nfrom catboost import CatBoostClassifier\nfrom matplotlib import pyplot as plt\nfrom matplotlib.lines import Line2D\nfrom sklearn import preprocessing\nfrom sklearn.base import BaseEstimator, TransformerMixin\nfrom sklearn.linear_model import LinearRegression, LogisticRegression\nfrom sklearn.metrics import roc_auc_score\nfrom tqdm import tqdm","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","papermill":{"duration":4.853147,"end_time":"2024-03-16T22:34:10.167515","exception":false,"start_time":"2024-03-16T22:34:05.314368","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:11.412125Z","iopub.execute_input":"2024-03-17T11:38:11.412819Z","iopub.status.idle":"2024-03-17T11:38:16.992474Z","shell.execute_reply.started":"2024-03-17T11:38:11.412779Z","shell.execute_reply":"2024-03-17T11:38:16.990893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"@dataclass\nclass NotebookRunConfig:\n    # if False, then feature selection and hypereparameters tuning is done from scratch\n    # else, this part is skipped and precalculated values are used\n    use_fixed_config: bool = True\n    \n    debug: bool = False\n\nNOTEBOOK_RUN_CONFIG = NotebookRunConfig(use_fixed_config=True, debug=False)","metadata":{"papermill":{"duration":0.024736,"end_time":"2024-03-16T22:34:10.207047","exception":false,"start_time":"2024-03-16T22:34:10.182311","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:16.995330Z","iopub.execute_input":"2024-03-17T11:38:16.996028Z","iopub.status.idle":"2024-03-17T11:38:17.004865Z","shell.execute_reply.started":"2024-03-17T11:38:16.995981Z","shell.execute_reply":"2024-03-17T11:38:17.003656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DATA_DIRPATH = Path('/kaggle/input/home-credit-credit-risk-model-stability')\nPARQUET_DATA_DIRPATH = DATA_DIRPATH / 'parquet_files'","metadata":{"papermill":{"duration":0.023139,"end_time":"2024-03-16T22:34:10.244823","exception":false,"start_time":"2024-03-16T22:34:10.221684","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.006533Z","iopub.execute_input":"2024-03-17T11:38:17.007816Z","iopub.status.idle":"2024-03-17T11:38:17.016895Z","shell.execute_reply.started":"2024-03-17T11:38:17.007773Z","shell.execute_reply":"2024-03-17T11:38:17.015654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_train_val_splits(val_weeks_cnt: int = 5) -> Tuple[pl.DataFrame, List[int], pl.DataFrame, List[int]]:\n    \"\"\"\n    Split data into train and validation.\n    The last `val_weeks_cnt` weeks are used for validation, others - for training\n    \"\"\"\n    train_base_df = pl.read_parquet(PARQUET_DATA_DIRPATH / 'train' / 'train_base.parquet')\n    first_val_week_num = train_base_df.select('WEEK_NUM').unique().sort('WEEK_NUM')[-val_weeks_cnt].rows()[0][0]\n    \n    train_data = train_base_df.filter(train_base_df['WEEK_NUM'] < first_val_week_num)\n    val_data = train_base_df.filter(train_base_df['WEEK_NUM'] >= first_val_week_num)\n    \n    train_fraction = len(train_data) / len(train_base_df)\n    print(f'train: {train_fraction * 100:.2f}%, val: {(1 - train_fraction) * 100:.2f}%')\n    \n    train_x, train_y = train_data.drop('target'), train_data['target']\n    val_x, val_y = val_data.drop('target'), val_data['target']\n    \n    return train_x, train_y.to_list(), val_x, val_y.to_list()","metadata":{"papermill":{"duration":0.028778,"end_time":"2024-03-16T22:34:10.288340","exception":false,"start_time":"2024-03-16T22:34:10.259562","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.020136Z","iopub.execute_input":"2024-03-17T11:38:17.021611Z","iopub.status.idle":"2024-03-17T11:38:17.034955Z","shell.execute_reply.started":"2024-03-17T11:38:17.021563Z","shell.execute_reply":"2024-03-17T11:38:17.033726Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Научимся вычислять score:","metadata":{"papermill":{"duration":0.014116,"end_time":"2024-03-16T22:34:10.316968","exception":false,"start_time":"2024-03-16T22:34:10.302852","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def eval_predicted_score(X, target: Sequence[int], pred_scores: Sequence[float], plot: bool = True) -> float:\n    X = X.with_columns([\n        pl.Series(target).alias('target'),\n        pl.Series(pred_scores).alias('pred_score'),\n    ])\n    \n    week_nums = []\n    gini_scores = []\n    \n    for week_num, data in X.group_by(['WEEK_NUM']):\n        roc_auc = roc_auc_score(data['target'], data['pred_score'])\n        week_nums.append(week_num[0])\n        gini_scores.append(2*roc_auc - 1)\n    \n    week_nums_x = np.array(week_nums)[:, None]\n    lin_model = LinearRegression()\n    lin_model.fit(week_nums_x, gini_scores)\n    falling_rate = min(0, lin_model.coef_[0])\n    \n    lin_model_preds = lin_model.predict(week_nums_x)\n    residuals = np.array(gini_scores) - lin_model_preds\n    residuals_std = residuals.std()\n    \n    loss = np.mean(gini_scores) + 88 * falling_rate - 0.5 * residuals.std()\n    \n    if plot:\n        fig, ax = plt.subplots()\n        \n        week_nums, gini_scores, lin_model_preds = list(zip(*sorted(zip(week_nums, gini_scores, lin_model_preds))))\n        ax.scatter(week_nums, gini_scores, label='gini score')\n        ax.plot(week_nums, lin_model_preds, label='lin model', color='orange')\n        ax.set_xlabel('week')\n        ax.set_ylabel('gini score')\n        ax.legend()\n        \n        # metrics values\n        metrics_handles = [Line2D([0], [0], color='white', lw=3)] * 3\n        \n        labels = []\n        labels.append(f'mean Gini score = {np.mean(gini_scores):.2f}')\n        labels.append(f'falling rate = {falling_rate:.2f}')\n        labels.append(f'residuals std = {residuals.std():.2f}')\n        # ----\n        \n        ax.legend(metrics_handles, labels, loc='lower right',\n                  fancybox=True, framealpha=0.7)\n        \n        plt.show()\n        \n    return loss","metadata":{"papermill":{"duration":0.032689,"end_time":"2024-03-16T22:34:10.365870","exception":false,"start_time":"2024-03-16T22:34:10.333181","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.039860Z","iopub.execute_input":"2024-03-17T11:38:17.042485Z","iopub.status.idle":"2024-03-17T11:38:17.063876Z","shell.execute_reply.started":"2024-03-17T11:38:17.042430Z","shell.execute_reply":"2024-03-17T11:38:17.062612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Проверим вычисление score при случайном предсказании для всех объектов:","metadata":{"papermill":{"duration":0.014212,"end_time":"2024-03-16T22:34:10.394589","exception":false,"start_time":"2024-03-16T22:34:10.380377","status":"completed"},"tags":[]}},{"cell_type":"code","source":"if not NOTEBOOK_RUN_CONFIG.use_fixed_config:\n    train_X, train_y, val_X, val_y = get_train_val_splits()","metadata":{"papermill":{"duration":0.022206,"end_time":"2024-03-16T22:34:10.431336","exception":false,"start_time":"2024-03-16T22:34:10.409130","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.068515Z","iopub.execute_input":"2024-03-17T11:38:17.072679Z","iopub.status.idle":"2024-03-17T11:38:17.078292Z","shell.execute_reply.started":"2024-03-17T11:38:17.072626Z","shell.execute_reply":"2024-03-17T11:38:17.077069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class SimpleScoringPredictor(BaseEstimator):\n    def fit(self, X, y):\n        return self\n    \n    def predict(self, X) -> Sequence[float]:\n        return np.random.uniform(0., 1., (len(X),))","metadata":{"papermill":{"duration":0.022862,"end_time":"2024-03-16T22:34:10.468538","exception":false,"start_time":"2024-03-16T22:34:10.445676","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.082522Z","iopub.execute_input":"2024-03-17T11:38:17.083658Z","iopub.status.idle":"2024-03-17T11:38:17.092221Z","shell.execute_reply.started":"2024-03-17T11:38:17.083609Z","shell.execute_reply":"2024-03-17T11:38:17.090642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if NOTEBOOK_RUN_CONFIG.debug:\n    simple_predictor = simple_predictor.fit(train_X, train_y)  # nothing happens here\n    pred_val_y_scores = simple_predictor.predict(val_X)\n\n    score = eval_predicted_score(val_X, val_y, pred_val_y_scores)\n    print(f'Total score: {score:.2f}')","metadata":{"papermill":{"duration":0.032796,"end_time":"2024-03-16T22:34:10.515754","exception":false,"start_time":"2024-03-16T22:34:10.482958","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.094601Z","iopub.execute_input":"2024-03-17T11:38:17.095725Z","iopub.status.idle":"2024-03-17T11:38:17.106934Z","shell.execute_reply.started":"2024-03-17T11:38:17.095680Z","shell.execute_reply":"2024-03-17T11:38:17.105646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Отбор признаков","metadata":{"papermill":{"duration":0.01579,"end_time":"2024-03-16T22:34:10.549477","exception":false,"start_time":"2024-03-16T22:34:10.533687","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"Кажется, что слишком сложно вручную отсмотреть все признаки и интуитивно оценить, какие из них будут полезны для предсказания.\nПоэтому будет пробовать отбирать потенциально полезные признаки некоторым эвристическим методом.","metadata":{"papermill":{"duration":0.014165,"end_time":"2024-03-16T22:34:10.581563","exception":false,"start_time":"2024-03-16T22:34:10.567398","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"### 1. Статические признаки","metadata":{"papermill":{"duration":0.014136,"end_time":"2024-03-16T22:34:10.610170","exception":false,"start_time":"2024-03-16T22:34:10.596034","status":"completed"},"tags":[]}},{"cell_type":"code","source":"FEATURES_TABLES_PATTERNS = {\n    'internal_static': '.*_static_\\d{1}_\\d{1}\\.parquet',\n    'external_static': '.*_static_cb_\\d{1}\\.parquet',\n}\n\n\ndef add_features(X: pl.DataFrame, dirpath: Path, pattern: str) -> pl.DataFrame:\n    pattern = re.compile(pattern)\n    \n    full_features_df = pl.concat([\n        pl.read_parquet(table_filepath)\n        for table_filepath in dirpath.iterdir()\n        if pattern.match(table_filepath.name)\n    ])\n    \n    return X.join(full_features_df, on='case_id', how='left')\n\n\ndef add_all_static_features(X: pl.DataFrame, dirpath: Path) -> pl.DataFrame:\n    for pattern_name in ['internal_static', 'external_static']:\n        X = add_features(X, dirpath, FEATURES_TABLES_PATTERNS[pattern_name])\n    return X","metadata":{"papermill":{"duration":0.026625,"end_time":"2024-03-16T22:34:10.651228","exception":false,"start_time":"2024-03-16T22:34:10.624603","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.109499Z","iopub.execute_input":"2024-03-17T11:38:17.110429Z","iopub.status.idle":"2024-03-17T11:38:17.122109Z","shell.execute_reply.started":"2024-03-17T11:38:17.110383Z","shell.execute_reply":"2024-03-17T11:38:17.120764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Поймём, какие признаки являются категориальными, а какие численными:","metadata":{"papermill":{"duration":0.015509,"end_time":"2024-03-16T22:34:10.685211","exception":false,"start_time":"2024-03-16T22:34:10.669702","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"Будем считать, что все признаки, имеющие тип float, не являются категориальными.\nДля оставшихся признаков посмотрим на отношение кол-ва различных значений к кол-ву не NaN значений.","metadata":{"papermill":{"duration":0.016343,"end_time":"2024-03-16T22:34:10.720797","exception":false,"start_time":"2024-03-16T22:34:10.704454","status":"completed"},"tags":[]}},{"cell_type":"code","source":"if not NOTEBOOK_RUN_CONFIG.use_fixed_config:\n    tmp_X = add_all_static_features(train_X, PARQUET_DATA_DIRPATH / 'train').drop('date_decision', 'MONTH', 'WEEK_NUM')\n    numeric_features = set(tmp_X.select(pl.col(pl.Float64)).columns)\n    tmp_X = tmp_X.drop(numeric_features)\n\n    not_null_counts = (tmp_X.null_count() - len(tmp_X)).select(pl.all().neg())\n    unique_to_not_null_fraction = tmp_X.select(pl.all().n_unique()) / not_null_counts\n    unique_to_not_null_fraction = np.array(unique_to_not_null_fraction.rows()[0])\n    unique_to_not_null_fraction = unique_to_not_null_fraction[unique_to_not_null_fraction < 0.2]\n\n    fig, ax = plt.subplots()\n\n    ax.hist(unique_to_not_null_fraction, bins=100)\n\n    ax.set_xlabel('#unique / #not null')\n    ax.set_title('Distribution for static non-float columns')\n\n    plt.show()","metadata":{"papermill":{"duration":0.037224,"end_time":"2024-03-16T22:34:10.778012","exception":false,"start_time":"2024-03-16T22:34:10.740788","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.129543Z","iopub.execute_input":"2024-03-17T11:38:17.130524Z","iopub.status.idle":"2024-03-17T11:38:17.143067Z","shell.execute_reply.started":"2024-03-17T11:38:17.130480Z","shell.execute_reply":"2024-03-17T11:38:17.141565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Пожалуй, в качестве порога можно взять 0.02.\nТ.е. считать признак категориальным, если отношение числа его уникальных значений к кол-ву не NaN значений не больше 0.02.","metadata":{"papermill":{"duration":0.014401,"end_time":"2024-03-16T22:34:10.809569","exception":false,"start_time":"2024-03-16T22:34:10.795168","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def get_features(data: pl.DataFrame) -> Tuple[List[str], List[str]]:\n    \"\"\"\n    :returns: names of numerical features, names of categorical features\n    \"\"\"\n    categorical_unique_to_non_null_threshold = 0.02\n    \n    numerical_features = data.select(pl.col(pl.Float64)).columns\n    data = data.drop(numerical_features)\n    \n    not_null_counts = (data.null_count() - len(data)).select(pl.all().neg())\n    unique_to_not_null_fraction = data.select(pl.all().n_unique()) / not_null_counts\n    categorical_features = unique_to_not_null_fraction.melt().filter(pl.col('value') < 0.03)['variable'].to_list()\n    \n    return numerical_features, categorical_features","metadata":{"papermill":{"duration":0.026057,"end_time":"2024-03-16T22:34:10.850313","exception":false,"start_time":"2024-03-16T22:34:10.824256","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.144837Z","iopub.execute_input":"2024-03-17T11:38:17.146360Z","iopub.status.idle":"2024-03-17T11:38:17.162904Z","shell.execute_reply.started":"2024-03-17T11:38:17.146311Z","shell.execute_reply":"2024-03-17T11:38:17.161548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def add_static_features_with_filtering(X: pl.DataFrame, dirpath: Path) -> Tuple[pl.DataFrame, List[str], List[str]]:\n    prev_cols = X.columns\n    X = add_all_static_features(X, dirpath)\n    new_cols = X.columns\n    print(f'{len(new_cols) - len(prev_cols)} static columns added')\n    numerical_features, categorical_features = get_features(X.drop(prev_cols))\n    X = X.select(prev_cols + numerical_features + categorical_features)\n    new_filtered_cols = X.columns\n    print(f'{len(new_filtered_cols) - len(prev_cols)} static columns left')\n    \n    return X, numerical_features, categorical_features","metadata":{"papermill":{"duration":0.024685,"end_time":"2024-03-16T22:34:10.889641","exception":false,"start_time":"2024-03-16T22:34:10.864956","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.165370Z","iopub.execute_input":"2024-03-17T11:38:17.166459Z","iopub.status.idle":"2024-03-17T11:38:17.178879Z","shell.execute_reply.started":"2024-03-17T11:38:17.166342Z","shell.execute_reply":"2024-03-17T11:38:17.177652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if not NOTEBOOK_RUN_CONFIG.use_fixed_config:\n    train_X_with_static, numerical_features, cat_features = add_static_features_with_filtering(train_X, PARQUET_DATA_DIRPATH / 'train')","metadata":{"papermill":{"duration":0.024454,"end_time":"2024-03-16T22:34:10.928574","exception":false,"start_time":"2024-03-16T22:34:10.904120","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.181495Z","iopub.execute_input":"2024-03-17T11:38:17.182385Z","iopub.status.idle":"2024-03-17T11:38:17.198657Z","shell.execute_reply.started":"2024-03-17T11:38:17.182349Z","shell.execute_reply":"2024-03-17T11:38:17.197198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Т.к. признаков довольно много (214), эвристически отсеем \"плохие\" признаки, значение которых \"слабо связано\" со значением целевой переменной.\n\nДля каждого категориального признака оставим значения, которые встречаются хотя бы в 1% строк.\nДля каждого этого значения вычислим среднее значение целевой переменной (т.е. долю значений, равных 1) и модуль разности этого значения с 1/2.\nВ качестве характеристики \"полезности\" признака возьмём максимальное значение полученной величины.","metadata":{"papermill":{"duration":0.016301,"end_time":"2024-03-16T22:34:10.961765","exception":false,"start_time":"2024-03-16T22:34:10.945464","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def calc_cat_features_usefulness(data: pl.DataFrame, value_fraction_threshold: float = 0.01) -> Dict[str, float]:\n    features_usefulness = {}\n    for feature_name in tqdm(data.columns):\n        if feature_name == 'target':\n            continue\n        feature_stats = data[[feature_name, 'target']].group_by(feature_name).agg(\n            pl.mean('target'),\n            pl.len() / len(train_X_with_static),\n        )\n        feature_stats = feature_stats.filter(pl.col('len') >= value_fraction_threshold)\n        feature_usefulness = feature_stats.select(pl.max('target')).item()\n        features_usefulness[feature_name] = feature_usefulness\n    return features_usefulness\n\n\ndef get_useful_cat_features(data: pl.DataFrame,\n                            target: List[int],\n                            value_fraction_threshold: float = 0.01,\n                            selected_features_fraction: float = 0.5) -> List[str]:\n    data = data.with_columns(pl.Series(target).alias('target'))\n    \n    cat_features_usefulness = calc_cat_features_usefulness(data, value_fraction_threshold)\n    cat_features, usefulness = zip(*reversed(sorted(list(cat_features_usefulness.items()), key=lambda p: p[1])))\n    selected_features_cnt = int(np.ceil(selected_features_fraction * len(cat_features)))\n    selected_features = cat_features[:selected_features_cnt]\n    print(f'Selected {len(selected_features)}/{len(cat_features)} categorical features based on their heuristic usefulness')\n    return list(selected_features)","metadata":{"papermill":{"duration":0.030003,"end_time":"2024-03-16T22:34:11.072770","exception":false,"start_time":"2024-03-16T22:34:11.042767","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.201223Z","iopub.execute_input":"2024-03-17T11:38:17.202291Z","iopub.status.idle":"2024-03-17T11:38:17.217191Z","shell.execute_reply.started":"2024-03-17T11:38:17.202245Z","shell.execute_reply":"2024-03-17T11:38:17.216327Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if not NOTEBOOK_RUN_CONFIG.use_fixed_config:\n    selected_cat_features = get_useful_cat_features(\n        train_X_with_static[cat_features], train_y\n    )","metadata":{"papermill":{"duration":0.023618,"end_time":"2024-03-16T22:34:11.110875","exception":false,"start_time":"2024-03-16T22:34:11.087257","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.218457Z","iopub.execute_input":"2024-03-17T11:38:17.219358Z","iopub.status.idle":"2024-03-17T11:38:17.235067Z","shell.execute_reply.started":"2024-03-17T11:38:17.219317Z","shell.execute_reply":"2024-03-17T11:38:17.234195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Для отбора численных признаков обучим линейную модель на всех (лог. регрессия) из них и отбросим те, при которых коэф-ты минимальны.\n\nПри этом, null заменим на нули.","metadata":{"papermill":{"duration":0.014463,"end_time":"2024-03-16T22:34:11.140309","exception":false,"start_time":"2024-03-16T22:34:11.125846","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def get_useful_numerical_features(data: pl.DataFrame,\n                                  target: List[int],\n                                  selected_features_fraction: float = 0.5) -> List[str]:\n    X = np.nan_to_num(data.to_numpy())\n    X = preprocessing.StandardScaler().fit_transform(X)\n    \n    log_reg = LogisticRegression(verbose=1, max_iter=10_000)\n    log_reg.fit(X, target)\n    \n    coef = log_reg.coef_[0]\n    features = data.columns\n    \n    coef, features = zip(*reversed(sorted(zip(np.abs(coef), features))))\n    selected_features_cnt = int(np.ceil(selected_features_fraction * len(features)))\n    selected_features = features[:selected_features_cnt]\n    \n    time.sleep(1)  # for correct output\n    \n    print(f'Selected {len(selected_features)}/{len(features)} numerical features based on their heuristic usefulness')\n    return list(selected_features)","metadata":{"papermill":{"duration":0.027374,"end_time":"2024-03-16T22:34:11.182922","exception":false,"start_time":"2024-03-16T22:34:11.155548","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.236667Z","iopub.execute_input":"2024-03-17T11:38:17.237370Z","iopub.status.idle":"2024-03-17T11:38:17.250219Z","shell.execute_reply.started":"2024-03-17T11:38:17.237330Z","shell.execute_reply":"2024-03-17T11:38:17.249241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if not NOTEBOOK_RUN_CONFIG.use_fixed_config:\n    selected_num_features = get_useful_numerical_features(\n        train_X_with_static[numerical_features], train_y\n    )","metadata":{"papermill":{"duration":0.023025,"end_time":"2024-03-16T22:34:11.220505","exception":false,"start_time":"2024-03-16T22:34:11.197480","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.251836Z","iopub.execute_input":"2024-03-17T11:38:17.252662Z","iopub.status.idle":"2024-03-17T11:38:17.264591Z","shell.execute_reply.started":"2024-03-17T11:38:17.252619Z","shell.execute_reply":"2024-03-17T11:38:17.263223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if not NOTEBOOK_RUN_CONFIG.use_fixed_config:\n    # save selected features in order to avoid recalculating\n    json.dump(selected_cat_features, open('selected_cat_features.json', 'w'))\n    json.dump(selected_num_features, open('selected_num_features.json', 'w'))\n\n# selected_cat_features = json.load(open('selected_cat_features.json', 'r'))\n# selected_num_features = json.load(open('selected_num_features.json', 'r'))","metadata":{"papermill":{"duration":0.023535,"end_time":"2024-03-16T22:34:11.258473","exception":false,"start_time":"2024-03-16T22:34:11.234938","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.267100Z","iopub.execute_input":"2024-03-17T11:38:17.268155Z","iopub.status.idle":"2024-03-17T11:38:17.280043Z","shell.execute_reply.started":"2024-03-17T11:38:17.267901Z","shell.execute_reply":"2024-03-17T11:38:17.279031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Полученные \"потенциально хорошие\" признаки дополнительно просеем с помощью алгоритма [recursive feature elimination](https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.RFE.html) над бустингом (до этого мы независимо фильтровали численные и категориальные признаки, а теперь будем учитывать возможные зависимости между ними).","metadata":{"papermill":{"duration":0.014128,"end_time":"2024-03-16T22:34:11.287108","exception":false,"start_time":"2024-03-16T22:34:11.272980","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def prepare_data_for_catboost(data: pl.DataFrame, cat_features: Sequence[str]) -> pd.DataFrame:\n    data = pl.concat([data[cat_features].cast(pl.String).fill_null('nan'),\n                      data.drop(cat_features).fill_null(np.nan)],\n                      how='horizontal')\n    return data.to_pandas()","metadata":{"papermill":{"duration":0.023791,"end_time":"2024-03-16T22:34:11.325428","exception":false,"start_time":"2024-03-16T22:34:11.301637","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.281550Z","iopub.execute_input":"2024-03-17T11:38:17.282671Z","iopub.status.idle":"2024-03-17T11:38:17.290274Z","shell.execute_reply.started":"2024-03-17T11:38:17.282632Z","shell.execute_reply":"2024-03-17T11:38:17.289046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if not NOTEBOOK_RUN_CONFIG.use_fixed_config:\n    # train_X_with_static = add_all_static_features(train_X, PARQUET_DATA_DIRPATH / 'train')\n    val_X_with_static = add_all_static_features(val_X, PARQUET_DATA_DIRPATH / 'train')\n\n    selection_train_X = prepare_data_for_catboost(train_X_with_static[selected_cat_features + selected_num_features],\n                                                  cat_features=selected_cat_features)\n    selection_val_X = prepare_data_for_catboost(val_X_with_static[selected_cat_features + selected_num_features],\n                                                cat_features=selected_cat_features)\n\n    selection_train_pool = catboost.Pool(selection_train_X, train_y, cat_features=selected_cat_features)\n    selection_val_pool = catboost.Pool(selection_val_X, val_y, cat_features=selected_cat_features)","metadata":{"papermill":{"duration":0.025376,"end_time":"2024-03-16T22:34:11.365722","exception":false,"start_time":"2024-03-16T22:34:11.340346","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.291731Z","iopub.execute_input":"2024-03-17T11:38:17.292121Z","iopub.status.idle":"2024-03-17T11:38:17.304813Z","shell.execute_reply.started":"2024-03-17T11:38:17.292091Z","shell.execute_reply":"2024-03-17T11:38:17.303711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if not NOTEBOOK_RUN_CONFIG.use_fixed_config:\n    model = CatBoostClassifier(iterations=200, random_seed=42)\n\n    summary = model.select_features(\n        selection_train_pool,\n        eval_set=selection_val_pool,\n        features_for_select=selection_train_X.columns,\n        num_features_to_select=10,\n        steps=10,\n        algorithm=catboost.EFeaturesSelectionAlgorithm.RecursiveByShapValues,\n        shap_calc_type=catboost.EShapCalcType.Regular,\n        train_final_model=False,\n        logging_level='Verbose',\n        plot=True,\n    )","metadata":{"papermill":{"duration":0.025042,"end_time":"2024-03-16T22:34:11.405870","exception":false,"start_time":"2024-03-16T22:34:11.380828","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.306191Z","iopub.execute_input":"2024-03-17T11:38:17.307277Z","iopub.status.idle":"2024-03-17T11:38:17.323733Z","shell.execute_reply.started":"2024-03-17T11:38:17.307236Z","shell.execute_reply":"2024-03-17T11:38:17.322586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if not NOTEBOOK_RUN_CONFIG.use_fixed_config:\n    fig, ax = plt.subplots()\n\n    ax.plot(summary['loss_graph']['removed_features_count'], summary['loss_graph']['loss_values'])\n    ax.set_xlabel('# of eliminated features')\n    ax.set_ylabel('test loss')\n\n    plt.show()","metadata":{"papermill":{"duration":0.024183,"end_time":"2024-03-16T22:34:11.444841","exception":false,"start_time":"2024-03-16T22:34:11.420658","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.325359Z","iopub.execute_input":"2024-03-17T11:38:17.326518Z","iopub.status.idle":"2024-03-17T11:38:17.335420Z","shell.execute_reply.started":"2024-03-17T11:38:17.326476Z","shell.execute_reply":"2024-03-17T11:38:17.334041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if not NOTEBOOK_RUN_CONFIG.use_fixed_config:\n    removed_features_cnt = summary['loss_graph']['removed_features_count'][np.argmin(summary['loss_graph']['loss_values'])]\n    removed_features = summary['eliminated_features_names'][:removed_features_cnt]\n    print(f'{removed_features_cnt}/{len(selected_cat_features) + len(selected_num_features)} features removed')","metadata":{"papermill":{"duration":0.026517,"end_time":"2024-03-16T22:34:11.486630","exception":false,"start_time":"2024-03-16T22:34:11.460113","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.337205Z","iopub.execute_input":"2024-03-17T11:38:17.337645Z","iopub.status.idle":"2024-03-17T11:38:17.345456Z","shell.execute_reply.started":"2024-03-17T11:38:17.337601Z","shell.execute_reply":"2024-03-17T11:38:17.344094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if not NOTEBOOK_RUN_CONFIG.use_fixed_config:\n    final_selected_cat_features = list(set(selected_cat_features).difference(removed_features))\n    final_selected_num_features = list(set(selected_num_features).difference(removed_features))\n    print(f'Left {len(final_selected_cat_features)}/{len(selected_cat_features)} categorical features')\n    print(f'Left {len(final_selected_num_features)}/{len(selected_num_features)} numerical features')\n    \n    print('\\nSelected categorical features:\\n', final_selected_cat_features)\n    print('\\nSelected numerical features:\\n', final_selected_num_features)","metadata":{"papermill":{"duration":0.024664,"end_time":"2024-03-16T22:34:11.525865","exception":false,"start_time":"2024-03-16T22:34:11.501201","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.347355Z","iopub.execute_input":"2024-03-17T11:38:17.348630Z","iopub.status.idle":"2024-03-17T11:38:17.358595Z","shell.execute_reply.started":"2024-03-17T11:38:17.348582Z","shell.execute_reply":"2024-03-17T11:38:17.357231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if NOTEBOOK_RUN_CONFIG.use_fixed_config:\n    # т.к. вычисления довольно долгие, я вставил сюда результат для применения на этапе предсказания\n    final_selected_cat_features = ['maritalst_385M', 'datefirstoffer_1144D', 'isbidproduct_1095L', 'education_1103M', 'lastrejectreason_759M', 'requesttype_4525192L']\n    final_selected_num_features = ['riskassesment_940T', 'numinstlswithdpd10_728L', 'mobilephncnt_593L', 'days360_512L', 'thirdquarter_1082L', 'pmtnum_254L', 'cntincpaycont9m_3716944L', 'numincomingpmts_3546848L', 'days120_123L', 'pmtaverage_4955615A', 'numinsttopaygr_769L', 'homephncnt_628L', 'numinstpaidearly3d_3546850L', 'numberofqueries_373L', 'days30_165L', 'pctinstlsallpaidlat10d_839L', 'fourthquarter_440L', 'price_1097A', 'avgdpdtolclosure24_3658938P', 'currdebt_22A', 'maxdpdtolerance_374P', 'monthsannuity_845L', 'maxdebt4_972A', 'numrejects9m_859L', 'maxdbddpdlast1m_3658939P', 'amtinstpaidbefduel24m_4187115A', 'maxdpdlast12m_727P']","metadata":{"papermill":{"duration":0.024931,"end_time":"2024-03-16T22:34:11.566742","exception":false,"start_time":"2024-03-16T22:34:11.541811","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.360463Z","iopub.execute_input":"2024-03-17T11:38:17.361398Z","iopub.status.idle":"2024-03-17T11:38:17.374282Z","shell.execute_reply.started":"2024-03-17T11:38:17.361353Z","shell.execute_reply":"2024-03-17T11:38:17.373204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Подбор гиперпараметров","metadata":{"papermill":{"duration":0.014331,"end_time":"2024-03-16T22:34:11.595774","exception":false,"start_time":"2024-03-16T22:34:11.581443","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"Будем подбирать гиперпараметры, максимизируя целевую метрику с помощью Optuna.","metadata":{"papermill":{"duration":0.014131,"end_time":"2024-03-16T22:34:11.624613","exception":false,"start_time":"2024-03-16T22:34:11.610482","status":"completed"},"tags":[]}},{"cell_type":"code","source":"if not NOTEBOOK_RUN_CONFIG.use_fixed_config:\n    hparams_train_X = prepare_data_for_catboost(train_X_with_static[final_selected_cat_features + final_selected_num_features],\n                                                cat_features=final_selected_cat_features)\n    hparams_val_X = prepare_data_for_catboost(val_X_with_static[final_selected_cat_features + final_selected_num_features],\n                                              cat_features=final_selected_cat_features)\n\n    hparams_train_pool = catboost.Pool(hparams_train_X, train_y, cat_features=final_selected_cat_features)\n    hparams_val_pool = catboost.Pool(hparams_val_X, val_y, cat_features=final_selected_cat_features)","metadata":{"papermill":{"duration":0.02479,"end_time":"2024-03-16T22:34:11.663861","exception":false,"start_time":"2024-03-16T22:34:11.639071","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.375956Z","iopub.execute_input":"2024-03-17T11:38:17.376855Z","iopub.status.idle":"2024-03-17T11:38:17.385244Z","shell.execute_reply.started":"2024-03-17T11:38:17.376821Z","shell.execute_reply":"2024-03-17T11:38:17.384251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def objective(trial):\n    kwargs = {\n        'objective': trial.suggest_categorical('objective', ['Logloss', 'CrossEntropy']),\n        'iterations': trial.suggest_int('iterations', 100, 1000, step=100),\n        'learning_rate': trial.suggest_float('learning_rate', 1e-3, 1, log=True),\n        'depth': trial.suggest_int('depth', 1, 10),\n        'subsample': trial.suggest_float('subsample', 0.01, 0.3),\n    }\n    \n    model = CatBoostClassifier(random_seed=42, **kwargs)\n    model.fit(hparams_train_pool, eval_set=hparams_val_pool)\n    pred_val_y_scores = model.predict(hparams_val_pool, prediction_type='Probability')[:, 1]\n    score = eval_predicted_score(val_X, val_y, pred_val_y_scores, plot=False)\n    return score","metadata":{"papermill":{"duration":0.027351,"end_time":"2024-03-16T22:34:11.705930","exception":false,"start_time":"2024-03-16T22:34:11.678579","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.386541Z","iopub.execute_input":"2024-03-17T11:38:17.387289Z","iopub.status.idle":"2024-03-17T11:38:17.398451Z","shell.execute_reply.started":"2024-03-17T11:38:17.387249Z","shell.execute_reply":"2024-03-17T11:38:17.397572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if not NOTEBOOK_RUN_CONFIG.use_fixed_config:\n    study = optuna.create_study(direction='maximize')\n    study.optimize(objective, n_trials=100)","metadata":{"papermill":{"duration":0.023977,"end_time":"2024-03-16T22:34:11.744598","exception":false,"start_time":"2024-03-16T22:34:11.720621","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.404678Z","iopub.execute_input":"2024-03-17T11:38:17.405450Z","iopub.status.idle":"2024-03-17T11:38:17.411373Z","shell.execute_reply.started":"2024-03-17T11:38:17.405406Z","shell.execute_reply":"2024-03-17T11:38:17.410213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if not NOTEBOOK_RUN_CONFIG.use_fixed_config:\n    print(f\"Best metric: {study.best_value} with the following  params\")\n    print(study.best_params)\n    best_model_params = study.best_params","metadata":{"papermill":{"duration":0.023389,"end_time":"2024-03-16T22:34:11.782989","exception":false,"start_time":"2024-03-16T22:34:11.759600","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.412620Z","iopub.execute_input":"2024-03-17T11:38:17.413153Z","iopub.status.idle":"2024-03-17T11:38:17.421990Z","shell.execute_reply.started":"2024-03-17T11:38:17.413109Z","shell.execute_reply":"2024-03-17T11:38:17.420628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if NOTEBOOK_RUN_CONFIG.use_fixed_config:\n    best_model_params = {\n        'objective': 'CrossEntropy',\n        'iterations': 600,\n        'learning_rate': 0.08,\n        'depth': 7,\n        'subsample': 0.255,\n    }","metadata":{"papermill":{"duration":0.023344,"end_time":"2024-03-16T22:34:11.820688","exception":false,"start_time":"2024-03-16T22:34:11.797344","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.423638Z","iopub.execute_input":"2024-03-17T11:38:17.424315Z","iopub.status.idle":"2024-03-17T11:38:17.433476Z","shell.execute_reply.started":"2024-03-17T11:38:17.424273Z","shell.execute_reply":"2024-03-17T11:38:17.432242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Обучение итоговой модели","metadata":{"papermill":{"duration":0.014337,"end_time":"2024-03-16T22:34:11.849312","exception":false,"start_time":"2024-03-16T22:34:11.834975","status":"completed"},"tags":[]}},{"cell_type":"code","source":"class FeaturesExtractor(BaseEstimator, TransformerMixin):\n    def fit(self, X: pl.DataFrame, tables_dirpath: Path):\n        return self\n    \n    def transform(self, X: pl.DataFrame, tables_dirpath: Path) -> pl.DataFrame:\n        \"\"\"\n        X must contain `case_id`\n        \n        :returns: X, categorical features\n        \"\"\"\n        X = add_all_static_features(X, tables_dirpath)\n        X = X[final_selected_cat_features + final_selected_num_features]\n        return X\n    \n    def get_categorical_features(self) -> List[str]:\n        return final_selected_cat_features","metadata":{"papermill":{"duration":0.024855,"end_time":"2024-03-16T22:34:11.888509","exception":false,"start_time":"2024-03-16T22:34:11.863654","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.435332Z","iopub.execute_input":"2024-03-17T11:38:17.436064Z","iopub.status.idle":"2024-03-17T11:38:17.445085Z","shell.execute_reply.started":"2024-03-17T11:38:17.436020Z","shell.execute_reply":"2024-03-17T11:38:17.444084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Теперь, после отбора признаков и подбора гиперпараметров, можем обучить итоговую модель на всей обучающей выборке и сделать предсказание для тестовой.","metadata":{"papermill":{"duration":0.014318,"end_time":"2024-03-16T22:34:11.918200","exception":false,"start_time":"2024-03-16T22:34:11.903882","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# сборка датасетов\ntrain_base_df = pl.read_parquet(PARQUET_DATA_DIRPATH / 'train' / 'train_base.parquet')\ntrain_X, train_y = train_base_df.drop('target'), train_base_df['target'].to_list()\ntest_base_df = pl.read_parquet(PARQUET_DATA_DIRPATH / 'test' / 'test_base.parquet')\ntest_X = test_base_df.clone()\n\nfeatures_extractor = FeaturesExtractor().fit(train_X, tables_dirpath=PARQUET_DATA_DIRPATH / 'train')\ntrain_X = features_extractor.transform(train_X, tables_dirpath=PARQUET_DATA_DIRPATH / 'train')\ntest_X = features_extractor.transform(test_X, tables_dirpath=PARQUET_DATA_DIRPATH / 'test')\ncat_features = features_extractor.get_categorical_features()\n\ntrain_X = prepare_data_for_catboost(train_X, cat_features)\ntest_X = prepare_data_for_catboost(test_X, cat_features)\n\ntrain_pool = catboost.Pool(train_X, train_y, cat_features=cat_features)\ntest_pool = catboost.Pool(test_X, cat_features=cat_features)","metadata":{"papermill":{"duration":12.619498,"end_time":"2024-03-16T22:34:24.552263","exception":false,"start_time":"2024-03-16T22:34:11.932765","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:17.446760Z","iopub.execute_input":"2024-03-17T11:38:17.447709Z","iopub.status.idle":"2024-03-17T11:38:29.962830Z","shell.execute_reply.started":"2024-03-17T11:38:17.447602Z","shell.execute_reply":"2024-03-17T11:38:29.961809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = CatBoostClassifier(random_seed=42, **best_model_params)\nmodel.fit(train_pool)\npred_test_y_scores = model.predict(test_pool, prediction_type='Probability')[:, 1]","metadata":{"papermill":{"duration":747.911508,"end_time":"2024-03-16T22:46:52.478500","exception":false,"start_time":"2024-03-16T22:34:24.566992","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:29.964339Z","iopub.execute_input":"2024-03-17T11:38:29.964715Z","iopub.status.idle":"2024-03-17T11:38:42.277306Z","shell.execute_reply.started":"2024-03-17T11:38:29.964683Z","shell.execute_reply":"2024-03-17T11:38:42.274561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions_df = test_base_df[['case_id']].with_columns(pl.Series(pred_test_y_scores).alias('score'))\npredictions_df.write_csv('submission.csv')","metadata":{"papermill":{"duration":0.104112,"end_time":"2024-03-16T22:46:52.671571","exception":false,"start_time":"2024-03-16T22:46:52.567459","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-03-17T11:38:42.278177Z","iopub.status.idle":"2024-03-17T11:38:42.278601Z","shell.execute_reply.started":"2024-03-17T11:38:42.278402Z","shell.execute_reply":"2024-03-17T11:38:42.278419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"papermill":{"duration":0.086959,"end_time":"2024-03-16T22:46:52.845644","exception":false,"start_time":"2024-03-16T22:46:52.758685","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]}]}