{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":96164,"databundleVersionId":11418275,"sourceType":"competition"}],"dockerImageVersionId":31040,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 1. Description\n\n**work in progress**\n\nThis notebook demonstrates the importance of selecting the appropriate cross-validation (**CV**) strategy when working with time-series data. A poor choice can result in misleading validation scores and models that perform well in testing but fail catastrophically in production. I will compare three CV approaches using DRW-Crypto competiion data:\n\n**1. K-Fold CV without Random Shuffling** \n- maintains original data order during splitting.\n- important when data has inherent structure or temporal components.\n- prevents information leakage from shuffling related observations.\n- ⚠️ **Still problematic for time series** - can train on future data to predict the past.\n- default behavior in `sklearn.model_selection.KFold` (**[documentation](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.KFold.html)**).\n\n**2. K-Fold CV with random shuffling**\n- random splitting of data into k folds.\n- assumes data points are independent and identically distributed.\n\n**3. Walk-forward validation (expanding window)**\n- respects temporal order by training only on past data.\n- progressive expansion of training set with fixed test windows.\n- essential for time series and sequential data problems.\n\n**4. Standard walk-forward validation**\n- same as above just both training and test windows slide forward through time.\n\n**5. Out-of-distribution (era-splitting)**\n- averages the impurity reduction across different eras (time periods/environments) instead of pooling all data together.\n\n## observations\n\n* LB using shuffled and un-shuffuladed data is almost identicall (0.09852 vs 0.09819)\n* Pearson correlation scores on validation samples vs tran are >2 worse whe suffling is not used.\n\n## experimental design\n\n1. **Consistent model setup**: all CV methods use identical hyperparameters and feature sets taken from https://www.kaggle.com/code/sadettinamilverdil/yat-r-m-tavsiyesi-de-ildir notebook by `sadettinamilverdil`.\n2. **Fixed random seed**: same seed ensures any performance differences stem from CV strategy, not random variation.\n3. **Standardized evaluation**: identical metrics (Pearson correlation) applied across all validation approaches.\n4. **\"Real-world\" testing**: public leaderboard score (LB) is used to analyze CV findings against unseen data.\n\n## changelog\n\n* **v1** build default outline of the notebook and added visualization for `K-fold` split without suffle.\n* **v2** added baseline `XGBoost` model. Train parameters were used on un-suffluled sample. LB score of 0.09819 was achieved.\n* **v3** added cross-validation analysis on different sample splits.\n* **v4** trained model on suffluled K-fold sample. LB score of 0.09852 was achieved.\n* **v5** enhanced documentation and methodology descriptions.\n* **v6** trained model on walk-Forward Validation (expanding window) sample. LB score of 0.07878 was achieved.\n* **v7** trained model on standard walk-Forward validation sample. LB score of 0.08845 was achieved.\n* **v8** changed XGBoost model to `WarpGBM` as https://www.kaggle.com/code/jefferythewind/warpgbm-invariant-example#Naive-GBM inspired by notebook. LB score of 0.08241 was achieved.\n* **v9** trained `WarpGBM` model on suffluled K-fold sample. LB score of 0.08022 was achieved.\n* **v10** added out-of-time sample to all validation methods.\n* **v11** added era-aware GBM model.\n* **v12** updated notebook text.","metadata":{}},{"cell_type":"code","source":"# data processing libraries\nimport numpy as np\nimport pandas as pd\n\nfrom datetime import datetime\nimport os\n\n# for monitoring progress\nfrom tqdm import tqdm\n\n# garbage collect\nimport gc\n\nimport seaborn as sns # plots for statistical analysis\nimport matplotlib.pyplot as plt # for data visualization\n\n# for spliting data into train and test\nfrom sklearn.model_selection import KFold\n\n# define default colors for plots in notebook\nfrom matplotlib import cycler\nfrom matplotlib.colors import LinearSegmentedColormap\ncolors = [\"#068D9D\", \"#53599A\", \"#607BB0\", \"#6D9DC5\", \"#77BECF\", \"#80DED9\", \"#AEECEF\"]\n\nplt.rc('axes', facecolor='#E6E6E6', edgecolor='none', axisbelow=True, grid=True, prop_cycle=cycler('color', colors))\n\n# for modeling\nfrom xgboost import XGBRegressor\nfrom scipy.stats import pearsonr\n\nSEED = 42","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-06-20T13:12:46.094636Z","iopub.execute_input":"2025-06-20T13:12:46.095139Z","iopub.status.idle":"2025-06-20T13:12:47.471688Z","shell.execute_reply.started":"2025-06-20T13:12:46.095116Z","shell.execute_reply":"2025-06-20T13:12:47.471095Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from IPython.display import clear_output\n\n# Upgrade Torch to 2.6.0+CUDA 12.4\n!pip install --upgrade torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu124\n\n# Confirm torch version\nimport torch\nprint(\"Torch version:\", torch.__version__)\nprint(\"Torch CUDA version:\", torch.version.cuda)\n\nimport torch\nprint(torch.__version__)\nprint(torch.version.cuda)\n\n!pip install warpgbm --no-build-isolation\n\nfrom sklearn.model_selection import TimeSeriesSplit\nfrom warpgbm import WarpGBM\n\nclear_output()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-20T13:12:47.472758Z","iopub.execute_input":"2025-06-20T13:12:47.473088Z","iopub.status.idle":"2025-06-20T13:16:00.953053Z","shell.execute_reply.started":"2025-06-20T13:12:47.473071Z","shell.execute_reply":"2025-06-20T13:16:00.952240Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 2. Load data","metadata":{}},{"cell_type":"code","source":"def reduce_mem_usage(dataframe, dataset):\n    \"\"\"\n    Function taken from: https://www.kaggle.com/code/ravaghi/drw-crypto-market-prediction-ensemble\n    \"\"\"\n    print('Reducing memory usage for:', dataset)\n    initial_mem_usage = dataframe.memory_usage().sum() / 1024**2\n    \n    for col in dataframe.columns:\n        col_type = dataframe[col].dtype\n\n        c_min = dataframe[col].min()\n        c_max = dataframe[col].max()\n        if str(col_type)[:3] == 'int':\n            if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                dataframe[col] = dataframe[col].astype(np.int8)\n            elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                dataframe[col] = dataframe[col].astype(np.int16)\n            elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                dataframe[col] = dataframe[col].astype(np.int32)\n            elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                dataframe[col] = dataframe[col].astype(np.int64)\n        else:\n            if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                dataframe[col] = dataframe[col].astype(np.float16)\n            elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                dataframe[col] = dataframe[col].astype(np.float32)\n            else:\n                dataframe[col] = dataframe[col].astype(np.float64)\n\n    final_mem_usage = dataframe.memory_usage().sum() / 1024**2\n    print('--- Memory usage before: {:.2f} MB'.format(initial_mem_usage))\n    print('--- Memory usage after: {:.2f} MB'.format(final_mem_usage))\n    print('--- Decreased memory usage by {:.1f}%\\n'.format(100 * (initial_mem_usage - final_mem_usage) / initial_mem_usage))\n\n    return dataframe ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-20T13:16:00.954100Z","iopub.execute_input":"2025-06-20T13:16:00.954846Z","iopub.status.idle":"2025-06-20T13:16:00.963582Z","shell.execute_reply.started":"2025-06-20T13:16:00.954814Z","shell.execute_reply":"2025-06-20T13:16:00.962893Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time\ndf_train = pd.read_parquet('/kaggle/input/drw-crypto-market-prediction/train.parquet')\ndf_test = pd.read_parquet('/kaggle/input/drw-crypto-market-prediction/test.parquet')\n\ndf_train = reduce_mem_usage(df_train, \"train\")\ndf_test = reduce_mem_usage(df_test, \"test\")\n\ndf_train = df_train.reset_index()\n\nprint(f\"Train dataset contains {df_train.shape[0]} rows and {df_train.shape[1]} columns.\" )\nprint(f\"Test dataset contains {df_test.shape[0]} rows and {df_test.shape[1]} columns.\" )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-20T13:16:00.965034Z","iopub.execute_input":"2025-06-20T13:16:00.965240Z","iopub.status.idle":"2025-06-20T13:16:58.672132Z","shell.execute_reply.started":"2025-06-20T13:16:00.965224Z","shell.execute_reply":"2025-06-20T13:16:58.671219Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time\n# Find and remove constant features (zero variance)\nconstant_features = []\nfor col in df_train.select_dtypes(include=[np.number]).columns:\n    if df_train[col].nunique(dropna=False) == 1:\n        constant_features.append(col)\n\nif constant_features:\n    print(f\"❌ Found {len(constant_features)} constant features (column 1 unique value).\")\n    df_train = df_train.drop(constant_features, axis=1)\n    test_constant_features = [col for col in constant_features if col in df_test.columns]\n    if test_constant_features:\n        df_test = df_test.drop(test_constant_features, axis=1)\n    print(f\"✅ Removed constant features. New shape: {df_train.shape}\")\nelse:\n    print(\"✅ No constant features found\")\n\ndel constant_features\ngc.collect()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-20T13:16:58.673026Z","iopub.execute_input":"2025-06-20T13:16:58.673282Z","iopub.status.idle":"2025-06-20T13:17:08.598437Z","shell.execute_reply.started":"2025-06-20T13:16:58.673261Z","shell.execute_reply":"2025-06-20T13:17:08.597676Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"selected_features = [\n    \"X863\", \"X856\", \"X344\", \"X598\", \"X862\", \"X385\", \"X852\", \"X603\", \"X860\", \"X674\",\n    \"X415\", \"X345\", \"X137\", \"X855\", \"X174\", \"X302\", \"X178\", \"X532\", \"X168\", \"X612\",\n    \"bid_qty\", \"ask_qty\", \"buy_qty\", \"sell_qty\", \"volume\"\n]\n\ndf_train = df_train[selected_features + [\"timestamp\", \"label\"]]\n# split into validation and train datasets\nidx = int(df_train.shape[0] * 0.8)\ndf_valid = df_train.iloc[idx:]\ndf_train = df_train.iloc[:idx]\ndf_test = df_test[selected_features]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-20T13:17:08.599145Z","iopub.execute_input":"2025-06-20T13:17:08.599353Z","iopub.status.idle":"2025-06-20T13:17:08.834244Z","shell.execute_reply.started":"2025-06-20T13:17:08.599337Z","shell.execute_reply":"2025-06-20T13:17:08.833646Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 3. K-fold validation | without suffle\n\n## 3.1. Visualize\n\nUse cumulative sum of `label` to visualize different market areas which are used for training and testing.","metadata":{}},{"cell_type":"code","source":"\nn_splits = 5\nn_cols = 2\n\ndf_dummy = df_train[[\"timestamp\", 'label']].copy()\ndf_dummy['value'] = df_dummy['label'].cumsum()\n\n# compare split results\nfig, ax = plt.subplots((n_splits+1) // n_cols, n_cols, figsize=(18, 6), sharex=True)\nax = ax.flatten()\n\n# all data\nax[0].plot(df_dummy['timestamp'], df_dummy['value'], label=\"all data\", color=colors[2])\nax[0].set_ylabel(\"cumsum of label\")\nax[0].set_title(\"All data\")\n\nkf = KFold(n_splits=n_splits)\nkf.get_n_splits(df_dummy)\n\n# add validation sections\nfor i in range(6):\n    ax[i].plot(df_valid['timestamp'], df_valid['label'].cumsum()+df_dummy['value'].iloc[-1], label=\"valid\", color='r')\n\nfor i, (train_idx, test_idx) in enumerate(kf.split(df_dummy)):\n    _df_train = df_dummy.loc[train_idx]\n    _df_test = df_dummy.loc[test_idx]\n    \n    if i == 0 or i == (n_splits - 1):\n        ax[i + 1].plot(_df_train['timestamp'], _df_train['value'], label=\"train\")\n    else:\n        _split_idx = (_df_train[\"timestamp\"] - _df_train[\"timestamp\"].shift(-1)).argmin() + 1\n        ax[i + 1].plot(_df_train.loc[:_split_idx, 'timestamp'], _df_train.loc[:_split_idx, 'value'], label=\"train\")\n        ax[i + 1].plot(_df_train.loc[_split_idx:, 'timestamp'], _df_train.loc[_split_idx:, 'value'], color=colors[0])\n    ax[i + 1].plot(_df_test['timestamp'], _df_test['value'], label=\"test\")\n    \n    ax[i + 1].set_ylabel(\"cumsum of label\")\n    ax[i + 1].set_title(f\"split {i+1}\")\n\n    # add white background to legend\n    legend = ax[i + 1].legend(frameon=1)\n    frame = legend.get_frame()\n    frame.set_facecolor('w')\n\nif n_cols == 2:\n    ax[-2].set_xlabel(\"timestamp\")\nax[-1].set_xlabel(\"timestamp\")\n\nplt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:36:35.400130Z","iopub.execute_input":"2025-06-19T13:36:35.400354Z","iopub.status.idle":"2025-06-19T13:36:40.405207Z","shell.execute_reply.started":"2025-06-19T13:36:35.400337Z","shell.execute_reply":"2025-06-19T13:36:40.404426Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3.2. Train model\n\n### XGBoost version\n\nParams taken from https://www.kaggle.com/code/sadettinamilverdil/yat-r-m-tavsiyesi-de-ildir notebook by `sadettinamilverdil`.","metadata":{}},{"cell_type":"code","source":"\nxgb_params = {\n    \"tree_method\": \"gpu_hist\",\n    \"colsample_bylevel\": 0.4778015829774066,\n    \"colsample_bynode\": 0.362764358742407,\n    \"colsample_bytree\": 0.7107423488010493,\n    \"gamma\": 1.7094857725240398,\n    \"learning_rate\": 0.02213323588455387,\n    \"max_depth\": 20,\n    \"max_leaves\": 12,\n    \"min_child_weight\": 16,\n    \"n_estimators\": 1667,\n    \"n_jobs\": -1,\n    \"random_state\": 42,\n    \"reg_alpha\": 39.352415706891264,\n    \"reg_lambda\": 75.44843704068275,\n    \"subsample\": 0.06566669853471274,\n    \"verbosity\": 0\n}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:36:40.406047Z","iopub.execute_input":"2025-06-19T13:36:40.406310Z","iopub.status.idle":"2025-06-19T13:36:40.411478Z","shell.execute_reply.started":"2025-06-19T13:36:40.406290Z","shell.execute_reply":"2025-06-19T13:36:40.410786Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\noof_preds = np.zeros(len(df_train))\ntest_preds = np.zeros(len(df_test))\n\ndf_cv_stats = list()\n\nfor i, (train_idx, valid_idx) in enumerate(kf.split(df_train)):\n    print(\"#\" * 25)\n    print(f\"### Fold {i + 1} \" + \"#\" * 14)\n    print(\"#\" * 25)\n\n    X_train = df_train.iloc[train_idx][selected_features]\n    y_train = df_train.iloc[train_idx][\"label\"]\n    X_valid = df_train.iloc[valid_idx][selected_features]\n    y_valid = df_train.iloc[valid_idx][\"label\"]\n    X_test = df_test[selected_features]\n\n    model = XGBRegressor(**xgb_params)\n\n    model.fit(\n        X_train, y_train,\n        eval_set=[(X_valid, y_valid)],\n        early_stopping_rounds=25,\n        verbose=200\n    )\n\n    oof_preds[valid_idx] = model.predict(X_valid)\n    test_preds += model.predict(X_test)\n\n    # calculate errors on train and validation samples\n    _train_cv = pearsonr(y_train, model.predict(X_train))[0]\n    _valid_cv = pearsonr(y_valid, oof_preds[valid_idx])[0]\n    # add out-of-time (validation) sample predictions\n    _oot_cv = pearsonr(df_valid['label'], model.predict(df_valid[selected_features]))[0]\n    df_cv_stats.append([f\"split {i+1}\", _train_cv, _valid_cv, _oot_cv])\n\n    print(\"\")\n\npearson_score = pearsonr(df_train[\"label\"], oof_preds)[0]\nprint(\"Final Pearson Correlation = \", pearson_score)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:36:40.412316Z","iopub.execute_input":"2025-06-19T13:36:40.412516Z","iopub.status.idle":"2025-06-19T13:36:53.141623Z","shell.execute_reply.started":"2025-06-19T13:36:40.412493Z","shell.execute_reply":"2025-06-19T13:36:53.140817Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### WrapGBM version\n\nCode example taken from https://www.kaggle.com/code/jefferythewind/warpgbm-invariant-example#Naive-GBM","metadata":{}},{"cell_type":"code","source":"\n# oof_preds = np.zeros(len(df_train))\n# test_preds = np.zeros(len(df_test))\n\n# df_cv_stats = list()\n\n# for i, (train_idx, valid_idx) in enumerate(kf.split(df_train)):\n#     print(\"#\" * 25)\n#     print(f\"### Fold {i + 1} \" + \"#\" * 14)\n#     print(\"#\" * 25)\n\n#     X_train = df_train.iloc[train_idx][selected_features].values\n#     y_train = df_train.iloc[train_idx][\"label\"].values\n#     X_valid = df_train.iloc[valid_idx][selected_features].values\n#     y_valid = df_train.iloc[valid_idx][\"label\"].values\n#     X_test = df_test[selected_features].values\n\n#     model = WarpGBM(\n#         max_depth=10,\n#         num_bins=100,\n#         n_estimators=100,\n#         learning_rate=0.1,\n#         colsample_bytree=1.0,\n#         min_child_weight=4\n#     )\n#     model.fit(\n#         X_train,\n#         y_train,\n#         X_eval=X_valid,\n#         y_eval=y_valid,\n#         eval_every_n_trees=1,\n#         early_stopping_rounds=10,\n#         eval_metric=\"rmsle\",\n#     )\n\n#     #keep best model\n#     best_i = int(np.argmin(model.eval_loss))\n#     opt_num_trees = best_i * model.eval_every_n_trees  # +1 since first eval is at tree N, not 0\n#     model.forest = model.forest[:( opt_num_trees + 1)]\n\n#     oof_preds[valid_idx] = model.predict(X_valid)\n#     test_preds += model.predict(X_test)\n\n#     # calculate errors on train and validation samples\n#     _train_cv = pearsonr(y_train, model.predict(X_train))[0]\n#     _valid_cv = pearsonr(y_valid, oof_preds[valid_idx])[0]\n#     df_cv_stats.append([f\"split {i+1}\", _train_cv, _valid_cv])\n\n#     print(\"\")\n\n# pearson_score = pearsonr(df_train[\"label\"], oof_preds)[0]\n# print(\"Final Pearson Correlation = \", pearson_score)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:36:53.142558Z","iopub.execute_input":"2025-06-19T13:36:53.142872Z","iopub.status.idle":"2025-06-19T13:36:53.147597Z","shell.execute_reply.started":"2025-06-19T13:36:53.142854Z","shell.execute_reply":"2025-06-19T13:36:53.146797Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3.3. visualize validation scores","metadata":{}},{"cell_type":"code","source":"df_cv_stats = pd.DataFrame(df_cv_stats, columns=[\"fold\", 'Pearson train', \"Pearson valid\", \"Pearson out-of-time\"])\ndf_cv_stats['diff'] = df_cv_stats['Pearson train'] - df_cv_stats['Pearson valid']\ndf_cv_stats['ratio'] = df_cv_stats['Pearson train'] / df_cv_stats['Pearson valid']\ndf_cv_stats","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:36:53.148431Z","iopub.execute_input":"2025-06-19T13:36:53.148622Z","iopub.status.idle":"2025-06-19T13:36:53.180978Z","shell.execute_reply.started":"2025-06-19T13:36:53.148598Z","shell.execute_reply":"2025-06-19T13:36:53.180439Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\nfig, ax = plt.subplots(1, 1, figsize=(10, 4))\n\nsns.scatterplot(data=df_cv_stats, x=\"Pearson train\", y=\"Pearson valid\", hue=\"fold\", style=\"fold\", s=200, ax=ax)\n\n# add white background to legend\nlegend = ax.legend(frameon=1)\nframe = legend.get_frame()\nframe.set_facecolor('w')\n\nplt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:36:53.181727Z","iopub.execute_input":"2025-06-19T13:36:53.181981Z","iopub.status.idle":"2025-06-19T13:36:53.483156Z","shell.execute_reply.started":"2025-06-19T13:36:53.181959Z","shell.execute_reply":"2025-06-19T13:36:53.482302Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\nfig, ax = plt.subplots(1, 1, figsize=(10, 4))\n\n# out-of-time sample\nsns.boxplot(x=df_cv_stats[\"Pearson out-of-time\"])\n\nplt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:36:53.484357Z","iopub.execute_input":"2025-06-19T13:36:53.485019Z","iopub.status.idle":"2025-06-19T13:36:53.608541Z","shell.execute_reply.started":"2025-06-19T13:36:53.484991Z","shell.execute_reply":"2025-06-19T13:36:53.607963Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3.4. submit predictions","metadata":{}},{"cell_type":"code","source":"# sample=pd.read_csv(\"/kaggle/input/drw-crypto-market-prediction/sample_submission.csv\")\n# sample[\"prediction\"] = test_preds / n_splits\n# sample.to_csv(\"submission.csv\", index=False)\n# sample.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:36:53.611461Z","iopub.execute_input":"2025-06-19T13:36:53.611873Z","iopub.status.idle":"2025-06-19T13:36:53.615044Z","shell.execute_reply.started":"2025-06-19T13:36:53.611855Z","shell.execute_reply":"2025-06-19T13:36:53.614337Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 4. K-fold validation | with random suffle\n\n## 4.1. Visualize\n\nSame as in 3.1. section.","metadata":{}},{"cell_type":"code","source":"n_splits = 5\nn_cols = 2\n\n# compare split results\nfig, ax = plt.subplots((n_splits+1) // n_cols, n_cols, figsize=(18, 6), sharex=True)\nax = ax.flatten()\n\n# all data\nax[0].plot(df_dummy['timestamp'], df_dummy['value'], label=\"all data\", color=colors[2])\nax[0].set_ylabel(\"cumsum of label\")\nax[0].set_title(\"All data\")\n\nkf = KFold(n_splits=n_splits, shuffle=True, random_state=SEED)\nkf.get_n_splits(df_dummy)\n\n# add validation sections\nfor i in range(6):\n    ax[i].plot(df_valid['timestamp'], df_valid['label'].cumsum()+df_dummy['value'].iloc[-1], label=\"valid\", color='r')\n\nfor i, (train_idx, test_idx) in enumerate(kf.split(df_dummy)):\n    # to not overload chart\n    _df_train = df_dummy.loc[train_idx].sample(500)\n    _df_test = df_dummy.loc[test_idx].sample(500)\n    \n    ax[i + 1].scatter(_df_train['timestamp'], _df_train['value'], label=\"train\", alpha=0.2)\n    ax[i + 1].scatter(_df_test['timestamp'], _df_test['value'], label=\"test\", alpha=0.2)\n    \n    ax[i + 1].set_ylabel(\"cumsum of label\")\n    ax[i + 1].set_title(f\"split {i+1}\")\n\n    # add white background to legend\n    legend = ax[i + 1].legend(frameon=1)\n    frame = legend.get_frame()\n    frame.set_facecolor('w')\n\nif n_cols == 2:\n    ax[-2].set_xlabel(\"timestamp\")\nax[-1].set_xlabel(\"timestamp\")\n\nplt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:36:53.615940Z","iopub.execute_input":"2025-06-19T13:36:53.616235Z","iopub.status.idle":"2025-06-19T13:36:56.405485Z","shell.execute_reply.started":"2025-06-19T13:36:53.616212Z","shell.execute_reply":"2025-06-19T13:36:56.404695Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4.2. Train model\n\n### XGBoost version","metadata":{}},{"cell_type":"code","source":"\noof_preds = np.zeros(len(df_train))\ntest_preds = np.zeros(len(df_test))\n\ndf_cv_stats = list()\n\nfor i, (train_idx, valid_idx) in enumerate(kf.split(df_train)):\n    print(\"#\" * 25)\n    print(f\"### Fold {i + 1} \" + \"#\" * 14)\n    print(\"#\" * 25)\n\n    X_train = df_train.iloc[train_idx][selected_features]\n    y_train = df_train.iloc[train_idx][\"label\"]\n    X_valid = df_train.iloc[valid_idx][selected_features]\n    y_valid = df_train.iloc[valid_idx][\"label\"]\n    X_test = df_test[selected_features]\n\n    model = XGBRegressor(**xgb_params)\n\n    model.fit(\n        X_train, y_train,\n        eval_set=[(X_valid, y_valid)],\n        early_stopping_rounds=25,\n        verbose=200\n    )\n\n    oof_preds[valid_idx] = model.predict(X_valid)\n    test_preds += model.predict(X_test)\n\n    # calculate errors on train and validation samples\n    _train_cv = pearsonr(y_train, model.predict(X_train))[0]\n    _valid_cv = pearsonr(y_valid, oof_preds[valid_idx])[0]\n    # add out-of-time (validation) sample predictions\n    _oot_cv = pearsonr(df_valid['label'], model.predict(df_valid[selected_features]))[0]\n    df_cv_stats.append([f\"split {i+1}\", _train_cv, _valid_cv, _oot_cv])\n\n    print(\"\")\n\npearson_score = pearsonr(df_train[\"label\"], oof_preds)[0]\nprint(\"Final Pearson Correlation = \", pearson_score)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:36:56.406303Z","iopub.execute_input":"2025-06-19T13:36:56.406539Z","iopub.status.idle":"2025-06-19T13:37:44.484114Z","shell.execute_reply.started":"2025-06-19T13:36:56.406521Z","shell.execute_reply":"2025-06-19T13:37:44.483225Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### WrapGBM version","metadata":{}},{"cell_type":"code","source":"\n# oof_preds = np.zeros(len(df_train))\n# test_preds = np.zeros(len(df_test))\n\n# df_cv_stats = list()\n\n# for i, (train_idx, valid_idx) in enumerate(kf.split(df_train)):\n#     print(\"#\" * 25)\n#     print(f\"### Fold {i + 1} \" + \"#\" * 14)\n#     print(\"#\" * 25)\n\n#     X_train = df_train.iloc[train_idx][selected_features].values\n#     y_train = df_train.iloc[train_idx][\"label\"].values\n#     X_valid = df_train.iloc[valid_idx][selected_features].values\n#     y_valid = df_train.iloc[valid_idx][\"label\"].values\n#     X_test = df_test[selected_features].values\n\n#     model = WarpGBM(\n#         max_depth=10,\n#         num_bins=100,\n#         n_estimators=100,\n#         learning_rate=0.1,\n#         colsample_bytree=1.0,\n#         min_child_weight=4\n#     )\n#     model.fit(\n#         X_train,\n#         y_train,\n#         X_eval=X_valid,\n#         y_eval=y_valid,\n#         eval_every_n_trees=1,\n#         early_stopping_rounds=10,\n#         eval_metric=\"rmsle\",\n#     )\n\n#     #keep best model\n#     best_i = int(np.argmin(model.eval_loss))\n#     opt_num_trees = best_i * model.eval_every_n_trees  # +1 since first eval is at tree N, not 0\n#     model.forest = model.forest[:( opt_num_trees + 1)]\n\n#     oof_preds[valid_idx] = model.predict(X_valid)\n#     test_preds += model.predict(X_test)\n\n#     # calculate errors on train and validation samples\n#     _train_cv = pearsonr(y_train, model.predict(X_train))[0]\n#     _valid_cv = pearsonr(y_valid, oof_preds[valid_idx])[0]\n#     df_cv_stats.append([f\"split {i+1}\", _train_cv, _valid_cv])\n\n#     print(\"\")\n\n# pearson_score = pearsonr(df_train[\"label\"], oof_preds)[0]\n# print(\"Final Pearson Correlation = \", pearson_score)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:37:44.485040Z","iopub.execute_input":"2025-06-19T13:37:44.485330Z","iopub.status.idle":"2025-06-19T13:37:44.489572Z","shell.execute_reply.started":"2025-06-19T13:37:44.485300Z","shell.execute_reply":"2025-06-19T13:37:44.488959Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4.3. visualize validation scores","metadata":{}},{"cell_type":"code","source":"df_cv_stats = pd.DataFrame(df_cv_stats, columns=[\"fold\", 'Pearson train', \"Pearson valid\", \"Pearson out-of-time\"])\ndf_cv_stats['diff'] = df_cv_stats['Pearson train'] - df_cv_stats['Pearson valid']\ndf_cv_stats['ratio'] = df_cv_stats['Pearson train'] / df_cv_stats['Pearson valid']\ndf_cv_stats","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:37:44.490357Z","iopub.execute_input":"2025-06-19T13:37:44.490574Z","iopub.status.idle":"2025-06-19T13:37:44.516783Z","shell.execute_reply.started":"2025-06-19T13:37:44.490552Z","shell.execute_reply":"2025-06-19T13:37:44.516005Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\nfig, ax = plt.subplots(1, 1, figsize=(10, 4))\n\nsns.scatterplot(data=df_cv_stats, x=\"Pearson train\", y=\"Pearson valid\", hue=\"fold\", style=\"fold\", s=200, ax=ax)\n\n# add white background to legend\nlegend = ax.legend(frameon=1)\nframe = legend.get_frame()\nframe.set_facecolor('w')\n\nplt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:37:44.517536Z","iopub.execute_input":"2025-06-19T13:37:44.517802Z","iopub.status.idle":"2025-06-19T13:37:44.795503Z","shell.execute_reply.started":"2025-06-19T13:37:44.517775Z","shell.execute_reply":"2025-06-19T13:37:44.794677Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\nfig, ax = plt.subplots(1, 1, figsize=(10, 4))\n\n# out-of-time sample\nsns.boxplot(x=df_cv_stats[\"Pearson out-of-time\"])\n\nplt.tight_layout()","metadata":{"trusted":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2025-06-19T13:37:44.796379Z","iopub.execute_input":"2025-06-19T13:37:44.796617Z","iopub.status.idle":"2025-06-19T13:37:44.932550Z","shell.execute_reply.started":"2025-06-19T13:37:44.796581Z","shell.execute_reply":"2025-06-19T13:37:44.931940Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4.4. submit predictions","metadata":{}},{"cell_type":"code","source":"# sample=pd.read_csv(\"/kaggle/input/drw-crypto-market-prediction/sample_submission.csv\")\n# sample[\"prediction\"] = test_preds / n_splits\n# sample.to_csv(\"submission.csv\", index=False)\n# sample.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:37:44.933238Z","iopub.execute_input":"2025-06-19T13:37:44.933435Z","iopub.status.idle":"2025-06-19T13:37:44.936619Z","shell.execute_reply.started":"2025-06-19T13:37:44.933420Z","shell.execute_reply":"2025-06-19T13:37:44.936045Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 5. Walk-Forward Validation (expanding window)\n\n## 5.1. Visualize","metadata":{}},{"cell_type":"code","source":"\ndf_dummy = df_train[[\"timestamp\", 'label']].copy()\ndf_dummy['value'] = df_dummy['label'].cumsum()\n\n# compare split results\nfig, ax = plt.subplots((n_splits+1) // n_cols, n_cols, figsize=(18, 6), sharex=True, sharey=True)\nax = ax.flatten()\n\n# all data\nax[0].plot(df_dummy['timestamp'], df_dummy['value'], label=\"all data\", color=colors[2])\nax[0].set_ylabel(\"cumsum of label\")\nax[0].set_title(\"All data\")\n\n# add validation sections\nfor i in range(6):\n    ax[i].plot(df_valid['timestamp'], df_valid['label'].cumsum()+df_dummy['value'].iloc[-1], label=\"valid\", color='r')\n\nfor i, _idx in enumerate(range(n_splits)):\n    # mannualy generate data splits\n    _range_length = len(df_dummy) // (n_splits + 1)\n    _start_idx = _range_length * (_idx + 1)\n    _end_idx = _start_idx + _range_length\n    \n    _df_train = df_dummy.loc[:_start_idx]\n    _df_test = df_dummy.loc[_start_idx:_end_idx]\n\n    ax[i + 1].plot(_df_train['timestamp'], _df_train['value'], label=\"train\")\n    ax[i + 1].plot(_df_test['timestamp'], _df_test['value'], label=\"test\")\n    \n    ax[i + 1].set_ylabel(\"cumsum of label\")\n    ax[i + 1].set_title(f\"split {i+1}\")\n\n    # add white background to legend\n    legend = ax[i + 1].legend(frameon=1)\n    frame = legend.get_frame()\n    frame.set_facecolor('w')\n\nif n_cols == 2:\n    ax[-2].set_xlabel(\"timestamp\")\nax[-1].set_xlabel(\"timestamp\")\n\nplt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:37:44.937431Z","iopub.execute_input":"2025-06-19T13:37:44.937708Z","iopub.status.idle":"2025-06-19T13:37:48.810166Z","shell.execute_reply.started":"2025-06-19T13:37:44.937684Z","shell.execute_reply":"2025-06-19T13:37:48.809366Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5.2. Train model\n\n### XGBoost version","metadata":{}},{"cell_type":"code","source":"\noof_preds = np.zeros(len(df_train))\ntest_preds = np.zeros(len(df_test))\n\ndf_cv_stats = list()\n\nfor i, _idx in enumerate(range(n_splits)):\n    # mannualy generate data splits\n    _range_length = len(df_dummy) // (n_splits + 1)\n    _start_idx = _range_length * (_idx + 1)\n    _end_idx = _start_idx + _range_length\n    \n    print(\"#\" * 25)\n    print(f\"### Fold {i + 1} \" + \"#\" * 14)\n    print(\"#\" * 25)\n\n    X_train = df_train.iloc[:_start_idx][selected_features]\n    y_train = df_train.iloc[:_start_idx][\"label\"]\n    X_valid = df_train.iloc[_start_idx:_end_idx][selected_features]\n    y_valid = df_train.iloc[_start_idx:_end_idx][\"label\"]\n    X_test = df_test[selected_features]\n\n    model = XGBRegressor(**xgb_params)\n\n    model.fit(\n        X_train, y_train,\n        eval_set=[(X_valid, y_valid)],\n        early_stopping_rounds=25,\n        verbose=200\n    )\n\n    oof_preds[_start_idx:_end_idx] = model.predict(X_valid)\n    test_preds += model.predict(X_test)\n\n    # calculate errors on train and validation samples\n    _train_cv = pearsonr(y_train, model.predict(X_train))[0]\n    _valid_cv = pearsonr(y_valid, oof_preds[_start_idx:_end_idx])[0]\n    # add out-of-time (validation) sample predictions\n    _oot_cv = pearsonr(df_valid['label'], model.predict(df_valid[selected_features]))[0]\n    df_cv_stats.append([f\"split {i+1}\", _train_cv, _valid_cv, _oot_cv])\n\n    print(\"\")\n\npearson_score = pearsonr(df_train[\"label\"], oof_preds)[0]\nprint(\"Final Pearson Correlation = \", pearson_score)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:37:48.811023Z","iopub.execute_input":"2025-06-19T13:37:48.811358Z","iopub.status.idle":"2025-06-19T13:37:56.771429Z","shell.execute_reply.started":"2025-06-19T13:37:48.811340Z","shell.execute_reply":"2025-06-19T13:37:56.770713Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5.3. visualize validation scores","metadata":{}},{"cell_type":"code","source":"df_cv_stats = pd.DataFrame(df_cv_stats, columns=[\"fold\", 'Pearson train', \"Pearson valid\", \"Pearson out-of-time\"])\ndf_cv_stats['diff'] = df_cv_stats['Pearson train'] - df_cv_stats['Pearson valid']\ndf_cv_stats['ratio'] = df_cv_stats['Pearson train'] / df_cv_stats['Pearson valid']\ndf_cv_stats","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:37:56.772157Z","iopub.execute_input":"2025-06-19T13:37:56.772380Z","iopub.status.idle":"2025-06-19T13:37:56.785251Z","shell.execute_reply.started":"2025-06-19T13:37:56.772364Z","shell.execute_reply":"2025-06-19T13:37:56.784476Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\nfig, ax = plt.subplots(1, 1, figsize=(10, 4))\n\nsns.scatterplot(data=df_cv_stats, x=\"Pearson train\", y=\"Pearson valid\", hue=\"fold\", style=\"fold\", s=200, ax=ax)\n\n# add white background to legend\nlegend = ax.legend(frameon=1)\nframe = legend.get_frame()\nframe.set_facecolor('w')\n\nplt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:37:56.786134Z","iopub.execute_input":"2025-06-19T13:37:56.786442Z","iopub.status.idle":"2025-06-19T13:37:57.190673Z","shell.execute_reply.started":"2025-06-19T13:37:56.786419Z","shell.execute_reply":"2025-06-19T13:37:57.189996Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\nfig, ax = plt.subplots(1, 1, figsize=(10, 4))\n\n# out-of-time sample\nsns.boxplot(x=df_cv_stats[\"Pearson out-of-time\"])\n\nplt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:37:57.191361Z","iopub.execute_input":"2025-06-19T13:37:57.191585Z","iopub.status.idle":"2025-06-19T13:37:57.306464Z","shell.execute_reply.started":"2025-06-19T13:37:57.191561Z","shell.execute_reply":"2025-06-19T13:37:57.305889Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5.4. submit predictions","metadata":{}},{"cell_type":"code","source":"# sample=pd.read_csv(\"/kaggle/input/drw-crypto-market-prediction/sample_submission.csv\")\n# sample[\"prediction\"] = test_preds / n_splits\n# sample.to_csv(\"submission.csv\", index=False)\n# sample.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:37:57.307160Z","iopub.execute_input":"2025-06-19T13:37:57.307424Z","iopub.status.idle":"2025-06-19T13:37:57.310578Z","shell.execute_reply.started":"2025-06-19T13:37:57.307397Z","shell.execute_reply":"2025-06-19T13:37:57.309895Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 6. Standard walk-forward validation\n\n## 6.1. Visualize","metadata":{}},{"cell_type":"code","source":"\ndf_dummy = df_train[[\"timestamp\", 'label']].copy()\ndf_dummy['value'] = df_dummy['label'].cumsum()\n\n# compare split results\nfig, ax = plt.subplots((n_splits+1) // n_cols, n_cols, figsize=(18, 6), sharex=True, sharey=True)\nax = ax.flatten()\n\n# all data\nax[0].plot(df_dummy['timestamp'], df_dummy['value'], label=\"all data\", color=colors[2])\nax[0].set_ylabel(\"cumsum of label\")\nax[0].set_title(\"All data\")\n\n# add validation sections\nfor i in range(6):\n    ax[i].plot(df_valid['timestamp'], df_valid['label'].cumsum()+df_dummy['value'].iloc[-1], label=\"valid\", color='r')\n\nfor i, _idx in enumerate(range(n_splits)):\n    # mannualy generate data splits\n    _range_length = len(df_dummy) // (n_splits + 1)\n    _start_idx = _range_length * (_idx + 1)\n    _end_idx = _start_idx + _range_length\n    \n    _df_train = df_dummy.loc[(_start_idx-_range_length):_start_idx]\n    _df_test = df_dummy.loc[_start_idx:_end_idx]\n\n    ax[i + 1].plot(_df_train['timestamp'], _df_train['value'], label=\"train\")\n    ax[i + 1].plot(_df_test['timestamp'], _df_test['value'], label=\"test\")\n    \n    ax[i + 1].set_ylabel(\"cumsum of label\")\n    ax[i + 1].set_title(f\"split {i+1}\")\n\n    # add white background to legend\n    legend = ax[i + 1].legend(frameon=1)\n    frame = legend.get_frame()\n    frame.set_facecolor('w')\n\nif n_cols == 2:\n    ax[-2].set_xlabel(\"timestamp\")\nax[-1].set_xlabel(\"timestamp\")\n\nplt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:37:57.311316Z","iopub.execute_input":"2025-06-19T13:37:57.311562Z","iopub.status.idle":"2025-06-19T13:38:00.417371Z","shell.execute_reply.started":"2025-06-19T13:37:57.311538Z","shell.execute_reply":"2025-06-19T13:38:00.416637Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 6.2. Train model\n\n### XGBoost version","metadata":{}},{"cell_type":"code","source":"\noof_preds = np.zeros(len(df_train))\ntest_preds = np.zeros(len(df_test))\n\ndf_cv_stats = list()\n\nfor i, _idx in enumerate(range(n_splits)):\n    # mannualy generate data splits\n    _range_length = len(df_dummy) // (n_splits + 1)\n    _start_idx = _range_length * (_idx + 1)\n    _end_idx = _start_idx + _range_length\n    \n    print(\"#\" * 25)\n    print(f\"### Fold {i + 1} \" + \"#\" * 14)\n    print(\"#\" * 25)\n\n    X_train = df_train.iloc[(_start_idx-_range_length):_start_idx][selected_features]\n    y_train = df_train.iloc[(_start_idx-_range_length):_start_idx][\"label\"]\n    X_valid = df_train.iloc[_start_idx:_end_idx][selected_features]\n    y_valid = df_train.iloc[_start_idx:_end_idx][\"label\"]\n    X_test = df_test[selected_features]\n\n    model = XGBRegressor(**xgb_params)\n\n    model.fit(\n        X_train, y_train,\n        eval_set=[(X_valid, y_valid)],\n        early_stopping_rounds=25,\n        verbose=200\n    )\n\n    oof_preds[_start_idx:_end_idx] = model.predict(X_valid)\n    test_preds += model.predict(X_test)\n\n    # calculate errors on train and validation samples\n    _train_cv = pearsonr(y_train, model.predict(X_train))[0]\n    _valid_cv = pearsonr(y_valid, oof_preds[_start_idx:_end_idx])[0]\n    # add out-of-time (validation) sample predictions\n    _oot_cv = pearsonr(df_valid['label'], model.predict(df_valid[selected_features]))[0]\n    df_cv_stats.append([f\"split {i+1}\", _train_cv, _valid_cv, _oot_cv])\n\n    print(\"\")\n\npearson_score = pearsonr(df_train[\"label\"], oof_preds)[0]\nprint(\"Final Pearson Correlation = \", pearson_score)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:38:00.418156Z","iopub.execute_input":"2025-06-19T13:38:00.418354Z","iopub.status.idle":"2025-06-19T13:38:05.056784Z","shell.execute_reply.started":"2025-06-19T13:38:00.418339Z","shell.execute_reply":"2025-06-19T13:38:05.055926Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 6.3. visualize validation scores","metadata":{}},{"cell_type":"code","source":"df_cv_stats = pd.DataFrame(df_cv_stats, columns=[\"fold\", 'Pearson train', \"Pearson valid\", \"Pearson out-of-time\"])\ndf_cv_stats['diff'] = df_cv_stats['Pearson train'] - df_cv_stats['Pearson valid']\ndf_cv_stats['ratio'] = df_cv_stats['Pearson train'] / df_cv_stats['Pearson valid']\ndf_cv_stats","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:38:05.057654Z","iopub.execute_input":"2025-06-19T13:38:05.057915Z","iopub.status.idle":"2025-06-19T13:38:05.069538Z","shell.execute_reply.started":"2025-06-19T13:38:05.057896Z","shell.execute_reply":"2025-06-19T13:38:05.068939Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\nfig, ax = plt.subplots(1, 1, figsize=(10, 4))\n\nsns.scatterplot(data=df_cv_stats, x=\"Pearson train\", y=\"Pearson valid\", hue=\"fold\", style=\"fold\", s=200, ax=ax)\n\n# add white background to legend\nlegend = ax.legend(frameon=1)\nframe = legend.get_frame()\nframe.set_facecolor('w')\n\nplt.tight_layout()","metadata":{"trusted":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2025-06-19T13:38:05.070235Z","iopub.execute_input":"2025-06-19T13:38:05.070434Z","iopub.status.idle":"2025-06-19T13:38:05.341191Z","shell.execute_reply.started":"2025-06-19T13:38:05.070419Z","shell.execute_reply":"2025-06-19T13:38:05.340495Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\nfig, ax = plt.subplots(1, 1, figsize=(10, 4))\n\n# out-of-time sample\nsns.boxplot(x=df_cv_stats[\"Pearson out-of-time\"])\n\nplt.tight_layout()","metadata":{"trusted":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2025-06-19T13:38:05.341922Z","iopub.execute_input":"2025-06-19T13:38:05.342146Z","iopub.status.idle":"2025-06-19T13:38:05.458396Z","shell.execute_reply.started":"2025-06-19T13:38:05.342121Z","shell.execute_reply":"2025-06-19T13:38:05.457909Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 6.4. submit predictions","metadata":{}},{"cell_type":"code","source":"# sample=pd.read_csv(\"/kaggle/input/drw-crypto-market-prediction/sample_submission.csv\")\n# sample[\"prediction\"] = test_preds / n_splits\n# sample.to_csv(\"submission.csv\", index=False)\n# sample.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:38:05.458923Z","iopub.execute_input":"2025-06-19T13:38:05.459104Z","iopub.status.idle":"2025-06-19T13:38:05.462283Z","shell.execute_reply.started":"2025-06-19T13:38:05.459091Z","shell.execute_reply":"2025-06-19T13:38:05.461589Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 7. out-of-distribution (era-splitting)\n\n## 7.1. Visualize","metadata":{}},{"cell_type":"code","source":"# add era column to train DataFrame\ndf_train['era'] = pd.qcut(range(len(df_train)), q=20, labels=False, duplicates='drop')\ndf_train.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-20T13:17:08.835170Z","iopub.execute_input":"2025-06-20T13:17:08.835432Z","iopub.status.idle":"2025-06-20T13:17:08.920201Z","shell.execute_reply.started":"2025-06-20T13:17:08.835384Z","shell.execute_reply":"2025-06-20T13:17:08.919594Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\nn_splits = 5\nn_cols = 2\n\n# compare split results\nfig, ax = plt.subplots((n_splits+1) // n_cols, n_cols, figsize=(18, 6), sharex=True)\nax = ax.flatten()\n\n# all data\ndf_dummy['era'] = df_train['era']\nax[0].plot(df_dummy['timestamp'], df_dummy['value'], label=\"all data\", color=colors[2])\nax[0].set_ylabel(\"cumsum of label\")\nax[0].set_title(\"All data\")\n\nkf = KFold(n_splits=n_splits, shuffle=True, random_state=SEED)\nkf.get_n_splits(df_dummy)\n\n# add validation sections\nfor i in range(6):\n    ax[i].plot(df_valid['timestamp'], df_valid['label'].cumsum()+df_dummy['value'].iloc[-1], label=\"valid\", color='r')\n\nfor i, (train_idx, test_idx) in enumerate(kf.split(df_dummy)):\n    # to not overload chart\n    _df_train = df_dummy.loc[train_idx].sample(500)\n    _df_test = df_dummy.loc[test_idx].sample(500)\n\n    sns.scatterplot(data=_df_test, x=\"timestamp\", y=\"value\", hue=\"era\", ax=ax[i + 1])\n        \n    ax[i + 1].set_ylabel(\"cumsum of label\")\n    ax[i + 1].set_title(f\"split {i+1}\")\n\n    # add white background to legend\n    legend = ax[i + 1].legend(frameon=1)\n    frame = legend.get_frame()\n    frame.set_facecolor('w')\n\nif n_cols == 2:\n    ax[-2].set_xlabel(\"timestamp\")\nax[-1].set_xlabel(\"timestamp\")\n\nplt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-20T13:23:27.478934Z","iopub.execute_input":"2025-06-20T13:23:27.479217Z","iopub.status.idle":"2025-06-20T13:23:32.610618Z","shell.execute_reply.started":"2025-06-20T13:23:27.479196Z","shell.execute_reply":"2025-06-20T13:23:32.609855Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 7.2. Train model\n\n### WrapGBM version","metadata":{}},{"cell_type":"code","source":"\noof_preds = np.zeros(len(df_train))\ntest_preds = np.zeros(len(df_test))\n\ndf_cv_stats = list()\n\nfor i, (train_idx, valid_idx) in enumerate(kf.split(df_train)):\n    print(\"#\" * 25)\n    print(f\"### Fold {i + 1} \" + \"#\" * 14)\n    print(\"#\" * 25)\n\n    X_train = df_train.iloc[train_idx][selected_features].values\n    y_train = df_train.iloc[train_idx][\"label\"].values\n    X_valid = df_train.iloc[valid_idx][selected_features].values\n    y_valid = df_train.iloc[valid_idx][\"label\"].values\n    X_test = df_test[selected_features].values\n\n    '''Define an Array of Integers to Define the Eras'''\n    eras = df_train.loc[ train_idx, 'era' ].values\n    print(\"Era Vector: \", eras)\n\n    model = WarpGBM(\n        max_depth=10,\n        num_bins=100,\n        n_estimators=100,\n        learning_rate=0.1,\n        colsample_bytree=1.0,\n        min_child_weight=4\n    )\n    model.fit(\n        X_train,\n        y_train,\n        eras, # use the eras here, in .fit()\n        X_eval=X_valid,\n        y_eval=y_valid,\n        eval_every_n_trees=1,\n        early_stopping_rounds=10,\n        eval_metric=\"rmsle\",\n    )\n\n    #keep best model\n    best_i = int(np.argmin(model.eval_loss))\n    opt_num_trees = best_i * model.eval_every_n_trees  # +1 since first eval is at tree N, not 0\n    model.forest = model.forest[:( opt_num_trees + 1)]\n\n    oof_preds[valid_idx] = model.predict(X_valid)\n    test_preds += model.predict(X_test)\n\n    # calculate errors on train and validation samples\n    _train_cv = pearsonr(y_train, model.predict(X_train))[0]\n    _valid_cv = pearsonr(y_valid, oof_preds[valid_idx])[0]\n    # add out-of-time (validation) sample predictions\n    _oot_cv = pearsonr(df_valid['label'], model.predict(df_valid[selected_features].values))[0]\n    df_cv_stats.append([f\"split {i+1}\", _train_cv, _valid_cv, _oot_cv])\n\n    print(\"\")\n\npearson_score = pearsonr(df_train[\"label\"], oof_preds)[0]\nprint(\"Final Pearson Correlation = \", pearson_score)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T13:54:30.401830Z","iopub.execute_input":"2025-06-19T13:54:30.402122Z","iopub.status.idle":"2025-06-19T14:22:42.131852Z","shell.execute_reply.started":"2025-06-19T13:54:30.402096Z","shell.execute_reply":"2025-06-19T14:22:42.131025Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 7.3. visualize validation scores","metadata":{}},{"cell_type":"code","source":"df_cv_stats = pd.DataFrame(df_cv_stats, columns=[\"fold\", 'Pearson train', \"Pearson valid\", \"Pearson out-of-time\"])\ndf_cv_stats['diff'] = df_cv_stats['Pearson train'] - df_cv_stats['Pearson valid']\ndf_cv_stats['ratio'] = df_cv_stats['Pearson train'] / df_cv_stats['Pearson valid']\ndf_cv_stats","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-19T14:22:42.133108Z","iopub.execute_input":"2025-06-19T14:22:42.133346Z","iopub.status.idle":"2025-06-19T14:22:42.152519Z","shell.execute_reply.started":"2025-06-19T14:22:42.133328Z","shell.execute_reply":"2025-06-19T14:22:42.151303Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\nfig, ax = plt.subplots(1, 1, figsize=(10, 4))\n\nsns.scatterplot(data=df_cv_stats, x=\"Pearson train\", y=\"Pearson valid\", hue=\"fold\", style=\"fold\", s=200, ax=ax)\n\n# add white background to legend\nlegend = ax.legend(frameon=1)\nframe = legend.get_frame()\nframe.set_facecolor('w')\n\nplt.tight_layout()","metadata":{"trusted":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2025-06-19T14:22:42.152964Z","iopub.status.idle":"2025-06-19T14:22:42.153201Z","shell.execute_reply.started":"2025-06-19T14:22:42.153097Z","shell.execute_reply":"2025-06-19T14:22:42.153107Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\nfig, ax = plt.subplots(1, 1, figsize=(10, 4))\n\n# out-of-time sample\nsns.boxplot(x=df_cv_stats[\"Pearson out-of-time\"])\n\nplt.tight_layout()","metadata":{"trusted":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2025-06-19T14:22:42.154424Z","iopub.status.idle":"2025-06-19T14:22:42.154714Z","shell.execute_reply.started":"2025-06-19T14:22:42.154554Z","shell.execute_reply":"2025-06-19T14:22:42.154564Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}