{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":96164,"databundleVersionId":11418275,"sourceType":"competition"}],"dockerImageVersionId":31040,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Reverse Engineering the original order of the Test rows\n\nThis notebook tries to predict the original order of rows in the test dataset. It also tries to predict month, season, and specific dates that can be important for bitcoin price.\n\nIn this challenge, the authors anonymized and shuffled the original order of the rows in test. This is unfortunate, because it means we cannot use any time-series approach.\n\nHere we explore whether the order of the Test rows can be reconstructed. The first approach is to train a model to predict the timestamp or the order in the train data, and look at the features that are more important. We can also train a model to predict month and season (TBD). The hypothesis is that some of the variables in train may be related to the previous time points - for example, they could be window-based averages or percentage increases compared to the previous. To do this, we train a model to predict the order, and look at variable importance. ","metadata":{}},{"cell_type":"markdown","source":"## Utility functions and parameters","metadata":{}},{"cell_type":"code","source":"%%capture\n\n#!pip install -q scikit-learn==1.3.2 autogluon==0.8.2\n!pip install -q autogluon","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-28T13:06:47.053689Z","iopub.execute_input":"2025-05-28T13:06:47.053944Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport pandas as pd\nimport polars as pl\nimport numpy as np\nfrom sklearn.ensemble import GradientBoostingRegressor\nfrom sklearn.model_selection import train_test_split\n\ndef is_interactive_session():\n    return os.environ.get('KAGGLE_KERNEL_RUN_TYPE','') == 'Interactive'\n\nis_interactive_session()\n\nconfig = {\n    \"autogluon_time\": 3600,\n    #\"reduce_features\": 0, # Set to >0 to use only the first n features\n    \"tail_rows\": 0 # Set to >0 to use only the last n rows in the file\n}\n\nif is_interactive_session():\n    print(\"Interactive session\")\n    config[\"autogluon_time\"] = 100\n    #config[\"reduce_features\"] = 200\n    config[\"tail_rows\"] = 2000\n    print(config)\nelse:\n    print(\"running as job\")\n    print(config)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Read Data","metadata":{}},{"cell_type":"code","source":"\n\n# Load the data\ntrain_df = pl.read_parquet(\"/kaggle/input/drw-crypto-market-prediction/train.parquet\")\ntrain_df = train_df.select(pl.all().shrink_dtype())\ntrain_df = train_df.to_pandas()\n#train_df = train_df.sort_values('timestamp').reset_index(drop=True)\ntrain_df['time_rank'] = np.arange(len(train_df)) / len(train_df)  # normalized 0–1\ntrain_df.head()\n\n# No need to read test yet\n#test_df = pl.read_parquet(\"/kaggle/input/drw-crypto-market-prediction/test.parquet#\")\n#test_df = test_df.select(pl.all().shrink_dtype())\n#test_df = test_df.to_pandas()\n#test_df.head()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Apply filters (if defined in config) (useful for debugging)","metadata":{}},{"cell_type":"code","source":"if config[\"tail_rows\"]>0:\n    train_df = train_df.tail(config[\"tail_rows\"])\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Train model","metadata":{}},{"cell_type":"code","source":"from autogluon.tabular import TabularPredictor\n\n# --- Step 3: Clean features ---\ntrain_df = train_df.replace([np.inf, -np.inf], np.nan).fillna(0)\ntrain_df = train_df.drop(columns=['timestamp', 'label'], errors='ignore')\n\n# Optional: reduce rows for memory\n# train_df = train_df.tail(200_000)\n\n# --- Step 4: Train AutoGluon model ---\npredictor = TabularPredictor(label='time_rank', eval_metric='r2').fit(\n    train_df,\n    presets='medium_quality',  # or 'best_quality' if memory allows\n    time_limit=config[\"autogluon_time\"],  # Optional: 30 minutes max\n    excluded_model_types=['NN_TORCH', 'CATBOOST']  # lighter memory\n)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fi = predictor.feature_importance(train_df)\nfi['importance'].head(20).plot(kind='barh', figsize=(8, 6), title='Top Temporal Features')\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- Step 3: Predict and evaluate ---\ny_true = val_data['time_rank'].values\ny_pred = predictor.predict(val_data.drop(columns='time_rank'))\n\nr2 = r2_score(y_true, y_pred)\nrmse = mean_squared_error(y_true, y_pred, squared=False)\nrank_corr, _ = spearmanr(y_true, y_pred)\n\nprint(f\"📈 R²: {r2:.4f}\")\nprint(f\"📉 RMSE: {rmse:.4f}\")\nprint(f\"📊 Spearman Rank Correlation: {rank_corr:.4f}\")\n\n# --- Step 4: Plot true vs. predicted time_rank ---\nplt.figure(figsize=(8, 6))\nplt.scatter(y_true, y_pred, alpha=0.2, s=10)\nplt.plot([0, 1], [0, 1], 'r--', lw=1)\nplt.xlabel(\"True time_rank\")\nplt.ylabel(\"Predicted time_rank\")\nplt.title(\"True vs Predicted time_rank\")\nplt.grid(True)\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Prediction on Test","metadata":{}},{"cell_type":"code","source":"# --- Load and shrink test set ---\ntest_df = pl.read_parquet(\"/kaggle/input/drw-crypto-market-prediction/test.parquet\")\ntest_df = test_df.select(pl.all().shrink_dtype()).to_pandas()\n\n# --- Clean test data ---\ntest_df = test_df.replace([np.inf, -np.inf], np.nan).fillna(0)\ntest_X = test_df.drop(columns='label', errors='ignore')\n\n# --- Predict time ---\ntest_df['predicted_time'] = predictor.predict(test_X)\n\n# --- Reorder test rows ---\ntest_df_ordered = test_df.sort_values('predicted_time').reset_index(drop=True)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}