{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":96164,"databundleVersionId":12993472,"sourceType":"competition"},{"sourceId":241449882,"sourceType":"kernelVersion"}],"dockerImageVersionId":31011,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **FOREWORD**\n\nThis is a baseline kernel for the [DRW - Crypto Market Prediction](https://www.kaggle.com/competitions/drw-crypto-market-prediction/code?competitionId=96164). This is a regression problem with pearson correlation eval-metric. This is a forecasting competition, with a public test set for syntax checks and a private test set for later periods. <br>\n\nFeature engineering is a key differentiator in such competitions. In this kernel, we use the last few timestamps as a cross-validation set and test out simple models to start the process. We then refit the model on the full train data and predict on the test set.\n\nWe borrow the model parameters from [here](https://www.kaggle.com/code/ravaghi/drw-crypto-market-prediction-ensemble)\n\n","metadata":{}},{"cell_type":"markdown","source":"# **IMPORTS**","metadata":{}},{"cell_type":"code","source":"!uv pip install -q --system -r /kaggle/input/drw2025-public-imports-v1/req_kaggle.txt\n\nexec( open(f\"/kaggle/input/drw2025-public-imports-v1/myimports.py\", \"r\").read() )\nexec( open(f\"/kaggle/input/drw2025-public-imports-v1/myutils.py\", \"r\").read() )\nexec( open(f\"/kaggle/input/drw2025-public-imports-v1/training.py\", \"r\").read() )\n\nprint()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-24T14:52:13.204135Z","iopub.execute_input":"2025-05-24T14:52:13.204328Z","iopub.status.idle":"2025-05-24T14:53:01.724482Z","shell.execute_reply.started":"2025-05-24T14:52:13.204308Z","shell.execute_reply":"2025-05-24T14:53:01.723412Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **CONFIGURATION**","metadata":{}},{"cell_type":"code","source":"%%time\n\nclass CFG:\n    \"\"\"\n    Configuration class for parameters and CV strategy for tuning and training\n    Some parameters may be unused here as this is a general configuration class\n    \"\"\";\n\n    # Data preparation:-\n    version_nb         = 2\n    model_id           = \"V2_2\"\n    model_label        = \"ML\"\n    test_req           = False\n    test_iter          = 200\n    gpu_switch         = \"ON\" if torch.cuda.is_available() else \"OFF\"\n    state              = 42\n    target             = f'label'\n    grouper            = f\"\"\n    tgt_mapper         = {}\n    ip_path            = f\"/kaggle/input/drw-crypto-market-prediction\"\n    op_path            = f\"/kaggle/working\"\n    orig_path          = f\"\"\n    data_path          = f\"\"\n    dtl_preproc_req    = True\n    ftre_plots_req     = True\n    ftre_imp_req       = True\n    nb_orig            = 0\n    orig_all_folds     = False\n\n    # Model Training:-\n    pstprcs_oof        = False\n    pstprcs_train      = False\n    pstprcs_test       = False\n    ML                 = True\n    test_preds_req     = True\n    n_splits           = 2\n    n_repeats          = 1\n    nbrnd_erly_stp     = 0\n    mdlcv_mthd         = 'TSS'\n    metric_obj         = 'maximize'\n\n    # Global variables for plotting:-\n    grid_specs = {'visible'  : True,\n                  'which'    : 'both',\n                  'linestyle': '--',\n                  'color'    : 'lightgrey',\n                  'linewidth': 0.75\n                 }\n\n    title_specs = {'fontsize'   : 9,\n                   'fontweight' : 'bold',\n                   'color'      : '#992600',\n                  }\n\ncollect()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-24T14:54:37.691447Z","iopub.execute_input":"2025-05-24T14:54:37.691936Z","iopub.status.idle":"2025-05-24T14:54:38.088426Z","shell.execute_reply.started":"2025-05-24T14:54:37.691884Z","shell.execute_reply":"2025-05-24T14:54:38.087369Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **KEY CONFIG PARAMETERS**\n\n|Configuration parameter| Explanation| Data type| Sample values |  \n| ---------------------- | ------------------------------- | --------------------- | --------------- |\n| version_nb    | Version Number | int | 1 | \n| model_id      | Model ID    | string | V1_1 | \n| model_label   | Model Label | string | ML | \n| test_req      | Test Required| bool | True / False | \n| test_iter     | Test case iterations for models | int | 50 |\n| gpu_switch      | Do we need GPU support | bool | True / False |\n| state           | Random state | int | 42 |\n| target          | Target column | str |  |\n| grouper         | CV grouper column | str |  |\n| ip_path, op_path | Data paths  | str | |\n| pstprcs_* | Do we need post-processing  | bool |True / False |\n| ML| Do we need machine learning models  | bool |True / False |\n| test_preds_req| Do we need test set predictions (training in inference kernel)  | bool |True / False |\n| n_splits/ n_repeats | N-splits and repeats for CV scheme | int | 3/5/10|\n| nbrnd_erly_stp | Early stopping rounds | int | 40|\n| mdlcv_mthd | Model CV method | str | RSKF|\n| ensemble_req | Do we need ensemble | bool | True / False |\n| metric_obj   | Metric direction | str | minimize/ maximize |","metadata":{}},{"cell_type":"markdown","source":"# **PREPROCESSING**\n\nIn this section, we load the dataset, drop useless columns and then initialize the CV scheme. <br>\nWe keep the last 10% data as development set here","metadata":{}},{"cell_type":"code","source":"%%time \n\ndrop_cols = \\\n[\n    'X697', 'X698', 'X699', 'X700', 'X701', 'X702', 'X703', 'X704', 'X705', 'X706', \n    'X707', 'X708', 'X709', 'X710', 'X711', 'X712', 'X713', 'X714', 'X715', 'X716',\n    'X717', 'X864', 'X867', 'X869', 'X870', 'X871', 'X872', 'X104', 'X110', 'X116',\n    'X122', 'X128', 'X134', 'X140', 'X146', 'X152', 'X158', 'X164', 'X170', 'X176',\n    'X182', 'X351', 'X357', 'X363', 'X369', 'X375', 'X381', 'X387', 'X393', 'X399',\n    'X405', 'X411', 'X417', 'X423', 'X429',\n    'timestamp', CFG.target,\n]\n\nXtrain = \\\n(\n    pl.scan_parquet(\n        os.path.join(CFG.ip_path, \"train.parquet\")\n    ).\n    drop( pl.col(drop_cols), strict = False)\n)\n\nytrain = \\\n(\n    pl.scan_parquet(\n        os.path.join(CFG.ip_path, \"train.parquet\")\n    ).\n    select(CFG.target).\n    collect(engine = \"streaming\").\n    to_pandas().\n    squeeze()   \n)\n\nlen_ = int(len(ytrain) * 0.90)\nytr  = ytrain.iloc[0: len_]\nydev = ytrain.iloc[len_ : ]\n\nXtr  = \\\n(\n    Xtrain.collect(engine = \"streaming\").\n    select(pl.all().shrink_dtype()).\n    to_pandas().\n    iloc[0 : len_]\n)\n\nXdev = \\\n(\n    Xtrain.collect(engine = \"streaming\").\n    select(pl.all().shrink_dtype()).\n    to_pandas().\n    iloc[len_ : ]\n)\n\nPrintColor(f\"---> Shape = {ytr.shape} {ydev.shape} {Xtr.shape} {Xdev.shape}\")\n\n_ = utils.CleanMemory()\nprint()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-24T14:55:53.313865Z","iopub.execute_input":"2025-05-24T14:55:53.315227Z","iopub.status.idle":"2025-05-24T14:56:01.695842Z","shell.execute_reply.started":"2025-05-24T14:55:53.315183Z","shell.execute_reply":"2025-05-24T14:56:01.694871Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **MODEL TRAINING**","metadata":{}},{"cell_type":"markdown","source":"## **OFFLINE CV**","metadata":{}},{"cell_type":"code","source":"%%time \n\nMdl_Master = \\\n{ \n f\"XGB1R\"   : XGBR(**{  \"objective\"              : \"reg:squarederror\",\n                        \"device\"                 : \"cuda\" if CFG.gpu_switch == \"ON\" else \"cpu\", \n                        \"n_estimators\"           : 679 if CFG.test_req == False else 25,\n                        \"learning_rate\"          : 0.096,\n                        \"colsample_bylevel\"      : 0.46,\n                        \"colsample_bynode\"       : 0.60,\n                        \"colsample_bytree\"       : 0.115,\n                        \"gamma\"                  : 1.039,\n                        \"max_depth\"              : 40,\n                        \"max_leaves\"             : 19,\n                        \"min_child_weight\"       : 76,\n                        \"n_jobs\"                 : -1,\n                        \"random_state\"           : CFG.state,\n                        \"reg_alpha\"              : 65.41,\n                        \"reg_lambda\"             : 19.907,\n                        \"subsample\"              : 0.0144,\n                        \"verbosity\"              : 0,\n                    }\n                  ),\n    \n f'LGBM1R'  : LGBMR(**{ \"objective\"          : \"regression_l2\",\n                        'device'             : \"gpu\" if CFG.gpu_switch == \"ON\" else \"cpu\",\n                        \"n_estimators\"       : 441 if CFG.test_req == False else 25,\n                        \"learning_rate\"      : 0.0145,\n                        \"colsample_bytree\"   : 0.524,               \n                        \"min_child_samples\"  : 47,\n                        \"min_child_weight\"   : 0.193,\n                        \"n_jobs\"             : -1,\n                        \"num_leaves\"         : 65,\n                        \"random_state\"       : CFG.state,\n                        \"reg_alpha\"          : 76.69,\n                        \"reg_lambda\"         : 78.57,\n                        \"subsample\"          : 0.35,\n                        \"verbosity\"          : -1\n                      }\n                   ),\n\n f'LGBM2R'  : LGBMR(**{ \"objective\"              : \"regression_l2\",\n                        \"data_sample_strategy\"   : \"goss\",\n                        \"n_estimators\"           : 268 if CFG.test_req == False else 25,\n                        'device'                 : \"gpu\" if CFG.gpu_switch == \"ON\" else \"cpu\",\n                        \"learning_rate\"          : 0.0136,\n                        \"colsample_bytree\"       : 0.32,\n                        \"min_child_samples\"      : 47,\n                        \"min_child_weight\"       : 0.652,\n                        \"n_jobs\"                 : -1,\n                        \"num_leaves\"             : 25,\n                        \"random_state\"           : CFG.state,\n                        \"reg_alpha\"              : 24.43,\n                        \"reg_lambda\"             : 39.82,\n                        \"subsample\"              : 0.21,\n                        \"verbosity\"              : -1\n                     }\n                   ),\n}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-24T14:56:32.824398Z","iopub.execute_input":"2025-05-24T14:56:32.824824Z","iopub.status.idle":"2025-05-24T14:56:32.836504Z","shell.execute_reply.started":"2025-05-24T14:56:32.824795Z","shell.execute_reply":"2025-05-24T14:56:32.835171Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time\n\nfor method, model in tqdm(Mdl_Master.items()):\n    PrintColor(\n        f\"\\n{'=' * 20} {method.upper()} MODEL TRAINING {'=' * 20}\\n\"\n    )\n\n    try:\n        model.fit(Xtr,ytr, verbose = 0) \n    except:\n        model.fit(Xtr , ytr)\n\n    dev_preds  = model.predict(Xdev)\n    score      = utils.ScoreMetric(ydev, dev_preds)\n    PrintColor(\n        f\"\\n---> OOF score = {score :,.8f}\\n\",\n        color = Fore.RED\n    )\n    ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-24T14:58:26.662158Z","iopub.execute_input":"2025-05-24T14:58:26.662568Z","iopub.status.idle":"2025-05-24T15:02:16.660997Z","shell.execute_reply.started":"2025-05-24T14:58:26.662539Z","shell.execute_reply":"2025-05-24T15:02:16.659948Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **FULL REFIT MODELS**","metadata":{}},{"cell_type":"code","source":"%%time \n\nMdl_Master = \\\n{ \n f\"XGB1R\"   : XGBR(**{  \"objective\"              : \"reg:squarederror\",\n                        \"device\"                 : \"cuda\" if CFG.gpu_switch == \"ON\" else \"cpu\", \n                        \"n_estimators\"           : 725 if CFG.test_req == False else 25,\n                        \"learning_rate\"          : 0.096,\n                        \"colsample_bylevel\"      : 0.46,\n                        \"colsample_bynode\"       : 0.60,\n                        \"colsample_bytree\"       : 0.115,\n                        \"gamma\"                  : 1.039,\n                        \"max_depth\"              : 40,\n                        \"max_leaves\"             : 19,\n                        \"min_child_weight\"       : 76,\n                        \"n_jobs\"                 : -1,\n                        \"random_state\"           : CFG.state,\n                        \"reg_alpha\"              : 65.41,\n                        \"reg_lambda\"             : 19.907,\n                        \"subsample\"              : 0.0144,\n                        \"verbosity\"              : 0,\n                    }\n                  ),\n    \n f'LGBM1R'  : LGBMR(**{ \"objective\"          : \"regression_l2\",\n                        'device'             : \"gpu\" if CFG.gpu_switch == \"ON\" else \"cpu\",\n                        \"n_estimators\"       : 500 if CFG.test_req == False else 25,\n                        \"learning_rate\"      : 0.0145,\n                        \"colsample_bytree\"   : 0.524,               \n                        \"min_child_samples\"  : 47,\n                        \"min_child_weight\"   : 0.193,\n                        \"n_jobs\"             : -1,\n                        \"num_leaves\"         : 65,\n                        \"random_state\"       : CFG.state,\n                        \"reg_alpha\"          : 76.69,\n                        \"reg_lambda\"         : 78.57,\n                        \"subsample\"          : 0.35,\n                        \"verbosity\"          : -1\n                      }\n                   ),\n\n f'LGBM2R'  : LGBMR(**{ \"objective\"              : \"regression_l2\",\n                        \"data_sample_strategy\"   : \"goss\",\n                        \"n_estimators\"           : 300 if CFG.test_req == False else 25,\n                        'device'                 : \"gpu\" if CFG.gpu_switch == \"ON\" else \"cpu\",\n                        \"learning_rate\"          : 0.0136,\n                        \"colsample_bytree\"       : 0.32,\n                        \"min_child_samples\"      : 47,\n                        \"min_child_weight\"       : 0.652,\n                        \"n_jobs\"                 : -1,\n                        \"num_leaves\"             : 25,\n                        \"random_state\"           : CFG.state,\n                        \"reg_alpha\"              : 24.43,\n                        \"reg_lambda\"             : 39.82,\n                        \"subsample\"              : 0.21,\n                        \"verbosity\"              : -1\n                     }\n                   ),\n}\n\nMdl_Preds = {}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-24T15:02:25.084472Z","iopub.execute_input":"2025-05-24T15:02:25.084827Z","iopub.status.idle":"2025-05-24T15:02:25.095879Z","shell.execute_reply.started":"2025-05-24T15:02:25.084801Z","shell.execute_reply":"2025-05-24T15:02:25.094206Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time \n\nXtest = \\\n(\n    pl.scan_parquet(\n        os.path.join(CFG.ip_path, \"test.parquet\")\n    ).\n    drop( pl.col(drop_cols), strict = False).\n    collect().\n    select(pl.all().shrink_dtype()).\n    to_pandas()\n)\n\nXtrain = \\\nXtrain.collect(engine = \"streaming\").select(pl.all().shrink_dtype()).to_pandas()\n\nfor method, model in tqdm(Mdl_Master.items()):\n    PrintColor(f\"---> {method.upper()} MODEL FULL-REFIT TRAINING\")\n\n    if \"XGB\" in method or \"CB\" in method:\n        model.fit(Xtrain, ytrain, verbose = 0)\n    else:\n        model.fit(Xtrain, ytrain,)\n    Mdl_Preds[method] = model.predict(Xtest)   \n\n_ = utils.CleanMemory()\nprint()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-24T15:02:34.600325Z","iopub.execute_input":"2025-05-24T15:02:34.601446Z","iopub.status.idle":"2025-05-24T15:07:56.243594Z","shell.execute_reply.started":"2025-05-24T15:02:34.601378Z","shell.execute_reply":"2025-05-24T15:07:56.242526Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **CLOSURE**","metadata":{}},{"cell_type":"code","source":"%%time\n\nsub_fl = pd.read_csv(os.path.join(CFG.ip_path, \"sample_submission.csv\"), index_col = [\"ID\"])\n\nsub_fl[\"prediction\"] = \\\nnp.average( \n    pd.DataFrame(Mdl_Preds).to_numpy(), \n    axis= 1, \n    weights = [0.50, 0.20, 0.30]\n)\n\nsub_fl.to_csv(\n    os.path.join(CFG.op_path, f\"submission.csv\")\n)\n\nprint()\n!ls\nprint()\n!head submission.csv\n\n_ = utils.CleanMemory()\nprint()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-24T15:08:07.455812Z","iopub.execute_input":"2025-05-24T15:08:07.456185Z","iopub.status.idle":"2025-05-24T15:08:14.563152Z","shell.execute_reply.started":"2025-05-24T15:08:07.456159Z","shell.execute_reply":"2025-05-24T15:08:14.561817Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null}]}