{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84493,"databundleVersionId":9849268,"sourceType":"competition"},{"sourceId":9634681,"sourceType":"datasetVersion","datasetId":5882430},{"sourceId":201377683,"sourceType":"kernelVersion"},{"sourceId":201255000,"sourceType":"kernelVersion"}],"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **FOREWORD**","metadata":{}},{"cell_type":"markdown","source":"This kernel is my opening model for the Jane Street 2024 Kaggle competition. We have a modified R-square metric as the eval metric here. This needs to be maximized. <br>\n\nIn this kernel, we explore a baseline approach using single lightgbm offline and online models using a simple train-test split CV strategy <br>\n\nLet's hope to grow this process as we move along the competition! All the best!","metadata":{}},{"cell_type":"markdown","source":"# **IMPORTS**\n\nWe import key packages from my [imports]() kernel and leverage the potential of Polars GPU here. Let's hope to expedite the data wrangling process with blazing fast Polars + 2x T4 GPUs on Kaggle! <br>","metadata":{}},{"cell_type":"code","source":"%%time \n\n!pip install polars[gpu]==1.9.0 -q --no-index --find-links=/kaggle/input/janestreet2024-imports-v1/polars\n!pip install lightgbm==4.5.0 -q --no-index --find-links=/kaggle/input/janestreet2024-imports-v1/packages\n!pip install scikit-learn==1.5.2 -q --no-index --find-links=/kaggle/input/janestreet2024-imports-v1/packages\n\nexec(\n    open(\"/kaggle/input/janestreet2024-imports-v1/myimports.py\", \"r\"\n        ).read()\n)","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **CONFIGURATION**","metadata":{}},{"cell_type":"code","source":"%%time\n\ntarget     = \"responder_6\"\nversion_nb = \"V1_1\"\nop_path    = f\"/kaggle/working\"\nip_path    = f\"/kaggle/input/janestreet2024-dataload-v1\"\nstate      = 42\nmethod     = \"LGBM1R\"\n\ntitle_specs     = {\"color\": \"maroon\", \"fontsize\" : 15, \"fontweight\" : \"bold\"}\n\n# Change this to use various start dates in the offline model\noffline_strt_dt = 500\n\n# Global variable for fitted models for the test set inference\nfitted_models = {}","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **PREPROCESSING**","metadata":{}},{"cell_type":"markdown","source":"This is a rather simple section used for data loads and basic features. We fixate the cv fold here for offline learning <br>\nTrain-test split usually works well in such competitions. We try and replicate the size of the public leaderboard here by placing the last 4.5 million rows in the offline CV evaluation set. This roughly translates to a date id of 1578. <br>\n\nWe place all date ids till 1578 in the training dataset. All dates after this date are part of offline CV scheme. <br> \n\nWe load the datasets here using lazy frame mode in Polars and shortlist features for the offline model. We then convert the data into pandas and prepare the Xtrain-ytrain/ Xdev-ydev for the model training section. In this simple baseline, we use stock-id and features_01 - feature_78 as features. <br>\n","metadata":{}},{"cell_type":"code","source":"%%time \n\n# Offline train set\ntrain    = pl.scan_parquet(os.path.join(ip_path, f\"XYtrain.parquet\"))\nsel_cols = train.collect_schema().names()\nsel_cols = \\\n[c for c in sel_cols\n  if c not in [\"id\", \"date_id\", \"time_id\", \"partition_id\"]\n]\ndrop_cols = [c for c in sel_cols if \"responder\" in c] + [\"weight\"]\n\nXtrain = \\\ntrain.filter(pl.col(\"date_id\").gt(offline_strt_dt)).\\\nselect(pl.col(sel_cols)).\\\ncollect(engine = \"gpu\").\\\nto_pandas()\n\nXtrain.index = range(len(Xtrain))\nytrain       = Xtrain[target]\nsw_tr        = Xtrain[\"weight\"].values.flatten()\nXtrain       = Xtrain.drop(drop_cols , axis=1, errors = \"ignore\")\n\n# Offline validation set\ndev  = pl.scan_parquet(os.path.join(ip_path, f\"XYdev.parquet\"))\nXdev = \\\ndev.select(pl.col(sel_cols)).\\\ncollect(engine = \"gpu\").\\\nto_pandas()\n\nXdev.index = range(len(Xdev))\nydev       = Xdev[target]\nsw_dev     = Xdev[\"weight\"].values.flatten()\nXdev       = Xdev.drop(drop_cols , axis=1, errors = \"ignore\")\n\ncollect();\nPrintColor(f\"\\n---> Shapes = {Xtrain.shape} {ytrain.shape} {sw_tr.shape} -- {Xdev.shape} {ydev.shape} {sw_dev.shape}\\n\")\n\nPrintColor(f\"---> Final features\")\nwith np.printoptions(linewidth = 100):\n    pprint(np.array(Xtrain.columns))\n\nprint()\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **OFFLINE MODEL TRAINING**","metadata":{}},{"cell_type":"code","source":"%%time \n\ndef ScoreMetric(ytrue, ypred, weight):\n    \"\"\"\n    This function is a modification of the ready-made R-square function with sample weight. \n    We have this as a column in the dataset\n    \"\"\"\n    \n    return r2_score(ytrue, ypred, sample_weight = weight)\n\nclass CustomMetricMaker:\n    \"This class makes the custom metric for LGBM and XGBoost early stopping\"\n\n    def __init__(self, method):\n        self.method = method\n\n    def make_metric(self, ytrue, ypred, weight):\n        \"\"\"\n        This method returns the relevant metric for LGBM and XGB. \n        Catboost has a slightly different signature for the same- will be provided in version 2\n        \"\"\"\n        \n        if \"LGB\" in self.method:\n            return 'Wgt_RSquare', ScoreMetric(ytrue, ypred, weight), True\n        else:\n            return ScoreMetric(ytrue, ypred, weight)\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **MODEL TRAINING**","metadata":{}},{"cell_type":"code","source":"%%time\n\nmymetric = CustomMetricMaker(method)\n\nmodel = \\\nLGBMR(**{\"device\"           : \"gpu\",\n         \"objective\"        : \"regression_l2\",\n         \"metrics\"          : \"custom\",\n         \"n_estimators\"     : 1000,\n         \"max_depth\"        : 10,\n         \"learning_rate\"    : 0.03,\n         \"colsample_bytree\" : 0.55,\n         \"subsample\"        : 0.80,\n         \"random_state\"     : state,\n         \"reg_lambda\"       : 1.25,\n         \"reg_alpha\"        : 0.001,\n         \"verbosity\"        : -1,\n         }\n      )\n\nmodel.fit(\n    Xtrain, ytrain,\n    eval_set           = [(Xdev, ydev)],\n    eval_names         = [(\"Dev\")],\n    eval_metric        = [mymetric.make_metric],\n    callbacks          = [log_evaluation(0), early_stopping(100, verbose = False)],\n)\n\nprint()\ndisplay(model)\n\n\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **OFFLINE INFERENCE**","metadata":{}},{"cell_type":"code","source":"%%time \n\ntr_preds   = model.predict(Xtrain)\ntr_score   = ScoreMetric(ytrain, tr_preds, sw_tr)\nbest_iter  = model.best_iteration_\ndev_preds  = model.predict(Xdev)\nscore      = ScoreMetric(ydev, dev_preds, sw_dev)\n\nfitted_models[f\"Offline\"] = model\n\n# Feature Importances:-\nftreimp = \\\npd.Series(data = model.feature_importances_, index = Xdev.columns)\n\nPrintColor(\n    f\"---> OOF = {score :.6f} Train = {tr_score :.6f} Iter = {best_iter :,.0f}\",\n    color = Fore.RED\n)\n\nftreimp.name = f\"{method}{version_nb}\"\nftreimp      = ftreimp.sort_values(ascending = False)\n\nfig, axes = \\\nplt.subplots(\n    2,1, figsize = (20, 20),\n    gridspec_kw = {\"wspace\" : 1.00, \"hspace\" : 0.75}\n)\n\nax = axes[0]\nftreimp.head(25).plot.bar(ax= ax, color = \"blue\")\nax.set_title(\n    f\"\\nFeature Importances - top 25\\n\", **title_specs\n)\nax.grid(visible = False)\n\nax = axes[1]\nftreimp.sort_values(ascending = True).\\\nhead(25).\\\nplot.bar(ax= ax, color = \"tab:blue\")\nax.set_title(\n    f\"\\nFeature Importances - bottom 25\\n\", **title_specs\n)\nax.grid(visible = False)\n\nplt.tight_layout()\nplt.show()\n\nftreimp.to_frame().\\\nto_csv(\n    os.path.join(op_path, f\"FtreImp_{method}{version_nb}.csv\")\n)\n\nprint()\ncollect();","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **ONLINE MODEL TRAINING**\n\nThis is a full-refit of the model from before. We estimate **n_estimators as 1.1 * best_iterations** from offline model inference, remove early stopping and refit the model to the complete train and OOF sections, after the chosen start date. So if we start the training from day 500, we refit the model from day 500 - last train date and use this for the test set predictions. <br> ","metadata":{}},{"cell_type":"code","source":"%%time\n\nX = pd.concat([Xtrain, Xdev], axis=0 , ignore_index = True)\ny = pd.concat([ytrain, ydev], axis=0 , ignore_index = True)\n\ndel Xtrain, Xdev, ytrain, ydev\n\nmodel = \\\nLGBMR(**{\"device\"           : \"gpu\",\n         \"objective\"        : \"regression_l2\",\n         \"n_estimators\"     : 450,\n         \"max_depth\"        : 10,\n         \"learning_rate\"    : 0.03,\n         \"colsample_bytree\" : 0.55,\n         \"subsample\"        : 0.80,\n         \"random_state\"     : state,\n         \"reg_lambda\"       : 1.25,\n         \"reg_alpha\"        : 0.001,\n         \"verbosity\"        : -1,\n         }\n      )\n\nmodel.fit(X, y, callbacks = [log_evaluation(0)])\nfitted_models[f\"Online\"] = model\n\njoblib.dump(\n    fitted_models,\n    os.path.join(op_path, f\"LGBM{version_nb}.joblib\")\n)\n\ndel X,y\ncollect()\nprint()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **NEXT STEP**\n\nThis is the end of my baseline train kernel. We now focus on infering this model on the provided test set API and submit using the provided submission structure <br>\n\nI have saved the trained model artefacts in the dataset [here](https://www.kaggle.com/datasets/ravi20076/janestreetpublicv1) <br>\nThis dataset contains the offline and online models for the 5 versions and the feature importances for all component features \n\nAll the best!","metadata":{}}]}