{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84493,"databundleVersionId":9849268,"sourceType":"competition"},{"sourceId":201188248,"sourceType":"kernelVersion"}],"dockerImageVersionId":30786,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **FOREWORD**","metadata":{}},{"cell_type":"markdown","source":"## **KERNEL OBJECTIVE** <br>\n\nThis is a small starter kernel where I explore the dataset, train a simple lightgbm model for starters without early stopping and fixate my cv-scheme for the model process. \n\n## **COMPETITION DETAILS** <br>\nThe below links may be of use to one and all looking to seek further information on the competition <br>\n1. https://www.kaggle.com/code/ryanholbrook/jane-street-rmf-demo-submission <br>\n2. https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting <br> \n3. https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/data <br> \n\nThese links describe the data, competition requirements and provide understanding of the forecasting API ","metadata":{}},{"cell_type":"markdown","source":"# **IMPORTS**","metadata":{}},{"cell_type":"code","source":"%%capture\n\nimport polars as pl, pandas as pd, numpy as np\nfrom gc import collect\nfrom pprint import pprint","metadata":{"execution":{"iopub.status.busy":"2024-10-15T08:02:42.087207Z","iopub.execute_input":"2024-10-15T08:02:42.087860Z","iopub.status.idle":"2024-10-15T08:02:42.095273Z","shell.execute_reply.started":"2024-10-15T08:02:42.087816Z","shell.execute_reply":"2024-10-15T08:02:42.093993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **CONFIGURATION**","metadata":{}},{"cell_type":"code","source":"%%time \n\ntarget      = \"responder_6\"\nstate       = 42\ndev_strt_id = 4_500_000\nversion_nb  = \"V1_1\"","metadata":{"execution":{"iopub.status.busy":"2024-10-15T08:02:44.298540Z","iopub.execute_input":"2024-10-15T08:02:44.299029Z","iopub.status.idle":"2024-10-15T08:02:44.306174Z","shell.execute_reply.started":"2024-10-15T08:02:44.298983Z","shell.execute_reply":"2024-10-15T08:02:44.304999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **PREPROCESSING**\n\nWe load the data and fixate the cv scheme here. We shall delve into the dataset a bit later <br>\nWe don't need to specify the sub-directory paths while importing the datasets. Polars knows to import all training components as this is a hive dataset. Specifying the train path is enough <br>\n\n**Weights parameter** is important here - this is a sample weight used in our custom eval-metric <br>","metadata":{}},{"cell_type":"code","source":"%%time \n\ntrain = \\\npl.scan_parquet(\n    f\"/kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet\"\n).\\\nselect(\n    pl.int_range(pl.len(), dtype=pl.UInt32).alias(\"id\"),\n    pl.all(),\n)\n\nall_cols      = train.collect_schema().names()\nsel_cols      = [c for c in all_cols \n                 if c.startswith(\n                     (\"responder\", \"weight\", 'id', 'date_id','time_id', 'partition_id')\n                 ) == False\n                ]\ntargets       = [c for c in all_cols if c.startswith(\"responder\") == True]\nsample_weight = train.select(pl.col(\"weight\")).collect().to_series()\n\nprint(f\"\\n---> Selected columns\\n\")\nwith np.printoptions(linewidth = 100):\n    pprint(np.array(sel_cols))\n    \ncollect();\nprint()","metadata":{"execution":{"iopub.status.busy":"2024-10-15T08:02:46.666771Z","iopub.execute_input":"2024-10-15T08:02:46.667223Z","iopub.status.idle":"2024-10-15T08:02:47.343435Z","shell.execute_reply.started":"2024-10-15T08:02:46.667179Z","shell.execute_reply":"2024-10-15T08:02:47.342203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **CV SCHEME**\n\nWe follow a simple train-test split here. We train the model on older data and infer on recent data. We shall try and maintain the same period (number of days estimated in the public test data as our offline dev-set) <br>\n\nWe estimate the last date from the train dataset corresponding to our proposed CV scheme and then split the data accordingly","metadata":{}},{"cell_type":"code","source":"%%time \n\nlen_train   = train.select(pl.col(\"date_id\")).collect().shape[0]\nlen_ofl_mdl = len_train - dev_strt_id\nlast_tr_dt  = train.select(pl.col(\"date_id\")).collect().row(len_ofl_mdl)[0]\n\nprint(f\"\\n---> Last offline train date = {last_tr_dt}\\n\")\n\nXYtrain = train.filter(pl.col(\"date_id\").le(last_tr_dt))\nXYdev   = train.filter(pl.col(\"date_id\").gt(last_tr_dt))\n","metadata":{"execution":{"iopub.status.busy":"2024-10-15T08:03:16.181874Z","iopub.execute_input":"2024-10-15T08:03:16.182380Z","iopub.status.idle":"2024-10-15T08:03:16.294546Z","shell.execute_reply.started":"2024-10-15T08:03:16.182334Z","shell.execute_reply":"2024-10-15T08:03:16.293343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **CLOSURE**","metadata":{}},{"cell_type":"markdown","source":"We simply store the data into 2 tables organized by date-ids. This will help us select relevant dates for subsequent model development processes. <br>\nThis will also help us infer with a streaming OOF data for code checks if needed. ","metadata":{}},{"cell_type":"code","source":"%%time \n\nXYdev.collect().\\\nwrite_parquet(\n    \"XYdev.parquet\", partition_by = \"date_id\",\n)\n\nXYtrain.\\\ncollect().\\\nwrite_parquet(\n    \"XYtrain.parquet\", \n    partition_by = \"date_id\"\n)\n\ncollect()","metadata":{"execution":{"iopub.status.busy":"2024-10-15T08:03:59.539988Z","iopub.execute_input":"2024-10-15T08:03:59.540449Z","iopub.status.idle":"2024-10-15T08:05:14.104597Z","shell.execute_reply.started":"2024-10-15T08:03:59.540405Z","shell.execute_reply":"2024-10-15T08:05:14.102183Z"},"trusted":true},"execution_count":null,"outputs":[]}]}