{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84493,"databundleVersionId":9871156,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Jane Street Real-Time Market Data Forecasting","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"This is the first in a series of notebooks that I'm preparing for the competition. We will go through the provided data and we will test some modelsfor predictions. We will understand the metric used in the competition and how to improve the results by including lagged responders.","metadata":{}},{"cell_type":"markdown","source":"# Import libraries","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport os\nimport numpy as np\nimport polars as pl\nfrom matplotlib import pyplot as plt\nimport seaborn as sns\nfrom pathlib import Path\nfrom tqdm import tqdm\nfrom functools import reduce\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.linear_model import LinearRegression\nfrom sklearn.metrics import mean_squared_error\n\nfrom scipy.stats import pearsonr, spearmanr, kendalltau\n\nDATA_DIR = Path('/kaggle/input/jane-street-real-time-market-data-forecasting')\nN_PARTITION = 10\n\nfeature_cols = [f'feature_{x:02}' for x in range(79)]\nresponder_cols = [f'responder_{i}' for i in range(9)]","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:39.839862Z","iopub.execute_input":"2024-11-23T01:14:39.840321Z","iopub.status.idle":"2024-11-23T01:14:43.434387Z","shell.execute_reply.started":"2024-11-23T01:14:39.840263Z","shell.execute_reply":"2024-11-23T01:14:43.432954Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Quick look","metadata":{}},{"cell_type":"code","source":"# let's look to the provided attachments file for the competition\nos.listdir(DATA_DIR)","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:43.437023Z","iopub.execute_input":"2024-11-23T01:14:43.437694Z","iopub.status.idle":"2024-11-23T01:14:43.447030Z","shell.execute_reply.started":"2024-11-23T01:14:43.437642Z","shell.execute_reply":"2024-11-23T01:14:43.445893Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We have several files within our data:\n\n- **train.parquet** - The training set, contains historical data and returns. For convenience, the training set has been partitioned into ten parts.\n- **test.parquet** - A mock test set which represents the structure of the unseen test set. This example set demonstrates a single batch served by the evaluation API, that is, data from a single date_id, time_id pair. The test set contains columns including date_id, time_id, symbol_id, weight, is_scored, and feature_{00...78}. You will not be directly using the test set or sample submission in this competition, as the evaluation API will get/set the test set and predictions.\nThe field *is_scored* indicates whether this row is included in the evaluation metric calculation.\n- **lags.parquet** - Values of responder_{0...8} lagged by one date_id. The evaluation API serves the entirety of the lagged responders for a date_id on that date_id's first time_id. In other words, all of the previous date's responders will be served at the first time step of the succeeding date.\n- **sample_submission.csv** - This file illustrates the format of the predictions your model should make.\n- **features.csv** - metadata pertaining to the anonymized features\n- **responders.csv** - metadata pertaining to the anonymized responders","metadata":{}},{"cell_type":"markdown","source":"## Train data","metadata":{}},{"cell_type":"markdown","source":"Let's start with the most important file, the train data","metadata":{}},{"cell_type":"code","source":"# train data\nos.listdir(DATA_DIR / \"train.parquet\")\nnb_batches = len(os.listdir(DATA_DIR / \"train.parquet\"))\nprint(f\"Nb of train batches {nb_batches}\")","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:43.448410Z","iopub.execute_input":"2024-11-23T01:14:43.448868Z","iopub.status.idle":"2024-11-23T01:14:43.462550Z","shell.execute_reply.started":"2024-11-23T01:14:43.448828Z","shell.execute_reply":"2024-11-23T01:14:43.461238Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_parquets = [f\"train.parquet/partition_id={i}/part-0.parquet\" for i in range(N_PARTITION)]\ntrain_parquets","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:43.465182Z","iopub.execute_input":"2024-11-23T01:14:43.465648Z","iopub.status.idle":"2024-11-23T01:14:43.474404Z","shell.execute_reply.started":"2024-11-23T01:14:43.465600Z","shell.execute_reply":"2024-11-23T01:14:43.473215Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We have in total 10 dataframes for the training data, we could look to the first one to check available fields.","metadata":{}},{"cell_type":"code","source":"# read train dataset, let's read only the first dataset and inspect provided fields\ntrain_parquets = [DATA_DIR / f\"train.parquet/partition_id={i}/part-0.parquet\" for i in range(N_PARTITION)]\n\npl_train = pl.concat([pl.read_parquet(_f) for _f in [train_parquets[0]]])","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:43.475846Z","iopub.execute_input":"2024-11-23T01:14:43.476619Z","iopub.status.idle":"2024-11-23T01:14:46.349421Z","shell.execute_reply.started":"2024-11-23T01:14:43.476565Z","shell.execute_reply":"2024-11-23T01:14:46.348195Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Fields description\n\nThe training set, contains historical data and returns. \nFor convenience, the training set has been partitioned into ten parts.\n\n- date_id: integer value with the date of the data\n- time_id - integer value to define the dte of the data during the day\n- symbol_id - indentifier of a unique financial instrument.\n- weight - the weighting used for calculating the scoring function.\n- feature_{00...78} - Anonymized market data.\n- responder_{0...8} - Anonymized responders clipped between -5 and 5. \n\n\nThe field that we want to predict is the **responder_6**","metadata":{}},{"cell_type":"code","source":"pl_train.shape","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:46.350856Z","iopub.execute_input":"2024-11-23T01:14:46.351189Z","iopub.status.idle":"2024-11-23T01:14:46.360138Z","shell.execute_reply.started":"2024-11-23T01:14:46.351155Z","shell.execute_reply":"2024-11-23T01:14:46.358807Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pl_train.head()","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:46.361854Z","iopub.execute_input":"2024-11-23T01:14:46.362685Z","iopub.status.idle":"2024-11-23T01:14:46.390695Z","shell.execute_reply.started":"2024-11-23T01:14:46.362630Z","shell.execute_reply":"2024-11-23T01:14:46.389395Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# we can easily get insights about the features with the describe function\npl_train.describe()","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:46.392428Z","iopub.execute_input":"2024-11-23T01:14:46.392792Z","iopub.status.idle":"2024-11-23T01:14:50.330463Z","shell.execute_reply.started":"2024-11-23T01:14:46.392756Z","shell.execute_reply":"2024-11-23T01:14:50.328976Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"When calling the describe function you have the following meaning for each row:\n\n- **\"count\"**\nThe number of non-null (non-NaN) values in the column.\n- **\"null_count\"**\nThe number of null (NaN or missing) values in the column.\n- **\"mean\"**\nThe average of all the values in the column (for numeric columns).\n- **\"std\"**\nThe standard deviation of the values (measure of spread) for numeric columns.\n- **\"min\"**\nThe smallest value in the column (for numeric, date, or categorical columns).\n- **\"25%\"**\nThe first quartile (Q1), representing the value below which 25% of the data lies (numeric columns).\n- **\"50%\"**\nThe median (Q2), representing the value below which 50% of the data lies.\n- **\"75%\"**\nThe third quartile (Q3), representing the value below which 75% of the data lies.\n- **\"max\"**\nThe largest value in the column (for numeric, date, or categorical columns).","metadata":{}},{"cell_type":"markdown","source":"## Lags","metadata":{}},{"cell_type":"markdown","source":"Together with the training data we're provided with also a lags.parquet file. \nAs it is explained in the description, they are the: ***values of responder_{0...8} lagged by one date_id. The evaluation API serves the entirety of the lagged responders for a date_id on that date_id's first time_id. In other words, all of the previous date's responders will be served at the first time step of the succeeding date.***\n\nFirst of all what is a *lag value*?\nA lag is the time difference between two observations in a sequence. In this case the lags are provided at date_id-1, this means that for every prediction of  a specific date we will the responders values of the previous date.\nIndeed it could happen that there's a correlation between the responders as it could happen in a real market.","metadata":{}},{"cell_type":"code","source":"lags_parquets = DATA_DIR / f\"lags.parquet/date_id=0/part-0.parquet\"\nlags = pl.read_parquet(lags_parquets)\nlags.head()","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:50.332249Z","iopub.execute_input":"2024-11-23T01:14:50.333161Z","iopub.status.idle":"2024-11-23T01:14:50.349980Z","shell.execute_reply.started":"2024-11-23T01:14:50.333106Z","shell.execute_reply":"2024-11-23T01:14:50.348702Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"lags.shape","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:50.353275Z","iopub.execute_input":"2024-11-23T01:14:50.353697Z","iopub.status.idle":"2024-11-23T01:14:50.360984Z","shell.execute_reply.started":"2024-11-23T01:14:50.353659Z","shell.execute_reply":"2024-11-23T01:14:50.359848Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"In the competition only a small sample is provided and if we want to play with the lags to train our model we will have to create the lags on our own from the traning data. Indeed in the training data we have the values of the responder at the date_id for a specific time_id and while when testing we will have the values at the date-1 being the lag_1 (i.e. from the previous date).","metadata":{}},{"cell_type":"markdown","source":"## Test dataset","metadata":{}},{"cell_type":"markdown","source":"Let's now look at the test set.","metadata":{}},{"cell_type":"code","source":"test_path = '/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet/date_id=0/part-0.parquet'\ntest_pl = pl.read_parquet(test_path)","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:50.362606Z","iopub.execute_input":"2024-11-23T01:14:50.363160Z","iopub.status.idle":"2024-11-23T01:14:50.394050Z","shell.execute_reply.started":"2024-11-23T01:14:50.363106Z","shell.execute_reply":"2024-11-23T01:14:50.392724Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_pl.shape","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:50.395580Z","iopub.execute_input":"2024-11-23T01:14:50.395960Z","iopub.status.idle":"2024-11-23T01:14:50.403064Z","shell.execute_reply.started":"2024-11-23T01:14:50.395925Z","shell.execute_reply":"2024-11-23T01:14:50.401713Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_pl.head()","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:50.404709Z","iopub.execute_input":"2024-11-23T01:14:50.405140Z","iopub.status.idle":"2024-11-23T01:14:50.430515Z","shell.execute_reply.started":"2024-11-23T01:14:50.405105Z","shell.execute_reply":"2024-11-23T01:14:50.429263Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The test set doesn't contain all the columns that are inside the train set. Indeed it represents as the data will be served at inference time. Data about the responders will be provided in a separete dataframe *lags.parquet* as explained in the previous section.","metadata":{}},{"cell_type":"code","source":"print(test_pl.columns)","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:50.431994Z","iopub.execute_input":"2024-11-23T01:14:50.432457Z","iopub.status.idle":"2024-11-23T01:14:50.442886Z","shell.execute_reply.started":"2024-11-23T01:14:50.432408Z","shell.execute_reply":"2024-11-23T01:14:50.441755Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(pl_train.columns)","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:50.444534Z","iopub.execute_input":"2024-11-23T01:14:50.445016Z","iopub.status.idle":"2024-11-23T01:14:50.454642Z","shell.execute_reply.started":"2024-11-23T01:14:50.444960Z","shell.execute_reply":"2024-11-23T01:14:50.453404Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Csv additional files","metadata":{}},{"cell_type":"code","source":"features = pd.read_csv(DATA_DIR / \"features.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:50.455979Z","iopub.execute_input":"2024-11-23T01:14:50.456382Z","iopub.status.idle":"2024-11-23T01:14:50.492920Z","shell.execute_reply.started":"2024-11-23T01:14:50.456321Z","shell.execute_reply":"2024-11-23T01:14:50.491737Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"features.shape","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:50.494283Z","iopub.execute_input":"2024-11-23T01:14:50.494662Z","iopub.status.idle":"2024-11-23T01:14:50.501563Z","shell.execute_reply.started":"2024-11-23T01:14:50.494628Z","shell.execute_reply":"2024-11-23T01:14:50.500423Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"features.head()","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:50.503152Z","iopub.execute_input":"2024-11-23T01:14:50.503587Z","iopub.status.idle":"2024-11-23T01:14:50.538704Z","shell.execute_reply.started":"2024-11-23T01:14:50.503533Z","shell.execute_reply":"2024-11-23T01:14:50.537578Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"responders = pd.read_csv(DATA_DIR / \"responders.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:50.539994Z","iopub.execute_input":"2024-11-23T01:14:50.540343Z","iopub.status.idle":"2024-11-23T01:14:50.559171Z","shell.execute_reply.started":"2024-11-23T01:14:50.540307Z","shell.execute_reply":"2024-11-23T01:14:50.558162Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"responders.shape","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:50.560469Z","iopub.execute_input":"2024-11-23T01:14:50.560767Z","iopub.status.idle":"2024-11-23T01:14:50.567629Z","shell.execute_reply.started":"2024-11-23T01:14:50.560737Z","shell.execute_reply":"2024-11-23T01:14:50.566425Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"responders.head()","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:50.568951Z","iopub.execute_input":"2024-11-23T01:14:50.569266Z","iopub.status.idle":"2024-11-23T01:14:50.584620Z","shell.execute_reply.started":"2024-11-23T01:14:50.569234Z","shell.execute_reply":"2024-11-23T01:14:50.583546Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"In the description provided by the competion host we don't have any further information about these additional .csv files. Their usage is not mandatory, from a first look we can see that ***features.csv*** establish a sort of relation of categories called \"tag_X\" with X=0,...,18.\nSimilarly for the responders we have ***respoders.csv*** file with a relation with \"tag_X\" with X=0,...,4. We don't know if the tag encountered in the responders are the same as in the features.\nIn a first baseline model I'm not going to consider those files but it is interesting starting to have a look and think to include somehow them.","metadata":{}},{"cell_type":"markdown","source":"## Sample Submission","metadata":{}},{"cell_type":"markdown","source":"The sample submission just provide an example of a the file format we will have to produce with our model.","metadata":{}},{"cell_type":"code","source":"sample_submission = pd.read_csv(DATA_DIR / \"sample_submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:50.586117Z","iopub.execute_input":"2024-11-23T01:14:50.586558Z","iopub.status.idle":"2024-11-23T01:14:50.603193Z","shell.execute_reply.started":"2024-11-23T01:14:50.586515Z","shell.execute_reply":"2024-11-23T01:14:50.602144Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sample_submission.shape","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:50.604643Z","iopub.execute_input":"2024-11-23T01:14:50.604935Z","iopub.status.idle":"2024-11-23T01:14:50.611710Z","shell.execute_reply.started":"2024-11-23T01:14:50.604906Z","shell.execute_reply":"2024-11-23T01:14:50.610628Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sample_submission.head()","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:50.613271Z","iopub.execute_input":"2024-11-23T01:14:50.613725Z","iopub.status.idle":"2024-11-23T01:14:50.636947Z","shell.execute_reply.started":"2024-11-23T01:14:50.613677Z","shell.execute_reply":"2024-11-23T01:14:50.633331Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Train data analysis","metadata":{}},{"cell_type":"markdown","source":"In this section we're going to get some insights from the data provided in the training set.","metadata":{}},{"cell_type":"markdown","source":"## Responders","metadata":{}},{"cell_type":"code","source":"# we're going to read all the data from the 10 train sets. \n# Reading all the columns at once will lead the notebook to got out of memory, \n# we therefore have to load only some fields. We will start with the main columns\n# and the responders\nfields = []\nfields = [\"date_id\", \"time_id\", \"symbol_id\"]\nfields.extend(responder_cols)\n\npl_train = pl.concat([pl.read_parquet(_f, columns=fields) for _f in train_parquets])","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:50.638426Z","iopub.execute_input":"2024-11-23T01:14:50.641307Z","iopub.status.idle":"2024-11-23T01:14:58.144578Z","shell.execute_reply.started":"2024-11-23T01:14:50.641247Z","shell.execute_reply":"2024-11-23T01:14:58.143504Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Time per day and symbols columns","metadata":{}},{"cell_type":"markdown","source":"Let's have a look to the timestamp and the symbols columns","metadata":{}},{"cell_type":"code","source":"time_id_per_day = pl_train.group_by(\"date_id\").agg([pl.col(\"time_id\").n_unique().alias(\"nb_times\")])\ntime_id_per_day = time_id_per_day.sort('date_id')\n\nprint(\"Number of dates: \",  time_id_per_day.select('date_id').n_unique())","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:14:58.145885Z","iopub.execute_input":"2024-11-23T01:14:58.146203Z","iopub.status.idle":"2024-11-23T01:15:00.642147Z","shell.execute_reply.started":"2024-11-23T01:14:58.146172Z","shell.execute_reply":"2024-11-23T01:15:00.640872Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"x = time_id_per_day[\"date_id\"].to_numpy()\ny = time_id_per_day[\"nb_times\"].to_numpy()\n\nplt.plot(x, y)\nplt.xlabel(\"Date ID\")\nplt.ylabel(\"Nb of times\")\nplt.title(\"Nb of times per day\")\nplt.grid()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:15:00.643689Z","iopub.execute_input":"2024-11-23T01:15:00.644159Z","iopub.status.idle":"2024-11-23T01:15:00.963293Z","shell.execute_reply.started":"2024-11-23T01:15:00.644107Z","shell.execute_reply":"2024-11-23T01:15:00.961826Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"Unique number of dates:\", np.unique(time_id_per_day[\"nb_times\"]))","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:15:00.969842Z","iopub.execute_input":"2024-11-23T01:15:00.970250Z","iopub.status.idle":"2024-11-23T01:15:00.977115Z","shell.execute_reply.started":"2024-11-23T01:15:00.970210Z","shell.execute_reply":"2024-11-23T01:15:00.975919Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Filter for rows where column 'X' is equal to a specific value, say 2\nvalue = 968\nfiltered_df = time_id_per_day.filter(pl.col(\"nb_times\") == value)\nfiltered_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:15:00.978324Z","iopub.execute_input":"2024-11-23T01:15:00.978764Z","iopub.status.idle":"2024-11-23T01:15:01.001803Z","shell.execute_reply.started":"2024-11-23T01:15:00.978722Z","shell.execute_reply":"2024-11-23T01:15:01.000442Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"From the graph above we can observe that every day we have 849 time_id before the date_id=677. After this date id we have 968 time_id per day.","metadata":{}},{"cell_type":"code","source":"symbols_per_day = pl_train.group_by(\"date_id\").agg([pl.col(\"symbol_id\").n_unique().alias(\"nb_symbols\")])\nsymbols_per_day = symbols_per_day.sort('date_id')","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:15:01.004080Z","iopub.execute_input":"2024-11-23T01:15:01.004599Z","iopub.status.idle":"2024-11-23T01:15:04.813871Z","shell.execute_reply.started":"2024-11-23T01:15:01.004544Z","shell.execute_reply":"2024-11-23T01:15:04.811373Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# concerning the symbols these are different \n# for every day with an increase for newer dates\nx = symbols_per_day[\"date_id\"].to_numpy()\ny = symbols_per_day[\"nb_symbols\"].to_numpy()\n\nplt.plot(x, y)\nplt.xlabel(\"Date ID\")\nplt.ylabel(\"Nb of symbols\")\nplt.title(\"Nb of symbols per day\")\nplt.grid()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:15:04.817343Z","iopub.execute_input":"2024-11-23T01:15:04.818168Z","iopub.status.idle":"2024-11-23T01:15:05.485606Z","shell.execute_reply.started":"2024-11-23T01:15:04.818109Z","shell.execute_reply":"2024-11-23T01:15:05.482093Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We can observe that for each date we don't have a constant number of symbols but they vary day by day.","metadata":{}},{"cell_type":"markdown","source":"Let's now have a look to the responders time series","metadata":{}},{"cell_type":"code","source":"pl_train.head()","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:15:05.493999Z","iopub.execute_input":"2024-11-23T01:15:05.498633Z","iopub.status.idle":"2024-11-23T01:15:05.527316Z","shell.execute_reply.started":"2024-11-23T01:15:05.498463Z","shell.execute_reply":"2024-11-23T01:15:05.522959Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, axes = plt.subplots(4, 2, figsize=(14, 20))\naxes = axes.flatten()\n\nfor i in range(0,8):\n    col_name = f'responder_{i}'\n    resp = pl_train.group_by(\"date_id\").agg(pl.col(f\"responder_{i}\").mean())\n    axes[i].plot(resp[col_name])\n    axes[i].set_title(col_name)  # Set title as the column name\n    axes[i].grid(True)  # Add grid lines\n    axes[i].set_xlabel(\"time series\")  # Optional: Label for x-axis\n    axes[i].set_ylabel(f\"responder_{i} value\")  # Optional: Label for y-axis","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:15:05.533138Z","iopub.execute_input":"2024-11-23T01:15:05.535972Z","iopub.status.idle":"2024-11-23T01:15:37.306004Z","shell.execute_reply.started":"2024-11-23T01:15:05.535639Z","shell.execute_reply":"2024-11-23T01:15:37.304655Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"From the previous graph we can observe as responder 3,4,5 has sharp negative/ positive variations. If we would like to get some additional information we could calculate the cumulative sum of each responder per datetime.","metadata":{}},{"cell_type":"code","source":"# let's now plot the cumulative sum of each responder (anonymized responders clipped between -5 and 5.)\nplt.figure(figsize=(12, 8))\nfor i in range(0,9):\n    resp = pl_train.group_by(\"date_id\").agg(pl.col(f\"responder_{i}\").sum().alias(\"responder_sum\"))\n    resp = resp.sort('date_id')\n    \n    x = resp[\"date_id\"].to_numpy()\n    y = np.cumsum(resp['responder_sum'].to_numpy())\n\n    plt.plot(x, y, label=f'responder_{i}')\n\nplt.xlabel(\"Date ID\")\nplt.ylabel(\"Cumulative sum responder\")\nplt.title(\"Cumulative sum responders/date id\")  \nplt.grid()\nplt.legend()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:15:37.307451Z","iopub.execute_input":"2024-11-23T01:15:37.307770Z","iopub.status.idle":"2024-11-23T01:15:59.825695Z","shell.execute_reply.started":"2024-11-23T01:15:37.307739Z","shell.execute_reply":"2024-11-23T01:15:59.824392Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We can observe the highly negative trend of some responders. The one we would like to predict is the responder 6. Let's now calculate the correlation matrix of the responders.\nThe correlation matrix is a table in which every cell contains a correlation coefficient, where 1 is considered a strong relationship between variables, 0 a neutral relationship and -1 an inverse relation, if a variable increases the other negative correlated decreases.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(8, 8))\nsns.heatmap(pl_train[[ f\"responder_{target}\" for target in range(9)]].corr(),  annot=True, square=True, cmap=\"jet\")\nplt.xlabel(\"responder_0  -  responder_8\")\nplt.ylabel(\"responder_0  -  responder_8\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-11-23T01:15:59.827059Z","iopub.execute_input":"2024-11-23T01:15:59.827419Z","iopub.status.idle":"2024-11-23T01:16:04.162253Z","shell.execute_reply.started":"2024-11-23T01:15:59.827344Z","shell.execute_reply":"2024-11-23T01:16:04.161123Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"From the correlation matrix we can observe as responder 6 is strongly correlated for each time id with responder 3 and weakly with  0,1,2.\n\nA proper model should therefore not only rely on the features but also on the other values of the responders.\n","metadata":{}},{"cell_type":"markdown","source":"## Features","metadata":{}},{"cell_type":"markdown","source":"Let's now have a general look at the ***features*** distribution","metadata":{}},{"cell_type":"code","source":"fields = []\nfeature_cols = [f'feature_{x:02}' for x in range(79)]\ntarget_column = \"responder_6\"\nfields.extend(feature_cols)\nfields.extend([target_column])\nfields.extend(['weight'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:16:04.163827Z","iopub.execute_input":"2024-11-23T01:16:04.164178Z","iopub.status.idle":"2024-11-23T01:16:04.170454Z","shell.execute_reply.started":"2024-11-23T01:16:04.164144Z","shell.execute_reply":"2024-11-23T01:16:04.169224Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"summaries = []\nfor batch_file in train_parquets:\n    # Load batch data\n    pl_train = pl.read_parquet(batch_file, columns=fields)\n    summaries.append(pl_train.describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:16:04.171880Z","iopub.execute_input":"2024-11-23T01:16:04.172327Z","iopub.status.idle":"2024-11-23T01:18:27.670579Z","shell.execute_reply.started":"2024-11-23T01:16:04.172279Z","shell.execute_reply":"2024-11-23T01:18:27.669389Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"len(summaries)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:18:27.671912Z","iopub.execute_input":"2024-11-23T01:18:27.672259Z","iopub.status.idle":"2024-11-23T01:18:27.679731Z","shell.execute_reply.started":"2024-11-23T01:18:27.672226Z","shell.execute_reply":"2024-11-23T01:18:27.678442Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We have the stats for each train set","metadata":{}},{"cell_type":"code","source":"summaries[1]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:18:27.681496Z","iopub.execute_input":"2024-11-23T01:18:27.682245Z","iopub.status.idle":"2024-11-23T01:18:27.700422Z","shell.execute_reply.started":"2024-11-23T01:18:27.682188Z","shell.execute_reply":"2024-11-23T01:18:27.699218Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Metric","metadata":{}},{"cell_type":"markdown","source":"Submissions are evaluated on a scoring function defined as the sample weighted zero-mean R-squared score (R2) of responder_6\n\n![image.png](attachment:d7204e3d-fc74-4e82-adc3-5831723ae5c2.png)\n\nThe metric should be interpreted as follow:\n\n- R2 =1: Perfect fit (all predictions match true values, accounting for weights)\n- R2 =0: Predictions are as good as using the weighted mean\n- R2<0: Model performs worse than using the weighted mean","metadata":{},"attachments":{"d7204e3d-fc74-4e82-adc3-5831723ae5c2.png":{"image/png":"iVBORw0KGgoAAAANSUhEUgAAASMAAABfCAIAAAAs1TZ0AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAboElEQVR4nO2df0xbV5bHz6y7L2L1ouw6E8keViaM7KQyUMUQhSUygxcPbDzrGWfoElEZEREFkUmGtATID2gbQgMhKYQUmkxYR6AgUNhYZeqpW8/AQqFmcJwFjIKxEmxtghViq8hvivy2iKs+sX/ww+anMRhDw/0oipL347577ff1Pffcc8/9ydTUFGAwmA3m7za7AhjMtgArDYMJBlhpGEwwwErDYIIBVhoGEwyw0jCYYICVhsEEA6w0DCYYYKVhMMHgjc2uAAbz4wONdt+/99BgtTtokvtmbEpGplxIrnwL7tMwGD8Z05Vc/2bvyYq6xs811Sf2O3TXc/NVFrTyTbhP23gGa47laxw+vggfsOUfq/NERIBq5DeMva2yxnzoUq6EvVlVCCZ0Z2XJY1FuXiKXtcTZEZ2my+J0G51JKTyCl5h78lHnhx3NDe3vlMtW6New0jYeoUzO16ksCICIzq4ulXMWXsBM/40QACDkHqdc3zpfPLeaB/SGJ06aAQCg9Lq+06K4kODWfKZ6zrarhSo4cfd1kBmyqss+dcou58Su0BgyXpnUcfbdK5OfXJYtITYWEAztGKWn/0f8PGwvC8wv7Q4GBEspc4YpTBAY+fysXCqWSMWynAcjftw3+fLRg7LMRIlULFFc+dq9YfVbqQrDquxEZUXfpjw8wEwOP8g/IpUlJsuU1/WuH1a81v3oZppMeXdocqlzr6wj7rnbh+6mSqSJuZ+7ViwPj9OCAk9x+XQsCQATFtVH9dZVW5JEaGzapbuNlxLDWHRX22N6A6u4NGiw/oqaSsjOjvYx4N/60OaG/LMNruT36zS3s/f2lp261uFglr+cjM06LXary26bFn9bJJfPI2e7L2uP0QFk3JH4lXt8rLQgwZblXZSyAQDZmq7cW+LbWx6Cm1z4cU4saWrvpDaqesvg1Ko0I+GK4/E/fp3pq4s0O45XVOdKOCRfcfV2YdxIdVGDbYVbyPj0d/iU9o56ZAVBjulUGjs7NvuM1JdpvaZeGLMm3I/K06RiiVQsVZb/1V9rzPVFvuKsDwslwEz+T8VvJLKCL4L60A3D7fpu/oHvXS5fX4Lri/xESWr5o6VMyKmpqSnHF4WpR3Iah7/3/XjcpwURMja3MDWMBcA4tZUVbf51UGz5perfHwym9xEZ/qyniIhfHH4NHCEAQLJ3zT8Qwmb76qrZ8YkxBNX5ZfdSdjuyNpQ1oNTq60rBKjxVWGlBhYjKvKzkEwBAdd+8pllpnLAYNk8QGkQrbuJxl5EG/qGY10Noa2OXKJoPtOmbvomFZ2h9zbWB2NKSVEEIAGNrudfhWLEkrLQgQwgyCrOiCACge2tL1PYAFUuZWypPvP0r6dET7zeY5v0A2ztuXy1rHpw5NtJSdOzffnXizioGikPGPhrYAj53qZPI3lF1Ll32q6PHzta0jc47RVs0VTdqu0aXui2IUD31548flf762ImrmvkuKDSiq71+S7fS6MsDZ5+ADbSl/+n8w3ZN+Z/I35fM9mbjQ4+eT+5csSCstKDD4qUV5sTtAgBkri9TPVvflDYAAIw0F5cYeOfqGnMF9q6GyvuWuTN0272K5vaOrqc0AABja9UYHQhZdbo+X++Zw2alAPaGhy1xbtxY9UEjOlLWVCKFQU15jc7LELa3VNS06HSPXq6/WWsHDdbm37H9S2Hd3RS2tb32ZovXLxqlu3lLrdV1Dy/qppZkb3gYADVsc3oOjRurPqgd/M52/0rB+YKCd8+d/V1u/Ys9nJXtDay0zYAju/ieeNoR2VBW279O571dXa5hZ11KjdxFU98BMJTLNateNNT/BAFwIt/kAACw+PIM8ZJxD4t5YXcCELvZi21HZKiteRZfmJvMIyfcbgDkotxzJ8eMBjsAIYh4c32NWg+MRVWh359zKWU/e5KmAdArp2vuJG0ymREAT7hvdZY4m/1PJMCrUY9taG6qabEjymbq7zUZek39JovZTu8OXbLv94CVtjmwJfkXp4NF7JryO8b1aM2s1bhiFQlsALu+8zkAEREdMes4eT7QRwOQwkj+zAGuND/jIAF7uD+bpze7tihdUaDxGmnQNIUA2Lt3L3reePsfe7hyBZ8AZOh6TANwI4RzHR9tGhhmAHgHojzuB2RuKjj2TrGfHqC1g4y6VpY07SAJjKWrxwlAxhwQzJ0c7LMgALZXnZdqvhf/uHMnAD32tznbI/J0o/7r/17w5w+pi0J/5oOVtlmQcacL08IJAHDoKsvb1/4a/kyaU5ohIgCs7e1WBoiDsrmoKYdpyMEA8WZElMc5Ru4NZbPfFM77BUaukW+JvVECr5EGciMELIJY3AESEb+9lCPfA0B3txlpYHESkkVzdw0ODC16j5Fr1EUKRHt99SEIITSxuj9oRZM7LPHihdQwFiCTrtUJsOtQUuzssxm7+QkFQEQdiFix+d7tJQkWIDS5Tisfxz1uHiHCrEJl35l6K6L6OoyUVLY2Jx97fywbABhbZ6cdgIxLPDT7WtHmQRsAhEWJvEqmX4zSUYkR86YLCNEZVd3Ccn8AANixeFohhBcXCwBA97QbaACeNHn/7CnGbn5CAxBRUV7vMZAJ5+sSfLXC0XI2vcZXPLwXgsy6ugzekqeIUFEcAADq6zBSAOx4aczcD82Yqd8JwPLq9mGZ5nvOAgDA5Mri9g1W2mZC8FOzJOrzTw5dzFujzDw813eNAoQcSPD8fg/1P0UAnCiR1xs5MdBvFUTn+RyjEMQOAIDlf8lpQ7cJAYQdjveE1c68x/zot/ye9+OmVLen+HvTiqChb4wUAHlY7FkDQT8ZGGYAwoVRq/+4pyW2Y73VwdbjZkL31Nw28rKK8xPWPWdFDVlGGAC+KGZORHbLs3GAEH4033MZMn3TFyaW7Jn9P+Nsqyx4Nyv9xI0F00HEzhACGJpeTmmMtX8IAbBjYjylz7zHPNHse4ysLWXvnjmRfqbWML7eBvqNfWBwHCDkgLfsZ43bAzO/Dss23wuEEAMEsXOdU5nrVBptbqk8f+qE4tdHFcfPljQY17kKa1uBnquLbpj2vVecsT8AkR/ucTcAkFyOR7PfUi4GYBd3t6d4uuvPA/uTPLGw1qaa/phLHyv5I60P58+AEbs5bADa48ZcWPu/uWgAVliY14BscMDq/R4jU71qVFb6kWL3M80fe4IeHe2iXAwAJ2zfXPMZm3mIBiCiZh0kyzffA3K5XADknn9a55e0HqUhc+2VBlr24d06zed1pdIdhvqiU1c7gh0E+yNl3Fj1QRNKLb4YoEVfu9k7CQBEUbNvNLIOWimYXfY2fciifvBSmjEXC4uMWoswTcIeHrCgEA53fkV+FsoFQK6xZb5PFkESADCJ5ublxjq+NHo7G+iuLx3/kiYinpiGgeT6mG3aAP6B2AkAXs4TulfTNjpt3JIAPpo/h2uMQgB7eUvNK/rDOpQ28bhFa+prbR+kAVjsyIycd/hA9dQ9sPi+dbuD7M1XrvVF5JdOR2YFAjJekcAGZFKrOp2IQY7Omis6iAwnYEzforXRDKIGNSXXjTE5mZFzj2QJUnIUYcjU1kORsb9csMyU/SY/jAUvbCNLP484JJdygLF1aS3UBO0waUpyK7ooAFbErLVGxGScku+hDV0D9J74pLcC1M7VI5QlhRPgNGqNFJqgrLrKs1d0DgaAJ4qZNp5XbP4cL2x2YHH2Raz3B3E9HpG/BxYgyj5rX/AEPyfBRo28pMFX9pLtDWW4VXj/B2V1nti/b4+mHIjkspfR5i7xxepibu3DrjunZNcQGS4+/lFFSqi9pb6+pSlfUYNIvvid9yrSvDMksNhhoYB6OzopMi7h0MJyww9F7VJrrRYrE7vUUmIi+nTZVaL2fmthqhqRvEMSAQdG7cATzrzHQLB5HKA72ow0Vy6NXN10eSBh8bOuX95xp157LV3LkNy34qPCCasFsd86MNM9rdz8aRib2UoDWxrHX/K0P6xmvcFyTLpGhh1zCwrcX1yQiSWpN/vXU+Rrz+SLBzlHlBU9fq9EcX1RqDj1wBHw+vRcTxUrSnu+n5p0OOYvQ57sq0gVS3MevFxNOe7OYoVYIk29PTTv6F8KE6XKT4emptyOV98td29Q+KH/ZqpULFFc+av3EpgVmj81NTU15Xh4SipNrR5adMJv1uURIdg8AWf2t2D88aMhBDypPPh2wjJQvU1VDX6tudxwKH1NQfOO4x/lxPlpjNA99Q29vIR4H4EIfjNh+qaHYh+WxoTYmyub5sdgEtFyWRhYWtsXhUGPmxounkg/V2+eCx2kH7f10sDiJSULvWvd2WFCofHJ+8HcUNkSxJhjZNO8n5X+uxpP8A0ydXSOAeyJl3uvPFqp+QAAI+3tZuDLZcKFJ/znjfO//KVhQbDpdEfPIgiC5PKEMYd/IZcnCnYtcbMXyNzc2IV4ae8pV0paEhRoyv5iyNTVptHq7fTBnOMg2iLLPpCtqajSFHP+7nRoyOqhB5uKrukc/OyE0EDXiXa+oomoGBGtrxkWKjIWjFX2K9IOaq63asxpOZFeVR7R1KqMdmDpDfbMyP0AACMtDw00kPHpb3tbWYzT8S1iv3UozKkrd4qzAvC6rhKqtba2y4ZgrHv4dGw0CwCo1v9qp4CITlVGe3/2KzcfWbQ6GxlfnLJ+0xHgjdzb1Rnfux1d9SUaGwCEKQrPSTkEC2DS7Rpzjgx2axvKWpoak04WnkvhLzf8QoP1NzVIfqn6jGhTR2jjunffruxngGDzuGAPftaNlaC6qz5QE8rqC4f9+YiQ09BcU9VkdCCIjI/3EcS6BvbEy6Wa+y1XbgqkWacXv1Bs+Wnll6fqVdrUT1I83anL6QQAIjQ2jgcAgJ6rb6ptiC2+eDpx3o8ai5+sEHe11L9PR/w2JzvwlV8OhnJ8iwCAGxu7jwUA4NBVqnoRKco+lzLfKFix+Q5tfcu4KPekODDv9IwV+U1pokQqlsgKvlq44Nv914pUiVQskWXWLZknaGrKpb+izCz/egusgf/B/cI68upb99TUVE+ZQiyRivODmw9gOSat908rlBWPfKc0mJx0uxwvnvZ3fvXwP8vylTKpWDKdEGGV46XA8+qLwt/I8z/zevqrz/MTkzNv/tU1OTX56lFjQapUnFr44OlyKQCCj7uzWCFOvdw6Mjn1g2vws8vKZOmR03V9fg0UX35eIE8t+ipgr88i3yNroWFDxiqSeLoGO7I2321Jrk5bYMAgW/PVxh0nKy5I2ABA99S3hCgzNisFKIsMW7bf3USorhvFKgsNliKZdq1l7I8PvOm4OrjyS6VjhUUf1e+ryowMmTlylapW3UqXXQFyDz9GmveHNFmkj/FFMCET3is7U3NXdeZoOUPs5kUkZH/8jkLEXv24hraoipsm08qurjtIbo5VePlZ5O7pDxHZzEMIQufZuYY7tcNJxR/OTr8O91nc0k3LtLtFcVqGWcIk6ZqGKbMvBzdOGjzrayFkZGZFxZ56wxM6cjqoksWOyyyOy9y0CvlmlzDt/eq0td5ND5mQouxjecBmO2F182mInnYxsdhhYfMePdJypaTdsXe08nwHTALApNvxfGeaMnC1ez3giLMuiTe7EuuEEMizBb4ve00gY5VnAl3mKpRmNxrsAABsyam393sdH9fdrjXRCMy9Xgu/d8n2biErAoPZKqykNETZzUbd/Xq1mSEj0/I/PDk/pmGX7MZfZGt8LGOqOlbQ4meIpOBkXZ1y6SVJGMwWZ4HSkOHGUWkleAdmklHKG9XKOE5AR18sUe6Dr7JWWFzHIhY/jwjxvw5+pXnDYDaMBUoj4s5/fiOZAABgaIdRU1VZbxhUq5oF+3PEfrhuVgNBkEuoafNgfCyZnweL2FJ1x2x9lrceWST3sPLqTvpErtqqKSvac/cPr7PlRrddOVaiX73U2Ck3H+aKfFwU/6+/XF+tMD9u9F//99y/fXhECKE0gadueI7MGk1/Wk70ZkdabRhkUslXSYEu1PuDxmxzfPkeWWzuPwIAAGUddkJ0oCZP1+YRWT5Jy7L84N/lGMwG4dPLT8zkBmLsIy8BAqU0nx6RJW4h1uIRwWC2BouUxqDZvFsz7AACAAHQL547IZYDADBmM09wInnri3vaah4RDGYj8bk+jSRnF/kOD5imzT1rS1mJNlBbNwQWRNlsVovFbOx49L8IAMD5uFVvsT6zWW3OrRXaj9lmvOF4ZqOR+0WfDQEAoFeD+v5/DttJsLl8zvT+onFpmXHGGgMFyKT5bDA+K3SoRY9i3gvEkp2AwwzdLyzSjgNBEAAESRIwMXT/Rr6KAWBEuZ+Vyrdg+DFmHaDR7vv3HhqsdgdNct+MTcnIlG/VzBo/KUiW9bGA8A7hZxBiIs48+DhlLiSEsmjVD9t6LMNOigZOnDLvw4zNXYiGwQCM6d7/yJRwIT8plED2jqoPKrROXkZVdZZwK45KfjI1NbXZdcBsEJS2IP26Ca0rUIYQ5jZWp+zxfWHwGWk4ld7gjD5d/UkKDwCQvkzxYQeKzdOUy7ZgN4Czhb/GsCW/Eat6OygA4MiuVmfHLE60xgAAQgwAA5MTFO1yvRq1Dz8bMPSYrNT0eMLS1u5MSQt0/pKAwAKCoR2jMwNw4udhe1lgfml3MLDpKTYWg/u01xu6/9apdzVOAODKS+vyYlf7Y8/Q1vamT++p+8cAwpWNdZnrTSy6IdAOG7UznDftUABL7bEzapcoR31TsUUyx3iD8/K/3pDR2bN7R2n92TuKRQqSsz9RVWdFkWDXtz7bwCquA5LLn5UZgLXH6AAy7kj8FpQZYKW9/oQIswqVkQQAUF23KrVOnzd4sUuYUXY5jedsa9/yianHdCqNnR2bfUa6NYWGlbYNIPjKiydFJADQxqoy9ep2Up+FFGXlygh9e/+WXn/k1N6qHwzPrLgsW+XewsEHj9O2CU7txVPXjTQAIciovpvpV4YM2vGM2hHOWy5P+WaDrA35RYPxpSWpgmVy628FcJ+2TeDIz59NYgMAsjaVqUx+BcyQ3P1bVmZA62uuDcTOyIyxtdxbfie0TQUrbdvATjx3XsZlATD25hubsXXgKpmwaW+cPfbrX8nSC1TG+S4cythwo7LB+6BdU/4n8vclypnebHzo0fPJpber3myw0rYRZGzOhyk8AACnrrwyMDvdIXu36mK67N+OHjtT2TYvGJbub64sqZntYcaNt7OOSt8u0vrIzu/UXi1uIZWf3M+PoUwN1+4aJjzn+hsqVTpd68BsxceNVR/UDn5nu3+l4HxBwbvnzv4ut/7FnuDv1LYqsNK2FUTkycKs/QQAUPrqco1fjsilGDdWFdZRRyrU5VKw6Mrv6DxWqU19856urd30ggEAQL1/arHRiDJ+qV/poZSuRuVWlGbHclmTbgaA/ptrTmmMrc9EAZD7hDNrFM1NNS12RNlM/b0mQ6+p32Qx2+ndoZuXF3NFcIzINoPgZxRm9p2q7Z+gDbVlzW9Vp4WvuSzaUFvzTFJ8V8KBTrcbALkoF8B0l+LoNY0wQAiEUSwAAOKwMm2/qWHxbi7zcLZqh6JSLnFZQOk7zAgg/IBn6/cxU/8oACGIjJgZMkaebtSfXnPlgw3u07YfvNSLObEkAExYVGX11jVvezWu/2MPV67gE4AM3Y9pAK5INBtKQpsHbAAQFhUxY8uFCLPOKsJY09tne6CNtb87ll5lnK4EGaO8fCaeBKA6W4cQQGSydC42ZWa7+lBhzBadMPMBVtp2hCvLm95fGz1vb7WttRQiIu1yjnwPAN3dZqSBxfNkREdD/U8RACdK5JWNIpT3MxY/UjDPiekes7t3CSP/efogKTgs4hIAo+1tFgQsYYJkLt4SDQ4MTW9XvyXjwnyDrcftCTshUyHQN+3IKF77tmYhvGgRAADd026gAcLjk+cWLT4fGKQBSGG09zLGUfsrTkTW/Fhlrry0Ub6wYEeP3swACOMT5i5m7OYh2mu7+h8fuE/bliB78y01LS0szVj/Jg+0oduEAMIOx4fNxmc4ngyNMEC8KYrymkoe6TUh0aFVRNkj86ANAMJEIo9zY3qQxuLPblc/++x5xueWBittG0IbbhXen0wt9XdL+yVhrOanCIAdc2Cu/0LDFhsAhAkjvMp3Gox0nNjTI9GmpvfPnU0/XtBgWaATyuVCAMRevsfynBmk8URR82s83/jc0mClbTfQSPOVEpPw4kdKQWDez/9zjwOw2FxPBzQtFfipt/PDpmul45PmNtYb7779X3C8ND+ZZXqgMc0vkCT/gQAAT+Y0xq79s2l6kLagS+TKSxtVhUmbtLOcX2ClbS/onpqC5snjH+UnBMyDx97NBmBo19xEOGXptwMATHq0QrU1tP80NTVyzrxs1cERhYA29Y0Cd+EMGHn4yCES0GCbxkoj2m5sKMq/bUI/6kEaYKVtK5BNXXTDFJNXnObf6AxRoxS9XCw/S5gs4xPg1DaozeMA47bma3eHQ4VsFpi1asMYAtredaf4PlLmyj3i3nk4OyuepPTdZoYviV+YLZctLazOU+z7Vn3q7aOpF5pecARcFgAREe211+zyxudWJVDb+GK2Oi59uVJx9oHV7+2on95VppT2fL/CFe7BzyoKTip/I5MlKjIL7upf/TD56uu6otPKIzKpWK4suN3+YomnOj7LkYmzG1/5rPjn+YkS6bwrv9OXX2gc/n7kfqb0SNkjfxu0KWAv//YA2ZqLK/reulTnZ3cGjL35ts598NISOUg8kJEpeTdS5h+TZF6VrLg/76i+zYIiT0q5DO0YA+5MuCJluHPttpFI+aA0ZcbDQnW2DyEgopM9+w87WnVw5JKAbv90FLiSLRp+tQBsPW4HqK7K4gdvKD9+b9V5RKZhnG03ClWDcDhBFHDvnkPfbgZhgoRD66ur5tIuDDZVqU0jdmPbTBZfQINNzRYEHOlxmWcmbmXjc2uC+7TXHmRtKi4fEn14OzXMr+WfFp2qurblGQ27ZL84GHg3umOUAp4sZpetpQ2S8mbUghxOFwCECBMOsqcroarWOYAjz8mO9upUyVAegLOlcwjxlUlrj9sMKlhprzlUZ0WRekdWVV6cr/3HEU3TlPOVw/7i2dCjHr3h2XQaOmAfFsdswHxVpEwRbWr/9Koj5j+ml6gCABA/F4YRph0pmXIeIHu3qqKi2c5Oyiu7cHhRZ7y08bl1wdkNXmfQs6Z3c+vNE76vXB5SXvrwwuGgTQ0jR3t9VUN7n5MmdnH2ieLlaalJ4UuoyNF86tg94kxjtfxpWclo+o0tv40m7tNeY9DwkJN7OHGNHoPpuS8i4t83wHRcHoIrzb4hzfZ53ZLG51YG92mYHyXI0lRQ1g5h/Jj/OPuj2CQCKw2DCQbYy4/BBAOsNAwmGGClYTDBACsNgwkGWGkYTDDASsNgggFWGgYTDLDSMJhggJWGwQQDrDQMJhhgpWEwwQArDYMJBlhpGEwwwErDYILB/wOiovQ7wk3dRQAAAABJRU5ErkJggg=="}}},{"cell_type":"markdown","source":"# Predictions","metadata":{}},{"cell_type":"markdown","source":"In this section we're going to test different models to do predictions.\nSeveral methods could be used to predict responsers values as:\n\n\n- linear models:\n\nLinear Regression\nRidge/Lasso/ElasticNet Regression\n\n- tree-based models:\n\nRandom Forest or Gradient Boosting (e.g., XGBoost, LightGBM, CatBoost).\n\n- non-linear methods:\n\nSupport Vector Regressor (SVR)\nK-Nearest Neighbors (KNN)\nNeural networks","metadata":{}},{"cell_type":"code","source":"def r2_score(y_true, y_pred, weights):\n    \"\"\"\n    Calculate the sample weighted zero-mean R-squared score (R2).\n\n    Parameters:\n    - y_true (pd.Series or np.array): Ground truth values.\n    - y_pred (pd.Series or np.array): Predicted values.\n    - weights (pd.Series or np.array): Sample weights.\n\n    Returns:\n    - float: R2 score.\n    \"\"\"\n    numerator = np.sum(weights * (y_true - y_pred) ** 2)\n    denominator = np.sum(weights * (y_true ** 2))\n    r2_score = 1 - (numerator / denominator)\n    return r2_score\n\ndef evaluate_model(model, test_data, cols=feature_cols):\n    y_pred = model.predict(test_data[cols])\n    y_true = test_data[TARGET].to_numpy() \n    weights = test_data['weight'].to_numpy()  \n\n    # Calculate R2 score\n    r2_score = calculate_r2(y_true, y_pred, weights)\n    print(f\"Sample weighted zero-mean R-squared score (R2) on test data: {r2_score}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:18:27.701871Z","iopub.execute_input":"2024-11-23T01:18:27.702282Z","iopub.status.idle":"2024-11-23T01:18:27.715703Z","shell.execute_reply.started":"2024-11-23T01:18:27.702242Z","shell.execute_reply":"2024-11-23T01:18:27.714477Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Constant model","metadata":{}},{"cell_type":"markdown","source":"As a first very simple model we could consider a constant model and compare how more advanced ones compare to it.","metadata":{}},{"cell_type":"code","source":"pl_train = pl.concat([pl.read_parquet(_f, columns=['weight', 'responder_6']) for _f in train_parquets])\nsubmission_constant = np.ones(len(pl_train))\ngt = pl_train['responder_6'].to_numpy()\nwg = pl_train['weight'].to_numpy()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:18:27.717276Z","iopub.execute_input":"2024-11-23T01:18:27.717769Z","iopub.status.idle":"2024-11-23T01:18:28.513738Z","shell.execute_reply.started":"2024-11-23T01:18:27.717719Z","shell.execute_reply":"2024-11-23T01:18:28.512445Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"R2_constant = r2_score(wg, submission_constant, gt)\nR2_constant","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:18:28.515182Z","iopub.execute_input":"2024-11-23T01:18:28.515694Z","iopub.status.idle":"2024-11-23T01:18:28.967317Z","shell.execute_reply.started":"2024-11-23T01:18:28.515620Z","shell.execute_reply":"2024-11-23T01:18:28.965935Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Linear model","metadata":{}},{"cell_type":"code","source":"# Train a linear model\nfrom sklearn.linear_model import LinearRegression\n# Initialize the model\nmodel = LinearRegression()\n# the column we want to predict\ntarget_column = \"responder_6\"","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:18:28.968928Z","iopub.execute_input":"2024-11-23T01:18:28.969272Z","iopub.status.idle":"2024-11-23T01:18:28.974404Z","shell.execute_reply.started":"2024-11-23T01:18:28.969238Z","shell.execute_reply":"2024-11-23T01:18:28.973159Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fields = []\nfields.extend(feature_cols)\nfields.extend([target_column])\nfields.extend(['weight'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:18:28.975653Z","iopub.execute_input":"2024-11-23T01:18:28.976081Z","iopub.status.idle":"2024-11-23T01:18:28.990074Z","shell.execute_reply.started":"2024-11-23T01:18:28.976033Z","shell.execute_reply":"2024-11-23T01:18:28.988709Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pl_train = pl.concat([pl.read_parquet(_f, columns=fields) for _f in train_parquets])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:18:28.991569Z","iopub.execute_input":"2024-11-23T01:18:28.991999Z","iopub.status.idle":"2024-11-23T01:18:45.604874Z","shell.execute_reply.started":"2024-11-23T01:18:28.991951Z","shell.execute_reply":"2024-11-23T01:18:45.603844Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# if we don't consider a sub sample (20%) it will run out of memory\nnum_rows_to_sample = len(pl_train) // 5\npl_train = pl_train.sample(n=num_rows_to_sample)\nX_train = pl_train.select(feature_cols).to_numpy()\ny_train = pl_train.select(target_column).to_numpy().flatten()\nX_train = np.nan_to_num(X_train)\ny_train = np.nan_to_num(y_train)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:18:45.606508Z","iopub.execute_input":"2024-11-23T01:18:45.606810Z","iopub.status.idle":"2024-11-23T01:19:27.350540Z","shell.execute_reply.started":"2024-11-23T01:18:45.606781Z","shell.execute_reply":"2024-11-23T01:19:27.349338Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.fit(X_train, y_train)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:19:27.352078Z","iopub.execute_input":"2024-11-23T01:19:27.352486Z","iopub.status.idle":"2024-11-23T01:20:18.196735Z","shell.execute_reply.started":"2024-11-23T01:19:27.352448Z","shell.execute_reply":"2024-11-23T01:20:18.195434Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# do prediction, train set per train set\nsubmission_linear = []\nfor batch_file in train_parquets:\n    # Load batch data\n    pl_train = pl.read_parquet(batch_file, columns=fields)\n    pl_train = pl_train.fill_null(0)\n    # Extract features and target\n    if feature_cols is None:\n        feature_cols = [col for col in pl_train.columns if col != target_column]\n    X = pl_train.select(feature_cols).to_numpy()\n    y_pred = model.predict(X).tolist()\n    submission_linear.extend(y_pred)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:20:18.198180Z","iopub.execute_input":"2024-11-23T01:20:18.198635Z","iopub.status.idle":"2024-11-23T01:20:45.498329Z","shell.execute_reply.started":"2024-11-23T01:20:18.198596Z","shell.execute_reply":"2024-11-23T01:20:45.497432Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"R2_linear = r2_score(wg, submission_linear, gt)\nR2_linear","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:20:45.499659Z","iopub.execute_input":"2024-11-23T01:20:45.500005Z","iopub.status.idle":"2024-11-23T01:20:48.214864Z","shell.execute_reply.started":"2024-11-23T01:20:45.499971Z","shell.execute_reply":"2024-11-23T01:20:48.213651Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It is interesting to observe how the simple linear model fitted on the 33% of the total data doesn't allow to achieve better results then a very simple constant model. The time series is highly no stationary. ","metadata":{}},{"cell_type":"markdown","source":"## XGBoost","metadata":{}},{"cell_type":"code","source":"import xgboost as xgb\nfrom sklearn.model_selection import train_test_split","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:20:48.216164Z","iopub.execute_input":"2024-11-23T01:20:48.216513Z","iopub.status.idle":"2024-11-23T01:20:48.948860Z","shell.execute_reply.started":"2024-11-23T01:20:48.216477Z","shell.execute_reply":"2024-11-23T01:20:48.947903Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fields = []\nfields.extend(feature_cols)\nfields.extend([target_column])\nfields.extend(['weight'])\npl_train = pl.concat([pl.read_parquet(_f, columns=fields) for _f in train_parquets])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:20:48.950274Z","iopub.execute_input":"2024-11-23T01:20:48.950629Z","iopub.status.idle":"2024-11-23T01:21:07.342117Z","shell.execute_reply.started":"2024-11-23T01:20:48.950595Z","shell.execute_reply":"2024-11-23T01:21:07.340838Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# let's consider a smaller example for demonstrative purposes\nnum_rows_to_sample = len(pl_train) // 5\npl_train = pl_train.sample(n=num_rows_to_sample)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:21:07.343947Z","iopub.execute_input":"2024-11-23T01:21:07.344425Z","iopub.status.idle":"2024-11-23T01:21:30.784167Z","shell.execute_reply.started":"2024-11-23T01:21:07.344375Z","shell.execute_reply":"2024-11-23T01:21:30.782806Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X = pl_train.select(feature_cols).to_numpy()\ny = pl_train.select(target_column).to_numpy().flatten()\nweights = pl_train.select(\"weight\").to_numpy().flatten()\nX = np.nan_to_num(X)\ny = np.nan_to_num(y)\nweights = np.nan_to_num(weights)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:21:30.785857Z","iopub.execute_input":"2024-11-23T01:21:30.786219Z","iopub.status.idle":"2024-11-23T01:21:49.849387Z","shell.execute_reply.started":"2024-11-23T01:21:30.786186Z","shell.execute_reply":"2024-11-23T01:21:49.848224Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)\nweights_train, weights_test = train_test_split(weights, test_size=0.2, random_state=42)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:21:49.850804Z","iopub.execute_input":"2024-11-23T01:21:49.851243Z","iopub.status.idle":"2024-11-23T01:22:15.556289Z","shell.execute_reply.started":"2024-11-23T01:21:49.851190Z","shell.execute_reply":"2024-11-23T01:22:15.555083Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"dtrain = xgb.DMatrix(X_train, label=y_train, weight=weights_train)  # Training set\ndtest = xgb.DMatrix(X_test, label=y_test, weight=weights_test)    # Test set","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:22:15.557863Z","iopub.execute_input":"2024-11-23T01:22:15.558237Z","iopub.status.idle":"2024-11-23T01:22:26.084168Z","shell.execute_reply.started":"2024-11-23T01:22:15.558203Z","shell.execute_reply":"2024-11-23T01:22:26.083134Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def r2_metric(predt: np.ndarray, dtrain: xgb.DMatrix):\n    '''Compute r2 metric for xgboost'''\n    y_true = dtrain.get_label()\n    weights = dtrain.get_weight()\n    numerator = np.sum(weights * (y_true - predt) ** 2)\n    denominator = np.sum(weights * (y_true ** 2))\n    r2_score = 1 - (numerator / denominator)\n    return ('R2', float(r2_score)) ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:22:26.085257Z","iopub.execute_input":"2024-11-23T01:22:26.085598Z","iopub.status.idle":"2024-11-23T01:22:26.091600Z","shell.execute_reply.started":"2024-11-23T01:22:26.085557Z","shell.execute_reply":"2024-11-23T01:22:26.090615Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Train with custom metric and capture the trained model\nparams = {\n    'objective': 'reg:squarederror',\n    'max_depth': 4,\n    'eta': 0.1,\n}\n\nmodel = xgb.train(\n    params,\n    dtrain,\n    num_boost_round=50,\n    evals=[(dtest, 'test')],  # Test R2 metric on test datasets\n    custom_metric=r2_metric,  # Custom metric function\n    verbose_eval=True\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:22:26.092345Z","iopub.execute_input":"2024-11-23T01:22:26.092732Z","iopub.status.idle":"2024-11-23T01:24:04.830395Z","shell.execute_reply.started":"2024-11-23T01:22:26.092698Z","shell.execute_reply":"2024-11-23T01:24:04.829514Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# do prediction, train set per train set\nsubmission_xgboost = []\nfor batch_file in train_parquets:\n    # Load batch data\n    pl_train = pl.read_parquet(batch_file, columns=fields)\n    pl_train = pl_train.fill_null(0)\n    # Extract features and target\n    if feature_cols is None:\n        feature_cols = [col for col in pl_train.columns if col != target_column]\n    X = pl_train.select(feature_cols).to_numpy()\n    X = xgb.DMatrix(X)  # Training set\n    y_pred = model.predict(X).tolist()\n    submission_xgboost.extend(y_pred)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:24:04.831710Z","iopub.execute_input":"2024-11-23T01:24:04.832721Z","iopub.status.idle":"2024-11-23T01:26:28.008168Z","shell.execute_reply.started":"2024-11-23T01:24:04.832681Z","shell.execute_reply":"2024-11-23T01:26:28.007099Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"R2_xgboost = r2_score(wg, submission_xgboost, gt)\nR2_xgboost","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:26:28.009791Z","iopub.execute_input":"2024-11-23T01:26:28.010789Z","iopub.status.idle":"2024-11-23T01:26:30.942755Z","shell.execute_reply.started":"2024-11-23T01:26:28.010736Z","shell.execute_reply":"2024-11-23T01:26:30.941410Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The R2 metric achieved with the xgboost is lower than the simple constant method. In order to improve the results we could add additional information as responders' lags. ","metadata":{}},{"cell_type":"markdown","source":"# Lags","metadata":{}},{"cell_type":"markdown","source":"So far we have limited ourself to consider only the features as values to train our model.\nWe have observed that responders are correlated and ideally we could improve the final results by involving also responders values in our model.\nThey are the: values of responder_{0...8} lagged by one date_id. The evaluation API serves the entirety of the lagged responders for a date_id on that date_id's first time_id. In other words, all of the previous date's responders will be served at the first time step of the succeeding date.\n","metadata":{}},{"cell_type":"markdown","source":"For creating the lags we will create some extra columns containg the lags i.e. the value of the responders of the previous day with respect to the considered responder 6 value.\nIn order to get the value of the previous day we will simply consider the last value of the responders for that day.\n\nCredit to the [notebook](https://www.kaggle.com/code/motono0223/js24-preprocessing-create-lags/notebook) for the proposed strategy","metadata":{}},{"cell_type":"code","source":"fields = []\nfields.extend(responder_cols)\nfields.extend(['date_id', 'symbol_id'])\nlag_cols_rename = { f\"responder_{idx}\" : f\"responder_{idx}_lag_1\" for idx in range(9)}\npl_train = pl.concat([pl.read_parquet(_f, columns=fields) for _f in train_parquets])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:26:30.944571Z","iopub.execute_input":"2024-11-23T01:26:30.945092Z","iopub.status.idle":"2024-11-23T01:26:37.401893Z","shell.execute_reply.started":"2024-11-23T01:26:30.945035Z","shell.execute_reply":"2024-11-23T01:26:37.400853Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"lag_cols_rename = { f\"responder_{idx}\" : f\"responder_{idx}_lag_1\" for idx in range(9)}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:26:37.403207Z","iopub.execute_input":"2024-11-23T01:26:37.403610Z","iopub.status.idle":"2024-11-23T01:26:37.408991Z","shell.execute_reply.started":"2024-11-23T01:26:37.403556Z","shell.execute_reply":"2024-11-23T01:26:37.407795Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(pl_train.columns)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:26:37.410248Z","iopub.execute_input":"2024-11-23T01:26:37.410598Z","iopub.status.idle":"2024-11-23T01:26:37.423015Z","shell.execute_reply.started":"2024-11-23T01:26:37.410564Z","shell.execute_reply":"2024-11-23T01:26:37.421880Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"lags = pl_train.rename(lag_cols_rename)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:26:37.424486Z","iopub.execute_input":"2024-11-23T01:26:37.425676Z","iopub.status.idle":"2024-11-23T01:26:37.437758Z","shell.execute_reply.started":"2024-11-23T01:26:37.425615Z","shell.execute_reply":"2024-11-23T01:26:37.436708Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(lags.columns)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:26:37.439329Z","iopub.execute_input":"2024-11-23T01:26:37.439844Z","iopub.status.idle":"2024-11-23T01:26:37.450455Z","shell.execute_reply.started":"2024-11-23T01:26:37.439799Z","shell.execute_reply":"2024-11-23T01:26:37.449378Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"lags = lags.with_columns(\n    date_id = pl.col('date_id') + 1,  # lagged by 1 day\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:26:37.451773Z","iopub.execute_input":"2024-11-23T01:26:37.452188Z","iopub.status.idle":"2024-11-23T01:26:37.484750Z","shell.execute_reply.started":"2024-11-23T01:26:37.452129Z","shell.execute_reply":"2024-11-23T01:26:37.483530Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"lags = lags.group_by([\"date_id\", \"symbol_id\"], maintain_order=True).last()  # pick up last record of previous date","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:26:37.486242Z","iopub.execute_input":"2024-11-23T01:26:37.486703Z","iopub.status.idle":"2024-11-23T01:26:39.641570Z","shell.execute_reply.started":"2024-11-23T01:26:37.486654Z","shell.execute_reply":"2024-11-23T01:26:39.639998Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"lags.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:26:39.643638Z","iopub.execute_input":"2024-11-23T01:26:39.644228Z","iopub.status.idle":"2024-11-23T01:26:39.659658Z","shell.execute_reply.started":"2024-11-23T01:26:39.644152Z","shell.execute_reply":"2024-11-23T01:26:39.658043Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pl_train.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:26:39.661713Z","iopub.execute_input":"2024-11-23T01:26:39.662069Z","iopub.status.idle":"2024-11-23T01:26:39.680576Z","shell.execute_reply.started":"2024-11-23T01:26:39.662035Z","shell.execute_reply":"2024-11-23T01:26:39.679039Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"lags.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:26:39.682265Z","iopub.execute_input":"2024-11-23T01:26:39.682726Z","iopub.status.idle":"2024-11-23T01:26:39.697105Z","shell.execute_reply.started":"2024-11-23T01:26:39.682690Z","shell.execute_reply":"2024-11-23T01:26:39.696108Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pl_train = pl_train.join(lags, on=[\"date_id\", \"symbol_id\"],  how=\"left\")\npl_train.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:26:39.698406Z","iopub.execute_input":"2024-11-23T01:26:39.698992Z","iopub.status.idle":"2024-11-23T01:26:43.313724Z","shell.execute_reply.started":"2024-11-23T01:26:39.698931Z","shell.execute_reply":"2024-11-23T01:26:43.312458Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pl_train = pl_train.drop(responder_cols)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:26:43.315132Z","iopub.execute_input":"2024-11-23T01:26:43.315651Z","iopub.status.idle":"2024-11-23T01:26:43.341298Z","shell.execute_reply.started":"2024-11-23T01:26:43.315586Z","shell.execute_reply":"2024-11-23T01:26:43.339708Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pl_train.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:26:43.343249Z","iopub.execute_input":"2024-11-23T01:26:43.343900Z","iopub.status.idle":"2024-11-23T01:26:43.366023Z","shell.execute_reply.started":"2024-11-23T01:26:43.343837Z","shell.execute_reply":"2024-11-23T01:26:43.364743Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"As expected the lag entries of the dataframe of the day 0 will be null since we don't have values for the responder for the day -1.\nLet's discard see the values for other days.","metadata":{}},{"cell_type":"code","source":"pl_train.filter(pl.col('date_id') != 0).tail()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:26:43.367449Z","iopub.execute_input":"2024-11-23T01:26:43.367796Z","iopub.status.idle":"2024-11-23T01:26:44.536581Z","shell.execute_reply.started":"2024-11-23T01:26:43.367764Z","shell.execute_reply":"2024-11-23T01:26:44.535300Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# save dataframe\npl_train.write_parquet(\"train_lags.parquet\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T01:26:44.537914Z","iopub.execute_input":"2024-11-23T01:26:44.538287Z","iopub.status.idle":"2024-11-23T01:26:48.229490Z","shell.execute_reply.started":"2024-11-23T01:26:44.538253Z","shell.execute_reply":"2024-11-23T01:26:48.228217Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We could now retrain a boosting based method with this additional information.","metadata":{}},{"cell_type":"markdown","source":"# Other possible solutions","metadata":{}},{"cell_type":"markdown","source":"Often in order to improve the final accuracy of the prediction, several models could be used and their results merged. This process takes the name of ensambling and the predictions from different models can be fused together with techniques like:\n\n- ***Bagging***: Combines predictions of models trained on different subsets of the training data (e.g., Random Forest).\n- ***Boosting***: Sequentially trains models, where each subsequent model focuses on correcting the errors of the previous ones (e.g., Gradient Boosting, XGBoost).\n- ***Stacking:*** Combines multiple models using a meta-model to make the final prediction.\n\nWe could train several models on the improved train set with the lagged responders and ensable their solution to improve the total score.","metadata":{}}]}