{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84493,"databundleVersionId":9871156,"sourceType":"competition"},{"sourceId":211702440,"sourceType":"kernelVersion"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<p style='font: 600 36px \"Lato\", \"Open Sans\", \"Helvetica Neue\", Helvetica, Arial, sans-serif'>\n    <img align=\"right\" width=\"300\" height=\"150\" style=\"border-radius: 12px;\" src=\"https://www.kaggle.com/competitions/84493/images/header\">\n    Jane Street Real-Time Market Data Forecasting [EDA]\n    <p style='font: 500 22px \"Lato\", \"Open Sans\", \"Helvetica Neue\", Helvetica, Arial, sans-serif'>\n        Predict financial market responders using real-world data\n    </p>\n</p>\n\n<p style='font: 300 18px/22px \"Lato\", \"Open Sans\", \"Helvetica Neue\", Helvetica, Arial, sans-serif'>\n    In this competition, hosted by Jane Street, you'll build a model using real-world data derived from production systems, which offers a glimpse into the daily challenges of successful trading. This challenge highlights the difficulties in modeling financial markets, including fat-tailed distributions, non-stationary time series, and sudden shifts in market behavior.\n</p>\n        \n","metadata":{}},{"cell_type":"code","source":"!pip install hvplot -q","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:41:41.355187Z","iopub.execute_input":"2024-12-19T00:41:41.355640Z","iopub.status.idle":"2024-12-19T00:41:53.493612Z","shell.execute_reply.started":"2024-12-19T00:41:41.355593Z","shell.execute_reply":"2024-12-19T00:41:53.492302Z"},"_kg_hide-input":true,"_kg_hide-output":false},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport polars as pl\nimport polars.selectors as cs\nimport hvplot\nimport holoviews as hv\nfrom holoviews import opts\nimport panel as pn\nimport hvplot.polars\nfrom hvplot.plotting import scatter_matrix\nimport os\nfrom os.path import split\nfrom glob import glob\nfrom itertools import combinations\nfrom typing import Callable\nfrom bokeh.models import NumeralTickFormatter\nimport matplotlib.pyplot as plt\nfrom tqdm import tqdm\nimport statsmodels.api as sm\n\nhv.extension('bokeh')\npl.Config(\n    fmt_str_lengths=80,\n    tbl_rows=80,\n    tbl_cols=95,\n    set_fmt_table_cell_list_len=50,\n    set_thousands_separator=' ',\n    float_precision=4,\n    set_fmt_float=\"full\",\n    tbl_cell_alignment = \"LEFT\",\n    tbl_cell_numeric_alignment=\"RIGHT\"\n)\nfrmt_big_numb = NumeralTickFormatter(format='0.0a')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-12-19T00:41:53.495652Z","iopub.execute_input":"2024-12-19T00:41:53.496048Z","iopub.status.idle":"2024-12-19T00:41:59.706846Z","shell.execute_reply.started":"2024-12-19T00:41:53.496012Z","shell.execute_reply":"2024-12-19T00:41:59.705687Z"},"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"MAIN_PATH = '/kaggle/input/jane-street-real-time-market-data-forecasting'\nTRAIN_PATH = MAIN_PATH + '/train.parquet'\nAUX_DATA_PATH = '/kaggle/input/janestreet-rtmdf-auxiliary-data'\nFEATURES_NUMBER = 79\nRESPONDERS_NUMBER = 9\nSYMBOL_NUMBER = 39","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:41:59.708258Z","iopub.execute_input":"2024-12-19T00:41:59.708848Z","iopub.status.idle":"2024-12-19T00:41:59.714501Z","shell.execute_reply.started":"2024-12-19T00:41:59.708780Z","shell.execute_reply":"2024-12-19T00:41:59.713271Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Volume and structure of data","metadata":{}},{"cell_type":"markdown","source":"We have the following train directory structure:\n<div style=\"margin-left: 5px;color: #6A5ACD;\">└── train.parquet </div>\n<div style=\"margin-left: 20px;color: #BA55D3;\">├── partition_id=0 </div>\n<div style=\"margin-left: 40px;color: #008B8B;\">├── part-0.parquet </div>\n<div style=\"margin-left: 20px;color: #BA55D3;\">├── ... </div>\n<div style=\"margin-left: 40px;color: #008B8B;\">├── ... </div>\n<div style=\"margin-left: 20px;color: #BA55D3;\">├── partition_id=9 </div>\n<div style=\"margin-left: 40px;color: #008B8B;\">├── part-0.parquet </div>\n<br>\nThe data is divided into 10 files, located in 10 directories, one in each. To process these files, we create a procedure that will read a given list of files in a loop, apply a given function to each of them, and return the collected dataframe.","metadata":{}},{"cell_type":"code","source":"train_files = sorted([path for path in glob(f'{TRAIN_PATH}/*/*.parquet')])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:41:59.716883Z","iopub.execute_input":"2024-12-19T00:41:59.717228Z","iopub.status.idle":"2024-12-19T00:41:59.762470Z","shell.execute_reply.started":"2024-12-19T00:41:59.717196Z","shell.execute_reply":"2024-12-19T00:41:59.761253Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def read_n_proc_parts(path_list: list,\n                      func: Callable|None = None,\n                      kwargs: dict|None = None, \n                      columns: str|list[str] = 'all', \n                      path_transit: bool = False,\n                     ) -> pl.DataFrame:\n    \n    \"\"\"\n    Reads files from a given list (`path_list`), applies a given function (`func`) \n    to individual parts and combines the result into a single dataframe \n\n    Parameters\n    ----------\n    path_list: list\n        List of files to be processed\n    func: callable, default None\n        Function to apply to each dataframe\n    kwargs: dict, default None\n        `func` arguments\n    columns: str|list, default 'all'\n        What columns to select from file when reading\n    path_transit: bool, default False\n        Transit the file path to the `func`\n    \"\"\"\n    \n    for i, path in tqdm(enumerate(path_list), total=(len(path_list)), ncols=80):\n        \n        if columns == 'all':\n            df = pl.read_parquet(path) if not path_transit else pl.read_parquet(path, include_file_paths='path')\n        else:\n            df = pl.read_parquet(path, columns=columns) if not path_transit else pl.read_parquet(path,\n                                                                                              columns=columns,\n                                                                                              include_file_paths='path')\n            \n        if isinstance(func, Callable):\n            if isinstance(kwargs, dict): \n                df = df.pipe(func, **kwargs)\n            else:                   \n                df = df.pipe(func)\n        \n        if i==0:\n            df_ret = df\n        else:\n            df_ret = df_ret.vstack(df)\n\n    return df_ret","metadata":{"trusted":true,"_kg_hide-input":false,"execution":{"iopub.status.busy":"2024-12-19T00:41:59.763906Z","iopub.execute_input":"2024-12-19T00:41:59.764227Z","iopub.status.idle":"2024-12-19T00:41:59.774153Z","shell.execute_reply.started":"2024-12-19T00:41:59.764198Z","shell.execute_reply":"2024-12-19T00:41:59.772922Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def read_n_proc_parts_lazy(path_list: list,\n                      func: Callable|None = None,\n                      kwargs: dict|None = None, \n                      columns: str|list[str] = 'all', \n                      path_transit: bool = False,\n                     ) -> pl.DataFrame:\n    \n    \"\"\"\n    Reads files from a given list (`path_list`), applies a given function (`func`) \n    to individual parts and combines the result into a single dataframe \n\n    Parameters\n    ----------\n    path_list: list\n        List of files to be processed\n    func: callable, default None\n        Function to apply to each dataframe\n    kwargs: dict, default None\n        `func` arguments\n    columns: str|list, default 'all'\n        What columns to select from file when reading\n    path_transit: bool, default False\n        Transit the file path to the `func`\n    \"\"\"\n    \n    for i, path in tqdm(enumerate(path_list), total=(len(path_list)), ncols=80):\n        \n        if columns == 'all':\n            df = pl.scan_parquet(path) if not path_transit else pl.scan_parquet(path, include_file_paths='path')\n        else:\n            df = pl.scan_parquet(path) if not path_transit else pl.scan_parquet(path, include_file_paths='path')\n            \n        if isinstance(func, Callable):\n            if isinstance(kwargs, dict): \n                df = df.pipe(func, **kwargs)\n            else:                   \n                df = df.pipe(func)\n        \n        if i==0:\n            df_ret = df\n        else:\n            df_ret = df_ret.collect().vstack(df)\n\n    return df_ret","metadata":{"trusted":true,"_kg_hide-input":false,"execution":{"iopub.status.busy":"2024-12-19T00:41:59.775558Z","iopub.execute_input":"2024-12-19T00:41:59.775951Z","iopub.status.idle":"2024-12-19T00:41:59.794345Z","shell.execute_reply.started":"2024-12-19T00:41:59.775917Z","shell.execute_reply":"2024-12-19T00:41:59.793185Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Analytical features","metadata":{}},{"cell_type":"markdown","source":"The description of the competition data states that the primary key consists of the features: `date_id`, `time_id`, `symbol_id`. Let's deal with this first.","metadata":{}},{"cell_type":"code","source":"def prkey_stat(df: pl.DataFrame) -> pl.DataFrame:\n    \"\"\"\n    Calculates the total number of rows and some totals by primary key details\n    \n    Return\n    ----------\n    pl.DataFrame with columns:\n        nrows: total number of rows\n        unique_date: number of unique date values\n        unique_time: number of unique time values\n        unique_symbol: number of unique sumbol values\n        unique_date_symb: number of unique combinations of date and symbol values\n        hypoth: a sign of the correctness of the hypothesis that \n                rows = uniq_time * uniq_date_symb\n    \"\"\"\n    rec = {}\n    rec['partition_id'] = df['path'][0].split('/')[-2][-1]\n    rec['nrows'] = df.shape[0]\n    rec['uniq_date'] = df['date_id'].n_unique()\n    rec['uniq_time'] = df['time_id'].n_unique()\n    rec['uniq_symbol'] = df['symbol_id'].n_unique()\n    rec['uniq_date_symb'] = df['date_id', 'symbol_id'].unique().shape[0]\n    rec['hypoth'] = rec['uniq_time'] * rec['uniq_date_symb'] == rec['nrows']\n    \n    return pl.DataFrame(rec)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:41:59.795710Z","iopub.execute_input":"2024-12-19T00:41:59.796184Z","iopub.status.idle":"2024-12-19T00:41:59.812839Z","shell.execute_reply.started":"2024-12-19T00:41:59.796129Z","shell.execute_reply":"2024-12-19T00:41:59.811580Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_stat = read_n_proc_parts(sorted([path for path in glob(f'{TRAIN_PATH}/*/*.parquet')]),\n                            prkey_stat,\n                            columns=['date_id', 'time_id', 'symbol_id'],\n                            path_transit = True,\n                           )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:41:59.814372Z","iopub.execute_input":"2024-12-19T00:41:59.814860Z","iopub.status.idle":"2024-12-19T00:42:03.465218Z","shell.execute_reply.started":"2024-12-19T00:41:59.814793Z","shell.execute_reply":"2024-12-19T00:42:03.464037Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"clmns_w_bars = ['nrows', 'uniq_time', 'uniq_symbol', 'uniq_date_symb']\nnumeric_clmns = df_stat.select(cs.numeric()).columns\ngrid = []\n\ngrid.append(df_stat.to_pandas().style\\\n                    .format('{:,d}', subset=numeric_clmns)\\\n                    .bar(cmap='RdYlGn_r', width=50, height=20, subset=clmns_w_bars)\\\n                    .highlight_min(color='red', subset=['hypoth']))\n\ngrid.append(pn.indicators.Number(value=df_stat['nrows'].sum(),\n                                 name='Total rows',\n                                 colors=[(np.inf, '#fff')],\n                                 title_size = '30pt',\n                                 font_size='30pt',\n                                 format = '{value:,}',\n                                 width=230,\n                                 styles={'background':'#2980B9', \n                                         'text-align':'center',\n                                         'padding':'80px 0 0 0',\n                                         'border-radius': '5px',\n                                         'height':'270px',\n                                        },\n                                )\n           )\n\npn.WidgetBox('### Statistics on key features', \n             pn.GridBox(*grid, styles={'background':'#add8e6'}, ncols=2)\n            )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:42:03.466939Z","iopub.execute_input":"2024-12-19T00:42:03.467372Z","iopub.status.idle":"2024-12-19T00:42:03.750665Z","shell.execute_reply.started":"2024-12-19T00:42:03.467324Z","shell.execute_reply":"2024-12-19T00:42:03.749589Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The tables contain *from two to six million records*. In total, all tables contain *47,1 million records*. Each contains data on 170 dates, with 169 dates in the last ones alone (***total dates 1699***). Each date is divided into ***869 or 968 time intervals***. There are from ***20 to 39 financial instruments***, i.e. not all of them are in each table.\n\nIt is especially worth explaining about the hypothesis. For all time intervals of a single date, the number of financial instruments is fixed. For example, **for a date with index 0** there is data on **8 instruments**, which means that for all **849 intervals** there will be 8 records, i.e. a total of **`8*849=6792`** rows. That is, the number of observed financial instruments changes from date to date, but cannot change within a date. To test this hypothesis, we calculate the sum of unique values ​​of `date_id & symbol_id`, multiply by the number of unique values ​​of `time id` and compare with the total number of rows. For example, for a table with index 0, we have **2290** unique combinations of `date_id & symbol_id` and **849** intervals per date. We multiply `2290*849` and get **1,944,210**, which is equal to the number of records in the table.\n\nEverything is great, but, as can be seen from the table above, this hypothesis is not met for the fourth file (index 3)! Let's figure it out... ","metadata":{}},{"cell_type":"markdown","source":"<div style='font: 300 16px/22px \"Lato\", \"Open Sans\", \"Helvetica Neue\", Helvetica, Arial, sans-serif'>\n    For the table above, we simply counted the number of unique values ​​of the <code>time_id</code> feature, assuming that all dates in the table are divided into the same number of intervals. Let's check if this is true.\n</div>","metadata":{}},{"cell_type":"code","source":"# let's open the file with index 3\ntmp_df = pl.read_parquet(TRAIN_PATH + '/partition_id=3/part-0.parquet', columns=['date_id', 'time_id', 'symbol_id'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:42:03.755472Z","iopub.execute_input":"2024-12-19T00:42:03.755933Z","iopub.status.idle":"2024-12-19T00:42:03.773250Z","shell.execute_reply.started":"2024-12-19T00:42:03.755900Z","shell.execute_reply":"2024-12-19T00:42:03.772020Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# let's count the number of intervals for each date\nn_uniq_time_per_date = tmp_df.group_by('date_id').agg(pl.col('time_id').unique().len())\n# for three of the 170 dates in the table, the number of time intervals differs\nn_uniq_time_per_date.group_by('time_id').agg(pl.col('date_id').len())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:42:03.774714Z","iopub.execute_input":"2024-12-19T00:42:03.775148Z","iopub.status.idle":"2024-12-19T00:42:03.849064Z","shell.execute_reply.started":"2024-12-19T00:42:03.775113Z","shell.execute_reply":"2024-12-19T00:42:03.847934Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# here are the indexes of these dates, and these are the last three dates\ndisplay(n_uniq_time_per_date.filter(pl.col('time_id')==968))\nprint(f\"min values of data_id - {n_uniq_time_per_date['date_id'].min()}\")\nprint(f\"max values of data_id - {n_uniq_time_per_date['date_id'].max()}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:42:03.850297Z","iopub.execute_input":"2024-12-19T00:42:03.850647Z","iopub.status.idle":"2024-12-19T00:42:03.863486Z","shell.execute_reply.started":"2024-12-19T00:42:03.850615Z","shell.execute_reply":"2024-12-19T00:42:03.862343Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style='font-size:16px'>In the table above, we noted that the provided files can contain dates divided into different numbers of time intervals - <strong>849</strong> or <strong>968</strong>. In files with <strong>indexes 0-2</strong> there are always <strong>849</strong> intervals, in files with <strong>indexes 4-9</strong> there are always <strong>968</strong> intervals, and only in the file with <strong>index 3</strong> there are <strong>both options</strong>. Thus, the total number of unique date-time combinations in the training dataset:\n\n`849*677 + 968*(5*170-1+3) = 1 564 069`\n\nLet's remember this number, we'll need it below.</div>","metadata":{}},{"cell_type":"markdown","source":"## Weights","metadata":{}},{"cell_type":"markdown","source":"Let's see what kind of beast this `weight` is. The organizer described this feature as follows: *\"The weighting used for calculating the scoring function\"*","metadata":{}},{"cell_type":"code","source":"df_wght = read_n_proc_parts(sorted([path for path in glob(f'{TRAIN_PATH}/*/*.parquet')]),\n                            columns=['date_id', 'time_id', 'symbol_id', 'weight'],\n                           )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:42:03.865062Z","iopub.execute_input":"2024-12-19T00:42:03.865515Z","iopub.status.idle":"2024-12-19T00:42:04.180961Z","shell.execute_reply.started":"2024-12-19T00:42:03.865479Z","shell.execute_reply":"2024-12-19T00:42:04.179891Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style='font: 300 16px/22px \"Lato\", \"Open Sans\", \"Helvetica Neue\", Helvetica, Arial, sans-serif'> First, a few observations:\n<ul>\n<li>the dataset contains <strong>50,382</strong> <code>date_id & symbol_id</code> combinations</li>\n<li>for a given <code>symbol_id</code> the <code>weight</code> is set on a given <code>date_id</code> and remains the same for all intervals (<code>time_id</code>) of that date</li>\n</ul></div>","metadata":{}},{"cell_type":"code","source":"df_agg_uniq_wght = df_wght.group_by(['date_id', 'symbol_id']).agg(pl.col('weight').unique())\nprint(f\"number of date/symbol combinations - {df_agg_uniq_wght.shape[0]:,d}\")\nprint(f\"number of unique weights per unique date/symbol combination is not more than - {df_agg_uniq_wght['weight'].list.len().unique().max()}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:42:04.182496Z","iopub.execute_input":"2024-12-19T00:42:04.182977Z","iopub.status.idle":"2024-12-19T00:42:07.397244Z","shell.execute_reply.started":"2024-12-19T00:42:04.182926Z","shell.execute_reply":"2024-12-19T00:42:07.395904Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let's look at the distribution of weights. All weights are positive, the bulk is concentrated in the range of 1-3, and there is a long tail up to 10","metadata":{}},{"cell_type":"code","source":"pn.WidgetBox('### Weights distribution',\n             pn.GridBox(*\n                        [\n                            (\n                                df_wght['weight'].hvplot\n                                .hist(line_color=None,\n                                      grid=True,\n                                      yformatter=frmt_big_numb,\n                                     ).opts(active_tools=['pan'])\n                            ),\n                            (\n                                df_wght['weight'].describe().to_pandas()\n                                .style\n                                .hide()\n                                .format(precision=2, \n                                        thousands=\",\",\n                                        subset=['value'])\n                            )\n                        ],\n                        ncols=2\n                       )\n            )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:42:07.398831Z","iopub.execute_input":"2024-12-19T00:42:07.399213Z","iopub.status.idle":"2024-12-19T00:42:12.160120Z","shell.execute_reply.started":"2024-12-19T00:42:07.399179Z","shell.execute_reply":"2024-12-19T00:42:12.158920Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let's see how the weights are distributed across financial instruments (`symbol_id`)","metadata":{}},{"cell_type":"code","source":"# to begin with, let's make a pivot table, which we will often return to in this section\ndf_wght_pivot = df_wght.sort('symbol_id').pivot('symbol_id', index=['date_id', 'time_id'], values='weight')\ndf_wght_pivot.head()\n# pay attention to the null values; here we are faced with the fact \n# that data for certain financial instruments is not available on every date","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:42:12.161587Z","iopub.execute_input":"2024-12-19T00:42:12.161973Z","iopub.status.idle":"2024-12-19T00:42:39.918010Z","shell.execute_reply.started":"2024-12-19T00:42:12.161938Z","shell.execute_reply":"2024-12-19T00:42:39.916612Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_wght_desc = df_wght_pivot.select(pl.exclude(['date_id', 'time_id'])).describe(percentiles=[0.05, 0.25, 0.5, 0.75, 0.95])\ncolumns = df_wght_desc['statistic'].to_list()\ndisplay(df_wght_desc.select(cs.numeric())\n                    .transpose(column_names=columns)\n                    .select(pl.all().exclude('count', 'null_count'))\n                    .to_pandas().style\n                    .format(precision=4)\n                    .bar(cmap='RdYlGn_r',\n                         width=50, \n                         height=40,\n                         props='width: 120px; border-left: 0.5px dotted gray;'\n                        )\n       )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:42:39.919705Z","iopub.execute_input":"2024-12-19T00:42:39.920262Z","iopub.status.idle":"2024-12-19T00:42:40.782843Z","shell.execute_reply.started":"2024-12-19T00:42:39.920206Z","shell.execute_reply":"2024-12-19T00:42:40.781711Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15, 5))\nflierprops = {'markerfacecolor':'red',\n              'markeredgecolor':'none',\n              'markersize':2\n             }\n\nax.set_title('Weights by financial instruments')\n(df_wght_pivot\n .select(pl.exclude(['date_id', 'time_id']))\n .to_pandas()\n .plot.box(ax=ax, flierprops=flierprops)\n);","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:42:40.784503Z","iopub.execute_input":"2024-12-19T00:42:40.784865Z","iopub.status.idle":"2024-12-19T00:42:44.355667Z","shell.execute_reply.started":"2024-12-19T00:42:40.784798Z","shell.execute_reply":"2024-12-19T00:42:44.354454Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Here we notice our obvious heavyweights - instruments 1 and 19, lightweights - instruments 6, 21, 26, 37. And outliers, outliers, outliers.... And only the representative of the lightest weight - instrument 6, never overate :)","metadata":{}},{"cell_type":"markdown","source":"Now let's try the correlation. To estimate the correlation of weights, it is necessary to remove rows with null values from our pivot table. Only a quarter of the data remains","metadata":{}},{"cell_type":"code","source":"print(f'number of unique date/time combinations - {df_wght_pivot.height:,d}')\nprint(f'remainder after nulls dropping - {df_wght_pivot.drop_nulls().height / df_wght_pivot.height:.1%}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:42:44.357163Z","iopub.execute_input":"2024-12-19T00:42:44.357525Z","iopub.status.idle":"2024-12-19T00:42:44.375318Z","shell.execute_reply.started":"2024-12-19T00:42:44.357491Z","shell.execute_reply":"2024-12-19T00:42:44.374213Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"However, let's see if there is a correlation of weights in this remaining quarter of the data (at least to make sure that there isn't one :) Here all the instruments are more or less going in the same direction (the minimum does not reach -0.4), although not in step. The first instrument stands out, although it is not exactly against everyone, but it plays its own game as much as possible","metadata":{}},{"cell_type":"code","source":"df_wght_pivot.select(cs.float()).drop_nulls().corr() \\\n.hvplot.heatmap(height=600,\n                width=800,\n                rot=90,\n               ).opts(active_tools=['pan'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:42:44.377107Z","iopub.execute_input":"2024-12-19T00:42:44.377576Z","iopub.status.idle":"2024-12-19T00:42:44.774114Z","shell.execute_reply.started":"2024-12-19T00:42:44.377525Z","shell.execute_reply":"2024-12-19T00:42:44.772729Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# if you try to filter out the features with the highest correlation (2 and 26 with corr. 0.83)\n# to look at them over a longer period, their correlation even decreases (0.73)\nsome_feat = df_wght_pivot.select(['2','26']).drop_nulls()\nprint(f'number of records - {some_feat.height:,d} ({some_feat.height / df_wght_pivot.height:.1%})')\nsome_feat.corr()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:42:44.775736Z","iopub.execute_input":"2024-12-19T00:42:44.776212Z","iopub.status.idle":"2024-12-19T00:42:44.810262Z","shell.execute_reply.started":"2024-12-19T00:42:44.776163Z","shell.execute_reply":"2024-12-19T00:42:44.809110Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Here's what it looks like on scatter_matrix for some features","metadata":{}},{"cell_type":"code","source":"gspec = pn.GridSpec(width=1200, height=1000)\n\ngspec[:, 0:8] = scatter_matrix(df_wght_pivot.select([str(col) for col in range(5)]).drop_nulls().to_pandas(),alpha=5,\n                               datashade= True,\n                               spread=True,\n                               tools=['pan'],\n                               diagonal_kwds=dict(line_color=None, yformatter=frmt_big_numb,), \n                              ).relabel('symbol_id 0-4')\n\ngspec[0:3, 8:12] = scatter_matrix(df_wght_pivot.select(['2','26']).drop_nulls().to_pandas(),\n                                  datashade= True,\n                                  spread=True,\n                                  tools=['pan'],\n                                  diagonal_kwds=dict(line_color=None, yformatter=frmt_big_numb,),\n                                 ).relabel('symbol_id 2 and 26')\n\ngspec[3:8, 8:12] = scatter_matrix(df_wght_pivot.select([str(col) for col in range(13,16)]).drop_nulls().to_pandas(),\n                                  datashade= True,\n                                  spread=True,\n                                  tools=['pan'],\n                                  diagonal_kwds=dict(line_color=None, yformatter=frmt_big_numb,),\n                                 ).relabel('symbol_id 13-15')\n\ngspec","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:42:44.811972Z","iopub.execute_input":"2024-12-19T00:42:44.812423Z","iopub.status.idle":"2024-12-19T00:43:13.552055Z","shell.execute_reply.started":"2024-12-19T00:42:44.812377Z","shell.execute_reply":"2024-12-19T00:43:13.550170Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"In general, the correlations did not work out, and this suggests that the weights clearly do not reflect any general economic trends.\n\nAnother thing we'll look at is how the weights are distributed over time across financial instruments. This chart also clearly shows the gaps in the data by date","metadata":{}},{"cell_type":"code","source":"# at the date level, the weights of each instrument are unchanged, so we use the average for aggregation\ndf_wght_short = df_wght.group_by(['date_id', 'symbol_id']).agg(pl.col('weight').mean())\n\n(df_wght_short.hvplot.scatter(x='date_id', \n                              y='symbol_id',\n                              c='weight',\n                              cmap='RdYlGn_r',\n                              width=1200,\n                              height=800,\n                              yticks=[col for col in range(SYMBOL_NUMBER)],\n                             ).opts(active_tools=['pan'])\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:43:13.553854Z","iopub.execute_input":"2024-12-19T00:43:13.554837Z","iopub.status.idle":"2024-12-19T00:43:15.888206Z","shell.execute_reply.started":"2024-12-19T00:43:13.554747Z","shell.execute_reply":"2024-12-19T00:43:15.887007Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"In the box plot above, we saw strong outliers. If these outliers are plotted separately, we can see a pattern - the bulk of the outliers are concentrated later in time. We can assume that the weights indicate the greater importance of recent events, i.e., they can change dynamically over time.","metadata":{}},{"cell_type":"code","source":"for symbol in range(SYMBOL_NUMBER):\n    q1, q3 = np.percentile(df_wght.filter(pl.col('symbol_id')==symbol)['weight'], (25, 75))\n    iqr = q3 - q1\n    curr_df = df_wght_short.filter((pl.col('symbol_id')==symbol) & (pl.col('weight') > q3 + iqr * 1.5))\n    if symbol==0: \n        df_outliers = curr_df\n    else:\n        df_outliers = df_outliers.vstack(curr_df) \n        \n(df_outliers.hvplot.scatter(x='date_id', \n                              y='symbol_id',\n                              c='weight',\n                              cmap='RdYlGn_r',\n                              width=1200,\n                              height=800,\n                              yticks=[col for col in range(39)],\n                             ).opts(active_tools=['pan'])\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:43:15.889706Z","iopub.execute_input":"2024-12-19T00:43:15.890090Z","iopub.status.idle":"2024-12-19T00:43:17.279269Z","shell.execute_reply.started":"2024-12-19T00:43:15.890055Z","shell.execute_reply":"2024-12-19T00:43:17.278167Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## What about nulls?","metadata":{}},{"cell_type":"markdown","source":"The absence of information can be of two types: \n 1. the absence of data on some financial instruments on certain dates\n 2. the presence of data, but in the null value\n\nWe have discussed the first case above, now let's consider the second one.","metadata":{}},{"cell_type":"code","source":"def null_counts(df: pl.DataFrame, \n                excl_columns: list[str]) -> pl.DataFrame:\n    \"\"\"\n    Counts nulls in the given dataframe across all columns, except those explicitly excluded. \n    Returns a dataframe with the following columns: features, null_numb, not_null_numb\n    \"\"\"   \n    rows = df.shape[0]\n    \n    df = (df.null_count()\n            .select(pl.all().exclude(excl_columns))\n            .unpivot(variable_name='features', value_name='null_numb')\n            .with_columns(not_null_numb = rows - pl.col('null_numb'))\n         )\n    \n    return df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:43:17.280756Z","iopub.execute_input":"2024-12-19T00:43:17.281250Z","iopub.status.idle":"2024-12-19T00:43:17.288575Z","shell.execute_reply.started":"2024-12-19T00:43:17.281200Z","shell.execute_reply":"2024-12-19T00:43:17.287211Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_nulls = read_n_proc_parts(sorted([path for path in glob(f'{TRAIN_PATH}/*/*.parquet')]),\n                             null_counts,\n                             {'excl_columns':['date_id', 'time_id', 'symbol_id', 'weight']}\n                            )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:43:17.289988Z","iopub.execute_input":"2024-12-19T00:43:17.290325Z","iopub.status.idle":"2024-12-19T00:44:19.579251Z","shell.execute_reply.started":"2024-12-19T00:43:17.290292Z","shell.execute_reply":"2024-12-19T00:44:19.578210Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# group by features and filter only features with nulls\ndf_nulls_bar = (df_nulls\n                .group_by('features').agg(pl.col('null_numb', 'not_null_numb').sum())\n                .filter(pl.col('null_numb')>0)\n                .sort('features')\n               )\n\ndf_nulls_bar.hvplot.bar(x='features',\n                        y=['not_null_numb', 'null_numb'],\n                        stacked=True,\n                        rot=60,\n                        height=600,\n                        width=1200,\n                        line_color=None,\n                        title=f'{df_nulls_bar.shape[0]} of {df_nulls[\"features\"].unique().shape[0]} features with nulls',\n                        ylabel='number'\n                       ).opts(active_tools=['pan'], yformatter=frmt_big_numb,)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:44:19.580669Z","iopub.execute_input":"2024-12-19T00:44:19.581032Z","iopub.status.idle":"2024-12-19T00:44:19.885084Z","shell.execute_reply.started":"2024-12-19T00:44:19.580999Z","shell.execute_reply":"2024-12-19T00:44:19.882985Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"A little more than half of the features contain null values (47 of 88), but in small proportions. The most affected are features 21, 26, 27, 31, where there are 18% nulls. It could have been worse...","metadata":{}},{"cell_type":"markdown","source":"# Features & responders","metadata":{}},{"cell_type":"markdown","source":"The organizer of the competition has provided us with characteristic descriptions of the features and responders. These descriptions are impersonal, simply designated as tags with numbers.","metadata":{}},{"cell_type":"code","source":"feat_props = pl.read_csv('/kaggle/input/jane-street-real-time-market-data-forecasting/features.csv')\nresp_props = pl.read_csv('/kaggle/input/jane-street-real-time-market-data-forecasting/responders.csv')\n\nfeat_props.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:44:19.894364Z","iopub.execute_input":"2024-12-19T00:44:19.894736Z","iopub.status.idle":"2024-12-19T00:44:19.921263Z","shell.execute_reply.started":"2024-12-19T00:44:19.894705Z","shell.execute_reply":"2024-12-19T00:44:19.919717Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# let's visualize\n(feat_props\n .select(cs.boolean()).cast(pl.Float64).transpose()\n .hvplot.heatmap(height=500,\n                 width=1000,\n                 rot=70,\n                 cmap=['#DCDCDC', '#696969'],\n                 colorbar=False,\n                ).opts(active_tools=['pan'])\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:44:19.922653Z","iopub.execute_input":"2024-12-19T00:44:19.923159Z","iopub.status.idle":"2024-12-19T00:44:20.145902Z","shell.execute_reply.started":"2024-12-19T00:44:19.923111Z","shell.execute_reply":"2024-12-19T00:44:20.144662Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let's check how similar the features are to each other according to this description. We use correlation","metadata":{}},{"cell_type":"code","source":"# similarity of features by tags\ncorr_mtrx = feat_props.select(cs.boolean()).transpose().corr()\n\n(corr_mtrx\n .hvplot.heatmap(height=700,\n                 width=1000,\n                 rot=70,\n                ).opts(active_tools=['pan'])\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:44:20.147266Z","iopub.execute_input":"2024-12-19T00:44:20.147628Z","iopub.status.idle":"2024-12-19T00:44:20.473192Z","shell.execute_reply.started":"2024-12-19T00:44:20.147585Z","shell.execute_reply":"2024-12-19T00:44:20.471897Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let's see if there are any that are absolutely identical","metadata":{}},{"cell_type":"code","source":"eqv_ft = (corr_mtrx\n          .hstack(pl.DataFrame({'feature':[f'feature_{numb:02d}' for numb in range(FEATURES_NUMBER)]}))\n          .rename({f'column_{numb}':f'feature_{numb:02d}' for numb in range(FEATURES_NUMBER)})\n          .unpivot(index='feature')\n          .filter((pl.col('value').round(5)==1) &\n                  (pl.col('feature')!=pl.col('variable'))\n                 )\n         )\n\n(eqv_ft\n .group_by('feature').agg('variable')\n .select(set_ident=pl.col('feature')\n         .list.concat(pl.col('variable'))\n         .list.sort())\n .unique()\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:44:20.474833Z","iopub.execute_input":"2024-12-19T00:44:20.475269Z","iopub.status.idle":"2024-12-19T00:44:20.499073Z","shell.execute_reply.started":"2024-12-19T00:44:20.475223Z","shell.execute_reply":"2024-12-19T00:44:20.497879Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We have six groups of identical details, let's take a closer look at them. We use [previously prepared](https://www.kaggle.com/code/pib73nl/janestreet-rtmdf-auxiliary-data) static descriptions of features. For each group, we will output a table built on the full volume of data, and a boxplot built on a 30% sample. Considering that 30% is 14 million rows, both the box and the whiskers will quite reliably reflect the full sample. Only single outliers may sometimes fall out, but you can always check with tables built on the full data.","metadata":{}},{"cell_type":"code","source":"df_desc = pl.read_parquet('/kaggle/input/janestreet-rtmdf-auxiliary-data/discribe.parquet')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:44:20.500537Z","iopub.execute_input":"2024-12-19T00:44:20.500946Z","iopub.status.idle":"2024-12-19T00:44:20.526410Z","shell.execute_reply.started":"2024-12-19T00:44:20.500910Z","shell.execute_reply":"2024-12-19T00:44:20.525170Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# in the future it will be necessary to work in lazy mode\ndef lazy_df_prep(parts: list) -> pl.LazyFrame:\n    \"\"\"\n    Prepare a LazyFrame from one or more parts (files)\n    \"\"\"\n    \n    list_for_concat = [pl.scan_parquet(x) for x in parts]\n    df = pl.concat(list_for_concat)\n    \n    return df\n\ndef display_description(features):\n    \"\"\"\n    Displays pretty table\n\n    Parametrs\n    -------------\n    feature: list of features for visualize\n    \"\"\"\n    # works only with this dataframe...\n    columns = df_desc['statistic'].to_list()\n    \n    display(df_desc\n            .select(features)\n            .to_pandas()\n            .transpose()\n            .rename(columns={i:name for i, name in enumerate(columns)})\n            .style\n            .format(precision=4)\n            .bar(cmap='RdYlGn_r',\n                 width=50, \n                 height=40,\n                 props='width: 120px; border-left: 0.5px dotted gray;'\n                )\n           )\n\ndef plot_box(df, features, smpl_frac=0.3):\n    \"\"\"\n    Displays a boxplot of the passed features\n\n    Parametrs\n    -------------\n    df: LazyFrame with the necessary features\n    feature: necessary features\n    smpl_frac: full sample proportion (def. - 0.3);\n               if you have a lot of time and RAM, \n               feel free to set smpl_frac=1\n    \n    \"\"\"\n    \n    df = (tmp_df\n          .select(features)\n          .collect()\n          .sample(fraction=smpl_frac)\n          .to_pandas()\n         )\n    \n    fig, axs = plt.subplots(nrows=1, ncols=2, figsize=(15, 4))\n    flierprops = {'markerfacecolor':'red',\n                  'markeredgecolor':'none',\n                  'markersize':2\n                 }\n    axs[0].set_title(f'{features[0]}-{features[-1]} (w/o outliers)')\n    df.plot.box(rot=70,\n                showfliers=False,\n                ax=axs[0],\n                flierprops=flierprops\n               )\n    axs[1].set_title(f'{features[0]}-{features[-1]} (with outliers)')\n    df.plot.box(rot=70,\n                ax=axs[1],\n                flierprops=flierprops\n               )","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-12-19T00:44:20.527596Z","iopub.execute_input":"2024-12-19T00:44:20.527942Z","iopub.status.idle":"2024-12-19T00:44:20.540882Z","shell.execute_reply.started":"2024-12-19T00:44:20.527910Z","shell.execute_reply":"2024-12-19T00:44:20.539494Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# let's gather all the data into one LazyFrame \ntmp_df = lazy_df_prep(train_files)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:44:20.542442Z","iopub.execute_input":"2024-12-19T00:44:20.542907Z","iopub.status.idle":"2024-12-19T00:44:20.559460Z","shell.execute_reply.started":"2024-12-19T00:44:20.542855Z","shell.execute_reply":"2024-12-19T00:44:20.558270Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Features 09-11**\n\nLooking at the heatmap above, you can see that for this set of features, only one tag is set - `tag_1`. Moreover, no other features include this tag. You can see that all the values ​​of these features are integers. Probably these are categorical features, although some have a large number of categories.","metadata":{}},{"cell_type":"code","source":"display_description([f'feature_{numb:02d}' for numb in range(9,12)])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:44:20.561148Z","iopub.execute_input":"2024-12-19T00:44:20.561900Z","iopub.status.idle":"2024-12-19T00:44:20.595589Z","shell.execute_reply.started":"2024-12-19T00:44:20.561850Z","shell.execute_reply":"2024-12-19T00:44:20.594468Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_box(tmp_df, [f'feature_{numb:02d}' for numb in range(9,12)])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:44:20.596968Z","iopub.execute_input":"2024-12-19T00:44:20.597300Z","iopub.status.idle":"2024-12-19T00:44:30.287857Z","shell.execute_reply.started":"2024-12-19T00:44:20.597268Z","shell.execute_reply":"2024-12-19T00:44:30.281735Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Features 21-31**\n\nHere is a similar situation - for a set of features only `tag_0` is \"turned on\", and this tag is not used by any other features. As for these features, it is difficult to guess what might distinguish them from others just by looking at the statistics.","metadata":{}},{"cell_type":"code","source":"display_description([f'feature_{numb:02d}' for numb in range(21,32)])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:44:30.295999Z","iopub.execute_input":"2024-12-19T00:44:30.297726Z","iopub.status.idle":"2024-12-19T00:44:30.422432Z","shell.execute_reply.started":"2024-12-19T00:44:30.297601Z","shell.execute_reply":"2024-12-19T00:44:30.419087Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_box(tmp_df, [f'feature_{numb:02d}' for numb in range(21,32)])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:44:30.425415Z","iopub.execute_input":"2024-12-19T00:44:30.426365Z","iopub.status.idle":"2024-12-19T00:45:01.303383Z","shell.execute_reply.started":"2024-12-19T00:44:30.426327Z","shell.execute_reply":"2024-12-19T00:45:01.302204Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Features 37-38**\n\nHere too, only one tag is turned on - `tag_3`, but, unlike the previous cases, this tag is also used by other features","metadata":{}},{"cell_type":"code","source":"display_description([f'feature_{numb:02d}' for numb in [37,38]])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:45:01.305036Z","iopub.execute_input":"2024-12-19T00:45:01.305386Z","iopub.status.idle":"2024-12-19T00:45:01.334128Z","shell.execute_reply.started":"2024-12-19T00:45:01.305353Z","shell.execute_reply":"2024-12-19T00:45:01.332915Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_box(tmp_df, [f'feature_{numb:02d}' for numb in range(37,39)])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:45:01.335356Z","iopub.execute_input":"2024-12-19T00:45:01.335655Z","iopub.status.idle":"2024-12-19T00:45:08.808778Z","shell.execute_reply.started":"2024-12-19T00:45:01.335626Z","shell.execute_reply":"2024-12-19T00:45:08.807595Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Features 73-78**\n\n\nHere we will consider three groups at once. They are united by `tag_8`, and they differ in the use of tags 12, 13 and 14, respectively.","metadata":{}},{"cell_type":"code","source":"display_description([f'feature_{numb:02d}' for numb in range(73,79)])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:45:08.810239Z","iopub.execute_input":"2024-12-19T00:45:08.810563Z","iopub.status.idle":"2024-12-19T00:45:08.842156Z","shell.execute_reply.started":"2024-12-19T00:45:08.810532Z","shell.execute_reply":"2024-12-19T00:45:08.840844Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_box(tmp_df, [f'feature_{numb:02d}' for numb in range(73,79)])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:45:08.844030Z","iopub.execute_input":"2024-12-19T00:45:08.845018Z","iopub.status.idle":"2024-12-19T00:45:35.912539Z","shell.execute_reply.started":"2024-12-19T00:45:08.844954Z","shell.execute_reply":"2024-12-19T00:45:35.911232Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now we will do approximately the same with **responders**","metadata":{}},{"cell_type":"code","source":"(resp_props\n .select(cs.boolean()).cast(pl.Float64).transpose()\n .hvplot.heatmap(height=400,\n                 width=600,\n                 rot=70,\n                 cmap=['#696969', '#DCDCDC'],\n                 colorbar=False,\n                 grid=True\n                ).opts(active_tools=['pan'])\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:45:35.914153Z","iopub.execute_input":"2024-12-19T00:45:35.914535Z","iopub.status.idle":"2024-12-19T00:45:36.085269Z","shell.execute_reply.started":"2024-12-19T00:45:35.914497Z","shell.execute_reply":"2024-12-19T00:45:36.084164Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# similarity of responders by tags\ncorr_mtrx = resp_props.select(cs.boolean()).transpose().corr()\n\n(corr_mtrx\n .hvplot.heatmap(height=400,\n                 width=600,\n                 rot=70,\n                ).opts(active_tools=['pan'])\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:45:36.086674Z","iopub.execute_input":"2024-12-19T00:45:36.087043Z","iopub.status.idle":"2024-12-19T00:45:36.244919Z","shell.execute_reply.started":"2024-12-19T00:45:36.087010Z","shell.execute_reply":"2024-12-19T00:45:36.243589Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There are no absolutely identical responders here. But there are not many of them in general, so we will show them all.","metadata":{}},{"cell_type":"code","source":"display_description([f'responder_{numb}' for numb in range(RESPONDERS_NUMBER)])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:45:36.246301Z","iopub.execute_input":"2024-12-19T00:45:36.246632Z","iopub.status.idle":"2024-12-19T00:45:36.278615Z","shell.execute_reply.started":"2024-12-19T00:45:36.246600Z","shell.execute_reply":"2024-12-19T00:45:36.277519Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_box(tmp_df, [f'responder_{numb}' for numb in range(RESPONDERS_NUMBER)])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:45:36.279945Z","iopub.execute_input":"2024-12-19T00:45:36.280265Z","iopub.status.idle":"2024-12-19T00:46:07.887361Z","shell.execute_reply.started":"2024-12-19T00:45:36.280235Z","shell.execute_reply":"2024-12-19T00:46:07.886161Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"An interesting picture... One and a half interquartile ranges fit into the interval [-1.5; 1.5] And in all directions there are densely distributed thin tails up to -5 and 5 respectively. What are we predicting, anyway? A change in some indicators?...","metadata":{}},{"cell_type":"markdown","source":"We've looked at correlation from the point of view of grouping tags, now let's look at correlation from the point of view of the values ​​themselves. To do this, we will load the [pre-calculated](https://www.kaggle.com/code/pib73nl/janestreet-rtmdf-auxiliary-data) data again.","metadata":{}},{"cell_type":"code","source":"feat_corr = np.load('/kaggle/input/janestreet-rtmdf-auxiliary-data/feat_corr.npy')\nresp_corr = np.load('/kaggle/input/janestreet-rtmdf-auxiliary-data/resp_corr.npy')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:46:07.889172Z","iopub.execute_input":"2024-12-19T00:46:07.889643Z","iopub.status.idle":"2024-12-19T00:46:07.920780Z","shell.execute_reply.started":"2024-12-19T00:46:07.889592Z","shell.execute_reply":"2024-12-19T00:46:07.918945Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This time let's start with responders","metadata":{}},{"cell_type":"code","source":"(pl.DataFrame(resp_corr)\n .hvplot.heatmap(height=400,\n                 width=500,\n                 rot=90,\n                 # xticks=[(numb, f'responder_{numb:02d}') for numb in range(RESPONDERS_NUMBER)],\n                 yticks=[(numb, f'responder_{numb:02d}') for numb in range(RESPONDERS_NUMBER)],\n                ).opts(active_tools=['pan'])\n)","metadata":{"execution":{"iopub.status.busy":"2024-12-19T00:46:07.922435Z","iopub.execute_input":"2024-12-19T00:46:07.923003Z","iopub.status.idle":"2024-12-19T00:46:08.092368Z","shell.execute_reply.started":"2024-12-19T00:46:07.922960Z","shell.execute_reply":"2024-12-19T00:46:08.091217Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There are no very high correlations. The desired `feature_06` has the highest correlation with `feature_03`","metadata":{}},{"cell_type":"code","source":"def plot_scattermatrix(df, features, smpl_frac=0.3):\n    df = (df\n          .select(features)\n          .collect()\n          .sample(fraction=smpl_frac)\n          .to_pandas()\n         )\n    display(scatter_matrix(df,\n                           alpha=0.2, \n                           rasterize=True,\n                           tools=['pan'],\n                           hist_kwds={'yformatter':frmt_big_numb}\n                          )\n           )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:46:08.094072Z","iopub.execute_input":"2024-12-19T00:46:08.094526Z","iopub.status.idle":"2024-12-19T00:46:08.101680Z","shell.execute_reply.started":"2024-12-19T00:46:08.094477Z","shell.execute_reply":"2024-12-19T00:46:08.100462Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_scattermatrix(tmp_df, [f'responder_{numb}' for numb in (3,6)])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:46:08.103125Z","iopub.execute_input":"2024-12-19T00:46:08.103495Z","iopub.status.idle":"2024-12-19T00:46:15.126891Z","shell.execute_reply.started":"2024-12-19T00:46:08.103439Z","shell.execute_reply":"2024-12-19T00:46:15.125252Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It seems like a cloud of positive correlation is forming, but the outliers live their own lives...\n\nAnd what about features?","metadata":{}},{"cell_type":"code","source":"(pl.DataFrame(feat_corr)\n .hvplot.heatmap(height=900,\n                 width=1000,\n                 rot=90,\n                 # xticks=[(numb, f'responder_{numb:02d}') for numb in range(FEATURES_NUMBER)],\n                 yticks=[(numb, f'responder_{numb:02d}') for numb in range(FEATURES_NUMBER)],\n                ).opts(active_tools=['pan'])\n)","metadata":{"execution":{"iopub.status.busy":"2024-12-19T00:46:15.129100Z","iopub.execute_input":"2024-12-19T00:46:15.129630Z","iopub.status.idle":"2024-12-19T00:46:15.419195Z","shell.execute_reply.started":"2024-12-19T00:46:15.129565Z","shell.execute_reply":"2024-12-19T00:46:15.417903Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This is an awesome meditation mat for true lovers of searching for secret signs. How about the repeating pattern of features 0-3 and 32-35? Let's take a closer look at the correlated features of these groups","metadata":{}},{"cell_type":"code","source":"plot_scattermatrix(tmp_df, [f'feature_{numb:02d}' for numb in (0, 2, 3)])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:46:15.420860Z","iopub.execute_input":"2024-12-19T00:46:15.421202Z","iopub.status.idle":"2024-12-19T00:46:30.888635Z","shell.execute_reply.started":"2024-12-19T00:46:15.421170Z","shell.execute_reply":"2024-12-19T00:46:30.886581Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_scattermatrix(tmp_df, [f'feature_{numb:02d}' for numb in (32, 34, 35)])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:46:30.891078Z","iopub.execute_input":"2024-12-19T00:46:30.891702Z","iopub.status.idle":"2024-12-19T00:46:45.170406Z","shell.execute_reply.started":"2024-12-19T00:46:30.891620Z","shell.execute_reply":"2024-12-19T00:46:45.167784Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The correlation is very high. Especially for the requisites 0,1,3, where even the outliers are correlated. You can safely get rid of something.\n\nTo avoid scrolling up, we will display the tags of these groups of features here, to understand the intra- and inter-group similarities and differences","metadata":{}},{"cell_type":"code","source":"# let's take a look at tagging\nfeat_props_t = feat_props.select(cs.boolean()).cast(pl.Float64).transpose()\nplot_list = []\ngroups = ((0,2,3), (32,34,35))\n\nfor numbers in groups:\n    plot_list.append(feat_props_t.select([f'column_{numb}' for numb in numbers])\n                     .hvplot.heatmap(height=300,\n                                     width=200,\n                                     rot=70,\n                                     cmap=['#DCDCDC', '#696969'],\n                                     colorbar=False,\n                                    ).opts(active_tools=['pan'], shared_axes=False)\n                    )\n\n# for i in range(len(groups)):     \n#     if i == 0:\n#         grid = plot_list[i].opts(shared_axes=False)\n#     else:\n#         grid += plot_list[i].opts(shared_axes=False)\n# grid\n\npn.GridBox(*plot_list, ncols=2)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:46:45.172709Z","iopub.execute_input":"2024-12-19T00:46:45.173859Z","iopub.status.idle":"2024-12-19T00:46:45.472538Z","shell.execute_reply.started":"2024-12-19T00:46:45.173782Z","shell.execute_reply":"2024-12-19T00:46:45.471314Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There is an intra-group difference in tags 12, 13, 14, and it is the same for both groups. The groups differ only in tag 16, apparently it strengthens the correlation in the first group :)\n\nFeel free to explore other suspicious groups :)","metadata":{}},{"cell_type":"markdown","source":"Well, just in case, let's check the correlation of the features with the `responder_6`. What if...","metadata":{}},{"cell_type":"code","source":"corr_feat_w_resp = [tmp_df.select(pl.corr(f'feature_{numb:02d}', 'responder_6')).collect().item() \n                    for numb \n                    in tqdm(range(FEATURES_NUMBER), total=FEATURES_NUMBER, ncols=80)\n                   ]\nprint(f'max negative correlation: {min(corr_feat_w_resp):7.4f}\\nmax positive correlation: {max(corr_feat_w_resp):7.4f}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:46:45.473980Z","iopub.execute_input":"2024-12-19T00:46:45.474296Z","iopub.status.idle":"2024-12-19T00:47:53.747651Z","shell.execute_reply.started":"2024-12-19T00:46:45.474265Z","shell.execute_reply":"2024-12-19T00:47:53.746524Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"No miracle happened. Not at all...","metadata":{}},{"cell_type":"markdown","source":"# A case from the lives of symbols","metadata":{}},{"cell_type":"markdown","source":"Well, and finally, let's take some random financial instrument and look at the behavior of its randomly selected details and responders. Although, to present the differences, we will take the features not random. But you can choose any others, to your taste","metadata":{}},{"cell_type":"code","source":"np.random.seed(42)\n\n# taking random symbol_id\nsymbol_id = np.random.randint(SYMBOL_NUMBER)\nprint(f'{symbol_id = }')\n\n# feel free to take random features\n# features = np.random.choice(range(FEATURES_NUMBER+1), size=(3), replace=False)\n# features = sorted([f'feature_{numb:02d}' for numb in features])\n\nfeatures = ['feature_00', 'feature_30', 'feature_44'] # check out these features to illustrate the diversity\nprint(f'{features = }')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:47:53.748936Z","iopub.execute_input":"2024-12-19T00:47:53.749253Z","iopub.status.idle":"2024-12-19T00:47:53.758560Z","shell.execute_reply.started":"2024-12-19T00:47:53.749223Z","shell.execute_reply":"2024-12-19T00:47:53.757416Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time\ndf_one_symbol = tmp_df.filter(pl.col('symbol_id')==symbol_id).collect()\nprint(df_one_symbol.shape)\ndf_one_symbol.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:47:53.759981Z","iopub.execute_input":"2024-12-19T00:47:53.760309Z","iopub.status.idle":"2024-12-19T00:48:07.734593Z","shell.execute_reply.started":"2024-12-19T00:47:53.760278Z","shell.execute_reply":"2024-12-19T00:48:07.733281Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_grid(df:pl.DataFrame, feature: str, filter_expr: pl.Expr=True):\n    \n    gspec = pn.GridSpec(width=1200, height=1000)\n    gspec[:1, :4] = daily_quotes_ohlc(df, feature=feature)\n    gspec[1, :2] = plot_tick_line(df, feature=feature, filter_expr=filter_expr)\n    gspec[2, :2] = lag_plot(df, feature=feature, filter_expr=filter_expr)\n    gspec[1:, 2:] = bivar_n_hist(df, x=feature, filter_expr=filter_expr)\n    \n    return gspec\n\ndef daily_quotes_ohlc(df:pl.DataFrame, feature: str):\n    fig = (df\n           .sort(['date_id', 'time_id'])\n           .group_by('date_id', maintain_order=True)\n           .agg(pl.col(feature).first().alias('open'),\n                pl.col(feature).max().alias('high'),\n                pl.col(feature).min().alias('low'),\n                pl.col(feature).last().alias('close')\n               )\n          ).hvplot.ohlc('date_id', \n                        ['open', 'high', 'low', 'close'],\n                        width=1200,\n                        title='Daily quotes for the entire period'\n                       )\n    return fig\n\ndef plot_tick_line(df:pl.DataFrame, feature: str, filter_expr: pl.Expr=True):\n    fig = (df\n           .filter(filter_expr)\n           .hvplot.line(y=feature, \n                        rasterize=True,\n                        colorbar=False,\n                        title='Tick quotes for the filtered period'\n                       ).opts(active_tools=['pan'], xformatter=frmt_big_numb,)\n          )\n    return fig\n\ndef lag_plot(df:pl.DataFrame, feature: str, filter_expr: pl.Expr=True, lag: int=1):\n    fig = hvplot.plotting.lag_plot(df\n                                    .filter(filter_expr)\n                                    .select(feature)\n                                    .to_pandas(),\n                                   legend=False,\n                                   alpha=0.7, \n                                   lag=lag,\n                                   title=f'Lag plot for the filtered period (lag={lag})'\n                                   ).opts(active_tools=['pan'])\n    return fig\n\ndef bivar_n_hist(df:pl.DataFrame, \n                 x: str,\n                 y: str='responder_6',\n                 filter_expr: pl.Expr=True):\n\n    bivar = (df\n             .filter(filter_expr)\n             .select([x, y])\n             .hvplot.bivariate(colorbar=False,\n                               width=500,\n                               height=500\n                              )\n            )\n    \n    fig = (bivar\n           .hist(dimension=[x, y])\n           .opts(opts.Histogram(line_color=None,\n                                xformatter=frmt_big_numb,\n                                yformatter=frmt_big_numb,\n                                active_tools=['pan']\n                               ),\n                 opts.Bivariate(active_tools=['pan'])\n                )\n           .opts(title='Relationship between the selected feature and the responder for the filtered period',\n                 fontsize={'title':'10pt'},\n                )\n          )\n    return fig\n    ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:48:07.739211Z","iopub.execute_input":"2024-12-19T00:48:07.740133Z","iopub.status.idle":"2024-12-19T00:48:07.829474Z","shell.execute_reply.started":"2024-12-19T00:48:07.740007Z","shell.execute_reply":"2024-12-19T00:48:07.823418Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# by dates we filter from the end, since many financial instruments are missing at the beginning\nlast_date = 1699 # number of dates in the entire period\nlen_per = 300\nfilter_date = pl.col('date_id').is_between(last_date-len_per, last_date)\n\nplot_grid(df_one_symbol, feature=features[0], filter_expr=filter_date)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:48:07.840132Z","iopub.execute_input":"2024-12-19T00:48:07.847140Z","iopub.status.idle":"2024-12-19T00:49:22.050162Z","shell.execute_reply.started":"2024-12-19T00:48:07.846657Z","shell.execute_reply":"2024-12-19T00:49:22.048410Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"`Feature_00` is a bright representative of features with pronounced seasonality. Monthly seasonality is visible to the naked eye, but there is also weekly seasonality, which is not so noticeable (see the last section in version 1 for more details). \n\nThe *lag-plot* convincingly illustrates the thesis that the weather tomorrow will be the same as today in 85 percent of cases. But there are still 15% that come as a surprise.\n\nConsidering that we have already seen the distribution of the `responder_6` and the magnitude of the correlation with it of all features, the relationship of their values ​​is quite expected. It can be described as follows - any value of the responder can correspond to any value of the requisite :)","metadata":{}},{"cell_type":"code","source":"plot_grid(df_one_symbol, feature=features[1], filter_expr=filter_date)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:49:22.052386Z","iopub.execute_input":"2024-12-19T00:49:22.053227Z","iopub.status.idle":"2024-12-19T00:50:35.609284Z","shell.execute_reply.started":"2024-12-19T00:49:22.053121Z","shell.execute_reply":"2024-12-19T00:50:35.608002Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"`Feature_30` does not change during the day, which is visible on the daily candlestick chart - the open, close, high and low prices are equal to each other. On the lag plot we see a noticeable straight line, which indicates that the indicator may not change for several days in a row, which is confirmed by the tick chart.","metadata":{}},{"cell_type":"code","source":"plot_grid(df_one_symbol, feature=features[2], filter_expr=filter_date)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:50:35.610932Z","iopub.execute_input":"2024-12-19T00:50:35.611336Z","iopub.status.idle":"2024-12-19T00:51:46.841936Z","shell.execute_reply.started":"2024-12-19T00:50:35.611299Z","shell.execute_reply":"2024-12-19T00:51:46.840532Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This is just some kind of noise... Well, at least the lag plot doesn't look like a rectangle, which gives some kind of illusory hope","metadata":{}},{"cell_type":"markdown","source":"## ...and responders","metadata":{}},{"cell_type":"markdown","source":"Well, how could we do without our `responder_6`! Let's take a look at its behavior for the selected financial instrument","metadata":{}},{"cell_type":"code","source":"last_date = 1699 # number of dates in the entire period\nlen_per = 300\nfilter_date = pl.col('date_id').is_between(last_date-len_per, last_date)\n\ngspec = pn.GridSpec(width=1200, height=700)\ngspec[0, :2] = daily_quotes_ohlc(df_one_symbol, feature='responder_6')\ngspec[1, 0] = plot_tick_line(df_one_symbol, feature='responder_6', filter_expr=filter_date)\ngspec[1, 1] = lag_plot(df_one_symbol, feature='responder_6', filter_expr=filter_date)\n\ngspec","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:51:46.843790Z","iopub.execute_input":"2024-12-19T00:51:46.844278Z","iopub.status.idle":"2024-12-19T00:51:55.193760Z","shell.execute_reply.started":"2024-12-19T00:51:46.844230Z","shell.execute_reply":"2024-12-19T00:51:55.191951Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It's also too noisy... But the lag plot is still holding on to high correlation. With a lag of one tick. Let's try to increase the lag.","metadata":{}},{"cell_type":"code","source":"lags_plot = []\nfor lag in [5, 15, 20, 25]: \n    lags_plot.append(lag_plot(df_one_symbol, feature='responder_6', filter_expr=filter_date, lag=lag))\n\npn.GridBox(*lags_plot, ncols=2)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:51:55.196757Z","iopub.execute_input":"2024-12-19T00:51:55.197561Z","iopub.status.idle":"2024-12-19T00:52:13.837449Z","shell.execute_reply.started":"2024-12-19T00:51:55.197479Z","shell.execute_reply":"2024-12-19T00:52:13.836109Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print('autocorrelation:\\n')\nfor i, lag in enumerate([5, 15, 20, 25]):\n    print(f' lag {lag:>2} = {lags_plot[i].data[\"y(t)\"].corr(lags_plot[i].data[f\"y(t + {lag})\"]):>7.4f}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:52:13.839214Z","iopub.execute_input":"2024-12-19T00:52:13.839616Z","iopub.status.idle":"2024-12-19T00:52:13.873332Z","shell.execute_reply.started":"2024-12-19T00:52:13.839579Z","shell.execute_reply":"2024-12-19T00:52:13.871983Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We see that by the 20th tick the correlation is completely lost. And in a day we have 900+ ticks...","metadata":{}},{"cell_type":"code","source":"# %%time\n# # If you have 10 minutes, you can run this and see how the autocorrelation decreases step by step\n# ax = sm.graphics.tsa.plot_acf(df_one_symbol['responder_6'], lags=25)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-19T00:52:13.874897Z","iopub.execute_input":"2024-12-19T00:52:13.875269Z","iopub.status.idle":"2024-12-19T01:01:37.003300Z","shell.execute_reply.started":"2024-12-19T00:52:13.875236Z","shell.execute_reply":"2024-12-19T01:01:37.002101Z"}},"outputs":[],"execution_count":null}]}