{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84493,"databundleVersionId":9871156,"sourceType":"competition"}],"dockerImageVersionId":30804,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"!pip install  --upgrade seaborn\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T18:54:44.182089Z","iopub.execute_input":"2024-12-04T18:54:44.182498Z","iopub.status.idle":"2024-12-04T18:54:57.188305Z","shell.execute_reply.started":"2024-12-04T18:54:44.182462Z","shell.execute_reply":"2024-12-04T18:54:57.186722Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!pip install  --upgrade matplotlib\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T18:54:57.190945Z","iopub.execute_input":"2024-12-04T18:54:57.191483Z","iopub.status.idle":"2024-12-04T18:55:29.800118Z","shell.execute_reply.started":"2024-12-04T18:54:57.191425Z","shell.execute_reply":"2024-12-04T18:55:29.798742Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom scipy.stats import jarque_bera, probplot\nfrom statsmodels.tsa.stattools import adfuller\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom matplotlib.ticker import FuncFormatter\nimport os\nfrom pathlib import Path","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-12-04T18:55:29.801891Z","iopub.execute_input":"2024-12-04T18:55:29.802288Z","iopub.status.idle":"2024-12-04T18:55:31.176121Z","shell.execute_reply.started":"2024-12-04T18:55:29.802248Z","shell.execute_reply":"2024-12-04T18:55:31.174734Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"root = Path('/kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet')\ndfs = []\nsymbol = 8 # Let's choose a symbol\n\nfor partition_number in np.arange(0, 10, 1):\n    train_path = root / f'partition_id={partition_number}'\n\n    filename = train_path / 'part-0.parquet' \n    print(f'Loading {filename}')\n    df = pd.read_parquet(filename)\n\n    mask = (df.symbol_id == symbol)\n    df = df.loc[mask].copy()\n    df['part_num'] = partition_number\n    dfs.append(df)\n\n\ndf = pd.concat(dfs)\ndf.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T18:55:31.179561Z","iopub.execute_input":"2024-12-04T18:55:31.180105Z","iopub.status.idle":"2024-12-04T18:57:14.178914Z","shell.execute_reply.started":"2024-12-04T18:55:31.180068Z","shell.execute_reply":"2024-12-04T18:57:14.177386Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n\ndef percentage_formatter(x, _):\n    return f\"{x*100:.0f}%\"\n\ntarget_variable = 'responder_6'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T18:57:14.180757Z","iopub.execute_input":"2024-12-04T18:57:14.181089Z","iopub.status.idle":"2024-12-04T18:57:14.187334Z","shell.execute_reply.started":"2024-12-04T18:57:14.181059Z","shell.execute_reply":"2024-12-04T18:57:14.186216Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%matplotlib inline\npd.set_option('display.max_rows', None)\npd.set_option('display.max_columns', None)\npd.set_option('use_inf_as_na', None)\nsns.set(rc={'figure.figsize': (10, 6)})\n\nimport warnings\nwarnings.filterwarnings(\"ignore\", category=FutureWarning, module=\"seaborn\")\nsns.set_context(\"paper\", rc={\"figure.figsize\": (3, 1.5)})\nsns.set_theme(style=\"whitegrid\")\ncustom_palette = sns.cubehelix_palette(start=.1, rot=-.7, reverse=True)\nsns.set_palette(palette=custom_palette)\nimage_path = Path('./analysis/images')\n\n# themes\nplt.rcParams['axes.labelcolor'] = 'gray' \nplt.rcParams['axes.titlecolor'] = 'gray'   # Disable the top spine\nplt.rcParams['axes.titlelocation'] = 'left'   # Disable the top spine\nplt.rcParams['axes.spines.top'] = False   # Disable the top spine\nplt.rcParams['axes.spines.right'] = False  # Disable the right spine\nplt.rcParams['axes.spines.left'] = True   # Enable the left spine\nplt.rcParams['axes.spines.bottom'] = True  # Enable the bottom spine\nplt.rcParams['axes.edgecolor'] = 'gray'  # Set spine color to black\nplt.rcParams['xtick.color'] = 'gray'     # Bottom tick color\nplt.rcParams['ytick.color'] = 'gray'     # Left tick color\nplt.rcParams['xtick.bottom'] = True       # Enable bottom ticks\nplt.rcParams['ytick.left'] = True         # Enable left ticks\nplt.rcParams['grid.color'] = 'lightgray'\nplt.rcParams['grid.linewidth'] = 0.5","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T18:57:14.188772Z","iopub.execute_input":"2024-12-04T18:57:14.189124Z","iopub.status.idle":"2024-12-04T18:57:14.216197Z","shell.execute_reply.started":"2024-12-04T18:57:14.189092Z","shell.execute_reply":"2024-12-04T18:57:14.214988Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Purpose\n\nThe purpose of this notebook is to explore the target variable of the competition, focusing on a single symbol in isolation (in this case, 8, chosen randomly). The goal is to apply relevant topics to other financial instruments in order to identify associations.\n\n# Target variable distribution\nHowever, this will be the first notebook to provide context for the analysis. Let’s begin by examining the distribution by partitioned dataframe.","metadata":{}},{"cell_type":"code","source":"useful = [target_variable, 'part_num']\n\n# plot utils\nmin_target = df[target_variable].min()\nmax_target = df[target_variable].max()\nquant_975 = df[target_variable].quantile(.975)\nquant_025 = df[target_variable].quantile(.025)\n\nxticks = np.arange(min_target-1, max_target+2, 1)\nyticks = np.arange(0.1, 1.1, 0.1)\nformatter = FuncFormatter(percentage_formatter)\n\n# plot\ng = sns.histplot(\n    x=target_variable, hue='part_num', data=df[useful],\n    multiple='stack', kde=True, bins=20, stat='density', palette='crest')\n\n# formatting\ng.set_title('Target variable distribution by partitioned dataframe'.upper())\ng.set_xticks(xticks)\ng.set_yticks(yticks)\ng.yaxis.set_major_formatter(formatter)\ng.axvline(x=quant_975, color='gray', linestyle='dotted')\ng.axvline(x=quant_025, color='gray', linestyle='dotted')\ng.text(quant_975*1.05, 0.5, f'q:97.5={quant_975.round(2)}', color='gray', ha='left')\ng.text(quant_025*1.05, 0.5, f'q:2.5={quant_025.round(2)}', color='gray', ha='right')\n\ng.grid(False, axis='x')\ng.grid(True, axis='y')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T18:57:14.217801Z","iopub.execute_input":"2024-12-04T18:57:14.218245Z","iopub.status.idle":"2024-12-04T18:57:22.223324Z","shell.execute_reply.started":"2024-12-04T18:57:14.218193Z","shell.execute_reply":"2024-12-04T18:57:22.222168Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Statistical test for normality\nstat, p_value = jarque_bera(df[target_variable].values)\nprint(f\"Jarque-Bera Test Statistic: {stat}, p-value: {p_value}\")\nif p_value > 0.05:\n    print(\"Fail to reject the null hypothesis: Data is normally distributed.\")\nelse:\n    print(\"Reject the null hypothesis: Data is not normally distributed.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T18:57:22.224917Z","iopub.execute_input":"2024-12-04T18:57:22.225268Z","iopub.status.idle":"2024-12-04T18:57:22.259971Z","shell.execute_reply.started":"2024-12-04T18:57:22.225234Z","shell.execute_reply":"2024-12-04T18:57:22.258898Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Distribution notes\n\nThe first thing we can observe is a certain symmetry, an average around zero, and that 95% of the values provided by responder_6 range from -1.67 to 1.93. Additionally, we notice long tails, which are characteristic of financial markets.\n\nMoreover, if we perform the Jarque-Bera test to evaluate whether the distribution is normal, we can note that it is not. It is important to mention that we use this test because it is better suited for financial markets, as it places special emphasis on kurtosis and skewness. For more information, you can visit: [Jarque-Bera Test - Wikipedia](https://en.wikipedia.org/wiki/Jarque%E2%80%93Bera_test)\n","metadata":{}},{"cell_type":"markdown","source":"# Target variable through date and time\n\nWe can observe that the target variable behaves similarly across all partitions, at least at first glance.\n","metadata":{}},{"cell_type":"code","source":"df.sort_values(['date_id', 'time_id'], inplace=True)\ndf['idx_col'] = np.arange(len(df))\n\ng = sns.lineplot(\n    data=df,\n    x='idx_col',\n    y=target_variable,\n    hue='part_num'\n)\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T18:57:22.261486Z","iopub.execute_input":"2024-12-04T18:57:22.261862Z","iopub.status.idle":"2024-12-04T18:57:33.049302Z","shell.execute_reply.started":"2024-12-04T18:57:22.261826Z","shell.execute_reply":"2024-12-04T18:57:33.047620Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Time\n\nLet’s examine the average of our target variable (y-axis) and its standard deviation by `time_id` (x-axis), considering all partitions.\n","metadata":{}},{"cell_type":"code","source":"t = sns.relplot(\n    data=df,\n    x='time_id',\n    y=target_variable,\n    legend=True,\n    kind='line',\n    errorbar='sd'\n)\n\nax = plt.gca()\nax.set_title('Target variable by time_id'.upper())\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T18:57:33.054237Z","iopub.execute_input":"2024-12-04T18:57:33.054623Z","iopub.status.idle":"2024-12-04T18:57:37.802405Z","shell.execute_reply.started":"2024-12-04T18:57:33.054580Z","shell.execute_reply":"2024-12-04T18:57:37.801242Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Apparently, there is more volatility at the beginning and end of the period. Additionally, the averages tend to be higher at the start and decrease toward the end.\n\nDoes this occur in all partitions? Let’s find out.","metadata":{}},{"cell_type":"code","source":"t1 = sns.relplot(\n    data=df,\n    x='time_id',\n    y=target_variable,\n    col='part_num',\n    hue='part_num',\n    legend=False,\n    kind='line',\n    col_wrap=4,\n    errorbar='sd',\n    palette='crest'\n)\nplt.show()\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T18:57:37.804274Z","iopub.execute_input":"2024-12-04T18:57:37.804654Z","iopub.status.idle":"2024-12-04T18:57:49.029607Z","shell.execute_reply.started":"2024-12-04T18:57:37.804598Z","shell.execute_reply":"2024-12-04T18:57:49.028435Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Alright, we can observe that:\n* The first three partitions have fewer records than the others.  \n* Partition number 3 shows significant volatility in its averages during the last 200 `time_ids`.  \n* The volatility pattern remains consistent across all partitions.  \n\nLet’s divide the `time_id` variable into 10 equal parts to see if this behavior becomes more evident.","metadata":{}},{"cell_type":"code","source":"df['time_id_bin'] = pd.cut(df['time_id'], 10)\ndf['time_id_bin_idx'] = df.time_id_bin.cat.codes\n\ndata = df.groupby('time_id_bin_idx')[target_variable].std()\ndata=data*2\n\nuseful = ['time_id_bin_idx', target_variable, 'part_num']\nfig, ax = plt.subplots()\n\ng = sns.boxenplot(\n    data=df[useful],\n    x='time_id_bin_idx',\n    y=target_variable,\n    hue='time_id_bin_idx',\n    palette='mako',\n    legend=False,\n    alpha=0.5,\n    ax=ax,\n)\ng1 = sns.lineplot(\n    data=data.reset_index(),\n    x='time_id_bin_idx',\n    y=target_variable,\n    ax=ax,\n    color='crimson',\n    linestyle='--',\n    legend=False,\n    linewidth=3\n)\ndata = data*-1\ng2 = sns.lineplot(\n    data=data.reset_index(),\n    x='time_id_bin_idx',\n    y=target_variable,\n    ax=ax,\n    color='crimson',\n    linestyle='--',\n    legend=False,\n    linewidth=3\n)\ng.set_title('Target variable per time_id bin')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T18:57:49.031139Z","iopub.execute_input":"2024-12-04T18:57:49.031482Z","iopub.status.idle":"2024-12-04T18:57:51.907532Z","shell.execute_reply.started":"2024-12-04T18:57:49.031448Z","shell.execute_reply":"2024-12-04T18:57:51.906316Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Indeed, we observe greater volatility at the beginning and end of the period.","metadata":{}},{"cell_type":"markdown","source":"## Adding atributes\n\nBefore proceeding with the analysis, let’s add some relevant variables:\n\n1. **cum_target**: As we noticed, the target variable seems to be differentiated. By calculating the cumulative sum, we can observe the trend of the responses across all partitions.  \n2. **daily_cum_target**: Similar to the previous attribute, but it only accumulates the target variable for each `date_id`.  \n3. **imbalance**: The sign (positive or negative) of each recorded value in the target variable.  \n4. **cum_imbalance**: The cumulative sum of `imbalance`.  \n5. **daily_cum_imbalance**: The cumulative sum of `imbalance` grouped by `date_id`.  \n","metadata":{}},{"cell_type":"code","source":"df['cum_target'] = df[target_variable].cumsum()\ndf['cum_daily_target'] = df.groupby(['date_id'])[target_variable].cumsum()\ndf['imbalance'] = np.sign(df[target_variable])\ndf['cum_imbalance'] = df['imbalance'].cumsum()\ndf['cum_daily_imbalace'] = df.groupby(['date_id'])['imbalance'].cumsum()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T18:57:51.909162Z","iopub.execute_input":"2024-12-04T18:57:51.909615Z","iopub.status.idle":"2024-12-04T18:57:51.994902Z","shell.execute_reply.started":"2024-12-04T18:57:51.909553Z","shell.execute_reply":"2024-12-04T18:57:51.993753Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"g = sns.lineplot(\n    data=df,\n    x='idx_col',\n    y='cum_target',\n)\ng.axhline(0, color='gray', linestyle='dashed', linewidth=3)\ng.set_title('Cumulative target value per period')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T18:57:51.996433Z","iopub.execute_input":"2024-12-04T18:57:51.996934Z","iopub.status.idle":"2024-12-04T18:57:57.667706Z","shell.execute_reply.started":"2024-12-04T18:57:51.996884Z","shell.execute_reply":"2024-12-04T18:57:57.666698Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"When analyzing the accumulation of the target variable, we can observe a bearish trend during the initial periods, followed by a bullish trend.\n\nThe total accumulation of all values ends up being positive. This could be because we have more positive imbalances than negative ones. Let’s see what the distribution shows:\n","metadata":{}},{"cell_type":"code","source":"data = df.imbalance.value_counts(normalize=True)\ndata.reset_index()\ng = sns.barplot(\n    data=data.reset_index(),\n    x='imbalance',\n    y='proportion',\n)\ng.set_title('Imbalance distribution')\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T18:57:57.669153Z","iopub.execute_input":"2024-12-04T18:57:57.669488Z","iopub.status.idle":"2024-12-04T18:57:57.933329Z","shell.execute_reply.started":"2024-12-04T18:57:57.669455Z","shell.execute_reply":"2024-12-04T18:57:57.932154Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"After reviewing the imbalances, we can conclude that there are more negative values than positive ones. How does the accumulation of imbalances look over time?\n","metadata":{}},{"cell_type":"code","source":"g = sns.lineplot(data=df, x='idx_col', y='cum_imbalance')\ng.axhline(0, color='gray', linestyle='dashed', linewidth=3)\ng.set_title('Cumulative Imbalance per period')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T18:57:57.934551Z","iopub.execute_input":"2024-12-04T18:57:57.934893Z","iopub.status.idle":"2024-12-04T18:58:03.086686Z","shell.execute_reply.started":"2024-12-04T18:57:57.934860Z","shell.execute_reply":"2024-12-04T18:58:03.085523Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Alright, we definitely see an accumulation driven by negative imbalances, despite having a bullish trend—interesting.\n\nHow does the imbalance relate to the target variable on a daily basis?\n","metadata":{}},{"cell_type":"code","source":"data = df.groupby(['date_id', 'imbalance'])[target_variable].agg(['mean', 'count'])\ndata = data.reset_index()\n\nxticks = np.arange(-2, 2.1, 0.5)\nyticks = np.arange(200, 801, 100)\nymed_line = data['count'].median()\nprint(ymed_line)\n\ng = sns.kdeplot(\n    data=data.reset_index(),\n    x='mean',\n    y='count',\n    hue='imbalance',\n    palette='crest',\n    fill=True,\n)\ng.set_xticks(xticks)\ng.set_yticks(yticks)\ng.axhline(ymed_line, color='gray', linestyle='dashed', linewidth=3)\ng.axvline(0.5, color='gray', linestyle='dashed', linewidth=3)\ng.axvline(-0.5, color='gray', linestyle='dashed', linewidth=3)\n\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T19:00:01.413014Z","iopub.execute_input":"2024-12-04T19:00:01.413473Z","iopub.status.idle":"2024-12-04T19:00:03.839165Z","shell.execute_reply.started":"2024-12-04T19:00:01.413437Z","shell.execute_reply":"2024-12-04T19:00:03.838038Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Alright, what do we see here? The positive imbalances observed in a day fall below 461 (the overall median of daily accumulation). However, these positive imbalances are farther from zero compared to the negative ones.\n\nThis suggests that while accumulation skews toward pessimism, the positive values tend to be larger.\n\nHow does this relationship look throughout the day (`date_id` attribute)?\n","metadata":{}},{"cell_type":"code","source":"g = sns.lineplot(df, x='time_id', y='cum_daily_imbalace', linewidth=3, legend=True)\ng = sns.lineplot(df, x='time_id', y='cum_daily_target', linewidth=3, legend=True)\ng.axhline(0, color='gray', linestyle='dashed', linewidth=3)\ng.set_title('Daily bases: Cumulative imbalance vs Cumulative Daily target')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T19:12:21.891188Z","iopub.execute_input":"2024-12-04T19:12:21.891657Z","iopub.status.idle":"2024-12-04T19:13:17.982489Z","shell.execute_reply.started":"2024-12-04T19:12:21.891585Z","shell.execute_reply":"2024-12-04T19:13:17.981435Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The target variable can accumulate, on average, between 0 and 10 units. However, the expected daily imbalance is -50.\n\nHow does it accumulate overall?\n","metadata":{}},{"cell_type":"code","source":"g = sns.lineplot(data=df, x='idx_col', y='cum_imbalance')\ng1 = sns.lineplot(data=df, x='idx_col', y='cum_target')\ng.set_title('All dataset: Cumulative imbalance vs Cumulative target')\ng.axhline(0, color='gray', linestyle='dashed', linewidth=3)\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T19:01:36.195173Z","iopub.execute_input":"2024-12-04T19:01:36.195589Z","iopub.status.idle":"2024-12-04T19:01:47.112162Z","shell.execute_reply.started":"2024-12-04T19:01:36.195549Z","shell.execute_reply":"2024-12-04T19:01:47.110877Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Considering all periods, this relationship becomes more pronounced.\n\nEarlier, we noticed higher volatility at the beginning and end of the day by dividing `time_id` into ten equal parts. Does this variable affect the relationship we just observed?\n","metadata":{}},{"cell_type":"code","source":"data = df.groupby(['date_id','time_id_bin_idx', 'imbalance'])[target_variable].agg(['mean', 'count'])\ndata = data.reset_index()\n\nxref = data.groupby('imbalance')['mean'].median()\nxref_line_1 = xref[-1]\nxref_line_2 = xref[1]\nyref_line = data['count'].median()\n\n\ng = sns.FacetGrid(data, col='time_id_bin_idx', hue='imbalance',\n                  col_wrap=4)\ng.map_dataframe(sns.kdeplot, x='mean', y='count', fill=True)\ng.add_legend()\ng.refline(x=xref_line_1)\ng.refline(x=xref_line_2)\ng.refline(y=yref_line)\ng.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T19:01:47.113875Z","iopub.execute_input":"2024-12-04T19:01:47.114226Z","iopub.status.idle":"2024-12-04T19:02:13.764300Z","shell.execute_reply.started":"2024-12-04T19:01:47.114192Z","shell.execute_reply":"2024-12-04T19:02:13.763115Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The wedge-like pattern amplifies the difference between imbalances and the target variable. Initially, there’s a tendency for positive imbalances to have a higher median. Over time, this behavior flattens, and in the final periods, the relationship reverses, with negative imbalances significantly outweighing positive ones.\n","metadata":{}},{"cell_type":"markdown","source":"### Summary:\n\n1. The target variable exhibits symmetry across all partitions, although it does not follow a normal distribution.  \n2. Throughout each day, we can observe higher volatility at the beginning and end of the period.  \n3. Partition number 3 shows significant volatility in its averages during the last 200 `time_ids`.  \n4. The total accumulation of all values ends up being positive.  \n5. There are more negative values than positive ones.  \n6. Daily accumulation skews toward pessimism, but the positive values tend to be larger.  \n7. The target variable can accumulate, on average, between 0 and 10 units. However, the expected daily imbalance is -50.  \n8. The wedge-like pattern amplifies the difference between imbalances and the target variable. Initially, there’s a tendency for positive imbalances to have a higher median. Over time, this behavior flattens, and in the final periods, the relationship reverses, with negative imbalances significantly outweighing positive ones.  \n\n---\n\n### What’s Next?\n\nIn the next notebook, we will analyze the relationship between different symbols while considering these variables. Afterward, we will start exploring attribute engineering mechanisms.\n\nIf you found this useful, please don’t forget to upvote—it motivates me greatly to keep publishing. Also, any feedback is always welcome!\n","metadata":{}}]}