{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84493,"databundleVersionId":9849268,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"A very interesting (large) dataset! Note that most of the analysis shared will be on a fraction of train data, since I wanted to share it early on to help others speed their EDA up.\n\nIn summary:\n> We have two types of features: mostly numerical, and a few categorical features.  \n> Some responder variables exhibit an interesting pattern of being more correlated.   \n> Volume of data requires a more sophisticated approach to training if one wants to rely on free tier provided by Kaggle.  \n> On top of standard data preprocessing, as some features exhibit high levels of correlation, regularisation and/or some kind of dimentionality reduction might prove useful, at least, for simpler (TSA) models. ","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n#         print(os.path.join(dirname, filename))\n        pass\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-10-15T12:55:53.959421Z","iopub.execute_input":"2024-10-15T12:55:53.959854Z","iopub.status.idle":"2024-10-15T12:55:55.273047Z","shell.execute_reply.started":"2024-10-15T12:55:53.959793Z","shell.execute_reply":"2024-10-15T12:55:55.272060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport math\n\n# plotting stuff\nfrom pandas.plotting import lag_plot\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.express as px\nimport plotly.graph_objects as go\n\n# system\nimport warnings\nwarnings.filterwarnings('ignore')\n# for the image import\n\nfrom IPython.display import Image\n# garbage collector to keep RAM in check\nimport gc \nimport polars as pl\n\n","metadata":{"execution":{"iopub.status.busy":"2024-10-15T12:55:55.275004Z","iopub.execute_input":"2024-10-15T12:55:55.275566Z","iopub.status.idle":"2024-10-15T12:55:58.114513Z","shell.execute_reply.started":"2024-10-15T12:55:55.275517Z","shell.execute_reply":"2024-10-15T12:55:58.113375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ROOT_DIR = \"/kaggle/input/jane-street-real-time-market-data-forecasting\"","metadata":{"execution":{"iopub.status.busy":"2024-10-15T12:55:58.116352Z","iopub.execute_input":"2024-10-15T12:55:58.117373Z","iopub.status.idle":"2024-10-15T12:55:58.122499Z","shell.execute_reply.started":"2024-10-15T12:55:58.117317Z","shell.execute_reply":"2024-10-15T12:55:58.121332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = (\n    pl.read_parquet(f\"{ROOT_DIR}/train.parquet\", n_rows=500000)\n)\ntrain.shape","metadata":{"execution":{"iopub.status.busy":"2024-10-15T12:55:58.124659Z","iopub.execute_input":"2024-10-15T12:55:58.125104Z","iopub.status.idle":"2024-10-15T12:55:59.505794Z","shell.execute_reply.started":"2024-10-15T12:55:58.125053Z","shell.execute_reply":"2024-10-15T12:55:59.504740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nimport math\n\n# Function to automatically deduce the plot type and create subplots\ndef create_auto_seaborn_subplots(data, columns, figsize=(12, 8), **kwargs):\n    \"\"\"\n    Automatically deduces the type of plot for each column or pair of columns \n    and creates subplots accordingly.\n\n    Parameters:\n    - data: DataFrame containing the data\n    - columns: Iterable of column names or tuples of column names (for pair plots)\n    - figsize: Size of the entire figure (width, height)\n    - **kwargs: Additional arguments passed to the Seaborn plotting functions\n    \"\"\"\n    # Number of plots\n    num_plots = len(columns)\n    \n    # Automatically determine number of rows and columns based on the number of plots\n    cols = min(math.ceil(math.sqrt(num_plots)), 4)\n    rows = math.ceil(num_plots / cols)\n\n    # Create a figure and axes for subplots\n    fig, axes = plt.subplots(rows, cols, figsize=figsize)\n    axes = axes.flatten()\n\n    # Loop through each column or pair of columns and deduce plot type\n    for i, col in enumerate(columns):\n        ax = axes[i]\n        col_dtype = data.schema[col]\n\n        if isinstance(col, tuple) and len(col) == 2:  # If two columns are provided, do scatter plot\n            sns.scatterplot(data=data, x=col[0], y=col[1], ax=ax, **kwargs)\n        else:  # If a single column, check if it's categorical or continuous\n#             if pd.api.types.is_numeric_dtype(data[col]):  # Numeric data: Histogram\n            if isinstance(col_dtype, (pl.Int8, pl.Int16, pl.Int32, pl.Int64, pl.Float32, pl.Float64)):\n    \n                sns.histplot(data=data, x=col, ax=ax, **kwargs)\n            else:  # Categorical data: Count plot\n                sns.countplot(data=data, x=col, ax=ax, **kwargs)\n\n#         ax.set_title(f\"{col}\")\n\n    # Hide any unused subplots\n    for ax in axes[num_plots:]:\n        ax.remove()\n    for ax in axes: \n        ax.set_yscale('log')\n\n    # Adjust layout\n    plt.tight_layout()\n    plt.title('Distribution of target variable candidates')\n    plt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-15T13:04:18.853013Z","iopub.status.idle":"2024-10-15T13:04:18.853456Z","shell.execute_reply.started":"2024-10-15T13:04:18.853254Z","shell.execute_reply":"2024-10-15T13:04:18.853275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"feature_cols = [col for col in train.columns if 'feature' in col]","metadata":{"execution":{"iopub.status.busy":"2024-10-15T12:56:34.969443Z","iopub.execute_input":"2024-10-15T12:56:34.970260Z","iopub.status.idle":"2024-10-15T12:56:34.974945Z","shell.execute_reply.started":"2024-10-15T12:56:34.970210Z","shell.execute_reply":"2024-10-15T12:56:34.973927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictor_cols = [col for col in train.columns if 'responder' in col]","metadata":{"execution":{"iopub.status.busy":"2024-10-15T13:48:09.795729Z","iopub.execute_input":"2024-10-15T13:48:09.796586Z","iopub.status.idle":"2024-10-15T13:48:09.801165Z","shell.execute_reply.started":"2024-10-15T13:48:09.796532Z","shell.execute_reply":"2024-10-15T13:48:09.800172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_train = train.sample(fraction=0.1)","metadata":{"execution":{"iopub.status.busy":"2024-10-15T12:56:29.644207Z","iopub.execute_input":"2024-10-15T12:56:29.644595Z","iopub.status.idle":"2024-10-15T12:56:29.704447Z","shell.execute_reply.started":"2024-10-15T12:56:29.644556Z","shell.execute_reply":"2024-10-15T12:56:29.703290Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here, we can observe that for some response variables, the clamping had a larger effect than for others (for ex, repsponder_1 esd less affected than all the others). ","metadata":{}},{"cell_type":"code","source":"create_auto_seaborn_subplots(sample_train, predictor_cols, figsize=(10, 8))","metadata":{"execution":{"iopub.status.busy":"2024-10-15T12:56:46.921918Z","iopub.execute_input":"2024-10-15T12:56:46.922924Z","iopub.status.idle":"2024-10-15T12:56:58.429204Z","shell.execute_reply.started":"2024-10-15T12:56:46.922883Z","shell.execute_reply":"2024-10-15T12:56:58.427984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import missingno as msno\nmsno.matrix(sample_train.to_pandas())  # Convert to pandas for visualization\n","metadata":{"execution":{"iopub.status.busy":"2024-10-15T14:06:06.191935Z","iopub.execute_input":"2024-10-15T14:06:06.192400Z","iopub.status.idle":"2024-10-15T14:06:09.813656Z","shell.execute_reply.started":"2024-10-15T14:06:06.192360Z","shell.execute_reply":"2024-10-15T14:06:09.812286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sample = sample_train.sample(fraction=0.1).to_pandas()","metadata":{"execution":{"iopub.status.busy":"2024-10-15T13:04:37.302356Z","iopub.execute_input":"2024-10-15T13:04:37.302749Z","iopub.status.idle":"2024-10-15T13:04:37.314356Z","shell.execute_reply.started":"2024-10-15T13:04:37.302714Z","shell.execute_reply":"2024-10-15T13:04:37.313387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # Create a PairGrid to customize the axes\n# g = sns.PairGrid(df_sample)\n\n# # Use scatter plot for off-diagonal, hist for diagonal\n# g.map_offdiag(sns.scatterplot)\n# g.map_diag(sns.histplot)\n\n# # Set y-axis to log scale for all axes\n# for ax in g.axes.flatten():\n#     ax.set_yscale('log')\n\n# plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-15T13:47:21.444672Z","iopub.execute_input":"2024-10-15T13:47:21.445078Z","iopub.status.idle":"2024-10-15T13:47:21.449821Z","shell.execute_reply.started":"2024-10-15T13:47:21.445038Z","shell.execute_reply":"2024-10-15T13:47:21.448839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Some columns are not very useful in our (small random) sample, let's remove those (either Null or show the partition number). \n* Create a composite id to function as a timestamp and sort the dataset\n* Look at (average as well as individual) response variable correlations","metadata":{}},{"cell_type":"markdown","source":"It is interesting how 3 days apart target variables are more correlated ","metadata":{}},{"cell_type":"code","source":"cleaned_sample_train = sample_train.select([col for col in df_sample.columns if sample_train.select(pl.col(col).n_unique())[0, 0] > 1])","metadata":{"execution":{"iopub.status.busy":"2024-10-15T13:46:53.419252Z","iopub.execute_input":"2024-10-15T13:46:53.420309Z","iopub.status.idle":"2024-10-15T13:46:53.629071Z","shell.execute_reply.started":"2024-10-15T13:46:53.420261Z","shell.execute_reply":"2024-10-15T13:46:53.628149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cleaned_sample_train1 = cleaned_sample_train.with_columns(\n    (pl.col('date_id').cast(pl.Utf8) + pl.lit(\" \") + pl.col('time_id').cast(pl.Utf8).str.zfill(6)).alias('composite_id')\n)","metadata":{"execution":{"iopub.status.busy":"2024-10-15T13:46:53.913789Z","iopub.execute_input":"2024-10-15T13:46:53.914235Z","iopub.status.idle":"2024-10-15T13:46:53.927532Z","shell.execute_reply.started":"2024-10-15T13:46:53.914194Z","shell.execute_reply":"2024-10-15T13:46:53.926532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cleaned_sample_train = cleaned_sample_train.with_columns(\n    (pl.col('date_id') * 1000000 + pl.col('time_id')).alias('composite_id')\n)\n","metadata":{"execution":{"iopub.status.busy":"2024-10-15T13:46:54.534646Z","iopub.execute_input":"2024-10-15T13:46:54.535405Z","iopub.status.idle":"2024-10-15T13:46:54.540285Z","shell.execute_reply.started":"2024-10-15T13:46:54.535362Z","shell.execute_reply":"2024-10-15T13:46:54.539318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cleaned_sample_train = cleaned_sample_train.sort(['composite_id', 'symbol_id'])","metadata":{"execution":{"iopub.status.busy":"2024-10-15T13:47:02.588541Z","iopub.execute_input":"2024-10-15T13:47:02.588977Z","iopub.status.idle":"2024-10-15T13:47:02.608643Z","shell.execute_reply.started":"2024-10-15T13:47:02.588934Z","shell.execute_reply":"2024-10-15T13:47:02.607455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cleaned_sample_train.sort(['date_id', 'time_id', 'symbol_id'])","metadata":{"execution":{"iopub.status.busy":"2024-10-15T13:11:26.448563Z","iopub.execute_input":"2024-10-15T13:11:26.448981Z","iopub.status.idle":"2024-10-15T13:11:26.481635Z","shell.execute_reply.started":"2024-10-15T13:11:26.448940Z","shell.execute_reply":"2024-10-15T13:11:26.480461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cleaned_sample_train.filter(cleaned_sample_train['symbol_id'] == 1).shape","metadata":{"execution":{"iopub.status.busy":"2024-10-15T13:21:57.018693Z","iopub.execute_input":"2024-10-15T13:21:57.019698Z","iopub.status.idle":"2024-10-15T13:21:57.031541Z","shell.execute_reply.started":"2024-10-15T13:21:57.019654Z","shell.execute_reply":"2024-10-15T13:21:57.030551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate the correlation matrix\ncorr = cleaned_sample_train.filter(cleaned_sample_train['symbol_id'] == 1)[predictor_cols].corr()\n\n# Create a mask for the upper triangle\nmask = np.triu(np.ones_like(corr, dtype=bool), k=1)\n\n# Create a clustermap with the mask applied\nsns.heatmap(corr, mask=mask, cmap='coolwarm', );","metadata":{"execution":{"iopub.status.busy":"2024-10-15T14:45:43.036139Z","iopub.execute_input":"2024-10-15T14:45:43.036573Z","iopub.status.idle":"2024-10-15T14:45:43.424935Z","shell.execute_reply.started":"2024-10-15T14:45:43.036533Z","shell.execute_reply":"2024-10-15T14:45:43.423873Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate the correlation matrix\ncorr = cleaned_sample_train.filter(cleaned_sample_train['symbol_id'] == 17)[predictor_cols].corr()\n\n# Create a mask for the upper triangle\nmask = np.triu(np.ones_like(corr, dtype=bool), k=1)\n\n# Create a clustermap with the mask applied\nsns.heatmap(corr, mask=mask, cmap='coolwarm', );","metadata":{"execution":{"iopub.status.busy":"2024-10-15T14:46:03.845964Z","iopub.execute_input":"2024-10-15T14:46:03.846407Z","iopub.status.idle":"2024-10-15T14:46:04.223602Z","shell.execute_reply.started":"2024-10-15T14:46:03.846365Z","shell.execute_reply":"2024-10-15T14:46:04.222383Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"correlations = []\n\n# Group by symbol_id and calculate correlation\nfor symbol_id, group in cleaned_sample_train.group_by('symbol_id'):\n    # Convert group to a DataFrame to calculate correlation\n    corr_matrix = group[predictor_cols].corr()\n    correlations.append(corr_matrix)","metadata":{"execution":{"iopub.status.busy":"2024-10-15T13:54:52.997960Z","iopub.execute_input":"2024-10-15T13:54:52.998501Z","iopub.status.idle":"2024-10-15T13:54:53.033058Z","shell.execute_reply.started":"2024-10-15T13:54:52.998457Z","shell.execute_reply":"2024-10-15T13:54:53.032146Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate the average correlation matrix\navg_corr_matrix = np.median(correlations, axis=0)\n\n# Create a mask for the upper triangle\nmask = np.triu(np.ones_like(avg_corr_matrix, dtype=bool), k=1)\n\n# Create a heatmap with the mask applied\nplt.figure(figsize=(8, 6))\nsns.heatmap(avg_corr_matrix, mask=mask, cmap='coolwarm', annot=True, \n            xticklabels=predictor_cols, yticklabels=predictor_cols)\nplt.title('Average Correlation Heatmap')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-15T14:03:59.512948Z","iopub.execute_input":"2024-10-15T14:03:59.513384Z","iopub.status.idle":"2024-10-15T14:04:01.253577Z","shell.execute_reply.started":"2024-10-15T14:03:59.513339Z","shell.execute_reply":"2024-10-15T14:04:01.252418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# from ydata_profiling import ProfileReport","metadata":{"execution":{"iopub.status.busy":"2024-10-15T12:03:56.536776Z","iopub.execute_input":"2024-10-15T12:03:56.537836Z","iopub.status.idle":"2024-10-15T12:03:56.543583Z","shell.execute_reply.started":"2024-10-15T12:03:56.537789Z","shell.execute_reply":"2024-10-15T12:03:56.541550Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nfor target in predictor_cols: \n    plt.plot((cleaned_sample_train.to_pandas().groupby(['date_id'])[target].mean()).cumsum())\n\nplt.xlabel('Trade days')\nplt.ylabel('Cumulative target')\nplt.title('Target distribution over trade days')\nplt.grid(visible=True)\nplt.legend(predictor_cols, loc='lower left', bbox_to_anchor=(0.05, -.5), title='Targets', ncol=3)\nsns.despine()\n\n","metadata":{"execution":{"iopub.status.busy":"2024-10-15T14:21:12.104555Z","iopub.execute_input":"2024-10-15T14:21:12.104974Z","iopub.status.idle":"2024-10-15T14:21:14.973561Z","shell.execute_reply.started":"2024-10-15T14:21:12.104935Z","shell.execute_reply":"2024-10-15T14:21:14.972548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # fig, ax = plt.subplots(figsize=(15, 5))\n# for target in predictor_cols:\n#     balance= pd.Series(train.to_pandas()[target]).cumsum()\n#     ax.set_xlabel (\"Trade\", fontsize=18)\n#     ax.set_ylabel (\"Cumulative responder 0\", fontsize=18);\n#     balance.plot(lw=3);\n#     del balance\n#     gc.collect();","metadata":{"execution":{"iopub.status.busy":"2024-10-15T14:45:21.849915Z","iopub.execute_input":"2024-10-15T14:45:21.850400Z","iopub.status.idle":"2024-10-15T14:45:21.855356Z","shell.execute_reply.started":"2024-10-15T14:45:21.850359Z","shell.execute_reply":"2024-10-15T14:45:21.854286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_feature_cols = [col for col in cleaned_sample_train.columns if 'feature' in col]","metadata":{"execution":{"iopub.status.busy":"2024-10-15T14:30:26.659648Z","iopub.execute_input":"2024-10-15T14:30:26.660648Z","iopub.status.idle":"2024-10-15T14:30:26.665472Z","shell.execute_reply.started":"2024-10-15T14:30:26.660603Z","shell.execute_reply":"2024-10-15T14:30:26.664270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_feature_cols_ticks = [f.split('_')[1] for f in new_feature_cols ]","metadata":{"execution":{"iopub.status.busy":"2024-10-15T14:31:59.000286Z","iopub.execute_input":"2024-10-15T14:31:59.000693Z","iopub.status.idle":"2024-10-15T14:31:59.005473Z","shell.execute_reply.started":"2024-10-15T14:31:59.000654Z","shell.execute_reply":"2024-10-15T14:31:59.004454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"correlations = []\n\n# Group by symbol_id and calculate correlation\nfor symbol_id, group in cleaned_sample_train.drop_nulls().group_by('symbol_id'):\n    # Convert group to a DataFrame to calculate correlation\n    corr_matrix = group[new_feature_cols].corr()\n    correlations.append(corr_matrix)\n    \n# Calculate the average correlation matrix\navg_corr_matrix = np.median(correlations, axis=0)\n\n# Create a mask for the upper triangle\nmask = np.triu(np.ones_like(avg_corr_matrix, dtype=bool), k=1)\n\n# Create a heatmap with the mask applied\nplt.figure(figsize=(18, 16))\nsns.heatmap(avg_corr_matrix, mask=mask, cmap='coolwarm', \n#             annot=True, \n            xticklabels=new_feature_cols_ticks, yticklabels=new_feature_cols_ticks)\nplt.title('Average Correlation Heatmap for Features')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-15T14:41:06.565177Z","iopub.execute_input":"2024-10-15T14:41:06.565583Z","iopub.status.idle":"2024-10-15T14:41:07.940513Z","shell.execute_reply.started":"2024-10-15T14:41:06.565546Z","shell.execute_reply":"2024-10-15T14:41:07.939114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here, we can see that feature_09, feature_10, feature_11 are _different_ - these are presumable categorical variables. There are also some clusters of variables which we might apply some regularisation of dimensionality reduction to when finetuning a model. ","metadata":{"execution":{"iopub.status.busy":"2024-10-15T14:42:07.811207Z","iopub.execute_input":"2024-10-15T14:42:07.811646Z","iopub.status.idle":"2024-10-15T14:42:07.830772Z","shell.execute_reply.started":"2024-10-15T14:42:07.811605Z","shell.execute_reply":"2024-10-15T14:42:07.829617Z"}}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}