{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":84493,"databundleVersionId":9871156,"sourceType":"competition"}],"dockerImageVersionId":30823,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"adb001c3-e0cf-499f-92a6-540735c4ada6","cell_type":"markdown","source":"## Jane Street Competition EDA - Understand your Data\n\nThe goal of this tutorial is to provide you with a comprehensive understanding of different techniques used to explore and analyze time-series data, which is crucial for building effective forecasting models.\n\nWe will cover a wide range of EDA techniques, including:\n\n* **Missing Value Analysis:** Understanding the extent and pattern of missing data in the dataset.\n\n* **Correlation Analysis:** Exploring relationships between different features and target variables.\n\n* **Distribution Analysis:** Visualizing the distribution of key variables to understand their behavior.\n\n* **Time-Series Decomposition:** Breaking down time-series data into trend, seasonality, and residual components.\n\n* **Frequency Analysis:** Using Fourier Transform to analyze the frequency components of the data.\n\n* **Wavelet Decomposition:** A more advanced technique to analyze time-series data at different scales.\n\nBy the end of this tutorial, you will have a solid foundation in EDA techniques that can be applied to any forecasting problem","metadata":{"tags":[],"execution":{"iopub.status.busy":"2025-01-14T00:18:12.015446Z","iopub.execute_input":"2025-01-14T00:18:12.015737Z","iopub.status.idle":"2025-01-14T00:18:14.972742Z","shell.execute_reply.started":"2025-01-14T00:18:12.015706Z","shell.execute_reply":"2025-01-14T00:18:14.972046Z"}}},{"id":"b7e04828-e744-4adf-97bd-f6b017708176","cell_type":"code","source":"#importing relevant libraries\n\nimport pandas as pd\nimport numpy as np\nimport polars as pl\nfrom matplotlib import pyplot as plt\nimport seaborn as sns\nfrom pathlib import Path","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"b6d464ea-35c0-461e-905c-95780f89d4ac","cell_type":"markdown","source":"Let's quickly peak through the feature, responders, and also the training set quickly to understand broad features of the data we're working with ","metadata":{}},{"id":"ab59d24c-725c-4d59-b42b-f69c6757da95","cell_type":"code","source":"data_path = Path(\"/kaggle/input/jane-street-real-time-market-data-forecasting\")\n\nfeatures = pd.read_csv(data_path/'features.csv')\nfeatures.head()","metadata":{"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-01-14T00:47:34.897906Z","iopub.execute_input":"2025-01-14T00:47:34.898219Z","iopub.status.idle":"2025-01-14T00:47:34.924360Z","shell.execute_reply.started":"2025-01-14T00:47:34.898197Z","shell.execute_reply":"2025-01-14T00:47:34.923652Z"}},"outputs":[],"execution_count":null},{"id":"3bc8ee67-13dc-4782-94ee-f936f7c548bd","cell_type":"code","source":"responders = pd.read_csv(data_path/'responders.csv')\nresponders.head()","metadata":{"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-01-14T00:48:12.807470Z","iopub.execute_input":"2025-01-14T00:48:12.807769Z","iopub.status.idle":"2025-01-14T00:48:12.821525Z","shell.execute_reply.started":"2025-01-14T00:48:12.807745Z","shell.execute_reply":"2025-01-14T00:48:12.820576Z"}},"outputs":[],"execution_count":null},{"id":"68463533-cdf3-4110-b5b9-3521bc58a9bd","cell_type":"code","source":"train_file = pd.read_parquet(data_path/'train.parquet'/'partition_id=0'/'part-0.parquet')\ntrain_file.head()","metadata":{"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-01-14T00:48:24.061312Z","iopub.execute_input":"2025-01-14T00:48:24.061604Z","iopub.status.idle":"2025-01-14T00:48:25.216384Z","shell.execute_reply.started":"2025-01-14T00:48:24.061579Z","shell.execute_reply":"2025-01-14T00:48:25.215460Z"}},"outputs":[],"execution_count":null},{"id":"1dd6faff-789a-4ec2-886a-999dc1a1c76a","cell_type":"markdown","source":"The training data has 4 types of columns:\n1. {date_id}, {time_id}, {symbol_id} - it becomes your primary key going into the data.\n2. weight: this column defines the relative contribution of each return on your forecast, which will primarily become a part of the loss, so we can calculate a weighted loss, as not all decisions are equally weighted.\n3. features: The known input features, which will be available at the test time.\n4. responder: Along with our actual column to be predicted: responder_6, we have addiitonal information to select the correlated features as additional info available for the predictor which is also called: constant_dynamic_features. ","metadata":{}},{"id":"983d4870-8ab8-4542-b93f-289bef1b4213","cell_type":"markdown","source":"### Missing Data Analysis\n\nFor this purpose, we use a library called: missingo, which is great for looking at the missing nature of the tabular data including time-series. \n\nWhen dealing with tabular data, missing data has a lot of information, that can indicate a lot about the nature of the underlying data. It may also indicate which features contribute to a higher degree towards target. \n\nIn the forecasting, these features along with any missing data need to be tackled with the right instruments, and treating missing data is a vast field, and treating missing values requires a separate discussion of its own, which we will discuss in future ","metadata":{}},{"id":"d0a8ecd7-477f-43dd-a278-6783601b08a8","cell_type":"markdown","source":"The missingno matrix provides a visual representation of missing values in the dataset. Each white line represents a missing value, while the black lines indicate non-missing data. This visualization helps us quickly identify which features have missing data and whether there are any patterns in the missingness (e.g., if certain features are missing together).","metadata":{}},{"id":"372b8dd4-a509-41b3-a610-11a4d988cf4d","cell_type":"code","source":"import missingno\nfeature_columns = [f\"feature_{feature_num:02d}\" for feature_num in range(79)]\nmissingno.matrix(train_file[feature_columns[0:50]].sample(1000))\nmissingno.matrix(train_file[feature_columns[50:]].sample(1000))\nplt.show()","metadata":{"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-01-14T00:18:31.601473Z","iopub.execute_input":"2025-01-14T00:18:31.601774Z","iopub.status.idle":"2025-01-14T00:18:33.665062Z","shell.execute_reply.started":"2025-01-14T00:18:31.601752Z","shell.execute_reply":"2025-01-14T00:18:33.664236Z"}},"outputs":[],"execution_count":null},{"id":"35d1f2fd-c997-4481-89a0-e04c1c2311c0","cell_type":"markdown","source":"We can see feature_21, feature_26, and others including some from the start are completely missing, and are NAs. We can remove these, as these can not be filled, nor can be tackled at the training time. ","metadata":{}},{"id":"7b71c20b-8178-4859-9ebb-8d3f7de1175a","cell_type":"markdown","source":"Understanding the pattern of missing data is crucial for deciding how to handle it (e.g., imputation, deletion). If missingness is random, simple imputation techniques may suffice. However, if missingness is systematic, more sophisticated methods may be required.","metadata":{}},{"id":"390ea10e-6b35-4128-a5c3-0994601f244d","cell_type":"code","source":"missingno.bar(train_file[feature_columns])","metadata":{"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-01-14T00:18:37.040321Z","iopub.execute_input":"2025-01-14T00:18:37.040682Z","iopub.status.idle":"2025-01-14T00:18:39.740468Z","shell.execute_reply.started":"2025-01-14T00:18:37.040651Z","shell.execute_reply":"2025-01-14T00:18:39.739310Z"}},"outputs":[],"execution_count":null},{"id":"31e5fea4-f1f7-44e0-999c-1042ed05e22a","cell_type":"markdown","source":"The bar plot shows the total number of non-missing values for each feature. Features with shorter bars have more missing data.\n\nThis plot helps prioritize which features may need more attention during data preprocessing. Features with a high percentage of missing values might need to be dropped or imputed carefully.","metadata":{}},{"id":"3608f734-2807-4475-9553-158d2a9c9b1b","cell_type":"code","source":"missingno.heatmap(train_file[feature_columns])","metadata":{"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-01-14T00:18:45.360097Z","iopub.execute_input":"2025-01-14T00:18:45.360399Z","iopub.status.idle":"2025-01-14T00:18:54.061343Z","shell.execute_reply.started":"2025-01-14T00:18:45.360377Z","shell.execute_reply":"2025-01-14T00:18:54.060427Z"}},"outputs":[],"execution_count":null},{"id":"e18eb0c4-e1c4-4103-a996-94f04335815e","cell_type":"markdown","source":"The heatmap shows the correlation of missingness between features. A high correlation (close to 1 or -1) indicates that missingness in one feature is strongly related to missingness in another.\n\nThis can help identify if missingness in certain features is related, which could indicate a systematic issue (e.g., data collection errors).","metadata":{}},{"id":"3c306930-7cea-4b22-8461-248c4b747d28","cell_type":"code","source":"missingno.dendrogram(train_file[feature_columns])","metadata":{"tags":[]},"outputs":[],"execution_count":null},{"id":"0353355b-a120-455d-b826-b7779150a285","cell_type":"markdown","source":"The dendrogram clusters features based on the similarity of their missingness patterns. Features that cluster together have similar patterns of missing data.\n\nThis can help identify groups of features that might be missing together, which could be useful for feature engineering or imputation strategies.","metadata":{}},{"id":"6bc2d8e1-0519-489f-99dd-03b97943a428","cell_type":"markdown","source":"## Correlation Analysis\n","metadata":{}},{"id":"e429b1ff-3d2b-40e5-b8d6-8a23f363dcdc","cell_type":"code","source":"feature_correlation = train_file[feature_columns].corr().fillna(0)\nplt.figure(figsize = (15, 15))\nsns.heatmap(feature_correlation, square = True, cmap = 'coolwarm')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-14T00:59:41.491413Z","iopub.execute_input":"2025-01-14T00:59:41.491746Z","iopub.status.idle":"2025-01-14T01:00:06.389688Z","shell.execute_reply.started":"2025-01-14T00:59:41.491718Z","shell.execute_reply":"2025-01-14T01:00:06.388826Z"}},"outputs":[],"execution_count":null},{"id":"88ae1b01-ee70-4259-b1ad-c275f95e22f1","cell_type":"markdown","source":"The heatmap shows the pairwise correlation between features. High positive or negative values indicate strong relationships between features.\n\nUnderstanding feature correlations is essential for feature selection and to avoid multicollinearity in predictive models. Features with high correlation might be redundant and can be removed to simplify the model.","metadata":{}},{"id":"0c68d338-6311-46fb-a7b0-22629d955f5e","cell_type":"code","source":"from scipy.cluster.hierarchy import dendrogram, linkage\n\ndistance_matrix = 1 - np.abs(feature_correlation)\nlinkage_matrix = linkage(distance_matrix, method = 'ward')\n\nplt.figure(figsize=(20, 10))\ndendrogram(\n    linkage_matrix,\n    labels=feature_correlation.columns,\n    leaf_rotation=90,\n    leaf_font_size=12\n)\nplt.title(\"Dendrogram of Feature Correlations\")\nplt.xlabel(\"Features\")\nplt.ylabel(\"Distance\")\nplt.show()","metadata":{"tags":[]},"outputs":[],"execution_count":null},{"id":"7f513ac5-039d-4079-85a5-edeb9d5f3cd1","cell_type":"markdown","source":"The dendrogram clusters features based on their correlation. Features that are close together in the dendrogram are highly correlated.\n\nThis visualization helps in identifying groups of correlated features, which can be useful for dimensionality reduction or feature engineering.","metadata":{}},{"id":"8a3c3e74-ecd9-48a3-95ff-7806c12cc632","cell_type":"code","source":"plt.figure(figsize = (20, 20))\nsns.clustermap(feature_correlation, cmap='coolwarm', annot=False, method='ward')\nplt.title(\"Heatmap with Dendrogram\")\nplt.show()\n","metadata":{"tags":[]},"outputs":[],"execution_count":null},{"id":"cf9d4013-b2c1-4dae-b13d-9603bbb5f04f","cell_type":"markdown","source":"The clustermap combines the heatmap and dendrogram, showing both the correlation values and the hierarchical clustering of features.\n\nThis provides a comprehensive view of feature relationships, helping to identify clusters of features that behave similarly.","metadata":{}},{"id":"dd228a92-a487-458f-81f1-65b05adc4b45","cell_type":"markdown","source":"### Target Relationship & Distribution","metadata":{"tags":[]}},{"id":"1deca931-801f-491f-92a5-df1795d5ea27","cell_type":"code","source":"plt.figure(figsize = (15, 10))\nsns.boxplot(train_file, x = 'symbol_id', y = 'responder_6', whis = 2.0, showfliers = False)\nplt.title('Distribution of the Different Stock Symbols')\nplt.show()","metadata":{"tags":[]},"outputs":[],"execution_count":null},{"id":"87763c83-bb95-4ff2-b7e9-4c3f1ffab79a","cell_type":"markdown","source":"The boxplot shows the distribution of the target variable (responder_6) across different stock symbols. The box represents the interquartile range (IQR), and the whiskers show the range of the data\n\nThis helps identify if there are significant differences in the target variable across different symbols. If certain symbols have consistently higher or lower values, this could be important for modeling.","metadata":{}},{"id":"3afb9d41-0332-4dd8-b984-3db1aa2f8a79","cell_type":"code","source":"import warnings\nwarnings.simplefilter(\"ignore\")\n\nplt.figure(figsize = (15, 10))\nsns.histplot(train_file,y = 'responder_6', kde = True, hue = 'symbol_id', color = 'grey')\nplt.title('Distribution of the Different Stock Symbols')\nplt.show()","metadata":{"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-01-14T01:01:20.313583Z","iopub.execute_input":"2025-01-14T01:01:20.313956Z","iopub.status.idle":"2025-01-14T01:01:58.945472Z","shell.execute_reply.started":"2025-01-14T01:01:20.313926Z","shell.execute_reply":"2025-01-14T01:01:58.944538Z"}},"outputs":[],"execution_count":null},{"id":"0350dbe5-75ad-42db-bbb5-e26200dc429f","cell_type":"markdown","source":"The histogram shows the distribution of responder_6 for each symbol, with a kernel density estimate (KDE) overlay.\n\nThis visualization helps understand the shape of the distribution (e.g., normal, skewed) and whether there are differences in the distribution across symbols.\n\n","metadata":{}},{"id":"0c1241ed-3604-43af-9a70-040091eb362a","cell_type":"code","source":"plt.figure(figsize = (15, 10))\nsns.histplot(train_file,y = 'responder_6', kde = True, hue = 'symbol_id', stat = 'probability')\nplt.title('Distribution of the Different Stock Symbols')\nplt.show()","metadata":{"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-01-14T01:01:58.946599Z","iopub.execute_input":"2025-01-14T01:01:58.946882Z","iopub.status.idle":"2025-01-14T01:02:38.573947Z","shell.execute_reply.started":"2025-01-14T01:01:58.946857Z","shell.execute_reply":"2025-01-14T01:02:38.573120Z"}},"outputs":[],"execution_count":null},{"id":"edc21abd-bbdc-4acd-9907-de0a514461b9","cell_type":"markdown","source":"This histogram is weighted by the weight column, which might represent the importance or volume of each data point.\n\nWeighting the histogram can provide a more accurate representation of the distribution, especially if some data points are more significant than others.","metadata":{}},{"id":"7f0015d1-f02d-4774-bfa0-99fca9889f8e","cell_type":"code","source":"import warnings\nwarnings.simplefilter('ignore')\n\nplt.figure(figsize = (15, 10))\nsns.histplot(train_file,y = 'responder_6', kde = True, hue = 'symbol_id', stat = 'probability', weights = 'weight')\nplt.title('Distribution of the Different Stock Symbols')\nplt.show()","metadata":{"tags":[]},"outputs":[],"execution_count":null},{"id":"8017fa4a-b1d6-4104-808f-e55d06f63956","cell_type":"markdown","source":"This heatmap shows the correlation of responder_6 across different symbols.\n\nThis can help identify if certain symbols move together, which could be useful for portfolio management or risk assessment.","metadata":{}},{"id":"62fb8a0f-c6d7-49d6-befa-1bf520701171","cell_type":"code","source":"plt.figure(figsize = (15, 10))\nsns.histplot(train_file,y = 'responder_6', kde = True, stat = 'probability', weights = 'weight')\nplt.title('Distribution of the Different Stock Symbols')\nplt.show()","metadata":{"tags":[]},"outputs":[],"execution_count":null},{"id":"8dff9d38-ad80-4250-ad53-69ac92b0105b","cell_type":"markdown","source":"Lastly, lets quickly look at the most probable returns when averaged by weights, and normalized. ","metadata":{}},{"id":"40d45726-8681-41c4-a170-91909e1fcb1f","cell_type":"code","source":"responder_column = [f'responder_{response_num}' for response_num in range(8)]\nresponder_corr = train_file[responder_column].corr()\n\nplt.figure(figsize = (10,10))\nsns.heatmap(responder_corr, annot = True, cmap = 'jet')\nplt.show()","metadata":{"tags":[]},"outputs":[],"execution_count":null},{"id":"1c3fc39c-6911-4a97-814b-9d74afabd979","cell_type":"markdown","source":" The heatmap shows the correlation between different responder variables.\n\nUnderstanding the relationships between different responders can help in feature engineering or in understanding the underlying structure of the data.","metadata":{}},{"id":"b0ba4532-c892-4949-a047-9964b433d8ad","cell_type":"code","source":"symbol_pivot = train_file.pivot(index = ['date_id','time_id'], columns = ['symbol_id'], values = ['responder_6']).reset_index()\nsymbol_pivot.columns = [f\"{first}_{second}\" for first, second in symbol_pivot.columns]\nsymbol_columns = [f\"responder_6_{symbol}\" for symbol in train_file['symbol_id'].unique()]\nsymbol_corr = symbol_pivot[symbol_columns].fillna(0).corr()\nplt.figure(figsize = (20, 20))\nsns.heatmap(symbol_corr, cmap = 'jet', annot=True)\nplt.title('Symbol Wise Corelation for Responder 6')\nplt.legend()\n","metadata":{"tags":[]},"outputs":[],"execution_count":null},{"id":"26e4d8de-fda8-4aaa-a3f0-a5fa969495f5","cell_type":"markdown","source":"This heatmap shows the correlation of responder_6 across different symbols.\n\nThis can help identify if certain symbols move together, which could be useful for portfolio management or risk assessment.","metadata":{}},{"id":"9b2cbd0e-98ed-478e-aebc-9ecb5570207e","cell_type":"code","source":"plt.figure(figsize = (5,10))\nsns.histplot(train_file, y = 'weight')\nplt.show()","metadata":{"tags":[]},"outputs":[],"execution_count":null},{"id":"e43b5262-c435-4629-9f95-b023d5ff46ed","cell_type":"markdown","source":"Let's quickly analyze the weight component to understand its distribution.","metadata":{}},{"id":"01ba315c-c5f9-491d-9bb9-2d96321e78ef","cell_type":"code","source":"plt.figure(figsize = (5,10))\nsns.histplot(train_file, y = 'weight', stat = 'probability')\nplt.show()","metadata":{"tags":[]},"outputs":[],"execution_count":null},{"id":"860b5806-b740-4838-8bcc-7641799c616b","cell_type":"markdown","source":"Also a quick preview of the weight component as a measure of standard probability. ","metadata":{}},{"id":"7b7ac091-c2e9-46e2-9df7-45f4b27fdf2f","cell_type":"markdown","source":"## Time Series Analysis","metadata":{}},{"id":"07e4ae99-b0a2-47ba-a581-c95fa1d342eb","cell_type":"code","source":"daily_avg = train_file.groupby(['date_id','symbol_id'])['responder_6'].mean().reset_index()\n\nfor symbol_id in train_file['symbol_id'].unique():\n    symbol_data = daily_avg[daily_avg['symbol_id']==symbol_id]\n    plt.figure(figsize=(18,6))\n    plt.plot(symbol_data['date_id'], symbol_data['responder_6'], label = f\"Symbol_{symbol_id}\")\n    plt.title(f\"Time Series for Symbol {symbol_id}\")\n    plt.xlabel(\"date_id\")\n    plt.ylabel(\"responder_6\")\n    plt.legend()\n    plt.show()","metadata":{"tags":[]},"outputs":[],"execution_count":null},{"id":"e1ebf04d-509b-4791-9289-3b901ca0ff15","cell_type":"markdown","source":"The time series plot shows the average value of responder_6 over time for each symbol.\n\nThis helps identify trends, seasonality, or anomalies in the data over time. It’s crucial for understanding the temporal behavior of the target variable.\n\n","metadata":{}},{"id":"d20ddcf5-c99e-40fb-9318-e96bdccf8402","cell_type":"code","source":"from statsmodels.tsa.seasonal import STL\n\nstl = STL(daily_avg['responder_6'], seasonal = 31, period = 7)\nresult = stl.fit()\nfig = result.plot()\nfig.set_size_inches(20,8)\nplt.tight_layout()\nplt.show()","metadata":{"tags":[]},"outputs":[],"execution_count":null},{"id":"47b8916e-6d34-4738-97a9-368838b51563","cell_type":"markdown","source":"The STL (Seasonal and Trend decomposition using Loess) decomposition breaks down the time series into trend, seasonal, and residual components.\n\nThis is useful for understanding the underlying patterns in the time series, such as long-term trends or recurring seasonal effects.","metadata":{}},{"id":"5643e809-7fa4-4fa2-9b55-225e9c008f60","cell_type":"code","source":"from scipy.fft import fft, fftfreq\n\n# Compute Fourier transform\nsymbol_zero = daily_avg[daily_avg['symbol_id']==0]\nyf = fft(symbol_zero[\"responder_6\"].values)\nxf = fftfreq(len(symbol_zero), 1)  # Assume unit frequency\n\n# Plot frequency spectrum\nplt.figure(figsize = (20,8))\nplt.plot(xf, np.abs(yf))\nplt.title(\"Frequency Spectrum\")\nplt.xticks(np.arange(min(xf), max(xf), step = xf.std()/5))\nplt.xlabel(\"Frequency\")\nplt.ylabel(\"Amplitude\")\nplt.show()","metadata":{"tags":[]},"outputs":[],"execution_count":null},{"id":"ad8978ff-fe1f-403b-96b7-25cca3cd53bc","cell_type":"markdown","source":"The frequency spectrum plot shows the amplitude of different frequency components in the time series, obtained using the Fourier Transform.\n\nThis helps identify dominant frequencies in the data, which could correspond to periodic patterns or cycles.","metadata":{}},{"id":"7ea64b18-0cdc-4d6d-a18b-8b158a444281","cell_type":"code","source":"","metadata":{},"outputs":[],"execution_count":null}]}