{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":96164,"databundleVersionId":11418275,"sourceType":"competition"}],"dockerImageVersionId":31040,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-05-29T17:52:16.520544Z","iopub.execute_input":"2025-05-29T17:52:16.520866Z","iopub.status.idle":"2025-05-29T17:52:16.864955Z","shell.execute_reply.started":"2025-05-29T17:52:16.520842Z","shell.execute_reply":"2025-05-29T17:52:16.864009Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <p style=\"background-color:#dff5e9; color:black; font-family:Verdana; font-size:100%; text-align:left; border: 3px solid #57c98a; border-radius:15px; padding:20px 20px;\"> ***This notebook performs an in-depth Exploratory Data Analysis (EDA) on the DRW Crypto Market Prediction dataset. We'll examine temporal patterns, feature relationships, and statistical properties to inform feature engineering and model selection.***</p>","metadata":{}},{"cell_type":"markdown","source":"<p style=\"background-color:#e8f4fc; color:black; font-family:Verdana; font-size:2em; text-align:left; border: 3px solid #3498db; border-radius:15px; padding:30px 30px;\">\n1. Setup & Configuration\n</p>","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom statsmodels.tsa.stattools import adfuller\nfrom statsmodels.graphics.tsaplots import plot_acf, plot_pacf\nimport warnings\nwarnings.filterwarnings('ignore')\n\n# Visualization configuration\nplt.style.use('seaborn')\nsns.set_palette(\"husl\")\nplt.rcParams['figure.figsize'] = (12, 6)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-29T17:52:16.866551Z","iopub.execute_input":"2025-05-29T17:52:16.867035Z","iopub.status.idle":"2025-05-29T17:52:18.230894Z","shell.execute_reply.started":"2025-05-29T17:52:16.866998Z","shell.execute_reply":"2025-05-29T17:52:18.230088Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"background-color:#e8f4fc; color:black; font-family:Verdana; font-size:2em; text-align:left; border: 3px solid #3498db; border-radius:15px; padding:30px 30px;\">\n2. Data Loading & Initial Inspection\n</p>\n\n","metadata":{}},{"cell_type":"code","source":"print(\"Loading data...\")\ntrain = pd.read_parquet('/kaggle/input/drw-crypto-market-prediction/train.parquet')\nprint(f\"Data loaded successfully. Shape: {train.shape}\")\ntrain = train[:262943]\n# Initial data overview\nprint(\"\\n=== Initial Data Overview ===\")\ndisplay(train.head())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-29T17:52:18.231677Z","iopub.execute_input":"2025-05-29T17:52:18.232027Z","iopub.status.idle":"2025-05-29T17:52:46.222950Z","shell.execute_reply.started":"2025-05-29T17:52:18.232007Z","shell.execute_reply":"2025-05-29T17:52:46.221732Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"\\nData types:\")\nprint(train.dtypes.value_counts())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-29T17:52:46.225097Z","iopub.execute_input":"2025-05-29T17:52:46.225381Z","iopub.status.idle":"2025-05-29T17:52:46.238956Z","shell.execute_reply.started":"2025-05-29T17:52:46.225358Z","shell.execute_reply":"2025-05-29T17:52:46.237566Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"\\nDescriptive statistics:\")\ntrain.describe().T # Transposed for better readability","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-29T17:55:34.225318Z","iopub.execute_input":"2025-05-29T17:55:34.225679Z","iopub.status.idle":"2025-05-29T17:55:45.461726Z","shell.execute_reply.started":"2025-05-29T17:55:34.225655Z","shell.execute_reply":"2025-05-29T17:55:45.460762Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"background-color:#e8f4fc; color:black; font-family:Verdana; font-size:2em; text-align:left; border: 3px solid #3498db; border-radius:15px; padding:30px 30px;\">\n3. Missing Value Analysis\n</p>\n\n ","metadata":{}},{"cell_type":"code","source":"print(\"\\n=== Missing Value Analysis ===\")\nmissing_data = train.isnull().sum().sort_values(ascending=False)\nmissing_percent = (missing_data / len(train) * 100).round(2)\nmissing_report = pd.concat([missing_data, missing_percent], axis=1)\nmissing_report.columns = ['Missing Count', 'Missing %']\ndisplay(missing_report[missing_report['Missing Count'] > 0])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-29T17:52:57.816190Z","iopub.execute_input":"2025-05-29T17:52:57.816486Z","iopub.status.idle":"2025-05-29T17:52:59.155710Z","shell.execute_reply.started":"2025-05-29T17:52:57.816458Z","shell.execute_reply":"2025-05-29T17:52:59.154481Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"background-color:#e8f4fc; color:black; font-family:Verdana; font-size:2em; text-align:left; border: 3px solid #3498db; border-radius:15px; padding:30px 30px;\">\n4. Target Variable Analysis\n</p>\n","metadata":{}},{"cell_type":"markdown","source":"My target variable, which represents cryptocurrency market price movements, is highly centered around zero with heavy tails, indicating that small changes are frequent while large swings are more rare. The ADF test confirmed that the series is stationary (very low p-value), which greatly simplifies time series modeling as no differencing is required. Descriptive statistics show a mean close to zero and a standard deviation around 1, consistent with this concentrated distribution. The presence of negative values in both the target and the anonymized Xₙ features is normal in financial data. My goal is to predict these values using a regression approach, with Pearson correlation as the evaluation metric, ensuring robust modeling and careful time-based validation.","metadata":{}},{"cell_type":"code","source":"print(\"\\n=== Target Variable Analysis ===\")\nfig, ax = plt.subplots(1, 2, figsize=(16, 6))\n\n# Distribution plot\nsns.histplot(train['label'], bins=100, kde=True, ax=ax[0])\nax[0].set_title('Target Distribution')\n\n# Stationarity test\nadf_test = adfuller(train['label'].dropna())\nprint(f\"\\nADF Test Results:\\n- ADF Statistic: {adf_test[0]:.4f}\\n- p-value: {adf_test[1]:.4f}\")\n\n# Descriptive stats\nprint(\"\\nTarget Statistics:\")\nprint(train['label'].describe().T)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-29T17:52:59.156785Z","iopub.execute_input":"2025-05-29T17:52:59.157198Z","iopub.status.idle":"2025-05-29T17:54:30.556321Z","shell.execute_reply.started":"2025-05-29T17:52:59.157057Z","shell.execute_reply":"2025-05-29T17:54:30.555240Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"background-color:#e8f4fc; color:black; font-family:Verdana; font-size:2em; text-align:left; border: 3px solid #3498db; border-radius:15px; padding:30px 30px;\">\n 5. Temporal Analysis\n</p>\n\n","metadata":{}},{"cell_type":"markdown","source":"In the first plot, the “Target Time Series,” I visualized the evolution of the label minute by minute over the entire training period. What I observe is a highly noisy series with constant fluctuations around zero. There are occasional extreme spikes (both positive and negative), such as a very pronounced one in late August / early September, which confirms the “heavy-tailed” nature I previously noticed in the distribution. The series doesn’t seem to exhibit a clear long-term trend, which is consistent with its stationarity.\n\nNext, to search for daily seasonal patterns, I grouped the average label values by hour of the day. The “Average Target by Hour of Day” chart reveals some interesting trends:\n\nI observe a tendency toward slightly negative average movements during the early hours (approximately between 0h and 8h), especially around 1–2 AM and 5–7 AM.\n\nThen, the market appears to show more pronounced positive average movements in the late morning and early afternoon (around 10–12 AM and 3–5 PM), with a notable peak around 10–11 AM.\n\nIn the evening (after 6 PM), averages turn negative again, with a sharp dip around 6–7 PM, followed by a slight recovery late at night.\n\nHours between 9–11 PM show positive averages once more.\n\nThese observations suggest that the time of day has an influence on the average movement of the crypto market, which is valuable information. This could reflect the activity cycles of major global players (Asia, Europe, the Americas) or periods of high liquidity. I could potentially use the hour of day, or binary indicators based on these time windows, as features in my model to capture these intra-day seasonal effects","metadata":{}},{"cell_type":"code","source":"print(\"\\n=== Temporal Analysis ===\")\ntrain['timestamp'] = pd.to_datetime(train.index)\n\n# Time series visualization\nfig, ax = plt.subplots(2, 1, figsize=(16, 10))\ntrain['label'].plot(ax=ax[0], title='Target Time Series')\n\n# Seasonality decomposition\ntrain['hour'] = train['timestamp'].dt.hour\ntrain['day_of_week'] = train['timestamp'].dt.dayofweek\ntrain['month'] = train['timestamp'].dt.month\n\ntrain.groupby('hour')['label'].mean().plot(kind='bar', ax=ax[1], \n                                          title='Average Target by Hour of Day')\nplt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-29T17:54:30.557703Z","iopub.execute_input":"2025-05-29T17:54:30.558097Z","iopub.status.idle":"2025-05-29T17:54:32.266592Z","shell.execute_reply.started":"2025-05-29T17:54:30.558060Z","shell.execute_reply":"2025-05-29T17:54:32.265390Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"background-color:#e8f4fc; color:black; font-family:Verdana; font-size:2em; text-align:left; border: 3px solid #3498db; border-radius:15px; padding:30px 30px;\">\n6. Feature Analysis\n\n\n</p>\n\n","metadata":{}},{"cell_type":"markdown","source":"<p style=\"background-color:#e8f4fc; color:black; font-family:Verdana; font-size:2em; text-align:left; border: 3px solid #3498db; border-radius:15px; padding:30px 30px;\">\n6.1 Explicit Features\n\n\n</p>\n","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"I analyzed the correlation between the explicit features (bid_qty, ask_qty, buy_qty, sell_qty, volume) and my target variable (label).\nI observed a strong positive correlation between buy volume (buy_qty), sell volume (sell_qty), and total volume (volume), which is expected.\nIn contrast, the standing order quantities (bid_qty, ask_qty) show negligible correlations with both other features and the target.\nThe executed volumes (buy_qty, sell_qty, volume) exhibit only a very weak positive correlation (~0.03) with the target, suggesting a faint linear signal.\nThis analysis confirms that executed volumes are the most relevant explicit features, but nonlinear models will likely be needed to capture more complex relationships.","metadata":{}},{"cell_type":"code","source":"print(\"\\n=== Explicit Feature Analysis ===\")\nexplicit_features = ['bid_qty', 'ask_qty', 'buy_qty', 'sell_qty', 'volume']\n\n# Correlation matrix\nplt.figure(figsize=(10, 8))\ncorr_matrix = train[explicit_features + ['label']].corr()\nsns.heatmap(corr_matrix, annot=True, cmap='coolwarm', center=0, \n            annot_kws={'size': 10}, fmt='.2f')\nplt.title('Feature Correlation Matrix')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-29T17:54:32.267708Z","iopub.execute_input":"2025-05-29T17:54:32.267951Z","iopub.status.idle":"2025-05-29T17:54:32.632013Z","shell.execute_reply.started":"2025-05-29T17:54:32.267933Z","shell.execute_reply":"2025-05-29T17:54:32.631031Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"background-color:#e8f4fc; color:black; font-family:Verdana; font-size:2em; text-align:left; border: 3px solid #3498db; border-radius:15px; padding:30px 30px;\">\n6.2 Anonymous Features\n\n\n</p>\n\n\n","metadata":{}},{"cell_type":"markdown","source":"The analysis of the anonymous features revealed that, despite their large number, their linear correlations with the target variable remain very weak (all below 0.1). The most correlated feature, X598, reaches only ~0.089. The scatter plots for the top features show high dispersion without a clear linear trend, suggesting that the relationships are likely non-linear or involve complex interactions with other features. This weak correlation confirms the low signal-to-noise ratio typical of the crypto market. Therefore, advanced models capable of capturing non-linearity — such as ensemble-based algorithms — will be necessary, along with targeted feature engineering to improve the quality of input signals.","metadata":{}},{"cell_type":"code","source":"print(\"\\n=== Anonymous Feature Analysis ===\")\nX_cols = [col for col in train.columns if col.startswith('X')]\n\n# Top correlated features\ncorr_with_target = train[X_cols].corrwith(train['label']).abs().sort_values(ascending=False)\nprint(\"\\nTop 10 Features by Correlation Magnitude:\")\nprint(corr_with_target.head(10))\n\n# Visualize top correlations\ntop_features = corr_with_target.head(5).index\nfig, axes = plt.subplots(1, 5, figsize=(20, 4))\nfor ax, feature in zip(axes, top_features):\n    sns.scatterplot(x=train[feature], y=train['label'], ax=ax, alpha=0.3)\n    ax.set_title(f'{feature} (corr: {corr_with_target[feature]:.2f})')\nplt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-29T17:54:32.634285Z","iopub.execute_input":"2025-05-29T17:54:32.634680Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n<p style=\"background-color:#e8f4fc; color:black; font-family:Verdana; font-size:2em; text-align:left; border: 3px solid #3498db; border-radius:15px; padding:30px 30px;\">\n7. Advanced Time Series Analysis\n\n\n</p>\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"Autocorrelation and Rolling Statistics Analysis\n\nThe ACF (Autocorrelation Function) and PACF (Partial Autocorrelation Function) plots provide critical insight into how the target time series depends on its own past values. I observe the following:\n\nACF: The autocorrelations are very high at the first lags (close to 1 at lag 1) and then gradually decline, yet remain significant across many lags. This suggests strong persistence or long-range dependencies in the series, even though it is stationary. This behavior is common in price movement series, where the impact of an event may last over time.\n\nPACF: The PACF shows a significant spike at lag 1, followed by quickly diminishing and insignificant correlations. This suggests that most of the autocorrelation is captured by immediate dependencies (lag 1).\n\nThese patterns — slowly decaying ACF and a sharp PACF cutoff — are typical of ARMA (AutoRegressive Moving Average) processes. For financial return series specifically, this often reflects “memory effects” or clustered volatility.\n\nAdditionally, my analysis of rolling statistics (a 60-minute moving average and moving standard deviation) superimposed on the raw target series (label) reveals key volatility insights:\n\nThe moving average (red) stays close to zero, confirming that on average, the market shows no consistent directional bias over 60-minute windows.\n\nThe moving standard deviation (green) displays clear periods of high volatility (visible as spikes) and low volatility. These “volatility clusters” indicate that large movements tend to be followed by other large movements, and small fluctuations by other small ones.\n\nConclusion:\nMy label exhibits strong first-order autocorrelation and volatility clustering. These insights are vital for modeling: I will likely incorporate lag-based features (e.g., the label value from the previous minute) and rolling volatility measures (e.g., moving standard deviation or mean absolute returns) as model inputs. Doing so will help the model better capture the temporal dynamics of the crypto market.","metadata":{}},{"cell_type":"code","source":"print(\"\\n=== Advanced Time Series Analysis ===\")\nfig, axes = plt.subplots(2, 1, figsize=(16, 10))\n\n# ACF/PACF plots\nplot_acf(train['label'].dropna(), lags=50, ax=axes[0], title='ACF')\nplot_pacf(train['label'].dropna(), lags=50, ax=axes[1], title='PACF')\nplt.tight_layout()\n\n# Rolling statistics\nrolling_window = 60  # 60-minute window\ntrain['rolling_mean'] = train['label'].rolling(window=rolling_window).mean()\ntrain['rolling_std'] = train['label'].rolling(window=rolling_window).std()\n\nplt.figure(figsize=(16, 6))\nplt.plot(train['label'], label='Raw Target', alpha=0.5)\nplt.plot(train['rolling_mean'], label=f'{rolling_window}-min MA', color='red')\nplt.plot(train['rolling_std'], label=f'{rolling_window}-min StdDev', color='green')\nplt.title('Rolling Statistics Analysis')\nplt.legend()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-29T17:54:47.724522Z","iopub.execute_input":"2025-05-29T17:54:47.724768Z","iopub.status.idle":"2025-05-29T17:55:02.983638Z","shell.execute_reply.started":"2025-05-29T17:54:47.724748Z","shell.execute_reply":"2025-05-29T17:55:02.982590Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n<p style=\"background-color:#e8f4fc; color:black; font-family:Verdana; font-size:2em; text-align:left; border: 3px solid #3498db; border-radius:15px; padding:30px 30px;\">\n8. Outlier Detection\n\n</p>\n\n\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"To complement my data exploration, I examined the presence of outliers in my target variable (label) and in some of my explicit features using boxplots.\n\nThe first chart, \"Target Outliers,\" visually confirms what my descriptive statistics and time series had already suggested:\n\nThe label shows a massive concentration of points around zero, as indicated by the very narrow body of the box (the area between the 25th and 75th percentiles).\nHowever, it exhibits a considerable number of outliers (individual points extending well beyond the \"whiskers\") in both positive and negative directions. These points represent extreme price movements that I identified as \"heavy tails\" in the distribution and spikes in the time series. Although statistically outlying, these values correspond to real and significant events in financial markets and should probably not be simply discarded. My model will need to be able to handle them.\n\nNext, I looked at boxplots for a few explicit features (bid_qty, ask_qty, buy_qty). What I observe is a similar pattern:\n\nEach feature shows a distribution with a strong concentration of low values, as indicated by the box (IQR) clustered toward the left.\nMost notably, they all exhibit a very large number of positive outliers. This means that there are moments when order quantities (either pending or executed) are exceptionally high. These \"spikes\" in volume or quantity may be linked to important market events or periods of high liquidity.\n\nIn summary: The outlier analysis confirms that my data, both the target and the features, are subject to extreme events. For the label, these are intense price movements. For the features, these are exceptionally high volumes or quantities. It is crucial that my model is not overly sensitive to these extreme values to the point of instability, but can still leverage them because they may contain valuable predictive information for these volatile markets. I will need to ensure that my modeling approach is robust against these outliers.","metadata":{}},{"cell_type":"code","source":"print(\"\\n=== Outlier Analysis ===\")\nfig, axes = plt.subplots(4, 1, figsize=(16, 12))  \n\n# Target outliers\nsns.boxplot(x=train['label'], ax=axes[0])\naxes[0].set_title('Target Outliers')\n\n# Feature outliers\nfor i, feature in enumerate(explicit_features[:3]):\n    sns.boxplot(x=train[feature], ax=axes[i + 1])\n    axes[i + 1].set_title(f'{feature} Outliers')\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-29T17:55:02.984652Z","iopub.execute_input":"2025-05-29T17:55:02.984917Z","iopub.status.idle":"2025-05-29T17:55:03.725725Z","shell.execute_reply.started":"2025-05-29T17:55:02.984896Z","shell.execute_reply":"2025-05-29T17:55:03.724567Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"background-color:#e8f4fc; color:black; font-family:Verdana; font-size:2em; text-align:left; border: 3px solid #3498db; border-radius:15px; padding:30px 30px;\">\n9. Key Insights & Recommendations\n\n</p>\n\n","metadata":{}},{"cell_type":"markdown","source":"**Target (Label) Characteristics:**\n\n* The `label` is a **stationary time series** (ADF p-value << 0.05), which simplifies direct modeling.\n* There are **strong autocorrelations** at lag 1 (ACF/PACF), indicating short-term temporal dependence.\n* The series exhibits a **heavy-tailed distribution** and **volatility clustering**, with many extreme outliers—this calls for robust modeling techniques.\n\n**Feature Engineering Opportunities:**\n\n* Explicit volume features (`buy_qty`, `sell_qty`, `volume`) show **low linear correlation** with the target, suggesting possible nonlinear relationships or interactions.\n* **Clear hourly patterns** are observed, encouraging the creation of time-based features (e.g., hour-of-day indicators).\n* Rolling statistics analysis reveals **changing volatility regimes**, justifying the use of volatility measures (e.g., rolling standard deviation) as features.\n* Despite weak linear correlations, the many anonymous features (`X_n`) should be explored for potential nonlinear interactions.\n\n**Modeling Considerations:**\n\n* Favor **models that can capture nonlinear relationships and complex interactions** (e.g., XGBoost, LightGBM, CatBoost) for tabular and time series data.\n* Use a **time-based validation strategy** (e.g., rolling or walk-forward validation) to simulate future prediction performance.\n* Consider techniques to **make the model robust to outliers**, rather than removing them altogether.\n\n**Next Steps:**\n\n* Perform **temporal feature engineering** (lags of the `label`, rolling means and standard deviations of `bid_qty`, `ask_qty`, `volume`, and the `X_n` features).\n* Conduct **feature selection and transformation** for the 890 anonymous features to improve performance and reduce complexity.\n* Develop and refine a **regression model** (tree-based boosting) optimized for Pearson correlation.\n\n---\n\nThis structured exploratory analysis provides a solid foundation for predictive model development. The insights gained will guide my feature engineering strategy and inform the choice of model for this financial time series forecasting task.\n","metadata":{}}]}