{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":96164,"databundleVersionId":11418275,"sourceType":"competition"}],"dockerImageVersionId":31040,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Key Insights\n* These are the features - bid/ask quantities, buy/sell quantities, volume and 890 anonimized features\n* Due to high number of columns in the dataset, this EDA is limited to only everyt 5th of the columns out of the 890 features\n* Sample size - 10% of the total rows\n* There is no clear trend over time for bid/ask qty or buy/sell qty.\n* There are some spikes at some intervals, otherwise a steady trend like how we see in any normal crypto bid/ask/buy/sell qty and volume\n* buy and sell_qty appear to be in top correlated features\n* There are many features that are showing strong postive and negative correlations.\n* Based on the boxplots, it seems like most x-features show compressed ranges near zero\n* Market metrics (bid_qty, ask_qty, volume) have much wider ranges (up to 30,000)\n* Market microstructure features (buy_qty, sell_qty, volume) show extreme outliers (dots beyond whiskers) -This suggests frequent \"shock\" events in market activity","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-05-30T02:20:08.948396Z","iopub.execute_input":"2025-05-30T02:20:08.948670Z","iopub.status.idle":"2025-05-30T02:20:09.283132Z","shell.execute_reply.started":"2025-05-30T02:20:08.948646Z","shell.execute_reply":"2025-05-30T02:20:09.282280Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Loading the libraries","metadata":{}},{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings('ignore')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-30T02:20:09.284610Z","iopub.execute_input":"2025-05-30T02:20:09.285044Z","iopub.status.idle":"2025-05-30T02:20:09.290259Z","shell.execute_reply.started":"2025-05-30T02:20:09.285021Z","shell.execute_reply":"2025-05-30T02:20:09.289354Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nsns.set_style('darkgrid')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-30T02:20:09.291214Z","iopub.execute_input":"2025-05-30T02:20:09.291494Z","iopub.status.idle":"2025-05-30T02:20:10.251102Z","shell.execute_reply.started":"2025-05-30T02:20:09.291469Z","shell.execute_reply":"2025-05-30T02:20:10.250308Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Import train dataset","metadata":{}},{"cell_type":"code","source":"# Read the parquet file\ntrain_df = pd.read_parquet('/kaggle/input/drw-crypto-market-prediction/train.parquet')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-30T02:20:10.252711Z","iopub.execute_input":"2025-05-30T02:20:10.253137Z","iopub.status.idle":"2025-05-30T02:20:32.989208Z","shell.execute_reply.started":"2025-05-30T02:20:10.253115Z","shell.execute_reply":"2025-05-30T02:20:32.987990Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Display basic info about the dataframe\nprint(train_df.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-30T02:20:32.990248Z","iopub.execute_input":"2025-05-30T02:20:32.990527Z","iopub.status.idle":"2025-05-30T02:20:33.012571Z","shell.execute_reply.started":"2025-05-30T02:20:32.990500Z","shell.execute_reply":"2025-05-30T02:20:33.011361Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Check if there is any null in the data","metadata":{}},{"cell_type":"code","source":"train_df.isnull().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-30T02:20:33.013647Z","iopub.execute_input":"2025-05-30T02:20:33.014067Z","iopub.status.idle":"2025-05-30T02:20:34.636515Z","shell.execute_reply.started":"2025-05-30T02:20:33.014032Z","shell.execute_reply":"2025-05-30T02:20:34.635674Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df.describe().transpose()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-30T02:20:34.637357Z","iopub.execute_input":"2025-05-30T02:20:34.637562Z","iopub.status.idle":"2025-05-30T02:20:52.677219Z","shell.execute_reply.started":"2025-05-30T02:20:34.637546Z","shell.execute_reply":"2025-05-30T02:20:52.676240Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Data Analysis","metadata":{}},{"cell_type":"code","source":"# Set the style for better visualization\nsns.set(style=\"whitegrid\")\nplt.figure(figsize=(20, 16))\n\n# Create subplots for each metric\nfig, axes = plt.subplots(5, 1, figsize=(20, 20))\n\n# 1. Bid Quantity Trend\naxes[0].plot(train_df.index, train_df['bid_qty'], color='blue', linewidth=1)\naxes[0].set_title('Bid Quantity Over Time', fontsize=14)\naxes[0].set_ylabel('Bid Quantity')\naxes[0].grid(True, linestyle='--', alpha=0.7)\n\n# 2. Ask Quantity Trend\naxes[1].plot(train_df.index, train_df['ask_qty'], color='red', linewidth=1)\naxes[1].set_title('Ask Quantity Over Time', fontsize=14)\naxes[1].set_ylabel('Ask Quantity')\naxes[1].grid(True, linestyle='--', alpha=0.7)\n\n# 3. Buy Quantity Trend\naxes[2].plot(train_df.index, train_df['buy_qty'], color='green', linewidth=1)\naxes[2].set_title('Buy Quantity Over Time', fontsize=14)\naxes[2].set_ylabel('Buy Quantity')\naxes[2].grid(True, linestyle='--', alpha=0.7)\n\n# 4. Sell Quantity Trend\naxes[3].plot(train_df.index, train_df['sell_qty'], color='purple', linewidth=1)\naxes[3].set_title('Sell Quantity Over Time', fontsize=14)\naxes[3].set_ylabel('Sell Quantity')\naxes[3].grid(True, linestyle='--', alpha=0.7)\n\n# 5. Volume Trend\naxes[4].plot(train_df.index, train_df['volume'], color='orange', linewidth=1)\naxes[4].set_title('Trading Volume Over Time', fontsize=14)\naxes[4].set_ylabel('Volume')\naxes[4].set_xlabel('Timestamp')\naxes[4].grid(True, linestyle='--', alpha=0.7)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-30T02:35:50.895998Z","iopub.execute_input":"2025-05-30T02:35:50.897611Z","iopub.status.idle":"2025-05-30T02:35:54.286597Z","shell.execute_reply.started":"2025-05-30T02:35:50.897552Z","shell.execute_reply":"2025-05-30T02:35:54.285372Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\n# Calculate full correlation matrix\ncorr_matrix = train_df[['bid_qty', 'ask_qty', 'buy_qty', 'sell_qty', 'volume', 'label']].corr()\n\n# Create a mask for the upper triangle\nmask = np.triu(np.ones_like(corr_matrix, dtype=bool))\n\nplt.figure(figsize=(10, 8))\nsns.heatmap(corr_matrix, mask=mask, annot=True, cmap='coolwarm', center=0,\n            fmt=\".2f\", linewidths=.5)\nplt.title('Lower Triangle Correlation Matrix', fontsize=16)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-30T02:20:52.678184Z","iopub.execute_input":"2025-05-30T02:20:52.678457Z","iopub.status.idle":"2025-05-30T02:20:53.181100Z","shell.execute_reply.started":"2025-05-30T02:20:52.678433Z","shell.execute_reply":"2025-05-30T02:20:53.180318Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from scipy.stats import pearsonr\nfrom math import ceil\n\n# Select features for analysis \nfocus_features = ['bid_qty', 'ask_qty', 'buy_qty', 'sell_qty', 'volume', 'label'] + \\\n                [f'X{i}' for i in range(1, 891, 5)]  \n\n# Downsample to 10% of data for faster computation\nsample_df = train_df[focus_features].iloc[::10].copy()\n\n# Calculate correlation matrix\ncorr_matrix = sample_df.corr()\n\n# Calculate significant correlations with p-values\nsignificant_corrs = []\nalpha = 0.05  # Significance threshold","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-30T02:20:53.182058Z","iopub.execute_input":"2025-05-30T02:20:53.182373Z","iopub.status.idle":"2025-05-30T02:20:58.028034Z","shell.execute_reply.started":"2025-05-30T02:20:53.182346Z","shell.execute_reply":"2025-05-30T02:20:58.027152Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for i, col1 in enumerate(sample_df.columns):\n    for col2 in sample_df.columns[i+1:]:  # Avoid duplicate pairs\n        r, p = pearsonr(sample_df[col1].dropna(), sample_df[col2].dropna())  # Added .dropna() for safety\n        if p < alpha:\n            significant_corrs.append({\n                'Feature1': col1,\n                'Feature2': col2,\n                'Correlation': r,\n                'P-value': p,\n                'Abs_Correlation': abs(r)\n            })\n\n# Convert to DataFrame\ncorr_results = pd.DataFrame(significant_corrs)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-30T02:20:58.030505Z","iopub.execute_input":"2025-05-30T02:20:58.030877Z","iopub.status.idle":"2025-05-30T02:21:44.706731Z","shell.execute_reply.started":"2025-05-30T02:20:58.030849Z","shell.execute_reply":"2025-05-30T02:21:44.705860Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Separate positive and negative correlations\npos_corrs = corr_results[corr_results['Correlation'] > 0].sort_values('Correlation', ascending=False)\nneg_corrs = corr_results[corr_results['Correlation'] < 0].sort_values('Correlation')\n\n# Visualization function\ndef plot_top_correlations(corr_df, title, n_top=20):\n    plt.figure(figsize=(12, 6))\n    top_df = corr_df.head(n_top)\n    sns.barplot(x='Correlation', y='Feature1', hue='Feature2', data=top_df, dodge=False)\n    plt.title(f'Top {n_top} {title} Correlations (p < 0.05)')\n    plt.xlabel(\"Pearson's r\")\n    plt.legend(bbox_to_anchor=(1.05, 1), loc='upper left')\n    plt.tight_layout()\n    plt.show()\n\n# Plot results\nplot_top_correlations(pos_corrs, 'Positive')\nplot_top_correlations(neg_corrs, 'Negative')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-30T02:21:44.707610Z","iopub.execute_input":"2025-05-30T02:21:44.707890Z","iopub.status.idle":"2025-05-30T02:21:46.632363Z","shell.execute_reply.started":"2025-05-30T02:21:44.707864Z","shell.execute_reply":"2025-05-30T02:21:46.631522Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plotting function with feature names\ndef plot_scatter_grid(corr_pairs, title, n_plots=20):\n    n_rows = ceil(n_plots / 4)\n    fig, axes = plt.subplots(n_rows, 4, figsize=(22, 5.5*n_rows))\n    fig.suptitle(f'{title} Correlations', y=1.02, fontsize=18, weight='bold')\n    \n    for idx, (_, row) in enumerate(corr_pairs.head(n_plots).iterrows()):\n        ax = axes[idx//4, idx%4] if n_rows > 1 else axes[idx%4]\n        \n        # Scatter plot with regression line\n        sns.regplot(x=row['Feature1'], y=row['Feature2'], data=sample_df, \n                    ax=ax, scatter_kws={'alpha':0.5}, line_kws={'color':'red'})\n        \n        # Enhanced title with feature names and stats\n        title_text = (f\"{row['Feature1']} vs {row['Feature2']}\\n\"\n                     f\"r = {row['Correlation']:.2f} (p = {row['P-value']:.1e})\")\n        ax.set_title(title_text, fontsize=12, pad=12)\n        \n        # Rotate x-labels if needed\n        ax.set_xlabel(row['Feature1'], fontsize=10)\n        ax.set_ylabel(row['Feature2'], fontsize=10)\n        \n        # Adjust tick parameters for readability\n        ax.tick_params(axis='both', which='major', labelsize=8)\n    \n    # Hide empty subplots\n    for idx in range(len(corr_pairs.head(n_plots)), n_rows*4):\n        if n_rows > 1:\n            axes[idx//4, idx%4].axis('off')\n        else:\n            axes[idx%4].axis('off')\n    \n    plt.tight_layout()\n    plt.show()\n\n# Generate plots with clear feature labels\nprint(f\"Found {len(pos_corrs)} significant positive correlations\")\nif len(pos_corrs) > 0:\n    plot_scatter_grid(pos_corrs, \"Top 20 Positive\")\n\nprint(f\"\\nFound {len(neg_corrs)} significant negative correlations\")\nif len(neg_corrs) > 0:\n    plot_scatter_grid(neg_corrs, \"Top 20 Negative\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-30T02:21:46.633229Z","iopub.execute_input":"2025-05-30T02:21:46.633463Z","iopub.status.idle":"2025-05-30T02:23:15.061362Z","shell.execute_reply.started":"2025-05-30T02:21:46.633445Z","shell.execute_reply":"2025-05-30T02:23:15.060234Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom matplotlib.dates import DateFormatter\nimport numpy as np\n\ndef plot_top_correlations_over_time(df, corr_df, time_col='timestamp', \n                                  n_pairs=20, samples_per_plot=2000,\n                                  fig_width=20, row_height=2.5):\n    # Prepare the top pairs\n    top_pairs = corr_df.head(n_pairs)\n    n_pairs = min(n_pairs, len(top_pairs))  # In case fewer than requested pairs exist\n    \n    # Calculate subplot layout (4 columns)\n    n_cols = 4\n    n_rows = int(np.ceil(n_pairs / n_cols))\n    \n    # Create figure with appropriate size\n    fig, axes = plt.subplots(n_rows, n_cols, figsize=(fig_width, row_height*n_rows),\n                           sharex=True, squeeze=False)\n    fig.suptitle(f'Top {n_pairs} Correlated Features Over Time', y=1.02, \n                fontsize=16, weight='bold')\n    \n    # Sample data for plotting (for performance)\n    plot_df = df.iloc[::max(1, len(df)//samples_per_plot)]\n    \n    # Plot each correlated pair\n    for idx, (_, row) in enumerate(top_pairs.iterrows()):\n        ax = axes[idx//n_cols, idx%n_cols]\n        \n        # Plot both features on same axes\n        sns.lineplot(data=plot_df, x=time_col, y=row['Feature1'], \n                    ax=ax, color='royalblue', label=row['Feature1'], alpha=0.7)\n        sns.lineplot(data=plot_df, x=time_col, y=row['Feature2'], \n                    ax=ax, color='crimson', label=row['Feature2'], alpha=0.7)\n        \n        # Set title with correlation info\n        corr_type = \"Positive\" if row['Correlation'] > 0 else \"Negative\"\n        title = (f\"{row['Feature1']} & {row['Feature2']}\\n\"\n                f\"{corr_type} r = {abs(row['Correlation']):.2f} (p = {row['P-value']:.1e})\")\n        ax.set_title(title, fontsize=10, pad=8)\n        \n        # Formatting\n        ax.xaxis.set_major_formatter(DateFormatter('%Y-%m-%d'))\n        ax.tick_params(axis='x', rotation=45)\n        ax.grid(True, alpha=0.3)\n        ax.legend(fontsize=8, loc='upper left')\n        \n        # Set y-axis scale based on feature ranges\n        y_min = min(plot_df[row['Feature1']].min(), plot_df[row['Feature2']].min())\n        y_max = max(plot_df[row['Feature1']].max(), plot_df[row['Feature2']].max())\n        margin = (y_max - y_min) * 0.1  # 10% margin\n        ax.set_ylim(y_min - margin, y_max + margin)\n    \n    # Hide any empty subplots\n    for idx in range(n_pairs, n_rows*n_cols):\n        axes[idx//n_cols, idx%n_cols].axis('off')\n    \n    plt.tight_layout()\n    plt.show()\n\n# Example usage:\n# For positive correlations\nplot_top_correlations_over_time(train_df, pos_corrs, n_pairs=20)\n\n# For negative correlations\nplot_top_correlations_over_time(train_df, neg_corrs, n_pairs=20)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-30T02:23:15.062613Z","iopub.execute_input":"2025-05-30T02:23:15.062997Z","iopub.status.idle":"2025-05-30T02:23:30.204655Z","shell.execute_reply.started":"2025-05-30T02:23:15.062967Z","shell.execute_reply":"2025-05-30T02:23:30.203890Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_feature_boxplots(df, features=None, n_top=20, figsize=(20, 10)):\n    \"\"\"\n    Plot horizontal boxplots for quick distribution comparison.\n    \"\"\"\n    if features is None:\n        # Select top numeric features by variance\n        numeric_cols = df.select_dtypes(include=[np.number]).columns\n        variances = df[numeric_cols].var().sort_values(ascending=False)\n        features = variances.head(n_top).index.tolist()\n    \n    plt.figure(figsize=figsize)\n    df[features].plot(kind='box', vert=False, patch_artist=True)\n    plt.title('Feature Distributions (Boxplots)')\n    plt.xlabel('Value Range')\n    plt.grid(True, alpha=0.3)\n    plt.tight_layout()\n    plt.show()\n\nplot_feature_boxplots(train_df)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-30T02:23:30.205585Z","iopub.execute_input":"2025-05-30T02:23:30.205901Z","iopub.status.idle":"2025-05-30T02:23:37.444147Z","shell.execute_reply.started":"2025-05-30T02:23:30.205877Z","shell.execute_reply":"2025-05-30T02:23:37.443269Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_order_flow_imbalance(df, window='1H', figsize=(16, 6)):\n    \"\"\"\n    Use index for rolling calculations\n    \"\"\"\n    df['imbalance'] = (df['buy_qty'] - df['sell_qty']) / (df['buy_qty'] + df['sell_qty'])\n    \n    plt.figure(figsize=figsize)\n    df['imbalance'].rolling(window).mean().plot()\n    plt.axhline(0, color='r', linestyle='--')\n    plt.title(f'Order Flow Imbalance ({window} rolling)')\n    plt.show()\n\nplot_order_flow_imbalance(train_df)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-30T02:23:47.186208Z","iopub.execute_input":"2025-05-30T02:23:47.186447Z","iopub.status.idle":"2025-05-30T02:23:49.022054Z","shell.execute_reply.started":"2025-05-30T02:23:47.186428Z","shell.execute_reply":"2025-05-30T02:23:49.021302Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Key Insights:\n* The chart shows cyclical fluctuations between positive and negative values, suggesting recurring shifts in market sentiment.\n* Notable spikes/dips (e.g., mid-2023, early 2024) may correlate with major market events or news.\n* 2023-03 to 2023-07: Prolonged negative imbalance (sustained selling pressure).\n* Late 2023: Sharp swings between extremes (high volatility).2024:\n* More balanced but volatile (quick reversals).","metadata":{}},{"cell_type":"code","source":"def plot_time_of_day_patterns(df, feature='volume', figsize=(12, 6)):\n    \"\"\"\n    Extract hour from index\n    \"\"\"\n    plt.figure(figsize=figsize)\n    df.groupby(df.index.hour)[feature].mean().plot(kind='bar')\n    plt.title(f'Hourly {feature} Pattern')\n    plt.xticks(rotation=0)\n    plt.show()\n\nplot_time_of_day_patterns(train_df)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-30T02:23:53.255519Z","iopub.execute_input":"2025-05-30T02:23:53.255741Z","iopub.status.idle":"2025-05-30T02:23:53.626975Z","shell.execute_reply.started":"2025-05-30T02:23:53.255723Z","shell.execute_reply":"2025-05-30T02:23:53.626158Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Key Insights:\n* Global Crypto Market Behavior:\n* The 24-hour pattern reflects the non-stop nature of crypto markets, but with clear activity peaks and troughs.\n* Unlike traditional markets, crypto trades 24/7, but human activity still creates cycles.\n* Highest Volume: Typically occurs during overlap of Asian, European, and US trading hours (~12:00-20:00 UTC).\n* Lowest Volume: Often during late US to early Asian hours (~0:00-8:00 UTC).\n* Liquidity Variations: Higher volume hours mean tighter spreads and better order execution.\n* Volatility Clusters: Peaks often coincide with news releases or market opens.","metadata":{}},{"cell_type":"markdown","source":"# If you find this noteboook useful, please consider upvoting. Thanks in Advance..!!","metadata":{}}]}