{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":96164,"databundleVersionId":11418275,"sourceType":"competition"}],"dockerImageVersionId":31040,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🚀 DRW Crypto Market Prediction - Comprehensive EDA\n\n### Exploring High-Frequency Crypto Trading Data\n\nThis notebook provides a thorough analysis of DRW's crypto market prediction dataset - a massive 3.7GB collection with ~900 features of high-frequency trading data.\n\n### Analysis Overview:\n- **Data Quality** - Clean and prepare the dataset\n- **Target Analysis** - Understand distribution and patterns\n- **Feature Exploration** - Identify most predictive signals  \n- **Temporal Patterns** - Time-based insights\n- **Dimensionality Reduction** - PCA","metadata":{}},{"cell_type":"markdown","source":"## 📚 Setup & Imports","metadata":{}},{"cell_type":"code","source":"# Core libraries\nimport numpy as np\nimport pandas as pd\nimport warnings\nwarnings.filterwarnings('ignore')\n\n# Visualization\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom datetime import datetime\n\n# ML & Statistics\nfrom sklearn.feature_selection import mutual_info_regression\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.decomposition import PCA\nfrom scipy import stats\nfrom statsmodels.graphics.tsaplots import plot_acf, plot_pacf\n\n# Utilities\nimport os\nimport gc\nimport time\n\n# Configure plotting\nplt.style.use('seaborn-whitegrid')\nsns.set_theme(style=\"whitegrid\", palette=\"muted\", color_codes=True)\nsns.set_context(\"notebook\", font_scale=1.2)\nplt.rcParams['figure.figsize'] = (12, 8)\npd.set_option('display.max_columns', 100)\n\nprint(\"✅ Libraries imported\")","metadata":{"trusted":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2025-05-23T20:29:44.652926Z","iopub.execute_input":"2025-05-23T20:29:44.653225Z","iopub.status.idle":"2025-05-23T20:29:47.015392Z","shell.execute_reply.started":"2025-05-23T20:29:44.653201Z","shell.execute_reply":"2025-05-23T20:29:47.014473Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📂 Data Loading and Overview\n\nIn this section, we'll load the dataset and examine its basic characteristics.\n\nThe dataset consists of:\n- A training set with target values (train.parquet)\n- A test set for prediction (test.parquet)\n- A sample submission file\n\nWe'll explore key metrics like dataset size, memory usage, and basic structure.","metadata":{}},{"cell_type":"code","source":"# Load datasets\nsample_submission = pd.read_csv(\"/kaggle/input/drw-crypto-market-prediction/sample_submission.csv\")\ntrain_df = pd.read_parquet(\"/kaggle/input/drw-crypto-market-prediction/train.parquet\")\ntest_df = pd.read_parquet(\"/kaggle/input/drw-crypto-market-prediction/test.parquet\")\n\nprint(f\"📊 Training Data: {train_df.shape}\")\nprint(f\"📊 Test Data: {test_df.shape}\")\nprint(f\"💾 Train Memory: {train_df.memory_usage(deep=True).sum() / 1024**3:.2f} GB\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:29:47.016980Z","iopub.execute_input":"2025-05-23T20:29:47.017404Z","iopub.status.idle":"2025-05-23T20:30:52.093926Z","shell.execute_reply.started":"2025-05-23T20:29:47.017381Z","shell.execute_reply":"2025-05-23T20:30:52.092729Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Dataset shape and structure overview\nprint(\"📊 DATASET OVERVIEW\")\nprint(\"=\"*50)\nprint(f\"🎯 Training Data Shape: {train_df.shape}\")\nprint(f\"🧪 Test Data Shape: {test_df.shape}\")\nprint(f\"📝 Sample Submission Shape: {sample_submission.shape}\")\n\nprint(f\"\\n📏 Dataset Dimensions:\")\nprint(f\"   📈 Total training samples: {train_df.shape[0]:,}\")\nprint(f\"   🔢 Total features: {train_df.shape[1]:,}\")\nprint(f\"   📊 Feature to sample ratio: 1:{train_df.shape[0]//train_df.shape[1]:,}\")\n\n# Check if datasets have datetime index\nprint(f\"\\n🕐 Index Information:\")\nprint(f\"   📅 Train index type: {type(train_df.index)}\")\nprint(f\"   📅 Test index type: {type(test_df.index)}\")\n\n# Quick peek at the data structure\nprint(f\"\\n👀 First few column names:\")\nprint(f\"   🏷️ {list(train_df.columns[:10])}\")\nif len(train_df.columns) > 10:\n    print(f\"   ... and {len(train_df.columns)-10} more columns\")\n\n# Check for target variable\nif 'label' in train_df.columns:\n    print(f\"\\n🎯 Target variable found: 'label'\")\n    print(f\"   📊 Target stats: min={train_df['label'].min():.4f}, max={train_df['label'].max():.4f}\")\n    print(f\"   📊 Target mean: {train_df['label'].mean():.6f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:30:52.095185Z","iopub.execute_input":"2025-05-23T20:30:52.095496Z","iopub.status.idle":"2025-05-23T20:30:52.124063Z","shell.execute_reply.started":"2025-05-23T20:30:52.095469Z","shell.execute_reply":"2025-05-23T20:30:52.122735Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔍 Data Quality Assessment\n\nThis section examines data quality issues that could affect model performance:\n\n<details>\n<summary><b>Click to show/hide implementation</b></summary>\n\nWe perform several key checks:\n1. **Missing Values** - Identify and quantify any null values\n2. **Infinite Values** - Detect and handle any infinite values that could cause computation errors\n3. **Constant Features** - Identify features with zero variance (no predictive power)\n4. **Data Shape and Structure** - Final validation after cleaning steps\n\nAddressing these issues early prevents downstream modeling problems and ensures reliable analysis.\n</details>","metadata":{}},{"cell_type":"code","source":"# Check for missing values\nmissing_counts = train_df.isnull().sum()\ntotal_missing = missing_counts.sum()\n\nif total_missing > 0:\n    missing_features = missing_counts[missing_counts > 0].sort_values(ascending=False)\n    print(f\"❌ Missing values found: {len(missing_features)} features\")\n    display(missing_features.head(10))\nelse:\n    print(\"✅ No missing values found\")\n\n# Check test set\ntest_missing = test_df.isnull().sum().sum()\nprint(f\"Test set missing values: {test_missing}\")\n\ndel missing_counts\nif 'missing_features' in locals():\n    del missing_features\ngc.collect()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:30:52.126320Z","iopub.execute_input":"2025-05-23T20:30:52.126603Z","iopub.status.idle":"2025-05-23T20:30:58.746026Z","shell.execute_reply.started":"2025-05-23T20:30:52.126581Z","shell.execute_reply":"2025-05-23T20:30:58.744377Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Check for infinite values\ninf_counts = np.isinf(train_df.select_dtypes(include=[np.number])).sum()\ninf_features = inf_counts[inf_counts > 0]\n\nif len(inf_features) > 0:\n    print(f\"❌ Found {len(inf_features)} features with infinite values\")\n    # Remove features with infinite values\n    train_df = train_df.drop(inf_features.index.tolist(), axis=1)\n    test_features_to_drop = [col for col in inf_features.index if col in test_df.columns]\n    if test_features_to_drop:\n        test_df = test_df.drop(test_features_to_drop, axis=1)\n    print(f\"✅ Removed infinite value features. New shape: {train_df.shape}\")\nelse:\n    print(\"✅ No infinite values found\")\n\ndel inf_counts\nif 'inf_features' in locals():\n    del inf_features\ngc.collect()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:30:58.747903Z","iopub.execute_input":"2025-05-23T20:30:58.748414Z","iopub.status.idle":"2025-05-23T20:31:13.288968Z","shell.execute_reply.started":"2025-05-23T20:30:58.748374Z","shell.execute_reply":"2025-05-23T20:31:13.287747Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Find and remove constant features (zero variance)\nconstant_features = []\nfor col in train_df.select_dtypes(include=[np.number]).columns:\n    if train_df[col].var() == 0:\n        constant_features.append(col)\n\nif constant_features:\n    print(f\"❌ Found {len(constant_features)} constant features\")\n    train_df = train_df.drop(constant_features, axis=1)\n    test_constant_features = [col for col in constant_features if col in test_df.columns]\n    if test_constant_features:\n        test_df = test_df.drop(test_constant_features, axis=1)\n    print(f\"✅ Removed constant features. New shape: {train_df.shape}\")\nelse:\n    print(\"✅ No constant features found\")\n\ndel constant_features\ngc.collect()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:31:13.290089Z","iopub.execute_input":"2025-05-23T20:31:13.290431Z","iopub.status.idle":"2025-05-23T20:31:20.895133Z","shell.execute_reply.started":"2025-05-23T20:31:13.290402Z","shell.execute_reply":"2025-05-23T20:31:20.894160Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Final data summary\nprint(f\"📊 Final Dataset Summary:\")\nprint(f\"   Training samples: {train_df.shape[0]:,}\")\nprint(f\"   Features: {train_df.shape[1]:,}\")\nprint(f\"   Memory usage: {train_df.memory_usage(deep=True).sum() / 1024**3:.2f} GB\")\n\nif 'label' in train_df.columns:\n    feature_cols = [col for col in train_df.columns if col != 'label']\n    print(f\"   Predictive features: {len(feature_cols)}\")\n    print(f\"   Target: label\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:31:20.897064Z","iopub.execute_input":"2025-05-23T20:31:20.898024Z","iopub.status.idle":"2025-05-23T20:31:20.933600Z","shell.execute_reply.started":"2025-05-23T20:31:20.897984Z","shell.execute_reply":"2025-05-23T20:31:20.932100Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Let's take a quick peek at our cleaned data\nprint(\"👀 QUICK PEEK AT CLEANED DATA\")\nprint(\"=\"*40)\n\n# Display basic info\nprint(f\"📊 Dataset shape: {train_df.shape}\")\nprint(f\"📊 Index range: {train_df.index.min()} to {train_df.index.max()}\")\n\n# Show first few rows\nprint(\"\\n🔝 First 3 rows of cleaned data:\")\ndisplay(train_df.head(3))\n\n# Show column names grouped by type\nif 'label' in train_df.columns:\n    feature_cols = [col for col in train_df.columns if col != 'label']\n    print(f\"\\n🏷️ Features ({len(feature_cols)}): {feature_cols[:10]} {'...' if len(feature_cols) > 10 else ''}\")\n    print(f\"🎯 Target: label\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:31:20.934793Z","iopub.execute_input":"2025-05-23T20:31:20.935117Z","iopub.status.idle":"2025-05-23T20:31:21.028858Z","shell.execute_reply.started":"2025-05-23T20:31:20.935091Z","shell.execute_reply":"2025-05-23T20:31:21.027753Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🎯 Target Variable Analysis\n\nUnderstanding the target variable is critical for model development and evaluation.\n\nThis section thoroughly examines the target variable through:\n\n1. **Statistical Analysis** - Descriptive statistics to understand central tendency and dispersion\n2. **Distribution Analysis** - Visualization of shape, skewness and kurtosis\n3. **Outlier Detection** - Multiple methods (IQR, Z-score) to identify extreme values\n4. **Temporal Patterns** - Time series analysis of target behavior\n\nInsights from this analysis guide modeling choices, including potential transformations and handling of extreme values.","metadata":{}},{"cell_type":"code","source":"# Target variable statistical analysis\ntarget = train_df['label']\n\nprint(f\"📊 Target Statistics:\")\nprint(f\"   Count: {target.count():,}\")\nprint(f\"   Mean: {target.mean():.8f}\")\nprint(f\"   Std: {target.std():.8f}\")\nprint(f\"   Min: {target.min():.8f}\")\nprint(f\"   Max: {target.max():.8f}\")\nprint(f\"   Skewness: {target.skew():.4f}\")\nprint(f\"   Kurtosis: {target.kurtosis():.4f}\")\n\n# Check for unique values\nunique_count = target.nunique()\nprint(f\"   Unique values: {unique_count:,}\")\nprint(f\"   Unique ratio: {unique_count / len(target):.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:31:21.030093Z","iopub.execute_input":"2025-05-23T20:31:21.030724Z","iopub.status.idle":"2025-05-23T20:31:21.181051Z","shell.execute_reply.started":"2025-05-23T20:31:21.030655Z","shell.execute_reply":"2025-05-23T20:31:21.180165Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ====================================================================\n# 📈 TARGET VARIABLE VISUALIZATIONS\n# ====================================================================\n\n# Comprehensive target variable visualization\nif 'label' in train_df.columns:\n    target = train_df['label']\n    \n    # Create a figure with 2x2 subplots\n    fig, axes = plt.subplots(2, 2, figsize=(14, 10))\n    \n    # 1. Distribution histogram\n    sns.histplot(target, bins=50, color='lightblue', ax=axes[0, 0], kde=True)\n    axes[0, 0].axvline(target.mean(), color='red', linestyle='--', label=f'Mean: {target.mean():.4f}')\n    axes[0, 0].axvline(target.median(), color='green', linestyle='--', label=f'Median: {target.median():.4f}')\n    axes[0, 0].set_title('Target Distribution')\n    axes[0, 0].legend()\n    \n    # 2. Box plot\n    sns.boxplot(y=target, color='lightcoral', ax=axes[0, 1])\n    axes[0, 1].set_title('Target Boxplot')\n    \n    # 3. Q-Q plot\n    from scipy import stats\n    stats.probplot(target.sample(min(5000, len(target))), plot=axes[1, 0])\n    axes[1, 0].set_title('Q-Q Plot vs Normal')\n    \n    # 4. Time series sample\n    sample_size = min(10000, len(target))\n    sample_target = target.iloc[-sample_size:]\n    axes[1, 1].plot(sample_target.index[-sample_size:], sample_target.values, color='purple', linewidth=1)\n    axes[1, 1].set_title('Time Series (Recent Sample)')\n    axes[1, 1].tick_params(axis='x', rotation=45)\n    \n    plt.tight_layout()\n    plt.show()\n    \n    # Additional statistical plots\n    fig, axes = plt.subplots(1, 3, figsize=(15, 5))\n    \n    # 1. Distribution with statistical annotations\n    sns.histplot(target, bins=50, kde=True, color='skyblue', ax=axes[0])\n    axes[0].axvline(target.mean(), color='red', linestyle='--', label=f'Mean: {target.mean():.6f}')\n    axes[0].axvline(target.median(), color='green', linestyle='--', label=f'Median: {target.median():.6f}')\n    axes[0].set_title('Target Distribution with Central Tendencies')\n    axes[0].legend()\n    \n    # 2. Cumulative distribution\n    sorted_values = np.sort(target)\n    cum_prob = np.arange(1, len(sorted_values) + 1) / len(sorted_values)\n    axes[1].plot(sorted_values, cum_prob, color='orange', linewidth=2)\n    axes[1].set_title('Cumulative Distribution Function')\n    axes[1].set_xlabel('Target Value')\n    axes[1].set_ylabel('Cumulative Probability')\n    \n    # 3. Rolling statistics (if enough data)\n    if len(target) > 1000:\n        window = len(target) // 100  # Dynamic window size\n        rolling_mean = target.rolling(window=window).mean()\n        rolling_std = target.rolling(window=window).std()\n        \n        axes[2].plot(rolling_mean.iloc[window:].index[::window], \n                     rolling_mean.iloc[window:].values[::window], \n                     label='Rolling Mean', color='blue')\n        \n        axes[2].fill_between(rolling_mean.iloc[window:].index[::window],\n                           (rolling_mean - rolling_std).iloc[window:].values[::window],\n                           (rolling_mean + rolling_std).iloc[window:].values[::window],\n                           alpha=0.3, label='±1 Std Dev')\n        axes[2].set_title('Rolling Statistics')\n        axes[2].legend()\n    else:\n        axes[2].text(0.5, 0.5, 'Insufficient data for rolling statistics',\n                   ha='center', va='center')\n        axes[2].set_title('Rolling Statistics (N/A)')\n    \n    plt.tight_layout()\n    plt.show()\nelse:\n    print(\"Target variable 'label' not found!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:31:21.183759Z","iopub.execute_input":"2025-05-23T20:31:21.184034Z","iopub.status.idle":"2025-05-23T20:31:29.016277Z","shell.execute_reply.started":"2025-05-23T20:31:21.184014Z","shell.execute_reply":"2025-05-23T20:31:29.014783Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ====================================================================\n# 🔍 OUTLIER DETECTION AND ANALYSIS\n# ====================================================================\n\nprint(\"\\n🔍 OUTLIER DETECTION ANALYSIS\")\nprint(\"=\"*60)\n\ntarget = train_df['label']\n\n# IQR Method\nQ1 = target.quantile(0.25)\nQ3 = target.quantile(0.75)\nIQR = Q3 - Q1\nlower_bound = Q1 - 1.5 * IQR\nupper_bound = Q3 + 1.5 * IQR\n\niqr_outliers = target[(target < lower_bound) | (target > upper_bound)]\n\nprint(f\"📊 IQR METHOD:\")\nprint(f\"   📉 Q1 (25th percentile): {Q1:.8f}\")\nprint(f\"   📉 Q3 (75th percentile): {Q3:.8f}\")\nprint(f\"   📉 IQR: {IQR:.8f}\")\nprint(f\"   📉 Lower bound: {lower_bound:.8f}\")\nprint(f\"   📉 Upper bound: {upper_bound:.8f}\")\nprint(f\"   🚨 Outliers found: {len(iqr_outliers):,} ({len(iqr_outliers)/len(target)*100:.2f}%)\")\n\n# Z-Score Method (using 3 standard deviations)\nz_scores = np.abs(stats.zscore(target))\nz_outliers = target[z_scores > 3]\n\nprint(f\"\\n📊 Z-SCORE METHOD (|z| > 3):\")\nprint(f\"   🚨 Outliers found: {len(z_outliers):,} ({len(z_outliers)/len(target)*100:.2f}%)\")\n\n# Modified Z-Score Method (more robust)\nmedian = target.median()\nmad = np.median(np.abs(target - median))  # Median Absolute Deviation\nmodified_z_scores = 0.6745 * (target - median) / mad\nmodified_z_outliers = target[np.abs(modified_z_scores) > 3.5]\n\nprint(f\"\\n📊 MODIFIED Z-SCORE METHOD (|mz| > 3.5):\")\nprint(f\"   🚨 Outliers found: {len(modified_z_outliers):,} ({len(modified_z_outliers)/len(target)*100:.2f}%)\")\n\n# Show extreme values\nprint(f\"\\n🔍 EXTREME VALUES:\")\nprint(f\"   🔴 Top 5 highest: {target.nlargest(5).values}\")\nprint(f\"   🔵 Top 5 lowest: {target.nsmallest(5).values}\")\n\n# Outlier impact assessment\ntarget_no_iqr_outliers = target[(target >= lower_bound) & (target <= upper_bound)]\n\nprint(f\"\\n📈 IMPACT OF OUTLIERS (IQR method):\")\nprint(f\"   📋 Original mean: {target.mean():.8f}\")\nprint(f\"   📋 Mean without outliers: {target_no_iqr_outliers.mean():.8f}\")\nprint(f\"   📋 Original std: {target.std():.8f}\")\nprint(f\"   📋 Std without outliers: {target_no_iqr_outliers.std():.8f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:31:29.017593Z","iopub.execute_input":"2025-05-23T20:31:29.017952Z","iopub.status.idle":"2025-05-23T20:31:29.146282Z","shell.execute_reply.started":"2025-05-23T20:31:29.017928Z","shell.execute_reply":"2025-05-23T20:31:29.145259Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🌌 Feature Analysis\n\nThis section examines the features and their relationships with the target variable.\n\nWith almost 900 features, computational efficiency is essential. We use the most recent ~5% of the data, which balances performance with statistical validity.\n\nThe analysis includes:\n\n1. **Feature Statistics** - Basic metrics (mean, std, skewness)\n2. **Mutual Information** - Nonlinear correlation measure to identify predictive features  \n3. **Correlation Analysis** - Linear relationships between features and with target\n4. **Feature Importance** - Ranking of features for potential selection\n\nThis comprehensive exploration helps identify the most valuable features for modeling.\n","metadata":{}},{"cell_type":"code","source":"# Create sample for intensive analyses (5% of data)\nsample_size = max(10000, len(train_df) // 20)  # At least 10k samples\nsample_df = train_df.iloc[-sample_size:].copy()\n\nprint(f\"📊 Analysis sample: {len(sample_df):,} rows ({len(sample_df)/len(train_df)*100:.1f}% of data)\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:31:29.147227Z","iopub.execute_input":"2025-05-23T20:31:29.147464Z","iopub.status.idle":"2025-05-23T20:31:29.225614Z","shell.execute_reply.started":"2025-05-23T20:31:29.147446Z","shell.execute_reply":"2025-05-23T20:31:29.224041Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Feature Statistics\n\nLet's start with a basic overview of the features:","metadata":{}},{"cell_type":"code","source":"# Feature statistics overview\nfeature_cols = [col for col in sample_df.columns if col != 'label'] if 'label' in sample_df.columns else list(sample_df.columns)\n\nprint(f\"📊 Feature Analysis Summary:\")\nprint(f\"   Total features: {len(feature_cols)}\")\n\n# Basic feature statistics\nfeature_stats = {}\nfor col in feature_cols:\n    feature_stats[col] = {\n        'mean': sample_df[col].mean(),\n        'std': sample_df[col].std(),\n        'skewness': sample_df[col].skew(),\n        'zeros_pct': (sample_df[col] == 0).sum() / len(sample_df) * 100\n    }\n\nfeature_stats_df = pd.DataFrame(feature_stats).T\n\nprint(f\"   High skew features (|skew| > 2): {(abs(feature_stats_df['skewness']) > 2).sum()}\")\nprint(f\"   Low variance features (std < 0.01): {(feature_stats_df['std'] < 0.01).sum()}\")\nprint(f\"   High zero features (>50% zeros): {(feature_stats_df['zeros_pct'] > 50).sum()}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:31:29.226760Z","iopub.execute_input":"2025-05-23T20:31:29.227172Z","iopub.status.idle":"2025-05-23T20:31:30.062653Z","shell.execute_reply.started":"2025-05-23T20:31:29.227130Z","shell.execute_reply":"2025-05-23T20:31:30.061661Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Mutual Information Analysis\n\n","metadata":{}},{"cell_type":"code","source":"# Mutual Information Analysis\nX_sample = sample_df.drop('label', axis=1)\ny_sample = sample_df['label']\n\n# Handle infinite values\nX_sample = X_sample.replace([np.inf, -np.inf], np.nan).fillna(0)\n\n# Compute mutual information scores\nmi_scores = mutual_info_regression(X_sample, y_sample, random_state=42, n_neighbors=3)\nmi_scores_series = pd.Series(mi_scores, index=X_sample.columns, name='MI_Score').sort_values(ascending=False)\n\nprint(f\"📊 Mutual Information Results:\")\nprint(f\"   Mean MI score: {mi_scores_series.mean():.6f}\")\nprint(f\"   Max MI score: {mi_scores_series.max():.6f}\")\nprint(f\"   Features with MI > 0.01: {(mi_scores_series > 0.01).sum()}\")\nprint(f\"   Features with MI > 0.05: {(mi_scores_series > 0.05).sum()}\")\n\nprint(f\"\\n🏆 Top 10 Most Important Features:\")\nfor i, (feature, score) in enumerate(mi_scores_series.head(10).items(), 1):\n    print(f\"   {i:2d}. {feature}: {score:.6f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:31:30.063872Z","iopub.execute_input":"2025-05-23T20:31:30.064220Z","iopub.status.idle":"2025-05-23T20:34:10.789936Z","shell.execute_reply.started":"2025-05-23T20:31:30.064191Z","shell.execute_reply":"2025-05-23T20:34:10.787010Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ====================================================================\n# 📈 MUTUAL INFORMATION VISUALIZATIONS\n# ====================================================================\n\n# Mutual Information Visualizations\nif 'label' in sample_df.columns and 'mi_scores_series' in locals():\n    # Create comprehensive MI visualizations using seaborn and matplotlib\n    fig, axes = plt.subplots(2, 2, figsize=(14, 10))\n    \n    # 1. MI Scores Distribution\n    sns.histplot(mi_scores_series.values, bins=50, kde=True, color='lightblue', ax=axes[0, 0])\n    axes[0, 0].axvline(mi_scores_series.mean(), color='red', linestyle='--', \n                      label=f'Mean: {mi_scores_series.mean():.4f}')\n    axes[0, 0].set_title('MI Scores Distribution')\n    axes[0, 0].set_xlabel('MI Score')\n    axes[0, 0].legend()\n    \n    # 2. Top 20 features\n    top_20_mi = mi_scores_series.head(20)\n    sns.barplot(y=top_20_mi.index, x=top_20_mi.values, orient='h', color='lightcoral', ax=axes[0, 1])\n    axes[0, 1].set_title('Top 20 Features by MI Score')\n    axes[0, 1].set_xlabel('MI Score')\n    \n    # 3. MI scores vs feature index (sorted)\n    axes[1, 0].plot(range(len(mi_scores_series)), mi_scores_series.values, color='green')\n    axes[1, 0].set_title('MI Score vs Feature Index')\n    axes[1, 0].set_xlabel('Feature Index (sorted by importance)')\n    axes[1, 0].set_ylabel('MI Score')\n    \n    # 4. Cumulative importance\n    cumulative_mi = np.cumsum(mi_scores_series.values)\n    cumulative_mi_pct = cumulative_mi / cumulative_mi[-1] * 100\n    axes[1, 1].plot(range(len(cumulative_mi_pct)), cumulative_mi_pct, color='purple')\n    axes[1, 1].axhline(50, color='red', linestyle='--', alpha=0.7, label='50%')\n    axes[1, 1].axhline(80, color='orange', linestyle='--', alpha=0.7, label='80%')\n    axes[1, 1].axhline(95, color='green', linestyle='--', alpha=0.7, label='95%')\n    axes[1, 1].set_title('Cumulative Feature Importance')\n    axes[1, 1].set_xlabel('Number of Features')\n    axes[1, 1].set_ylabel('Cumulative Importance %')\n    axes[1, 1].legend()\n    \n    plt.tight_layout()\n    plt.show()\n    \n    # Additional visualization focused on top features\n    plt.figure(figsize=(12, 6))\n    \n    # Bar chart of top 30 features\n    top_30_mi = mi_scores_series.head(30)\n    sns.barplot(y=top_30_mi.index, x=top_30_mi.values, orient='h', palette='viridis')\n    plt.title('Top 30 Features by Mutual Information')\n    plt.xlabel('MI Score')\n    plt.tight_layout()\n    plt.show()\n    \n    print(f\"✅ MI analysis: {len(mi_scores_series)} features, top score: {mi_scores_series.iloc[0]:.6f}\")\nelse:\n    print(\"Cannot create MI visualizations - missing required data!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:34:10.791106Z","iopub.execute_input":"2025-05-23T20:34:10.791394Z","iopub.status.idle":"2025-05-23T20:34:12.891907Z","shell.execute_reply.started":"2025-05-23T20:34:10.791373Z","shell.execute_reply":"2025-05-23T20:34:12.890824Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Correlation Analysis","metadata":{}},{"cell_type":"code","source":"# Feature correlation analysis\ncorr_sample_size = min(len(sample_df), 5000)\ncorr_sample = sample_df.iloc[-corr_sample_size:]\n\n# Select top features for correlation analysis\nif 'mi_scores_series' in locals():\n    top_features = mi_scores_series.head(50).index.tolist()\n    analysis_cols = top_features + (['label'] if 'label' in corr_sample.columns else [])\nelse:\n    analysis_cols = corr_sample.columns[:50].tolist()\n\n# Compute correlation matrix\ncorr_data = corr_sample[analysis_cols]\ncorr_matrix = corr_data.corr()\n\n# Analyze correlation patterns\nif 'label' in corr_matrix.columns:\n    target_correlations = corr_matrix['label'].drop('label').abs().sort_values(ascending=False)\n    \n    print(f\"Target correlation - Max: {target_correlations.iloc[0]:.4f}, Strong correlations (>0.1): {(target_correlations > 0.1).sum()}\")\n\n# Find highly correlated feature pairs\nhigh_correlation_threshold = 0.9\nhigh_corr_pairs = []\nfor i in range(len(corr_matrix.columns)):\n    for j in range(i+1, len(corr_matrix.columns)):\n        if abs(corr_matrix.iloc[i, j]) > high_correlation_threshold:\n            high_corr_pairs.append((\n                corr_matrix.columns[i], \n                corr_matrix.columns[j], \n                corr_matrix.iloc[i, j]\n            ))\n\nif high_corr_pairs:\n    print(f\"Found {len(high_corr_pairs)} highly correlated pairs (|corr| > {high_correlation_threshold})\")\nelse:\n    print(\"No highly correlated feature pairs found\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:34:12.893248Z","iopub.execute_input":"2025-05-23T20:34:12.893605Z","iopub.status.idle":"2025-05-23T20:34:12.979810Z","shell.execute_reply.started":"2025-05-23T20:34:12.893577Z","shell.execute_reply":"2025-05-23T20:34:12.978602Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create correlation heatmap for top features\nplt.figure(figsize=(12, 8))\n\n# Select subset for visualization\nif 'mi_scores_series' in locals() and 'label' in corr_matrix.columns:\n    viz_features = mi_scores_series.head(25).index.tolist() + ['label']\n    viz_corr_matrix = corr_matrix.loc[viz_features, viz_features]\n    title_suffix = \"(Top 25 Features + Target)\"\nelse:\n    viz_corr_matrix = corr_matrix.iloc[:30, :30]\n    title_suffix = \"(Top 30 Features)\"\n\n# Create heatmap\nmask = np.triu(np.ones_like(viz_corr_matrix, dtype=bool))\nplt.figure(figsize=(12, 12))\n\nsns.heatmap(\n    viz_corr_matrix,\n    mask=mask,\n    annot=True,\n    annot_kws={\"size\": 8},     # Smaller font size for annotations\n    cmap='RdYlBu_r',\n    center=0,\n    square=True,\n    fmt='.2f',\n    cbar_kws={\"shrink\": 0.8},\n    linewidths=0.5,             # Optional: adds separation lines\n    linecolor='lightgrey'      # Optional: cleaner grid look\n)\n\nplt.title(f'Feature Correlation Matrix {title_suffix}', fontsize=14)\nplt.xticks(rotation=45, ha='right')  # Rotate x-axis labels\nplt.yticks(rotation=0)               # Keep y-axis labels horizontal\nplt.tight_layout()\nplt.show()\n\n# Target correlation analysis\nplt.figure(figsize=(12, 5))\n\nplt.subplot(1, 2, 1)\ntarget_correlations.head(15).plot(kind='barh', color='steelblue', alpha=0.7)\nplt.title('Top 15 Target Correlations')\nplt.xlabel('Absolute Correlation')\nplt.grid(True, alpha=0.3)\n\nplt.subplot(1, 2, 2)\nall_target_corrs = corr_matrix['label'].drop('label')\nplt.hist(all_target_corrs, bins=30, alpha=0.7, color='lightgreen', edgecolor='black')\nplt.axvline(all_target_corrs.mean(), color='red', linestyle='--', \n            label=f'Mean: {all_target_corrs.mean():.4f}')\nplt.title('Target Correlation Distribution')\nplt.xlabel('Correlation with Target')\nplt.ylabel('Frequency')\nplt.legend()\nplt.grid(True, alpha=0.3)\n\nplt.tight_layout()\nplt.show()\n\nprint(f\"✅ Correlation analysis complete - {len(analysis_cols)} features, {len(high_corr_pairs)} high-corr pairs\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:34:12.981048Z","iopub.execute_input":"2025-05-23T20:34:12.981384Z","iopub.status.idle":"2025-05-23T20:34:15.411893Z","shell.execute_reply.started":"2025-05-23T20:34:12.981354Z","shell.execute_reply":"2025-05-23T20:34:15.408605Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The analysis shows that there are 95 highly correlated pairs with ($|\\text{corr}| > 0.9$). \nIndeed the correlation heatmap even shows features or even groups of features which are correlated with 1 or -1.\n","metadata":{}},{"cell_type":"markdown","source":"## ⏰ Temporal Patterns Analysis\n\nThis analysis explores the target variable:\n\n1. **Target Behavior** - Analyzing volatility, trends, etc.\n2. **Autocorrelation** - Identifying time-lagged relationships","metadata":{}},{"cell_type":"code","source":"# Temporal patterns analysis\ntemporal_sample_size = min(50000, len(train_df))\ntemporal_sample = train_df.iloc[-temporal_sample_size:].copy()\n\nprint(f\"Analyzing temporal patterns on {len(temporal_sample):,} recent observations\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:34:15.413474Z","iopub.execute_input":"2025-05-23T20:34:15.414861Z","iopub.status.idle":"2025-05-23T20:34:15.636952Z","shell.execute_reply.started":"2025-05-23T20:34:15.414824Z","shell.execute_reply":"2025-05-23T20:34:15.634316Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ====================================================================\n# 📊 TEMPORAL VISUALIZATIONS\n# ====================================================================\n\n# Temporal Visualizations with Seaborn\ntarget_temporal = temporal_sample['label']\n\n# Create comprehensive temporal visualizations using matplotlib/seaborn\nfig, axes = plt.subplots(2, 2, figsize=(14, 10))\n\n# 1. Target time series (sample for visualization)\nsample_size = min(5000, len(target_temporal))\nsample_step = max(1, len(target_temporal) // sample_size)\nsample_target = target_temporal.iloc[::sample_step]\n\naxes[0, 0].plot(range(len(sample_target)), sample_target.values, color='blue', linewidth=1)\naxes[0, 0].set_title('Target Time Series (Recent Data)')\naxes[0, 0].set_xlabel('Time Index')\naxes[0, 0].set_ylabel('Target Value')\n\n# 2. Rolling statistics\nif len(target_temporal) > 1000:\n    window = len(target_temporal) // 100\n    rolling_mean = target_temporal.rolling(window=window).mean()\n    rolling_std = target_temporal.rolling(window=window).std()\n    \n    # Sample for visualization\n    r_sample_step = max(1, len(rolling_mean) // 1000)\n    x_values = range(window, len(rolling_mean), r_sample_step)\n    \n    axes[0, 1].plot(x_values, rolling_mean.iloc[window::r_sample_step], label='Rolling Mean', color='green')\n    axes[0, 1].fill_between(\n        x_values,\n        rolling_mean.iloc[window::r_sample_step] - rolling_std.iloc[window::r_sample_step],\n        rolling_mean.iloc[window::r_sample_step] + rolling_std.iloc[window::r_sample_step],\n        alpha=0.3, label='±1 Std Dev'\n    )\n    axes[0, 1].set_title('Rolling Mean & Volatility')\n    axes[0, 1].legend()\n\n# 3. Autocorrelation function\nmax_lags = min(30, len(target_temporal) // 10)\nif max_lags > 0:\n    # Using statsmodels plot_acf function\n    plot_pacf(target_temporal, lags=max_lags, ax=axes[1, 0], alpha=0.05, title='Autocorrelation Function')\n    axes[1, 0].set_xlabel('Lag')\n    axes[1, 0].set_ylabel('Correlation')\n    # Store autocorrelation values for later reference\n    #acf_result = plot_pacf(target_temporal, lags=max_lags, alpha=0.05)\n    #autocorrs = acf_result.correlations[1:]  # Skip lag 0 which is always 1.0\n\n# 4. Target distribution over time periods\nn_periods = 5\nperiod_size = len(target_temporal) // n_periods\nperiod_data = []\nperiod_names = []\n\nfor i in range(n_periods):\n    start_idx = i * period_size\n    end_idx = (i + 1) * period_size if i < n_periods - 1 else len(target_temporal)\n    period_data.append(target_temporal.iloc[start_idx:end_idx].values)\n    period_names.append(f'P{i+1}')\n\nsns.boxplot(data=period_data, ax=axes[1, 1], palette='viridis')\naxes[1, 1].set_xticklabels(period_names)\naxes[1, 1].set_title('Target Distribution Over Time Periods')\naxes[1, 1].set_xlabel('Time Period')\naxes[1, 1].set_ylabel('Target Value')\n\nplt.tight_layout()\nplt.show()\n\n# Additional time-based visualizations\nif len(target_temporal) > 1000:\n    fig, axes = plt.subplots(1, 2, figsize=(14, 5))\n    \n    # 1. Volatility pattern\n    vol_window = max(50, len(target_temporal) // 200)\n    volatility = target_temporal.rolling(window=vol_window).std()\n    vol_sample_step = max(1, len(volatility) // 1000)\n    \n    # Skip NaN values at the beginning due to rolling window\n    x_vol = range(vol_window, len(volatility), vol_sample_step)\n    axes[0].plot(x_vol, volatility.iloc[vol_window::vol_sample_step], color='purple')\n    axes[0].set_title('Volatility Pattern')\n    axes[0].set_xlabel('Time Index')\n    axes[0].set_ylabel('Target Volatility')\n    \n    # 2. Cumulative sum (trend analysis)\n    cum_target = np.cumsum(target_temporal)\n    cum_sample_step = max(1, len(cum_target) // 1000)\n    axes[1].plot(range(0, len(cum_target), cum_sample_step), \n                 cum_target.iloc[::cum_sample_step], \n                 color='darkgreen')\n    axes[1].set_title('Cumulative Sum')\n    axes[1].set_xlabel('Time Index')\n    axes[1].set_ylabel('Cumulative Sum')\n    \n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:34:15.639024Z","iopub.execute_input":"2025-05-23T20:34:15.639522Z","iopub.status.idle":"2025-05-23T20:34:18.555240Z","shell.execute_reply.started":"2025-05-23T20:34:15.639495Z","shell.execute_reply":"2025-05-23T20:34:18.553965Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The target demonstrates significant autocorrelation at lag 1 ($\\rho = 0.9822$), implying it represents a long-window forward-looking metric, such as an aggregated or cumulative return over multiple time steps - probably Bitcoin. Such long-term returns overlap across consecutive time steps, which naturally induces strong autocorrelation.","metadata":{}},{"cell_type":"markdown","source":"## 🧬 Dimensionality Reduction with PCA\n\nWith 800+ features, dimensionality reduction is essential for efficient modeling and avoiding curse of dimensionality.\n","metadata":{}},{"cell_type":"code","source":"# Principal Component Analysis\n# Use sample data for computational efficiency\nX_pca = sample_df.drop('label', axis=1)\ny_pca = sample_df['label']\n\n# Clean and standardize data\nX_pca = X_pca.replace([np.inf, -np.inf], np.nan).fillna(0)\nscaler = StandardScaler()\nX_scaled = scaler.fit_transform(X_pca)\n\nprint(f\"PCA input: {X_pca.shape[0]:,} samples, {X_pca.shape[1]:,} features\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:34:18.556386Z","iopub.execute_input":"2025-05-23T20:34:18.556709Z","iopub.status.idle":"2025-05-23T20:34:19.663340Z","shell.execute_reply.started":"2025-05-23T20:34:18.556662Z","shell.execute_reply":"2025-05-23T20:34:19.662302Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Fit PCA directly with variance threshold of 95%\ntarget_variance = 0.95\npca = PCA(n_components=target_variance, random_state=42)\npca.fit(X_scaled)\n\n# Get the number of components selected to achieve 95% variance\noptimal_components = pca.n_components_\ntotal_explained_variance = np.sum(pca.explained_variance_ratio_)\n\n# Store results in a dictionary for visualization purposes\npca_results = {\n    optimal_components: {\n        'pca_model': pca,\n        'total_explained_variance': total_explained_variance,\n        'explained_variance_ratio': pca.explained_variance_ratio_,\n        'transformed_data': pca.transform(X_scaled)\n    }\n}\n\nprint(f\"PCA results: {optimal_components} components for {total_explained_variance:.3f} variance\")\nprint(f\"\\n🎯 OPTIMAL COMPONENTS: {optimal_components} (for {target_variance:.1%} variance explained)\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:34:19.664804Z","iopub.execute_input":"2025-05-23T20:34:19.666000Z","iopub.status.idle":"2025-05-23T20:34:23.958231Z","shell.execute_reply.started":"2025-05-23T20:34:19.665970Z","shell.execute_reply":"2025-05-23T20:34:23.957073Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Additional visualizations with seaborn and matplotlib\nplt.figure(figsize=(15, 10))\n\n# Plot 1: Scree plot (elbow method)\nplt.subplot(2, 3, 1)\nif optimal_components in pca_results:\n    explained_var = pca_results[optimal_components]['explained_variance_ratio']\n    sns.lineplot(x=range(1, min(21, len(explained_var)+1)), y=explained_var[:20], \n                marker='o', color='blue', linewidth=2)\n    plt.title('Scree Plot (First 20 Components)')\n    plt.xlabel('Principal Component')\n    plt.ylabel('Explained Variance Ratio')\n    plt.grid(True, alpha=0.3)\n\n# Plot 2: Cumulative variance\nplt.subplot(2, 3, 2)\ncumulative_var = np.cumsum(pca_results[optimal_components]['explained_variance_ratio'])\nsns.lineplot(x=range(1, len(cumulative_var)+1), y=cumulative_var, \n         marker='o', color='darkgreen', linewidth=2)\nplt.axhline(y=target_variance, color='red', linestyle='--', alpha=0.7, \n           label=f'{target_variance:.1%} target')\nplt.title('Cumulative Explained Variance')\nplt.xlabel('Number of Components')\nplt.ylabel('Cumulative Variance Explained')\nplt.legend()\nplt.grid(True, alpha=0.3)\n\n# Plot 3: Component importance visualization\nplt.subplot(2, 3, 3)\n# With just one PCA result, show bars for different variance thresholds\nvariance_thresholds = [0.7, 0.8, 0.9, 0.95, 0.99]\ncomponents_needed = []\nfor threshold in variance_thresholds:\n    # Find how many components needed for this threshold\n    cumsum = np.cumsum(pca_results[optimal_components]['explained_variance_ratio'])\n    n_needed = np.argmax(cumsum >= threshold) + 1\n    components_needed.append(n_needed)\n\nsns.barplot(x=[f\"{int(t*100)}%\" for t in variance_thresholds], y=components_needed, palette='viridis')\nplt.title('Components for Variance Thresholds')\nplt.xlabel('Variance Threshold')\nplt.ylabel('Components Required')\nplt.grid(True, alpha=0.3)\n\n# Plot 4: Top component loadings with seaborn\nplt.subplot(2, 3, 4)\nif optimal_components in pca_results:\n    pca_model = pca_results[optimal_components]['pca_model']\n    pc1_loadings = pca_model.components_[0]\n    \n    # Get top 15 features by absolute loading\n    top_indices = np.argsort(np.abs(pc1_loadings))[-15:]\n    top_loadings = pc1_loadings[top_indices]\n    top_features = [X_pca.columns[i] for i in top_indices]\n    \n    # Create DataFrame for seaborn\n    loading_df = pd.DataFrame({\n        'Feature': [f\"...{feat[-10:]}\" for feat in top_features],\n        'Loading': top_loadings\n    })\n    \n    # Use seaborn barplot with color mapped to loading value\n    sns.barplot(y='Feature', x='Loading', data=loading_df, \n                palette='RdYlGn', orient='h')\n    plt.title('Top PC1 Feature Loadings')\n    plt.xlabel('Loading Value')\n    plt.grid(True, alpha=0.3)\n\n# Plot 5: 2D PCA projection with seaborn\nplt.subplot(2, 3, 5)\nif optimal_components in pca_results:\n    transformed_data = pca_results[optimal_components]['transformed_data']\n    \n    # Sample for visualization\n    sample_size = min(2000, transformed_data.shape[0])\n    sample_indices = np.random.choice(transformed_data.shape[0], sample_size, replace=False)\n    \n    pc1_sample = transformed_data[sample_indices, 0]\n    pc2_sample = transformed_data[sample_indices, 1] if transformed_data.shape[1] > 1 else np.zeros(len(sample_indices))\n    \n    if y_pca is not None:\n        # Create DataFrame for seaborn\n        vis_df = pd.DataFrame({\n            'PC1': pc1_sample,\n            'PC2': pc2_sample,\n            'Target': y_pca.iloc[sample_indices].values\n        })\n        \n        # Use seaborn scatterplot\n        sns.scatterplot(data=vis_df, x='PC1', y='PC2', hue='Target', \n                       palette='viridis', alpha=0.6, s=10, legend=False)\n        plt.colorbar(plt.cm.ScalarMappable(cmap='viridis'), \n                    ax=plt.gca(), label='Target Value')\n    else:\n        # Simple scatter without target coloring\n        sns.scatterplot(x=pc1_sample, y=pc2_sample, color='blue', alpha=0.6, s=10)\n    \n    plt.title('First Two Principal Components')\n    plt.xlabel(f'PC1 ({pca_model.explained_variance_ratio_[0]:.3f} variance)')\n    plt.ylabel(f'PC2 ({pca_model.explained_variance_ratio_[1]:.3f} variance)' if len(pca_model.explained_variance_ratio_) > 1 else 'PC2')\n    plt.grid(True, alpha=0.3)\n\n# Plot 6: Component importance distribution\nplt.subplot(2, 3, 6)\nif optimal_components in pca_results:\n    # Show distribution of variance explained by components\n    explained_importance = pca_results[optimal_components]['explained_variance_ratio']\n    component_df = pd.DataFrame({\n        'Component': range(1, len(explained_importance)+1),\n        'Variance Explained': explained_importance\n    })\n    \n    # Calculate components needed for different variance levels\n    cumsum = np.cumsum(explained_importance)\n    key_thresholds = [50, 75, 90, 95, 99]\n    components_needed = []\n    \n    for threshold in key_thresholds:\n        n_needed = np.argmax(cumsum >= threshold/100) + 1\n        components_needed.append(n_needed)\n    \n    # Show annotations for key variance thresholds\n    for i, (threshold, n_comp) in enumerate(zip(key_thresholds, components_needed)):\n        plt.axvline(x=n_comp, color=f'C{i}', linestyle='--', alpha=0.7, \n                  label=f'{threshold}% variance: {n_comp} components')\n    \n    # Plot variance distribution\n    sns.histplot(component_df['Variance Explained'], bins=20, kde=True, \n               color='lightgreen', alpha=0.7)\n    plt.title('Component Importance Distribution')\n    plt.xlabel('Variance Explained per Component')\n    plt.ylabel('Count')\n    plt.legend(fontsize='x-small')\n    plt.grid(True, alpha=0.3)\n\nplt.tight_layout()\nplt.show()\n\nprint(f\"\\n✅ PCA analysis completed!\")\nif optimal_components:\n    print(f\"   🎯 Recommended components: {optimal_components}\")\n    print(f\"   📊 Variance explained: {pca_results[optimal_components]['total_explained_variance']:.4f}\")\n    print(f\"   📊 Dimensionality reduction: {X_pca.shape[1]} → {optimal_components} ({(1-optimal_components/X_pca.shape[1])*100:.1f}% reduction)\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-23T20:34:23.960007Z","iopub.execute_input":"2025-05-23T20:34:23.960434Z","iopub.status.idle":"2025-05-23T20:34:25.983539Z","shell.execute_reply.started":"2025-05-23T20:34:23.960397Z","shell.execute_reply":"2025-05-23T20:34:25.982446Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🎯 Summary\n\n### Analysis Completed\n- **Data Quality Assessment** - Verified clean, analysis-ready dataset\n- **Target Variable Analysis** - Characterized distribution and temporal behavior  \n- **Feature Exploration** - Identified predictive features using mutual information\n- **Dimensionality Analysis** - Evaluated PCA compression opportunities\n\n### Key Findings\n- High-quality dataset with no missing values but some columns with inf values and constant values\n- The target probably represents a multi-step forward return\n- Significant feature redundancy - 125 features already explain 95% of variance\n- (Groups of) highly correlated features - dimension reduction techniques should be beneficial\n\n### 🙌 Feedback Welcome\n\nIf you found this analysis useful or have suggestions for improvement, I'd love to hear your thoughts in the comments!","metadata":{}}]}