{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84493,"databundleVersionId":9871156,"sourceType":"competition"}],"dockerImageVersionId":30839,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"The dataset provided by Jane Street presents a unique challenge as all parameters have been anonymized and assigned random names. This makes it difficult to discern the exact nature of the data or fully comprehend the underlying problem. To bridge this gap, I will first provide a comprehensive explanation of the problem using relatable real-world examples.\n\nThe goal is to ensure a clear understanding of the dataset and the task at hand. Once the problem is thoroughly understood, I will proceed to detail my approach for solving it.\n\nThe explanation is quite detailed, so let’s dive into it step by step. Get ready for a long hustle 🦾","metadata":{}},{"cell_type":"markdown","source":"# 📊 Problem Breakdown and Explanation\nThis notebook provides a detailed explanation of the problem in the **Jane Street Real-Time Market Data Forecasting** competition hosted on Kaggle. The goal is to forecast financial market responders using anonymized real-world data. Let's break down the problem step by step.\n\n### Understanding the Jane Street Market Prediction Dataset\n\nThe dataset provided by **Jane Street** for market prediction consists of **numerical values** representing various **market-related attributes**.  \nHere I explain each component of the dataset in a way that even someone with **zero knowledge** of **finance or stock markets** can understand. \n\n\n### The Problems Faced in Real Life (that we try to solve in this notebook)\nIn real-world financial trading, the problem mirrors this competition:\n\n- **Prediction Goal**: Traders want to predict key metrics to make better trading decisions, such as buying or selling financial instruments.\n- **Challenges Faced**:\n  1. **Fat-Tailed Distributions**: Financial data often exhibits fat-tailed distributions, meaning extreme values (outliers) occur more frequently than in a normal distribution. These outliers can significantly impact predictions, making risk estimation and financial modeling more complex.\n  2. **Non-Stationary Data**: A dataset is non-stationary when its statistical properties (such as mean, variance, and correlation structure) change over time. In finance, stock prices, interest rates, and market trends fluctuate due to evolving economic conditions, making traditional time series models less reliable.\n  3. **Market Shifts**: Financial markets are highly dynamic, and external factors—such as economic crises, regulatory changes, technological innovations, or geopolitical events—can cause sudden and unpredictable shifts. These shifts can render historical data less useful for future predictions.\n  4. **High Dimensionality**: Financial models must process a vast number of features (e.g., stock prices, macroeconomic indicators, social sentiment, news events) in real time. Handling such large, complex datasets efficiently requires advanced techniques like dimensionality reduction and feature engineering.\n  5. **Data Anonymization**: Real-world financial datasets often undergo data anonymization, where sensitive information is masked, encrypted, or replaced to maintain confidentiality. While this protects privacy, it can also reduce data interpretability and model accuracy, making it harder to extract meaningful insights.\n\n\n\n\n\n\n\n\n---\n\n## 📌 1. What is This Dataset About?  \n\nImagine you are trying to **predict** whether the **price of a stock** will **go up or down** based on past data.  \n\n- This dataset contains **financial indicators** (features) that describe the **stock market** at a given time.\n- Your goal is to **use these indicators** to predict `responder_6`, which represents **some kind of market reaction**. eg. Increase in price, decrease in price, increase in liquidity(i.e eas of buying an asset), etc. \n\n### 🔍 **Real-Life Analogy**  \nThink of it like predicting if a **cricket player** will **score more than 50 runs** in a match based on past statistics:  \n\n🏏 **Factors to Consider in Cricket Prediction:**  \n- **Pitch condition**  \n- **Weather**  \n- **Player's form**  \n- **Opponent team strength**  \n\n📈 **Similarly, in the stock market, we use:**  \n- **Stock price movements**  \n- **Trading volume**  \n- **Market trends**  \n\n---\n\n## 📌 2. Breaking Down the Dataset  \n\n## **Files in the Dataset**  \n\nThe dataset comprises several files:\n\n#### Main Files:\n- `train.parquet`\n- `test.parquet`\n- `lags.parquet`\n#### Metadata Files:\n- `features.csv`\n- `responders.csv`\n- `sample_submission.csv`\n\n\n\n| File Name | Description |\n|-----------|-------------|\n| `train.parquet` | The training dataset containing features, responders, and additional columns. |\n| `test.parquet` | The test dataset used for making predictions. |\n| `lags.parquet` | Contains lagged versions of features for time-series analysis. |\n| `features.csv` | Provides metadata about each feature, including tags that classify them. |\n| `responders.csv` | Provides metadata about responder columns, including classification tags. |\n| `sample_submission.csv` | Example format for the competition submission. |\n| `kaggle_evaluation` | Contains evaluation metrics and competition scoring details. |\n\n\n\n## **1. `train.parquet` (Main Training Dataset)**  \n\nThis file contains the main dataset used for training models. It consists of:  \n\n- **Features (`feature_00` to `feature_78`)**  \n  - These represent different market indicators (anonymized).  \n  - They could include price movements, trading volume, volatility, etc.  \n  - A total of **79 features** are present.  \n\n- **Responders (`responder_0` to `responder_8`)**  \n  - These are the target values representing market reactions.  \n  - **So we use these Features to predict the Responders. The main task is to predict only the `responder_6` and not the others** .  \n\n- **Other Important Columns:**  \n  \n  | Column | Description |\n  |--------|-------------|\n  | `date_id` | The day when the data was recorded (time-series aspect). |\n  | `time_id` | The specific timestamp when the data was recorded. |\n  | `symbol_id` | Represents different financial instruments (e.g., different stocks/assets). |\n  | `weight` | Represents the importance of each data point in evaluation. |\n\n- **We have 92 columns in total**\n\n\n## **2. `test.parquet` (Test Dataset)**  \n\n- Contains **only features** and additional columns like `date_id`, `time_id`, and `symbol_id`.  \n- Does not contain responder values (`responder_6` is missing since this is what we are predicting).  \n- Predictions made on this data will be submitted in `sample_submission.csv`.  \n\n\n\n## **3. `lags.parquet` (Lagged Features File)**  \n\n- Provides **lagged versions of features**, which are useful for **time-series forecasting**.  \n- Example: If `feature_01` represents the stock price at time `t`, a lagged version might store its value from `t-1`.  \n- Helps capture **temporal dependencies** in the data.  \n\n\nThe `lags.parquet` file contains **lagged values** of responders (`responder_0` to `responder_8`) from the **previous day (`date_id`)**, available only at the **first time step (`time_id = 0`)** of each day.\n\n**---**\n\n#### 📊 Why Use `lags.parquet`?\n1. **Historical Context:** Past responder values help identify patterns over time.\n2. **Feature Engineering:** Enables the creation of features like:\n   - **Lag Differences:** Differences between current and lagged values, e.g., `responder_6 - lag_responder_6`.\n   - **Rolling Statistics:** Moving averages or variances of lagged values.\n   - **Momentum Indicators:** Features that capture trends or reversals in responders.\n3. **Market Dynamics:** Captures temporal dependencies such as autocorrelation or mean reversion.\n\n**---**\n\n#### 🔍 How to Use `lags.parquet`\n1. **Merge with Training Data:** Combine the lagged data with the training dataset using `date_id` as the key.\n2. **Feature Engineering:** Use the lagged values to create new features like differences, ratios, or rolling statistics.\n3. **Incorporate in Model:** Include these lagged features in your model training to capture temporal dependencies effectively.\n\n**---**\n\n#### 🚩 Limitations\n1. Lagged values are only available at **`time_id = 0`** for each day.\n2. Only **1-day lag** is provided. Longer lagged features require additional preprocessing.\n\n**---**\n\n#### 🌟 Conclusion\n`lags.parquet` is highly valuable for capturing temporal dependencies and improving predictions. It allows the model to learn patterns based on past responder values while respecting the sequence of data. Use it thoughtfully to enhance performance. 🚀\n\n\n\n## **4. `features.csv` (Feature Metadata File)**  \n\nThis file contains metadata about each feature, including **classification tags**.  \n\n### **Columns in `features.csv`:**  \n| Column | Description |\n|--------|-------------|\n| `feature` | Name of the feature (e.g., `feature_00`, `feature_01`, ..., `feature_78`). |\n| `tag_0` to `tag_8` | Boolean values (`True`/`False`) indicating feature categories. |\n\n- There are **79 features**.  \n- Each feature has 9 **tags** (binary indicators) that classify them.\n- Each Responder also has 5 different**tags**\n\n### **Tag Interpretation for features:**  \n| Tag | Possible Meaning |\n|-----|-----------------|\n| `tag_0` | Is this feature related to stock price movements? (yes/no) |\n| `tag_1` | Related to trading volume. (yes/no)|\n| `tag_2` | Linked to market volatility. (yes/no) |\n| `tag_3` | Based on historical data (lagging indicator). (yes/no) |\n| `tag_4` | Leading indicator (helps predict future trends). (yes/no) |\n| `tag_5` | Could represent short-term patterns. (yes/no) |\n| `tag_6` | May be linked to economic indicators. (yes/no) |\n| `tag_7` | Could indicate macroeconomic conditions. (yes/no) |\n| `tag_8` | Possibly related to sector-based trends. (yes/no) |\n\n💡 **Example:**  \n- `feature_02` has `tag_2 = True`, so it might be related to market volatility.  \n- `feature_05` has `tag_3 = True`, so it may be a historical lagging indicator.  \n\n\n\n## **5. `responders.csv` (Responder Metadata File)**  \n\nThis file contains metadata about the **9 responder values (`responder_0` to `responder_8`)**, including classification tags.  \n\n### **Columns in `responders.csv`:**  \n| Column | Description |\n|--------|-------------|\n| `responder` | Name of the responder (`responder_0` to `responder_8`). |\n| `tag_0` to `tag_4` | Boolean values (`True`/`False`) classifying responders. |\n\n### **Tag Interpretation for Responders:**  \n| Tag | Possible Meaning |\n|-----|-----------------|\n| `tag_0` | Does this responder represent immediate market reaction. (yes/no)|\n| `tag_1` | Measures long-term impact. (yes/no)|\n| `tag_2` | Could indicate correlation with volatility. (yes/no)|\n| `tag_3` | May track fundamental economic signals. (yes/no)|\n| `tag_4` | Might be linked to institutional trading activity. (yes/no)|\n\n💡 **Example:**  \n- `responder_2` has `tag_0 = True` and `tag_1 = True`, meaning it might measure both short-term and long-term reactions.  \n- `responder_6` has `tag_2 = True`, suggesting it is linked to market volatility.  \n\n\n\n## **6. `sample_submission.csv` (Submission Format File)**  \n\nThis file shows the expected **format for submissions**:  \n| Column | Description |\n|--------|-------------|\n| `row_id` | Unique identifier for each row in the test dataset. |\n| `responder_6` | Predicted value for responder_6 (the main target). |\n\n## **7. `kaggle_evaluation` (Evaluation Metrics File)**  \n\n- Contains information about how predictions will be scored in the competition.  \n- Jane Street likely uses a weighted metric where **some data points (weights) matter more** than others.  \n\n\n\n---\n\n## **What are these features, responders, and tags?**\n\n### 🔹 **A. Features (`feature_00` to `feature_78`)**  \nFeatures are **numerical values** that describe different aspects of the **stock market** at a given time.  \nEven though the dataset doesn’t tell us exactly what each feature represents (they are **anonymized**), they could include:  \n\n✅ **Stock price movements** (e.g., how much a stock’s price changed in the last 10 minutes)  \n✅ **Trading volume** (e.g., how many people are buying/selling that stock)  \n✅ **Market trends** (e.g., is the stock following an upward or downward trend?)  \n\n#### 🔍 **Real-Life Analogy**  \nIf you were predicting **house prices**, the **features** would be:  \n🏠 **Size of the house (square feet)**  \n🛏 **Number of bedrooms**  \n📍 **Location of the house**  \n🚨 **Crime rate in the area**  \n🏫 **Nearby schools and hospitals**  \n\nSimilarly, in this dataset, **features** are different **attributes that describe the stock market**.  \n\n**---**\n\n\n### 🔹 **B. Responders (`responder_0` to `responder_8`)**  \nThe responders are the **values we are trying to predict**.  \nThey represent **how the market reacts** based on the given features.  \n\nIn financial markets, \"responders\" could represent various metrics or signals derived from market activities, such as:\n\n- **Price Movements**: Changes in the price of financial instruments.\n- **Liquidity Indicators**: Measures of market depth or ease of transaction.\n- **Volatility Metrics**: Fluctuations in asset prices over time.\n- **Order Book Signals**: Information derived from buy/sell orders in the market.\n- **Market Sentiment Scores**: Indicators derived from sentiment analysis on financial news or social media.\n  \n#### 🔍 **Real-Life Analogy**  \nContinuing with the **house price example**, a **responder** would be:  \n- 💵 **The actual sale price of the house**  \n- 🏠 **Whether the house will sell within 30 days**  \n\nSimilarly, `responder_6` is a **financial outcome** that traders might use to **make decisions**.  \n\n\n\n\n#### ✅ **You need to predict `responder_6`, but what does it mean?**  \nSince it is **anonymized**, it could represent:  \n- 📈 **Future stock price movements** (whether a stock will go up or down in the next time step)  \n- 📊 **Market sentiment** (how traders are reacting to the stock)  \n- 💰 **Profitability of a trading strategy**\n\n\n\n#### ✅ **But if we only need to predict `responder_6`, why are all other responders given in the dataset?**\n\nThe other responders (`responder_1` to `responder_5`) may serve several purposes in improving the prediction of `responder_6`:\n\n\n   - Other responders may be **strongly correlated** with `responder_6`, acting as useful predictors.  \n     *Example:* In customer surveys, earlier responses might indicate patterns that help predict the final response.\n\n \n   - Training on multiple responders may help the model learn **general patterns**, improving accuracy.  \n     *Example:* In finance, multiple stock movements can provide insight into a specific stock’s future price.\n\n\n   - New features can be created from other responders, such as:  \n     Averaging past responders: `mean(responder_1, ..., responder_5)`  \n     Finding trends: Standard deviation, differences, or moving averages.\n\n\n   - All responders may be included to ensure **data integrity and provide context**, even if some are not directly used.\n\n  \n   - Some responders **may not** contribute to predicting `responder_6`.  \n     Techniques like **correlation analysis, PCA, or SHAP values** can help remove unnecessary features.\n\n\n**---**\n\n\n\n### 🔹 **C. Tags (`tag_0` to `tag_8` for features, `tag_0` to `tag_4` for responders)**  \nTags help **categorize** features and responders.  \nThey provide **metadata** that helps us understand **which features belong to which category**.  \n\n| Tag Name | Possible Meaning |\n|---------|------------------|\n| 📉 `tag_0` | Feature is related to **stock price movements**. |\n| 📊 `tag_1` | Feature is related to **trading volume**. |\n| 🔀 `tag_2` | Feature is linked to **market volatility**. |\n| ⏳ `tag_3` | Feature is based on **historical data** (**lagging indicator**). |\n| 🔮 `tag_4` | Feature is a **leading indicator** (helps predict **future trends**). |\n\n#### 🔍 **Real-Life Analogy**  \nImagine you're predicting **how well a student** will perform in an exam. Different tags can indicate:  \n\n✅ `tag_0` → **How many hours the student studied**  \n✅ `tag_1` → **The difficulty level of the exam**  \n✅ `tag_2` → **How well the student performed in past exams**  \n\nSimilarly, **tags help classify financial indicators** based on their **importance and role**.  \n\n\n**---**\n\n\n### 🔹 **D. Additional Columns**  \nThese columns provide **context** about the data.  \n\n| Column Name | Meaning |\n|------------|---------|\n| 📅 `date_id` | The **day** when the data was recorded. Helps maintain **time order**. |\n| ⏰ `time_id` | The **exact time** when the data was recorded. Useful for **time-series forecasting**. |\n| 📌 `symbol_id` | Represents different **financial instruments** (e.g., different stocks, assets, or trading instruments). |\n| ⚖ `weight` | A special value that **indicates the importance** of each data point. Some predictions **matter more than others** for evaluation. |\n\n#### 🔍 **Real-Life Analogy**  \n\n| Dataset Column | Real-Life Equivalent |\n|---------------|----------------------|\n| 📅 `date_id` | The **day** you checked the weather. |\n| ⏰ `time_id` | The **exact time** you checked the weather. |\n| 📌 `symbol_id` | Different **cricket players** (you analyze each player's stats separately). |\n| ⚖ `weight` | How **important** a cricket match is (e.g., **World Cup final vs. a local league match**). |\n\n\n\n\n\n**---**\n\n### 🔹 **E. Summary**\n\n| Concept | Explanation |\n|---------|------------|\n| 📊 **Features (`feature_00` to `feature_78`)** | Various **market indicators** (e.g., **price movements, volume, trends**). |\n| 🎯 **Responders (`responder_0` to `responder_8`)** | **Market reaction values** (**we predict `responder_6`**). |\n| 📅 **`date_id`, `time_id`** | When the data was recorded. Important for **time-series analysis**. |\n| 📌 **`symbol_id`** | Represents different **stocks or financial instruments**. |\n| ⚖ **Weight (`weight`)** | **Importance** of each data point in the final evaluation. |\n| 🏷 **Tags (`tag_0` to `tag_8`)** | Classify features into **categories like price, volume, volatility**. |\n\n\n\n\n---\n\n## 📌 3. What is Our Goal?  \n\n### ✅ **Objective:**  \nBuild a **machine learning model** that takes **features (market indicators)** as input and predicts **`responder_6`** as output.  \n\n### 🔴 **Challenges:**  \n🚨 The data is **anonymized**, so we **don’t know exactly** what each feature represents.  \n📉 The **stock market is highly unpredictable** (affected by news, economic events, trader psychology).  \n⚖ Some features might be **more important than others**, and we need to **find them**.  \n\n### 🛠 **Plan to Solve the Problem:**  \n1️⃣ **Explore the data** – Check for **missing values, distributions, and correlations**.  \n2️⃣ **Feature engineering** – Create **new meaningful features** using existing ones.  \n3️⃣ **Choose a model** – Try different algorithms like **XGBoost, LSTMs, or Transformers**.  \n4️⃣ **Validate the model** – Use techniques like **cross-validation** to ensure **reliability**.  \n5️⃣ **Submit predictions** – Use **Jane Street’s evaluation system** to check our model’s performance.  \n\n---\n\n## 📌 4. What Does the Predicted **responder_6** Look Like?\n\n### ✅ **Understanding the Output:**\n\nThe output of our model will be the predicted values for **`responder_6`**, a numerical variable that reflects a specific market response. Here's a detailed explanation of how this output will look and how to interpret it.\n\n**---**\n\n### 🔍 **Characteristics of Predicted `responder_6`:**\n1. **Numerical Range:**\n   - The values of **`responder_6`** are clipped between **-5** and **5**, based on the dataset constraints.\n   - Example: If the model predicts `3.2`, it indicates a **moderate positive market response**.\n\n2. **Structure of Predictions:**\n   - Predictions will align with the **test set** structure, producing one value for **`responder_6`** per row (i.e., per **symbol_id**, **date_id**, and **time_id**).\n   - The prediction will be made sequentially, as the **evaluation API** serves data one timestep at a time.\n\n3. **Weighted Evaluation:**\n   - Predictions will contribute to the **weighted zero-mean R² score**, where the `weight` column determines how much each prediction influences the final evaluation metric.\n\n**---**\n\n### 📊 **Example of Predicted Output:**\n\nThe predicted output for **`responder_6`** will look like this in the final submission:\n\n| **date_id** | **time_id** | **symbol_id** | **responder_6 (Predicted)** |\n|-------------|-------------|---------------|-----------------------------|\n| 1001        | 15          | 200           | 3.42                        |\n| 1001        | 15          | 201           | -1.87                       |\n| 1001        | 15          | 202           | 0.12                        |\n| 1001        | 16          | 200           | 4.98                        |\n| 1001        | 16          | 203           | -3.45                       |\n\n**---**\n\n### 📈 **Interpreting Predicted Values:**\n1. **Positive Values:** \n   - Suggest a **positive market impact** or **gain** for the respective symbol at the given timestep.\n   - Example: A value of `3.42` indicates a moderately positive response.\n\n2. **Negative Values:** \n   - Indicate a **negative market impact** or **loss**.\n   - Example: A value of `-3.45` reflects a strong negative response.\n\n3. **Values Close to Zero:**\n   - Implies a **neutral** or **insignificant** market response.\n   - Example: A value of `0.12` shows little expected change.\n\n**---**\n\n### 🌍 **Real-Life Implications:**\nIn real-world financial trading, the predicted **`responder_6`** could be used in several ways:\n- **Portfolio Adjustments:** Deciding whether to increase or decrease positions for specific financial instruments.\n- **Market Making:** Determining whether to buy or sell certain assets at specific times.\n- **Risk Management:** Identifying high-risk assets with extreme predictions (positive or negative).\n\nThis predictive output is critical for making informed trading decisions and managing financial strategies effectively. 🚀\n---\n\n\n## 📌 5. Evaluation\n\n\n\n### How the Evaluation API Works\nThe competition uses an evaluation API that:\n\n- **Serves Test Data Incrementally**: Data for a single `date_id` and `time_id` pair is served at a time.\n- **Requires Real-Time Predictions**: Your model must predict responder_6 at each timestep without looking ahead.\n- **Handles Large Data**: The test set is expected to be as large as the training set (~4.5 million rows).\n\n🚀 **Let’s build a strong machine learning solution for the Jane Street competition!**  \n\n\n### Evaluation Metric\nSubmissions are evaluated using a **weighted R-squared score** for responder_6. The formula is:\n\n$$\nR^2 = 1 - \\frac{\\sum w_i (y_i - \\hat{y}_i )^2}{\\sum w_i y_i^2}\n$$\n\nWhere:\n\n\n- $ y_i $: Ground truth values of `responder_6`.\n\n- $ \\hat{y}_i $: Predicted values.\n\n- $ w_i $: Sample weights provided in the dataset.\n\n\n---\n\n## 📌 6. Real-Life Impact\nSuccessfully predicting responder_6 in this competition reflects solving real-life challenges in financial trading, where accurate forecasting can:\n\n- Optimize trading strategies.\n- Enhance risk management.\n- Improve profitability in high-frequency trading environments.\n\n---\n\n## 📌 7. Next Steps\nUse this breakdown as a guide to:\n- Dive deeper into feature analysis.\n- Develop robust time-series models.\n- Experiment with innovative ideas to tackle the challenges of non-stationarity and fat-tailed distributions.\n","metadata":{}},{"cell_type":"markdown","source":"# IMPLEMENTATION","metadata":{}},{"cell_type":"markdown","source":"## 1. Visualizations and Understanding ","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv, pd.read_parquet )\nimport polars as pl\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom matplotlib import pyplot as plt\nfrom matplotlib.ticker import MaxNLocator, FormatStrFormatter, PercentFormatter\n\nimport os, gc\nfrom tqdm.auto import tqdm\nimport pickle # module to serialize and deserialize objects\nimport re # for Regular expression operations \n\nimport tensorflow as tf\nfrom tensorflow.keras import layers, models\nfrom tensorflow.keras.optimizers import Adam\n\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom torch.utils.data  import Dataset, DataLoader\nfrom pytorch_lightning import (LightningDataModule, LightningModule, Trainer)\nfrom pytorch_lightning.callbacks import EarlyStopping, ModelCheckpoint, Timer\n\nfrom sklearn.metrics import r2_score\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.ensemble import VotingRegressor\n\nimport lightgbm as lgb\nfrom lightgbm import LGBMRegressor\n\nfrom xgboost import XGBRegressor\nfrom catboost import CatBoostRegressor\n\nimport warnings\nwarnings.filterwarnings('ignore')\npd.options.display.max_columns = None\n\nimport kaggle_evaluation.jane_street_inference_server\ngridColor = 'lightgrey'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:53:22.213258Z","iopub.execute_input":"2025-01-23T04:53:22.213677Z","iopub.status.idle":"2025-01-23T04:53:56.752700Z","shell.execute_reply.started":"2025-01-23T04:53:22.213636Z","shell.execute_reply":"2025-01-23T04:53:56.751516Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time\npath = \"/kaggle/input/jane-street-real-time-market-data-forecasting\"\nsamples = [] \n\n# Load a data from each file:\nr = range(2)\nfor i in r:\n    file_path = f\"{path}/train.parquet/partition_id={i}/part-0.parquet\"\n    part = pd.read_parquet(file_path)\n    samples.append(part)\n    \nsample_df = pd.concat(samples, ignore_index=True) # Concatenate all samples into one DataFrame if needed\n\nsample_df.round(1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:53:56.753790Z","iopub.execute_input":"2025-01-23T04:53:56.754697Z","iopub.status.idle":"2025-01-23T04:54:09.920681Z","shell.execute_reply.started":"2025-01-23T04:53:56.754659Z","shell.execute_reply":"2025-01-23T04:54:09.919427Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### Simplified example to illustrate these columns with real world examples.\n\n### Example Data Row:\n- **date_id**: 10 (e.g., January 10th)\n- **time_id**: 3 (e.g., 10:30 AM)\n- **symbol_id**: 1 (Stock of Apple)\n\n   NOTE: In other rows symbol_id 2 could be stock of Amazon ; symbol_id 3 Google stocks; .......and other companies stocks upto symbol_18 then it could be followed by Bonds, Crypto currencies, Commodities(Gold and Crude oil, etc), etc..\n- **weight**: 1.5\n- **feature_00**: 5000 (e.g., Trading Volume)\n- **feature_01**: 25.0 (e.g., P/E Ratio)\n- ...\n- **feature_78**: 0.8 (e.g., Volatility Index)\n- **responder_0**: +0.7 (eg. An indicator of the market impact of trading the stock or its liquidity. i.e it could reflect how easily the stock can be bought or sold without affecting its price.)\n- ...\n- **feature_8**: -0.1 (eg. the stocks price is expected to decrease by 0.1% on the next day)\n- **responder_6**: +1.2 (eg. the stocks price increased by 1.2% during next time period)\n\n### Explanation:\n#### Date and Time:\nOn **January 10th** at **10:30 AM**, we have an observation for **Stock A**.\n\n#### Stock Identifier:\n- **symbol_id 1** refers to **Stock of Apple**.\n\n#### Weight:\n- The **weight** of **1.5** suggests this data point is moderately important.\n\n#### Features:\n- **feature_00** (5000) indicates that **5,000 shares** were traded during this time.\n- **feature_01** (25.0) indicates a **P/E ratio of 25**, suggesting how the stock is valued compared to its earnings.\n- Other features provide additional data about the stock's condition.\n\n#### Responder:\n- **responder_6 (+1.2)** signifies that the stock's price increased by **1.2%** during the next time period. This is the target value you aim to predict using the features.\n","metadata":{}},{"cell_type":"markdown","source":"### TIME SERIES DATA ANALYSIS","metadata":{}},{"cell_type":"code","source":"train = sample_df\ntrain['N'] = train.index.values\ntrain['id'] = train.index.values\n\nxx = sample_df[(sample_df.symbol_id == 1)]['id']\nyy = sample_df[(sample_df.symbol_id == 1)]['responder_6']\n\nplt.figure(figsize=(16, 5))\nplt.plot(xx, yy, color='green', linewidth=0.05)\nplt.suptitle('Returns of responder_6 for Financial Instrument (symbol_id = 1)', weight='bold', fontsize=16)\nplt.xlabel(\"Time (Sequential ID)\", fontsize=12)\nplt.ylabel(\"Returns (responder_6)\", fontsize=12)\nplt.grid(color='lightgray', linewidth=0.8)\nplt.axhline(0, color='red', linestyle='-', linewidth=1.2)\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:54:09.921904Z","iopub.execute_input":"2025-01-23T04:54:09.922328Z","iopub.status.idle":"2025-01-23T04:54:11.247125Z","shell.execute_reply.started":"2025-01-23T04:54:09.922291Z","shell.execute_reply":"2025-01-23T04:54:11.245851Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This graph visualizes the behavior of `responder_6`'s returns for a specific financial instrument identified by `symbol_id = 1`.\n\n\n#### **Key Definitions**  \n\n**What are returns?**: Returns are a way to quantify how much money is made or lost on a financial instrument (like a stock, bond, or asset) over a given period.\nA positive return means a profit, while a negative return means a loss.\neg:\nIf you bought a stock for 100 and its value rose to 105, your return is 5 or 5%.\nIf the value fell to 95, your return is -$5 or -5%.\n\n---\n\nNow, Here's how to interpret the graph:\n\n- **X-Axis (Time):** Represents sequential **time** using the `id` column from the dataset.\n\nHow does the dataset track time? \nWe use the **row index (`id`)** for simplicity in this graph.  \nThe **row index (`id`)** indirectly represents a chronological order since rows are sorted by `date_id` and `time_id`.  \n(Note: For more detailed analysis, we could combine `date_id` and `time_id` to create precise timestamps, which we'll do in future)\n\n\n- **Y-Axis (Returns):** Represents the `responder_6` values, which measure the financial returns for the selected instrument.( eg. stock prices of Amazon( which we call symbol_id 1)\n\n---\n\n\n**Graph Features**\n1. **Green Line:** Shows how the returns (`responder_6`) change over time for `symbol_id = 1`. eg. How \"Amazon's\" \"stock prices\" change over time\n2. **Red Line (Baseline):** Marks the zero-return baseline. This helps distinguish between:\n   - Positive returns (above the red line).\n   - Negative returns (below the red line).\n\n**Why Does This Matter?**\nThe dataset contains features and responders (`responder_6`) for multiple financial instruments. Each `symbol_id` corresponds to one such instrument. \nNow by selecting rows with `symbol_id = 1` and plotting their `responder_6` values, we can observe the returns for this specific instrument over time.\n\nEven though the dataset doesn’t explicitly track time, the row order (`id`) lets us derive temporal behavior, enabling this analysis.\n\n**Key Insights**\n- This graph helps us analyze the returns trend for a single instrument (`symbol_id = 1`) over the dataset's timeframe.\n- Positive and negative fluctuations are easily identified with respect to the zero-return line (red).\n\n","metadata":{}},{"cell_type":"code","source":"# Plotting cumulative responder_6 for symbol_id=1\nplt.figure(figsize=(14, 4))\nplt.plot(xx, yy.cumsum(), color='green', linewidth=0.6)\nplt.suptitle('Cumulative Responder_6 (for Symbol ID = 1)', weight='bold', fontsize=16)\nplt.xlabel(\"Time (which we got from Sequential ID column)\", fontsize=12)\nplt.ylabel(\"Cumulative Returns\", fontsize=12)\nplt.yticks(np.arange(-500, 1000, 250))\nplt.grid(color='lightgray', linewidth=0.7)\nplt.axhline(0, color='red', linestyle='-', linewidth=0.7)  # Zero baseline\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:54:11.249266Z","iopub.execute_input":"2025-01-23T04:54:11.249599Z","iopub.status.idle":"2025-01-23T04:54:11.470698Z","shell.execute_reply.started":"2025-01-23T04:54:11.249569Z","shell.execute_reply":"2025-01-23T04:54:11.469337Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\nThis graph visualizes the **cumulative returns** for a specific financial instrument identified by `symbol_id = 1`. Unlike the previous graph that shows individual returns (`responder_6`), this one illustrates the **accumulated trend** of returns over time.  \n\n---\n\n#### **Key Definitions**  \n\n1. **Cumulative Returns**:  \n   Cumulative returns measure the running total of returns (`responder_6`) for an asset over time.  \n   For example:\n   - If the first few returns are 2, -1, and 3, the cumulative values will be 2, 1, and 4, respectively. i.e. 2, 2-1, 2-1+4,........etc cumilative sum\n   - Positive values indicate an overall gain, while negative values indicate an overall loss.\n\n2. **Symbol ID**:  \n   `symbol_id` refers to a specific financial instrument (e.g., a stock or asset). Here, we focus on `symbol_id = 1`.eg. Amazon's stock prices\n\n---\n\n\n#### **Interpreting the Graph**\n\n- **X-Axis (Time):**  \n  Represents the sequential **time** using the `id` column from the dataset. Each `id` corresponds to a specific moment in the dataset's timeframe, reflecting the chronological order of events.  \n  Since the dataset is sorted by `date_id` and `time_id`, the `id` provides an implicit temporal sequence. This means we can interpret the changes in `responder_6` over time using the row index (`id`) as a proxy for time.\n\n- **Y-Axis (Cumulative Returns):**  \n  Shows the cumulative sum of `responder_6` values for `symbol_id = 1`.  \n  Positive values (above zero) indicate an overall profit, while negative values (below zero) indicate an overall loss.\n\n---\n\n\n#### **Why Does This Matter?**\nBy analyzing cumulative returns:\n- We can observe the long-term performance of `symbol_id = 1`. eg. long-term performance of Amazon stocks  \n- The graph provides insight into the sustainability of gains or losses for this financial instrument over time.\n\n---\n\n#### **Key Insights**\n- **Positive Trends:**  \n  If the green line steadily moves upward, it indicates consistent profits over time.  \n\n- **Negative Trends:**  \n  If the green line trends downward, it reflects sustained losses.  \n\n- **Fluctuations:**  \n  Sharp peaks or dips indicate periods of high volatility for the instrument.\n\n","metadata":{}},{"cell_type":"code","source":"# for symbol_id == 0\nplt.figure(figsize=(18, 7))\npredictor_cols = [col for col in sample_df.columns if 'responder' in col]\nfor i in predictor_cols: \n    if i == 'responder_6': \n        c='red'\n        lw=2.5\n        plt.plot((sample_df[sample_df.symbol_id == 0].groupby(['date_id'])[i].mean()).cumsum(), linewidth = lw, color = c)\n    else: \n        lw=1\n        plt.plot((sample_df[sample_df.symbol_id == 0].groupby(['date_id'])[i].mean()).cumsum(), linewidth = lw)\n\nplt.xlabel('Trade days (from date_id column)')\nplt.ylabel('Cumulative response (from responder values)')\nplt.title('Response time series for symbol_id 0 over trade days  \\n (Responder 6 (red) and other responders in diff colors)', weight='bold')\nplt.grid(visible=True, color = gridColor, linewidth = 0.7)\nplt.axhline(0, color='green', linestyle='-', linewidth=2)\nplt.legend(predictor_cols)\nsns.despine()\n#plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:54:11.472557Z","iopub.execute_input":"2025-01-23T04:54:11.472914Z","iopub.status.idle":"2025-01-23T04:54:13.742133Z","shell.execute_reply.started":"2025-01-23T04:54:11.472888Z","shell.execute_reply":"2025-01-23T04:54:13.740869Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport re\n\n# Parameters\ndf_train = sample_df\ns_id = 0  # Change params to take a look at other symbols\nres_columns = [col for col in df_train.columns if re.match(\"responder_\", col)]\nn_rows = len(res_columns)  # Number of responders\n\n# Plot configuration\nfig, axs = plt.subplots(figsize=(18, 4 * n_rows))\ngrid_color = '#cccccc'  # Define gridline color\n\n# Loop through responders to plot\nfor j, responder in enumerate(res_columns):\n    xx = df_train[df_train.symbol_id == s_id]['N']\n    yy = df_train[df_train.symbol_id == s_id][responder]\n    color = 'red' if j == 6 else 'black'  # Highlight responder_6 in red\n    \n    # Cumulative sum plot\n    ax1 = plt.subplot(n_rows, 3, j * 3 + 1)\n    ax1.plot(xx, yy.cumsum(), color=color, linewidth=0.8)\n    ax1.axhline(0, color='blue', linestyle='-', linewidth=0.9)\n    ax1.grid(color=grid_color)\n    ax1.set_title(f\"Cumulative Sum: {responder}\", fontsize=12)\n    \n    # Line plot\n    ax2 = plt.subplot(n_rows, 3, j * 3 + 2)\n    ax2.plot(xx, yy, color=color, linewidth=0.5)\n    ax2.axhline(0, color='blue', linestyle='-', linewidth=0.9)\n    ax2.grid(color=grid_color)\n    ax2.set_title(f\"Raw Values: {responder}\", fontsize=12)\n    \n    # Histogram plot\n    ax3 = plt.subplot(n_rows, 3, j * 3 + 3)\n    bins = 100  # Adjusted for better clarity\n    ax3.hist(yy, bins=bins, color=color, density=True, histtype=\"step\", linewidth=1.2)\n    ax3.hist(yy, bins=bins, color='lightgrey', density=True, alpha=0.7)\n    ax3.grid(color=grid_color)\n    ax3.set_title(f\"Histogram: {responder}\", fontsize=12)\n    ax3.set_xlim([-2.5, 2.5])\n    ax3.set_ylim([0, 3.5])\n\n# Global styling\nfig.suptitle(f\"Responder Analysis for Symbol ID {s_id}\", fontsize=16, weight='bold')\nfig.tight_layout(pad=3)\nfig.patch.set_linewidth(2)\nfig.patch.set_edgecolor('#000000')\nfig.patch.set_facecolor('#f9f9f9')\nplt.show()\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:54:13.743215Z","iopub.execute_input":"2025-01-23T04:54:13.743506Z","iopub.status.idle":"2025-01-23T04:54:25.968458Z","shell.execute_reply.started":"2025-01-23T04:54:13.743482Z","shell.execute_reply":"2025-01-23T04:54:25.967124Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#responder_6 for all different symbol ID's \n\nres_columns = [col for col in df_train.columns if re.match(\"responder_\", col)]\nrow=10\nfig, axs = plt.subplots(figsize=(18, 5*row))\nb=300\nj = 0\nfor i in range(1, 3 * row + 1, 3):\n    xx= sample_df[(sample_df.symbol_id==j)] ['N']\n    yy= sample_df[(sample_df.symbol_id==j)]['responder_6']\n    c='black'\n        \n    ax1 = plt.subplot(row, 3, i)\n    ax1.plot(   xx,yy.cumsum()   , color = c, linewidth =0.8 )\n    plt.axhline(0, color='red', linestyle='-', linewidth=0.7)\n    plt.grid(color = gridColor)\n    plt.xlabel('Time')\n    \n    ax2 = plt.subplot(row, 3, i+1)\n    ax2.plot(xx,yy   , color = c, linewidth =0.05)\n    plt.axhline(0, color='red', linestyle='-', linewidth=0.7)\n    ax2.set_title(f\"symbol_id={j}\", fontsize = '14')\n    plt.grid(color = gridColor)\n    plt.xlabel('Time')\n    \n    ax3 = plt.subplot(row, 3, i+2)\n    ax3.hist(yy, bins=b, color = c, density=True, histtype=\"step\" )\n    ax3.hist(yy, bins=b, color = 'lightgrey',density=True)\n    plt.grid(color = gridColor)\n    ax3.set_xlim([-2.5, 2.5])\n    ax3.set_ylim([0, 1.5])\n    plt.xlabel('Time')\n    \n    j = j + 1\n    \nfig.patch.set_linewidth(3)\nfig.patch.set_edgecolor('#000000')\nfig.patch.set_facecolor('#eeeeee') \nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:54:25.969903Z","iopub.execute_input":"2025-01-23T04:54:25.970257Z","iopub.status.idle":"2025-01-23T04:54:38.323363Z","shell.execute_reply.started":"2025-01-23T04:54:25.970230Z","shell.execute_reply":"2025-01-23T04:54:38.321937Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#missing values\n\ndf_train = sample_df\nplt.figure(figsize=(20, 3))    # Plot missing values\nplt.bar(x=df_train.isna().sum().index, height=df_train.isna().sum().values, color=\"red\", label='missing')   # analog: using missingno\nplt.xticks(rotation=90)\nplt.title(f'Missing values over the {len(df_train)} samples which have a target')\nplt.grid()\nplt.legend()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:54:38.324634Z","iopub.execute_input":"2025-01-23T04:54:38.325066Z","iopub.status.idle":"2025-01-23T04:54:41.209983Z","shell.execute_reply.started":"2025-01-23T04:54:38.325017Z","shell.execute_reply":"2025-01-23T04:54:41.208262Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#features\n\nfeatures = pd.read_csv(f\"{path}/features.csv\")\nfeatures","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:54:41.211258Z","iopub.execute_input":"2025-01-23T04:54:41.211698Z","iopub.status.idle":"2025-01-23T04:54:41.252392Z","shell.execute_reply.started":"2025-01-23T04:54:41.211655Z","shell.execute_reply":"2025-01-23T04:54:41.251071Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Which features have which all tags\n\nplt.figure(figsize=(18, 6))\nplt.imshow(features.iloc[:, 1:].T.values, cmap=\"gray_r\")\nplt.xlabel(\"feature_00 - feature_78\")\nplt.ylabel(\"tag_0 - tag_16\")\nplt.yticks(np.arange(17))\nplt.xticks(np.arange(79))\nplt.grid(color = 'lightgrey')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:54:41.253438Z","iopub.execute_input":"2025-01-23T04:54:41.253850Z","iopub.status.idle":"2025-01-23T04:54:41.908872Z","shell.execute_reply.started":"2025-01-23T04:54:41.253815Z","shell.execute_reply":"2025-01-23T04:54:41.907718Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#correlation bw all features based on their tags\n\nplt.figure(figsize=(11, 11))\nmatrix = features[[ f\"tag_{no}\" for no in range(0,17,1) ] ].T.corr()\nsns.heatmap(matrix, square=True, cmap=\"coolwarm\", alpha =0.9, vmin=-1, vmax=1, center= 0, linewidths=0.5, linecolor='white')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:54:41.909652Z","iopub.execute_input":"2025-01-23T04:54:41.909931Z","iopub.status.idle":"2025-01-23T04:54:42.659307Z","shell.execute_reply.started":"2025-01-23T04:54:41.909907Z","shell.execute_reply":"2025-01-23T04:54:42.658015Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#weights\n\nsample_df['weight'].describe().round(1)\n\n\nplt.figure(figsize=(8,3))\nplt.hist(sample_df['weight'], bins=30, color='grey', edgecolor = 'white',density=True )\nplt.title('Distribution of weights')\nplt.grid(color = 'lightgrey', linewidth=0.5)\nplt.axvline(1.7, color='red', linestyle='-', linewidth=0.7)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:54:42.663184Z","iopub.execute_input":"2025-01-23T04:54:42.663520Z","iopub.status.idle":"2025-01-23T04:54:43.116706Z","shell.execute_reply.started":"2025-01-23T04:54:42.663495Z","shell.execute_reply":"2025-01-23T04:54:43.115268Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#responders\n\nresponders = pd.read_csv(f\"{path}/responders.csv\")\nresponders","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:54:43.117970Z","iopub.execute_input":"2025-01-23T04:54:43.118263Z","iopub.status.idle":"2025-01-23T04:54:43.137791Z","shell.execute_reply.started":"2025-01-23T04:54:43.118238Z","shell.execute_reply":"2025-01-23T04:54:43.136539Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#correlation matrix for all responders based on their tags\nplt.figure(figsize=(6, 6))\nresponders = pd.read_csv(f\"{path}/responders.csv\")\nmatrix = responders[[ f\"tag_{no}\" for no in range(0,5,1) ] ].T.corr()\nsns.heatmap(matrix, square=True, cmap=\"coolwarm\", alpha =0.9, vmin=-1, vmax=1, center= 0, linewidths=0.5, \n            linecolor='white', annot=True, fmt='.2f')\nplt.xlabel(\"Responder_0 - Responder_8\")\nplt.ylabel(\"Responder_0 - Responder_8\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:54:43.138934Z","iopub.execute_input":"2025-01-23T04:54:43.139257Z","iopub.status.idle":"2025-01-23T04:54:43.584797Z","shell.execute_reply.started":"2025-01-23T04:54:43.139231Z","shell.execute_reply":"2025-01-23T04:54:43.583688Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### What This Heatmap Shows:\nIt displays the correlation between responders (responder_0 to responder_8) based on their tag distributions (tag_0 to tag_4).  \nEach cell represents how similar or opposite two responders are in terms of their tag values.  \n- **Red** (positive correlation): Responders have similar tag patterns.  \n- **Blue** (negative correlation): Responders have opposite tag patterns.  \n- **Diagonal values** (1.00): Each responder is perfectly correlated with itself.\n\n#### How It's Useful:\n- **Identify similar responders**: Group them for analysis or recommendations.\n- **Detect outliers**: Responders behaving very differently from others.\n- **Feature selection**: Remove redundant responders with high correlation.\n- **Analyze behavioral trends**: Identify who reacts similarly or oppositely.\n","metadata":{}},{"cell_type":"markdown","source":"### Responders: analysis","metadata":{}},{"cell_type":"code","source":"col =[]\nfor i in range(9):\n    col.append(f\"responder_{i}\") \n\nsample_df[col].describe().round(1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:54:43.585926Z","iopub.execute_input":"2025-01-23T04:54:43.586332Z","iopub.status.idle":"2025-01-23T04:54:45.414560Z","shell.execute_reply.started":"2025-01-23T04:54:43.586294Z","shell.execute_reply":"2025-01-23T04:54:45.413378Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#responders distribution against other responders\n\nnumerical_features = []\nnumerical_features = sample_df.filter(regex='^responder_').columns.tolist()  # Separate responders\nnumerical_features.remove('responder_6')\n\ngs = 600  # Grid size for hexbin\nk = 1\ncol = 3\nrow = 3\n\nfig, axs = plt.subplots(row, col, figsize=(6*col, 5*row))\n\nfor i in numerical_features:\n    ax = plt.subplot(col, row, k)\n    \n    # Hexbin plot\n    hb = ax.hexbin(sample_df[i], sample_df['responder_6'], \n                   gridsize=gs, cmap='CMRmap', bins='log', alpha=0.8)\n    \n    # Labels and title for each subplot\n    ax.set_xlabel(f'{i}', fontsize=12)\n    ax.set_ylabel('responder_6', fontsize=12)\n    ax.tick_params(axis='x', labelsize=8)\n    ax.tick_params(axis='y', labelsize=8)\n    \n    # Adding a colorbar for the hexbin\n    cb = fig.colorbar(hb, ax=ax, orientation='vertical')\n    cb.set_label('Log Density', fontsize=10)\n    \n    k += 1\n\n# Add a main title for the figure\nfig.suptitle('Responder Relationships: Hexbin Plot with Responder_6', \n             fontsize=18, weight='bold', color='black')\n\n# Beautify the figure's border and background\nfig.patch.set_linewidth(3)\nfig.patch.set_edgecolor('#000000')\nfig.patch.set_facecolor('#eeeeee')   \n\nplt.tight_layout(rect=[0, 0, 1, 0.96])  # Adjust layout to fit the title\nplt.show()\n\n\n\n\n\n#Each subplot is a hexagonal binning plot showing the density of data points where values of responder_6 and another responder (responder_i) overlap.\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:54:45.415614Z","iopub.execute_input":"2025-01-23T04:54:45.415929Z","iopub.status.idle":"2025-01-23T04:55:12.070242Z","shell.execute_reply.started":"2025-01-23T04:54:45.415902Z","shell.execute_reply":"2025-01-23T04:55:12.068278Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The hexbin plots highlight whether responder_6 is correlated with the other responders. \nFor example:\n\nPositive Correlation: A diagonal pattern from bottom-left to top-right.\n\nNegative Correlation: A diagonal pattern from top-left to bottom-right.\n\nNo Correlation: A scattered pattern with no discernible shape.","metadata":{}},{"cell_type":"code","source":"## Responder_6 Distribution Against Selected Features\n\n# Responder_6 Distribution Against Selected Features\nnumerical_features = [f'feature_{i}' for i in ['05', '06', '07', '08', '12', '15', '19', '32', '38', '39', '50', '51', '65', '66', '67']]\n\ngs = 600  # Grid size for hexbin\nk = 1\ncol = 3\nrow = int(np.ceil(len(numerical_features) / col))  # Dynamically calculate rows\nsz = 5\nw = sz * col\nh = w / col * row\n\nfig, axs = plt.subplots(row, col, figsize=(w, h))\n\nfor i in numerical_features:\n    ax = plt.subplot(row, col, k)\n    \n    # Hexbin plot\n    hb = ax.hexbin(sample_df['responder_6'], sample_df[i], gridsize=gs, cmap='CMRmap', bins='log', alpha=0.5)\n    ax.set_xlabel(f'{i}', fontsize=10)\n    ax.set_ylabel('responder_6', fontsize=10)\n    ax.tick_params(axis='x', labelsize=8)\n    ax.tick_params(axis='y', labelsize=8)\n    \n    # Add colorbar\n    cb = fig.colorbar(hb, ax=ax, orientation='vertical', shrink=0.8)\n    cb.set_label('Log Density', fontsize=8)\n    \n    k += 1\n\n# Add a main title\nfig.suptitle('Responder_6 Distribution Against Selected Features', fontsize=16, weight='bold')\n\n# Beautify figure\nfig.patch.set_linewidth(3)\nfig.patch.set_edgecolor('#000000')\nfig.patch.set_facecolor('#eeeeee')   \n\nplt.tight_layout(rect=[0, 0, 1, 0.95])  # Adjust for title\nplt.show()\n  ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:55:12.071705Z","iopub.execute_input":"2025-01-23T04:55:12.072248Z","iopub.status.idle":"2025-01-23T04:56:00.321457Z","shell.execute_reply.started":"2025-01-23T04:55:12.072190Z","shell.execute_reply":"2025-01-23T04:56:00.319646Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Relationship between selected features from the dataset.\n\nnumerical_features = [f'feature_0{i}' for i in range(5, 9)] + [f'feature_{i}' for i in range(15, 20)]\n\na = 0\nk = 1\nn = 3  # Number of subplots per row\n\nfig_num = 1  # Track the figure number\nplt.figure(figsize=(15, 4))  # Initialize first figure\n\nfor i in numerical_features[:-1]:\n    a += 1\n    for j in numerical_features[a:]:\n        plt.subplot(1, n, k)\n        \n        # Create hexbin plot\n        hb = plt.hexbin(sample_df[i], sample_df[j], gridsize=200, cmap='CMRmap', bins='log', alpha=0.8)\n        plt.grid()\n        plt.xlabel(f'{i}', fontsize=10)\n        plt.ylabel(f'{j}', fontsize=10)\n        plt.tick_params(axis='x', labelsize=6)\n        plt.tick_params(axis='y', labelsize=6)\n        plt.title(f'{i} vs {j}', fontsize=12)\n        \n        # Add colorbar\n        if k == n:\n            cb = plt.colorbar(hb, ax=plt.gca(), shrink=0.8)\n            cb.set_label('Log Density', fontsize=8)\n        \n        k += 1\n        \n        # If row limit is reached, reset the figure\n        if k > n:\n            plt.suptitle(f'Pairwise Hexbin Plots - Figure {fig_num}', fontsize=14, weight='bold')\n            plt.show()\n            \n            k = 1\n            fig_num += 1\n            plt.figure(figsize=(15, 4))  # Start new figure\n\n# Handle the final figure\nif k != 1:\n    plt.tight_layout()\n    plt.suptitle(f'Pairwise Hexbin Plots - Figure {fig_num}', fontsize=14, weight='bold')\n    plt.show()\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:56:00.322681Z","iopub.execute_input":"2025-01-23T04:56:00.323046Z","iopub.status.idle":"2025-01-23T04:56:39.329843Z","shell.execute_reply.started":"2025-01-23T04:56:00.323015Z","shell.execute_reply":"2025-01-23T04:56:39.328579Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Assuming df_train is your DataFrame (e.g., sample_df)\ndf_train = sample_df\n\n# Define feature columns\nfeature_cols = [f'feature_{i:02d}' for i in range(79)]\n\n# Compute statistical summaries for feature columns\nfeature_stats = df_train[feature_cols].describe().transpose()\n\n# Reset index and rename columns\nfeature_stats = feature_stats.reset_index().rename(columns={'index': 'Feature'})\n\n# Select the desired statistics\ndesired_stats = ['Feature', 'mean', 'std', 'min', 'max']\n\n# Display the DataFrame\nprint(feature_stats[desired_stats])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:58:17.936380Z","iopub.execute_input":"2025-01-23T04:58:17.936780Z","iopub.status.idle":"2025-01-23T04:58:33.922283Z","shell.execute_reply.started":"2025-01-23T04:58:17.936749Z","shell.execute_reply":"2025-01-23T04:58:33.921184Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## INFERENCE FROM ABOVE EXAMINATIONS OF DATASET","metadata":{}},{"cell_type":"markdown","source":"### **Feature Distribution Insights**\n\n#### Columns with 100% Missing Values:\n- **feature_21**, **feature_26**, **feature_27**, and **feature_31** have no usable data.\n\n#### High Missingness (≥ 67%):\n- Features like **feature_00** through **feature_04** are missing in approximately **67%** of records.\n\n#### Moderate Missingness (~15%):\n- Features such as **feature_39**, **feature_42**, **feature_50**, and **feature_53** have **15%** missing values.\n\n#### abs Low Missingness (< 5%):\n- Features such as **feature_17**, **feature_08**, **feature_51**, and others have negligible missing values (**< 5%**).\n\n#### Imputation Actions:\n- Missing values in **feature_00** were successfully imputed using the **median**.\n\n#### Remaining Missing Data:\n- Features with significant missing data post-imputation include:\n  - **feature_01** to **feature_04** (67% missing).\n  - **feature_39**, **feature_42** (~15% missing).\n\n\n---\n\n### Responder Feature Summary\n\nThe **responder_*** columns exhibit the following statistical properties:\n- **Mean**: Close to 0 for all responders.\n- **Standard Deviation (Std)**: Values range from **0.8 to 1.1**.\n- **Min and Max**: Range spans from **-5 to 5**.\n- **Quartiles**: Median (50%) is consistently around **0** for all responders, indicating a central tendency near zero.\n\n---\n\n### Next Steps for Data Cleaning and Analysis\n\n#### Columns with 100% Missing Values:\n- Consider removing these features (**feature_21**, **feature_26**, **feature_27**, **feature_31**) as they provide no value.\n\n#### High-Missing Features (e.g., feature_00 to feature_04):\n- Impute missing values with domain-appropriate methods, or assess if these columns can be excluded based on feature importance analysis.\n\n#### Adam Responder Data:\n- Explore relationships between responder columns and the target to determine predictive relevance.\n\n#### Feature Engineering:\n- Address skewed distributions and normalize or scale features where necessary, especially for columns like **feature_39** or **feature_42**.\n","metadata":{}},{"cell_type":"markdown","source":"## DATASET CLEANING BASED ON ABOVE OBERVATIONS","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom sklearn.preprocessing import StandardScaler\nimport os\nfrom scipy.stats import skew\nimport warnings\nwarnings.filterwarnings('ignore')\n\ndef analyze_missing_data(df):\n    \"\"\"\n    Analyze missing data patterns and categorize features based on missing percentages\n    \"\"\"\n    # Calculate missing percentages for each column\n    missing_percentages = (df.isnull().sum() / len(df)) * 100\n    \n    # Categorize features based on missing percentages\n    completely_missing = missing_percentages[missing_percentages == 100].index.tolist()\n    high_missing = missing_percentages[(missing_percentages >= 50) & (missing_percentages < 100)].index.tolist()\n    moderate_missing = missing_percentages[(missing_percentages >= 25) & (missing_percentages < 50)].index.tolist()\n    low_missing = missing_percentages[(missing_percentages > 0) & (missing_percentages < 25)].index.tolist()\n    \n    print(\"\\nMissing Data Analysis:\")\n    print(f\"Completely Missing (100%): {len(completely_missing)} features\")\n    print(f\"High Missing (50-99%): {len(high_missing)} features\")\n    print(f\"Moderate Missing (25-49%): {len(moderate_missing)} features\")\n    print(f\"Low Missing (1-24%): {len(low_missing)} features\")\n    \n    return {\n        'completely_missing': completely_missing,\n        'high_missing': high_missing,\n        'moderate_missing': moderate_missing,\n        'low_missing': low_missing,\n        'missing_percentages': missing_percentages\n    }\n\ndef handle_missing_data(df, missing_analysis, target_col='responder_6'):\n    \"\"\"\n    Handle missing data based on different categories and their relationship with target\n    \"\"\"\n    print(\"\\nHandling Missing Data:\")\n    \n    # 1. Remove completely missing columns\n    df.drop(columns=missing_analysis['completely_missing'], inplace=True)\n    print(f\"Removed {len(missing_analysis['completely_missing'])} completely missing columns\")\n    \n    # 2. Handle high missing features\n    high_missing_correlations = {}\n    for feature in missing_analysis['high_missing']:\n        if feature.startswith('feature_'):  # Only process feature columns\n            # Calculate correlation with target using available data\n            valid_data = df[[feature, target_col]].dropna()\n            if len(valid_data) > 0:\n                corr = valid_data[feature].corr(valid_data[target_col])\n                high_missing_correlations[feature] = corr\n    \n    # Keep high-missing features with significant correlation (|corr| >= 0.01)\n    high_missing_to_keep = [f for f, c in high_missing_correlations.items() if abs(c) >= 0.01]\n    high_missing_to_drop = [f for f in missing_analysis['high_missing'] \n                           if f.startswith('feature_') and f not in high_missing_to_keep]\n    \n    # Drop high-missing features with low correlation\n    df.drop(columns=high_missing_to_drop, inplace=True)\n    print(f\"Dropped {len(high_missing_to_drop)} high-missing features with low correlation\")\n    \n    # 3. Handle moderate missing features\n    moderate_missing_correlations = {}\n    for feature in missing_analysis['moderate_missing']:\n        if feature.startswith('feature_'):\n            valid_data = df[[feature, target_col]].dropna()\n            if len(valid_data) > 0:\n                corr = valid_data[feature].corr(valid_data[target_col])\n                moderate_missing_correlations[feature] = corr\n    \n    # For moderate-missing features, use more sophisticated imputation based on correlation\n    for feature in missing_analysis['moderate_missing']:\n        if feature.startswith('feature_'):\n            if abs(moderate_missing_correlations.get(feature, 0)) >= 0.05:\n                # For features with higher correlation, use more sophisticated imputation\n                # Here we'll use interpolation with forward and backward fill\n                df[feature].interpolate(method='linear', limit_direction='both', inplace=True)\n                # Fill any remaining missing values with median\n                df[feature].fillna(df[feature].median(), inplace=True)\n            else:\n                # For features with lower correlation, use median imputation\n                df[feature].fillna(df[feature].median(), inplace=True)\n    \n    print(f\"Handled {len(missing_analysis['moderate_missing'])} moderate-missing features\")\n    \n    # 4. Handle low missing features\n    for feature in missing_analysis['low_missing']:\n        if feature.startswith('feature_'):\n            # For low missing values, use interpolation\n            df[feature].interpolate(method='linear', limit_direction='both', inplace=True)\n            # Fill any remaining missing values with median\n            df[feature].fillna(df[feature].median(), inplace=True)\n    \n    print(f\"Handled {len(missing_analysis['low_missing'])} low-missing features\")\n    \n    return df\n\ndef load_and_clean_data(path, num_partitions=2):\n    \"\"\"\n    Load and clean data from specified partitions\n    \"\"\"\n    print(\"Loading data...\")\n    samples = []\n    \n    # Load data from specified partitions\n    for i in range(num_partitions):\n        file_path = f\"{path}/train.parquet/partition_id={i}/part-0.parquet\"\n        part = pd.read_parquet(file_path)\n        samples.append(part)\n    \n    # Combine samples\n    df = pd.concat(samples, ignore_index=True)\n    print(f\"Loaded {len(df)} rows from {num_partitions} partitions\")\n    \n    # Analyze missing data\n    missing_analysis = analyze_missing_data(df)\n    \n    # Handle missing data\n    df = handle_missing_data(df, missing_analysis)\n    \n    # Step 3: Analyze responder relationships\n    responder_cols = [f'responder_{i}' for i in range(9) if i != 6]\n    responder_corrs = {col: df[col].corr(df['responder_6']) for col in responder_cols}\n    significant_responders = [r for r, c in responder_corrs.items() if abs(c) >= 0.1]\n    \n    # Keep only significant responders and target\n    responders_to_drop = [r for r in responder_cols if r not in significant_responders]\n    df.drop(columns=responders_to_drop, inplace=True)\n    print(\"\\nKept significant responders:\", significant_responders)\n    \n    # Step 4: Feature Engineering\n    # Get remaining feature columns\n    feature_cols = [col for col in df.columns if col.startswith('feature_')]\n    \n    # Calculate skewness for features\n    skewed_features = [col for col in feature_cols if abs(skew(df[col].dropna())) > 0.5]\n    \n    # Log transform skewed features (adding 1 to handle zeros/negative values)\n    for col in skewed_features:\n        min_val = df[col].min()\n        if min_val <= 0:\n            df[col] = np.log1p(df[col] - min_val + 1)\n    print(\"\\nApplied log transformation to skewed features:\", len(skewed_features))\n    \n    # Standardize features\n    scaler = StandardScaler()\n    df[feature_cols] = scaler.fit_transform(df[feature_cols])\n    print(\"\\nStandardized all features\")\n    \n    return df, scaler, missing_analysis\n\ndef save_processed_files(df, scaler, missing_analysis, output_path):\n    \"\"\"\n    Save all necessary processed files\n    \"\"\"\n    # Create output directory if it doesn't exist\n    os.makedirs(output_path, exist_ok=True)\n    \n    # Save cleaned dataset\n    df.to_parquet(f\"{output_path}/cleaned_train.parquet\", index=False)\n    \n    # Save missing data analysis as separate files or pad the lists to make them equal length\n    missing_info = {}\n    max_length = max(len(v) for v in [\n        missing_analysis['completely_missing'],\n        missing_analysis['high_missing'],\n        missing_analysis['moderate_missing'],\n        missing_analysis['low_missing']\n    ])\n    \n    # Pad each list with None to make them equal length\n    missing_info = {\n        'completely_missing': missing_analysis['completely_missing'] + [None] * (max_length - len(missing_analysis['completely_missing'])),\n        'high_missing': missing_analysis['high_missing'] + [None] * (max_length - len(missing_analysis['high_missing'])),\n        'moderate_missing': missing_analysis['moderate_missing'] + [None] * (max_length - len(missing_analysis['moderate_missing'])),\n        'low_missing': missing_analysis['low_missing'] + [None] * (max_length - len(missing_analysis['low_missing']))\n    }\n    \n    # Save as DataFrame\n    pd.DataFrame(missing_info).to_csv(f\"{output_path}/missing_data_info.csv\", index=False)\n    \n    # Save scaler parameters\n    scaler_params = {\n        'mean': scaler.mean_,\n        'scale': scaler.scale_\n    }\n    pd.DataFrame(scaler_params).to_csv(f\"{output_path}/scaler_params.csv\", index=False)\n    \n    # Also save a summary of missing data analysis\n    missing_summary = {\n        'category': ['Completely Missing', 'High Missing', 'Moderate Missing', 'Low Missing'],\n        'count': [\n            len(missing_analysis['completely_missing']),\n            len(missing_analysis['high_missing']),\n            len(missing_analysis['moderate_missing']),\n            len(missing_analysis['low_missing'])\n        ]\n    }\n    pd.DataFrame(missing_summary).to_csv(f\"{output_path}/missing_data_summary.csv\", index=False)\n    \n    print(f\"\\nSaved processed files to {output_path}\")\n    \n    # Print summary of missing data\n    print(\"\\nMissing Data Summary:\")\n    print(pd.DataFrame(missing_summary))\n\n# Main execution\nif __name__ == \"__main__\":\n    # Set paths\n    input_path = \"/kaggle/input/jane-street-real-time-market-data-forecasting\"\n    output_path = \"/kaggle/working/processed_data\"\n    \n    # Process data\n    print(\"Starting data processing...\")\n    df_cleaned, scaler, missing_analysis = load_and_clean_data(input_path, num_partitions=2)\n    \n    # Save processed files\n    save_processed_files(df_cleaned, scaler, missing_analysis, output_path)\n    \n    # Display sample of cleaned data\n    print(\"\\nSample of cleaned data:\")\n    print(df_cleaned.round(1).head())\n    \n    # Display basic information about the cleaned dataset\n    print(\"\\nCleaned dataset info:\")\n    print(f\"Shape: {df_cleaned.shape}\")\n    print(f\"Features: {len([col for col in df_cleaned.columns if col.startswith('feature_')])}\")\n    print(f\"Responders: {len([col for col in df_cleaned.columns if col.startswith('responder_')])}\")\n    print(f\"Missing values: {df_cleaned.isnull().sum().sum()}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T05:34:13.067661Z","iopub.execute_input":"2025-01-23T05:34:13.068790Z","iopub.status.idle":"2025-01-23T05:35:32.650715Z","shell.execute_reply.started":"2025-01-23T05:34:13.068725Z","shell.execute_reply":"2025-01-23T05:35:32.649645Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time\n# Load the cleaned data\ncleaned_df = pd.read_parquet(\"/kaggle/working/processed_data/cleaned_train.parquet\")\n\ndf_train = cleaned_df\n\n\n\n# Calculate missing values percentage\nmissing_percentages = (df_train.isna().sum() / len(df_train) * 100).round(2)\n\n# Create a DataFrame with the percentages\nmissing_df = pd.DataFrame({\n    'Column': missing_percentages.index,\n    'Missing_Percentage': missing_percentages.values\n})\n\n# Sort by percentage in descending order\nmissing_df = missing_df.sort_values('Missing_Percentage', ascending=False)\n\n# Display only columns with missing values (optional)\nmissing_df = missing_df[missing_df['Missing_Percentage'] > 0]\n\nprint(f\"Missing values percentage over {len(df_train)} samples:\")\nprint(missing_df.to_string(index=False))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T05:53:11.957199Z","iopub.execute_input":"2025-01-23T05:53:11.957632Z","iopub.status.idle":"2025-01-23T05:53:16.150064Z","shell.execute_reply.started":"2025-01-23T05:53:11.957603Z","shell.execute_reply":"2025-01-23T05:53:16.148747Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"So now only feature_04 has missing values so lets just fill all these with median","metadata":{}},{"cell_type":"code","source":"# Load the current cleaned dataset\ndf = pd.read_parquet(\"/kaggle/working/processed_data/cleaned_train.parquet\")\n\n# Handle missing values in feature_04\nmedian_value = df['feature_04'].median()\ndf['feature_04'] = df['feature_04'].fillna(median_value)\n\n# Verify no missing values remain\nmissing_check = (df.isna().sum() / len(df) * 100).round(2)\nprint(\"Missing values percentage after cleaning:\")\nprint(missing_check[missing_check > 0])\n\n# Save the updated dataset, replacing the old one\ndf.to_parquet(\"/kaggle/working/processed_data/cleaned_train.parquet\")\n\nprint(\"\\nDataset updated and saved. No missing values remain.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T05:55:55.507629Z","iopub.execute_input":"2025-01-23T05:55:55.508303Z","iopub.status.idle":"2025-01-23T05:56:25.284306Z","shell.execute_reply.started":"2025-01-23T05:55:55.508262Z","shell.execute_reply":"2025-01-23T05:56:25.283128Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let's take a final check","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom scipy.stats import skew\n\n# Load the cleaned data\ncleaned_df = pd.read_parquet(\"/kaggle/working/processed_data/cleaned_train.parquet\")\n\n# Step 1: Check if columns with 100% missing values are removed\nremoved_columns = ['feature_21', 'feature_26', 'feature_27', 'feature_31']\nmissing_columns = [col for col in removed_columns if col in cleaned_df.columns]\nif not missing_columns:\n    print(\"✅ Columns with 100% missing values were successfully removed.\")\nelse:\n    print(f\"❌ These columns were not removed: {missing_columns}\")\n\n# Step 2: Check high-missing features for imputation or removal\nhigh_missing_features = ['feature_00', 'feature_01', 'feature_02', 'feature_03', 'feature_04']\nfor feature in high_missing_features:\n    if feature in cleaned_df.columns:\n        missing_percentage = cleaned_df[feature].isna().mean() * 100\n        if missing_percentage == 0:\n            print(f\"✅ Missing values in {feature} were handled correctly.\")\n        else:\n            print(f\"❌ {feature} still has {missing_percentage:.2f}% missing values.\")\n    else:\n        print(f\"✅ {feature} was removed from the dataset.\")\n\n# Step 3: Check responder columns and their correlation with the target\n\ntarget_column = \"responder_6\"\nresponder_columns = [col for col in cleaned_df.columns if \"responder\" in col]\n\nif responder_columns:\n    for col in responder_columns:\n        correlation = cleaned_df[col].corr(cleaned_df[target_column])\n        print(f\"Correlation of {col} with {target_column}: {correlation:.4f}\")\nelse:\n    print(\"✅ No responder columns are present or they were already validated.\")\n\n# Step 4: Check skewness and normalization/scaling of specific features\n# Replace with the actual list of skewed features\nskewed_features = ['feature_39', 'feature_42']\nfor feature in skewed_features:\n    if feature in cleaned_df.columns:\n        skewness = skew(cleaned_df[feature].dropna())\n        print(f\"Skewness of {feature}: {skewness:.4f}\")\n        if abs(skewness) < 1:\n            print(f\"✅ {feature} has been normalized/scaled successfully.\")\n        else:\n            print(f\"❌ {feature} might still be skewed (Skewness: {skewness:.4f}).\")\n    else:\n        print(f\"✅ {feature} was removed from the dataset.\")\n\n# Final missing value check (overall validation)\nmissing_percentages = (cleaned_df.isna().sum() / len(cleaned_df) * 100).round(2)\nremaining_missing = missing_percentages[missing_percentages > 0]\n\nif remaining_missing.empty:\n    print(\"✅ No missing values found in the cleaned dataset.\")\nelse:\n    print(f\"❌ Missing values still exist in these columns:\\n{remaining_missing}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T06:01:57.486253Z","iopub.execute_input":"2025-01-23T06:01:57.489747Z","iopub.status.idle":"2025-01-23T06:02:02.487482Z","shell.execute_reply.started":"2025-01-23T06:01:57.489549Z","shell.execute_reply":"2025-01-23T06:02:02.486022Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time\n# Load the cleaned data\ncleaned_df = pd.read_parquet(\"/kaggle/working/processed_data/cleaned_train.parquet\")\n\ndf_train = cleaned_df\n\n\n\n# Calculate missing values percentage\nmissing_percentages = (df_train.isna().sum() / len(df_train) * 100).round(2)\n\n# Create a DataFrame with the percentages\nmissing_df = pd.DataFrame({\n    'Column': missing_percentages.index,\n    'Missing_Percentage': missing_percentages.values\n})\n\n# Sort by percentage in descending order\nmissing_df = missing_df.sort_values('Missing_Percentage', ascending=False)\n\n# Display only columns with missing values (optional)\nmissing_df = missing_df[missing_df['Missing_Percentage'] > 0]\n\nprint(f\"Missing values percentage over {len(df_train)} samples:\")\nprint(missing_df.to_string(index=False))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T05:56:43.046972Z","iopub.execute_input":"2025-01-23T05:56:43.047424Z","iopub.status.idle":"2025-01-23T05:56:47.493616Z","shell.execute_reply.started":"2025-01-23T05:56:43.047389Z","shell.execute_reply":"2025-01-23T05:56:47.491810Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"So now no missing values remain","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\n\nsns.kdeplot(cleaned_df['feature_39'], label=\"Feature 39\", fill=True)\nsns.kdeplot(cleaned_df['feature_42'], label=\"Feature 42\", fill=True)\nplt.legend()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T06:03:12.047805Z","iopub.execute_input":"2025-01-23T06:03:12.048258Z","iopub.status.idle":"2025-01-23T06:03:55.455557Z","shell.execute_reply.started":"2025-01-23T06:03:12.048226Z","shell.execute_reply":"2025-01-23T06:03:55.454297Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Compute correlation of all features with responder_6\ncorrelations_with_target = cleaned_df.corr()['responder_6'].sort_values(ascending=False)\n\n# Display correlations (excluding responder_6 itself)\ncorrelations_with_target = correlations_with_target.drop('responder_6')\n\n# Thresholds for strong, moderate, and weak correlations\nstrong_threshold = 0.5\nmoderate_threshold = 0.3\n\n# Identify strong, moderate, and weak correlations\nstrong_corr = correlations_with_target[correlations_with_target >= strong_threshold]\nmoderate_corr = correlations_with_target[(correlations_with_target < strong_threshold) & (correlations_with_target >= moderate_threshold)]\nweak_corr = correlations_with_target[correlations_with_target < moderate_threshold]\n\nprint(\"Strong Correlations with responder_6:\")\nprint(strong_corr)\n\nprint(\"\\nModerate Correlations with responder_6:\")\nprint(moderate_corr)\n\nprint(\"\\nWeak Correlations with responder_6:\")\nprint(weak_corr)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T06:13:57.314427Z","iopub.status.idle":"2025-01-23T06:13:57.314816Z","shell.execute_reply":"2025-01-23T06:13:57.314671Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"investigate correlations among the responder columns where the correlation coefficient exceeds 0.5","metadata":{}},{"cell_type":"code","source":"# Correlation of each responder column with responder_6\ntarget_column = 'responder_6'\ncorrelation_with_target = {}\n\nfor col in responder_columns:\n    if col != target_column:\n        correlation_with_target[col] = cleaned_df[col].corr(cleaned_df[target_column])\n\n# Sort by correlation strength\nsorted_correlation = sorted(correlation_with_target.items(), key=lambda x: abs(x[1]), reverse=True)\n\nprint(\"Correlation of responders with responder_6:\")\nfor feature, corr in sorted_correlation:\n    print(f\"{feature}: {corr:.4f}\")\n\n\n\n\n\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T06:07:59.236888Z","iopub.execute_input":"2025-01-23T06:07:59.237414Z","iopub.status.idle":"2025-01-23T06:07:59.865005Z","shell.execute_reply.started":"2025-01-23T06:07:59.237381Z","shell.execute_reply":"2025-01-23T06:07:59.863731Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"When strong correlations are identified among features (e.g., responder columns), multicollinearity can negatively impact your model's performance by inflating variance and making feature importance unreliable. ","metadata":{}},{"cell_type":"markdown","source":"Use a simple model (e.g., Random Forest) to rank the predictors by their importance in predicting responder_6.","metadata":{}},{"cell_type":"code","source":"'''from sklearn.ensemble import RandomForestRegressor\nfrom sklearn.model_selection import train_test_split\n\n# Define features (excluding responder_6) and target\nX = cleaned_df[responder_columns].drop(columns=['responder_6'])\ny = cleaned_df['responder_6']\n\n# Train-test split\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)\n\n# Train a Random Forest model\nrf_model = RandomForestRegressor(random_state=42)\nrf_model.fit(X_train, y_train)\n\n# Feature importance\nfeature_importance = pd.DataFrame({\n    'Feature': X.columns,\n    'Importance': rf_model.feature_importances_\n}).sort_values(by='Importance', ascending=False)\n\nprint(\"Feature Importance for predicting responder_6:\")\nprint(feature_importance)\n'''","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T06:08:23.238586Z","iopub.execute_input":"2025-01-23T06:08:23.239113Z","iopub.status.idle":"2025-01-23T06:13:57.313286Z","shell.execute_reply.started":"2025-01-23T06:08:23.239080Z","shell.execute_reply":"2025-01-23T06:13:57.310875Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"ENSEMBLE_SOLUTIONS = ['SOLUTION_14','SOLUTION_5']\nOPTION,__WTS = 'option 91',[0.899, 0.28]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:56:39.542569Z","iopub.status.idle":"2025-01-23T04:56:39.543071Z","shell.execute_reply":"2025-01-23T04:56:39.542837Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def predict(test:pl.DataFrame, lags:pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:    \n    pdB = predict_14(test,lags).to_pandas()\n    pdC = predict_5 (test,lags).to_pandas()\n\n    pdB = pdB.rename(columns={'responder_6':'responder_B'})\n    pdC = pdC.rename(columns={'responder_6':'responder_C'})\n    pds = pd.merge(pdB,pdC, on=['row_id'])\n    pds['responder_6'] =\\\n        pds['responder_B'] *__WTS[0] +\\\n        pds['responder_C'] *__WTS[1] \n\n    display(pds)\n    predictions = test.select('row_id', pl.lit(0.0).alias('responder_6'))\n    pred = pds['responder_6'].to_numpy()\n    predictions = predictions.with_columns(pl.Series('responder_6', pred.ravel()))\n    return predictions","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:56:39.544176Z","iopub.status.idle":"2025-01-23T04:56:39.544625Z","shell.execute_reply":"2025-01-23T04:56:39.544433Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if 'SOLUTION_14' in ENSEMBLE_SOLUTIONS:    \n    \n    lags_ : pl.DataFrame | None = None\n\n    def predict_14(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n        global lags_\n        if lags is not None:\n            lags_ = lags\n\n        predictions_14 = test.select(\n            'row_id',\n            pl.lit(0.0).alias('responder_6'),\n        )\n        symbol_ids = test.select('symbol_id').to_numpy()[:, 0]\n\n        if not lags is None:\n            lags = lags.group_by([\"date_id\", \"symbol_id\"], maintain_order=True).last() # pick up last record of previous date\n            test = test.join(lags, on=[\"date_id\", \"symbol_id\"],  how=\"left\")\n        else:\n            test = test.with_columns(\n                ( pl.lit(0.0).alias(f'responder_{idx}_lag_1') for idx in range(9) )\n            )\n\n        preds = np.zeros((test.shape[0],))\n        preds += xgb_model.predict(test[xgb_feature_cols].to_pandas()) / 2\n        test_input = test[CONFIG.feature_cols].to_pandas()\n        test_input = test_input.fillna(method = 'ffill').fillna(0)\n        test_input = torch.FloatTensor(test_input.values).to(\"cuda:0\")\n        with torch.no_grad():\n            for i, nn_model in enumerate(tqdm(models)):\n                nn_model.eval()\n                preds += nn_model(test_input).cpu().numpy() / 10\n        print(f\"predict> preds.shape =\", preds.shape)\n\n        predictions_14 = \\\n        test.select('row_id').\\\n        with_columns(\n            pl.Series(\n                name   = 'responder_6', \n                values = np.clip(preds, a_min = -5, a_max = 5),\n                dtype  = pl.Float64,\n            )\n        )\n\n        # The predict function must return a DataFrame\n        #assert isinstance(predictions, pl.DataFrame | pd.DataFrame)\n        # with columns 'row_id', 'responer_6'\n        #assert list(predictions.columns) == ['row_id', 'responder_6']\n        # and as many rows as the test data.\n        #assert len(predictions) == len(test)\n\n        return predictions_14","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-23T04:56:39.545540Z","iopub.status.idle":"2025-01-23T04:56:39.546023Z","shell.execute_reply":"2025-01-23T04:56:39.545810Z"}},"outputs":[],"execution_count":null}]}