{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":84493,"databundleVersionId":9871156,"sourceType":"competition"},{"sourceId":9756372,"sourceType":"datasetVersion","datasetId":5973876},{"sourceId":9801075,"sourceType":"datasetVersion","datasetId":6006872},{"sourceId":9806342,"sourceType":"datasetVersion","datasetId":6010899},{"sourceId":203900450,"sourceType":"kernelVersion"}],"dockerImageVersionId":30823,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"\n## **1. Importing Required Libraries**\n\nThis section imports the necessary libraries used in the code.\n\n```python\nimport numpy as np\nimport pandas as pd\nimport torch\nfrom torch.utils.data import DataLoader, TensorDataset\nimport gc\nfrom sklearn.metrics import r2_score\nfrom tqdm import tqdm\n```\n\n- **NumPy (`np`)**: Fundamental package for numerical computations in Python.\n- **Pandas (`pd`)**: Library for data manipulation and analysis.\n- **Torch (`torch`)**: PyTorch is an open-source machine learning library.\n- **DataLoader and TensorDataset**: Utilities from PyTorch for loading data.\n- **Garbage Collector (`gc`)**: Module to manage memory in Python.\n- **R2 Score**: A metric from scikit-learn for evaluating model performance.\n- **TQDM**: A library for showing progress bars.\n\n## **2. Preprocessing Data**\n\nThis function prepares the input data by filling missing values and converting it to a PyTorch tensor.\n\n```python\ndef preprocess_data(data: pd.DataFrame, feature_cols: list) -> torch.FloatTensor:\n    \"\"\"Prepare input data by filling missing values.\"\"\"\n    data = data[feature_cols].fillna(method='ffill').fillna(0)\n    return torch.FloatTensor(data.values)\n```\n\n- **Input**: A Pandas DataFrame (`data`) and a list of feature column names (`feature_cols`).\n- **Process**: \n  - Fill missing values using forward fill (`ffill`) method.\n  - Fill any remaining missing values with `0`.\n  - Convert the DataFrame to a PyTorch FloatTensor.\n- **Output**: A PyTorch FloatTensor containing the processed data.\n\n## **3. Predicting with Multiple Models**\n\nThis function makes predictions using multiple PyTorch models and averages their outputs.\n\n```python\ndef predict_with_models(models, input_data: torch.FloatTensor) -> np.ndarray:\n    \"\"\"Predict using multiple PyTorch models.\"\"\"\n    predictions = np.zeros((input_data.shape[0],))\n    with torch.no_grad():\n        for model in models:\n            model.eval()\n            predictions += model(input_data).cpu().numpy() / len(models)\n    return predictions\n```\n\n- **Input**: \n  - A list of PyTorch models (`models`).\n  - Input data in the form of a PyTorch FloatTensor (`input_data`).\n- **Process**:\n  - Initialize an array to store predictions.\n  - Iterate over each model:\n    - Set the model to evaluation mode using `model.eval()`.\n    - Make predictions with the model.\n    - Convert predictions to a NumPy array and accumulate them.\n  - Average the accumulated predictions by dividing by the number of models.\n- **Output**: A NumPy array containing the averaged predictions.\n\n## **4. Clearing Memory**\n\nThis function clears variables from memory and forces garbage collection to free up space.\n\n```python\ndef clear_memory(*variables):\n    \"\"\"Clear variables from memory and force garbage collection.\"\"\"\n    for var in variables:\n        del var\n    gc.collect()\n```\n\n- **Input**: A list of variables to be cleared from memory.\n- **Process**:\n  - Iterate over the provided variables and delete them using the `del` keyword.\n  - Force garbage collection using `gc.collect()` to free up memory.\n- **Output**: This function does not return anything; it clears memory to optimize performance.\n","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport torch\nfrom torch.utils.data import DataLoader, TensorDataset\nimport gc\nfrom sklearn.metrics import r2_score\nfrom tqdm import tqdm\n\ndef preprocess_data(data: pd.DataFrame, feature_cols: list) -> torch.FloatTensor:\n    \"\"\"Prepare input data by filling missing values.\"\"\"\n    data = data[feature_cols].fillna(method='ffill').fillna(0)\n    return torch.FloatTensor(data.values)\n\ndef predict_with_models(models, input_data: torch.FloatTensor) -> np.ndarray:\n    \"\"\"Predict using multiple PyTorch models.\"\"\"\n    predictions = np.zeros((input_data.shape[0],))\n    with torch.no_grad():\n        for model in models:\n            model.eval()\n            predictions += model(input_data).cpu().numpy() / len(models)\n    return predictions\n\ndef clear_memory(*variables):\n    \"\"\"Clear variables from memory and force garbage collection.\"\"\"\n    for var in variables:\n        del var\n    gc.collect()","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-12-23T12:13:14.544720Z","iopub.execute_input":"2024-12-23T12:13:14.545049Z","iopub.status.idle":"2024-12-23T12:13:18.136647Z","shell.execute_reply.started":"2024-12-23T12:13:14.545017Z","shell.execute_reply":"2024-12-23T12:13:18.135737Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Data Manipulation and Visualization Libraries\n\n1. **Pandas (`pd`)**:\n    - `import pandas as pd`\n    - Used for data manipulation and analysis.\n2. **Polars (`pl`)**:\n    - `import polars as pl`\n    - An efficient DataFrame library for data manipulation.\n3. **NumPy (`np`)**:\n    - `import numpy as np`\n    - Fundamental package for numerical computations.\n4. **OS and Garbage Collection (`os`, `gc`)**:\n    - `import os, gc`\n    - Used for file handling and memory management.\n5. **TQDM**:\n    - `from tqdm.auto import tqdm`\n    - Used to display progress bars.\n6. **Matplotlib**:\n    - `from matplotlib import pyplot as plt`\n    - For creating static, animated, and interactive visualizations.\n7. **Pickle**:\n    - `import pickle`\n    - Used for serializing and deserializing Python objects.\n\n### PyTorch Libraries\n\n1. **PyTorch**:\n    - `import torch`\n    - `import torch.nn as nn`\n    - `import torch.nn.functional as F`\n    - Used for building and training neural networks.\n2. **PyTorch Lightning**:\n    - `from pytorch_lightning import (LightningDataModule, LightningModule, Trainer)`\n    - Simplifies training code by handling boilerplate.\n    - `from pytorch_lightning.callbacks import EarlyStopping, ModelCheckpoint, Timer`\n    - Provides callbacks for managing training processes.\n\n### Data Handling and Metrics\n\n1. **Scikit-learn**:\n    - `import pandas as pd`\n    - `import numpy as np`\n    - `from sklearn.metrics import r2_score`\n    - `from sklearn.model_selection import train_test_split`\n    - Used for data manipulation, evaluation metrics, and splitting datasets.\n2. **Torch Utilities**:\n    - `from torch.utils.data import Dataset, DataLoader`\n    - Handles datasets and data loading.\n\n### Machine Learning Models\n\n1. **LightGBM**:\n    - `from lightgbm import LGBMRegressor`\n    - `import lightgbm as lgb`\n    - Gradient boosting framework that uses tree-based learning algorithms.\n2. **XGBoost**:\n    - `from xgboost import XGBRegressor`\n    - An optimized distributed gradient boosting library.\n3. **CatBoost**:\n    - `from catboost import CatBoostRegressor`\n    - A high-performance open-source library for gradient boosting.\n4. **Voting Regressor**:\n    - `from sklearn.ensemble import VotingRegressor`\n    - Combines multiple models for improved prediction performance.\n\n### Warnings and Display Settings\n\n1. **Warnings**:\n    - `import warnings`\n    - Manages warning messages.\n    - `warnings.filterwarnings('ignore')`\n    - Suppresses warnings to keep the output clean.\n2. **Pandas Display Options**:\n    - `pd.options.display.max_columns = None`\n    - Ensures all columns are displayed when printing DataFrames.\n\n### Kaggle Evaluation\n\n1. **Kaggle Evaluation**:\n    - `import kaggle_evaluation.jane_street_inference_server`\n    - Used for evaluating models in a Kaggle competition.\n\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport polars as pl\nimport numpy as np\nimport os, gc\nfrom tqdm.auto import tqdm\nfrom matplotlib import pyplot as plt\nimport pickle\n\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom pytorch_lightning import (LightningDataModule, LightningModule, Trainer)\nfrom pytorch_lightning.callbacks import EarlyStopping, ModelCheckpoint, Timer\n\nimport pandas as pd\nimport numpy as np\nfrom sklearn.metrics import r2_score\nfrom sklearn.model_selection import train_test_split\nfrom torch.utils.data import Dataset, DataLoader\n\n\nfrom sklearn.metrics import r2_score\nfrom lightgbm import LGBMRegressor\nimport lightgbm as lgb\nfrom xgboost import XGBRegressor\nfrom catboost import CatBoostRegressor\nfrom sklearn.ensemble import VotingRegressor\n\nimport warnings\nwarnings.filterwarnings('ignore')\npd.options.display.max_columns = None\n\nimport kaggle_evaluation.jane_street_inference_server","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T12:13:18.137628Z","iopub.execute_input":"2024-12-23T12:13:18.138067Z","iopub.status.idle":"2024-12-23T12:13:24.795845Z","shell.execute_reply.started":"2024-12-23T12:13:18.138043Z","shell.execute_reply":"2024-12-23T12:13:24.795162Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n## **5. Configuration Class**\n\nThis class holds various configuration settings used throughout the script.\n\n```python\nclass CONFIG:\n    seed = 42\n    target_col = \"responder_6\"\n    feature_cols = [f\"feature_{idx:02d}\" for idx in range(79)] + [f\"responder_{idx}_lag_1\" for idx in range(9)]\n    \n    model_paths = [\n        \"/kaggle/input/js-xs-nn-trained-model\",\n        \"/kaggle/input/js-with-lags-trained-xgb/result.pkl\",\n    ]\n```\n\n- **`seed`**: The random seed for reproducibility.\n- **`target_col`**: The target column name for predictions.\n- **`feature_cols`**: A list of feature columns, including 79 features and 9 lagged responder features.\n- **`model_paths`**: Paths to pre-trained models.\n\n## **6. Load Validation Data**\n\nThis part loads validation data from a Parquet file into a Pandas DataFrame.\n\n```python\nvalid = pl.scan_parquet(\n    f\"/kaggle/input/js24-preprocessing-create-lags/validation.parquet/\"\n).collect().to_pandas()\n```\n\n- **`pl.scan_parquet`**: Reads the Parquet file using Polars.\n- **`.collect().to_pandas()`**: Converts the Polars DataFrame to a Pandas DataFrame.\n\n## **7. Load Pre-trained Model**\n\nThis section loads a pre-trained XGBoost model from a pickle file.\n\n```python\nxgb_model = None\nmodel_path = CONFIG.model_paths[1]\nwith open(model_path, \"rb\") as fp:\n    result = pickle.load(fp)\n    xgb_model = result[\"model\"]\n```\n\n- **`model_path`**: Path to the pre-trained model.\n- **`pickle.load`**: Loads the model from the file.\n- **`xgb_model`**: The loaded XGBoost model.\n\n## **8. Define Feature Columns for XGBoost Model**\n\nThis line defines the feature columns to be used for the XGBoost model.\n\n```python\nxgb_feature_cols = [\"symbol_id\", \"time_id\"] + CONFIG.feature_cols\n```\n\n- **`xgb_feature_cols`**: List of feature columns including `symbol_id`, `time_id`, and features specified in the `CONFIG` class.\n\n## **9. Display Model**\n\nThis part displays the loaded model, typically for verification purposes.\n\n```python\ndisplay(xgb_model)\n```\n\n- **`display(xgb_model)`**: Displays the XGBoost model's details (works in a Jupyter Notebook environment).\n","metadata":{}},{"cell_type":"code","source":"class CONFIG:\n    seed = 42\n    target_col = \"responder_6\"\n    feature_cols = [f\"feature_{idx:02d}\" for idx in range(79)]+ [f\"responder_{idx}_lag_1\" for idx in range(9)]\n    \n    model_paths = [\n        \"/kaggle/input/js-xs-nn-trained-model\",\n        \"/kaggle/input/js-with-lags-trained-xgb/result.pkl\",\n    ]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T12:13:24.797301Z","iopub.execute_input":"2024-12-23T12:13:24.797959Z","iopub.status.idle":"2024-12-23T12:13:24.802369Z","shell.execute_reply.started":"2024-12-23T12:13:24.797933Z","shell.execute_reply":"2024-12-23T12:13:24.801436Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"valid = pl.scan_parquet(\n    f\"/kaggle/input/js24-preprocessing-create-lags/validation.parquet/\"\n).collect().to_pandas()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T12:13:24.803690Z","iopub.execute_input":"2024-12-23T12:13:24.803995Z","iopub.status.idle":"2024-12-23T12:13:27.260477Z","shell.execute_reply.started":"2024-12-23T12:13:24.803973Z","shell.execute_reply":"2024-12-23T12:13:27.259408Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"xgb_model = None\nmodel_path = CONFIG.model_paths[1]\nwith open( model_path, \"rb\") as fp:\n    result = pickle.load(fp)\n    xgb_model = result[\"model\"]\n\nxgb_feature_cols = [\"symbol_id\", \"time_id\"] + CONFIG.feature_cols\n\n# Show model\ndisplay(xgb_model)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T12:13:27.261500Z","iopub.execute_input":"2024-12-23T12:13:27.261746Z","iopub.status.idle":"2024-12-23T12:13:27.368993Z","shell.execute_reply.started":"2024-12-23T12:13:27.261727Z","shell.execute_reply":"2024-12-23T12:13:27.368215Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n## **10. Custom R2 Metric for Validation**\n\nThis function calculates a custom R2 metric for validation purposes, considering sample weights.\n\n```python\ndef r2_val(y_true, y_pred, sample_weight):\n    r2 = 1 - np.average((y_pred - y_true) ** 2, weights=sample_weight) / (np.average((y_true) ** 2, weights=sample_weight) + 1e-41)\n    return r2\n```\n\n- **Input**:\n  - `y_true`: Actual values.\n  - `y_pred`: Predicted values.\n  - `sample_weight`: Weights for each sample.\n- **Process**:\n  - Calculate the weighted average of squared differences between actual and predicted values.\n  - Calculate the weighted average of squared actual values.\n  - Compute the R2 value using the formula provided.\n- **Output**: The R2 value.\n\n## **11. Neural Network Class (NN)**\n\nThis class defines a neural network using PyTorch Lightning, including initialization, forward pass, training step, validation step, and optimizer configuration.\n\n##### **2.1 Initialization**\n\nThe `__init__` method initializes the neural network with given hyperparameters and layers.\n\n```python\nclass NN(LightningModule):\n    def __init__(self, input_dim, hidden_dims, dropouts, lr, weight_decay):\n        super().__init__()\n        self.save_hyperparameters()\n        layers = []\n        in_dim = input_dim\n        for i, hidden_dim in enumerate(hidden_dims):\n            layers.append(nn.BatchNorm1d(in_dim))\n            if i > 0:\n                layers.append(nn.SiLU())\n            if i < len(dropouts):\n                layers.append(nn.Dropout(dropouts[i]))\n            layers.append(nn.Linear(in_dim, hidden_dim))\n            in_dim = hidden_dim\n        layers.append(nn.Linear(in_dim, 1))  # Output layer\n        layers.append(nn.Tanh())\n        self.model = nn.Sequential(*layers)\n        self.lr = lr\n        self.weight_decay = weight_decay\n        self.validation_step_outputs = []\n```\n\n- **Input Parameters**:\n  - `input_dim`: Dimension of input features.\n  - `hidden_dims`: List of hidden layer dimensions.\n  - `dropouts`: List of dropout rates.\n  - `lr`: Learning rate.\n  - `weight_decay`: Weight decay for regularization.\n\n##### **2.2 Forward Pass**\n\nThe `forward` method defines the forward pass of the neural network.\n\n```python\ndef forward(self, x):\n    return 5 * self.model(x).squeeze(-1)  # Output as a 1D tensor\n```\n\n- **Input**: Input tensor `x`.\n- **Output**: Scaled output tensor.\n\n##### **2.3 Training Step**\n\nThe `training_step` method defines the training step.\n\n```python\ndef training_step(self, batch):\n    x, y, w = batch\n    y_hat = self(x)\n    loss = F.mse_loss(y_hat, y, reduction='none') * w  # Consider sample weights\n    loss = loss.mean()\n    self.log('train_loss', loss, on_step=False, on_epoch=True, batch_size=x.size(0))\n    return loss\n```\n\n- **Input**: Batch of data (`x`, `y`, `w`).\n- **Output**: Loss value.\n\n##### **2.4 Validation Step**\n\nThe `validation_step` method defines the validation step.\n\n```python\ndef validation_step(self, batch):\n    x, y, w = batch\n    y_hat = self(x)\n    loss = F.mse_loss(y_hat, y, reduction='none') * w\n    loss = loss.mean()\n    self.log('val_loss', loss, on_step=False, on_epoch=True, batch_size=x.size(0))\n    self.validation_step_outputs.append((y_hat, y, w))\n    return loss\n```\n\n- **Input**: Batch of data (`x`, `y`, `w`).\n- **Output**: Loss value.\n\n##### **2.5 Validation Epoch End**\n\nThe `on_validation_epoch_end` method calculates the validation R2 metric at the end of an epoch.\n\n```python\ndef on_validation_epoch_end(self):\n    y = torch.cat([x[1] for x in self.validation_step_outputs]).cpu().numpy()\n    if self.trainer.sanity_checking:\n        prob = torch.cat([x[0] for x in self.validation_step_outputs]).cpu().numpy()\n    else:\n        prob = torch.cat([x[0] for x in self.validation_step_outputs]).cpu().numpy()\n        weights = torch.cat([x[2] for x in self.validation_step_outputs]).cpu().numpy()\n        # r2_val\n        val_r_square = r2_val(y, prob, weights)\n        self.log(\"val_r_square\", val_r_square, prog_bar=True, on_step=False, on_epoch=True)\n    self.validation_step_outputs.clear()\n```\n\n- **Process**:\n  - Concatenate the validation outputs.\n  - Calculate the validation R2 metric using `r2_val`.\n  - Log the validation R2 metric.\n\n##### **2.6 Optimizer Configuration**\n\nThe `configure_optimizers` method sets up the optimizer and learning rate scheduler.\n\n```python\ndef configure_optimizers(self):\n    optimizer = torch.optim.Adam(self.parameters(), lr=self.lr, weight_decay=self.weight_decay)\n    scheduler = torch.optim.lr_scheduler.ReduceLROnPlateau(optimizer, mode='min', factor=0.5, patience=5, verbose=True)\n    return {\n        'optimizer': optimizer,\n        'lr_scheduler': {\n            'scheduler': scheduler,\n            'monitor': 'val_loss',\n        }\n    }\n```\n\n- **Optimizer**: Adam optimizer.\n- **Scheduler**: ReduceLROnPlateau scheduler.\n\n##### **2.7 Training Epoch End**\n\nThe `on_train_epoch_end` method logs the metrics at the end of each training epoch.\n\n```python\ndef on_train_epoch_end(self):\n    if self.trainer.sanity_checking:\n        return\n    epoch = self.trainer.current_epoch\n    metrics = {k: v.item() if isinstance(v, torch.Tensor) else v for k, v in self.trainer.logged_metrics.items()}\n    formatted_metrics = {k: f\"{v:.5f}\" for k, v in metrics.items()}\n    print(f\"Epoch {epoch}: {formatted_metrics}\")\n```\n\n- **Process**:\n  - Retrieve and format logged metrics.\n  - Print metrics for the epoch.\n","metadata":{}},{"cell_type":"code","source":"# Custom R2 metric for validation\ndef r2_val(y_true, y_pred, sample_weight):\n    r2 = 1 - np.average((y_pred - y_true) ** 2, weights=sample_weight) / (np.average((y_true) ** 2, weights=sample_weight) + 1e-41)\n    return r2\n\n\nclass NN(LightningModule):\n    def __init__(self, input_dim, hidden_dims, dropouts, lr, weight_decay):\n        super().__init__()\n        self.save_hyperparameters()\n        layers = []\n        in_dim = input_dim\n        for i, hidden_dim in enumerate(hidden_dims):\n            layers.append(nn.BatchNorm1d(in_dim))\n            if i > 0:\n                layers.append(nn.SiLU())\n            if i < len(dropouts):\n                layers.append(nn.Dropout(dropouts[i]))\n            layers.append(nn.Linear(in_dim, hidden_dim))\n            # layers.append(nn.ReLU())\n            in_dim = hidden_dim\n        layers.append(nn.Linear(in_dim, 1)) \n        layers.append(nn.Tanh())\n        self.model = nn.Sequential(*layers)\n        self.lr = lr\n        self.weight_decay = weight_decay\n        self.validation_step_outputs = []\n\n    def forward(self, x):\n        return 5 * self.model(x).squeeze(-1) \n    def training_step(self, batch):\n        x, y, w = batch\n        y_hat = self(x)\n        loss = F.mse_loss(y_hat, y, reduction='none') * w  \n        loss = loss.mean()\n        self.log('train_loss', loss, on_step=False, on_epoch=True, batch_size=x.size(0))\n        return loss\n\n    def validation_step(self, batch):\n        x, y, w = batch\n        y_hat = self(x)\n        loss = F.mse_loss(y_hat, y, reduction='none') * w\n        loss = loss.mean()\n        self.log('val_loss', loss, on_step=False, on_epoch=True, batch_size=x.size(0))\n        self.validation_step_outputs.append((y_hat, y, w))\n        return loss\n\n    def on_validation_epoch_end(self):\n        \"\"\"Calculate validation WRMSE at the end of the epoch.\"\"\"\n        y = torch.cat([x[1] for x in self.validation_step_outputs]).cpu().numpy()\n        if self.trainer.sanity_checking:\n            prob = torch.cat([x[0] for x in self.validation_step_outputs]).cpu().numpy()\n        else:\n            prob = torch.cat([x[0] for x in self.validation_step_outputs]).cpu().numpy()\n            weights = torch.cat([x[2] for x in self.validation_step_outputs]).cpu().numpy()\n            # r2_val\n            val_r_square = r2_val(y, prob, weights)\n            self.log(\"val_r_square\", val_r_square, prog_bar=True, on_step=False, on_epoch=True)\n        self.validation_step_outputs.clear()\n\n    def configure_optimizers(self):\n        optimizer = torch.optim.Adam(self.parameters(), lr=self.lr, weight_decay=self.weight_decay)\n        scheduler = torch.optim.lr_scheduler.ReduceLROnPlateau(optimizer, mode='min', factor=0.5, patience=5,\n                                                               verbose=True)\n        return {\n            'optimizer': optimizer,\n            'lr_scheduler': {\n                'scheduler': scheduler,\n                'monitor': 'val_loss',\n            }\n        }\n\n    def on_train_epoch_end(self):\n        if self.trainer.sanity_checking:\n            return\n        epoch = self.trainer.current_epoch\n        metrics = {k: v.item() if isinstance(v, torch.Tensor) else v for k, v in self.trainer.logged_metrics.items()}\n        formatted_metrics = {k: f\"{v:.5f}\" for k, v in metrics.items()}\n        print(f\"Epoch {epoch}: {formatted_metrics}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T12:13:27.369673Z","iopub.execute_input":"2024-12-23T12:13:27.369901Z","iopub.status.idle":"2024-12-23T12:13:27.383670Z","shell.execute_reply.started":"2024-12-23T12:13:27.369881Z","shell.execute_reply":"2024-12-23T12:13:27.382760Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n## **11. Defining Number of Folds**\n\nThis section sets the number of folds for cross-validation.\n\n```python\nN_folds = 5\n```\n\n- **`N_folds`**: Number of folds for cross-validation (5 in this case).\n\n## **12. Loading Pre-trained Models**\n\nThis loop loads the pre-trained models from checkpoint files and moves them to the GPU.\n\n```python\nmodels = []\nfor fold in range(N_folds):\n    checkpoint_path = f\"{CONFIG.model_paths[0]}/nn_{fold}.model\"\n    model = NN.load_from_checkpoint(checkpoint_path)\n    models.append(model.to(\"cuda:0\"))\n```\n\n- **`models`**: A list to store the loaded models.\n- **Loop**:\n  - **`fold`**: Iterates over the number of folds.\n  - **`checkpoint_path`**: Path to the model checkpoint file for each fold.\n  - **`NN.load_from_checkpoint`**: Loads the model from the checkpoint file.\n  - **`model.to(\"cuda:0\")`**: Moves the model to the GPU (CUDA device).\n  - **`models.append`**: Adds the loaded model to the `models` list.\n\n## **13. Preparing Validation Data**\n\nThis section preprocesses the validation data and prepares the target and weight arrays.\n\n```python\nX_valid = preprocess_data(valid, CONFIG.feature_cols)\ny_valid = valid[CONFIG.target_col].values\nw_valid = valid[\"weight\"].values\n\nprint(\"Validation data prepared:\", X_valid.shape, y_valid.shape, w_valid.shape)\n```\n\n- **`X_valid`**: Preprocessed validation input features.\n- **`y_valid`**: Validation target values.\n- **`w_valid`**: Sample weights for the validation data.\n- **`print`**: Outputs the shapes of the prepared validation data.\n\n## **14. Making Predictions and Calculating Validation Score**\n\nThis section makes predictions using the loaded models and calculates the R2 score for the validation data.\n\n```python\ny_pred_valid_nn = predict_with_models(models, X_valid.to(\"cuda:0\"))\nvalid_score = r2_score(y_valid, y_pred_valid_nn, sample_weight=w_valid)\nprint(f\"Validation R2 score (NN): {valid_score:.5f}\")\n```\n\n- **`y_pred_valid_nn`**: Predictions for the validation data.\n- **`predict_with_models`**: Function that makes predictions using the loaded models.\n- **`X_valid.to(\"cuda:0\")`**: Moves the validation data to the GPU.\n- **`r2_score`**: Calculates the R2 score for the validation predictions.\n- **`print`**: Outputs the validation R2 score.\n","metadata":{}},{"cell_type":"code","source":"N_folds = 5\n\nmodels = []\nfor fold in range(N_folds):\n    checkpoint_path = f\"{CONFIG.model_paths[0]}/nn_{fold}.model\"\n    model = NN.load_from_checkpoint(checkpoint_path)\n    models.append(model.to(\"cuda:0\"))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T12:13:27.384464Z","iopub.execute_input":"2024-12-23T12:13:27.384666Z","iopub.status.idle":"2024-12-23T12:13:28.320850Z","shell.execute_reply.started":"2024-12-23T12:13:27.384648Z","shell.execute_reply":"2024-12-23T12:13:28.319945Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X_valid = preprocess_data(valid, CONFIG.feature_cols)\ny_valid = valid[CONFIG.target_col].values\nw_valid = valid[\"weight\"].values\n\nprint(\"Validation data prepared:\", X_valid.shape, y_valid.shape, w_valid.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T12:13:28.323134Z","iopub.execute_input":"2024-12-23T12:13:28.323378Z","iopub.status.idle":"2024-12-23T12:13:31.401518Z","shell.execute_reply.started":"2024-12-23T12:13:28.323358Z","shell.execute_reply":"2024-12-23T12:13:31.400713Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"y_pred_valid_nn = predict_with_models(models, X_valid.to(\"cuda:0\"))\nvalid_score = r2_score(y_valid, y_pred_valid_nn, sample_weight=w_valid)\nprint(f\"Validation R2 score (NN): {valid_score:.5f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T12:13:31.402467Z","iopub.execute_input":"2024-12-23T12:13:31.402767Z","iopub.status.idle":"2024-12-23T12:13:33.259972Z","shell.execute_reply.started":"2024-12-23T12:13:31.402743Z","shell.execute_reply":"2024-12-23T12:13:33.259008Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n## **15. Global Lags DataFrame**\n\nThis section initializes a global variable for lags DataFrame.\n\n```python\nlags_ : pl.DataFrame | None = None\n```\n\n- **`lags_`**: A global variable to store the lags DataFrame, initially set to `None`.\n\n## **16. Predict Function**\n\nThis function makes predictions based on the input test DataFrame and optional lags DataFrame.\n\n```python\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n    global lags_\n    if lags is not None:\n        lags_ = lags\n```\n\n- **Input**: \n  - `test`: A Polars DataFrame containing test data.\n  - `lags`: An optional Polars DataFrame containing lag data.\n- **Process**:\n  - Updates the global `lags_` variable if `lags` is provided.\n\n##### **2.1 Initial Predictions DataFrame**\n\nCreates an initial predictions DataFrame with a constant value for 'responder_6'.\n\n```python\npredictions = test.select(\n    'row_id',\n    pl.lit(0.0).alias('responder_6'),\n)\nsymbol_ids = test.select('symbol_id').to_numpy()[:, 0]\n```\n\n- **`predictions`**: DataFrame initialized with 'row_id' and 'responder_6' set to `0.0`.\n\n##### **2.2 Handling Lags Data**\n\nJoins the test DataFrame with the lags DataFrame or initializes lag columns if lags are not provided.\n\n```python\nif not lags is None:\n    lags = lags.group_by([\"date_id\", \"symbol_id\"], maintain_order=True).last() # pick up last record of previous date\n    test = test.join(lags, on=[\"date_id\", \"symbol_id\"], how=\"left\")\nelse:\n    test = test.with_columns(\n        (pl.lit(0.0).alias(f'responder_{idx}_lag_1') for idx in range(9))\n    )\n```\n\n- **`lags.group_by(...).last()`**: Groups and selects the last record of the previous date.\n- **`test.join(...)`**: Joins the test DataFrame with the lags DataFrame.\n- **`test.with_columns(...)`**: Initializes lag columns with `0.0` if lags are not provided.\n\n##### **2.3 Making Predictions**\n\nMakes predictions using the XGBoost model and neural network models, and combines their outputs.\n\n```python\npreds = np.zeros((test.shape[0],))\npreds += xgb_model.predict(test[xgb_feature_cols].to_pandas()) / 2\ntest_input = test[CONFIG.feature_cols].to_pandas()\ntest_input = test_input.fillna(method='ffill').fillna(0)\ntest_input = torch.FloatTensor(test_input.values).to(\"cuda:0\")\nwith torch.no_grad():\n    for i, nn_model in enumerate(tqdm(models)):\n        nn_model.eval()\n        preds += nn_model(test_input).cpu().numpy() / 10\nprint(f\"predict> preds.shape =\", preds.shape)\n```\n\n- **`preds`**: Initializes predictions array.\n- **`xgb_model.predict(...)`**: Makes predictions using the XGBoost model.\n- **`test_input.fillna(...)`**: Fills missing values in the test input.\n- **`torch.FloatTensor(test_input.values).to(\"cuda:0\")`**: Converts test input to a PyTorch tensor and moves it to the GPU.\n- **`nn_model.eval()`**: Sets the neural network model to evaluation mode.\n- **`nn_model(test_input).cpu().numpy()`**: Makes predictions with the neural network models and combines them.\n\n##### **2.4 Creating Predictions DataFrame**\n\nCreates the final predictions DataFrame and performs validation checks.\n\n```python\npredictions = \\\ntest.select('row_id').\\\nwith_columns(\n    pl.Series(\n        name='responder_6', \n        values=np.clip(preds, a_min=-5, a_max=5),\n        dtype=pl.Float64,\n    )\n)\n\n# The predict function must return a DataFrame\nassert isinstance(predictions, pl.DataFrame | pd.DataFrame)\n# with columns 'row_id', 'responer_6'\nassert list(predictions.columns) == ['row_id', 'responder_6']\n# and as many rows as the test data.\nassert len(predictions) == len(test)\n\nreturn predictions\n```\n\n- **`predictions`**: DataFrame with clipped prediction values.\n- **`assert`**: Validates the predictions DataFrame structure and length.\n\n#### **3. Inference Server**\n\nThis section sets up the inference server for the Kaggle competition.\n\n```python\ninference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)\n\nif os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n    inference_server.serve()\nelse:\n    inference_server.run_local_gateway(\n        (\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet',\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/lags.parquet',\n        )\n    )\n```\n\n- **`JSInferenceServer`**: Initializes the inference server with the `predict` function.\n- **`os.getenv('KAGGLE_IS_COMPETITION_RERUN')`**: Checks if the environment is a Kaggle competition rerun.\n- **`inference_server.serve()`**: Serves the inference server.\n- **`inference_server.run_local_gateway(...)`**: Runs the local gateway with specified file paths.\n\n## Summary\n\nThis code snippet defines a prediction function that handles test data, optional lags data, makes predictions using both XGBoost and neural network models, and sets up an inference server. The final predictions are validated and returned as a DataFrame. \n","metadata":{}},{"cell_type":"code","source":"lags_ : pl.DataFrame | None = None\n    \ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n    global lags_\n    if lags is not None:\n        lags_ = lags\n\n    predictions = test.select(\n        'row_id',\n        pl.lit(0.0).alias('responder_6'),\n    )\n    symbol_ids = test.select('symbol_id').to_numpy()[:, 0]\n\n    if not lags is None:\n        lags = lags.group_by([\"date_id\", \"symbol_id\"], maintain_order=True).last() # pick up last record of previous date\n        test = test.join(lags, on=[\"date_id\", \"symbol_id\"],  how=\"left\")\n    else:\n        test = test.with_columns(\n            ( pl.lit(0.0).alias(f'responder_{idx}_lag_1') for idx in range(9) )\n        )\n    \n    preds = np.zeros((test.shape[0],))\n    preds += xgb_model.predict(test[xgb_feature_cols].to_pandas()) / 2\n    test_input = test[CONFIG.feature_cols].to_pandas()\n    test_input = test_input.fillna(method = 'ffill').fillna(0)\n    test_input = torch.FloatTensor(test_input.values).to(\"cuda:0\")\n    with torch.no_grad():\n        for i, nn_model in enumerate(tqdm(models)):\n            nn_model.eval()\n            preds += nn_model(test_input).cpu().numpy() / 10\n    print(f\"predict> preds.shape =\", preds.shape)\n    \n    predictions = \\\n    test.select('row_id').\\\n    with_columns(\n        pl.Series(\n            name   = 'responder_6', \n            values = np.clip(preds, a_min = -5, a_max = 5),\n            dtype  = pl.Float64,\n        )\n    )\n\n    # The predict function must return a DataFrame\n    assert isinstance(predictions, pl.DataFrame | pd.DataFrame)\n    # with columns 'row_id', 'responer_6'\n    assert list(predictions.columns) == ['row_id', 'responder_6']\n    # and as many rows as the test data.\n    assert len(predictions) == len(test)\n\n    return predictions","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T12:13:33.260911Z","iopub.execute_input":"2024-12-23T12:13:33.261250Z","iopub.status.idle":"2024-12-23T12:13:33.270912Z","shell.execute_reply.started":"2024-12-23T12:13:33.261212Z","shell.execute_reply":"2024-12-23T12:13:33.269897Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"inference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)\n\nif os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n    inference_server.serve()\nelse:\n    inference_server.run_local_gateway(\n        (\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet',\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/lags.parquet',\n        )\n    )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-23T12:16:59.174253Z","iopub.execute_input":"2024-12-23T12:16:59.174557Z","iopub.status.idle":"2024-12-23T12:16:59.356280Z","shell.execute_reply.started":"2024-12-23T12:16:59.174536Z","shell.execute_reply":"2024-12-23T12:16:59.355444Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## References & Useful notebooks:\n\nPreprocessing : https://www.kaggle.com/code/motono0223/js24-preprocessing-create-lags\n\nTraining (XGB) : https://www.kaggle.com/code/motono0223/js24-train-gbdt-model-with-lags-singlemodel\n\ntrained XGB model : https://www.kaggle.com/datasets/motono0223/js24-trained-gbdt-model\n\nTraining (NN): https://www.kaggle.com/code/voix97/jane-street-rmf-training-nn\n\ntrained NN model : https://www.kaggle.com/datasets/voix97/js-xs-nn-trained-model\n\nInference of NN : https://www.kaggle.com/code/voix97/jane-street-rmf-nn-with-pytorch-lightning\n\nInference of NN+XGB: this notebook https://www.kaggle.com/code/voix97/jane-street-rmf-nn-xgb\n\nEDA(1) : https://www.kaggle.com/code/motono0223/eda-jane-street-real-time-market-data-forecasting\n\nEDA(2) : https://www.kaggle.com/code/motono0223/eda-v2-jane-street-real-time-market-forecasting","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}