{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceType":"competition","sourceId":98450,"databundleVersionId":11749951,"isSourceIdPinned":false}],"dockerImageVersionId":31089,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Competition Overview","metadata":{}},{"cell_type":"markdown","source":"The \"Beyond Visible Spectrum: AI for Agriculture 2025\" challenge is a competition focused on leveraging hyperspectral imagery and advanced AI techniques for sustainable farming. It's organized by Manchester Metropolitan University and the Chinese Academy of Sciences and is part of the IEEE International Conference on Image Processing (ICIP) 2025.","metadata":{}},{"cell_type":"markdown","source":"The challenge has two main tasks:\n\n- **Task 1: Hyperspectral Data Analysis for Crop Disease Classification:** The core task is to develop deep learning models to classify crop diseases (specifically, wheat stripe rust) from hyperspectral images captured by UAVs. The goal is to detect subtle physiological changes in crops before they're visible to the human eye. This aims to improve early disease detection, which can reduce yield loss and economic impact for farmers.\n\n- **Task 2: Synthetic Hyperspectral Data Generation:** Participants are also encouraged to explore generative models (like GANs, VAEs, and diffusion models) to create realistic, synthetic hyperspectral datasets. This is intended to overcome the problem of data scarcity and help build more robust and generalizable AI models for agriculture.\n\nThe competition provides pre-processed datasets with wavelengths ranging from 450 to 950 nm. The evaluation for the classification task requires participants to submit a probability for a TARGET variable for each ID in the test set.\n\nPrizes are awarded to the top three participants based on a private leaderboard: £100 for 1st place, £60 for 2nd place, and £40 for 3rd place. The competition ran from April 20, 2025, to May 20, 2025.","metadata":{}},{"cell_type":"markdown","source":"**[Competition Link](https://www.kaggle.com/competitions/beyond-visible-spectrum-ai-for-agriculture-2025/overview)**","metadata":{}},{"cell_type":"markdown","source":"# Dataset Description","metadata":{}},{"cell_type":"markdown","source":"**Dataset Description**\n\nThe dataset consists of high-resolution hyperspectral imagery of wheat fields, captured via a DJI M600 Pro UAV system equipped with a S185 snapshot hyperspectral sensor. The data was collected on two dates, 3 May 2019 and 8 May 2019, to capture different stages of wheat growth and downy mildew progression.\n- Spectral Range: The sensor operates in the visible to near-infrared spectrum, from 450 to 950 nm, with a spectral resolution of 4 nm.\n- Data Resolution: Raw data includes a 1000x1000 panchromatic image and a 50x50 hyperspectral image. After pre-processing to remove noisy bands at the spectral edges, 100 spectral bands were retained for the study.\n- Spatial Resolution: Images were captured at an altitude of 60 meters, resulting in a high spatial resolution of approximately 4cm per pixel.\n- Associated Research: The dataset is referenced in the paper titled, \"A Deep Learning-Based Approach for Automated Yellow Rust Disease Detection from High-Resolution Hyperspectral UAV Images\".\n\n**Task and Submission Format**\n\nThe task is to predict the percentage of disease (from 1 to 100) present within each image patch. This is a regression task, where the model is required to output a continuous numerical value representing the disease severity.\n- Input: Hyperspectral image patches from the test dataset.\n- Output: An integer value between 1 and 100, representing the estimated percentage of diseased crop area.\n- Submission File: A two-column CSV file with the following format:\n- Id: The unique file name of the image patch.\n- label: The predicted disease percentage.","metadata":{}},{"cell_type":"markdown","source":"# Data Dictionaries","metadata":{}},{"cell_type":"markdown","source":"**`train.csv` Data Dictionary**\n\nThe `train.csv` file serves as the ground truth for the training dataset, linking each hyperspectral image patch to its corresponding disease percentage.","metadata":{}},{"cell_type":"markdown","source":"<style type=\"text/css\">\n.tg  {border-collapse:collapse;border-spacing:0;}\n.tg td{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px;\n  overflow:hidden;padding:10px 5px;word-break:normal;}\n.tg th{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px;\n  font-weight:normal;overflow:hidden;padding:10px 5px;word-break:normal;}\n.tg .tg-k9u1{border-color:inherit;color:#1B1C1D;font-size:100%;text-align:left;vertical-align:bottom}\n.tg .tg-7zrl{text-align:left;vertical-align:bottom}\n.tg .tg-0lax{text-align:left;vertical-align:top}\n</style>\n<table class=\"tg\"><thead>\n  <tr>\n    <th class=\"tg-k9u1\">Column</th>\n    <th class=\"tg-7zrl\">Data Type</th>\n    <th class=\"tg-7zrl\">Description</th>\n  </tr></thead>\n<tbody>\n  <tr>\n    <td class=\"tg-7zrl\">Id</td>\n    <td class=\"tg-7zrl\">String</td>\n    <td class=\"tg-0lax\">A unique identifier for each sample, which corresponds to a hyperspectral image file (.npy). For example, sample697.npy.</td>\n  </tr>\n  <tr>\n    <td class=\"tg-7zrl\">label</td>\n    <td class=\"tg-7zrl\">Integer</td>\n    <td class=\"tg-0lax\">The target variable, representing the percentage of disease in the image patch. The value is an integer between 1 and 100.</td>\n  </tr>\n</tbody>\n</table>","metadata":{}},{"cell_type":"markdown","source":"**Hyperspectral Image Data Dictionary (`.npy files`)**\n\nThe core of this dataset is the hyperspectral image data, which is stored in individual `.npy files`. Each file represents a single image patch.","metadata":{}},{"cell_type":"markdown","source":"<style type=\"text/css\">\n.tg  {border-collapse:collapse;border-spacing:0;}\n.tg td{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px;\n  overflow:hidden;padding:10px 5px;word-break:normal;}\n.tg th{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px;\n  font-weight:normal;overflow:hidden;padding:10px 5px;word-break:normal;}\n.tg .tg-k9u1{border-color:inherit;color:#1B1C1D;font-size:100%;text-align:left;vertical-align:bottom}\n.tg .tg-7zrl{text-align:left;vertical-align:bottom}\n.tg .tg-0lax{text-align:left;vertical-align:top}\n</style>\n<table class=\"tg\"><thead>\n  <tr>\n    <th class=\"tg-k9u1\">Field</th>\n    <th class=\"tg-7zrl\">Data Type</th>\n    <th class=\"tg-7zrl\">Description</th>\n  </tr></thead>\n<tbody>\n  <tr>\n    <td class=\"tg-7zrl\">Data Format</td>\n    <td class=\"tg-7zrl\">NumPy Array (.npy)</td>\n    <td class=\"tg-0lax\">Each file contains a NumPy array representing a hyperspectral image. These are likely 3D tensors of shape (width, height, channels).</td>\n  </tr>\n  <tr>\n    <td class=\"tg-7zrl\">Spatial Dimensions</td>\n    <td class=\"tg-7zrl\">50x50 pixels</td>\n    <td class=\"tg-0lax\">The width and height of each image patch. The spatial resolution is approximately 4 cm per pixel.</td>\n  </tr>\n  <tr>\n    <td class=\"tg-7zrl\">Spectral Bands</td>\n    <td class=\"tg-7zrl\">100 Bands</td>\n    <td class=\"tg-0lax\">The number of spectral bands for each pixel. The raw data originally had 125 bands, but the first 10 and last 14 noisy bands were removed, leaving 100 bands for analysis.</td>\n  </tr>\n  <tr>\n    <td class=\"tg-7zrl\">Spectral Range</td>\n    <td class=\"tg-7zrl\">450-950 nm</td>\n    <td class=\"tg-0lax\">The wavelength range of the spectral data, covering the visible and near-infrared spectrum.</td>\n  </tr>\n</tbody></table>","metadata":{}},{"cell_type":"markdown","source":"**Test and Submission Data Dictionary**\n\nThe test dataset contains the same hyperspectral image files as the training set, but without the labels. Participants must predict the percentage of disease for these files and format their predictions according to the following dictionary.","metadata":{}},{"cell_type":"markdown","source":"<style type=\"text/css\">\n.tg  {border-collapse:collapse;border-spacing:0;}\n.tg td{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px;\n  overflow:hidden;padding:10px 5px;word-break:normal;}\n.tg th{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px;\n  font-weight:normal;overflow:hidden;padding:10px 5px;word-break:normal;}\n.tg .tg-k9u1{border-color:inherit;color:#1B1C1D;font-size:100%;text-align:left;vertical-align:bottom}\n.tg .tg-7zrl{text-align:left;vertical-align:bottom}\n.tg .tg-0lax{text-align:left;vertical-align:top}\n</style>\n<table class=\"tg\"><thead>\n  <tr>\n    <th class=\"tg-k9u1\">Column</th>\n    <th class=\"tg-7zrl\">Data Type</th>\n    <th class=\"tg-7zrl\">Description</th>\n  </tr></thead>\n<tbody>\n  <tr>\n    <td class=\"tg-7zrl\">Id</td>\n    <td class=\"tg-7zrl\">String</td>\n    <td class=\"tg-0lax\">The unique file name of a given image patch from the test set (e.g., sample1.npy). This is used to match predictions to the correct image.</td>\n  </tr>\n  <tr>\n    <td class=\"tg-7zrl\">label</td>\n    <td class=\"tg-7zrl\">Integer or Float</td>\n    <td class=\"tg-0lax\">The predicted percentage of disease for the corresponding image patch. This should be a value between 1 and 100.</td>\n  </tr>\n</tbody>\n</table>","metadata":{}},{"cell_type":"markdown","source":"**[Dataset Link](https://www.kaggle.com/competitions/beyond-visible-spectrum-ai-for-agriculture-2025/data)**","metadata":{}},{"cell_type":"markdown","source":"# Pipeline overview","metadata":{}},{"cell_type":"markdown","source":"**1. Data Preparation and Loading**\n\nThe pipeline begins by loading and preparing the data. The core component is the HyperspectralDataset class, a custom PyTorch Dataset.\n\n- **Loading:** It reads a CSV file containing image IDs and their corresponding disease percentages, then loads the actual hyperspectral images from `.npy files`.\n- **Preprocessing:** The code handles inconsistent image dimensions and band counts. It normalizes the pixel values to a 0-1 range and permutes the tensor dimensions to be in the format required by PyTorch (C x H x W, i.e., Channels x Height x Width).\n- **Data Augmentation:** For the training set, the pipeline uses Kornia to apply real-time data augmentation. This includes random horizontal/vertical flips, affine transformations, and crops. This is a critical step to increase the dataset size and improve the model's generalization capabilities.\n- **Target Normalization:** The disease percentage labels (1-100) are normalized to a 0-1 range for stable training.\n\n**2. Model Architecture**\n\nThe central element of the pipeline is a custom-built Convolutional Neural Network (`HyperspectralCNN`). This model is designed to process the unique multi-channel hyperspectral data.\n\n- **Feature Extraction:** The model uses a series of 2D convolutional layers (`nn.Conv2d`) followed by batch normalization and ReLU activations. These layers progressively extract features from the hyperspectral image patches. Max pooling layers reduce the spatial dimensions.\n- **Attention Mechanism:** A key feature of this model is the incorporation of Channel and Spatial Attention modules.\n- **Channel Attention:** This module focuses on which spectral bands are most relevant for disease detection, effectively weighting the importance of different wavelengths.\n- **Spatial Attention:** This module helps the model identify the most informative spatial regions within an image, such as diseased areas.\n- **Regression Head:** The features extracted by the convolutional and attention layers are passed to a series of fully connected layers. The final layer outputs a single value, representing the predicted disease percentage, making this a regression model.\n\n**3. Training and Evaluation**\n\nThis part of the pipeline handles the model's learning process and performance assessment.\n\n- **Loss Function and Optimizer:** The `nn.MSELoss` (Mean Squared Error) function is used to quantify the difference between the model's predictions and the true labels. This is appropriate for the regression task. The Adam optimizer is used to minimize this loss by adjusting the model's weights.\n- **Training Loop:** The `train_model` function iteratively trains the model on the training data. For each epoch, it performs a forward pass, calculates the loss, and updates the weights using backpropagation. It also evaluates the model on a separate validation set after each epoch.\n- **Model Checkpointing:** To save the best-performing model, the pipeline monitors the validation loss and saves the model's state dictionary (`state_dict`) whenever a new minimum loss is achieved.\n- **Evaluation:** The evaluate_model function calculates the average loss on the validation set, providing a metric to track the model's performance on unseen data.\n\n**4. Prediction and Submission**\n\nThe final step of the pipeline is to generate the submission file required by the competition.\n\n- **Test Data Loading:** A `TestHyperspectralDataset` class loads the test data, which contains only the image files without labels.\n- **Prediction:** The trained model is set to evaluation mode (`model.eval()`) and used to generate predictions for each image in the test set.\n- **Post-processing:** The model's raw output (a value between 0 and 1) is scaled back to the 1-100 range and rounded to the nearest integer. The values are also clipped to ensure they fall within the required range.\n- **Submission File:** The final predictions and corresponding image IDs are compiled into a pandas DataFrame and saved as a `submission.csv` file in the specified format.","metadata":{}},{"cell_type":"markdown","source":"**Additional Information**","metadata":{}},{"cell_type":"markdown","source":"**Kornia**\n\nKornia is an open-source, differentiable computer vision library that extends the functionality of PyTorch. It provides a comprehensive collection of operators and routines for common computer vision tasks, all implemented as PyTorch modules. Its key feature is that all operations are differentiable, which means they can be seamlessly integrated into a deep learning model's computational graph. This allows for end-to-end training where image processing tasks like geometric transformations or color corrections are learned alongside the model's primary objective.\n\nKornia is particularly valuable for the following reasons:\n\n**1. Advanced Data Augmentation**\n\nKornia provides a wide range of powerful, GPU-accelerated data augmentation techniques. In the context of a machine learning project, data augmentation is a crucial strategy to combat overfitting and enhance a model's ability to generalize. It artificially increases the size and diversity of the training dataset by applying random transformations to the input data. Common augmentations include:\n\n- Geometric Transformations: Random rotation, scaling, translation, and cropping.\n- Photometric Transformations: Random adjustments to brightness, contrast, saturation, and hue.\n- Spatial Operations: Random perspective shifts, elastic deformations, or affine transformations.\n\nFor your project on hyperspectral images, Kornia's data augmentation capabilities are essential. Since hyperspectral datasets are often limited in size, using augmentation can help your model learn to recognize disease patterns regardless of variations in the image, such as changes in the field of view or lighting conditions.\n\n**2. Integration with PyTorch**\n\nBecause Kornia is built on PyTorch, its modules function exactly like any other PyTorch layer. This allows you to easily incorporate complex computer vision operations directly into your neural network architecture. For instance, you could design a custom data preprocessing pipeline that includes real-time augmentation as part of the model's forward pass. This tight integration simplifies the development and deployment of complex models.\n\n**3. Differentiable Operations**\n\nThe \"differentiable\" nature of Kornia's functions is its core strength. This means that a model can not only learn from the data but can also learn to apply the most effective transformations during training. For example, a model could be trained to apply a specific color correction that helps it better distinguish between healthy and diseased wheat, a capability that is not possible with traditional, non-differentiable augmentation libraries.\n\nIn summary, Kornia acts as a vital tool for both accelerating development and improving the performance of your deep learning model. It provides the building blocks for robust data pipelines and enables the application of advanced techniques that make your model more accurate and reliable.","metadata":{}},{"cell_type":"markdown","source":"# Setup Environment","metadata":{}},{"cell_type":"code","source":"%%capture\n!pip install kornia","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nimport torch.optim as optim\nimport torchvision.transforms as T\nfrom torch.utils.data import Dataset, DataLoader\nimport kornia.augmentation as K\nimport numpy as np\nimport pandas as pd\nfrom sklearn.model_selection import train_test_split\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom tqdm import tqdm","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.execute_input":"2025-09-03T17:25:47.128466Z","iopub.status.busy":"2025-09-03T17:25:47.128231Z","iopub.status.idle":"2025-09-03T17:25:57.824039Z","shell.execute_reply":"2025-09-03T17:25:57.823468Z","shell.execute_reply.started":"2025-09-03T17:25:47.128443Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Data Visualization","metadata":{}},{"cell_type":"markdown","source":"This section demonstrates how to handle the hyperspectral data.\n\n- A single sample file (`.npy`) is loaded using `np.load()`. These files contain 3D data cubes of shape `(50, 50, 100)`, representing the width, height, and number of spectral bands for each image patch.\n- The code then displays a single channel (`[:, :, 0]`) of this data cube as a grayscale image using Matplotlib. This step is crucial for an initial visual check to ensure the data is loaded correctly. It also highlights the key difference from a standard RGB image, which only has three channels.","metadata":{}},{"cell_type":"code","source":"sample = np.load('/kaggle/input/beyond-visible-spectrum-ai-for-agriculture-2025/ot/ot/sample1002.npy')  # (128, 128, 125)\nplt.imshow(sample[:, :, 0])\nplt.title('First channel')\nplt.colorbar()\nplt.show()","metadata":{"execution":{"iopub.execute_input":"2025-09-03T17:25:57.825097Z","iopub.status.busy":"2025-09-03T17:25:57.824791Z","iopub.status.idle":"2025-09-03T17:25:58.240121Z","shell.execute_reply":"2025-09-03T17:25:58.239492Z","shell.execute_reply.started":"2025-09-03T17:25:57.825079Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Data Loading and Data Augmentation","metadata":{}},{"cell_type":"markdown","source":"The goal of the pipeline is to predict the percentage of crop disease from a hyperspectral image. The code is built using the PyTorch framework and leverages key concepts in computer vision and deep learning, including data augmentation, attention mechanisms, and regression.","metadata":{}},{"cell_type":"markdown","source":"**Data Handling (`HyperspectralDataset Class`)**\n\nThe HyperspectralDataset class serves as the data pipeline for the model. It inherits from PyTorch's Dataset class, which is a standard practice for preparing data for training deep learning models.\n\n- **Data Abstraction:** The Dataset class allows for abstracting the complexities of data loading and preprocessing. The `__len__` method returns the total number of samples, and the `__getitem__` method loads a single sample (an image and its corresponding label) by its index.\n- **Preprocessing:** The `__getitem__` method performs crucial preprocessing steps:\n- **Normalization:** The raw image data, typically stored as high-bit integer values, is converted to a float32 tensor and normalized to a 0-1 range.\n- **Dimension Standardization:** The code standardizes the number of spectral bands to 100 and resizes the image patches to 64x64 pixels, ensuring a consistent input format for the neural network.\n- **Target Normalization:** The disease percentage labels (from 1-100) are normalized to a 0.0-1.0 range, which helps to stabilize the training process of the regression model.\n- **Data Augmentation:** The constructor uses the Kornia library to define a series of on-the-fly, differentiable data augmentations. These transformations, such as random horizontal/vertical flips, affine transformations (rotations, scaling, translation), and random cropping, are applied to the training data. This is a fundamental technique for machine learning that helps to prevent overfitting and improves the model's ability to generalize to new, unseen data by artificially increasing the size and diversity of the training set.","metadata":{}},{"cell_type":"markdown","source":"**Parameters Setup**","metadata":{}},{"cell_type":"markdown","source":"**`BANDS = 100 and NUM_BANDS = 100`:** These two variables refer to the same concept: the number of spectral bands in the hyperspectral images. Hyperspectral data captures light across a wide range of wavelengths, with each wavelength represented as a \"band\" or channel. In this case, each image has 100 spectral bands.","metadata":{}},{"cell_type":"code","source":"BANDS = 100\nNUM_BANDS = 100\n\nclass HyperspectralDataset(Dataset):\n    def __init__(self, df, base_path, patch_size=64, augment=False, normalize_target=True):\n        self.df = df\n        self.base_path = base_path\n        self.patch_size = patch_size\n        self.augment = augment\n        self.normalize_target = normalize_target  \n        self.transform = nn.Sequential(\n            K.RandomHorizontalFlip(p=0.3),     \n            K.RandomVerticalFlip(p=0.3),\n            K.RandomAffine(degrees=5, translate=(0.05, 0.05), scale=(0.95, 1.05), p=0.5),\n            K.RandomCrop((patch_size, patch_size), padding=4, p=0.5)\n        )\n    def __len__(self):\n        return len(self.df)\n\n    def __getitem__(self, idx):\n        row = self.df.iloc[idx]\n        img_path = f\"{self.base_path}/{row['id']}\"\n\n        try:\n            img = np.load(img_path)\n\n            if len(img.shape) == 2:\n                img = np.repeat(img[:, :, np.newaxis], NUM_BANDS, axis=2)\n            elif len(img.shape) == 3:\n                if img.shape[2] > NUM_BANDS:\n                    img = img[:, :, :NUM_BANDS]\n                elif img.shape[2] < NUM_BANDS:\n                    pad_width = ((0, 0), (0, 0), (0, NUM_BANDS - img.shape[2]))\n                    img = np.pad(img, pad_width, mode='constant')\n\n            img = img.astype(np.float32) / 65535.0  \n\n            img = torch.tensor(img, dtype=torch.float32).permute(2, 0, 1)\n\n            if self.augment:\n                img = self.transform(img.unsqueeze(0)).squeeze(0)\n\n            if img.shape[1] != self.patch_size or img.shape[2] != self.patch_size:\n                img = F.interpolate(img.unsqueeze(0), size=(self.patch_size, self.patch_size), mode='bilinear').squeeze(0)\n\n            label = torch.tensor(row['label'], dtype=torch.float32)\n            if self.normalize_target:\n                label = label / 100.0\n\n            return img, label\n\n        except Exception as e:\n            print(f\"Error loading {img_path}: {str(e)}\")\n            dummy_img = torch.zeros(NUM_BANDS, self.patch_size, self.patch_size)\n            dummy_label = torch.tensor(0.0, dtype=torch.float32)\n            return dummy_img, dummy_label","metadata":{"execution":{"iopub.execute_input":"2025-09-03T17:25:58.247047Z","iopub.status.busy":"2025-09-03T17:25:58.246794Z","iopub.status.idle":"2025-09-03T17:25:58.261339Z","shell.execute_reply":"2025-09-03T17:25:58.260663Z","shell.execute_reply.started":"2025-09-03T17:25:58.247028Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Model Training","metadata":{}},{"cell_type":"markdown","source":"**Setup Training Parameters**","metadata":{}},{"cell_type":"code","source":"BATCH_SIZE = 32\nEPOCHS = 50\nLEARNING_RATE = 0.001\nDEVICE = torch.device('cuda' if torch.cuda.is_available() else 'cpu')","metadata":{"execution":{"iopub.execute_input":"2025-09-03T17:25:58.262135Z","iopub.status.busy":"2025-09-03T17:25:58.261958Z","iopub.status.idle":"2025-09-03T17:25:58.356655Z","shell.execute_reply":"2025-09-03T17:25:58.355913Z","shell.execute_reply.started":"2025-09-03T17:25:58.262122Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**`BATCH_SIZE = 32`:** This is a core hyperparameter in deep learning training. It represents the number of data samples (in this case, hyperspectral image patches) that are processed together in a single forward and backward pass. Using a batch size greater than 1 allows for more efficient computation on modern hardware like GPUs.\n\n**`EPOCHS = 50`:** An epoch signifies one full pass of the entire training dataset through the neural network. The model will go through the entire dataset 50 times to learn from the data.\n\n**`LEARNING_RATE = 0.001`:** This parameter controls how much the model's parameters (weights) are adjusted during training with respect to the loss gradient. A smaller learning rate means smaller adjustments, which can lead to more stable training but may be slower.\n\n**`DEVICE = torch.device('cuda' if torch.cuda.is_available() else 'cpu')`:** This determines the hardware used for computation. If a NVIDIA GPU with CUDA support is available, the model will be trained on the GPU, which is significantly faster for deep learning tasks. Otherwise, the training will default to the CPU.","metadata":{}},{"cell_type":"markdown","source":"## Model Architecture","metadata":{}},{"cell_type":"markdown","source":"**Model Architecture (HyperspectralCNN Class)**\n\nThe HyperspectralCNN is a custom Convolutional Neural Network designed to process the unique multi-channel nature of hyperspectral data.\n\n- **Convolutional Layers:** The model uses standard 2D convolutional layers (`nn.Conv2d`) to extract spatial features from the 64x64 image patches. Batch normalization (`nn.BatchNorm2d`) is used to normalize activations within the network, which accelerates training and improves model stability.\n- **Attention Mechanisms:** The model incorporates two types of attention modules to enhance its performance. These modules are a core theoretical component of modern deep learning and allow the model to dynamically focus on the most important features.\n- **Channel Attention (`ChannelAttention`):** This mechanism helps the model determine which of the 100 spectral bands are most relevant for predicting disease severity. It calculates and applies a weighting to each channel, allowing the network to prioritize the most informative wavelengths (e.g., those related to plant stress or chlorophyll content).\n- **Spatial Attention (`SpatialAttention`):** This mechanism helps the model focus on the most critical spatial regions within the image patch, effectively highlighting areas that contain diseased tissue while downplaying less relevant background or healthy crop.\n- **Regression Head:** Unlike a classifier that predicts a category, this network is a regressor. The features extracted by the convolutional and attention layers are passed to a series of fully connected layers (`nn.Linear`) that are tasked with mapping these features to a single, continuous numerical output: the predicted disease percentage. The use of dropout (`nn.Dropout`) helps prevent overfitting in the fully connected layers.","metadata":{}},{"cell_type":"code","source":"class ChannelAttention(nn.Module):\n    def __init__(self, in_channels, reduction_ratio=16):\n        super(ChannelAttention, self).__init__()\n        self.avg_pool = nn.AdaptiveAvgPool2d(1)\n        self.max_pool = nn.AdaptiveMaxPool2d(1)\n        \n        self.fc = nn.Sequential(\n            nn.Linear(in_channels, in_channels // reduction_ratio),\n            nn.ReLU(inplace=True),\n            nn.Linear(in_channels // reduction_ratio, in_channels),\n            nn.Sigmoid()\n        )\n\n    def forward(self, x):\n        b, c, _, _ = x.size()\n        avg_out = self.fc(self.avg_pool(x).view(b, c))\n        max_out = self.fc(self.max_pool(x).view(b, c))\n        out = avg_out + max_out\n        return out.view(b, c, 1, 1)\n\n\nclass SpatialAttention(nn.Module):\n    def __init__(self, kernel_size=7):\n        super(SpatialAttention, self).__init__()\n        self.conv = nn.Sequential(\n            nn.Conv2d(2, 8, kernel_size, padding=kernel_size//2),\n            nn.ReLU(),\n            nn.Conv2d(8, 1, kernel_size=1)\n        )\n        self.sigmoid = nn.Sigmoid()\n    \n    def forward(self, x):\n        avg_out = torch.mean(x, dim=1, keepdim=True)\n        max_out, _ = torch.max(x, dim=1, keepdim=True)\n        concat = torch.cat([avg_out, max_out], dim=1)\n        attention = self.sigmoid(self.conv(concat))\n        return x * attention\n\nclass HyperspectralCNN(nn.Module):\n    def __init__(self, in_channels=NUM_BANDS):\n        super().__init__()\n        \n        self.conv1 = nn.Sequential(\n            nn.Conv2d(in_channels, 64, kernel_size=3, padding=1),\n            nn.BatchNorm2d(64),\n            nn.ReLU(),\n            nn.MaxPool2d(2)\n        )\n        \n        self.ca1 = ChannelAttention(64)\n        self.sa1 = SpatialAttention()\n        \n        self.conv2 = nn.Sequential(\n            nn.Conv2d(64, 128, kernel_size=3, padding=1),\n            nn.BatchNorm2d(128),\n            nn.ReLU(),\n            nn.MaxPool2d(2)\n        )\n        \n        self.ca2 = ChannelAttention(128)\n        self.sa2 = SpatialAttention()\n        \n        self.conv3 = nn.Sequential(\n            nn.Conv2d(128, 256, kernel_size=3, padding=1),\n            nn.BatchNorm2d(256),\n            nn.ReLU(),\n            nn.AdaptiveAvgPool2d(1)\n        )\n        \n        self.regressor = nn.Sequential(\n            nn.Linear(256, 128),\n            nn.ReLU(),\n            nn.Dropout(0.5),\n            nn.Linear(128, 64),\n            nn.ReLU(),\n            nn.Linear(64, 1)\n        )\n        \n    def forward(self, x):\n        x = self.conv1(x)\n        x = self.ca1(x) * x\n        x = self.sa1(x) * x\n        \n        x = self.conv2(x)\n        x = self.ca2(x) * x\n        x = self.sa2(x) * x\n        \n        x = self.conv3(x)\n        x = x.view(x.size(0), -1)\n        return self.regressor(x)","metadata":{"execution":{"iopub.execute_input":"2025-09-03T17:25:58.357839Z","iopub.status.busy":"2025-09-03T17:25:58.357620Z","iopub.status.idle":"2025-09-03T17:25:58.370590Z","shell.execute_reply":"2025-09-03T17:25:58.370001Z","shell.execute_reply.started":"2025-09-03T17:25:58.357822Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Model Training and Validation","metadata":{}},{"cell_type":"markdown","source":"The `train_model` and `evaluate_model` functions define the learning process and performance assessment of the model.\n\n- **Loss Function:** The `nn.MSELoss` (Mean Squared Error) is used as the loss function. This is an appropriate choice for a regression task as it measures the average squared difference between the model's predictions and the ground truth values. The goal of training is to minimize this loss.\n\n- **Optimizer:** The Adam optimizer is used to update the model's weights during training. It is an adaptive learning rate algorithm that is highly effective for a wide range of deep learning problems.\n\n- **Validation and Overfitting:** The dataset is split into training and validation sets. The validation set is used to evaluate the model's performance on data it has not seen before. This allows for monitoring for overfitting, a state where the model performs well on the training data but poorly on new data. The best model, based on the lowest validation loss, is saved for later use.\n\n- **Model Checkpointing:** By saving the model's `state_dict` at each new best validation loss, the code ensures that the final model is the one that performed optimally on the validation set, not necessarily the one from the final training epoch.","metadata":{}},{"cell_type":"markdown","source":"**Model Validation Function**","metadata":{}},{"cell_type":"code","source":"def evaluate_model(model, loader, criterion):\n    model.eval()\n    total_loss = 0.0\n    all_preds = []\n    all_labels = []\n    \n    with torch.no_grad():\n        for inputs, labels in loader:\n            inputs, labels = inputs.to(DEVICE), labels.to(DEVICE)\n            outputs = model(inputs)\n            loss = criterion(outputs.squeeze(), labels)\n            total_loss += loss.item() * inputs.size(0)\n            \n            all_preds.extend(outputs.squeeze().cpu().numpy())\n            all_labels.extend(labels.cpu().numpy())\n    \n    return total_loss / len(loader.dataset), np.array(all_preds), np.array(all_labels)","metadata":{"execution":{"iopub.execute_input":"2025-09-03T17:25:58.371595Z","iopub.status.busy":"2025-09-03T17:25:58.371332Z","iopub.status.idle":"2025-09-03T17:25:58.384111Z","shell.execute_reply":"2025-09-03T17:25:58.383613Z","shell.execute_reply.started":"2025-09-03T17:25:58.371572Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Model Training**","metadata":{}},{"cell_type":"code","source":"def train_model(model, train_loader, val_loader, epochs, criterion, optimizer):\n    best_loss = float('inf')\n    train_losses = []\n    val_losses = []\n    \n    for epoch in range(epochs):\n        model.train()\n        train_loss = 0.0\n        valid_samples = 0\n        \n        for inputs, labels in tqdm(train_loader, desc=f\"Epoch {epoch+1}/{epochs}\"):\n            inputs, labels = inputs.to(DEVICE), labels.to(DEVICE)\n            \n            if torch.isnan(inputs).any() or torch.isnan(labels).any():\n                continue\n                \n            optimizer.zero_grad()\n            outputs = model(inputs)\n            \n            if torch.isnan(outputs).any():\n                continue\n                \n            loss = criterion(outputs.squeeze(), labels) \n            \n            if not torch.isnan(loss):\n                loss.backward()\n                optimizer.step()\n                train_loss += loss.item() * inputs.size(0)\n                valid_samples += inputs.size(0)\n        \n        if valid_samples > 0:\n            train_loss /= valid_samples\n            val_loss, val_preds, val_labels = evaluate_model(model, val_loader, criterion)\n            train_losses.append(train_loss)\n            val_losses.append(val_loss)\n            \n            print(f\"Epoch {epoch+1}: Train Loss: {train_loss:.4f}, Val Loss: {val_loss:.4f}\")\n            print(f\"Sample predictions: {val_preds[:5]}, True labels: {val_labels[:5]}\")\n            \n            if val_loss < best_loss:\n                best_loss = val_loss\n                torch.save(model.state_dict(), 'Spectrum_CNN.pth')\n        else:\n            print(f\"Epoch {epoch+1}: No valid training samples\")\n    \n    plt.figure(figsize=(8, 5))\n    plt.plot(train_losses, label='Train Loss', marker='o')\n    plt.plot(val_losses, label='Validation Loss', marker='o')\n    plt.xlabel('Epoch')\n    plt.ylabel('Loss')\n    plt.title('Training and Validation Loss')\n    plt.legend()\n    plt.grid(True)\n    plt.tight_layout()\n    plt.show()\n    \n    return model","metadata":{"execution":{"iopub.execute_input":"2025-09-03T17:25:58.385116Z","iopub.status.busy":"2025-09-03T17:25:58.384876Z","iopub.status.idle":"2025-09-03T17:25:58.399287Z","shell.execute_reply":"2025-09-03T17:25:58.398626Z","shell.execute_reply.started":"2025-09-03T17:25:58.385101Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def main():\n    train_df = pd.read_csv('/kaggle/input/beyond-visible-spectrum-ai-for-agriculture-2025/train.csv')\n    base_path = '/kaggle/input/beyond-visible-spectrum-ai-for-agriculture-2025/ot/ot'\n    \n    train_df, val_df = train_test_split(train_df, test_size=0.2, random_state=42)\n    \n    train_dataset = HyperspectralDataset(train_df, base_path, augment=True)\n    val_dataset = HyperspectralDataset(val_df, base_path, augment=False)\n    \n    train_loader = DataLoader(train_dataset, batch_size=BATCH_SIZE, shuffle=True, num_workers=4)\n    val_loader = DataLoader(val_dataset, batch_size=BATCH_SIZE, shuffle=False, num_workers=4)\n    \n    model = HyperspectralCNN().to(DEVICE)\n    criterion = nn.MSELoss()\n    optimizer = optim.Adam(model.parameters(), lr=LEARNING_RATE, weight_decay=1e-5)\n    \n    model = train_model(model, train_loader, val_loader, EPOCHS, criterion, optimizer)\n    \n    model.load_state_dict(torch.load('Spectrum_CNN.pth'))\n    \n    return model\n\nmodel = main()","metadata":{"execution":{"iopub.execute_input":"2025-09-03T17:25:58.401294Z","iopub.status.busy":"2025-09-03T17:25:58.401079Z","iopub.status.idle":"2025-09-03T17:44:26.684003Z","shell.execute_reply":"2025-09-03T17:44:26.683218Z","shell.execute_reply.started":"2025-09-03T17:25:58.401279Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Model Inference and Submission Generation","metadata":{}},{"cell_type":"markdown","source":"This section of the code outlines the process of using the fully trained model to make predictions on a new, unseen test dataset and generate a submission file.","metadata":{}},{"cell_type":"markdown","source":"**Model Loading and Evaluation Mode:** \n\nThe `model.load_state_dict(torch.load('Spectrum_CNN.pth'))` command is used to load the saved, best-performing model weights from the training phase. Immediately after loading, the `model.eval()` method is called. This is a critical step in PyTorch, as it switches the model from training mode to evaluation mode. In evaluation mode, layers like Dropout are deactivated, and Batch Normalization uses its learned running statistics instead of the batch-specific statistics. This ensures consistent and deterministic predictions, which is essential for accurate evaluation.","metadata":{}},{"cell_type":"code","source":"model = HyperspectralCNN(in_channels=100).to(DEVICE)\n\nmodel.load_state_dict(torch.load('Spectrum_CNN.pth'))\nmodel.eval()\nprint(\"Model weights:\", list(model.parameters())[0][0, 0, :5])\ntest_input = torch.randn(1, 100, 64, 64).to(DEVICE)\nprint(\"Test output:\", model(test_input).item())","metadata":{"execution":{"iopub.execute_input":"2025-09-03T17:44:26.685578Z","iopub.status.busy":"2025-09-03T17:44:26.685287Z","iopub.status.idle":"2025-09-03T17:44:27.039177Z","shell.execute_reply":"2025-09-03T17:44:27.038447Z","shell.execute_reply.started":"2025-09-03T17:44:26.685550Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Model Inference Pipeline:** \n\nA new TestHyperspectralDataset class is created specifically for the test data. This class is similar to the training dataset but notably lacks data augmentation, as augmentation is only for training to improve generalization. The TestHyperspectralDataset also includes logic for normalization and dimension standardization to prepare the test data for the model in the same format as the training data.\n\n**Inference Loop:** \n\nThe code iterates through the test data using a DataLoader. The context manager (`with torch.no_grad():`) is used to disable gradient calculation during this loop. Since we are only making predictions and not updating the model's weights, this significantly reduces memory usage and speeds up the process.\n\n**Prediction Post-processing:** \n\nAfter the model generates an output, several post-processing steps are applied to convert the raw model output into a final prediction:\n\n- The output tensor is moved from the GPU to the CPU (`.cpu()`) and converted to a NumPy array (`.numpy()`).\n- The output is denormalized by multiplying by 100 to return the prediction to its original scale (1-100 percentage).\n- The `np.clip` function is used to ensure the final predictions fall within the valid range of 1 to 100, and `round().astype(int)` rounds the values to integers as required for the final submission.","metadata":{}},{"cell_type":"code","source":"class TestHyperspectralDataset(Dataset):\n    def __init__(self, test_csv, base_path, patch_size=64, num_bands=100):\n        self.df = pd.read_csv(test_csv)\n        self.base_path = base_path\n        self.patch_size = patch_size\n        self.num_bands = num_bands\n        \n    def __len__(self):\n        return len(self.df)\n    \n    def __getitem__(self, idx):\n        row = self.df.iloc[idx]\n        img_path = os.path.join(self.base_path, row['id'])\n        \n        try:\n            img = np.load(img_path)\n            \n            if len(img.shape) == 2:\n                img = np.repeat(img[:, :, np.newaxis], self.num_bands, axis=2)\n            elif len(img.shape) == 3:\n                if img.shape[2] > self.num_bands:\n                    img = img[:, :, :self.num_bands] \n                elif img.shape[2] < self.num_bands:\n                    pad_width = ((0, 0), (0, 0), (0, self.num_bands - img.shape[2]))\n                    img = np.pad(img, pad_width, mode='constant')\n            \n            normalized_img = np.zeros_like(img)\n            for band in range(img.shape[2]):\n                band_data = img[:, :, band]\n                if np.max(band_data) > 0:  \n                    normalized_img[:, :, band] = (band_data - np.min(band_data)) / (np.max(band_data) - np.min(band_data))\n            \n            img_tensor = torch.tensor(normalized_img, dtype=torch.float32).permute(2, 0, 1)\n            \n            if img_tensor.shape[1] != self.patch_size or img_tensor.shape[2] != self.patch_size:\n                img_tensor = F.interpolate(img_tensor.unsqueeze(0), \n                                         size=(self.patch_size, self.patch_size),\n                                         mode='bilinear').squeeze(0)\n            \n            return img_tensor, row['id']\n        \n        except Exception as e:\n            print(f\"Error loading {img_path}: {str(e)}\")\n            dummy_img = torch.zeros(self.num_bands, self.patch_size, self.patch_size)\n            return dummy_img, row['id']\n\n\ntest_csv_path = '/kaggle/input/beyond-visible-spectrum-ai-for-agriculture-2025/test.csv'\nbase_path = '/kaggle/input/beyond-visible-spectrum-ai-for-agriculture-2025/ot/ot'\n\ntest_dataset = TestHyperspectralDataset(test_csv_path, base_path, num_bands=100)\ntest_loader = DataLoader(test_dataset, batch_size=BATCH_SIZE, shuffle=False, num_workers=4)","metadata":{"execution":{"iopub.execute_input":"2025-09-03T17:44:27.040874Z","iopub.status.busy":"2025-09-03T17:44:27.040210Z","iopub.status.idle":"2025-09-03T17:44:27.063567Z","shell.execute_reply":"2025-09-03T17:44:27.062903Z","shell.execute_reply.started":"2025-09-03T17:44:27.040848Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Submission File Generation:** \n\nFinally, the processed predictions and corresponding image IDs are compiled into a pandas DataFrame, which is then saved as a `submission.csv` file. This is the standard format for submitting results to many machine learning competitions.","metadata":{}},{"cell_type":"code","source":"predictions = []\nids = []\n\nwith torch.no_grad():\n    for inputs, img_ids in test_loader:\n        inputs = inputs.to(DEVICE)\n        \n        if torch.isnan(inputs).any():\n            print(f\"Skipping batch with NaN values\")\n            predictions.extend([50] * len(img_ids))  \n            ids.extend(img_ids)\n            continue\n            \n        outputs = model(inputs)\n        preds = outputs.squeeze().cpu().numpy()\n        preds = preds * 100  \n\n        preds = np.clip(preds, 1, 100).round().astype(int)\n        \n        if preds.ndim == 0:  \n            preds = [preds.item()]\n        else:\n            preds = preds.tolist()\n        \n        predictions.extend(preds)\n        ids.extend(img_ids)\n\nsubmission_df = pd.DataFrame({'ID': ids, 'TARGET': predictions})\nsubmission_df.to_csv('submission.csv', index=False)\nprint(\"Submission created successfully\")\nprint(\"\\nSubmission preview:\")\nprint(submission_df.head())","metadata":{"execution":{"iopub.execute_input":"2025-09-03T17:44:27.065012Z","iopub.status.busy":"2025-09-03T17:44:27.064377Z","iopub.status.idle":"2025-09-03T17:44:45.544710Z","shell.execute_reply":"2025-09-03T17:44:45.543673Z","shell.execute_reply.started":"2025-09-03T17:44:27.064985Z"},"trusted":true},"outputs":[],"execution_count":null}]}