{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":92399,"databundleVersionId":11038207,"sourceType":"competition"}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Nexar Dashcam Crash Prediction: Solution Approach\n\n## 1. Competition Overview\n\n**Competition:** 🚗💥 Nexar Dashcam Crash Prediction Challenge (Kaggle)\n**Objective:** Develop a machine learning model to predict vehicle collisions or near-misses from dashcam video footage *before* they happen.\n\n## 2. Problem Statement\n\nRoad safety systems like Advanced Driver Assistance Systems (ADAS) and autonomous vehicles rely on accurately perceiving the driving environment and anticipating potential hazards. Predicting an imminent collision from dashcam footage allows these systems to take proactive measures (alerting the driver, initiating braking) potentially preventing accidents or mitigating their severity.\n\nThe challenge lies in:\n*   Detecting subtle pre-crash cues within complex and dynamic video sequences.\n*   Making predictions significantly *before* the event occurs to allow reaction time.\n*   Handling real-world variations like diverse weather conditions, lighting changes, occlusions, and unexpected road events.\n*   Differentiating between truly dangerous situations (collisions, near-misses) and normal driving maneuvers.\n\n## 3. Goal\n\nThe primary goal of this project is to **build and train a deep learning model that accurately predicts the likelihood of an imminent collision or near-miss event based on dashcam video input, aiming to maximize the competition's evaluation metric (mAP@TTA).**\n\nThis involves:\n*   **Accuracy:** Correctly classifying sequences leading to an event (positive class: collision or near-miss) vs. normal driving (negative class).\n*   **Timeliness:** Making the prediction as early as possible before the actual event time, ideally around or before the annotated `time_of_alert`.\n\n## 4. Data\n\n*   **Input:** Sequences of dashcam video frames (MP4 format, 1280x720 resolution, 30 FPS).\n*   **Training Data:** ~1500 videos (~40 seconds each), balanced between positive (collision/near-miss) and negative (normal driving) cases.\n    *   **Annotations:** `id`, `target` (0/1), `time_of_event` (for positives), `time_of_alert` (for positives).\n*   **Test Data:** ~1344 videos (~10 seconds each), ending at specific times before a potential event. No labels provided.\n\n## 5. Evaluation Metric\n\n*   **Mean Average Precision over different Times To Accident (mAP@TTA):** Submissions are evaluated based on the model's predicted probability score for each test video. The metric calculates Average Precision (AP) from Precision-Recall curves at three specific time-to-accident thresholds (500ms, 1000ms, 1500ms before the event). The final score is the mean of these three AP values.\n*   **Significance:** This metric explicitly rewards both **classification accuracy (Precision/Recall)** and **prediction timeliness (Anticipation Time)**. Higher scores require predicting events earlier while maintaining high precision and recall.\n\n## 6. Planned Approach & Strategies\n\nThis notebook implements an end-to-end deep learning pipeline using PyTorch to tackle the challenge. The core architecture chosen is a **Convolutional Neural Network (CNN) combined with a Long Short-Term Memory (LSTM) network**, leveraging the strengths of each for spatial and temporal feature extraction.\n\n**Pipeline Breakdown:**\n\n1.  **Environment Setup:**\n    *   Install necessary libraries (including `decord` for efficient video loading).\n    *   Configure device (prioritizing GPU if available).\n    *   Define paths and hyperparameters.\n\n2.  **Data Loading & Preprocessing:**\n    *   **Video Reading:** Use `decord` (with OpenCV as a fallback in `sample_frames_robust_stride`) for potentially faster video decoding compared to pure OpenCV.\n    *   **Frame Sampling:** Implement **frame skipping** (using `FRAME_STRIDE`) within the sampling function to reduce data redundancy and computational load, creating sequences with an effective FPS lower than the original 30.\n    *   **PyTorch Dataset:** A custom `DashcamDataset` class handles loading metadata, constructing video paths, sampling frame sequences dynamically (on-the-fly processing), and applying initial transformations.\n    *   **Transforms:** Utilize `torchvision.transforms` for:\n        *   Resizing frames to a consistent input size (`IMG_SIZE`).\n        *   Normalizing pixel values (using ImageNet statistics as a baseline).\n    *   **Data Augmentation (Training Only):** Apply augmentations to improve model robustness and reduce overfitting:\n        *   `RandomHorizontalFlip`: Applied consistently across the entire sequence within the `Dataset`.\n        *   `ColorJitter`: Applied frame-wise within the `Compose` pipeline to simulate varying lighting conditions.\n    *   **Train/Validation Split:** Perform a **stratified, time-based split** (if split files don't exist) using the original `train.csv` to ensure the validation set represents data chronologically after the training set, simulating real-world prediction. *Correction: The current split is random stratified, not time-based. A time-based split using `issue_d` would be a better approach for future iterations.*\n    *   **DataLoader:** Use PyTorch `DataLoader` with `num_workers` for parallel loading, `pin_memory` for faster GPU transfer, and a custom `collate_fn` to handle potential errors during video loading.\n\n3.  **Model Architecture (CNN+LSTM):**\n    *   **CNN Backbone:** Use a pre-trained ResNet18 (`pretrained=True`) as the feature extractor, removing its final classification layer. This leverages transfer learning for spatial feature extraction from individual frames.\n    *   **LSTM Network:** An LSTM layer processes the sequence of frame features extracted by the CNN to capture temporal dependencies and patterns indicative of pre-crash scenarios.\n    *   **Classifier Head:** Fully connected layers with ReLU activation and Dropout map the LSTM output to a single output logit for binary classification.\n\n4.  **Training Strategy:**\n    *   **Loss Function:** `nn.BCEWithLogitsLoss` (Binary Cross-Entropy with Logits) suitable for binary classification with a single logit output (combines Sigmoid + BCELoss for stability).\n    *   **Optimizer:** `optim.Adam` with `weight_decay` for L2 regularization.\n    *   **Learning Rate Scheduler:** `optim.lr_scheduler.ReduceLROnPlateau` monitors validation loss and reduces the learning rate if performance stagnates.\n    *   **Automatic Mixed Precision (AMP):** Use `torch.cuda.amp` (`autocast` and `GradScaler`) to speed up training and reduce GPU memory usage on compatible hardware.\n    *   **Checkpointing:** Save the model (including model state, optimizer state, epoch, best validation loss) whenever the validation loss improves. Load the best checkpoint to resume training or for final inference.\n    *   **Early Stopping:** Monitor validation loss and stop training if it doesn't improve for a defined `EARLY_STOPPING_PATIENCE` number of epochs, preventing excessive overfitting and saving computation time.\n\n5.  **Overfitting Mitigation:** Explicitly address overfitting (observed in previous runs) using:\n    *   Data Augmentation (Horizontal Flip, Color Jitter).\n    *   Regularization (Weight Decay in optimizer, Dropout in classifier head).\n    *   Early Stopping based on validation loss.\n\n6.  **Inference & Submission:**\n    *   Load the *best* model checkpoint saved during training (based on validation loss).\n    *   Set the model to evaluation mode (`model.eval()`).\n    *   Create a DataLoader for the `test.csv` data.\n    *   Perform inference on the test set, applying sigmoid to model outputs to get probabilities [0, 1].\n    *   Format predictions into a `submission.csv` file with `id` and `score` columns.\n\n## 7. Achieving the Goal\n\nBy executing the cells in this notebook sequentially, the pipeline will: load and prepare the data efficiently using Decord and optimized DataLoaders; train the CNN+LSTM model using AMP for speed; employ regularization, augmentation, and early stopping to find a well-generalized model state based on validation performance; load the best model; and finally generate predictions on the test set to create the submission file. The aim is to achieve an mAP@TTA score significantly better than the baseline by leveraging GPU acceleration and robust training practices.","metadata":{}},{"cell_type":"code","source":"!pip install decord","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T07:49:33.489837Z","iopub.execute_input":"2025-04-25T07:49:33.490114Z","iopub.status.idle":"2025-04-25T07:49:40.128107Z","shell.execute_reply.started":"2025-04-25T07:49:33.490094Z","shell.execute_reply":"2025-04-25T07:49:40.127141Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Cell: Imports and Configuration\nimport os\n# import cv2 # Keep if needed for other potential ops, but not core sampling/resize now\nimport pandas as pd\nimport numpy as np\nimport torch\nfrom torch.utils.data import Dataset, DataLoader # Moved DataLoader import here\nfrom torchvision import transforms\nfrom tqdm.notebook import tqdm\nfrom sklearn.model_selection import train_test_split\nimport torch.optim as optim\nimport torch.nn as nn\nimport torchvision.models as models\nimport gc # Import garbage collector\n\n# --- Decord Import ---\nimport decord\nfrom decord import VideoReader, cpu\n# Initialize Decord context (do this once globally)\ndecord.bridge.set_bridge('torch') # Let Decord return PyTorch tensors directly\n\n# --- AMP Imports ---\nfrom torch.cuda.amp import GradScaler, autocast\n\n# --- Collate Function Import ---\nfrom torch.utils.data.dataloader import default_collate\n\n\nprint(\"--- Initial Setup ---\")\n#--- Basic Configuration ---\nDEVICE = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nprint(f\"Using device: {DEVICE}\")\n\nBASE_DATA_PATH = '/kaggle/input/nexar-collision-prediction'\nTRAIN_VIDEO_DIR = os.path.join(BASE_DATA_PATH, 'train')\nTEST_VIDEO_DIR = os.path.join(BASE_DATA_PATH, 'test')\nORIGINAL_TRAIN_CSV = os.path.join(BASE_DATA_PATH, 'train.csv')\nTEST_CSV_PATH = os.path.join(BASE_DATA_PATH, 'test.csv') # Corrected variable name\nSAMPLE_SUB_CSV = os.path.join(BASE_DATA_PATH, 'sample_submission.csv')\n\n\n# --- Working Directory Paths ---\nWORKING_DIR = '/kaggle/working/'\nTRAIN_CSV_SPLIT = os.path.join(WORKING_DIR, 'train_split.csv')\nVAL_CSV_SPLIT = os.path.join(WORKING_DIR, 'val_split.csv')\nSUBMISSION_CSV = os.path.join(WORKING_DIR, 'submission.csv')\nCHECKPOINT_PATH = os.path.join(WORKING_DIR, \"best_model_cnn_lstm.pth\")\n\n\n# --- Model/Training Hyperparameters ---\n\nIMG_SIZE = 224\nSEQ_LEN = 16 # Number of frames per sample\nN_CHANNELS = 3\nNUM_CLASSES = 1    # Binary classification (crash/no-crash)\nLSTM_HIDDEN_SIZE = 512\nLSTM_LAYERS = 2\nPRETRAINED = True # Use pretrained ResNet\nLEARNING_RATE = 5e-5\nMAX_EPOCHS = 100     # Set high, let early stopping decide\nEARLY_STOPPING_PATIENCE = 5 # Example: Stop after 10 epochs with no val loss improvement\n\n# !! CRITICAL TUNING PARAMETERS for GPU !!\nBATCH_SIZE = 32     # << START HERE for GPU (e.g., 16, 32, 64) - TUNE THIS\nNUM_WORKERS = 2     # << START HERE (e.g., 2, 4) - TUNE THIS\nPREFETCH_FACTOR = 2 # For DataLoader prefetching\n\n\n# --- Set final paths for train/val CSVs (will use splits if they exist) ---\nTRAIN_CSV_PATH_FINAL = TRAIN_CSV_SPLIT\nVAL_CSV_PATH_FINAL = VAL_CSV_SPLIT\n\nprint(\"Initial configuration and paths defined.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T07:49:43.575399Z","iopub.execute_input":"2025-04-25T07:49:43.575714Z","iopub.status.idle":"2025-04-25T07:49:43.693144Z","shell.execute_reply.started":"2025-04-25T07:49:43.575692Z","shell.execute_reply":"2025-04-25T07:49:43.692390Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Initialize Decord context (do this once globally if preferred)\ndecord.bridge.set_bridge('torch') # Let Decord return PyTorch tensors directly\n\nFRAME_STRIDE = 3 # Sample every 3rd frame (effective 10 FPS) - TUNE THIS (2, 3, 4)\n\ndef sample_frames_robust_stride(video_path, seq_len, stride):\n    \"\"\" Samples seq_len frames using stride. Falls back to OpenCV. \"\"\"\n    frames_tensor = None\n    # --- Try Decord ---\n    try:\n        vr = VideoReader(video_path, ctx=cpu(0))\n        total_frames = len(vr)\n        if total_frames == 0: return None\n\n        # Calculate indices using stride\n        end_frame = total_frames - 1\n        # Start sampling potentially earlier to ensure seq_len frames are available\n        # Max possible start index: end_frame - (seq_len - 1) * stride\n        effective_len = (seq_len -1) * stride + 1\n        if effective_len > total_frames:\n            # Not enough frames even with stride=1, sample all available\n            indices = np.linspace(0, total_frames - 1, seq_len, dtype=int)\n        else:\n            # Sample a starting point randomly\n            start_max = end_frame - effective_len + 1\n            start_index = np.random.randint(0, start_max + 1)\n            indices = np.arange(start_index, start_index + effective_len, stride)\n            # Ensure we don't exceed total_frames due to rounding/edge cases\n            indices = indices[indices <= end_frame][:seq_len] # Take only up to seq_len\n\n        indices = np.clip(indices, 0, total_frames - 1).astype(int)\n\n        # Ensure we have exactly seq_len indices, pad if necessary due to sampling near end\n        if len(indices) < seq_len:\n            padding_needed = seq_len - len(indices)\n            last_idx = indices[-1] if len(indices) > 0 else total_frames - 1\n            padding_indices = np.full(padding_needed, last_idx)\n            indices = np.concatenate((indices, padding_indices)).astype(int)\n\n        frames_tensor = vr.get_batch(indices) # (T, H, W, C), torch.uint8\n\n        # Final check for shape (should match seq_len now)\n        if frames_tensor.shape[0] != seq_len:\n             print(f\"WARN: Frame count mismatch after Decord+Stride ({frames_tensor.shape[0]} vs {seq_len}) for {video_path}. Trying OpenCV.\")\n             frames_tensor = None # Force OpenCV fallback\n\n    except Exception as e_decord:\n        print(f\"Decord failed for {video_path}: {e_decord}. Trying OpenCV...\")\n        frames_tensor = None\n\n    # --- Fallback to OpenCV (with stride) ---\n    if frames_tensor is None:\n        try:\n            cap = cv2.VideoCapture(video_path)\n            if not cap.isOpened(): return None\n            total_frames_cv = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n            if total_frames_cv == 0:\n                 cap.release(); return None\n\n            frames_list = []\n            effective_len_cv = (seq_len -1) * stride + 1\n            if effective_len_cv > total_frames_cv:\n                # Fallback: Sample uniformly if stride doesn't fit\n                indices_cv = np.linspace(0, total_frames_cv - 1, seq_len, dtype=int)\n            else:\n                start_max_cv = total_frames_cv - effective_len_cv\n                start_index_cv = np.random.randint(0, start_max_cv + 1)\n                indices_cv = np.arange(start_index_cv, start_index_cv + effective_len_cv, stride)\n                indices_cv = indices_cv[indices_cv < total_frames_cv][:seq_len]\n\n            indices_cv = np.clip(indices_cv, 0, total_frames_cv - 1).astype(int)\n\n            # Read specific frames (more efficient than reading all)\n            for idx_to_read in indices_cv:\n                cap.set(cv2.CAP_PROP_POS_FRAMES, idx_to_read)\n                ret, frame = cap.read()\n                if ret:\n                    frame_rgb = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)\n                    frames_list.append(torch.from_numpy(frame_rgb))\n                else:\n                    # If read fails, append last good frame (or handle differently)\n                    if frames_list: frames_list.append(frames_list[-1].clone())\n                    else: break # Cannot read even first frame\n\n            cap.release()\n\n            if len(frames_list) < seq_len: # Pad if necessary\n                if not frames_list: return None\n                padding = frames_list[-1].unsqueeze(0).repeat(seq_len - len(frames_list), 1, 1, 1)\n                frames_tensor = torch.cat((torch.stack(frames_list, dim=0), padding), dim=0)\n            else:\n                frames_tensor = torch.stack(frames_list[:seq_len], dim=0)\n\n        except Exception as e_cv:\n            print(f\"OpenCV also failed for {video_path}: {e_cv}\")\n            return None\n\n    # Final check: ensure tensor is not None and has correct shape\n    if frames_tensor is None or frames_tensor.shape[0] != seq_len:\n         print(f\"ERROR: Final frame tensor shape incorrect for {video_path}. Got {frames_tensor.shape if frames_tensor is not None else 'None'}\")\n         return None\n\n    return frames_tensor # Shape (T, H, W, C) uint8","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-23T15:38:59.530357Z","iopub.execute_input":"2025-04-23T15:38:59.530710Z","iopub.status.idle":"2025-04-23T15:38:59.537494Z","shell.execute_reply.started":"2025-04-23T15:38:59.530681Z","shell.execute_reply":"2025-04-23T15:38:59.536583Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Create a PyTorch Dataset","metadata":{}},{"cell_type":"code","source":"import os\nimport cv2\nimport pandas as pd\nimport numpy as np\nimport torch\nfrom torch.utils.data import Dataset \nfrom torchvision import transforms\n\n\n# --- PyTorch Dataset ---\n\n\nclass DashcamDataset(Dataset):\n    def __init__(self, csv_file, video_dir, seq_len, transform=None, mode='train', apply_hflip=False, flip_p=0.5):\n        \"\"\"\n        Args:\n            csv_file (string): Path to the csv file (train, val, or test).\n            video_dir (string): Directory with all the video files.\n            seq_len (int): Number of frames per sequence.\n            transform (callable, optional): Transform applied AFTER potential HFlip.\n            mode (string): 'train', 'val', or 'test'.\n            apply_hflip (bool): Whether to apply consistent RandomHorizontalFlip (train mode only).\n            flip_p (float): Probability of applying the horizontal flip.\n        \"\"\"\n        self.metadata = pd.read_csv(csv_file)\n        self.video_dir = video_dir\n        self.seq_len = seq_len\n        self.transform = transform\n        self.mode = mode\n        self.apply_hflip = apply_hflip and self.mode == 'train'\n        self.flip_p = flip_p\n\n        # --- Corrected Target Column Handling ---\n        # The target is *already* in train.csv (0 or 1). For val split, it's copied.\n        # For test.csv, there is no target.\n        if self.mode in ['train', 'val']:\n            if 'target' not in self.metadata.columns:\n                 raise ValueError(f\"'target' column not found in {csv_file}. Ensure the CSV is correct.\")\n            # Ensure target is numeric (it should be 0/1 already)\n            self.metadata['target'] = pd.to_numeric(self.metadata['target'], errors='coerce')\n            if self.metadata['target'].isnull().any():\n                 print(f\"Warning: Found NaN values in target column of {csv_file}. Check data integrity.\")\n                 # Decide on handling: dropna, fillna(0)? For now, keep going.\n\n    def __len__(self):\n        return len(self.metadata)\n\n    def __getitem__(self, idx):\n        if torch.is_tensor(idx):\n            idx = idx.tolist()\n\n        row = self.metadata.iloc[idx]\n        video_id_raw = row['id'] # Get the ID as it is in the CSV\n\n        # --- Determine Correct Filename ---\n        # Try converting to int first for padding, fallback to float/string if needed\n        try:\n            video_id_int = int(video_id_raw)\n            video_id_padded = str(video_id_int).zfill(5)\n        except ValueError:\n            # Handle cases where ID might be float (like 667.0) or already a padded string\n            video_id_str = str(video_id_raw)\n            if '.' in video_id_str: # Handle float IDs like '667.0'\n                video_id_padded = str(int(float(video_id_str))).zfill(5)\n            else: # Assume it might already be padded or just needs zfill\n                 video_id_padded = video_id_str.zfill(5)\n                \n        video_path = os.path.join(self.video_dir, f\"{video_id_padded}.mp4\")\n\n        # Use Decord to sample frames -> (T, H, W, C), uint8, CPU\n        frames_tensor_hwc = sample_frames_robust_stride(video_path, self.seq_len,FRAME_STRIDE)\n\n        if frames_tensor_hwc is None:\n            print(f\"Error processing video {video_id_padded}, returning None.\")\n            return None # collate_fn handles this\n\n        # Permute HWC -> CHW and convert to float [0.0, 1.0]\n        frames_tensor_chw = frames_tensor_hwc.permute(0, 3, 1, 2).float() / 255.0\n\n        # Apply Consistent Augmentations (like HFlip) BEFORE Compose\n        if self.apply_hflip and torch.rand(1) < self.flip_p:\n             frames_tensor_chw = transforms.functional.hflip(frames_tensor_chw)\n\n        # Apply Compose Transforms (Resize, Normalize, etc.)\n        if self.transform:\n            frames_processed = self.transform(frames_tensor_chw)\n        else:\n            frames_processed = frames_tensor_chw\n\n        # Get Label/ID\n        if self.mode in ['train', 'val']:\n            label = row['target'] # Get target directly from the loaded CSV\n            # Handle potential NaNs read from CSV if any issues occurred\n            if pd.isna(label):\n                print(f\"Warning: NaN label encountered for video {video_id_padded} at index {idx}. Returning None.\")\n                return None # Or return a default label like 0.0? None is safer for collate_fn.\n            label = torch.tensor(label, dtype=torch.float32)\n            return frames_processed, label\n        else: # 'test' mode\n            # Return the original, non-padded ID as string\n            return frames_processed, str(video_id_raw)\n\nprint(\"DashcamDataset class defined.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-23T15:49:31.954314Z","iopub.execute_input":"2025-04-23T15:49:31.954694Z","iopub.status.idle":"2025-04-23T15:49:31.968161Z","shell.execute_reply.started":"2025-04-23T15:49:31.954660Z","shell.execute_reply":"2025-04-23T15:49:31.967277Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- Transforms ---\n# Define normalization (using ImageNet stats as a standard starting point)\nnormalize = transforms.Normalize(mean=[0.485, 0.456, 0.406],\n                                 std=[0.229, 0.224, 0.225])\n\n# Define Resize transform - works directly on (T, C, H, W)\n# Use antialias=True for better quality when downsampling\nresize = transforms.Resize((IMG_SIZE, IMG_SIZE), antialias=True)\n\n# --- Validation/Test Transforms ---\n# Input: Tensor (T, C, H, W), float [0.0, 1.0]\n# Output: Tensor (T, C, IMG_SIZE, IMG_SIZE), normalized\nval_test_transform = transforms.Compose([\n    resize,      # Resize each frame in the sequence\n    normalize    # Normalize the sequence tensor\n])\n\n# --- Training Transforms ---\n# Input: Tensor (T, C, H, W), float [0.0, 1.0]\n# Output: Tensor (T, C, IMG_SIZE, IMG_SIZE), augmented, normalized\n# Note: RandomHorizontalFlip is applied *before* this Compose in Dataset.__getitem__ for consistency\n\n# Optional: Define RandomCrop if desired (applied independently per frame)\n# If using RandomCrop, Resize should output slightly larger first\n# RESIZE_FOR_CROP_SIZE = IMG_SIZE + 32 # Example padding\n# resize_for_crop = transforms.Resize((RESIZE_FOR_CROP_SIZE, RESIZE_FOR_CROP_SIZE), antialias=True)\n# random_crop = transforms.RandomCrop((IMG_SIZE, IMG_SIZE))\n\ntrain_transform = transforms.Compose([resize,transforms.ColorJitter(brightness=0.5, contrast=0.5, saturation=0.2, hue=0.1),transforms.RandomAffine(degrees=10, translate=(0.05, 0.05), scale=(0.9, 1.1)),normalize])\n\n# Assign to variables used later (if needed, e.g., when creating datasets)\ntrain_transform_final = train_transform\nval_transform_final = val_test_transform\ntest_transform_final = val_test_transform\n\nprint(\"Transform pipelines defined (Decord compatible).\")\nprint(\"Note: RandomHorizontalFlip handled in Dataset.__getitem__ for sequence consistency.\")\n\n# --- Collate Function ---\ndef collate_fn(batch):\n    \"\"\"\n    Custom collate function to handle batches from DashcamDataset.\n    \"\"\"\n    # 1. Filter out None values\n    batch = [item for item in batch if item is not None]\n\n    # 2. Handle empty batch case\n    if not batch:\n        print(\"Warning: Collate received an empty batch after filtering.\")\n        # Return structure suitable for either mode if possible, or handle in loop\n        # Returning structure potentially suitable for test mode prediction loop\n        return torch.empty((0, SEQ_LEN, N_CHANNELS, IMG_SIZE, IMG_SIZE)), []\n\n    # 3. Determine mode\n    is_test_mode = isinstance(batch[0][1], str)\n\n    # 4. Unpack and stack\n    if is_test_mode: # Test mode\n        frames, ids = zip(*batch)\n        frames_batch = default_collate(frames) # Use default_collate for tensors\n        return frames_batch, list(ids)\n    else: # Train/Val mode\n        # Default collate works for (tensor, tensor) tuples\n        return default_collate(batch)\n\nprint(\"Collate function defined.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-23T15:49:45.005963Z","iopub.execute_input":"2025-04-23T15:49:45.006266Z","iopub.status.idle":"2025-04-23T15:49:45.013946Z","shell.execute_reply.started":"2025-04-23T15:49:45.006242Z","shell.execute_reply":"2025-04-23T15:49:45.013206Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import torch.optim as optim\nfrom sklearn.model_selection import train_test_split\nimport os # Make sure os is imported\nimport pandas as pd # Make sure pandas is imported\n\n# --- (Previous code: DEVICE, Model Hyperparams, Training Hyperparams) ---\nDEVICE = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nprint(f\"Using device: {DEVICE}\")\nNUM_CLASSES = 1\nLSTM_HIDDEN_SIZE = 512\nLSTM_LAYERS = 2\nPRETRAINED = True\nLEARNING_RATE = 1e-4\n\n# --- Paths ---\nBASE_DATA_PATH = '/kaggle/input/nexar-collision-prediction'\nORIGINAL_TRAIN_CSV = os.path.join(BASE_DATA_PATH, 'train.csv') # Keep original path separate\n\n# Define paths for the split files in the working directory\nTRAIN_CSV_SPLIT = '/kaggle/working/train_split.csv'\nVAL_CSV_SPLIT = '/kaggle/working/val_split.csv'\n\n# --- Perform Split ONLY if files don't exist ---\n# --- Perform Split ONLY if files don't exist ---\nif not os.path.exists(TRAIN_CSV_SPLIT) or not os.path.exists(VAL_CSV_SPLIT):\n    print(f\"Split files not found in /kaggle/working/. Performing train_test_split on {ORIGINAL_TRAIN_CSV}...\")\n    try:\n        full_train_df = pd.read_csv(ORIGINAL_TRAIN_CSV)\n\n        # --- Use the EXISTING 'target' column for stratification ---\n        if 'target' not in full_train_df.columns:\n             print(\"ERROR: Cannot perform stratified split because 'target' column is missing from train.csv.\")\n             # Handle error: maybe raise exception or proceed without stratification\n             stratify_col = None\n        else:\n            # Ensure target is suitable for stratification (handle potential NaNs if any)\n            full_train_df['target'] = pd.to_numeric(full_train_df['target'], errors='coerce').fillna(-1).astype(int) # Temp fillna for stratify\n            if (full_train_df['target'] == -1).any():\n                print(\"Warning: Found missing/non-numeric targets in train.csv. Stratification might be affected.\")\n            stratify_col = full_train_df['target']\n            print(f\"Stratifying split using 'target' column. Distribution:\\n{stratify_col.value_counts(normalize=True)}\")\n\n        train_df, val_df = train_test_split(full_train_df,\n                                            test_size=0.15,          # Validation set size\n                                            random_state=42,       # For reproducibility\n                                            stratify=stratify_col) # Stratify based on target\n\n        # Before saving, potentially revert the fillna if it was temporary\n        # train_df['target'] = train_df['target'].replace(-1, np.nan)\n        # val_df['target'] = val_df['target'].replace(-1, np.nan)\n\n        train_df.to_csv(TRAIN_CSV_SPLIT, index=False)\n        val_df.to_csv(VAL_CSV_SPLIT, index=False)\n        print(f\"Created {TRAIN_CSV_SPLIT} and {VAL_CSV_SPLIT}\")\n        del full_train_df, train_df, val_df # Clean up memory\n        gc.collect()\n\n    except FileNotFoundError:\n        print(f\"ERROR: Original train CSV not found at {ORIGINAL_TRAIN_CSV}. Cannot perform split.\")\n        TRAIN_CSV_PATH_FINAL = None # Indicate failure\n        VAL_CSV_PATH_FINAL = None\n    except Exception as e:\n         print(f\"An error occurred during train/test split: {e}\")\n         TRAIN_CSV_PATH_FINAL = None\n         VAL_CSV_PATH_FINAL = None\nelse:\n    print(f\"Using existing split files: {TRAIN_CSV_PATH_FINAL} and {VAL_CSV_PATH_FINAL}\")\n\n# Check if paths are valid before proceeding\nif TRAIN_CSV_PATH_FINAL is None or VAL_CSV_PATH_FINAL is None:\n    raise RuntimeError(\"Train/Validation split failed. Cannot proceed.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-23T15:49:51.400079Z","iopub.execute_input":"2025-04-23T15:49:51.400368Z","iopub.status.idle":"2025-04-23T15:49:51.410102Z","shell.execute_reply.started":"2025-04-23T15:49:51.400346Z","shell.execute_reply":"2025-04-23T15:49:51.409322Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- Create Datasets ---\nprint(\"\\nCreating Datasets using split files...\")\n\ntrain_dataset = DashcamDataset(csv_file=TRAIN_CSV_PATH_FINAL,\n                               video_dir=TRAIN_VIDEO_DIR,\n                               seq_len=SEQ_LEN,\n                               transform=train_transform_final,\n                               mode='train',\n                               apply_hflip=True,\n                               flip_p=0.5)\nprint(f\"Training dataset size: {len(train_dataset)}\")\n\nval_dataset = DashcamDataset(csv_file=VAL_CSV_PATH_FINAL,\n                             video_dir=TRAIN_VIDEO_DIR,\n                             seq_len=SEQ_LEN,\n                             transform=val_transform_final,\n                             mode='val',\n                             apply_hflip=False)\nprint(f\"Validation dataset size: {len(val_dataset)}\")\n\n# --- Create DataLoaders ---\nprint(\"\\nCreating DataLoaders...\")\ntrain_loader = DataLoader(train_dataset,\n                          batch_size=BATCH_SIZE,\n                          shuffle=True,\n                          num_workers=NUM_WORKERS,\n                          pin_memory=True,\n                          prefetch_factor=PREFETCH_FACTOR,\n                          collate_fn=collate_fn)\n\nval_loader = DataLoader(val_dataset,\n                        batch_size=BATCH_SIZE * 2, # Can use larger BS for validation\n                        shuffle=False,\n                        num_workers=NUM_WORKERS,\n                        pin_memory=True,\n                        prefetch_factor=PREFETCH_FACTOR,\n                        collate_fn=collate_fn)\n\nprint(\"Train and Validation DataLoaders created.\")\n\n# --- Optional: Sanity Check Iteration (before training) ---\nprint(\"\\nIterating through one batch of Train Loader for sanity check...\")\ntry:\n    if len(train_loader) == 0:\n         print(\"Train loader is empty. Cannot iterate.\")\n    else:\n        first_batch_data = next(iter(train_loader))\n        if first_batch_data is None or not first_batch_data[0].numel():\n            print(f\"First batch is empty or None.\")\n        else:\n            frames_batch, labels_batch = first_batch_data\n            print(f\"Example Batch 1:\")\n            print(\"  Frames batch shape:\", frames_batch.shape) # Should be (B, T, C, H, W)\n            print(\"  Labels batch shape:\", labels_batch.shape) # Should be (B,)\n            print(\"  Labels:\", labels_batch)\n            print(\"\\nIteration Example Complete.\")\n            del first_batch_data, frames_batch, labels_batch # Clean up memory\n            gc.collect()\nexcept StopIteration:\n     print(\"Train loader is empty or could not fetch the first batch.\")\nexcept Exception as e:\n    print(f\"\\nError during DataLoader iteration check: {e}\")\n    import traceback\n    traceback.print_exc()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-23T15:49:59.265274Z","iopub.execute_input":"2025-04-23T15:49:59.265643Z","iopub.status.idle":"2025-04-23T15:52:21.413530Z","shell.execute_reply.started":"2025-04-23T15:49:59.265617Z","shell.execute_reply":"2025-04-23T15:52:21.412659Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Define the Model Architecture.\n\nYou need a neural network model that can take your batches of video sequences (with shape like [Batch Size, Sequence Length, Channels, Height, Width], e.g., [8, 16, 3, 224, 224]) and output a single prediction score (between 0 and 1) for each sequence in the batch (shape [Batch Size]).\n\nBased on our earlier discussion and the nature of the task (video classification/event prediction), a common and effective starting point is the CNN + LSTM architecture:\n\nCNN Feature Extractor: Use a pre-trained 2D Convolutional Neural Network (like ResNet, EfficientNet, MobileNet, etc.) designed for image classification. You'll remove its final classification layer.\n\nRole: Processes each frame individually to extract spatial features. It converts each (C, H, W) frame into a fixed-size feature vector (e.g., size 512, 1024, or 2048 depending on the CNN).\n\nSequence Model (LSTM/GRU): Use a Recurrent Neural Network layer (like LSTM or GRU).\n\nRole: Takes the sequence of feature vectors (one vector per frame, output by the CNN) and models the temporal relationships between them. It learns patterns over time within the sequence.\n\nClassifier Head: Usually one or more fully connected (Linear) layers.\n\nRole: Takes the final output from the LSTM/GRU and maps it to the single output score. A Sigmoid activation function is used at the very end to ensure the output is between 0 and 1, representing the probability of a collision/near-miss.","metadata":{}},{"cell_type":"markdown","source":"# Key Decisions in this Code:\n\nCNN Backbone: Uses ResNet18. You can easily swap this for resnet34, resnet50, efficientnet_b0, etc. Using pretrained=True leverages knowledge learned from ImageNet.\n\nLSTM Input Size: Automatically determined from the chosen CNN backbone (num_cnn_features).\n\nLSTM Output: Uses the output features from the very last time step (lstm_output[:, -1, :]) as the summary of the sequence for classification. This is a common choice.\n\nOutput Layer: Outputs a single value (num_classes=1). The Sigmoid activation, needed to get the 0-1 probability, is often combined with the loss function (BCEWithLogitsLoss) for better numerical stability during training.","metadata":{}},{"cell_type":"code","source":"import torch\nimport torch.nn as nn\nimport torchvision.models as models\n\nclass VideoClassifierCNN_LSTM(nn.Module):\n    def __init__(self, num_classes=1, lstm_hidden_size=512, lstm_layers=2, pretrained=True):\n        \"\"\"\n        Args:\n            num_classes (int): Number of output classes (1 for binary classification with sigmoid).\n            lstm_hidden_size (int): Number of features in the LSTM hidden state.\n            lstm_layers (int): Number of recurrent layers in LSTM.\n            pretrained (bool): Whether to use a pretrained CNN backbone.\n        \"\"\"\n        super().__init__()\n\n        self.lstm_hidden_size = lstm_hidden_size\n        self.lstm_layers = lstm_layers\n\n        # --- CNN Backbone ---\n        # Load a pretrained ResNet (or another model like EfficientNet)\n        # We'll remove the final fully connected layer (the original classifier)\n        base_model = models.resnet18(pretrained=pretrained) # Example: ResNet18\n        # Get the number of features output by the ResNet's pooling layer\n        num_cnn_features = base_model.fc.in_features\n        # Remove the final layer\n        modules = list(base_model.children())[:-1]\n        self.cnn_backbone = nn.Sequential(*modules)\n        # Freeze backbone layers if desired (transfer learning)\n        # for param in self.cnn_backbone.parameters():\n        #     param.requires_grad = False\n\n\n        # --- LSTM Layer ---\n        # Input features to LSTM will be the output features from the CNN backbone\n        self.lstm = nn.LSTM(input_size=num_cnn_features,\n                            hidden_size=lstm_hidden_size,\n                            num_layers=lstm_layers,\n                            batch_first=True, # Input shape: (batch, seq_len, features)\n                            dropout=0.6 if lstm_layers > 1 else 0) # Add dropout if multiple layers\n\n        # --- Classifier Head ---\n        self.fc1 = nn.Linear(lstm_hidden_size, lstm_hidden_size // 2)\n        self.relu = nn.ReLU()\n        self.dropout = nn.Dropout(0.65)\n        self.fc2 = nn.Linear(lstm_hidden_size // 2, num_classes)\n        # Sigmoid activation will be applied later (usually with BCEWithLogitsLoss for stability)\n        # or explicitly here if using BCELoss\n\n    def forward(self, x):\n        # x shape: (batch_size, seq_len, C, H, W)\n\n        batch_size, seq_len, C, H, W = x.shape\n\n        # --- Pass through CNN ---\n        # Reshape input for CNN: Treat sequence dimension as part of the batch\n        cnn_input = x.view(batch_size * seq_len, C, H, W)\n        # Get features from CNN\n        cnn_output = self.cnn_backbone(cnn_input) # Shape: (batch*seq_len, num_cnn_features, 1, 1)\n        # Remove spatial dimensions (squeeze)\n        cnn_features = cnn_output.view(batch_size * seq_len, -1) # Shape: (batch*seq_len, num_cnn_features)\n        # Reshape back into sequence: (batch_size, seq_len, num_cnn_features)\n        lstm_input = cnn_features.view(batch_size, seq_len, -1)\n\n        # --- Pass through LSTM ---\n        # Initialize hidden and cell states (optional, defaults to zeros)\n        # h0 = torch.zeros(self.lstm_layers, batch_size, self.lstm_hidden_size).to(x.device)\n        # c0 = torch.zeros(self.lstm_layers, batch_size, self.lstm_hidden_size).to(x.device)\n        # Get LSTM output (output features for each time step + final hidden/cell states)\n        # lstm_output shape: (batch_size, seq_len, lstm_hidden_size)\n        # hn shape: (num_layers, batch_size, lstm_hidden_size)\n        # cn shape: (num_layers, batch_size, lstm_hidden_size)\n        lstm_output, (hn, cn) = self.lstm(lstm_input) #, (h0, c0))\n\n        # --- Classifier ---\n        # We usually use the output of the *last* time step from the LSTM\n        last_lstm_output = lstm_output[:, -1, :] # Shape: (batch_size, lstm_hidden_size)\n\n        # Pass through classifier head\n        out = self.fc1(last_lstm_output)\n        out = self.relu(out)\n        out = self.dropout(out)\n        out = self.fc2(out) # Shape: (batch_size, num_classes)\n\n        # Remove the last dimension if num_classes is 1\n        if out.shape[1] == 1:\n            out = out.squeeze(1) # Shape: (batch_size,)\n\n        return out\n\n# --- Example Usage (how to create the model) ---\n# model = VideoClassifierCNN_LSTM(num_classes=1, lstm_hidden_size=512, lstm_layers=2)\n# print(model)\n\n# --- Move to GPU if available ---\n# device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n# model.to(device)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-23T15:54:34.430893Z","iopub.execute_input":"2025-04-23T15:54:34.431236Z","iopub.status.idle":"2025-04-23T15:54:34.439620Z","shell.execute_reply.started":"2025-04-23T15:54:34.431201Z","shell.execute_reply":"2025-04-23T15:54:34.438827Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# -------------------------------------------------------------\n# --- Step 2: Instantiate Model, Loss, Optimizer ---\n# -------------------------------------------------------------\nprint(\"\\nInstantiating Model, Loss, and Optimizer...\")\n\nmodel = VideoClassifierCNN_LSTM(\n    num_classes=NUM_CLASSES,\n    lstm_hidden_size=LSTM_HIDDEN_SIZE,\n    lstm_layers=LSTM_LAYERS,\n    pretrained=PRETRAINED\n).to(DEVICE)\n\n# Loss Function - BCEWithLogitsLoss is recommended for binary classification\n# as it combines Sigmoid and BCELoss for numerical stability.\ncriterion = nn.BCEWithLogitsLoss()\n\n# Optimizer - Adam is a popular choice\noptimizer = optim.AdamW(model.parameters(), lr=LEARNING_RATE, weight_decay=1e-3)\n\n# Learning Rate Scheduler (e.g., reduce LR if validation loss plateaus)\nscheduler = optim.lr_scheduler.ReduceLROnPlateau(optimizer, mode='min', factor=0.1, patience=3, verbose=True)\n\nprint(\"Model, Loss, Optimizer instantiated.\")\n# You can print the model summary if needed:\n# print(model)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-23T16:06:04.383064Z","iopub.execute_input":"2025-04-23T16:06:04.383498Z","iopub.status.idle":"2025-04-23T16:06:05.084291Z","shell.execute_reply.started":"2025-04-23T16:06:04.383464Z","shell.execute_reply":"2025-04-23T16:06:05.083424Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import torch\nfrom torch.utils.data import DataLoader\nimport torch.nn as nn\nimport torch.optim as optim\nfrom tqdm.notebook import tqdm # Or regular tqdm\n\n# --- AMP Imports ---\nfrom torch.cuda.amp import GradScaler, autocast\n\n# --- Dummy model for example ---\n# Assume your VideoClassifierCNN_LSTM model is defined elsewhere\nclass DummyModel(nn.Module):\n    def __init__(self):\n        super().__init__()\n        self.fc = nn.Linear(16*3*224*224, 1) # Example structure\n    def forward(self, x):\n        return self.fc(x.view(x.size(0), -1)).squeeze(1) # Output (B,)\n\n# -------------------------------------------------------------\n# --- Step 4 & 5: Implement Training and Evaluation Loops ---\n# -------------------------------------------------------------\n\n# --- Initialize GradScaler ONCE, typically before the main training loop ---\n# It should only be enabled when using CUDA\n# DEVICE should be defined globally (e.g., DEVICE = torch.device('cuda' if torch.cuda.is_available() else 'cpu'))\n# We pass the scaler into the training function\n# scaler = GradScaler(enabled=(DEVICE.type == 'cuda'))\n\ndef train_one_epoch(model, loader, criterion, optimizer, scaler, device, epoch_num, max_epochs):\n    \"\"\"Trains the model for one epoch using AMP.\"\"\"\n    model.train() # Set model to training mode\n    running_loss = 0.0\n    correct_predictions = 0\n    total_samples = 0\n\n    # Use tqdm for progress bar\n    progress_bar = tqdm(loader, desc=f\"Epoch {epoch_num+1}/{max_epochs} [Train]\", leave=False, unit=\"batch\")\n\n    for batch_idx, batch_data in enumerate(progress_bar):\n        # Check for empty batch from collate_fn\n        if batch_data is None or not batch_data[0].numel():\n            print(f\"Warning: Skipping empty training batch {batch_idx}\")\n            continue\n\n        videos, labels = batch_data\n        videos = videos.to(device) # (B, T, C, H, W)\n        labels = labels.to(device) # (B,) - ensure criterion expects this shape\n\n        # Zero gradients BEFORE the forward pass in this batch\n        optimizer.zero_grad()\n\n        # --- Automatic Mixed Precision Context ---\n        # Runs the forward pass under autocast\n        with torch.amp.autocast(device_type=DEVICE.type, enabled=(DEVICE.type == 'cuda')):\n            outputs = model(videos) # Should be (B,) for BCEWithLogitsLoss\n            loss = criterion(outputs, labels)\n\n        # --- Scaled Backward Pass ---\n        # Scales loss. Calls backward() on scaled loss to create scaled gradients.\n        scaler.scale(loss).backward()\n\n        # Optional: Gradient Clipping (apply *before* scaler.step)\n        # If you clip, unscale the gradients first\n        # scaler.unscale_(optimizer) # Unscale gradients back to fp32\n        # torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)\n\n        # --- Scaler Step and Update ---\n        # scaler.step() first unscales the gradients of the optimizer's assigned params.\n        # If these gradients do not contain infs or NaNs, optimizer.step() is then called.\n        # Otherwise, optimizer.step() is skipped.\n        scaler.step(optimizer)\n\n        # Updates the scale for next iteration.\n        scaler.update()\n\n        # --- Update Statistics ---\n        running_loss += loss.item() * videos.size(0) # Loss per batch * batch size\n        total_samples += labels.size(0)\n\n        # Calculate accuracy (outside autocast context)\n        with torch.no_grad():\n            preds = torch.sigmoid(outputs) > 0.5 # Shape: (B, 1) if output was (B,1), or (B,) if output was (B,)\n            # Ensure labels are boolean and match pred shape for comparison\n            correct_predictions += (preds == labels.bool().view_as(preds)).sum().item() # Use view_as for safety\n        # Update progress bar postfix\n        current_loss = loss.item()\n        current_acc = correct_predictions / total_samples if total_samples > 0 else 0\n        progress_bar.set_postfix(loss=f\"{current_loss:.4f}\", acc=f\"{current_acc:.4f}\")\n\n    # End of Epoch Calculations\n    epoch_loss = running_loss / total_samples if total_samples > 0 else 0\n    epoch_acc = correct_predictions / total_samples if total_samples > 0 else 0\n    progress_bar.close() # Close the tqdm bar for this epoch\n\n    return epoch_loss, epoch_acc\n\n\ndef evaluate(model, loader, criterion, device, epoch_num, max_epochs):\n    model.eval() # Set model to evaluation mode\n    running_loss = 0.0\n    correct_predictions = 0\n    total_samples = 0\n    all_outputs = []\n    all_labels = []\n\n    progress_bar = tqdm(loader, desc=f\"Epoch {epoch_num+1}/{max_epochs} [Val.] \", leave=False, unit=\"batch\")\n\n    with torch.no_grad():\n        for batch_idx, batch_data in enumerate(progress_bar):\n            if batch_data is None or not batch_data[0].numel():\n                print(f\"Warning: Skipping empty validation batch {batch_idx}\")\n                continue\n\n            videos, labels = batch_data\n            videos = videos.to(device)\n            labels = labels.to(device) # Shape (B,)\n\n            with torch.amp.autocast(device_type=DEVICE.type, enabled=(DEVICE.type == 'cuda')):\n                outputs = model(videos) # Shape (B,)\n                loss = criterion(outputs, labels)\n\n            running_loss += loss.item() * videos.size(0)\n            total_samples += labels.size(0)\n\n            # --- CORRECTED Accuracy Calculation ---\n            preds = torch.sigmoid(outputs) > 0.5 # Get probabilities and threshold -> Shape (B,) boolean\n            correct_predictions += (preds == labels.bool()).sum().item() # Direct comparison if labels are (B,)\n\n            all_outputs.append(torch.sigmoid(outputs).cpu())\n            all_labels.append(labels.cpu())\n\n            # Display *batch* accuracy (optional, but useful for debugging)\n            batch_acc = (preds == labels.bool()).float().mean().item()\n            progress_bar.set_postfix(loss=f\"{loss.item():.4f}\", batch_acc=f\"{batch_acc:.4f}\") # Show batch acc\n\n    # --- CORRECTED Epoch Accuracy ---\n    epoch_loss = running_loss / total_samples if total_samples > 0 else 0\n    epoch_acc = correct_predictions / total_samples if total_samples > 0 else 0\n    progress_bar.close()\n\n    # --- Optional: Calculate AUC here ---\n    # try:\n    #     all_outputs_cat = torch.cat(all_outputs).numpy()\n    #     all_labels_cat = torch.cat(all_labels).numpy()\n    #     val_auc = roc_auc_score(all_labels_cat, all_outputs_cat)\n    #     print(f\"  Val AUC: {val_auc:.4f}\")\n    # except ValueError:\n    #     print(\"  Val AUC: Could not calculate (likely only one class present in batch)\")\n    #     val_auc = -1.0 # Or some indicator\n\n    return epoch_loss, epoch_acc #, val_auc","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-23T15:58:34.419316Z","iopub.execute_input":"2025-04-23T15:58:34.419668Z","iopub.status.idle":"2025-04-23T15:58:34.432800Z","shell.execute_reply.started":"2025-04-23T15:58:34.419638Z","shell.execute_reply":"2025-04-23T15:58:34.431897Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Cell [13]: Main Training Loop (Corrected Checkpoint Paths)\n\n# -------------------------------------------------------------\n# --- Step 6: Run the Training Process ---\n# -------------------------------------------------------------\n\n# --- Initialize variables BEFORE checking for checkpoint ---\nstart_epoch = 0\nbest_val_loss = float('inf')\nhistory = {'train_loss': [], 'train_acc': [], 'val_loss': [], 'val_acc': []}\n\n# Initialize scaler (using updated torch.amp syntax)\nscaler = torch.amp.GradScaler('cuda', enabled=(DEVICE.type == 'cuda'))\n\n# --- Early Stopping Parameters ---\n\npatience_counter = 0\n\n# --- Check for and load existing checkpoint ---\nif os.path.exists(CHECKPOINT_PATH): # <-- Use CHECKPOINT_PATH variable\n    print(f\"Loading checkpoint from {CHECKPOINT_PATH}...\")\n    try:\n        checkpoint = torch.load(CHECKPOINT_PATH, map_location=DEVICE)\n        if 'model_state_dict' in checkpoint and 'optimizer_state_dict' in checkpoint and 'epoch' in checkpoint:\n            model.load_state_dict(checkpoint['model_state_dict'])\n            optimizer.load_state_dict(checkpoint['optimizer_state_dict'])\n            start_epoch = checkpoint['epoch']\n            best_val_loss = checkpoint.get('best_val_loss', float('inf'))\n            if 'scheduler_state_dict' in checkpoint and 'scheduler' in locals() and scheduler is not None: # Check if scheduler exists\n                try: # Add extra try-except for scheduler loading\n                    scheduler.load_state_dict(checkpoint['scheduler_state_dict'])\n                    print(\"Loaded scheduler state.\")\n                except Exception as e_sched:\n                    print(f\"Warning: Could not load scheduler state: {e_sched}\")\n            # Load patience counter if saved, otherwise reset\n            patience_counter = checkpoint.get('patience_counter', 0)\n            print(f\"Resuming training from Epoch {start_epoch}\")\n            print(f\"Previous best validation loss: {best_val_loss:.4f}\")\n            print(f\"Resuming patience counter: {patience_counter}\")\n        else:\n            print(\"Checkpoint file seems incomplete or only model state_dict. Loading weights only.\")\n            model.load_state_dict(checkpoint)\n            print(\"Starting training from epoch 0, optimizer/scheduler state not resumed.\")\n            start_epoch = 0\n            best_val_loss = float('inf')\n            patience_counter = 0 # Reset patience\n\n    except Exception as e:\n        print(f\"Error loading checkpoint: {e}. Starting training from scratch.\")\n        start_epoch = 0\n        best_val_loss = float('inf')\n        patience_counter = 0\nelse:\n    print(\"No checkpoint found. Starting training from scratch.\")\n\n\nprint(\"\\n--- Starting Training ---\")\n\n# MAX_EPOCHS should be defined in your config cell\nfor epoch in range(start_epoch, MAX_EPOCHS):\n    # --- Training Phase ---\n    train_loss, train_acc = train_one_epoch(model, train_loader, criterion, optimizer, scaler, DEVICE, epoch, MAX_EPOCHS) # Pass scaler\n\n    # --- Validation Phase ---\n    if 'val_loader' in locals() and val_loader is not None: # Check if val_loader was created\n        val_loss, val_acc = evaluate(model, val_loader, criterion, DEVICE, epoch, MAX_EPOCHS)\n    else:\n        print(\"Warning: No validation loader available. Skipping validation and early stopping.\")\n        val_loss, val_acc = -1.0, -1.0\n\n    # --- Log Epoch Results ---\n    print(f\"Epoch {epoch+1}/{MAX_EPOCHS} Summary:\")\n    print(f\"  Train Loss: {train_loss:.4f} | Train Acc: {train_acc:.4f}\")\n    if 'val_loader' in locals() and val_loader is not None:\n        print(f\"  Val. Loss:  {val_loss:.4f} | Val. Acc:  {val_acc:.4f}\") # Ensure acc calc is correct\n\n    # --- Store history ---\n    history['train_loss'].append(train_loss)\n    history['train_acc'].append(train_acc)\n    if 'val_loader' in locals() and val_loader is not None:\n        history['val_loss'].append(val_loss)\n        history['val_acc'].append(val_acc)\n    else: # Append placeholders if no validation\n        history['val_loss'].append(None)\n        history['val_acc'].append(None)\n\n\n    # --- Learning Rate Scheduling ---\n    if 'val_loader' in locals() and val_loader is not None and 'scheduler' in locals() and scheduler is not None:\n       scheduler.step(val_loss) # Step scheduler based on validation loss\n\n    # --- Early Stopping Logic & Checkpoint Saving ---\n    if 'val_loader' in locals() and val_loader is not None:\n        if val_loss < best_val_loss:\n            best_val_loss = val_loss\n            patience_counter = 0\n            # *** CORRECTED PATH ***\n            model_save_path = CHECKPOINT_PATH # Use the full path variable\n            checkpoint = {\n                'epoch': epoch + 1,\n                'model_state_dict': model.state_dict(),\n                'optimizer_state_dict': optimizer.state_dict(),\n                'best_val_loss': best_val_loss,\n                'scheduler_state_dict': scheduler.state_dict() if 'scheduler' in locals() and scheduler is not None else None,\n                'patience_counter': patience_counter # Optional: Save patience counter state\n            }\n            print(f\"  * Validation loss improved to {best_val_loss:.4f}. Saving checkpoint to {model_save_path}. Patience reset.\")\n            # Add a disk space check before saving (optional but helpful)\n            # !df -h /kaggle/working/\n            print(f\"DEBUG: Attempting to save checkpoint to {model_save_path} for epoch {epoch+1}\") # Add debug print\n            try:\n                torch.save(checkpoint, model_save_path)\n                print(f\"DEBUG: Save successful for epoch {epoch+1}\") # Confirm success\n            except Exception as e:\n                print(f\"DEBUG: Error saving checkpoint: {e}\") # Catch specific save errors\n        else:\n            # Validation loss did not improve\n            patience_counter += 1\n            print(f\"  * Validation loss ({val_loss:.4f}) did not improve from best ({best_val_loss:.4f}). Patience: {patience_counter}/{EARLY_STOPPING_PATIENCE}\")\n            if patience_counter >= EARLY_STOPPING_PATIENCE:\n                print(f\"--- Early stopping triggered after {EARLY_STOPPING_PATIENCE} epochs without improvement. ---\")\n                break # Exit the training loop\n    else:\n        # --- Optional Periodic Saving (if no validation) ---\n        if (epoch + 1) % 5 == 0: # Example: Save every 5 epochs\n             # *** CORRECTED PATH ***\n             periodic_save_path = os.path.join(WORKING_DIR, f\"model_epoch_{epoch+1}.pth\")\n             checkpoint = {\n                'epoch': epoch + 1,\n                'model_state_dict': model.state_dict(),\n                'optimizer_state_dict': optimizer.state_dict(),\n                'best_val_loss': best_val_loss,\n                'scheduler_state_dict': scheduler.state_dict() if 'scheduler' in locals() and scheduler is not None else None,\n                'patience_counter': patience_counter # Still save current patience state\n             }\n             print(f\"Saving periodic checkpoint to {periodic_save_path}\")\n             try:\n                 torch.save(checkpoint, periodic_save_path)\n             except Exception as e:\n                 print(f\"ERROR saving periodic checkpoint: {e}\")\n\n\nprint(\"\\n--- Training Finished ---\")\nif 'val_loader' in locals() and val_loader is not None:\n    print(f\"Best Validation Loss achieved: {best_val_loss:.4f}\") # This holds the best loss found\n    # Determine the actual epoch number where training stopped\n    final_epoch_completed = epoch + 1 # Add 1 because epoch is 0-indexed\n    print(f\"Training stopped after completing epoch: {final_epoch_completed}\")\nelse:\n    print(\"Training finished (fixed epochs or periodic saves).\")\n\n# --- IMPORTANT: Load BEST model before inference ---\nprint(\"\\nLoading best model weights for evaluation/submission...\")\ntry:\n    # Ensure CHECKPOINT_PATH points to the file saved for the best validation loss\n    checkpoint = torch.load(CHECKPOINT_PATH, map_location=DEVICE)\n    model.load_state_dict(checkpoint['model_state_dict'])\n    model.eval() # Set to evaluation mode\n    print(\"Successfully loaded best model weights.\")\nexcept FileNotFoundError:\n    print(f\"ERROR: Could not find best checkpoint at {CHECKPOINT_PATH} after training.\")\n    # Handle error - maybe use the model state from the very last epoch if needed, but it's not the best\nexcept Exception as e:\n     print(f\"Error loading best model weights after training: {e}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-23T16:06:21.461479Z","iopub.execute_input":"2025-04-23T16:06:21.461846Z","iopub.status.idle":"2025-04-23T16:23:35.175977Z","shell.execute_reply.started":"2025-04-23T16:06:21.461817Z","shell.execute_reply":"2025-04-23T16:23:35.171786Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The training loop completed, but the ultimate goal is to create a submission file for the competition based on predictions on the test set.\n\nThe previous code focused only on training and validation. \n\nObjective of the next section :-\n\nLoad the Test Data: \nInstantiate the Model: Create an instance of your VideoClassifierCNN_LSTM model.\nLoad the Best Weights: Load the state_dict from the best model saved during training (best_model_cnn_lstm.pth).\nSet to Evaluation Mode: Put the model in evaluation mode (model.eval()).\nCreate a Test DataLoader: Similar to the training/validation loaders, but using the test dataset and shuffle=False.\nPerform Inference: Iterate through the test DataLoader, pass the video sequences through the model, get the output scores (remember to apply sigmoid to the logits!), and store these scores along with their corresponding video IDs.\nFormat and Save the Submission File: Create a Pandas DataFrame with 'id' and 'score' columns and save it as submission.csv.","metadata":{}},{"cell_type":"code","source":"# --- Step 2 & 3: Instantiate Model and Load Weights ---\nprint(\"Loading trained model...\")\nMODEL_PATH='best_model_cnn_lstm.pth' # Path to your saved CHECKPOINT file\nDEVICE = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\n\n# 1. Instantiate the model architecture EXACTLY as used during training\nmodel = VideoClassifierCNN_LSTM(\n    num_classes=1,             # Ensure this matches training\n    lstm_hidden_size=512,      # Ensure this matches training\n    lstm_layers=2,             # Ensure this matches training\n    pretrained=False           # Usually False when loading fine-tuned weights\n                               # Set based on how you initialized the model\n                               # *before* loading the checkpoint during training/saving.\n                               # If you loaded pretrained weights *then* trained and saved,\n                               # you still initialize with pretrained=False here because\n                               # the state_dict contains the fine-tuned weights.\n).to(DEVICE)\n\ntry:\n    # 2. Load the entire checkpoint dictionary\n    # Use weights_only=False (default) because the checkpoint contains more than just weights\n    checkpoint = torch.load(MODEL_PATH, map_location=DEVICE)\n\n    # 3. Extract the model's state dictionary from the checkpoint\n    # Check if the expected key exists for robustness\n    if 'model_state_dict' in checkpoint:\n        model_weights = checkpoint['model_state_dict']\n    else:\n        # Handle case where the file might *only* contain the state dict\n        # (e.g., from an older saving method or different script)\n        print(\"Warning: Checkpoint dictionary does not contain 'model_state_dict'. Assuming the file IS the state dict.\")\n        model_weights = checkpoint # Assume the loaded object IS the state dict\n\n    # 4. Load the extracted weights into the model instance\n    model.load_state_dict(model_weights)\n\n    print(\"Model weights loaded successfully.\")\n\nexcept FileNotFoundError:\n    print(f\"ERROR: Model checkpoint not found at {MODEL_PATH}. Cannot proceed.\")\n    # Decide how to handle - exit(), raise error, return None, etc.\n    exit() # Or raise FileNotFoundError(\"Checkpoint not found\")\nexcept Exception as e:\n    print(f\"Error loading model weights: {e}\")\n    # Log the full traceback for debugging complex errors\n    import traceback\n    traceback.print_exc()\n    # Decide how to handle - exit(), raise error, return None, etc.\n    exit() # Or raise e\n\n# 5. Set the model to evaluation mode (important!)\nmodel.eval()\nprint(\"Model set to evaluation mode.\")\n\n# --- Now you can proceed with inference using the loaded model ---","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- Step 7: Inference on Test Set ---\n# -------------------------------------------------------------\nprint(\"\\n--- Starting Inference on Test Set ---\")\n\n# --- Load Best Model Weights (Use your corrected loading logic) ---\nprint(\"Loading best trained model weights...\")\n\ntry:\n    checkpoint = torch.load(CHECKPOINT_PATH, map_location=DEVICE) # Load the whole checkpoint\n    if 'model_state_dict' in checkpoint:\n        model.load_state_dict(checkpoint['model_state_dict'])\n        print(\"Best model weights loaded successfully from checkpoint.\")\n    else:\n        # Fallback if only state dict was saved (less likely with your saving code)\n        model.load_state_dict(checkpoint)\n        print(\"Checkpoint contained only state_dict, loaded successfully.\")\n    # Load best val loss achieved (optional, for reference)\n    loaded_best_val_loss = checkpoint.get('best_val_loss', float('inf'))\n    print(f\"(Best validation loss during training was: {loaded_best_val_loss:.4f})\")\n\nexcept FileNotFoundError:\n    print(f\"ERROR: Best model checkpoint not found at {CHECKPOINT_PATH}. Using model from last epoch (if available) or initial state.\")\n    # Handle appropriately - maybe exit or raise error if inference requires the best model\nexcept Exception as e:\n    print(f\"Error loading best model weights: {e}\")\n    # Handle appropriately\n\nmodel.eval() # Set model to evaluation mode\n\n# --- Create Test DataLoader HERE (after training) ---\nprint(\"\\nCreating Test Dataset and DataLoader...\")\ntry:\n    test_dataset = DashcamDataset(csv_file=TEST_CSV_PATH,\n                                  video_dir=TEST_VIDEO_DIR,\n                                  seq_len=SEQ_LEN,\n                                  transform=test_transform_final, # Use test transforms\n                                  mode='test',\n                                  apply_hflip=False) # No flip for test\n\n    test_loader = DataLoader(test_dataset,\n                             batch_size=BATCH_SIZE * 2, # Inference batch size\n                             shuffle=False,\n                             num_workers=NUM_WORKERS,\n                             pin_memory=True,\n                             prefetch_factor=PREFETCH_FACTOR,\n                             collate_fn=collate_fn)\n    print(f\"Test dataset size: {len(test_dataset)}\")\nexcept FileNotFoundError:\n    print(f\"ERROR: Test CSV not found at {TEST_CSV_PATH}. Cannot run inference.\")\n    test_loader = None\nexcept Exception as e:\n    print(f\"Error creating test dataset/loader: {e}\")\n    test_loader = None\n\n# --- Run Inference (only if test_loader exists) ---\npredictions = {}\nif test_loader:\n    with torch.no_grad():\n        test_progress_bar = tqdm(test_loader, desc=\"Inference\", leave=False)\n        for inputs_test, ids_test in test_progress_bar:\n            # Check for empty batch from collate_fn\n            if inputs_test is None or not inputs_test.numel():\n                print(f\"Warning: Skipping empty test batch.\")\n                continue\n\n            inputs_test = inputs_test.to(DEVICE)\n\n            with autocast(enabled=(DEVICE.type == 'cuda')):\n                outputs_test = model(inputs_test)\n\n            probs = torch.sigmoid(outputs_test).cpu().numpy().flatten()\n\n            for video_id, prob in zip(ids_test, probs):\n                predictions[str(video_id)] = prob\n    print(f\"Inference complete. Generated predictions for {len(predictions)} videos.\")\nelse:\n    print(\"Skipping inference as test loader was not created.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-16T06:39:39.209101Z","iopub.status.idle":"2025-04-16T06:39:39.209428Z","shell.execute_reply":"2025-04-16T06:39:39.209248Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- Step 8: Create Submission File ---\n# -------------------------------------------------------------\nprint(\"\\nCreating submission file...\")\ntry:\n    # Load sample submission to get all required test IDs in the correct order\n    sample_sub = pd.read_csv(os.path.join(BASE_DATA_PATH, 'sample_submission.csv'))\n    submission_df = pd.DataFrame({'id': sample_sub['id']}) # Use IDs from sample submission\n\n    # Map predictions - handle cases where a test video might have failed processing\n    submission_df['score'] = submission_df['id'].astype(str).map(predictions).fillna(0.5) # Fill missing with 0.5? Or maybe 0? Check competition baseline.\n\n    # Ensure scores are within [0, 1]\n    submission_df['score'] = submission_df['score'].clip(0.0, 1.0)\n\n    # Save submission file\n    submission_df.to_csv(SUBMISSION_CSV, index=False)\n    print(f\"Submission file saved to {SUBMISSION_CSV}\")\n    print(submission_df.head())\n\nexcept FileNotFoundError:\n     print(\"ERROR: sample_submission.csv not found. Cannot create submission file in correct format.\")\nexcept Exception as e:\n     print(f\"Error creating submission file: {e}\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#torch.save({'test': 1}, '/kaggle/working/test_save.pth')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-25T07:49:50.070147Z","iopub.execute_input":"2025-04-25T07:49:50.070461Z","iopub.status.idle":"2025-04-25T07:49:50.074791Z","shell.execute_reply.started":"2025-04-25T07:49:50.070437Z","shell.execute_reply":"2025-04-25T07:49:50.073890Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}