{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":92399,"databundleVersionId":11038207,"sourceType":"competition"}],"dockerImageVersionId":30919,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Below is a sample submission code that first runs inference on the training videos (to check model behavior on known data) and then on the test videos for final submission. This example uses a simple custom 3D CNN built from scratch.","metadata":{}},{"cell_type":"code","source":"import cv2\nimport numpy as np\nimport os\nimport pandas as pd\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-02-24T16:50:07.464833Z","iopub.execute_input":"2025-02-24T16:50:07.465070Z","iopub.status.idle":"2025-02-24T16:50:11.988923Z","shell.execute_reply.started":"2025-02-24T16:50:07.465047Z","shell.execute_reply":"2025-02-24T16:50:11.988243Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Define a Simple Custom 3D CNN Model\n\nThis simple model has two 3D convolutional layers with batch normalization and global pooling, ending in a fully connected layer with sigmoid activation.","metadata":{}},{"cell_type":"code","source":"# class CustomCrashModel(nn.Module):\n    \n#     def __init__(self):\n#         super(CustomCrashModel, self).__init__()\n\n#         # Input shape - (B, 3, T, H, W)\n#         self.conv1 = nn.Conv3d(in_channels = 3, out_channels = 8, kernel_size = 3, padding = 1)\n#         self.bn1 = nn.BatchNorm3d(8)\n#         self.pool1 = nn.MaxPool3d(kernel_size = (1, 2, 2))  # Pool spatial dimensions\n\n#         self.conv2 = nn.Conv3d(in_channels = 8, out_channels = 16, kernel_size = 3, padding = 1)\n#         self.bn2 = nn.BatchNorm3d(16)\n#         self.pool2 = nn.AdaptiveAvgPool3d(1)  # Global Pooling\n\n#         self.fc = nn.Linear(16, 1)\n\n#     def forward(self, x):\n#         # x - (B, 3, T, H, W)\n#         x = F.relu(self.bn1(self.conv1(x)))\n#         x = self.pool1(x)  # shape: (B, 8, T, H/2, W/2)\n#         x = F.relu(self.bn2(self.conv2(x)))\n#         x = self.pool2(x)  # shape: (B, 16, 1, 1, 1)\n#         x = x.view(x.size(0), -1)  # flatten to (B, 16)\n#         x = self.fc(x)\n\n#         return torch.sigmoid(x)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-24T16:50:11.989674Z","iopub.execute_input":"2025-02-24T16:50:11.990169Z","iopub.status.idle":"2025-02-24T16:50:11.993821Z","shell.execute_reply.started":"2025-02-24T16:50:11.990136Z","shell.execute_reply":"2025-02-24T16:50:11.992894Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"class SimpleCrashNN(nn.Module):\n    def __init__(self):\n        super(SimpleCrashNN, self).__init__()\n        self.conv1 = nn.Conv3d(in_channels=3, out_channels=8, kernel_size=3, padding=1)\n        self.bn1 = nn.BatchNorm3d(8)\n        self.conv2 = nn.Conv3d(in_channels=8, out_channels=16, kernel_size=3, padding=1)\n        self.bn2 = nn.BatchNorm3d(16)\n        self.pool = nn.AdaptiveAvgPool3d(1)\n        self.fc = nn.Linear(16, 1)\n    \n    def forward(self, x):\n        x = F.relu(self.bn1(self.conv1(x)))  # shape: (B, 8, T, H, W)\n        x = F.relu(self.bn2(self.conv2(x)))  # shape: (B, 16, T, H, W)\n        x = self.pool(x)                     # shape: (B, 16, 1, 1, 1)\n        x = x.view(x.size(0), -1)            # flatten to (B, 16)\n        x = self.fc(x)                       # shape: (B, 1)\n        return torch.sigmoid(x)              # probability in [0, 1]\n\n# Instantiate the model and move it to device\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nmodel = SimpleCrashNN().to(device)\nmodel.eval()  # Set model to evaluation mode\nprint(\"Simple custom NN initialized and set to eval mode.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-24T16:50:11.994716Z","iopub.execute_input":"2025-02-24T16:50:11.995037Z","iopub.status.idle":"2025-02-24T16:50:12.332317Z","shell.execute_reply.started":"2025-02-24T16:50:11.995005Z","shell.execute_reply":"2025-02-24T16:50:12.331427Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Load CSV Data (Training & Test)\n\nWe load the training CSV to verify our model on known data and the test CSV for final submission.","metadata":{}},{"cell_type":"code","source":"# Cell 3: Load training and test CSV files\ntrain_df = pd.read_csv('/kaggle/input/nexar-collision-prediction/train.csv')\ntest_df = pd.read_csv('/kaggle/input/nexar-collision-prediction/test.csv')\nprint(\"Training data loaded. Number of training videos:\", len(train_df))\nprint(\"Test data loaded. Number of test videos:\", len(test_df))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-24T16:50:12.333171Z","iopub.execute_input":"2025-02-24T16:50:12.333415Z","iopub.status.idle":"2025-02-24T16:50:12.352473Z","shell.execute_reply.started":"2025-02-24T16:50:12.333395Z","shell.execute_reply":"2025-02-24T16:50:12.351656Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Define Frame Extraction Function\n\nThis function extracts a fixed number of frames (e.g. 16) uniformly from a video, resizes them to 224×224, converts from BGR to RGB, normalizes the pixel values, and outputs a tensor of shape (B, C, T, H, W).","metadata":{}},{"cell_type":"code","source":"def extract_frames(video_path, num_frames=16, resize=(224, 224)):\n    cap = cv2.VideoCapture(video_path)\n    if not cap.isOpened():\n        print(f\"Error opening video file: {video_path}\")\n        return None\n\n    total_frames = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    frame_indices = np.linspace(0, total_frames - 1, num_frames, dtype=int)\n    \n    frames = []\n    for idx in range(total_frames):\n        ret, frame = cap.read()\n        if not ret:\n            break\n        if idx in frame_indices:\n            frame = cv2.resize(frame, resize)\n            frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)\n            frame = frame.astype(np.float32) / 255.0  # Normalize pixels to [0, 1]\n            frames.append(frame)\n    cap.release()\n    \n    if len(frames) < num_frames:\n        # Duplicate last frame if video is too short\n        while len(frames) < num_frames:\n            frames.append(frames[-1])\n    \n    # Convert frames to tensor and rearrange: (T, H, W, C) -> (B, C, T, H, W)\n    frames = np.stack(frames, axis=0)           # (T, H, W, C)\n    frames = np.transpose(frames, (3, 0, 1, 2))   # (C, T, H, W)\n    frames_tensor = torch.from_numpy(frames).unsqueeze(0)\n    return frames_tensor","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-24T16:50:12.353327Z","iopub.execute_input":"2025-02-24T16:50:12.353622Z","iopub.status.idle":"2025-02-24T16:50:12.359744Z","shell.execute_reply.started":"2025-02-24T16:50:12.353594Z","shell.execute_reply":"2025-02-24T16:50:12.359016Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Define Inference Function\n\nThis function extracts frames from a video file, runs the model, and returns a crash probability.","metadata":{}},{"cell_type":"code","source":"def predict_video(video_path):\n    frames_tensor = extract_frames(video_path, num_frames=16, resize=(224,224))\n    if frames_tensor is None:\n        return 0.0  # Default probability if video processing fails\n    frames_tensor = frames_tensor.to(device)\n    with torch.no_grad():\n        output = model(frames_tensor)  # Model output shape: (B, 1)\n        prob = output.item()\n    return prob","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-24T16:50:12.361817Z","iopub.execute_input":"2025-02-24T16:50:12.362066Z","iopub.status.idle":"2025-02-24T16:50:12.375729Z","shell.execute_reply.started":"2025-02-24T16:50:12.362046Z","shell.execute_reply":"2025-02-24T16:50:12.375061Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Run Inference on Training Data\n\nWe use the training CSV and corresponding video files (assumed to be in a \"train\" folder) to generate predictions. This step is useful for sanity checking and debugging.","metadata":{}},{"cell_type":"code","source":"# Cell 6: Generate predictions for training videos\ntrain_predictions = []\n\n# Training videos are stored in \"train/\" folder with filenames \"<id>.mp4\"\nfor idx, row in train_df.iterrows():\n    # Convert video ID to an integer and format with leading zeros (5 digits)\n    video_id = int(float(row['id']))\n    video_filename = f\"{video_id:05d}.mp4\"  # e.g., 01924.mp4\n    video_path = os.path.join(\"/kaggle/input/nexar-collision-prediction/train\", video_filename)\n    prob = predict_video(video_path)\n    train_predictions.append(prob)\n    if idx % 50 == 0:\n        print(f\"Processed {idx} training videos...\")\n\ntrain_df['predicted_score'] = train_predictions\nprint(\"Training predictions generated.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-24T16:51:42.057771Z","iopub.execute_input":"2025-02-24T16:51:42.058099Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Run Inference on Test Data\n\nWe now run the same inference function on the test videos (assumed to be in a \"test\" folder) to generate final predictions.","metadata":{}},{"cell_type":"code","source":"# Cell 7: Generate predictions for test videos\ntest_predictions = []\n\n# Test videos are stored in \"test/\" folder with filenames \"<id>.mp4\"\nfor idx, row in test_df.iterrows():\n    video_id = int(float(row['id']))\n    video_filename = f\"{video_id:05d}.mp4\"  # Format with 5 digits\n    video_path = os.path.join(\"/kaggle/input/nexar-collision-prediction/test\", video_filename)\n    prob = predict_video(video_path)\n    test_predictions.append(prob)\n    if idx % 50 == 0:\n        print(f\"Processed {idx} test videos...\")\n\ntest_df['score'] = test_predictions\nprint(\"Test predictions generated.\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Save Submission File\n\nWe create a submission CSV file from the test predictions with columns \"id\" and \"score\".","metadata":{}},{"cell_type":"code","source":"submission = test_df[['id', 'score']]\nsubmission.to_csv('submission.csv', index=False)\nprint(\"Submission file 'submission.csv' created successfully.\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Final Explanation & Next Steps\n\n1. **Training Data Inference:**  \n   - We first run our simple custom 3D CNN on training videos to generate predictions. This helps validate that our model pipeline works as expected.\n\n2. **Test Data Inference:**  \n   - We then run the same inference on test videos and create a submission file with \"id\" and \"score\" columns.\n\n3. **Submission:**  \n   - The generated submission CSV is ready to be uploaded to the leaderboard.\n\nReplace the simple custom model with your actual crash prediction model and adjust preprocessing as needed for improved performance. Happy Kaggling!","metadata":{}}]}