{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# RSNA Pneumonia Detection Challenge\n## Deep Learning Object Detection using Fast R-CNN and Faster R-CNN\n\nThis notebook implements and compares two state-of-the-art object detection models—**Fast R-CNN** and **Faster R-CNN**—for localizing lung opacities (pneumonia) in chest X-rays. We use the dataset provided by the **RSNA Pneumonia Detection Challenge**.\n\n### Workflow Overview:\n1. **Setup & Dependencies**: Load required medical and deep learning libraries.\n2. **Exploratory Data Analysis (EDA)**: Understand class distributions, label structures, and visualize DICOM X-ray scans with bounding box annotations.\n3. **Dataset & Preprocessing**: Develop a custom PyTorch `Dataset` that reads medical DICOM data, resizes images, and scales bounding boxes.\n4. **Fast R-CNN Model**: Simulate Fast R-CNN by freezing the Region Proposal Network (RPN) and backbone parameters, training only the classification/regression heads (RoI Heads). We also demonstrate OpenCV Selective Search on X-rays.\n5. **Faster R-CNN Model**: Instantiate and fine-tune a pre-trained Faster R-CNN model end-to-end.\n6. **Evaluation Metric**: Implement the official Kaggle RSNA evaluation metric—**Mean Average Precision (mAP) at IoU thresholds of 0.4 to 0.75**.\n7. **Training & Comparison**: Train both models, monitor training losses and validation mAP, and compare their performance.\n8. **Visualizations**: Overlay predicted bounding boxes vs. ground truth on validation images.\n9. **Final Report**: Analyze architectural trade-offs, training progress, and results.\n","metadata":{}},{"cell_type":"code","source":"# Install pydicom if not already present (pre-installed in Kaggle)\n!pip install -q pydicom\n\nimport os\nimport time\nimport random\nimport pydicom\nimport cv2\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport matplotlib.patches as patches\nfrom PIL import Image\nfrom sklearn.model_selection import train_test_split","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-25T11:22:00.395132Z","iopub.execute_input":"2026-06-25T11:22:00.395787Z","iopub.status.idle":"2026-06-25T11:22:03.532024Z","shell.execute_reply.started":"2026-06-25T11:22:00.395749Z","shell.execute_reply":"2026-06-25T11:22:03.531185Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import torch\nimport torch.nn as nn\nimport torch.optim as optim\nfrom torch.utils.data import Dataset, DataLoader\n\nimport torchvision\nfrom torchvision.models.detection import fasterrcnn_resnet50_fpn\nfrom torchvision.models.detection.faster_rcnn import FastRCNNPredictor\n\n# Set seeds for reproducibility\ndef set_seed(seed=42):\n    random.seed(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    if torch.cuda.is_available():\n        torch.cuda.manual_seed(seed)\n        torch.cuda.manual_seed_all(seed)\n        torch.backends.cudnn.deterministic = True\n        torch.backends.cudnn.benchmark = False\n\nset_seed(42)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-25T11:22:03.533859Z","iopub.execute_input":"2026-06-25T11:22:03.534212Z","iopub.status.idle":"2026-06-25T11:22:03.542560Z","shell.execute_reply.started":"2026-06-25T11:22:03.534182Z","shell.execute_reply":"2026-06-25T11:22:03.541716Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"class Config:\n    # Kaggle directory paths\n    DATA_DIR = '/kaggle/input/competitions/rsna-pneumonia-detection-challenge'\n    TRAIN_IMAGES_DIR = os.path.join(DATA_DIR, 'stage_2_train_images')\n    TRAIN_CSV = os.path.join(DATA_DIR, 'stage_2_train_labels.csv')\n    \n    # Model configuration\n    IMG_SIZE = 512  # Resize from 1024x1024 to save memory and train faster\n    BATCH_SIZE = 8\n    EPOCHS = 8\n    LR = 0.005\n    MOMENTUM = 0.9\n    WEIGHT_DECAY = 0.0005\n    DEVICE = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\n    \n    # Dataset subset size for demonstrating the pipeline\n    # Set to a number (e.g. 1500) to train and evaluate fast. Set to None to use full dataset.\n    SUBSET_SIZE = 1500\n    VAL_SPLIT = 0.2\n\nprint(f\"Using device: {Config.DEVICE}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-25T11:22:03.543740Z","iopub.execute_input":"2026-06-25T11:22:03.544072Z","iopub.status.idle":"2026-06-25T11:22:03.554150Z","shell.execute_reply.started":"2026-06-25T11:22:03.544035Z","shell.execute_reply":"2026-06-25T11:22:03.553355Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Load train labels CSV\ndf = pd.read_csv(Config.TRAIN_CSV)\nprint(f\"Total entries in CSV: {len(df)}\")\nprint(f\"Unique patients: {df['patientId'].nunique()}\")\nprint(df.head())\n\n# Count classes\n# Note: Target = 1 means pneumonia bounding box. Target = 0 means no pneumonia.\nclass_counts = df.groupby('patientId')['Target'].first().value_counts()\nprint(\"\\nClass Distribution (Unique Patients):\")\nprint(f\"Normal / No Lung Opacity (0): {class_counts.get(0, 0)}\")\nprint(f\"Lung Opacity / Pneumonia (1): {class_counts.get(1, 0)}\")\n\n# Plotting class distribution\nplt.figure(figsize=(6, 4))\nclass_counts.plot(kind='bar', color=['#3498db', '#e74c3c'])\nplt.title('Patient Class Distribution')\nplt.xlabel('Pneumonia Presence (Target)')\nplt.ylabel('Count')\nplt.xticks([0, 1], ['Normal/Other (0)', 'Pneumonia (1)'], rotation=0)\nplt.grid(axis='y', linestyle='--', alpha=0.7)\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-25T11:22:03.555747Z","iopub.execute_input":"2026-06-25T11:22:03.556032Z","iopub.status.idle":"2026-06-25T11:22:03.740712Z","shell.execute_reply.started":"2026-06-25T11:22:03.556011Z","shell.execute_reply":"2026-06-25T11:22:03.740157Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def visualize_sample(patient_id, df, img_dir):\n    dcm_path = os.path.join(img_dir, f\"{patient_id}.dcm\")\n    if not os.path.exists(dcm_path):\n        print(f\"File {dcm_path} not found. Skipping visualization.\")\n        return\n        \n    dcm_data = pydicom.dcmread(dcm_path)\n    img = dcm_data.pixel_array\n    \n    # Get bounding boxes for this patient\n    boxes = df[df['patientId'] == patient_id]\n    \n    fig, ax = plt.subplots(1, 1, figsize=(8, 8))\n    ax.imshow(img, cmap=plt.cm.bone)\n    ax.set_title(f\"Patient: {patient_id} | Pneumonia: {boxes.iloc[0]['Target']}\")\n    \n    for idx, row in boxes.iterrows():\n        if row['Target'] == 1:\n            x, y, w, h = row['x'], row['y'], row['width'], row['height']\n            # Draw rectangle\n            rect = patches.Rectangle((x, y), w, h, linewidth=2, edgecolor='r', facecolor='none')\n            ax.add_patch(rect)\n            ax.text(x, y-10, 'Pneumonia', color='r', weight='bold', fontsize=10)\n            \n    plt.axis('off')\n    plt.show()\n\n# Select a patient with lung opacity\npneumonia_patients = df[df['Target'] == 1]['patientId'].unique()\nif len(pneumonia_patients) > 0:\n    visualize_sample(pneumonia_patients[0], df, Config.TRAIN_IMAGES_DIR)\nelse:\n    print(\"No pneumonia cases available in the training images folder.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-25T11:22:03.741424Z","iopub.execute_input":"2026-06-25T11:22:03.741622Z","iopub.status.idle":"2026-06-25T11:22:04.182934Z","shell.execute_reply.started":"2026-06-25T11:22:03.741604Z","shell.execute_reply":"2026-06-25T11:22:04.181982Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"class RSNADataset(Dataset):\n    def __init__(self, patient_ids, df, img_dir, img_size=512, transforms=None):\n        self.patient_ids = list(patient_ids)\n        self.df = df\n        self.img_dir = img_dir\n        self.img_size = img_size\n        self.transforms = transforms\n        \n        # Pre-group boxes by patientId for faster retrieval\n        self.grouped = df.groupby('patientId')\n        \n    def __len__(self):\n        return len(self.patient_ids)\n        \n    def __getitem__(self, idx):\n        patient_id = self.patient_ids[idx]\n        dcm_path = os.path.join(self.img_dir, f\"{patient_id}.dcm\")\n        \n        # Read DICOM image\n        dcm_data = pydicom.dcmread(dcm_path)\n        img = dcm_data.pixel_array.astype(np.float32)\n        \n        # Normalize pixel values to [0, 1]\n        img = (img - img.min()) / (img.max() - img.min() + 1e-5)\n        \n        # Get original image shape\n        orig_h, orig_w = img.shape\n        \n        # Resize image\n        if self.img_size != orig_h or self.img_size != orig_w:\n            img_resized = cv2.resize(img, (self.img_size, self.img_size))\n            scale_x = self.img_size / orig_w\n            scale_y = self.img_size / orig_h\n        else:\n            img_resized = img\n            scale_x = 1.0\n            scale_y = 1.0\n            \n        # Convert image to 3 channels (required by torchvision pre-trained backbones)\n        img_3ch = np.stack([img_resized, img_resized, img_resized], axis=0) # [3, H, W]\n        img_tensor = torch.as_tensor(img_3ch, dtype=torch.float32)\n        \n        # Get bounding box labels\n        patient_df = self.grouped.get_group(patient_id)\n        \n        boxes = []\n        labels = []\n        \n        for _, row in patient_df.iterrows():\n            if row['Target'] == 1:\n                x = row['x'] * scale_x\n                y = row['y'] * scale_y\n                w = row['width'] * scale_x\n                h = row['height'] * scale_y\n                \n                # Convert to xmin, ymin, xmax, ymax\n                xmin = x\n                ymin = y\n                xmax = x + w\n                ymax = y + h\n                \n                # Ensure coordinates are within image boundaries\n                xmin = max(0.0, min(xmin, self.img_size - 1.0))\n                ymin = max(0.0, min(ymin, self.img_size - 1.0))\n                xmax = max(xmin + 1.0, min(xmax, self.img_size))\n                ymax = max(ymin + 1.0, min(ymax, self.img_size))\n                \n                boxes.append([xmin, ymin, xmax, ymax])\n                labels.append(1) # Class 1 for lung opacity\n                \n        if len(boxes) == 0:\n            # Background image (no pneumonia)\n            boxes = torch.zeros((0, 4), dtype=torch.float32)\n            labels = torch.zeros((0,), dtype=torch.int64)\n        else:\n            boxes = torch.as_tensor(boxes, dtype=torch.float32)\n            labels = torch.as_tensor(labels, dtype=torch.int64)\n            \n        target = {\n            'boxes': boxes,\n            'labels': labels,\n            'patient_id': patient_id\n        }\n        \n        return img_tensor, target\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-25T11:22:04.184052Z","iopub.execute_input":"2026-06-25T11:22:04.184294Z","iopub.status.idle":"2026-06-25T11:22:04.195126Z","shell.execute_reply.started":"2026-06-25T11:22:04.184273Z","shell.execute_reply":"2026-06-25T11:22:04.194371Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Get unique patient IDs and their targets\npatient_targets = df.groupby('patientId')['Target'].first().reset_index()\n\n# Apply subset size if defined to enable rapid training/verification\nif Config.SUBSET_SIZE is not None and Config.SUBSET_SIZE < len(patient_targets):\n    print(f\"Subsampling dataset to {Config.SUBSET_SIZE} unique patients...\")\n    patient_targets = patient_targets.sample(n=Config.SUBSET_SIZE, random_state=42).reset_index(drop=True)\nelse:\n    print(\"Using the complete RSNA dataset.\")\n    \n# Split into train and validation sets (stratified by target)\ntrain_ids, val_ids = train_test_split(\n    patient_targets['patientId'].values,\n    test_size=Config.VAL_SPLIT,\n    stratify=patient_targets['Target'].values,\n    random_state=42\n)\n\nprint(f\"Train set size: {len(train_ids)} | Validation set size: {len(val_ids)}\")\n\ntrain_dataset = RSNADataset(train_ids, df, Config.TRAIN_IMAGES_DIR, img_size=Config.IMG_SIZE)\nval_dataset = RSNADataset(val_ids, df, Config.TRAIN_IMAGES_DIR, img_size=Config.IMG_SIZE)\n\ndef collate_fn(batch):\n    return tuple(zip(*batch))\n\ntrain_loader = DataLoader(\n    train_dataset,\n    batch_size=Config.BATCH_SIZE,\n    shuffle=True,\n    num_workers=2,\n    collate_fn=collate_fn\n)\n\nval_loader = DataLoader(\n    val_dataset,\n    batch_size=Config.BATCH_SIZE,\n    shuffle=False,\n    num_workers=2,\n    collate_fn=collate_fn\n)\n\nprint(\"DataLoaders successfully configured!\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-25T11:22:04.196109Z","iopub.execute_input":"2026-06-25T11:22:04.196443Z","iopub.status.idle":"2026-06-25T11:22:04.236454Z","shell.execute_reply.started":"2026-06-25T11:22:04.196414Z","shell.execute_reply":"2026-06-25T11:22:04.235888Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def demonstrate_selective_search(patient_id, df, img_dir):\n    dcm_path = os.path.join(img_dir, f\"{patient_id}.dcm\")\n    if not os.path.exists(dcm_path):\n        print(\"DCM file not found locally. Skipping OpenCV Selective Search demo.\")\n        return\n        \n    dcm_data = pydicom.dcmread(dcm_path)\n    img = dcm_data.pixel_array.copy()\n    img = ((img - img.min()) / (img.max() - img.min() + 1e-5) * 255).astype(np.uint8)\n    \n    # Convert to 3-channel RGB\n    img_rgb = cv2.cvtColor(img, cv2.COLOR_GRAY2RGB)\n    \n    # OpenCV Selective Search\n    try:\n        ss = cv2.ximgproc.segmentation.createSelectiveSearchSegmentation()\n        ss.setBaseImage(img_rgb)\n        ss.switchToSelectiveSearchFast()\n        rects = ss.process()\n        \n        # Draw top 100 proposals\n        img_out = img_rgb.copy()\n        for i, rect in enumerate(rects[:100]):\n            x, y, w, h = rect\n            cv2.rectangle(img_out, (x, y), (x+w, y+h), (0, 255, 0), 1, cv2.LINE_AA)\n            \n        plt.figure(figsize=(10, 5))\n        plt.subplot(1, 2, 1)\n        plt.imshow(img_rgb)\n        plt.title(\"Original Gray Image\")\n        plt.axis('off')\n        \n        plt.subplot(1, 2, 2)\n        plt.imshow(img_out)\n        plt.title(\"Top 100 Selective Search Proposals\")\n        plt.axis('off')\n        plt.show()\n        print(f\"Selective Search generated {len(rects)} total proposals.\")\n        \n    except Exception as e:\n        print(\"OpenCV Selective Search is not pre-compiled in this environment.\")\n        print(\"Simulating proposal extraction: generating fixed grid proposals (like in anchor generation).\")\n        # Simple fallback mockup\n        img_out = img_rgb.copy()\n        h, w = img_rgb.shape[:2]\n        proposals = []\n        for grid_y in range(100, h-100, 150):\n            for grid_x in range(100, w-100, 150):\n                cv2.rectangle(img_out, (grid_x, grid_y), (grid_x+120, grid_y+150), (255, 165, 0), 2)\n                proposals.append([grid_x, grid_y, 120, 150])\n        \n        plt.figure(figsize=(10, 5))\n        plt.subplot(1, 2, 1)\n        plt.imshow(img_rgb)\n        plt.title(\"Original Gray Image\")\n        plt.axis('off')\n        \n        plt.subplot(1, 2, 2)\n        plt.imshow(img_out)\n        plt.title(\"Simulated Region Proposals (Grid anchors)\")\n        plt.axis('off')\n        plt.show()\n        print(f\"Simulated proposal generator created {len(proposals)} proposals.\")\n\nif len(pneumonia_patients) > 0:\n    demonstrate_selective_search(pneumonia_patients[0], df, Config.TRAIN_IMAGES_DIR)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-25T11:22:04.237237Z","iopub.execute_input":"2026-06-25T11:22:04.237497Z","iopub.status.idle":"2026-06-25T11:22:13.897239Z","shell.execute_reply.started":"2026-06-25T11:22:04.237468Z","shell.execute_reply":"2026-06-25T11:22:13.896520Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Fast R-CNN Setup\n\nIn **Fast R-CNN**, region proposals are computed offline (using methods like Selective Search) and fed directly into the model. The model does not optimize proposal generation.\n\nTo implement **Fast R-CNN** within the modern PyTorch/torchvision modular framework, we:\n1. Load the `fasterrcnn_resnet50_fpn` model.\n2. Modify the output classes to 2 (Background + Pneumonia).\n3. **Freeze the backbone** and **freeze the RPN (Region Proposal Network)** weights. This forces the RPN to act as a fixed, non-learnable proposal generator (analogous to the offline Selective Search algorithm in Fast R-CNN).\n4. Train only the final classification and bounding box regression layers (`roi_heads`), optimizing only the detection head.\n","metadata":{}},{"cell_type":"code","source":"def get_fast_rcnn_model():\n    # Load pretrained Faster R-CNN\n    try:\n        model = fasterrcnn_resnet50_fpn(weights='DEFAULT')\n    except:\n        model = fasterrcnn_resnet50_fpn(pretrained=True)\n        \n    # Get number of input features for the classifier\n    in_features = model.roi_heads.box_predictor.cls_score.in_features\n    \n    # Replace the box predictor with one that outputs 2 classes (Background + Pneumonia)\n    model.roi_heads.box_predictor = FastRCNNPredictor(in_features, num_classes=2)\n    \n    # FREEZE BACKBONE AND RPN (to simulate Fast R-CNN's fixed/external proposals)\n    for param in model.backbone.parameters():\n        param.requires_grad = False\n        \n    for param in model.rpn.parameters():\n        param.requires_grad = False\n        \n    return model\n\nfast_rcnn = get_fast_rcnn_model().to(Config.DEVICE)\n# Verify trainable parameters\ntrainable_params = sum(p.numel() for p in fast_rcnn.parameters() if p.requires_grad)\ntotal_params = sum(p.numel() for p in fast_rcnn.parameters())\nprint(f\"Fast R-CNN loaded successfully!\")\nprint(f\"Trainable Parameters: {trainable_params:,} / {total_params:,} ({trainable_params/total_params:.2%})\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-25T11:29:36.262684Z","iopub.execute_input":"2026-06-25T11:29:36.263488Z","iopub.status.idle":"2026-06-25T11:29:36.921199Z","shell.execute_reply.started":"2026-06-25T11:29:36.263450Z","shell.execute_reply":"2026-06-25T11:29:36.920317Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Faster R-CNN Setup\n\nIn **Faster R-CNN**, proposal generation is fully integrated into the network via the **Region Proposal Network (RPN)**, sharing convolutional features with the detection heads.\n\nTo train **Faster R-CNN**, we optimize the backbone, the RPN, and the RoI heads end-to-end, allowing the proposal network to learn specifically how to identify lung opacities.\n","metadata":{}},{"cell_type":"code","source":"def get_faster_rcnn_model():\n    # Load pretrained Faster R-CNN\n    try:\n        model = fasterrcnn_resnet50_fpn(weights='DEFAULT')\n    except:\n        model = fasterrcnn_resnet50_fpn(pretrained=True)\n        \n    # Get number of input features for the classifier\n    in_features = model.roi_heads.box_predictor.cls_score.in_features\n    \n    # Replace the box predictor with one that outputs 2 classes (Background + Pneumonia)\n    model.roi_heads.box_predictor = FastRCNNPredictor(in_features, num_classes=2)\n    \n    # Keep backbone and RPN trainable (Faster R-CNN end-to-end learning)\n    for param in model.parameters():\n        param.requires_grad = True\n        \n    return model\n\nfaster_rcnn = get_faster_rcnn_model().to(Config.DEVICE)\n# Verify trainable parameters\ntrainable_params_faster = sum(p.numel() for p in faster_rcnn.parameters() if p.requires_grad)\ntotal_params_faster = sum(p.numel() for p in faster_rcnn.parameters())\nprint(f\"Faster R-CNN loaded successfully!\")\nprint(f\"Trainable Parameters: {trainable_params_faster:,} / {total_params_faster:,} ({trainable_params_faster/total_params_faster:.2%})\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-25T11:30:25.312732Z","iopub.execute_input":"2026-06-25T11:30:25.313037Z","iopub.status.idle":"2026-06-25T11:30:25.901342Z","shell.execute_reply.started":"2026-06-25T11:30:25.313014Z","shell.execute_reply":"2026-06-25T11:30:25.900516Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Custom RSNA Evaluation Metric: Mean Average Precision (mAP)\n\nThe RSNA Pneumonia Detection Challenge calculates the **mean Average Precision (mAP)** over Intersection over Union (IoU) thresholds ranging from **0.4 to 0.75** with a step size of **0.05**.\nThe score is the average AP calculated for each image.\n\nWe implement the metric matching the competition specifications:\n","metadata":{}},{"cell_type":"code","source":"def calculate_iou(box1, box2):\n    # Bounding boxes format: [xmin, ymin, xmax, ymax]\n    x1 = max(box1[0], box2[0])\n    y1 = max(box1[1], box2[1])\n    x2 = min(box1[2], box2[2])\n    y2 = min(box1[3], box2[3])\n    \n    intersection = max(0.0, x2 - x1) * max(0.0, y2 - y1)\n    area1 = (box1[2] - box1[0]) * (box1[3] - box1[1])\n    area2 = (box2[2] - box2[0]) * (box2[3] - box2[1])\n    union = area1 + area2 - intersection\n    \n    if union == 0:\n        return 0.0\n    return intersection / union\n\ndef calculate_image_ap(gt_boxes, pred_boxes, pred_scores, iou_threshold):\n    # Calculate average precision for a single image at a single IoU threshold\n    num_gt = len(gt_boxes)\n    num_pred = len(pred_boxes)\n    \n    if num_gt == 0 and num_pred == 0:\n        return 1.0  # Perfect match for healthy patient (no boxes predicted, no boxes present)\n    if num_gt == 0 or num_pred == 0:\n        return 0.0  # False alarm or completely missed detection\n        \n    # Sort predictions by score descending\n    sorted_indices = sorted(range(num_pred), key=lambda k: pred_scores[k], reverse=True)\n    pred_boxes = [pred_boxes[idx] for idx in sorted_indices]\n    \n    matched_gt = set()\n    tp = 0\n    ap = 0.0\n    \n    for k, p_box in enumerate(pred_boxes):\n        best_iou = -1.0\n        best_gt_idx = -1\n        \n        for i, gt_box in enumerate(gt_boxes):\n            if i in matched_gt:\n                continue\n            iou = calculate_iou(p_box, gt_box)\n            if iou > best_iou:\n                best_iou = iou\n                best_gt_idx = i\n                \n        if best_iou >= iou_threshold:\n            matched_gt.add(best_gt_idx)\n            tp += 1\n            precision = tp / (k + 1)\n            ap += precision / num_gt\n            \n    return ap\n\ndef calculate_image_map(gt_boxes, pred_boxes, pred_scores):\n    # Mean AP across thresholds 0.4 to 0.75 in steps of 0.05\n    thresholds = [0.4, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75]\n    ap_sum = 0.0\n    for t in thresholds:\n        ap_sum += calculate_image_ap(gt_boxes, pred_boxes, pred_scores, t)\n    return ap_sum / len(thresholds)\n\ndef calculate_dataset_map(all_gt_boxes, all_pred_boxes, all_pred_scores):\n    # Calculate the mean AP over the entire validation dataset\n    map_scores = []\n    for gt, pred, scores in zip(all_gt_boxes, all_pred_boxes, all_pred_scores):\n        map_scores.append(calculate_image_map(gt, pred, scores))\n    return np.mean(map_scores)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-25T11:30:40.156485Z","iopub.execute_input":"2026-06-25T11:30:40.157097Z","iopub.status.idle":"2026-06-25T11:30:40.167096Z","shell.execute_reply.started":"2026-06-25T11:30:40.157065Z","shell.execute_reply":"2026-06-25T11:30:40.166401Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def train_one_epoch(model, optimizer, loader, device):\n    model.train()\n    epoch_loss = 0.0\n    classifier_loss_sum = 0.0\n    box_reg_loss_sum = 0.0\n    rpn_loss_sum = 0.0\n    \n    for images, targets in loader:\n        # Move tensors to device\n        images = list(image.to(device) for image in images)\n        targets_dev = []\n        for t in targets:\n            targets_dev.append({\n                'boxes': t['boxes'].to(device),\n                'labels': t['labels'].to(device)\n            })\n            \n        optimizer.zero_grad()\n        loss_dict = model(images, targets_dev)\n        losses = sum(loss for loss in loss_dict.values())\n        \n        losses.backward()\n        optimizer.step()\n        \n        epoch_loss += losses.item()\n        classifier_loss_sum += loss_dict.get('loss_classifier', torch.tensor(0.0)).item()\n        box_reg_loss_sum += loss_dict.get('loss_box_reg', torch.tensor(0.0)).item()\n        rpn_loss_sum += loss_dict.get('loss_objectness', torch.tensor(0.0)).item() + \\\n                        loss_dict.get('loss_rpn_box_reg', torch.tensor(0.0)).item()\n                        \n    num_batches = len(loader)\n    return {\n        'total': epoch_loss / num_batches,\n        'classifier': classifier_loss_sum / num_batches,\n        'box_reg': box_reg_loss_sum / num_batches,\n        'rpn': rpn_loss_sum / num_batches\n    }\n\n@torch.no_grad()\ndef evaluate_model(model, loader, device, score_threshold=0.05):\n    model.eval()\n    all_gt_boxes = []\n    all_pred_boxes = []\n    all_pred_scores = []\n    \n    for images, targets in loader:\n        images_dev = list(image.to(device) for image in images)\n        predictions = model(images_dev)\n        \n        for target, pred in zip(targets, predictions):\n            # Get Ground Truth\n            gt_boxes = target['boxes'].numpy().tolist()\n            all_gt_boxes.append(gt_boxes)\n            \n            # Get Predictions\n            pred_boxes = pred['boxes'].cpu().numpy()\n            pred_labels = pred['labels'].cpu().numpy()\n            pred_scores = pred['scores'].cpu().numpy()\n            \n            # Filter by confidence threshold and label (class 1: lung opacity)\n            valid_mask = (pred_labels == 1) & (pred_scores >= score_threshold)\n            filtered_boxes = pred_boxes[valid_mask].tolist()\n            filtered_scores = pred_scores[valid_mask].tolist()\n            \n            all_pred_boxes.append(filtered_boxes)\n            all_pred_scores.append(filtered_scores)\n            \n    # Compute Competition mAP\n    val_map = calculate_dataset_map(all_gt_boxes, all_pred_boxes, all_pred_scores)\n    return val_map\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-25T11:30:43.120747Z","iopub.execute_input":"2026-06-25T11:30:43.121034Z","iopub.status.idle":"2026-06-25T11:30:43.131868Z","shell.execute_reply.started":"2026-06-25T11:30:43.121010Z","shell.execute_reply":"2026-06-25T11:30:43.131055Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"=== Training Fast R-CNN (Frozen RPN & Backbone) ===\")\noptimizer_fast = optim.SGD(\n    [p for p in fast_rcnn.parameters() if p.requires_grad],\n    lr=Config.LR,\n    momentum=Config.MOMENTUM,\n    weight_decay=Config.WEIGHT_DECAY\n)\nlr_scheduler_fast = optim.lr_scheduler.StepLR(optimizer_fast, step_size=3, gamma=0.33)\n\nfast_history = []\nfor epoch in range(Config.EPOCHS):\n    start_time = time.time()\n    losses = train_one_epoch(fast_rcnn, optimizer_fast, train_loader, Config.DEVICE)\n    lr_scheduler_fast.step()\n    \n    # Evaluate\n    val_map = evaluate_model(fast_rcnn, val_loader, Config.DEVICE)\n    elapsed = time.time() - start_time\n    \n    fast_history.append({'epoch': epoch+1, 'loss': losses['total'], 'map': val_map, 'time': elapsed})\n    print(f\"Epoch {epoch+1}/{Config.EPOCHS} | Loss: {losses['total']:.4f} | Validation mAP: {val_map:.4f} | Time: {elapsed:.1f}s\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-25T11:30:56.087019Z","iopub.execute_input":"2026-06-25T11:30:56.087783Z","iopub.status.idle":"2026-06-25T11:48:31.487421Z","shell.execute_reply.started":"2026-06-25T11:30:56.087740Z","shell.execute_reply":"2026-06-25T11:48:31.486373Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"\\n=== Training Faster R-CNN (End-to-End Joint Optimization) ===\")\noptimizer_faster = optim.SGD(\n    [p for p in faster_rcnn.parameters() if p.requires_grad],\n    lr=Config.LR,\n    momentum=Config.MOMENTUM,\n    weight_decay=Config.WEIGHT_DECAY\n)\nlr_scheduler_faster = optim.lr_scheduler.StepLR(optimizer_faster, step_size=3, gamma=0.33)\n\nfaster_history = []\nfor epoch in range(Config.EPOCHS):\n    start_time = time.time()\n    losses = train_one_epoch(faster_rcnn, optimizer_faster, train_loader, Config.DEVICE)\n    lr_scheduler_faster.step()\n    \n    # Evaluate\n    val_map = evaluate_model(faster_rcnn, val_loader, Config.DEVICE)\n    elapsed = time.time() - start_time\n    \n    faster_history.append({'epoch': epoch+1, 'loss': losses['total'], 'map': val_map, 'time': elapsed})\n    print(f\"Epoch {epoch+1}/{Config.EPOCHS} | Loss: {losses['total']:.4f} | Validation mAP: {val_map:.4f} | Time: {elapsed:.1f}s\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-25T11:48:31.489435Z","iopub.execute_input":"2026-06-25T11:48:31.489687Z","iopub.status.idle":"2026-06-25T12:33:52.785497Z","shell.execute_reply.started":"2026-06-25T11:48:31.489659Z","shell.execute_reply":"2026-06-25T12:33:52.784503Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fast_df = pd.DataFrame(fast_history)\nfaster_df = pd.DataFrame(faster_history)\n\n# Plot Loss comparison\nplt.figure(figsize=(12, 5))\nplt.subplot(1, 2, 1)\nplt.plot(fast_df['epoch'], fast_df['loss'], label='Fast R-CNN (Frozen RPN)', color='#e74c3c', marker='o')\nplt.plot(faster_df['epoch'], faster_df['loss'], label='Faster R-CNN (Full Train)', color='#2ecc71', marker='s')\nplt.title('Training Loss Curves')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\nplt.grid(True)\n\n# Plot mAP comparison\nplt.subplot(1, 2, 2)\nplt.plot(fast_df['epoch'], fast_df['map'], label='Fast R-CNN (Frozen RPN)', color='#e74c3c', marker='o')\nplt.plot(faster_df['epoch'], faster_df['map'], label='Faster R-CNN (Full Train)', color='#2ecc71', marker='s')\nplt.title('Validation Mean Average Precision (mAP @ 0.4-0.75)')\nplt.xlabel('Epoch')\nplt.ylabel('mAP')\nplt.legend()\nplt.grid(True)\n\nplt.tight_layout()\nplt.show()\n\n# Display comparison table\ncomparison_data = {\n    'Metric': ['Final Loss', 'Best Validation mAP', 'Avg Epoch Time (s)', 'Total Parameters', 'Trainable Parameters'],\n    'Fast R-CNN': [\n        f\"{fast_df['loss'].iloc[-1]:.4f}\",\n        f\"{fast_df['map'].max():.4f}\",\n        f\"{fast_df['time'].mean():.1f}s\",\n        f\"{total_params:,}\",\n        f\"{trainable_params:,}\"\n    ],\n    'Faster R-CNN': [\n        f\"{faster_df['loss'].iloc[-1]:.4f}\",\n        f\"{faster_df['map'].max():.4f}\",\n        f\"{faster_df['time'].mean():.1f}s\",\n        f\"{total_params_faster:,}\",\n        f\"{trainable_params_faster:,}\"\n    ]\n}\ncomparison_table = pd.DataFrame(comparison_data)\nprint(\"\\n=== Comparative Performance Metrics ===\")\nprint(comparison_table.to_string(index=False))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-25T12:33:52.786852Z","iopub.execute_input":"2026-06-25T12:33:52.787575Z","iopub.status.idle":"2026-06-25T12:33:53.119433Z","shell.execute_reply.started":"2026-06-25T12:33:52.787544Z","shell.execute_reply":"2026-06-25T12:33:53.118817Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def visualize_predictions(model, dataset, device, num_samples=3, threshold=0.15):\n    model.eval()\n    fig, axes = plt.subplots(num_samples, 2, figsize=(12, 5 * num_samples))\n    if num_samples == 1:\n        axes = np.expand_dims(axes, axis=0)\n        \n    # Filter patient list for samples that contain pneumonia boxes to make comparison interesting\n    has_pneumonia = []\n    for i in range(len(dataset)):\n        _, target = dataset[i]\n        if len(target['boxes']) > 0:\n            has_pneumonia.append(i)\n            if len(has_pneumonia) >= num_samples * 2:\n                break\n                \n    # Pick a random selection of patients with pneumonia\n    if len(has_pneumonia) >= num_samples:\n        sample_indices = random.sample(has_pneumonia, num_samples)\n    else:\n        sample_indices = random.sample(range(len(dataset)), num_samples)\n        \n    for row_idx, dataset_idx in enumerate(sample_indices):\n        img, target = dataset[dataset_idx]\n        patient_id = target['patient_id']\n        \n        # Run inference\n        with torch.no_grad():\n            pred = model([img.to(device)])[0]\n            \n        pred_boxes = pred['boxes'].cpu().numpy()\n        pred_labels = pred['labels'].cpu().numpy()\n        pred_scores = pred['scores'].cpu().numpy()\n        \n        # Filter by threshold and class\n        valid_mask = (pred_labels == 1) & (pred_scores >= threshold)\n        pred_boxes = pred_boxes[valid_mask]\n        pred_scores = pred_scores[valid_mask]\n        \n        # Convert image tensor back to displayable numpy\n        img_np = img[0].numpy()\n        \n        # Plot Ground Truth\n        ax_gt = axes[row_idx, 0]\n        ax_gt.imshow(img_np, cmap='gray')\n        ax_gt.set_title(f\"GT: {patient_id}\")\n        ax_gt.axis('off')\n        for box in target['boxes'].numpy():\n            x1, y1, x2, y2 = box\n            rect = patches.Rectangle((x1, y1), x2-x1, y2-y1, linewidth=2, edgecolor='g', facecolor='none')\n            ax_gt.add_patch(rect)\n            \n        # Plot Predictions\n        ax_pred = axes[row_idx, 1]\n        ax_pred.imshow(img_np, cmap='gray')\n        ax_pred.set_title(f\"Prediction (Thresh={threshold})\")\n        ax_pred.axis('off')\n        for box, score in zip(pred_boxes, pred_scores):\n            x1, y1, x2, y2 = box\n            rect = patches.Rectangle((x1, y1), x2-x1, y2-y1, linewidth=2, edgecolor='r', facecolor='none')\n            ax_pred.add_patch(rect)\n            ax_pred.text(x1, y1-5, f\"{score:.2f}\", color='r', weight='bold', fontsize=8)\n            \n    plt.tight_layout()\n    plt.show()\n\nprint(\"=== Fast R-CNN Predictions Visualization ===\")\nvisualize_predictions(fast_rcnn, val_dataset, Config.DEVICE, num_samples=2)\n\nprint(\"=== Faster R-CNN Predictions Visualization ===\")\nvisualize_predictions(faster_rcnn, val_dataset, Config.DEVICE, num_samples=2)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-25T12:33:53.121075Z","iopub.execute_input":"2026-06-25T12:33:53.121290Z","iopub.status.idle":"2026-06-25T12:33:54.931788Z","shell.execute_reply.started":"2026-06-25T12:33:53.121269Z","shell.execute_reply":"2026-06-25T12:33:54.931007Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# RSNA Pneumonia Detection Report\n\nThis report discusses the architectural differences, training dynamics, performance outcomes, and design considerations comparing Fast R-CNN and Faster R-CNN for the RSNA Pneumonia Detection Challenge.\n\n---\n\n### 1. Model Architecture Comparison\n\n#### Fast R-CNN\n- **Region Proposal Generator**: Relies on an external, offline algorithm (e.g., Selective Search). Proposals are computed beforehand on the CPU, which creates a major performance bottleneck.\n- **Feature Sharing**: Extracted features are computed once per image using a deep CNN backbone. Region proposals are then projected onto the feature map to extract RoI features via RoI Pooling/Align.\n- **Training Strategy**: While the detection head (classification and regression layers) is trained end-to-end, the proposal generator is fixed and cannot adapt to the target dataset.\n- **Our Implementation**: In this notebook, we simulate Fast R-CNN by freezing the pre-trained backbone and Region Proposal Network (RPN) weights. This treats the RPN as a fixed, offline region proposal mechanism, optimizing only the downstream classification and localization heads (RoI heads) on pneumonia cases.\n\n#### Faster R-CNN\n- **Region Proposal Generator**: Integrates the **Region Proposal Network (RPN)** directly into the network. The RPN is a fully convolutional network that shares feature representation with the detection head.\n- **Anchor Boxes**: Uses multi-scale and multi-aspect ratio anchor boxes at each sliding window position on the feature map to generate candidate regions of interest.\n- **Joint Optimization**: The backbone, RPN, and RoI heads are optimized jointly using a multi-task loss function. This allows the proposal generator to learn features specifically relevant to lung opacities, rather than generic edges or textures.\n- **Our Implementation**: We fine-tuned the entire Faster R-CNN model end-to-end, enabling simultaneous optimization of proposal generation and object classification.\n\n---\n\n### 2. Training Process & Parameters\n\n- **Optimization**: Stochastic Gradient Descent (SGD) with a learning rate of `0.005`, momentum of `0.9`, and weight decay of `0.0005`.\n- **Learning Rate Scheduler**: Step learning rate scheduler decaying the learning rate by a factor of 0.33 every 3 epochs.\n- **Input Resolution**: Resized DICOM images from original `1024x1024` to `512x512` pixels to allow faster processing and prevent out-of-memory errors on GPU while scaling the coordinates of bounding boxes correspondingly.\n- **Batch Size**: 8 images per batch.\n- **Validation Split**: 20% of the dataset was reserved for validation, stratified by pneumonia targets to maintain class proportions.\n\n---\n\n### 3. Results, Comparison, and Key Takeaways\n\n#### Quantitative Results (Demonstration on Subset):\n- **Fast R-CNN (Frozen RPN)**: Fast R-CNN reaches a plateau in validation mAP quickly. Because the RPN weights are frozen, the proposal generator cannot adapt to target chest X-rays, leading to lower recall on complex, faint opacities.\n- **Faster R-CNN (Trainable RPN)**: Shows a steady increase in validation mAP across epochs, outperforming Fast R-CNN. Training the RPN allows the network to specifically suggest bounding boxes around lung opacity boundaries.\n- **Training Speed**: Fast R-CNN trains significantly faster per epoch because a large percentage of parameters (backbone + RPN) are frozen, meaning gradients are only calculated for the RoI Head. Faster R-CNN is slower to train but yields a substantially higher detection quality.\n\n#### Key Takeaways:\n1. **End-to-End Joint Learning**: Joint optimization of the RPN and RoI heads is crucial for domain-specific tasks like medical imaging where target bounding boxes represent diffuse visual anomalies (lung opacity) rather than sharp objects (e.g., cars or dogs).\n2. **Training Complexity vs. Accuracy**: If computational resources are severely constrained, using a pre-trained frozen proposal generator (Fast R-CNN) allows for extremely lightweight training of only the box predictor. However, for clinical reliability, Faster R-CNN's localization accuracy makes it the superior choice.\n","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}