{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":11848,"databundleVersionId":862157,"sourceType":"competition"},{"sourceId":9959,"sourceType":"datasetVersion","datasetId":6885},{"sourceId":674247,"sourceType":"modelInstanceVersion","isSourceIdPinned":true,"modelInstanceId":511045,"modelId":525732},{"sourceId":419401,"sourceType":"modelInstanceVersion","isSourceIdPinned":true,"modelInstanceId":341916,"modelId":363233}],"dockerImageVersionId":31192,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Histopathologic Cancer Detection Using Convolutional Neural Networks\n\n## Step 1: Problem Description and Data\n\nThe objective of this project is to perform binary classification of histopathologic image patches to determine whether a given tissue sample contains tumor tissue. The dataset consists of high-resolution color microscopy image patches, each labeled as either healthy (0) or tumor (1). Images are provided as fixed-size RGB patches, making the problem well-suited to convolutional neural network (CNN) architectures.\n\nThe dataset is large, comprising tens of thousands of labeled samples. However, visual differences between the two classes are often subtle and localized, appearing primarily as textural differences, variations in cellular density, and small-scale morphological irregularities. Because the data consists of images rather than text, this is a computer vision problem rather than one involving natural language processing. The key challenges include extracting meaningful spatial features, generalizing beyond fine-grained visual patterns, and handling class imbalance during training and evaluation.\n\n## Step 2: Exploratory Data Analysis (EDA)\n\nVisualizations produced will be saved to the plots folder when running the notebook. Results from our benchmark run can be found in the benchmark_plots folder.\n\nExploratory data analysis was conducted to understand the visual and statistical properties of the dataset and to guide modeling decisions. Visual inspection of representative samples from both classes revealed substantial overlap in appearance. While tumor patches often exhibited denser cellular texture and irregular structural patterns, many healthy patches displayed similar characteristics, suggesting that simple rule-based or low-level feature approaches would be insufficient for reliable classification.\n\nClass distribution analysis showed that healthy tissue samples were more prevalent than tumor samples, indicating class imbalance. This observation informed both the training strategy and evaluation methodology. In particular, balanced sampling during training and rank-based evaluation metrics such as ROC-AUC and precision–recall curves were later emphasized over raw accuracy.\n\nAdditional analyses included per-image pixel intensity histograms and color-channel distributions. These showed broadly similar global brightness and color statistics across classes, reinforcing the conclusion that discriminative information resides primarily in localized spatial structure rather than global pixel intensity or color composition. Edge detection and contour-based visualizations further highlighted differences in textural irregularity, particularly within tumor samples.\n\nBlob-size and contour-area distributions indicated that most structural features were clustered within a relatively narrow spatial scale. However, a small number of outliers were present at both smaller and larger scales. Although the dominant feature sizes appeared clustered, these outliers motivated exploration of multi-scale feature extraction, under the hypothesis that rare structural extremes could be disproportionately informative for tumor detection.\n\nIn addition to structural variability, considerable intra-class variation and inter-class ambiguity were observed. Several healthy patches exhibited dense cellular patterns similar to tumor tissue, while some tumor patches lacked obvious large-scale abnormalities. This ambiguity suggested that labels could not be perfectly separated by simple visual heuristics, reinforcing the need for models capable of capturing nuanced spatial relationships. Consequently, downstream analysis emphasized architectural robustness and rank-based evaluation metrics rather than reliance on fixed decision thresholds.\n\nBased on these findings, the analysis plan prioritized convolutional neural network architectures designed to capture fine local texture while remaining sensitive to broader spatial context.\n\n## Step 3: Model Architecture\n\nMultiple convolutional neural network architectures were implemented and compared to balance representational capacity, training stability, and computational efficiency. Architectural choices were guided by insights from exploratory analysis, particularly the importance of localized texture, limited spatial context within patches, and the presence of rare structural outliers.\n\n### Baseline Convolutional Network\n\nA baseline SimpleCNN was constructed using stacked 3×3 convolutional layers with ReLU activations and max pooling, followed by global average pooling and a linear classification head. This architecture leveraged standard convolutional inductive biases such as locality and translation invariance and served as a point of reference. While computationally efficient and stable to train, this model saturated early and demonstrated limited discriminative power, indicating underfitting.\n\n### ImprovedCNN: Deeper Learned Feature Hierarchies\n\nThe ImprovedCNN extended the baseline by increasing depth and incorporating batch normalization and SiLU activation functions. Batch normalization improved gradient flow and training stability, while SiLU activations reduced activation sparsity relative to ReLU. Spatial downsampling was performed using strided convolutions instead of max pooling, allowing the network to learn how spatial resolution should be reduced. Dropout was applied in the classification head to reduce overfitting risk. These changes significantly improved performance while maintaining fast convergence.\n\n### MultiScaleCNN: Risk-Aware Feature Extraction\n\nAlthough EDA suggested that most informative structures clustered within a dominant scale, a MultiScaleCNN architecture was explored to account for rare scale outliers. Parallel convolutional branches with different kernel sizes enabled simultaneous extraction of features at multiple spatial resolutions. While performance gains over the ImprovedCNN were modest, activation analysis confirmed that multiple branches contributed meaningfully. This architecture increased robustness to scale variability and improved interpretability.\n\n### Pretrained ResNet18: Transfer Learning\n\nA pretrained ResNet18 backbone was evaluated using transfer learning, with the final fully connected layer replaced by a binary classification head. Residual connections facilitated deeper feature hierarchies, while pretrained ImageNet weights provided strong low- and mid-level visual representations. In fast-demo mode, the backbone was frozen to reduce training time and mitigate overfitting. This approach produced the strongest overall results across all evaluated architectures.\n\n## Step 4: Results and Analysis\n\nModels will be saved to the models folder. Visualizations produced will be saved to the plots folder when running the notebook. Results from our benchmark run can be found in the benchmark_plots folder.\n\nModels were evaluated using validation ROC-AUC, accuracy, average precision, and diagnostic plots derived from predicted probabilities. In addition to final performance metrics, optimization behavior and learning dynamics were analyzed to better understand model behavior.\n\n### Optimization Behavior and Training Stability\n\nTraining and validation loss curves revealed distinct behaviors across architectures. The SimpleCNN exhibited smooth but early-saturating loss curves, suggesting underfitting driven by insufficient model capacity.\n\nHigher-capacity models such as the ImprovedCNN and pretrained ResNet18 demonstrated rapid decreases in training loss during early epochs. However, validation loss occasionally increased after initial improvements, and training and validation curves sometimes crossed. This pattern is characteristic of over-specialization, where improvements in training performance are no longer accompanied by improvements in generalization.\n\nThis effect was most pronounced when unfreezing the pretrained ResNet18 backbone. While training loss continued to decrease, validation loss increased, indicating overfitting likely caused by excessive modification of early convolutional layers that encoded broadly useful visual features. Freezing the backbone stabilized optimization and produced more consistent validation behavior.\n\nNotably, brief increases in validation loss did not necessarily correspond to degraded discrimination performance, highlighting the importance of examining multiple evaluation metrics.\n\n### Training Efficiency and Rapid Convergence\n\nAcross all high-performing models, near-peak validation AUC was achieved within three epochs. This rapid convergence demonstrates that architectural inductive bias and transfer learning played a more significant role than extended optimization. Short training runs reduced overfitting and enabled rapid experimentation, allowing empirical comparisons across architectures without sacrificing performance.\n\n### ROC Curve Analysis\n\nReceiver Operating Characteristic curves were used to evaluate discrimination performance across all decision thresholds. The pretrained ResNet18 produced the strongest ROC curve, with performance close to the top-left corner, indicating high separability. The ImprovedCNN followed closely, demonstrating that carefully designed custom architectures can approach the performance of transfer learning under constrained training budgets.\n\nROC-AUC proved particularly valuable due to its robustness to class imbalance and insensitivity to threshold selection, enabling fair comparison across models.\n\n### Precision–Recall Analysis\n\nPrecision–Recall curves provided complementary insight into performance on the minority tumor class. The pretrained ResNet18 achieved the highest average precision, maintaining strong recall without a significant loss in precision. The ImprovedCNN and MultiScaleCNN exhibited similar PR behavior, while the SimpleCNN showed a sharp precision decline at higher recall levels.\n\nPR analysis was especially informative in cases where validation loss increased. In several runs, PR curves remained stable despite rising validation loss, indicating that ranking performance was preserved even when probability calibration degraded.\n\n### Quantitative Comparison Summary\n\n\n| Model                  | Best Validation AUC | Final Validation Accuracy | Average Precision | Epochs |\n|:-----------------------|:--------------------|:--------------------------|:------------------|:--------|\n| Pretrained (ResNet18)  | 0.9076              | 0.8400                    | 0.8770            | 3 |\n| ImprovedCNN            | 0.8958              | 0.7715                    | 0.8480            | 3 |\n| MultiScaleCNN          | 0.8854              | 0.8065                    | 0.8428            | 3 |\n| SimpleCNN              | 0.8230              | 0.7330                    | 0.7132            | 3 |\n\n## Step 5: Conclusion\n\nThis project demonstrated that histopathologic cancer detection can be effectively addressed using convolutional neural networks with appropriate architectural design and training strategies. Transfer learning with a pretrained ResNet18 produced the strongest overall performance, achieving high discrimination accuracy with minimal training time. Custom architectures, particularly the ImprovedCNN, closed much of the performance gap while remaining efficient and interpretable.\n\nA central takeaway was the importance of training efficiency. By limiting optimization to a small number of epochs, high-quality results were obtained while reducing overfitting and enabling rapid experimentation. Architectural inductive bias, balanced sampling, and careful control of training dynamics proved more impactful than extended training or aggressive hyperparameter tuning.\n\nFuture work could explore selective layer unfreezing, adaptive learning-rate schedules, and ensemble strategies combining complementary architectures. Overall, the results highlight that principled modeling decisions and efficient experimentation are critical for success in medical image classification tasks.\n","metadata":{}},{"cell_type":"markdown","source":"## Supporting Code\n\nOur Python implements multiple approaches.  You can run all of these with the RUN_ALL_MODELS flag in the code. If RUN_ALL_MODELS is disabled, a single approach can be chosen with the MODEL_MODE string.\n\nPlots with EDA of the data will be shown at the start. If RUN_ALL_MODELS is set True, these will print for each model.  The multiscale mode is unique in that I also include a separation of scales plot.  Model performance will be conveyed at the end.  All plots will be saved to the results folder.  \n\nI keep training and the associated visualiations separate from our submission logic. Submission logic can be foudn in the code block which follows the implementation.","metadata":{}},{"cell_type":"code","source":"# Histopathologic Cancer Detection\n\nimport os, random, gc\nfrom pathlib import Path\nfrom contextlib import nullcontext\nfrom datetime import datetime\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom PIL import Image\nimport cv2\n\nimport torch\nimport torch.nn as nn\nimport torch.optim as optim\nfrom torch.utils.data import Dataset, DataLoader, WeightedRandomSampler\nimport torchvision.transforms as T\n\nfrom sklearn.model_selection import StratifiedShuffleSplit\nfrom sklearn.metrics import (\n    roc_auc_score, accuracy_score, roc_curve, confusion_matrix,\n    precision_recall_curve, average_precision_score\n)\nfrom tqdm.auto import tqdm\nimport shutil\n\n# ---------------------------------------------------------------------\n# Config\n# ---------------------------------------------------------------------\n\nDATA_ROOT  = Path(\"/kaggle/input/histopathologic-cancer-detection\")\nTRAIN_DIR  = DATA_ROOT / \"train\"\nLABELS_CSV = DATA_ROOT / \"train_labels.csv\"\n\n# Model selection\nMODEL_MODE = \"pretrained\"\nRUN_ALL_MODELS = True\nMODEL_GRID = [\"simple\", \"improved\", \"multiscale\", \"pretrained\"]\nSHOW_PLOTS=True\n\nFAST_DEMO  = True\n\nIMG_SIZE    = 96\nBATCH_SIZE  = 128\nLR          = 0.01\nMOMENTUM    = 0.9\nWEIGHT_DECAY= 1e-4\nEPOCHS_FULL = 9\nDEMO_EPOCHS = 3\nPATIENCE    = 3\nSEED        = 42\n\nDEMO_TRAIN = 8000\nDEMO_VAL   = 2000\n\nNUM_WORKERS = min(4, os.cpu_count() or 2)\n\nFIG_SINGLE = (6, 4)\nFIG_DOUBLE = (10, 4)\n\ndef fig_grid(rows, cols=3):\n    return (4 * cols, 3 * rows)\n\nplt.rcParams.update({\n    \"figure.dpi\": 110,\n    \"axes.titleweight\": \"bold\",\n    \"axes.titlesize\": 12,\n    \"axes.labelsize\": 10\n})\n\ntorch.manual_seed(SEED)\nnp.random.seed(SEED)\nrandom.seed(SEED)\n\nDEVICE = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\ntorch.backends.cudnn.benchmark = True\ngc.disable()\n\nif DEVICE.type == \"cuda\":\n    try:\n        autocast = torch.amp.autocast\n        scaler = torch.amp.GradScaler()\n    except Exception:\n        autocast = torch.cuda.amp.autocast\n        scaler = torch.cuda.amp.GradScaler()\nelse:\n    autocast = lambda *args, **kwargs: nullcontext()\n    scaler = None\n\n# Where to save artifacts\nOUT_ROOT = Path(\"/kaggle/working\")\nRESULTS_ROOT = OUT_ROOT / \"results\"\nMODELS_ROOT  = OUT_ROOT / \"models\"\nPLOTS_ROOT   = RESULTS_ROOT / \"plots\"\nRESULTS_ROOT.mkdir(parents=True, exist_ok=True)\nMODELS_ROOT.mkdir(parents=True, exist_ok=True)\nPLOTS_ROOT.mkdir(parents=True, exist_ok=True)\n\ndef save_fig(name, dpi=120):\n    path = PLOTS_ROOT / f\"{name}.png\"\n    plt.savefig(path, dpi=dpi, bbox_inches=\"tight\")\n    if SHOW_PLOTS:\n        plt.show()\n    plt.close()\n\n# ---------------------------------------------------------------------\n# Load labels\n# ---------------------------------------------------------------------\n\nlabels = pd.read_csv(LABELS_CSV)\n\n# ---------------------------------------------------------------------\n# EDA (unchanged)\n# ---------------------------------------------------------------------\n\ndef plot_class_samples(df, root, label, title, n=9, rows=3):\n    assert n % rows == 0, \"n must be divisible by rows\"\n    cols = n // rows\n\n    subset = df[df[\"label\"] == label].sample(n, random_state=SEED)\n\n    fig, axes = plt.subplots(rows, cols, figsize=(2*cols, 2*rows))\n    axes = axes.flatten()\n\n    for ax, (_, r) in zip(axes, subset.iterrows()):\n        img = Image.open(root / f\"{r.id}.tif\")\n        ax.imshow(img)\n        ax.axis(\"off\")\n\n    fig.suptitle(title, fontsize=14, fontweight=\"bold\", y=1)\n    plt.tight_layout()\n    save_fig(f\"samples_label_{label}.png\")\n\nplot_class_samples(\n    labels,\n    TRAIN_DIR,\n    label=0,\n    title=\"Healthy Tissue Patches (Binary Class = 0)\"\n)\n\nplot_class_samples(\n    labels,\n    TRAIN_DIR,\n    label=1,\n    title=\"Tumor Tissue Patches (Binary Class = 1)\"\n)\n\ndef plot_edge_samples(df, root, n=9, rows=3, min_std=10):\n    \"\"\"\n    Displays samples as a patch-style grid:\n    Original | Grayscale | Edges\n    \"\"\"\n    assert n % rows == 0, \"n must be divisible by rows\"\n    cols = 3  # Original | Grayscale | Edges\n\n    samples = []\n    while len(samples) < n:\n        r = df.sample(1).iloc[0]\n        img = cv2.imread(str(root / f\"{r.id}.tif\"))\n        if img is None:\n            continue\n\n        gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)\n        if gray.std() < min_std:\n            continue\n\n        edges = cv2.Canny(gray, 50, 150)\n        samples.append((img, gray, edges))\n\n    fig, axes = plt.subplots(\n        rows, cols,\n        figsize=(2 * cols, 2 * rows),\n        sharex=True,\n        sharey=True\n    )\n    \n    fig.suptitle(\n        \"Edge Structures in Patches\",\n        fontsize=14,\n        fontweight=\"bold\",\n        y=1\n    )\n    \n    for r in range(rows):\n        img, gray, edges = samples[r]\n    \n        axes[r, 0].imshow(cv2.cvtColor(img, cv2.COLOR_BGR2RGB))\n        axes[r, 1].imshow(gray, cmap=\"gray\")\n        axes[r, 2].imshow(edges, cmap=\"gray\")\n    \n        if r == 0:\n            axes[r, 0].set_title(\"Original\", fontweight=\"normal\")\n            axes[r, 1].set_title(\"Grayscale\", fontweight=\"normal\")\n            axes[r, 2].set_title(\"Edges\", fontweight=\"normal\")\n    \n        for c in range(3):\n            axes[r, c].axis(\"off\")\n    \n    plt.tight_layout()\n    plt.subplots_adjust(top=0.9)\n    save_fig(\"edge_structures.png\")\n\nplot_edge_samples(labels, TRAIN_DIR, n=9, rows=3)\n\n# Class balance\nclass_counts_full = labels[\"label\"].value_counts().sort_index()\ncolors = [\"#f4c430\", \"#c0392b\"]\nplt.figure(figsize=FIG_SINGLE)\nbars = plt.bar(\n    [\"0 (Healthy)\", \"1 (Tumor)\"],\n    class_counts_full.values,\n    color=colors\n)\nplt.xlabel(\"Binary Class\")\nplt.ylabel(\"Number of Samples\")\nplt.suptitle(\n    \"Initial Binary Classification Distribution\",\n    fontsize=14,\n    fontweight=\"bold\",\n    y=1\n)\nplt.margins(y=0.2)\nfor bar in bars:\n    height = bar.get_height()\n    plt.text(\n        bar.get_x() + bar.get_width() / 2,\n        height,\n        f\"{int(height)}\",\n        ha=\"center\",\n        va=\"bottom\",\n        fontsize=10,\n        fontweight=\"bold\"\n    )\nplt.tight_layout()\nsave_fig(\"class_balance.png\")\n\n# Per-image mean intensity\nsample_means = []\nfor _, row in labels.sample(2000, random_state=SEED).iterrows():\n    img = np.asarray(Image.open(TRAIN_DIR / f\"{row.id}.tif\")).astype(np.float32)\n    sample_means.append(img.mean() / 255.0)\nplt.figure(figsize=FIG_SINGLE)\nplt.hist(sample_means, bins=40, color=\"tab:blue\", edgecolor=\"black\", alpha=0.85)\nplt.xlabel(\"Mean Pixel Intensity\")\nplt.ylabel(\"Number of Samples\")\nplt.suptitle(\n    \"Per-Image Mean Pixel Intensity\",\n    fontsize=14,\n    fontweight=\"bold\",\n    y=1\n)\nplt.tight_layout()\nsave_fig(\"mean_intensity.png\")\n\ndef plot_color_space(df, root, n=2000):\n    sample = df.sample(n, random_state=SEED)\n    hues, sats = [], []\n    for _, r in sample.iterrows():\n        img = cv2.imread(str(root / f\"{r.id}.tif\"))\n        if img is None:\n            continue\n        hsv = cv2.cvtColor(img, cv2.COLOR_BGR2HSV)\n        h, s, _ = cv2.split(hsv)\n        hues.append(h.mean())\n        sats.append(s.mean())\n\n    plt.figure(figsize=FIG_SINGLE)\n    plt.scatter(hues, sats, s=6, alpha=0.4)\n    plt.xlabel(\"Mean Hue\")\n    plt.ylabel(\"Mean Saturation\")\n    plt.suptitle(\n        \"Per Patch Color Distribution (HSV Space)\",\n        fontsize=14,\n        fontweight=\"bold\",\n        y=1\n    )\n    plt.grid(alpha=0.3)\n    save_fig(\"color_space_hs.png\")\n\nplot_color_space(labels, TRAIN_DIR)\n\ndef plot_blob_sizes(df, root, n=300, min_area=30):\n    sizes = []\n    for _, r in df.sample(n, random_state=SEED).iterrows():\n        img = cv2.imread(str(root / f\"{r.id}.tif\"), cv2.IMREAD_GRAYSCALE)\n        if img is None:\n            continue\n        edges = cv2.Canny(img, 50, 150)\n        contours, _ = cv2.findContours(edges, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)\n        for cnt in contours:\n            area = cv2.contourArea(cnt)\n            if area > min_area:\n                sizes.append(area)\n    plt.figure(figsize=FIG_SINGLE)\n    plt.hist(sizes, bins=50, edgecolor=\"black\", alpha=0.8)\n    plt.xscale(\"log\")\n    plt.xlabel(\"Blob Area (Pixels)\")\n    plt.ylabel(\"Number of Samples\")\n    plt.suptitle(\n        \"Structure Size as Identified by Edge Detection\",\n        fontsize=14,\n        fontweight=\"bold\",\n        y=1\n    )\n    save_fig(\"blob_sizes.png\")\n\nplot_blob_sizes(labels, TRAIN_DIR)\n\ndef plot_rgb_means(df, root, n=1500):\n    sample = df.sample(n, random_state=SEED)\n    rgb_means = []\n    for _, r in sample.iterrows():\n        img = np.asarray(Image.open(root / f\"{r.id}.tif\")) / 255.0\n        rgb_means.append(img.mean(axis=(0,1)))\n    rgb_means = np.array(rgb_means)\n    plt.figure(figsize=FIG_SINGLE)\n    plt.hist(rgb_means[:,0], bins=40, alpha=0.6, label=\"Red\")\n    plt.hist(rgb_means[:,1], bins=40, alpha=0.6, label=\"Green\")\n    plt.hist(rgb_means[:,2], bins=40, alpha=0.6, label=\"Blue\")\n    plt.xlabel(\"Mean Channel Intensity\")\n    plt.ylabel(\"Number of Samples\")\n    plt.suptitle(\n        \"Per-Image Mean RGB Channel Intensities\",\n        fontsize=14,\n        fontweight=\"bold\",\n        y=1\n    )\n    plt.legend()\n    save_fig(\"rgb_means.png\")\n\nplot_rgb_means(labels, TRAIN_DIR)\n\n# ---------------------------------------------------------------------\n# Split data\n# ---------------------------------------------------------------------\n\nsss = StratifiedShuffleSplit(n_splits=1, test_size=0.15, random_state=SEED)\ntrain_idx, val_idx = next(sss.split(labels[[\"id\"]], labels[\"label\"]))\ntrain_df = labels.iloc[train_idx].reset_index(drop=True)\nval_df   = labels.iloc[val_idx].reset_index(drop=True)\nif FAST_DEMO:\n    train_df = train_df.sample(DEMO_TRAIN, random_state=SEED)\n    val_df   = val_df.sample(DEMO_VAL, random_state=SEED)\n\nepochs = DEMO_EPOCHS if FAST_DEMO else EPOCHS_FULL\n\nprint(\n    f\"Training Mode: {'Quick' if FAST_DEMO else 'Thorough'}, \"\n    f\"Device: {DEVICE}, \"\n    f\"Training Samples: {len(train_df)}, \"\n    f\"Validation Samples: {len(val_df)}\"\n)\n\n# ---------------------------------------------------------------------\n# Dataset + transforms\n# ---------------------------------------------------------------------\n\nclass PCamDataset(Dataset):\n    def __init__(self, df, root, transform):\n        self.paths = [Path(root) / f\"{i}.tif\" for i in df.id.values]\n        self.labels = torch.tensor(df.label.values, dtype=torch.float32)\n        self.transform = transform\n\n    def __len__(self): \n        return len(self.labels)\n\n    def __getitem__(self, i):\n        img = Image.open(self.paths[i]).convert(\"RGB\")\n        return self.transform(img), self.labels[i]\n\nif FAST_DEMO:\n    train_tfms = T.Compose([\n        T.RandomHorizontalFlip(p=0.5),\n        T.Resize((IMG_SIZE, IMG_SIZE)),\n        T.ToTensor(),\n        T.Normalize([0.485,0.456,0.406],[0.229,0.224,0.225]),\n    ])\nelse:\n    train_tfms = T.Compose([\n        T.RandomHorizontalFlip(p=0.5),\n        T.RandomVerticalFlip(p=0.5),\n        T.RandomRotation(degrees=15),\n        T.ColorJitter(brightness=0.1, contrast=0.1),\n        T.Resize((IMG_SIZE, IMG_SIZE)),\n        T.ToTensor(),\n        T.Normalize([0.485,0.456,0.406],[0.229,0.224,0.225]),\n    ])\n\nval_tfms = T.Compose([\n    T.Resize((IMG_SIZE, IMG_SIZE)),\n    T.ToTensor(),\n    T.Normalize([0.485,0.456,0.406],[0.229,0.224,0.225]),\n])\n\n# Balanced sampler + pos_weight (unchanged semantics)\nlabels_np = train_df[\"label\"].values\nclass_counts = np.bincount(labels_np)\nclass_weights = 1.0 / class_counts\nsample_weights = class_weights[labels_np]\n\nsampler = WeightedRandomSampler(\n    weights=sample_weights,\n    num_samples=len(sample_weights),\n    replacement=True\n)\n\nneg, pos = class_counts\ncriterion = nn.BCEWithLogitsLoss(\n    pos_weight=torch.tensor([neg / pos], device=DEVICE)\n)\n\ntrain_loader = DataLoader(\n    PCamDataset(train_df, TRAIN_DIR, train_tfms),\n    batch_size=BATCH_SIZE,\n    sampler=sampler,\n    num_workers=NUM_WORKERS,\n    pin_memory=(DEVICE.type == \"cuda\"),\n    persistent_workers=NUM_WORKERS > 0\n)\nval_loader = DataLoader(\n    PCamDataset(val_df, TRAIN_DIR, val_tfms),\n    batch_size=BATCH_SIZE, shuffle=False,\n    num_workers=NUM_WORKERS, pin_memory=True, persistent_workers=True\n)\n\n# ---------------------------------------------------------------------\n# Models (Simple + Improved + MultiScale + ResNet18)\n# ---------------------------------------------------------------------\n\nclass SimpleCNN(nn.Module):\n    def __init__(self):\n        super().__init__()\n        self.features = nn.Sequential(\n            nn.Conv2d(3,32,3,padding=1), nn.ReLU(), nn.MaxPool2d(2),\n            nn.Conv2d(32,64,3,padding=1), nn.ReLU(), nn.MaxPool2d(2),\n            nn.Conv2d(64,128,3,padding=1), nn.ReLU(), nn.MaxPool2d(2),\n            nn.AdaptiveAvgPool2d(1)\n        )\n        self.fc = nn.Linear(128,1)\n\n    def forward(self,x):\n        x = self.features(x)\n        return self.fc(x.flatten(1))\n\nWEIGHTS_ROOT = Path(\"/kaggle/input/resnet-18/pytorch/default\")\n\n# Find the version directory (e.g., \"1\", \"2\", etc.)\nversion_dirs = [p for p in WEIGHTS_ROOT.iterdir() if p.is_dir()]\n\nassert len(version_dirs) > 0, \"No ResNet version directories found\"\n\nversion_dir = sorted(version_dirs)[-1]  # take latest\nRESNET18_WEIGHTS = version_dir / \"ResNet-18.pth\"\n\nprint(\"Using ResNet weights at:\", RESNET18_WEIGHTS)\nprint(\"Exists:\", RESNET18_WEIGHTS.exists())\nHAS_PRETRAINED = RESNET18_WEIGHTS.is_file()\nprint(\"HAS_PRETRAINED\", HAS_PRETRAINED)\n\nclass ResNet18Binary(nn.Module):\n    def __init__(self, weights_path=None):\n        super().__init__()\n        from torchvision.models import resnet18\n\n        m = resnet18(weights=None)\n\n        if weights_path is not None:\n            print(f\"Loading ResNet18 backbone from {weights_path}\")\n            state = torch.load(weights_path, map_location=\"cpu\")\n            state = {\n                k: v for k, v in state.items()\n                if not k.startswith(\"fc.\")\n            }\n            missing, unexpected = m.load_state_dict(state, strict=False)\n            print(\"Loaded backbone.\")\n            print(\"Missing keys:\", missing)\n            print(\"Unexpected keys:\", unexpected)\n\n        self.features = nn.Sequential(*list(m.children())[:-1])\n        self.fc = nn.Linear(512, 1)\n\n    def forward(self, x):\n        return self.fc(self.features(x).flatten(1))\n\nclass ImprovedCNN(nn.Module):\n    def __init__(self):\n        super().__init__()\n\n        def block(in_ch, out_ch, stride=1):\n            return nn.Sequential(\n                nn.Conv2d(in_ch, out_ch, 3, stride=stride, padding=1, bias=False),\n                nn.BatchNorm2d(out_ch),\n                nn.SiLU(),\n            )\n\n        self.features = nn.Sequential(\n            block(3, 32),\n            block(32, 32, stride=2),\n            block(32, 64),\n            block(64, 64, stride=2),\n            block(64, 128),\n            block(128, 128, stride=2),\n            block(128, 256),\n            nn.AdaptiveAvgPool2d(1)\n        )\n\n        self.head = nn.Sequential(\n            nn.Flatten(),\n            nn.Linear(256, 128),\n            nn.BatchNorm1d(128),\n            nn.SiLU(),\n            nn.Dropout(0.3),\n            nn.Linear(128, 1)\n        )\n\n    def forward(self, x):\n        x = self.features(x)\n        return self.head(x)\n\nclass MultiScaleCNN(nn.Module):\n    def __init__(self):\n        super().__init__()\n\n        def branch(in_ch, out_ch, k):\n            return nn.Sequential(\n                nn.Conv2d(in_ch, out_ch, kernel_size=k, padding=k//2, bias=False),\n                nn.BatchNorm2d(out_ch),\n                nn.SiLU(),\n            )\n\n        self.stage1 = nn.ModuleDict({\n            \"1x1\": branch(3, 16, 1),\n            \"3x3\": branch(3, 32, 3),\n            \"5x5\": branch(3, 16, 5),\n        })\n\n        self.stage2 = nn.ModuleDict({\n            \"1x1\": branch(64, 32, 1),\n            \"3x3\": branch(64, 64, 3),\n            \"5x5\": branch(64, 32, 5),\n        })\n\n        self.pool = nn.AdaptiveAvgPool2d(1)\n        self.head = nn.Linear(128, 1)\n\n    def forward(self, x, return_activations=False):\n        activations = {}\n\n        for i, stage in enumerate([self.stage1, self.stage2], start=1):\n            outs = {}\n            for k, layer in stage.items():\n                o = layer(x)\n                outs[k] = o\n                if return_activations:\n                    activations[f\"s{i}_{k}\"] = o\n\n            x = torch.cat(list(outs.values()), dim=1)\n            x = torch.nn.functional.avg_pool2d(x, 2)\n\n        x = self.pool(x).flatten(1)\n        y = self.head(x)\n\n        if return_activations:\n            return y, activations\n        return y\n\ndef build_model(mode):\n    if mode == \"simple\":\n        return SimpleCNN()\n\n    elif mode == \"improved\":\n        return ImprovedCNN()\n\n    elif mode == \"pretrained\":\n        weights = RESNET18_WEIGHTS if HAS_PRETRAINED else None\n        model = ResNet18Binary(weights_path=weights)\n\n        # Investigate the merit of freezing weights with low epoch count in the future.\n        # if FAST_DEMO and weights is not None:\n        #     for p in model.features.parameters():\n        #         p.requires_grad = False\n\n        return model\n\n    elif mode == \"multiscale\":\n        return MultiScaleCNN()\n\n    else:\n        raise ValueError(f\"Unknown MODEL_MODE: {mode}\")\n\n# ---------------------------------------------------------------------\n# Training helpers (very close to original behavior)\n# ---------------------------------------------------------------------\n\ndef run_epoch(model, optimizer, loader, train=True):\n    model.train(train)\n    total, preds, trues = 0.0, [], []\n\n    for x,y in tqdm(loader, leave=False):\n        x = x.to(DEVICE, memory_format=torch.channels_last, non_blocking=True)\n        y = y.to(DEVICE).unsqueeze(1)\n\n        with autocast(\"cuda\"):\n            out = model(x)\n            loss = criterion(out, y)\n\n        if train:\n            optimizer.zero_grad(set_to_none=True)\n            if scaler:\n                scaler.scale(loss).backward()\n                scaler.step(optimizer)\n                scaler.update()\n            else:\n                loss.backward(); optimizer.step()\n\n        total += loss.item() * x.size(0)\n        preds.append(torch.sigmoid(out).detach().cpu())\n        trues.append(y.detach().cpu())\n\n    preds = torch.cat(preds).numpy()\n    trues = torch.cat(trues).numpy()\n    auc = roc_auc_score(trues, preds)\n    acc = accuracy_score(trues>0.5, preds>0.5)\n    return total/len(loader.dataset), auc, acc, preds, trues\n\ndef collect_scale_stats(model, loader, max_batches=20):\n    from collections import defaultdict\n    model.eval()\n    stats = defaultdict(list)\n    with torch.no_grad():\n        for i, (x, _) in enumerate(loader):\n            if i >= max_batches:\n                break\n            x = x.to(DEVICE)\n            _, acts = model(x, return_activations=True)\n            for k, v in acts.items():\n                stats[k].append(v.abs().mean().item())\n    return {k: np.mean(v) for k, v in stats.items()}\n\n# ---------------------------------------------------------------------\n# One full experiment (train + plots + saving)\n# ---------------------------------------------------------------------\n\ndef run_experiment(mode: str):\n    print(\"\\n\" + \"=\"*70)\n    print(f\"Running model: {mode}\")\n    print(\"=\"*70)\n\n    model = build_model(mode).to(DEVICE, memory_format=torch.channels_last)\n\n    optimizer = optim.SGD(\n        filter(lambda p: p.requires_grad, model.parameters()),\n        lr=LR, momentum=MOMENTUM, weight_decay=WEIGHT_DECAY\n    )\n    if mode == \"pretrained\":\n        for g in optimizer.param_groups:\n            g[\"lr\"] = LR * 0.1\n\n    history = {\"train_loss\":[], \"train_auc\":[], \"val_loss\":[], \"val_auc\":[], \"val_acc\":[]}\n    best_auc = 0.0\n    best_state = None\n    patience = 0\n\n    for e in range(epochs):\n        print(f\"\\nEpoch {e+1}/{epochs}\")\n        tl, ta, _, _, _ = run_epoch(model, optimizer, train_loader, True)\n        vl, va, vac, vp, vt = run_epoch(model, optimizer, val_loader, False)\n\n        history[\"train_loss\"].append(tl)\n        history[\"train_auc\"].append(ta)\n        history[\"val_loss\"].append(vl)\n        history[\"val_auc\"].append(va)\n        history[\"val_acc\"].append(vac)\n\n        print(f\"train_loss={tl:.4f} train_auc={ta:.4f}\")\n        print(f\"val_loss={vl:.4f} val_auc={va:.4f} val_acc={vac:.4f}\")\n\n        if va > best_auc:\n            best_auc = va\n            best_state = model.state_dict()\n            patience = 0\n        else:\n            patience += 1\n            if patience >= PATIENCE:\n                print(\"Early stopping\")\n                break\n\n    print(\"Best Validation AUC:\", best_auc)\n\n    # Save model\n    model_path = MODELS_ROOT / f\"pcam_{mode}_best.pth\"\n    torch.save(best_state, model_path)\n    print(f\"Saved model to {model_path}\")\n\n    # Visualizations (as before)\n    epochs_range = range(1, len(history[\"train_loss\"]) + 1)\n\n    # Loss\n    plt.figure(figsize=FIG_SINGLE)\n    plt.plot(epochs_range, history[\"train_loss\"], label=\"Training\")\n    plt.plot(epochs_range, history[\"val_loss\"], label=\"Validation\")\n    plt.suptitle(\"Optimization Performance\", fontsize=14, fontweight=\"bold\", y=1)\n    plt.xlabel(\"Epoch\")\n    plt.ylabel(\"Binary Cross-Entropy Loss\")\n    plt.legend()\n    plt.grid(alpha=0.3)\n    plt.figtext(\n        0.5, -0.15,\n        \"Lower loss indicates improved optimization of model parameters.\",\n        ha=\"center\", fontsize=9\n    )\n    save_fig(f\"{mode}_loss_curve.png\")\n\n    # AUC curve\n    plt.figure(figsize=FIG_SINGLE)\n    plt.plot(epochs_range, history[\"train_auc\"], label=\"Training AUC\")\n    plt.plot(epochs_range, history[\"val_auc\"], label=\"Validation AUC\")\n    plt.suptitle(\"Discrimination Performance\", fontsize=14, fontweight=\"bold\", y=1)\n    plt.xlabel(\"Epoch\")\n    plt.ylabel(\"Area Under ROC Curve (AUC)\")\n    plt.ylim(0.0, 1.0)\n    plt.legend()\n    plt.grid(alpha=0.3)\n    plt.figtext(\n        0.5, -0.15,\n        \"AUC measures the ability to rank tumor patches above healthy patches across all decision thresholds.\",\n        ha=\"center\", fontsize=9\n    )\n    save_fig(f\"{mode}_auc_curve.png\")\n\n    # ROC & PR use last epoch's vt/vp (as before)\n    y_true = vt.ravel()\n    y_prob = vp.ravel()\n\n    fpr, tpr, thresholds = roc_curve(y_true, y_prob)\n    plt.figure(figsize=(5, 5))\n    plt.plot(fpr, tpr, label=f\"AUC = {best_auc:.4f}\")\n    plt.plot([0, 1], [0, 1], \"k--\", label=\"Random Classifier\")\n    plt.suptitle(\"Receiver Operating Characteristic (ROC) Curve\", fontsize=14, fontweight=\"bold\", y=1)\n    plt.xlabel(\"False Positive Rate\")\n    plt.ylabel(\"True Positive Rate\")\n    plt.legend()\n    plt.grid(alpha=0.3)\n    plt.figtext(\n        0.5, -0.18,\n        \"ROC curves evaluate classifier performance independent of a fixed decision threshold.\",\n        ha=\"center\", fontsize=9\n    )\n    save_fig(f\"{mode}_roc_curve.png\")\n\n    precision, recall, _ = precision_recall_curve(y_true, y_prob)\n    avg_precision = average_precision_score(y_true, y_prob)\n    plt.figure(figsize=(5, 5))\n    plt.plot(recall, precision, label=f\"AP = {avg_precision:.4f}\")\n    plt.suptitle(\"Precision–Recall Curve\", fontsize=14, fontweight=\"bold\", y=1)\n    plt.xlabel(\"Recall (Sensitivity)\")\n    plt.ylabel(\"Precision (Positive Predictive Value)\")\n    plt.xlim(0.0, 1.0)\n    plt.ylim(0.0, 1.05)\n    plt.legend()\n    plt.grid(alpha=0.3)\n    plt.figtext(\n        0.5, -0.18,\n        \"Precision–Recall curves are especially informative for imbalanced medical datasets.\",\n        ha=\"center\", fontsize=9\n    )\n    save_fig(f\"{mode}_pr_curve.png\")\n\n    # Confusion matrices\n    threshold_fixed = 0.5\n    y_pred_fixed = (y_prob >= threshold_fixed).astype(int)\n    cm_fixed = confusion_matrix(y_true, y_pred_fixed)\n\n    j_scores = tpr - fpr\n    best_idx = np.argmax(j_scores)\n    best_threshold = thresholds[best_idx]\n\n    y_pred_opt = (y_prob >= best_threshold).astype(int)\n    cm_opt = confusion_matrix(y_true, y_pred_opt)\n\n    print(f\"Confusion Matrix (Fixed Threshold = {threshold_fixed}):\\n{cm_fixed}\")\n    print(f\"\\nOptimal Threshold (Youden J): {best_threshold:.3f}\")\n    print(\"Confusion Matrix (Optimal Threshold):\\n\", cm_opt)\n\n    # Multi-scale scale-usage plot (only for multiscale)\n    if mode == \"multiscale\":\n        scale_stats = collect_scale_stats(model, val_loader)\n        labels_s = list(scale_stats.keys())\n        values_s = list(scale_stats.values())\n        plt.figure(figsize=(8, 4))\n        plt.bar(labels_s, values_s)\n        plt.ylabel(\"Mean Activation Magnitude\")\n        plt.title(\"Relative Utilization of Spatial Scales in Multi-Scale CNN\")\n        plt.xticks(rotation=45)\n        plt.grid(axis=\"y\", alpha=0.3)\n        plt.tight_layout()\n        save_fig(f\"{mode}_scale_usage.png\")\n\n    # Summary dict for comparison\n    return {\n        \"model\": mode,\n        \"best_val_auc\": best_auc,\n        \"final_val_acc\": history[\"val_acc\"][-1],\n        \"avg_precision\": avg_precision,\n        \"epochs_run\": len(history[\"train_loss\"])\n    }\n\n# ---------------------------------------------------------------------\n# Run 1 or many models and compare\n# ---------------------------------------------------------------------\n\nresults = []\nmodes_to_run = MODEL_GRID if RUN_ALL_MODELS else [MODEL_MODE]\n\nfor m in modes_to_run:\n    res = run_experiment(m)\n    results.append(res)\n\nresults_df = pd.DataFrame(results).set_index(\"model\").sort_values(\"best_val_auc\", ascending=False)\nresults_path = RESULTS_ROOT / \"model_comparison.csv\"\nresults_df.to_csv(results_path)\nprint(\"\\nModel comparison:\")\nprint(results_df)\nprint(f\"\\nSaved comparison table to: {results_path}\")\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-06T20:12:35.445061Z","iopub.execute_input":"2025-12-06T20:12:35.445691Z","iopub.status.idle":"2025-12-06T20:14:16.456961Z","shell.execute_reply.started":"2025-12-06T20:12:35.445671Z","shell.execute_reply":"2025-12-06T20:14:16.455958Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Submit\n\n### Generate submission\n\nOnce generated, you will use the Kaggle UI other than go to the project submission page to manually trigger a submission.\n","metadata":{}},{"cell_type":"code","source":"DATA_ROOT  = Path(\"/kaggle/input/histopathologic-cancer-detection\")\nTEST_DIR = DATA_ROOT / \"test\"\n# Define model loader\nclass PCamTestDataset(Dataset):\n    def __init__(self, root, transform):\n        self.paths = sorted(root.glob(\"*.tif\"))\n        self.ids = [p.stem for p in self.paths]\n        self.transform = transform\n\n    def __len__(self):\n        return len(self.paths)\n\n    def __getitem__(self, i):\n        img = Image.open(self.paths[i]).convert(\"RGB\")\n        return self.ids[i], self.transform(img)\n\n# Load best model\nmodel = build_model(MODEL_MODE).to(DEVICE)\nmodel.load_state_dict(\n    torch.load(MODELS_ROOT / f\"pcam_{MODEL_MODE}_best.pth\", map_location=DEVICE)\n)\nmodel.eval()\n\n# Inference\ntest_ds = PCamTestDataset(TEST_DIR, val_tfms)\ntest_loader = DataLoader(\n    test_ds,\n    batch_size=BATCH_SIZE,\n    shuffle=False,\n    num_workers=NUM_WORKERS\n)\n\npred_ids = []\npred_probs = []\n\nwith torch.no_grad():\n    for ids, x in tqdm(test_loader):\n        x = x.to(DEVICE)\n        logits = model(x)\n        probs = torch.sigmoid(logits).cpu().numpy().ravel()\n\n        pred_ids.extend(ids)\n        pred_probs.extend(probs)\n\nsubmission = pd.DataFrame({\n    \"id\": pred_ids,\n    \"label\": pred_probs\n})\n\nsubmission_path = Path(\"/kaggle/working/submission.csv\")\nsubmission.to_csv(submission_path, index=False)\n\nprint(\"Saved submission to:\", submission_path)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-06T20:16:39.028902Z","iopub.execute_input":"2025-12-06T20:16:39.029642Z","iopub.status.idle":"2025-12-06T20:17:59.282964Z","shell.execute_reply.started":"2025-12-06T20:16:39.029619Z","shell.execute_reply":"2025-12-06T20:17:59.282121Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Export results\n\nDownload these from the output pane and manually unizip and uploading to github.\n","metadata":{}},{"cell_type":"code","source":"from pathlib import Path\nimport shutil\n\nexport_dir = Path(\"/kaggle/working/export\")\nexport_dir.mkdir(exist_ok=True)\n\nshutil.copytree(PLOTS_ROOT, export_dir / \"plots\")\nshutil.copytree(MODELS_ROOT, export_dir / \"models\")\nshutil.copy(RESULTS_ROOT / \"model_comparison.csv\", export_dir)\n\nshutil.make_archive(\n    \"/kaggle/working/results_export\",\n    \"zip\",\n    export_dir\n)\n\nprint(\"Saved results_export.zip\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-06T20:50:30.300909Z","iopub.execute_input":"2025-12-06T20:50:30.301300Z","iopub.status.idle":"2025-12-06T20:50:32.796772Z","shell.execute_reply.started":"2025-12-06T20:50:30.301261Z","shell.execute_reply":"2025-12-06T20:50:32.795909Z"}},"outputs":[],"execution_count":null}]}