{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":11848,"databundleVersionId":862157,"sourceType":"competition"}],"dockerImageVersionId":31192,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **CNN Cancer Detection**","metadata":{}},{"cell_type":"markdown","source":"---\n---\n# **1. Brief description of the problem and data**\n\n---\n## Overview\n\nThis project uses deep learning models to identify the presence of metastatic cancer from  histopathology images.","metadata":{}},{"cell_type":"markdown","source":"<div style = \"background-color: GhostWhite;\n             padding: 10px;\n             border: 4px solid Blue;\n             margin-left: 2em;\">\n\nThis Kaggle competition is a binary image classification problem where you will identify metastatic cancer in small image patches taken from larger digital pathology scans.\n\n</div>\n\n<div style = \"background-color: AliceBlue;\n             padding: 10px;\n             border: 4px solid DeepSkyBlue;\n             margin-left: 2em;\">\n\nIn this competition, you must create an algorithm to identify metastatic cancer in small image patches taken from larger digital pathology scans. The data for this competition is a slightly modified version of the PatchCamelyon (PCam) benchmark dataset (the original PCam dataset contains duplicate images due to its probabilistic sampling, however, the version presented on Kaggle does not contain duplicates).\n\nPCam is highly interesting for both its size, simplicity to get started on, and approachability. In the authors' words:\n\n> [PCam] packs the clinically-relevant task of metastasis detection into a straight-forward binary image classification task, akin to CIFAR-10 and MNIST. Models can easily be trained on a single GPU in a couple hours, and achieve competitive scores in the Camelyon16 tasks of tumor detection and whole-slide image diagnosis. Furthermore, the balance between task-difficulty and tractability makes it a prime suspect for fundamental machine learning research on topics as active learning, model uncertainty, and explainability.\n\n</div>","metadata":{}},{"cell_type":"code","source":"# -----------------------------\n# Import libraries\n# -----------------------------\n\n# Silence Warnings\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\n# General libraries\nimport os\nimport random\nimport shutil\nimport numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\nfrom IPython.display import display, HTML, Markdown\n\n# Visualizations\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.express as px\nfrom matplotlib.patches import Patch, Rectangle\n\n# Image Processing\nimport cv2 as cv\nfrom skimage import io\nfrom skimage.transform import rotate\nfrom tifffile import imread\n\n# TensorFlow\nimport tensorflow as tf\nfrom tensorflow.keras import layers, models, optimizers\nfrom tensorflow.keras.utils import image_dataset_from_directory\n\n# Common Keras Layers\nfrom tensorflow.keras.layers import (\n    RandomFlip, RandomRotation, RandomZoom,\n    Conv2D, MaxPooling2D, AveragePooling2D,\n    Flatten, Dense, Dropout, BatchNormalization\n)\n\n# Profiling\nimport pandas_profiling as pp\n\n# Scikit-learn\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import roc_auc_score, accuracy_score","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T22:06:08.673353Z","iopub.execute_input":"2025-12-03T22:06:08.674170Z","iopub.status.idle":"2025-12-03T22:06:16.141543Z","shell.execute_reply.started":"2025-12-03T22:06:08.674139Z","shell.execute_reply":"2025-12-03T22:06:16.140927Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Set random_state for reproducibility\nrandom_state = 86","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T22:06:21.392096Z","iopub.execute_input":"2025-12-03T22:06:21.393088Z","iopub.status.idle":"2025-12-03T22:06:21.396691Z","shell.execute_reply.started":"2025-12-03T22:06:21.393063Z","shell.execute_reply":"2025-12-03T22:06:21.395828Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<br>\n\n---\n---\n# **2. Exploratory Data Analysis (EDA)**\n---\n\n\n### Import Data\n\n???","metadata":{}},{"cell_type":"markdown","source":"<div style = \"background-color: GhostWhite;\n             padding: 10px;\n             border: 4px solid Blue;\n             margin-left: 2em;\">\n\nBriefly describe the challenge problem and NLP. Describe the size, dimension, structure, etc., of the data.\n\n</div>\n\n<div style = \"background-color: AliceBlue;\n             padding: 10px;\n             border: 4px solid DeepSkyBlue;\n             margin-left: 2em;\">\n\nIn this dataset, you are provided with a large number of small pathology images to classify. Files are named with an image id. The train_labels.csv file provides the ground truth for the images in the train folder. You are predicting the labels for the images in the test folder. A positive label indicates that the center 32x32px region of a patch contains at least one pixel of tumor tissue. Tumor tissue in the outer region of the patch does not influence the label. This outer region is provided to enable fully-convolutional models that do not use zero-padding, to ensure consistent behavior when applied to a whole-slide image.\n\nThe original PCam dataset contains duplicate images due to its probabilistic sampling, however, the version presented on Kaggle does not contain duplicates. We have otherwise maintained the same data and splits as the PCam benchmark.\n\n</div>","metadata":{}},{"cell_type":"code","source":"# -----------------------------\n# Load files\n# -----------------------------\nsample_submission = pd.read_csv(\"../input/histopathologic-cancer-detection/sample_submission.csv\")\ntrain_raw = pd.read_csv(\"../input/histopathologic-cancer-detection/train_labels.csv\")\n\ntrain_path = \"../input/histopathologic-cancer-detection/train/\"\ntest_path = \"../input/histopathologic-cancer-detection/test/\"","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T22:06:25.641724Z","iopub.execute_input":"2025-12-03T22:06:25.642041Z","iopub.status.idle":"2025-12-03T22:06:25.891669Z","shell.execute_reply.started":"2025-12-03T22:06:25.642018Z","shell.execute_reply":"2025-12-03T22:06:25.891057Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n## Inspect Data\n\n???","metadata":{}},{"cell_type":"markdown","source":"<div style = \"background-color: GhostWhite;\n             padding: 10px;\n             border: 4px solid Blue;\n             margin-left: 2em;\">\n\nDescribe any data cleaning procedures. Based on your EDA, what is your plan of analysis?\n\n</div>","metadata":{}},{"cell_type":"code","source":"# -----------------------------\n# Initial look at train_raw\n# -----------------------------\ntrain_raw.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T22:06:32.577404Z","iopub.execute_input":"2025-12-03T22:06:32.577695Z","iopub.status.idle":"2025-12-03T22:06:32.588280Z","shell.execute_reply.started":"2025-12-03T22:06:32.577673Z","shell.execute_reply":"2025-12-03T22:06:32.587533Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# -----------------------------\n# Statistical summary of train_raw\n# -----------------------------\ntrain_raw.describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T22:06:35.168913Z","iopub.execute_input":"2025-12-03T22:06:35.169186Z","iopub.status.idle":"2025-12-03T22:06:35.190213Z","shell.execute_reply.started":"2025-12-03T22:06:35.169166Z","shell.execute_reply":"2025-12-03T22:06:35.189605Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# -----------------------------\n# Check for null data in train_raw\n# -----------------------------\ntrain_raw.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T22:06:37.209410Z","iopub.execute_input":"2025-12-03T22:06:37.209710Z","iopub.status.idle":"2025-12-03T22:06:37.238308Z","shell.execute_reply.started":"2025-12-03T22:06:37.209687Z","shell.execute_reply":"2025-12-03T22:06:37.237426Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# -----------------------------\n# Get number of images in train set and test set\n# -----------------------------\nn_train_raw = len(os.listdir(\"../input/histopathologic-cancer-detection/train\"))\ndisplay(HTML(f\"<strong>Number of images in Train Set:</strong> {n_train_raw}\"))\n\nn_test = len(os.listdir(\"../input/histopathologic-cancer-detection/test\"))\ndisplay(HTML(f\"<strong>Number of images in Test Set:</strong>  {n_test}\"))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T22:06:40.795751Z","iopub.execute_input":"2025-12-03T22:06:40.796073Z","iopub.status.idle":"2025-12-03T22:06:46.284794Z","shell.execute_reply.started":"2025-12-03T22:06:40.796049Z","shell.execute_reply":"2025-12-03T22:06:46.283977Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n## **Visualize Data**\n\n???","metadata":{}},{"cell_type":"markdown","source":"<div style = \"background-color: GhostWhite;\n             padding: 10px;\n             border: 4px solid Blue;\n             margin-left: 2em;\">\n\nShow a few visualizations like histograms. Describe any data cleaning procedures. Based on your EDA, what is your plan of analysis?\n\n</div>","metadata":{}},{"cell_type":"code","source":"# -----------------------------\n# Get label counts for train_raw\n# -----------------------------\ndisplay(pd.DataFrame(data={\"Counts\": train_raw[\"label\"].value_counts()}))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T22:06:53.886152Z","iopub.execute_input":"2025-12-03T22:06:53.886711Z","iopub.status.idle":"2025-12-03T22:06:53.897333Z","shell.execute_reply.started":"2025-12-03T22:06:53.886686Z","shell.execute_reply":"2025-12-03T22:06:53.896674Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# -----------------------------\n# Pie chart\n# -----------------------------\ncolors = sns.color_palette(\"seismic\", 2).as_hex()\n\nfig = px.pie(\n    train_raw, \n    values=train_raw[\"label\"].value_counts().values,\n    names=train_raw[\"label\"].unique(),\n    color_discrete_sequence=colors\n)\n\nfig.update_layout(\n    title={\n        \"text\": \"Label Distribution (Pie Chart)\",\n        \"y\": 0.95,\n        \"x\": 0.5,\n        \"xanchor\": \"center\",\n        \"yanchor\": \"top\",\n        \"font\": dict(size=18, weight=\"bold\")\n    },\n    legend_title=\"Label Meaning\",\n    legend=dict(\n        orientation=\"v\",\n        yanchor=\"top\",\n        y=0.95,\n        xanchor=\"left\",\n        x=1.02,\n        bordercolor=\"Gainsboro\",\n        borderwidth=2\n    )\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T22:06:57.230225Z","iopub.execute_input":"2025-12-03T22:06:57.230533Z","iopub.status.idle":"2025-12-03T22:07:00.147513Z","shell.execute_reply.started":"2025-12-03T22:06:57.230510Z","shell.execute_reply":"2025-12-03T22:07:00.146963Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# -----------------------------\n# Histogram\n# -----------------------------\ncolors = sns.color_palette(\"seismic\", 2).as_hex()\nax = sns.countplot(\n    x=train_raw[\"label\"],\n    palette=\"seismic\"\n)\nax.set(title=\"Label Distribution (Histogram)\")\nax.title.set_fontweight(\"bold\")\n\nlegend_patches = [\n    Patch(color=colors[0], label=\"0 = Negative for Cancer\"),\n    Patch(color=colors[1], label=\"1 = Positive for Cancer\")\n]\n\nplt.legend(handles=legend_patches, title=\"Label Meaning\", loc=\"upper right\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T22:07:04.235446Z","iopub.execute_input":"2025-12-03T22:07:04.236165Z","iopub.status.idle":"2025-12-03T22:07:04.464841Z","shell.execute_reply.started":"2025-12-03T22:07:04.236139Z","shell.execute_reply":"2025-12-03T22:07:04.464056Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# -----------------------------\n# Visualize some images from train_raw\n# -----------------------------\nlabel_colors = sns.color_palette(\"seismic\", 2)\n\nfig, ax = plt.subplots(5, 5, figsize=(15, 15))\n\nfor i, axis in enumerate(ax.flat):\n    file = str(train_path + train_raw.id[i] + \".tif\")\n    image = io.imread(file)\n    axis.imshow(image,\n                #cmap=\"gray\"\n               )\n    label = train_raw.label[i]\n    color = label_colors[label]\n    box = Rectangle((32, 32), 32, 32,\n                     linewidth=2,\n                     edgecolor=color,\n                     facecolor=\"none\")\n    axis.add_patch(box)\n    \n    # Set label below image\n    axis.set(\n        xlabel=f\"{label} ({'Negative' if label==0 else 'Positive'})\",\n        xticks=[], yticks=[]\n    )\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T22:07:08.519367Z","iopub.execute_input":"2025-12-03T22:07:08.520142Z","iopub.status.idle":"2025-12-03T22:07:10.036630Z","shell.execute_reply.started":"2025-12-03T22:07:08.520115Z","shell.execute_reply":"2025-12-03T22:07:10.035807Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<br>\n\n---\n---\n# **3. Model Architecture**\n---\n\n???\n\n<div style = \"background-color: LavenderBlush;\n             padding: 10px;\n             border: 4px solid LightPink;\n             margin-left: 2em;\">\n\nA bitch is gonna build two models. Describe 'em both.\n\n</div>","metadata":{}},{"cell_type":"code","source":"# -----------------------------\n# Parameters\n# -----------------------------\nimg_size = (96, 96)\nbatch_size = 32\nAUTOTUNE = tf.data.AUTOTUNE","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T22:07:20.266125Z","iopub.execute_input":"2025-12-03T22:07:20.266858Z","iopub.status.idle":"2025-12-03T22:07:20.272127Z","shell.execute_reply.started":"2025-12-03T22:07:20.266820Z","shell.execute_reply":"2025-12-03T22:07:20.270759Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# -----------------------------\n# File paths\n# -----------------------------\ntrain_files = [os.path.join(train_path, f\"{fid}.tif\") for fid in train_raw.id]\ntrain_targets = train_raw.label.values\n\n# Grab all TIFF files in the folder\ntest_files = sorted([\n    os.path.join(test_path, f) for f in os.listdir(test_path) if f.endswith(\".tif\")\n])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T22:07:20.837528Z","iopub.execute_input":"2025-12-03T22:07:20.838162Z","iopub.status.idle":"2025-12-03T22:07:21.125443Z","shell.execute_reply.started":"2025-12-03T22:07:20.838134Z","shell.execute_reply":"2025-12-03T22:07:21.124835Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# -----------------------------\n# Dataset generator functions\n# -----------------------------\ndef load_img_train(path, label):\n    path = path.numpy().decode(\"utf-8\")\n    img = io.imread(path)\n    # Convert to tensor early to avoid dtype mismatch\n    img = tf.convert_to_tensor(img)\n    \n    if img.ndim == 2:                               # Grayscale\n        img = tf.stack([img, img, img], axis=-1)    # Force RGB\n    elif img.shape[-1] == 4:                        # RGBA\n        img = img[..., :3]\n\n    img = tf.image.resize(img, img_size)\n    img = tf.cast(img, tf.float32) / 255.0\n    return img, label\n\n\ndef set_shape_train(img, label):\n    img.set_shape((*img_size, 3))\n    label.set_shape(())\n    return img, label\n\n\ndef load_img_test(path):\n    path = path.numpy().decode(\"utf-8\")\n    img = io.imread(path)\n\n    img = tf.convert_to_tensor(img)\n\n    # Grayscale\n    if img.ndim == 2:\n        img = tf.stack([img, img, img], axis=-1)\n\n    # Weird TIFFs shaped (H, 3)\n    if img.ndim == 2 and img.shape[-1] == 3:\n        img = tf.expand_dims(img, axis=1)\n\n    # RGBA\n    if img.shape[-1] == 4:\n        img = img[..., :3]\n\n    img = tf.image.resize(img, img_size)\n    img = tf.cast(img, tf.float32) / 255.0\n    return img\n\n\ndef load_img_test_wrapper(path):\n    img = tf.py_function(load_img_test, [path], tf.float32)\n    img.set_shape((*img_size, 3))\n    return img","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T22:07:23.688144Z","iopub.execute_input":"2025-12-03T22:07:23.688900Z","iopub.status.idle":"2025-12-03T22:07:23.703917Z","shell.execute_reply.started":"2025-12-03T22:07:23.688865Z","shell.execute_reply":"2025-12-03T22:07:23.703088Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# -----------------------------\n# Train + Val Set\n# -----------------------------\ntrain_raw_data = tf.data.Dataset.from_tensor_slices((train_files, train_targets))\ntrain_raw_data = train_raw_data.map(\n    lambda x, y: tf.py_function(load_img_train, [x, y], [tf.float32, tf.int64]),\n    num_parallel_calls=AUTOTUNE\n)\ntrain_raw_data = train_raw_data.map(set_shape_train)\ntrain_raw_data = train_raw_data.shuffle(1000)\n\nval_split = 0.2\nn_val = int(len(train_files) * val_split)\n\n# Final split: Take/skip AFTER batching\nval_set   = train_raw_data.take(n_val).batch(batch_size).cache().prefetch(AUTOTUNE)\ntrain_set = train_raw_data.skip(n_val).batch(batch_size).cache().prefetch(AUTOTUNE)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T22:07:26.712925Z","iopub.execute_input":"2025-12-03T22:07:26.713292Z","iopub.status.idle":"2025-12-03T22:07:28.150979Z","shell.execute_reply.started":"2025-12-03T22:07:26.713272Z","shell.execute_reply":"2025-12-03T22:07:28.150062Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# -----------------------------\n# Test Set\n# -----------------------------\ntest_set = tf.data.Dataset.from_tensor_slices(test_files)\ntest_set = test_set.map(load_img_test_wrapper, num_parallel_calls=AUTOTUNE)\ntest_set = test_set.batch(batch_size).prefetch(AUTOTUNE)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T22:08:06.073240Z","iopub.execute_input":"2025-12-03T22:08:06.073953Z","iopub.status.idle":"2025-12-03T22:08:06.213248Z","shell.execute_reply.started":"2025-12-03T22:08:06.073919Z","shell.execute_reply":"2025-12-03T22:08:06.212461Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n## Model 1: Sequential CNN\n\n???","metadata":{}},{"cell_type":"markdown","source":"<div style = \"background-color: GhostWhite;\n             padding: 10px;\n             border: 4px solid Blue;\n             margin-left: 2em;\">\n\nDescribe your model architecture and reasoning for why you believe that specific architecture would be suitable for this problem. Compare multiple architectures and tune hyperparameters.\n\n</div>","metadata":{}},{"cell_type":"code","source":"# -----------------------------\n# Sequential CNN Model\n# -----------------------------\ndef seq_cnn(dropout_rate=0.2, lr=0.001, input_shape=(*img_size, 3)):\n    model = models.Sequential([\n        layers.Conv2D(32, (3,3), activation=\"relu\", input_shape=input_shape),\n        layers.MaxPooling2D(),\n        layers.BatchNormalization(),\n        \n        layers.Conv2D(32, (3,3), activation=\"relu\"),\n        layers.MaxPooling2D(),\n        layers.BatchNormalization(),\n        \n        layers.Flatten(),\n        layers.Dropout(dropout_rate),\n        layers.Dense(32, activation=\"relu\"),\n        layers.Dense(1, activation=\"sigmoid\")\n    ])\n    \n    model.compile(\n        optimizer=optimizers.Adam(learning_rate=lr),\n        loss=\"binary_crossentropy\",\n        metrics = [\"accuracy\", tf.keras.metrics.AUC(name=\"roc_auc\", curve=\"ROC\")]\n    )\n    return model","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T22:21:35.212294Z","iopub.execute_input":"2025-12-03T22:21:35.212608Z","iopub.status.idle":"2025-12-03T22:21:35.218176Z","shell.execute_reply.started":"2025-12-03T22:21:35.212586Z","shell.execute_reply":"2025-12-03T22:21:35.217573Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n## Model 2: VGGNet\n\n???","metadata":{}},{"cell_type":"markdown","source":"<div style = \"background-color: GhostWhite;\n             padding: 10px;\n             border: 4px solid Blue;\n             margin-left: 2em;\">\n\nDescribe your model architecture and reasoning for why you believe that specific architecture would be suitable for this problem. Compare multiple architectures and tune hyperparameters.\n\n</div>","metadata":{}},{"cell_type":"code","source":"# -----------------------------\n# VGGNet Model\n# -----------------------------\ndef vgg_net(dropout_rate=0.2, lr=0.001, input_shape=(*img_size, 3)):\n    model = models.Sequential([\n        layers.Conv2D(32, (3,3), activation=\"relu\", padding=\"same\",\n                      input_shape=input_shape),\n        layers.Conv2D(32, (3,3), activation=\"relu\", padding=\"same\"),\n        layers.MaxPooling2D(),\n        layers.BatchNormalization(),\n        \n        layers.Conv2D(32, (3,3), activation=\"relu\", padding=\"same\"),\n        layers.Conv2D(32, (3,3), activation=\"relu\", padding=\"same\"),\n        layers.MaxPooling2D(),\n        layers.BatchNormalization(),\n        \n        layers.Flatten(),\n        layers.Dropout(dropout_rate),\n        layers.Dense(32, activation=\"relu\"),\n        layers.Dense(1, activation=\"sigmoid\")\n    ])\n    \n    model.compile(\n        optimizer=optimizers.Adam(learning_rate=lr),\n        loss=\"binary_crossentropy\",\n        metrics = [\"accuracy\", tf.keras.metrics.AUC(name=\"roc_auc\", curve=\"ROC\")]\n    )\n    return model","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T22:08:15.579008Z","iopub.execute_input":"2025-12-03T22:08:15.579756Z","iopub.status.idle":"2025-12-03T22:08:15.585644Z","shell.execute_reply.started":"2025-12-03T22:08:15.579730Z","shell.execute_reply":"2025-12-03T22:08:15.584868Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<br>\n\n---\n---\n# **4. Results and Analysis**\n\n---\n## Model 1: Sequential CNN (Tuning on Train + Valid)\n\n???","metadata":{}},{"cell_type":"markdown","source":"<div style = \"background-color: GhostWhite;\n             padding: 10px;\n             border: 4px solid Blue;\n             margin-left: 2em;\">\n\nRun hyperparameter tuning, try different architectures for comparison, apply techniques to improve training or performance, and discuss what helped.\n\nIncludes results with tables and figures. There is an analysis of why or why not something worked well, troubleshooting, and a hyperparameter optimization procedure summary.\n\n</div>","metadata":{}},{"cell_type":"code","source":"# -----------------------------\n# Tuning and Metrics\n# -----------------------------\nhyperparams = [\n    {\"dropout_rate\": 0.2, \"lr\": 0.001},\n    {\"dropout_rate\": 0.2, \"lr\": 0.0005},\n]\n\nbest_val_auc = 0\nbest_model_seq_cnn = None\nseq_cnn_summary = []\n\ndef seq_cnn_predict(model, dataset, labeled=True):\n    all_probs = []\n    all_labels = []\n    for batch in tqdm(dataset, desc=\"Predicting\"):\n        if labeled:\n            imgs, labels = batch\n            all_labels.append(labels.numpy())\n        else:\n            imgs = batch\n        probs = model.predict(imgs, verbose=0)\n        all_probs.append(probs)\n    all_probs = np.concatenate(all_probs, axis=0).flatten()\n    if labeled:\n        all_labels = np.concatenate(all_labels, axis=0).flatten()\n        return all_probs, all_labels\n    else:\n        return all_probs\n\nfor params in hyperparams:\n    print(f\"\\nTraining model with params: {params}\")\n    model = seq_cnn(**params)\n    \n    history = model.fit(\n        train_set,\n        validation_data=val_set,\n        epochs=10,\n        callbacks=[tf.keras.callbacks.EarlyStopping(\n            monitor=\"val_loss\", patience=3, restore_best_weights=True\n        )]\n    )\n    \n    # Predict train + val\n    train_probs, train_labels = seq_cnn_predict(model, train_set, labeled=True)\n    val_probs, val_labels     = seq_cnn_predict(model, val_set, labeled=True)\n\n    # Convert probabilities to binary predictions\n    train_preds = (train_probs > 0.5).astype(int)\n    val_preds   = (val_probs > 0.5).astype(int)\n\n    # Compute metrics\n    train_auc = roc_auc_score(train_labels, train_probs)\n    val_auc   = roc_auc_score(val_labels, val_probs)\n    train_acc = accuracy_score(train_labels, train_preds)\n    val_acc   = accuracy_score(val_labels, val_preds)\n\n    print(f\"Train Acc: {train_acc:.4f} | Train ROC-AUC: {train_auc:.4f}\")\n    print(f\"Val Acc: {val_acc:.4f} | Val   ROC-AUC: {val_auc:.4f}\")\n    \n    # Save metrics for seq_cnn_summary\n    seq_cnn_summary.append({\n        \"dropout_rate\": params[\"dropout_rate\"],\n        \"lr\": params[\"lr\"],\n        \"train_acc\": train_acc,\n        \"train_auc\": train_auc,\n        \"val_acc\": val_acc,\n        \"val_auc\": val_auc\n    })\n    \n    # Track best model by val ROC-AUC\n    if val_auc > best_val_auc:\n        best_val_auc = val_auc\n        best_model_seq_cnn = model\n        print(\"Best model updated!\")\n\n\n# -----------------------------\n# Metrics summary table\n# -----------------------------\nseq_cnn_summary = pd.DataFrame(seq_cnn_summary)\nprint(\"\\nHyperparameter tuning summary:\")\nprint(seq_cnn_summary)\n\n\n# -----------------------------\n# Best model metrics\n# -----------------------------\nbest_row = seq_cnn_summary.loc[seq_cnn_summary[\"val_auc\"].idxmax()]\n\nprint(\"\\nBest model metrics (based on val ROC-AUC):\")\nprint(f\"Dropout rate: {best_row['dropout_rate']}\")\nprint(f\"Learning rate: {best_row['lr']}\")\nprint(f\"Train Accuracy: {best_row['train_acc']:.4f}\")\nprint(f\"Train ROC-AUC: {best_row['train_auc']:.4f}\")\nprint(f\"Val Accuracy:   {best_row['val_acc']:.4f}\")\nprint(f\"Val ROC-AUC:   {best_row['val_auc']:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T22:21:38.799179Z","iopub.execute_input":"2025-12-03T22:21:38.799945Z","iopub.status.idle":"2025-12-03T22:59:00.944431Z","shell.execute_reply.started":"2025-12-03T22:21:38.799920Z","shell.execute_reply":"2025-12-03T22:59:00.943653Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ------------------------\n# Plot Results\n# ------------------------\nseq_cnn_summary[\"label\"] = seq_cnn_summary.apply(\n    lambda row: f\"drop={row['dropout_rate']}, lr={row['lr']}\", axis=1\n)\n\nx = range(len(seq_cnn_summary))\nplt.figure(figsize=(10,6))\nplt.plot(x, seq_cnn_summary[\"train_auc\"], marker=\"o\", linestyle=\"-\", \n         color=\"RoyalBlue\", label=\"Train ROC-AUC\")\nplt.plot(x, seq_cnn_summary[\"train_acc\"], marker=\"s\", linestyle=\"--\", \n         color=\"RoyalBlue\", label=\"Train Accuracy\")\n\nplt.plot(x, seq_cnn_summary[\"val_auc\"], marker=\"o\", linestyle=\"-\",\n         color=\"Crimson\", label=\"Val ROC-AUC\")\nplt.plot(x, seq_cnn_summary[\"val_acc\"], marker=\"s\", linestyle=\"--\",\n         color=\"Crimson\", label=\"Val Accuracy\")\n\n# Highlight best model\nhighlight_x = len(seq_cnn_summary) - 1\nplt.axvspan(\n    highlight_x - 0.1,  # left boundary\n    highlight_x + 0.1,  # right boundary\n    color=\"Plum\",\n    alpha=0.2,\n    label=\"Best Model\"\n)\n\n# Labels & style\nplt.xticks(x, seq_cnn_summary[\"label\"], rotation=45)\nplt.title(\"Sequential CNN Model: Hyperparameter Tuning\", weight=\"bold\")\nplt.ylabel(\"Score\")\nplt.grid(True, linestyle='--', linewidth=0.5, alpha=0.7)\nplt.legend()\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-04T00:05:06.317451Z","iopub.execute_input":"2025-12-04T00:05:06.317686Z","iopub.status.idle":"2025-12-04T00:05:06.550526Z","shell.execute_reply.started":"2025-12-04T00:05:06.317670Z","shell.execute_reply":"2025-12-04T00:05:06.549842Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# -----------------------------\n# Sequential CNN on test set\n# -----------------------------\ntest_probs_seq_cnn = seq_cnn_predict(best_model_seq_cnn, test_set, labeled=False)\ntest_files_sorted = sorted([f for f in os.listdir(test_path) if f.endswith(\".tif\")])\ntest_ids = [os.path.splitext(f)[0] for f in test_files_sorted]\n\nsubmission_seq_cnn = pd.DataFrame({\n    \"id\": test_ids,\n    \"label\": test_probs_seq_cnn\n})\nsubmission_seq_cnn.to_csv(\"submission_seq_cnn.csv\", index=False)\nprint(\"Submission CSV saved for Sequential CNN model.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T22:59:36.211046Z","iopub.execute_input":"2025-12-03T22:59:36.211887Z","iopub.status.idle":"2025-12-03T23:03:15.744154Z","shell.execute_reply.started":"2025-12-03T22:59:36.211847Z","shell.execute_reply":"2025-12-03T23:03:15.743360Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n## Model 2: VGGNet (Tuning on Train + Valid)\n\n???","metadata":{}},{"cell_type":"markdown","source":"<div style = \"background-color: GhostWhite;\n             padding: 10px;\n             border: 4px solid Blue;\n             margin-left: 2em;\">\n\nRun hyperparameter tuning, try different architectures for comparison, apply techniques to improve training or performance, and discuss what helped.\n\nIncludes results with tables and figures. There is an analysis of why or why not something worked well, troubleshooting, and a hyperparameter optimization procedure summary.\n\n</div>","metadata":{}},{"cell_type":"code","source":"# -----------------------------\n# Tuning and Metrics\n# -----------------------------\nhyperparams = [\n    {\"dropout_rate\": 0.2, \"lr\": 0.001},\n    {\"dropout_rate\": 0.2, \"lr\": 0.0005},\n]\n\nbest_val_auc = 0\nbest_model_vgg = None\nvgg_summary = []\n\ndef vgg_net_predict(model, dataset, labeled=True):\n    all_probs = []\n    all_labels = []\n    for batch in tqdm(dataset, desc=\"Predicting\"):\n        if labeled:\n            imgs, labels = batch\n            all_labels.append(labels.numpy())\n        else:\n            imgs = batch\n        probs = model.predict(imgs, verbose=0)\n        all_probs.append(probs)\n    all_probs = np.concatenate(all_probs, axis=0).flatten()\n    if labeled:\n        all_labels = np.concatenate(all_labels, axis=0).flatten()\n        return all_probs, all_labels\n    else:\n        return all_probs\n\nfor params in hyperparams:\n    print(f\"\\nTraining model with params: {params}\")\n    model = vgg_net(**params)\n    \n    history = model.fit(\n        train_set,\n        validation_data=val_set,\n        epochs=10,\n        callbacks=[tf.keras.callbacks.EarlyStopping(\n            monitor=\"val_loss\", patience=3, restore_best_weights=True\n        )]\n    )\n    \n    # Predict train + val\n    train_probs, train_labels = vgg_net_predict(model, train_set, labeled=True)\n    val_probs, val_labels     = vgg_net_predict(model, val_set, labeled=True)\n\n    # Convert probabilities to binary predictions\n    train_preds = (train_probs > 0.5).astype(int)\n    val_preds   = (val_probs > 0.5).astype(int)\n\n    # Compute metrics\n    train_auc = roc_auc_score(train_labels, train_probs)\n    val_auc   = roc_auc_score(val_labels, val_probs)\n    train_acc = accuracy_score(train_labels, train_preds)\n    val_acc   = accuracy_score(val_labels, val_preds)\n\n    print(f\"Train Acc: {train_acc:.4f} | Train ROC-AUC: {train_auc:.4f}\")\n    print(f\"Val Acc:   {val_acc:.4f}   | Val   ROC-AUC: {val_auc:.4f}\")\n    \n    # Save metrics for vgg_summary\n    vgg_summary.append({\n        \"dropout_rate\": params[\"dropout_rate\"],\n        \"lr\": params[\"lr\"],\n        \"train_acc\": train_acc,\n        \"train_auc\": train_auc,\n        \"val_acc\": val_acc,\n        \"val_auc\": val_auc\n    })\n    \n    # Track best model by val ROC-AUC\n    if val_auc > best_val_auc:\n        best_val_auc = val_auc\n        best_model_vgg = model\n        print(\"Best model updated!\")\n        \n\n# -----------------------------\n# Metrics summary table\n# -----------------------------\nvgg_summary = pd.DataFrame(vgg_summary)\nprint(\"\\nHyperparameter tuning vgg_summary:\")\nprint(vgg_summary)\n\n\n# -----------------------------\n# Best model metrics\n# -----------------------------\nbest_row = vgg_summary.loc[vgg_summary[\"val_auc\"].idxmax()]\n\nprint(\"\\nBest model metrics (based on val ROC-AUC):\")\nprint(f\"Dropout rate: {best_row['dropout_rate']}\")\nprint(f\"Learning rate: {best_row['lr']}\")\nprint(f\"Train Accuracy: {best_row['train_acc']:.4f}\")\nprint(f\"Train ROC-AUC: {best_row['train_auc']:.4f}\")\nprint(f\"Val Accuracy:   {best_row['val_acc']:.4f}\")\nprint(f\"Val ROC-AUC:   {best_row['val_auc']:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T23:33:44.925288Z","iopub.execute_input":"2025-12-03T23:33:44.925613Z","iopub.status.idle":"2025-12-04T00:05:06.315939Z","shell.execute_reply.started":"2025-12-03T23:33:44.925591Z","shell.execute_reply":"2025-12-04T00:05:06.315010Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ------------------------\n# Plot Results\n# ------------------------\nvgg_summary[\"label\"] = vgg_summary.apply(\n    lambda row: f\"drop={row['dropout_rate']}, lr={row['lr']}\", axis=1\n)\n\nx = range(len(vgg_summary))\nplt.figure(figsize=(10,6))\nplt.plot(x, vgg_summary[\"train_auc\"], marker=\"o\", linestyle=\"-\",\n         color=\"RoyalBlue\", label=\"Train ROC-AUC\")\nplt.plot(x, vgg_summary[\"train_acc\"], marker=\"s\", linestyle=\"--\",\n         color=\"RoyalBlue\", label=\"Train Accuracy\")\n\nplt.plot(x, vgg_summary[\"val_auc\"], marker=\"o\", linestyle=\"-\",\n         color=\"Crimson\", label=\"Val ROC-AUC\")\nplt.plot(x, vgg_summary[\"val_acc\"], marker=\"s\", linestyle=\"--\",\n         color=\"Crimson\", label=\"Val Accuracy\")\n\n# Highlight best model\nhighlight_x = len(vgg_summary) - 1\nplt.axvspan(\n    highlight_x - 0.1,  # left boundary\n    highlight_x + 0.1,  # right boundary\n    color=\"Plum\",\n    alpha=0.2,\n    label=\"Best Model\"\n)\n\n# Labels & style\nplt.xticks(x, vgg_summary[\"label\"], rotation=45)\nplt.title(\"VGGNet Model: Hyperparameter Tuning\", weight=\"bold\")\nplt.ylabel(\"Score\")\nplt.grid(True, linestyle='--', linewidth=0.5, alpha=0.7)\nplt.legend()\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-04T00:05:16.027074Z","iopub.execute_input":"2025-12-04T00:05:16.027331Z","iopub.status.idle":"2025-12-04T00:05:16.270416Z","shell.execute_reply.started":"2025-12-04T00:05:16.027314Z","shell.execute_reply":"2025-12-04T00:05:16.269645Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# -----------------------------\n# VGGNet on test set\n# -----------------------------\ntest_probs_vgg = vgg_net_predict(best_model_vgg, test_set, labeled=False)\ntest_files_sorted = sorted([f for f in os.listdir(test_path) if f.endswith(\".tif\")])\ntest_ids = [os.path.splitext(f)[0] for f in test_files_sorted]\n\nsubmission_vgg = pd.DataFrame({\n    \"id\": test_ids,\n    \"label\": test_probs_vgg\n})\nsubmission_vgg.to_csv(\"submission_vgg.csv\", index=False)\nprint(\"CSV saved for VGGNet model.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-04T00:06:51.744773Z","iopub.execute_input":"2025-12-04T00:06:51.745109Z","iopub.status.idle":"2025-12-04T00:10:10.321255Z","shell.execute_reply.started":"2025-12-04T00:06:51.745087Z","shell.execute_reply":"2025-12-04T00:10:10.320565Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<br>\n\n---\n---\n# **5. Conclusion**\n\n---\n## Final Results\n\n???","metadata":{}},{"cell_type":"markdown","source":"<div style = \"background-color: GhostWhite;\n             padding: 10px;\n             border: 4px solid Blue;\n             margin-left: 2em;\">\n\nDiscuss and interpret results as well as learnings and takeaways.\n\n</div>","metadata":{}},{"cell_type":"markdown","source":"---\n## Discussion\n\n???","metadata":{}},{"cell_type":"markdown","source":"<div style = \"background-color: GhostWhite;\n             padding: 10px;\n             border: 4px solid Blue;\n             margin-left: 2em;\">\n\nWhat did and did not help improve the performance of your models? What improvements could you try in the future?\n\n</div>","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}