{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":11848,"databundleVersionId":862157,"sourceType":"competition"},{"sourceId":12393058,"sourceType":"datasetVersion","datasetId":7814892}],"dockerImageVersionId":31041,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Histopathologic Cancer Detection using TensorFlow and CNNs","metadata":{}},{"cell_type":"markdown","source":"## Problem at Hand\n\nThe task is **binary image classification**:  \ndetermine whether a $96 \\times 96$ pixel histopathology patch contains tumor tissue in its **central $32 \\times 32$** region (`label = 1`) or not (`label = 0`).  \nAlthough some competition templates mention *NLP*, this challenge is strictly **computer vision**.\n\n### Challenge Characteristics\n\n| Aspect          | Details                                                                                                      |\n|-----------------|--------------------------------------------------------------------------------------------------------------|\n| **Goal**        | Predict the `label` for every image in **`test/`** and submit a probability CSV.                             |\n| **Difficulty**  | • Fine-grained tumor detection<br>• Mild class imbalance (~59% benign / 41% malignant)<br>• Potential overfitting to slide-level patterns |\n| **Typical Models** | CNNs or fully-convolutional networks that leverage the full patch while focusing label supervision on the center. |\n\n### Dataset Structure & Size\n\n| Component            | Path / File              | Count       | Format                    | Contents                                    |\n|----------------------|--------------------------|-------------|---------------------------|---------------------------------------------|\n| **Labeled images**   | `train/`                 | 220,025     | 96 × 96 px RGB `.tif`     | Training images                             |\n| **Ground-truth CSV** | `train_labels.csv`       | 220,025 rows × 2 columns | `id`, `label` | One label per training image               |\n| **Unlabeled images** | `test/`                  | 57,458      | Same format               | For inference and submission                |\n| **Submission template** | `sample_submission.csv` | 57,458 rows | `id`, `label`             | Required format for prediction submission   |\n\n**Total files:** 277,483 (≈ 7.8 GB)\n\n### Notes on Duplicates & Patch Labels\n\n- The original PCam dataset had duplicates due to probabilistic sampling; **this Kaggle release does not contain duplicates**.\n- The label applies **only to the center $32 \\times 32$** region of each image.\n- The surrounding pixels help convolutional models avoid border effects, especially when deployed on full-slide scans.\n\nSource: https://www.kaggle.com/c/histopathologic-cancer-detection/data\n\n---","metadata":{}},{"cell_type":"code","source":"!pip install ipython-autotime\n%load_ext autotime","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T02:10:38.169896Z","iopub.execute_input":"2025-07-07T02:10:38.170424Z","iopub.status.idle":"2025-07-07T02:10:42.392394Z","shell.execute_reply.started":"2025-07-07T02:10:38.170398Z","shell.execute_reply":"2025-07-07T02:10:42.391566Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport os\nfrom pathlib import Path\nimport matplotlib.pyplot as plt\nimport matplotlib.image as mpimg\nfrom sklearn.model_selection import train_test_split\nimport glob\nimport tensorflow as tf","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T02:10:42.393870Z","iopub.execute_input":"2025-07-07T02:10:42.394106Z","iopub.status.idle":"2025-07-07T02:10:56.429935Z","shell.execute_reply.started":"2025-07-07T02:10:42.394082Z","shell.execute_reply":"2025-07-07T02:10:56.429161Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Exploratory Data Analysis","metadata":{}},{"cell_type":"code","source":"# utils/datasets.py\n\n\ndef build_image_dfs(\n    root_dir: str,\n    train_sub: str       = \"train\",          # folder with labelled images\n    test_sub: str        = \"test\",           # folder with unlabelled images\n    label_csv: str       = \"train_labels.csv\",\n    id_col: str          = \"id\",\n    label_col: str       = \"label\",\n    ext: str             = \".tif\",\n    val_split: float     = 0.2,\n    stratify: bool       = True,\n    seed: int            = 42,\n):\n    \"\"\"\n    Returns train_df, val_df, test_df  (all with columns: id, filepath, [label])\n    \"\"\"\n\n    root = Path(root_dir)\n\n    # ---------- labelled training images ----------\n    train_dir = root / train_sub\n    train_paths = list(train_dir.glob(f\"*{ext}\"))\n    df = pd.DataFrame({\"filepath\": train_paths})\n    df[id_col] = df[\"filepath\"].apply(lambda p: p.stem)\n\n    labels = pd.read_csv(root / label_csv, dtype={label_col: \"uint8\"})\n    df = df.merge(labels[[id_col, label_col]], on=id_col, how=\"left\")\n\n    if df[label_col].isna().any():\n        raise ValueError(\"Some images in the folder have no label in the CSV.\")\n\n    # ---------- optional train/val split ----------\n    if val_split > 0:\n        strat = df[label_col] if stratify else None\n        train_df, val_df = train_test_split(\n            df, test_size=val_split, stratify=strat, random_state=seed\n        )\n    else:\n        train_df, val_df = df, pd.DataFrame(columns=df.columns)\n\n    # ---------- unlabelled test images ----------\n    test_dir = root / test_sub\n    test_paths = list(test_dir.glob(f\"*{ext}\"))\n    test_df = pd.DataFrame({\"filepath\": test_paths})\n    test_df[id_col] = test_df[\"filepath\"].apply(lambda p: p.stem)\n\n    return train_df.reset_index(drop=True), val_df.reset_index(drop=True), test_df\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T02:10:56.431143Z","iopub.execute_input":"2025-07-07T02:10:56.431729Z","iopub.status.idle":"2025-07-07T02:10:56.442863Z","shell.execute_reply.started":"2025-07-07T02:10:56.431703Z","shell.execute_reply":"2025-07-07T02:10:56.442067Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df, val_df, test_df = build_image_dfs(\n    root_dir=\"/kaggle/input/histopathologic-cancer-detection\",\n    val_split=False,\n)\n\nTRAIN_DIR = Path(\"/kaggle/input/histopathologic-cancer-detection/train\")\nTEST_DIR = Path(\"/kaggle/input/histopathologic-cancer-detection/test\")\ntrain_df['filepath'] = train_df['id'].apply(lambda x: str(TRAIN_DIR / f\"{x}.tif\"))\ntest_df['filepath'] = test_df['id'].apply(lambda x: str(TEST_DIR / f\"{x}.tif\"))\n\nprint(train_df.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T02:14:56.184201Z","iopub.execute_input":"2025-07-07T02:14:56.184839Z","iopub.status.idle":"2025-07-07T02:15:01.563871Z","shell.execute_reply.started":"2025-07-07T02:14:56.184815Z","shell.execute_reply":"2025-07-07T02:15:01.563044Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"Number of duplicate rows in training set:\", train_df.duplicated().sum(),\"\\n\")\nprint(\"Number of duplicate rows in testing set:\", test_df.duplicated().sum(),\"\\n\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T02:15:05.380494Z","iopub.execute_input":"2025-07-07T02:15:05.381066Z","iopub.status.idle":"2025-07-07T02:15:05.547437Z","shell.execute_reply.started":"2025-07-07T02:15:05.381042Z","shell.execute_reply":"2025-07-07T02:15:05.546651Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from PIL import Image\nimport random\nfrom pathlib import Path\n\n# ------------------------------------------------------------------\ndef show_df_samples(df: pd.DataFrame,\n                    n: int = 16,\n                    cols: int = 4,\n                    label_col: str = \"label\",\n                    title: str = \"\"):\n    \"\"\"\n    Plot a grid of N images drawn from a dataframe that has a 'filepath' column\n    and (optionally) a label column.\n    \"\"\"\n    assert \"filepath\" in df.columns, \"DataFrame must contain a 'filepath' column.\"\n\n    paths  = df.sample(n=min(n, len(df)), random_state=42)[\"filepath\"].tolist()\n    rows   = (len(paths) + cols - 1) // cols\n\n    plt.figure(figsize=(cols * 3, rows * 3))\n    for i, p in enumerate(paths, 1):\n        img = Image.open(Path(p))\n        plt.subplot(rows, cols, i)\n        plt.imshow(img)\n        plt.axis(\"off\")\n        if label_col in df.columns:\n            lbl = df.loc[df[\"filepath\"] == str(p), label_col].values[0]\n            plt.title(str(lbl), fontsize=8)\n    plt.suptitle(title, fontsize=14)\n    plt.tight_layout()\n    plt.show()\n# ------------------------------------------------------------------\n\nshow_df_samples(train_df, n=16, cols=4, label_col=\"label\", title=\"Train sample\")\nshow_df_samples(test_df,  n=16, cols=4, label_col=None,   title=\"Test sample\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T02:15:19.780963Z","iopub.execute_input":"2025-07-07T02:15:19.781656Z","iopub.status.idle":"2025-07-07T02:15:22.790659Z","shell.execute_reply.started":"2025-07-07T02:15:19.781631Z","shell.execute_reply":"2025-07-07T02:15:22.789778Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Checking the distribution of classes\ncounts = train_df['label'].value_counts()\nprint(counts)\nsns.barplot(x=counts.index, y=counts.values)\nplt.xlabel('Label')\nplt.ylabel('Number of Observations')\nplt.title('Distribution of Labels in the Training Data')\nplt.xticks(ticks=[0, 1], labels=['0: No Cancer', '1: Cancer'])\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T02:15:34.010558Z","iopub.execute_input":"2025-07-07T02:15:34.011348Z","iopub.status.idle":"2025-07-07T02:15:34.436605Z","shell.execute_reply.started":"2025-07-07T02:15:34.011300Z","shell.execute_reply":"2025-07-07T02:15:34.435973Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_bal = (\n    pd.concat([\n        train_df.query('label == 1'),\n        train_df.query('label == 0').sample(n=len(train_df[train_df['label'] == 1]), random_state=42)\n    ])\n    .sample(frac=1, random_state=42)      # shuffle\n    .reset_index(drop=True)\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T02:15:36.467619Z","iopub.execute_input":"2025-07-07T02:15:36.467895Z","iopub.status.idle":"2025-07-07T02:15:36.603870Z","shell.execute_reply.started":"2025-07-07T02:15:36.467875Z","shell.execute_reply":"2025-07-07T02:15:36.602863Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Checking the distribution of classes\nbal_counts = train_bal['label'].value_counts()\nprint(bal_counts)\nsns.barplot(x=bal_counts.index, y=bal_counts.values)\nplt.xlabel('Label')\nplt.ylabel('Number of Observations')\nplt.title('Distribution of Labels in the Balanced Training Data')\nplt.xticks(ticks=[0, 1], labels=['0: No Cancer', '1: Cancer'])\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T02:15:38.379210Z","iopub.execute_input":"2025-07-07T02:15:38.379522Z","iopub.status.idle":"2025-07-07T02:15:38.491583Z","shell.execute_reply.started":"2025-07-07T02:15:38.379501Z","shell.execute_reply":"2025-07-07T02:15:38.490744Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Model Architecture\n\nFor this task, I built a custom convolutional neural network (CNN) designed to balance strong performance with computational efficiency. Below, I explain the architecture, justify the design choices, and compare it against alternatives.\n\n\n* The CNN architecture consists of 3 convolutional blocks. Each block includes two 3×3 convolutional layers followed by batch normalization and LeakyReLU activation. This structure captures local textures like cell boundaries and tissue structure while keeping training stable and fast.\n\n* After each convolutional block, a max pooling layer is used to downsample the feature maps, and a SpatialDropout2D layer with 20% dropout is added to reduce overfitting. Spatial dropout helps prevent co-adaptation of entire feature maps, which is particularly helpful when working with medical images.\n\n* The number of filters starts at 32 and doubles in each block (32 → 64 → 128), allowing the model to learn increasingly abstract features.\n\n* After the convolutional blocks, the model flattens the output and passes it through a fully connected dense layer with 256 units. Batch normalization, LeakyReLU activation, and a dropout of 30% are used here to promote generalization.\n\n* The final layer is a single dense neuron with a sigmoid activation to produce a binary cancer vs. non-cancer probability.\n\n* L2 regularization is applied to all convolutional and dense layers to penalize large weights and further prevent overfitting.\n\n* The optimizer used is Adam with a learning rate of 1e-3, chosen for its fast convergence and adaptability.\n\n","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"# Unified, clean import block\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import (\n    Dense, Dropout, Flatten, BatchNormalization, Activation,\n    Conv2D, MaxPooling2D, LeakyReLU, SpatialDropout2D\n)\nfrom tensorflow.keras import regularizers, layers,models\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.callbacks import EarlyStopping, ReduceLROnPlateau, ModelCheckpoint\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator \n\n\n\n\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T02:22:56.061468Z","iopub.execute_input":"2025-07-07T02:22:56.062161Z","iopub.status.idle":"2025-07-07T02:22:56.066655Z","shell.execute_reply.started":"2025-07-07T02:22:56.062137Z","shell.execute_reply":"2025-07-07T02:22:56.065775Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"IMG_SIZE    = (96, 96)\nBATCH_SIZE  = 64\nVAL_SPLIT   = 0.20\nRANDOM_SEED = 42\nEPOCHS      = 50               # upper bound, EarlyStopping will stop sooner\n\nbal_train_df, val_df = train_test_split(\n    train_bal,\n    test_size   = VAL_SPLIT,\n    stratify    = train_bal['label'],\n    random_state= RANDOM_SEED\n)\n\n\n\n\ntrain_gen = ImageDataGenerator(\n    rescale            = 1/255.,\n    rotation_range     = 20,\n    width_shift_range  = 0.2,\n    height_shift_range = 0.2,\n    horizontal_flip    = True,\n    vertical_flip      = True,\n    zoom_range         = 0.2,\n    shear_range        = 0.2,\n    fill_mode          = 'nearest'\n)\nval_gen = ImageDataGenerator(rescale=1/255.)\ntest_gen = ImageDataGenerator(rescale=1/255.)\n\ntrain_flow = train_gen.flow_from_dataframe(\n    dataframe   = bal_train_df,              \n    x_col       = 'filepath',\n    y_col       = 'label',\n    target_size = IMG_SIZE,\n    batch_size  = BATCH_SIZE,\n    class_mode  = 'raw',\n    shuffle     = True,\n    validate_filenames=False #Comment out if you want to check file validation\n)\n\nval_flow = test_gen.flow_from_dataframe(\n    dataframe   = val_df,\n    x_col       = 'filepath',\n    y_col       = 'label',\n    target_size = IMG_SIZE,\n    batch_size  = BATCH_SIZE,\n    class_mode  = 'raw',\n    shuffle     = False,\n    validate_filenames=False #Comment out if you want to check file validation\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-07T02:16:13.699639Z","iopub.execute_input":"2025-07-07T02:16:13.700342Z","iopub.status.idle":"2025-07-07T02:16:13.950927Z","shell.execute_reply.started":"2025-07-07T02:16:13.700318Z","shell.execute_reply":"2025-07-07T02:16:13.950361Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def build_cnn(num_blocks    = 3,\n              start_filters = 32,\n              dense_units   = 256,\n              dropout_conv  = 0.2,\n              dropout_dense = 0.3,\n              l2_reg        = 1e-4,\n              lr            = 1e-3):\n    \"\"\"\n    Build & compile a small-to-medium CNN.\n    Args control width/depth so you can grid-search later.\n    \"\"\"\n    inputs = layers.Input(shape=(*IMG_SIZE, 3))\n    \n    x = inputs\n    filters = start_filters\n    for b in range(num_blocks):\n        x = layers.Conv2D(filters, 3, padding='same',\n                          kernel_regularizer=regularizers.l2(l2_reg))(x)\n        x = layers.BatchNormalization()(x)\n        x = layers.LeakyReLU()(x)\n        x = layers.Conv2D(filters, 3, padding='same',\n                          kernel_regularizer=regularizers.l2(l2_reg))(x)\n        x = layers.BatchNormalization()(x)\n        x = layers.LeakyReLU()(x)\n        x = layers.MaxPooling2D()(x)\n        x = layers.SpatialDropout2D(dropout_conv)(x)\n        filters *= 2                       # double filters each block\n\n    x = layers.Flatten()(x)               # or GlobalAveragePooling2D()\n    x = layers.Dense(dense_units, kernel_regularizer=regularizers.l2(l2_reg))(x)\n    x = layers.BatchNormalization()(x)\n    x = layers.LeakyReLU()(x)\n    x = layers.Dropout(dropout_dense)(x)\n\n    outputs = layers.Dense(1, activation='sigmoid')(x)\n\n    model = models.Model(inputs, outputs)\n    model.compile(\n        optimizer = Adam(learning_rate=lr),\n        loss      = 'binary_crossentropy',\n        metrics   = ['accuracy', tf.keras.metrics.AUC(name='auc')]\n    )\n    return model\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-06T18:38:34.445605Z","iopub.execute_input":"2025-07-06T18:38:34.445920Z","iopub.status.idle":"2025-07-06T18:38:34.453313Z","shell.execute_reply.started":"2025-07-06T18:38:34.445899Z","shell.execute_reply":"2025-07-06T18:38:34.452691Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# callbacks = [\n#     EarlyStopping(patience=3, min_delta=1e-3, restore_best_weights=True, monitor='val_loss'),\n#     ReduceLROnPlateau(factor=0.5, patience=3, min_lr=1e-6, monitor='val_loss'),\n#     ModelCheckpoint('best_cnn.h5', save_best_only=True, monitor='val_loss')\n# ]\n\n# model = build_cnn(num_blocks=3, start_filters=32)\n\n# history = model.fit(\n#     train_flow,\n#     epochs        = 25,\n#     validation_data = val_flow,\n#     callbacks     = callbacks,\n#     verbose       = 1\n# )\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-06T18:38:42.141863Z","iopub.execute_input":"2025-07-06T18:38:42.142144Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import json, os, shutil\nfrom pathlib import Path\nimport tensorflow as tf\n\nSTATE_PATH   = Path(\"/kaggle/working/train_state.json\")\n\n# 👉  point to the file you just added as a dataset input\nWEIGHTS_PATH = Path(\"/kaggle/input/best-cnn-for-cancer-detection/best_cnn.h5\")\n\n# ─── Build model ───────────────────────────────────────────────────────────────\ntf.keras.backend.clear_session()\nmodel = build_cnn(num_blocks=3, start_filters=32)\n\ninitial_epoch = 0\nif WEIGHTS_PATH.exists():\n    model.load_weights(str(WEIGHTS_PATH))\n\n    if STATE_PATH.exists():                 # resume with stored epoch (2nd run+)\n        with open(STATE_PATH) as f:\n            initial_epoch = json.load(f).get(\"epoch\", 0)\n    else:                                   # first run after uploading .h5\n        initial_epoch = 12                  # <-- last completed epoch\n    print(f\"Resuming from epoch {initial_epoch}\")\n\n# ─── Callback that writes train_state.json each epoch ──────────────────────────\nclass EpochTracker(tf.keras.callbacks.Callback):\n    def __init__(self, path):\n        super().__init__()\n        self.path = Path(path)\n\n    def on_epoch_end(self, epoch, logs=None):\n        with open(self.path, \"w\") as f:\n            json.dump({\"epoch\": epoch + 1}, f)\n\ntracker_cb = EpochTracker(STATE_PATH)\n\n# ─── Standard callbacks (EarlyStopping, LR plateau, checkpoint) ────────────────\ncallbacks = [\n    tf.keras.callbacks.EarlyStopping(\n        monitor=\"val_loss\", patience=3, min_delta=1e-3, restore_best_weights=True),\n    tf.keras.callbacks.ReduceLROnPlateau(\n        monitor=\"val_loss\", factor=0.5, patience=3, min_lr=1e-6),\n    tf.keras.callbacks.ModelCheckpoint(\n        \"/kaggle/working/best_cnn.h5\", save_best_only=True, monitor=\"val_loss\"),\n    tracker_cb,\n]\n\n# ─── Train (will stop early after ≤3 flat epochs) ──────────────────────────────\nhistory = model.fit(\n    train_flow,\n    validation_data=val_flow,\n    epochs=25,\n    initial_epoch=initial_epoch,\n    callbacks=callbacks,\n    verbose=1,\n)\n\n# ─── Move large checkpoint out of /kaggle/working before commit packs files ────\n# if Path(\"/kaggle/working/best_cnn.h5\").exists():\n#     shutil.move(\"/kaggle/working/best_cnn.h5\", \"/kaggle/temp/best_cnn.h5\")\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Results and Analysis\n\nI used early stopping with a patience of 3 and a minimum delta of 1e-3 to stop training if validation loss did not improve.\n\nA learning rate scheduler reduced the learning rate by half if the validation loss plateaued for 3 epochs.\n\nThe best model weights (based on validation loss) were saved automatically using a model checkpoint callback.\n\nI also implemented a custom callback to track the last completed epoch and save it in a JSON file so that training could resume in case of interruptions.\n\nI experimented with a deeper VGG-style model (4 convolutional blocks without batch normalization), but it overfitted quickly and performed worse on validation data.\n\nI also tried transfer learning using ResNet-18 and EfficientNet-B0. These models achieved slightly higher AUC scores, but they required larger input sizes (e.g., 224x224), significantly more GPU memory, and 2-3× longer training times.\n\nFor the purposes of this assignment and within Kaggle's 30 GPU-hour quota, the custom 3-block CNN offered the best trade-off between performance and efficiency.\n\nI tuned key hyperparameters including the number of convolutional blocks, number of filters, dropout rates, L2 regularization strength, and learning rate using random search on a 10k image subset.\n\nThe final configuration (3 blocks, starting with 32 filters, 256 dense units, dropout rates of 0.2 and 0.3, L2 = 1e-4, and Adam optimizer with lr = 1e-3) consistently gave strong validation performance while training in under 10 minutes per epoch on a P100 GPU.","metadata":{}},{"cell_type":"code","source":"plt.plot(history.history['val_auc'], label='val AUC')\nplt.plot(history.history['val_accuracy'], label='val Acc')\nplt.legend(); plt.title('Validation metrics'); plt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-06T18:26:03.319733Z","iopub.status.idle":"2025-07-06T18:26:03.319986Z","shell.execute_reply.started":"2025-07-06T18:26:03.319871Z","shell.execute_reply":"2025-07-06T18:26:03.319883Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def summarize_run(name, model, history):\n    best_idx = np.argmin(history.history['val_loss'])\n    return {\n        'model'      : name,\n        'val_loss'   : history.history['val_loss'][best_idx],\n        'val_acc'    : history.history['val_accuracy'][best_idx],\n        'val_auc'    : history.history['val_auc'][best_idx],\n        'params'     : model.count_params()\n    }\n\nresults = []\nresults.append(summarize_run(\"3-block-32f\", model, history))\npd.DataFrame(results)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-06T18:26:03.320760Z","iopub.status.idle":"2025-07-06T18:26:03.321054Z","shell.execute_reply.started":"2025-07-06T18:26:03.320890Z","shell.execute_reply":"2025-07-06T18:26:03.320904Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- test generator -------------------------------------------------\ntest_flow = test_gen.flow_from_dataframe(\n    dataframe  = test_df,\n    x_col      = \"filepath\",\n    y_col      = None,\n    class_mode = None,\n    target_size= IMG_SIZE,\n    batch_size = BATCH_SIZE,\n    shuffle    = False\n)\n\n# --- predict & write submission ------------------------------------\npreds = model.predict(test_flow, verbose=1).ravel()\nsubmission = pd.DataFrame({\"id\": test_df[\"id\"], \"label\": preds})\nsubmission.to_csv(\"submission.csv\", index=False)\nsubmission.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-06T21:55:10.237980Z","iopub.execute_input":"2025-07-06T21:55:10.238250Z","iopub.status.idle":"2025-07-06T21:57:26.494628Z","shell.execute_reply.started":"2025-07-06T21:55:10.238229Z","shell.execute_reply":"2025-07-06T21:57:26.493944Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Conclusion:\n\nThis architecture was selected because it provides a strong balance between accuracy, regularization, and efficiency. It generalizes well without requiring deep pretrained models, and the data augmentation strategy further enhances its robustness. For histopathologic cancer detection on small image tiles, this approach is highly suitable both in terms of learning capacity and resource constraints.\n\nIssues that I ran into were number of samples leading into millions of parameters which took time to process.  The callback helped by reducing the number of epochs to do, but I easily spent more than 10 hours building the CNN.  Saving the best model helped for future run and recommend saving datasets in the future due to the large time to compile.  ","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-06T21:54:50.942155Z","iopub.execute_input":"2025-07-06T21:54:50.942446Z","iopub.status.idle":"2025-07-06T21:54:50.946549Z","shell.execute_reply.started":"2025-07-06T21:54:50.942423Z","shell.execute_reply":"2025-07-06T21:54:50.945761Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}