{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":11848,"databundleVersionId":862157,"sourceType":"competition"}],"dockerImageVersionId":31012,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\n\n# Note: Ignore the file walk as there are thousands of files.\n#for dirname, _, filenames in os.walk('/kaggle/input'):\n#    for filename in filenames:\n#        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-05-18T01:17:14.789051Z","iopub.execute_input":"2025-05-18T01:17:14.789326Z","iopub.status.idle":"2025-05-18T01:17:15.109329Z","shell.execute_reply.started":"2025-05-18T01:17:14.789302Z","shell.execute_reply":"2025-05-18T01:17:15.108574Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# CNN Cancer Detection\n\nThis notebook explores the Histopathologic Cancer Detection from Cukierski (2018) to identify metastatic tissue in histopathologic lymph node-section scans. It is a modified subset of PatchCamelyon (PCam) without duplicates (later, we do not need to check for duplicates). This project aims to output that the center 32x32px region of a patch contains at least one pixel of tumor tissue.","metadata":{}},{"cell_type":"markdown","source":"# Exploratory Data Analysis (EDA)","metadata":{}},{"cell_type":"markdown","source":"Firstly, let us understand the data we have been given. Let us explore the labels.","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\ndf = pd.read_csv('/kaggle/input/histopathologic-cancer-detection/train_labels.csv')\ndf.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T05:26:58.360761Z","iopub.execute_input":"2025-05-20T05:26:58.361058Z","iopub.status.idle":"2025-05-20T05:26:58.730118Z","shell.execute_reply.started":"2025-05-20T05:26:58.361036Z","shell.execute_reply":"2025-05-20T05:26:58.729458Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let us check: What is the distribution of the labels (i.e., count)? Are there any null values (just checking)?","metadata":{}},{"cell_type":"code","source":"print(df['label'].value_counts())\nprint(df.isnull().sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-19T07:29:49.777223Z","iopub.execute_input":"2025-05-19T07:29:49.777702Z","iopub.status.idle":"2025-05-19T07:29:49.800804Z","shell.execute_reply.started":"2025-05-19T07:29:49.777680Z","shell.execute_reply":"2025-05-19T07:29:49.800061Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let us visualize that distribution to see if there is any imbalance.","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\n\nlabel_counts = df['label'].value_counts()\ntotal = len(df)\n\n# Plot\nplt.figure(figsize=(6,4))\nax = sns.countplot(data=df, x='label')\n\nfor p in ax.patches:\n    count = p.get_height()\n    percent = f'{100 * count / total:.2f}%'\n    x = p.get_x() + p.get_width() / 2\n    y = p.get_height()\n    ax.text(x, y + total * 0.01, percent, ha='center', va='bottom', fontsize=12)\n\n# Format\nplt.title(\"Label Distribution\")\nplt.xlabel(\"Label\")\nplt.ylabel(\"Count\")\nplt.ylim(0, label_counts.max() * 1.1)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-18T01:17:15.349119Z","iopub.execute_input":"2025-05-18T01:17:15.349743Z","iopub.status.idle":"2025-05-18T01:17:16.146016Z","shell.execute_reply.started":"2025-05-18T01:17:15.349718Z","shell.execute_reply":"2025-05-18T01:17:16.145122Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This chart shows that 60% are non-cancerous and 40% are cancerous. There is a slight imbalance of data, but not too much. We'll continue as-is, but we may do some additional monitoring of the results (not just accuracy).\n\nMoving on to the actual images, let us see how they look.","metadata":{}},{"cell_type":"code","source":"import cv2\nimport numpy as np\nimport os\n\ndef load_image(image_id, base_path='/kaggle/input/histopathologic-cancer-detection/train'):\n    path = os.path.join(base_path, f\"{image_id}.tif\")\n    return cv2.imread(path)\n\ndef show_samples(df, label, n=5):\n    samples = df[df['label'] == label].sample(n)\n    fig, axes = plt.subplots(1, n, figsize=(15, 5))\n    for img_id, ax in zip(samples['id'], axes):\n        img = load_image(img_id)\n        ax.imshow(cv2.cvtColor(img, cv2.COLOR_BGR2RGB))\n        ax.axis('off')\n    plt.suptitle(f\"Label: {label}\")\n    plt.show()\n\nshow_samples(df, label=0)\nshow_samples(df, label=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-18T01:17:16.146936Z","iopub.execute_input":"2025-05-18T01:17:16.147330Z","iopub.status.idle":"2025-05-18T01:17:16.754634Z","shell.execute_reply.started":"2025-05-18T01:17:16.147304Z","shell.execute_reply":"2025-05-18T01:17:16.753830Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"From this, the images are all the same size. We can get the actual size of an image by using 'shape'. This will tell us what is the width, height, and channels (RGB).","metadata":{}},{"cell_type":"code","source":"img = load_image(df['id'][0])\nprint(\"Image shape:\", img.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-18T01:17:16.755550Z","iopub.execute_input":"2025-05-18T01:17:16.755822Z","iopub.status.idle":"2025-05-18T01:17:16.761332Z","shell.execute_reply.started":"2025-05-18T01:17:16.755804Z","shell.execute_reply":"2025-05-18T01:17:16.760522Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"So, it's a 96x96-pixel image with RGB (3 channels). Next, we want to check the distribution of pixels' colors in a sample image (also known as intensity). Can we see any pattern differences?","metadata":{}},{"cell_type":"code","source":"def plot_color_distribution_histogram(image):\n    colors = ['r', 'g', 'b']\n    for i, color in enumerate(colors):\n        plt.hist(image[..., i].ravel(), bins=256, color=color, alpha=0.5)\n    plt.title(\"Pixel Intensity Distribution\")\n    plt.xlabel(\"Pixel value\")\n    plt.ylabel(\"Frequency\")\n    plt.show()\n\nimg_tumor = load_image(df[df['label'] == 1].sample(1).iloc[0]['id'])\nimg_normal = load_image(df[df['label'] == 0].sample(1).iloc[0]['id'])\n\nplot_color_distribution_histogram(img_tumor)\nplot_color_distribution_histogram(img_normal)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-18T01:17:16.762221Z","iopub.execute_input":"2025-05-18T01:17:16.762524Z","iopub.status.idle":"2025-05-18T01:17:18.607157Z","shell.execute_reply.started":"2025-05-18T01:17:16.762505Z","shell.execute_reply":"2025-05-18T01:17:18.606504Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"From this, we can see that when cancer is present, there should be deeper colors—whereas when there is not, the colors are more spread and less concentrated. But this is only one sample, so we shouldn't extrapolate too much meaning right now. Since we know there are no duplicates or nulls, we do not need to clean the data.","metadata":{}},{"cell_type":"code","source":"def tissue_ratio(img, threshold=200):\n    gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)\n    return np.mean(gray < threshold)\n\n# Use only a sample for EDA\nsample_ids = df['id'].sample(2000, random_state=42)\n\nratios = sample_ids.apply(lambda id_: tissue_ratio(load_image(id_)))\ndf['tissue_ratio'] = ratios\n\nsns.histplot(data=df, x='tissue_ratio', hue='label', bins=30, kde=True)\nplt.title(\"Tissue Area Ratio Distribution\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-18T01:17:18.607904Z","iopub.execute_input":"2025-05-18T01:17:18.608194Z","iopub.status.idle":"2025-05-18T01:17:22.896851Z","shell.execute_reply.started":"2025-05-18T01:17:18.608167Z","shell.execute_reply":"2025-05-18T01:17:22.896233Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We have taken a sample of 2000 images here. Looking at the tissue area, we can see that the images with cancer are more distributed, with a higher ratio of tissue. In contrast, the images for non-cancerous tissue are spread more. This chart supports the previous finding from the pixel intensity graph.","metadata":{}},{"cell_type":"markdown","source":"**Set up for importing images**\n\nTo import the images, we need to add additional columns such as filename (which includes the ext) and convert the label to a string (required for ImageDataGenerator).","metadata":{}},{"cell_type":"code","source":"df['label_str'] = df['label'].astype(str)\ndf['filename'] = df['id'] + '.tif'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T05:27:05.935258Z","iopub.execute_input":"2025-05-20T05:27:05.935547Z","iopub.status.idle":"2025-05-20T05:27:06.020054Z","shell.execute_reply.started":"2025-05-20T05:27:05.935525Z","shell.execute_reply":"2025-05-20T05:27:06.019504Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Split the dataset**\n\nThe last thing we want to do is split the dataset so we can train the models. Here, we first take a sample of the dataset df_sampled and then use this sample to create a test/train split.\n\nWe only took a sample because if we loaded the whole dataset, Kaggle would crash. I've chosen 40K here, as this allows for faster iterations over the models.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\ndf_sampled, _ = train_test_split(\n    df,\n    train_size=40000,\n    stratify=df['label'],\n    random_state=42\n)\n\ntrain_df, val_df = train_test_split(\n    df_sampled,\n    test_size=0.2,                # 20% for validation\n    stratify=df_sampled['label'],         # maintain class balance\n    random_state=42               # reproducibility\n)\n\nprint(\"Train label distribution:\")\nprint(train_df['label'].value_counts(normalize=True))\n\nprint(\"\\nValidation label distribution:\")\nprint(val_df['label'].value_counts(normalize=True))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T05:27:08.822866Z","iopub.execute_input":"2025-05-20T05:27:08.823229Z","iopub.status.idle":"2025-05-20T05:27:08.981300Z","shell.execute_reply.started":"2025-05-20T05:27:08.823207Z","shell.execute_reply":"2025-05-20T05:27:08.980471Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now, we will load all the images (data) and normalize them. Normalize means changing it from 0-255 to 0-1, so we divide by 255.\n\nWhat is different here is that we load it all into a dataset. This means that if it's already cached and prefetched, it will make the GPU more efficient.","metadata":{}},{"cell_type":"code","source":"import cv2\nimport numpy as np\nimport tensorflow as tf\n\nBATCH_SIZE = 32\n\ndef load_image_cv2(path):\n    img = cv2.imread(path)\n    img = cv2.resize(img, (96, 96))\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n    return img.astype(np.float32) / 255.0\n\ndef data_generator(paths, labels):\n    for path, label in zip(paths, labels):\n        img = load_image_cv2(path)\n        yield img, label\n\ntrain_paths = [f\"/kaggle/input/histopathologic-cancer-detection/train/{f}\" for f in train_df['filename']]\ntrain_labels = train_df['label'].values.astype(np.float32)\n\nval_paths = [f\"/kaggle/input/histopathologic-cancer-detection/train/{f}\" for f in val_df['filename']]\nval_labels = val_df['label'].values.astype(np.float32)\n\ntrain_data = tf.data.Dataset.from_generator(\n    lambda: data_generator(train_paths, train_labels),\n    output_signature=(\n        tf.TensorSpec(shape=(96, 96, 3), dtype=tf.float32),\n        tf.TensorSpec(shape=(), dtype=tf.float32)\n    )\n).shuffle(1024).repeat().batch(BATCH_SIZE, drop_remainder=True).prefetch(tf.data.AUTOTUNE)\n\nval_data = tf.data.Dataset.from_generator(\n    lambda: data_generator(val_paths, val_labels),\n    output_signature=(\n        tf.TensorSpec(shape=(96, 96, 3), dtype=tf.float32),\n        tf.TensorSpec(shape=(), dtype=tf.float32)\n    )\n).batch(BATCH_SIZE, drop_remainder=True).cache().prefetch(tf.data.AUTOTUNE)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T05:27:13.421030Z","iopub.execute_input":"2025-05-20T05:27:13.421746Z","iopub.status.idle":"2025-05-20T05:27:13.526708Z","shell.execute_reply.started":"2025-05-20T05:27:13.421720Z","shell.execute_reply":"2025-05-20T05:27:13.525993Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The code below ensures we've loaded the dataset correctly. Previously, I had problems with data labelled as all zero or all one.","metadata":{}},{"cell_type":"code","source":"for images, labels in val_data.take(1):\n    print(np.unique(labels.numpy(), return_counts=True))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T03:30:09.803413Z","iopub.execute_input":"2025-05-20T03:30:09.803955Z","iopub.status.idle":"2025-05-20T03:30:10.603175Z","shell.execute_reply.started":"2025-05-20T03:30:09.803932Z","shell.execute_reply":"2025-05-20T03:30:10.602412Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Model Architecture","metadata":{}},{"cell_type":"markdown","source":"I will try two architectures with this CNN problem. The first is a basic model with a few layers and batch normalization. The second is the VGGNet that we learned about in class.","metadata":{}},{"cell_type":"markdown","source":"# Architecture 1: Basic Model with Batch Normalization\n\nHere, we will create a model with only a few layers and normalization between each layer. This should be easy to train and shouldn't overfit too much with the batch normalization between the layers. This architecture is similar to the VGGNet architecture, but without as many layers.","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras import layers, models, callbacks\n\ndef build_basic(input_shape=(96, 96, 3)):\n    model = models.Sequential(name=\"Basic_Model\")\n    model.add(layers.Input(shape=input_shape))\n\n    model.add(layers.Conv2D(32, (3,3), activation='relu', padding='same'))\n    model.add(layers.BatchNormalization())\n\n    model.add(layers.Conv2D(64, (3,3), activation='relu', padding='same'))\n    model.add(layers.BatchNormalization())\n\n    model.add(layers.MaxPooling2D(pool_size=(2,2)))\n    model.add(layers.BatchNormalization())\n\n    model.add(layers.Flatten())\n    model.add(layers.Dense(1, activation='sigmoid'))\n    return model\n\nbasic_model = build_basic()\nbasic_model.summary()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-19T04:55:00.943485Z","iopub.execute_input":"2025-05-19T04:55:00.943791Z","iopub.status.idle":"2025-05-19T04:55:02.209029Z","shell.execute_reply.started":"2025-05-19T04:55:00.943768Z","shell.execute_reply":"2025-05-19T04:55:02.208405Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Architecture 2: VGGNet\n\nThis follows the VGGNet architecture, which has multiple convolutional layers followed by a max-pool layer. This is then repeated N times (for simplicity, I've repeated it only 3 times). Inside the hidden layers, we will use relu and sigmoid to output the final result. We'll use dropout to help regulate the final output and ensure no overfitting.","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras import layers, models\n\ndef build_vgg_like(input_shape=(96, 96, 3)):\n    model = models.Sequential(name=\"VGGNet\")\n    model.add(layers.Input(shape=input_shape))\n    \n    model.add(layers.Conv2D(32, (3,3), activation='relu', padding='same'))\n    model.add(layers.Conv2D(32, (3,3), activation='relu', padding='same'))\n    model.add(layers.MaxPooling2D(pool_size=(2,2)))\n    \n    model.add(layers.Conv2D(64, (3,3), activation='relu', padding='same'))\n    model.add(layers.Conv2D(64, (3,3), activation='relu', padding='same'))\n    model.add(layers.MaxPooling2D(pool_size=(2,2)))\n    \n    model.add(layers.Conv2D(128, (3,3), activation='relu', padding='same'))\n    model.add(layers.Conv2D(128, (3,3), activation='relu', padding='same'))\n    model.add(layers.MaxPooling2D(pool_size=(2,2)))\n\n    model.add(layers.Flatten())\n    model.add(layers.Dense(256, activation='relu'))\n    model.add(layers.Dense(256, activation='relu'))\n    model.add(layers.Dropout(0.5))\n    model.add(layers.Dense(1, activation='sigmoid'))\n    \n    return model\n\nvgg_model = build_vgg_like()\nvgg_model.summary()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T03:30:20.535337Z","iopub.execute_input":"2025-05-20T03:30:20.535610Z","iopub.status.idle":"2025-05-20T03:30:21.894595Z","shell.execute_reply.started":"2025-05-20T03:30:20.535588Z","shell.execute_reply":"2025-05-20T03:30:21.894044Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Results and Analysis\n\nNow, it's time to run our models and see if we can improve them. Firstly, I'll run the model for the basic architecture. We will then use the confusion matrix and the Area Under the Curve (AUC) to check for the model's accuracy. Since this is about medical data, I'll also want to know the false negative rate.\n\nAfter this, we will run the second model (VGGNet) and perform the same evaluation. Finally, I'll perform some hyperparameter tuning on the best model, such as deciding what optimizer or loss to use.\n\nAll my models need to know how many steps per epoch to run. Knowing the number of steps prevents them from running out of data.","metadata":{}},{"cell_type":"code","source":"steps_per_epoch = len(train_df) // BATCH_SIZE","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T03:30:27.518173Z","iopub.execute_input":"2025-05-20T03:30:27.518446Z","iopub.status.idle":"2025-05-20T03:30:27.522032Z","shell.execute_reply.started":"2025-05-20T03:30:27.518424Z","shell.execute_reply":"2025-05-20T03:30:27.521408Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Basic Model","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras import layers, models\n\nbasic_model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])\n\ncallbacks_list = [\n    # This helps stop the model early\n    callbacks.EarlyStopping(\n        monitor='val_loss',\n        patience=3,\n        restore_best_weights=True\n    ),\n    # Reduces learning rate when plateued\n    callbacks.ReduceLROnPlateau(\n        monitor='val_loss',\n        factor=0.2,\n        patience=3,\n        min_lr=1e-6,\n        verbose=1\n    ),\n    # save best model\n    callbacks.ModelCheckpoint(\n        filepath='best_basic_model.keras',\n        monitor='val_accuracy',\n        save_best_only=True,\n        verbose=1\n    )\n]\n\nbm_history = basic_model.fit(\n    train_data, \n    validation_data = val_data, \n    epochs=10, \n    steps_per_epoch=steps_per_epoch,\n    callbacks=callbacks_list\n)\nbasic_model.evaluate(val_data)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-19T05:23:15.110513Z","iopub.execute_input":"2025-05-19T05:23:15.110918Z","iopub.status.idle":"2025-05-19T05:27:32.393807Z","shell.execute_reply.started":"2025-05-19T05:23:15.110892Z","shell.execute_reply":"2025-05-19T05:27:32.393001Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now that we have run our first model, let us plot the history for accuracy and loss.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# Accuracy plot\nplt.figure(figsize=(12, 5))\n\nplt.subplot(1, 2, 1)\nplt.plot(bm_history.history['accuracy'], label='Train Accuracy')\nplt.plot(bm_history.history['val_accuracy'], label='Val Accuracy')\nplt.title('Model Accuracy')\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.legend()\n\n# Loss plot\nplt.subplot(1, 2, 2)\nplt.plot(bm_history.history['loss'], label='Train Loss')\nplt.plot(bm_history.history['val_loss'], label='Val Loss')\nplt.title('Model Loss')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-19T05:27:39.177476Z","iopub.execute_input":"2025-05-19T05:27:39.178077Z","iopub.status.idle":"2025-05-19T05:27:39.537615Z","shell.execute_reply.started":"2025-05-19T05:27:39.178051Z","shell.execute_reply":"2025-05-19T05:27:39.536683Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"From this graph, we can see that the val accuracy initially increased but then decreased. The validation loss increases sharply when the accuracy decreases. The charts show that the model is likely overfitting and may not be good.\n\nThe next function helps get the predicted and true values for the labels. This is useful for creating the confusion matrix and AUC.","metadata":{}},{"cell_type":"code","source":"import numpy as np\n\ndef get_y_true_and_y_pred(val_data, model, threshold):\n\n    # Predict on validation set\n    y_true = []\n    y_pred_probs = []\n    \n    for batch_x, batch_y in val_data:\n        preds = model.predict(batch_x, verbose=0)\n        y_pred_probs.extend(preds.ravel())  # Flatten to 1D list\n        y_true.extend(batch_y.numpy().ravel())  # Convert from tensor to NumPy array\n    \n    # Convert to NumPy arrays\n    y_true = np.array(y_true)\n    y_pred_probs = np.array(y_pred_probs)\n    \n    # Convert probabilities to binary predictions\n    y_pred = (y_pred_probs > threshold).astype(int)\n    return y_true, y_pred","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T05:25:53.682415Z","iopub.execute_input":"2025-05-20T05:25:53.682750Z","iopub.status.idle":"2025-05-20T05:25:53.688962Z","shell.execute_reply.started":"2025-05-20T05:25:53.682726Z","shell.execute_reply":"2025-05-20T05:25:53.687925Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"y_true_bm, y_pred_bm = get_y_true_and_y_pred(val_data, basic_model, 0.5)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-19T05:27:49.432175Z","iopub.execute_input":"2025-05-19T05:27:49.432824Z","iopub.status.idle":"2025-05-19T05:28:04.036025Z","shell.execute_reply.started":"2025-05-19T05:27:49.432799Z","shell.execute_reply":"2025-05-19T05:28:04.035089Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix, ConfusionMatrixDisplay\n\n# Confusion Matrix\ncm = confusion_matrix(y_true_bm, y_pred_bm)\ndisp = ConfusionMatrixDisplay(confusion_matrix=cm)\ndisp.plot(cmap='Blues')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-19T05:28:07.414760Z","iopub.execute_input":"2025-05-19T05:28:07.415478Z","iopub.status.idle":"2025-05-19T05:28:07.640064Z","shell.execute_reply.started":"2025-05-19T05:28:07.415451Z","shell.execute_reply":"2025-05-19T05:28:07.639251Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The confusion matrix shows that there are still many false negatives, which is not great for medical projects.","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import roc_curve, auc\n\nfpr, tpr, _ = roc_curve(y_true_bm, y_pred_probs_bm)\nroc_auc = auc(fpr, tpr)\n\nplt.figure()\nplt.plot(fpr, tpr, label=f'ROC curve (AUC = {roc_auc:.2f})')\nplt.plot([0, 1], [0, 1], linestyle='--', color='gray')\nplt.xlabel('False Positive Rate')\nplt.ylabel('True Positive Rate (Recall)')\nplt.title('Receiver Operating Characteristic')\nplt.legend(loc='lower right')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-19T05:28:15.300051Z","iopub.execute_input":"2025-05-19T05:28:15.300395Z","iopub.status.idle":"2025-05-19T05:28:15.490467Z","shell.execute_reply.started":"2025-05-19T05:28:15.300369Z","shell.execute_reply":"2025-05-19T05:28:15.489517Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The ROC curve plot shows that the AUC is 0.84, which is not bad but also not great.","metadata":{}},{"cell_type":"markdown","source":"## VGGNet Model","metadata":{}},{"cell_type":"code","source":"train_data = tf.data.Dataset.from_generator(\n    lambda: data_generator(train_paths, train_labels),\n    output_signature=(\n        tf.TensorSpec(shape=(96, 96, 3), dtype=tf.float32),\n        tf.TensorSpec(shape=(), dtype=tf.float32)\n    )\n).shuffle(1024).repeat().batch(BATCH_SIZE, drop_remainder=True).prefetch(tf.data.AUTOTUNE)\n\nval_data = tf.data.Dataset.from_generator(\n    lambda: data_generator(val_paths, val_labels),\n    output_signature=(\n        tf.TensorSpec(shape=(96, 96, 3), dtype=tf.float32),\n        tf.TensorSpec(shape=(), dtype=tf.float32)\n    )\n).batch(BATCH_SIZE, drop_remainder=True).cache().prefetch(tf.data.AUTOTUNE)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T03:33:21.625136Z","iopub.execute_input":"2025-05-20T03:33:21.625423Z","iopub.status.idle":"2025-05-20T03:33:21.688522Z","shell.execute_reply.started":"2025-05-20T03:33:21.625401Z","shell.execute_reply":"2025-05-20T03:33:21.687762Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras import layers, models, callbacks\n\nvgg_model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])\n\ncallbacks_list = [\n    callbacks.EarlyStopping(\n        monitor='val_loss',\n        patience=3,\n        restore_best_weights=True\n    ),\n    callbacks.ReduceLROnPlateau(\n        monitor='val_loss',\n        factor=0.2,\n        patience=3,\n        min_lr=1e-6\n    ),\n    # save best model\n    callbacks.ModelCheckpoint(\n        filepath='best_vgg_model.keras',\n        monitor='val_accuracy',\n        save_best_only=True,\n        verbose=1\n    )\n]\n\nvgg_history = vgg_model.fit(\n    train_data, \n    validation_data = val_data, \n    epochs=10,\n    steps_per_epoch=steps_per_epoch,\n    callbacks=callbacks_list\n)\nvgg_model.evaluate(val_data)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T03:33:25.063627Z","iopub.execute_input":"2025-05-20T03:33:25.063918Z","iopub.status.idle":"2025-05-20T03:47:42.404572Z","shell.execute_reply.started":"2025-05-20T03:33:25.063895Z","shell.execute_reply":"2025-05-20T03:47:42.404010Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# Accuracy plot\nplt.figure(figsize=(12, 5))\n\nplt.subplot(1, 2, 1)\nplt.plot(vgg_history.history['accuracy'], label='Train Accuracy')\nplt.plot(vgg_history.history['val_accuracy'], label='Val Accuracy')\nplt.title('Model Accuracy')\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.legend()\n\n# Loss plot\nplt.subplot(1, 2, 2)\nplt.plot(vgg_history.history['loss'], label='Train Loss')\nplt.plot(vgg_history.history['val_loss'], label='Val Loss')\nplt.title('Model Loss')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T03:48:11.164822Z","iopub.execute_input":"2025-05-20T03:48:11.165121Z","iopub.status.idle":"2025-05-20T03:48:11.574093Z","shell.execute_reply.started":"2025-05-20T03:48:11.165099Z","shell.execute_reply":"2025-05-20T03:48:11.573233Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We notice that the training loss decreases, but the validation loss increases. This suggests overfitting on the training data. Also, the value loss is not going anywhere, which could suggest that the learning rate is too high and overshot.","metadata":{}},{"cell_type":"code","source":"y_true_vgg, y_pred_vgg = get_y_true_and_y_pred(val_data, vgg_model, 0.5)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T03:48:33.674578Z","iopub.execute_input":"2025-05-20T03:48:33.675322Z","iopub.status.idle":"2025-05-20T03:48:47.965417Z","shell.execute_reply.started":"2025-05-20T03:48:33.675295Z","shell.execute_reply":"2025-05-20T03:48:47.964773Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix, ConfusionMatrixDisplay\n\n# Confusion Matrix\ncm = confusion_matrix(y_true_vgg, y_pred_vgg)\ndisp = ConfusionMatrixDisplay(confusion_matrix=cm)\ndisp.plot(cmap='Blues')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T03:48:50.151588Z","iopub.execute_input":"2025-05-20T03:48:50.151862Z","iopub.status.idle":"2025-05-20T03:48:50.321930Z","shell.execute_reply.started":"2025-05-20T03:48:50.151840Z","shell.execute_reply":"2025-05-20T03:48:50.321295Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.metrics import roc_curve, auc\n\nfpr, tpr, _ = roc_curve(y_true_vgg, y_pred_probs_vgg)\nroc_auc = auc(fpr, tpr)\n\nplt.figure()\nplt.plot(fpr, tpr, label=f'ROC curve (AUC = {roc_auc:.2f})')\nplt.plot([0, 1], [0, 1], linestyle='--', color='gray')\nplt.xlabel('False Positive Rate')\nplt.ylabel('True Positive Rate (Recall)')\nplt.title('Receiver Operating Characteristic')\nplt.legend(loc='lower right')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T03:48:53.649225Z","iopub.execute_input":"2025-05-20T03:48:53.649473Z","iopub.status.idle":"2025-05-20T03:48:53.816640Z","shell.execute_reply.started":"2025-05-20T03:48:53.649456Z","shell.execute_reply":"2025-05-20T03:48:53.815897Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now, let us tune the model's hyperparameters. It's already very good, but considering this is medical data, we want to avoid false negatives and minimize them.\n\nFirstly, we want to tune the learning rate. The default is 1e-3, so we will take it down to 1e-4.","metadata":{}},{"cell_type":"markdown","source":"## Tuning: VGGNet Model - Learning Rate","metadata":{}},{"cell_type":"markdown","source":"This block of code resets the iterators.","metadata":{}},{"cell_type":"code","source":"train_data = tf.data.Dataset.from_generator(\n    lambda: data_generator(train_paths, train_labels),\n    output_signature=(\n        tf.TensorSpec(shape=(96, 96, 3), dtype=tf.float32),\n        tf.TensorSpec(shape=(), dtype=tf.float32)\n    )\n).shuffle(1024).repeat().batch(BATCH_SIZE, drop_remainder=True).prefetch(tf.data.AUTOTUNE)\n\nval_data = tf.data.Dataset.from_generator(\n    lambda: data_generator(val_paths, val_labels),\n    output_signature=(\n        tf.TensorSpec(shape=(96, 96, 3), dtype=tf.float32),\n        tf.TensorSpec(shape=(), dtype=tf.float32)\n    )\n).batch(32, drop_remainder=True).cache().prefetch(tf.data.AUTOTUNE)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T04:05:52.467249Z","iopub.execute_input":"2025-05-20T04:05:52.467528Z","iopub.status.idle":"2025-05-20T04:05:52.536672Z","shell.execute_reply.started":"2025-05-20T04:05:52.467506Z","shell.execute_reply":"2025-05-20T04:05:52.535903Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras import layers, models, callbacks\n\nnew_optimizer = tf.keras.optimizers.Adam(learning_rate=1e-4)\n\nvgg_model.compile(optimizer=new_optimizer, loss='binary_crossentropy', metrics=['accuracy'])\n\ncallbacks_list = [\n    callbacks.EarlyStopping(\n        monitor='val_loss',\n        patience=3,\n        restore_best_weights=True\n    ),\n    callbacks.ReduceLROnPlateau(\n        monitor='val_loss',\n        factor=0.2,\n        patience=3,\n        min_lr=1e-6\n    ),\n    # save best model\n    callbacks.ModelCheckpoint(\n        filepath='best_vgg_model_v1.keras',\n        monitor='val_accuracy',\n        save_best_only=True,\n        verbose=1\n    )\n]\n\nvgg_history_v2 = vgg_model.fit(\n    train_data, \n    validation_data = val_data, \n    epochs=10,\n    steps_per_epoch=steps_per_epoch,\n    callbacks=callbacks_list\n)\nvgg_model.evaluate(val_data)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T04:05:56.115952Z","iopub.execute_input":"2025-05-20T04:05:56.116633Z","iopub.status.idle":"2025-05-20T04:10:24.269354Z","shell.execute_reply.started":"2025-05-20T04:05:56.116610Z","shell.execute_reply":"2025-05-20T04:10:24.268771Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# Accuracy plot\nplt.figure(figsize=(12, 5))\n\nplt.subplot(1, 2, 1)\nplt.plot(vgg_history_v2.history['accuracy'], label='Train Accuracy')\nplt.plot(vgg_history_v2.history['val_accuracy'], label='Val Accuracy')\nplt.title('Model Accuracy')\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.legend()\n\n# Loss plot\nplt.subplot(1, 2, 2)\nplt.plot(vgg_history_v2.history['loss'], label='Train Loss')\nplt.plot(vgg_history_v2.history['val_loss'], label='Val Loss')\nplt.title('Model Loss')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T04:27:44.981722Z","iopub.execute_input":"2025-05-20T04:27:44.982494Z","iopub.status.idle":"2025-05-20T04:27:45.302470Z","shell.execute_reply.started":"2025-05-20T04:27:44.982467Z","shell.execute_reply":"2025-05-20T04:27:45.301727Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We can see that the validation loss is still increasing while the train loss is still decreasing. The loss in different directions suggests potential overfitting, so we will adapt the model to add a dropout layer.","metadata":{}},{"cell_type":"markdown","source":"## VGGNet Model Tuning: Additional Dropout Layer","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras import layers, models\n\ndef build_vgg_like_v2(input_shape=(96, 96, 3)):\n    model = models.Sequential(name=\"VGGNetv2\")\n    model.add(layers.Input(shape=input_shape))\n    \n    model.add(layers.Conv2D(32, (3,3), activation='relu', padding='same'))\n    model.add(layers.Conv2D(32, (3,3), activation='relu', padding='same'))\n    model.add(layers.MaxPooling2D(pool_size=(2,2)))\n    \n    model.add(layers.Conv2D(64, (3,3), activation='relu', padding='same'))\n    model.add(layers.Conv2D(64, (3,3), activation='relu', padding='same'))\n    model.add(layers.MaxPooling2D(pool_size=(2,2)))\n    \n    model.add(layers.Conv2D(128, (3,3), activation='relu', padding='same'))\n    model.add(layers.Conv2D(128, (3,3), activation='relu', padding='same'))\n    model.add(layers.MaxPooling2D(pool_size=(2,2)))\n\n    model.add(layers.Flatten())\n    model.add(layers.Dense(256, activation='relu'))\n    model.add(layers.Dropout(0.3))\n    model.add(layers.Dense(256, activation='relu'))\n    model.add(layers.Dropout(0.5))\n    model.add(layers.Dense(1, activation='sigmoid'))\n    \n    return model\n\nvgg_model_v2 = build_vgg_like_v2()\nvgg_model_v2.summary()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T04:04:36.687108Z","iopub.execute_input":"2025-05-20T04:04:36.687401Z","iopub.status.idle":"2025-05-20T04:04:36.831471Z","shell.execute_reply.started":"2025-05-20T04:04:36.687377Z","shell.execute_reply":"2025-05-20T04:04:36.830914Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_data = tf.data.Dataset.from_generator(\n    lambda: data_generator(train_paths, train_labels),\n    output_signature=(\n        tf.TensorSpec(shape=(96, 96, 3), dtype=tf.float32),\n        tf.TensorSpec(shape=(), dtype=tf.float32)\n    )\n).shuffle(1024).repeat().batch(BATCH_SIZE, drop_remainder=True).prefetch(tf.data.AUTOTUNE)\n\nval_data = tf.data.Dataset.from_generator(\n    lambda: data_generator(val_paths, val_labels),\n    output_signature=(\n        tf.TensorSpec(shape=(96, 96, 3), dtype=tf.float32),\n        tf.TensorSpec(shape=(), dtype=tf.float32)\n    )\n).batch(BATCH_SIZE, drop_remainder=True).cache().prefetch(tf.data.AUTOTUNE)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T04:04:43.026166Z","iopub.execute_input":"2025-05-20T04:04:43.026642Z","iopub.status.idle":"2025-05-20T04:04:43.095616Z","shell.execute_reply.started":"2025-05-20T04:04:43.026620Z","shell.execute_reply":"2025-05-20T04:04:43.095041Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"new_optimizer = tf.keras.optimizers.Adam(learning_rate=1e-4)\n\nvgg_model_v2.compile(optimizer=new_optimizer, loss='binary_crossentropy', metrics=['accuracy'])\n\ncallbacks_list = [\n    callbacks.EarlyStopping(\n        monitor='val_loss',\n        patience=3,\n        restore_best_weights=True\n    ),\n    callbacks.ReduceLROnPlateau(\n        monitor='val_loss',\n        factor=0.2,\n        patience=3,\n        min_lr=1e-6\n    ),\n    # save best model\n    callbacks.ModelCheckpoint(\n        filepath='best_vgg_model_v2.keras',\n        monitor='val_accuracy',\n        save_best_only=True,\n        verbose=1\n    )\n]\n\nvgg_history_v3 = vgg_model_v2.fit(\n    train_data, \n    validation_data = val_data, \n    epochs=10,\n    steps_per_epoch=steps_per_epoch,\n    callbacks=callbacks_list\n)\nvgg_model_v2.evaluate(val_data)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-19T06:18:25.988322Z","iopub.execute_input":"2025-05-19T06:18:25.989005Z","iopub.status.idle":"2025-05-19T06:27:08.034057Z","shell.execute_reply.started":"2025-05-19T06:18:25.988978Z","shell.execute_reply":"2025-05-19T06:27:08.033447Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# Accuracy plot\nplt.figure(figsize=(12, 5))\n\nplt.subplot(1, 2, 1)\nplt.plot(vgg_history_v3.history['accuracy'], label='Train Accuracy')\nplt.plot(vgg_history_v3.history['val_accuracy'], label='Val Accuracy')\nplt.title('Model Accuracy')\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.legend()\n\n# Loss plot\nplt.subplot(1, 2, 2)\nplt.plot(vgg_history_v3.history['loss'], label='Train Loss')\nplt.plot(vgg_history_v3.history['val_loss'], label='Val Loss')\nplt.title('Model Loss')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-19T06:27:30.497843Z","iopub.execute_input":"2025-05-19T06:27:30.498397Z","iopub.status.idle":"2025-05-19T06:27:30.833928Z","shell.execute_reply.started":"2025-05-19T06:27:30.498369Z","shell.execute_reply":"2025-05-19T06:27:30.833139Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Finally, we are on the right track. The accuracy for both val and train is heading upwards, and the model loss is heading downwards. For the final tuning, we want to add some L2 regularization to the final dense layers.","metadata":{}},{"cell_type":"markdown","source":"## VGGNet Model Tuning: L2 Regularization","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras import layers, models\n\ndef build_vgg_like_v3(input_shape=(96, 96, 3)):\n    model = models.Sequential(name=\"VGGNetv3\")\n    model.add(layers.Input(shape=input_shape))\n    \n    model.add(layers.Conv2D(32, (3,3), activation='relu', padding='same'))\n    model.add(layers.Conv2D(32, (3,3), activation='relu', padding='same'))\n    model.add(layers.MaxPooling2D(pool_size=(2,2)))\n    \n    model.add(layers.Conv2D(64, (3,3), activation='relu', padding='same'))\n    model.add(layers.Conv2D(64, (3,3), activation='relu', padding='same'))\n    model.add(layers.MaxPooling2D(pool_size=(2,2)))\n    \n    model.add(layers.Conv2D(128, (3,3), activation='relu', padding='same'))\n    model.add(layers.Conv2D(128, (3,3), activation='relu', padding='same'))\n    model.add(layers.MaxPooling2D(pool_size=(2,2)))\n\n    model.add(layers.Flatten())\n    model.add(layers.Dense(256, activation='relu', kernel_regularizer=tf.keras.regularizers.l2(0.001)))\n    model.add(layers.Dropout(0.3))\n    model.add(layers.Dense(256, activation='relu', kernel_regularizer=tf.keras.regularizers.l2(0.001)))\n    model.add(layers.Dropout(0.5))\n    model.add(layers.Dense(1, activation='sigmoid'))\n    \n    return model\n\nvgg_model_v3 = build_vgg_like_v3()\nvgg_model_v3.summary()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-19T08:57:18.671150Z","iopub.execute_input":"2025-05-19T08:57:18.671414Z","iopub.status.idle":"2025-05-19T08:57:20.060396Z","shell.execute_reply.started":"2025-05-19T08:57:18.671393Z","shell.execute_reply":"2025-05-19T08:57:20.059788Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_data = tf.data.Dataset.from_generator(\n    lambda: data_generator(train_paths, train_labels),\n    output_signature=(\n        tf.TensorSpec(shape=(96, 96, 3), dtype=tf.float32),\n        tf.TensorSpec(shape=(), dtype=tf.float32)\n    )\n).shuffle(1024).repeat().batch(BATCH_SIZE, drop_remainder=True).prefetch(tf.data.AUTOTUNE)\n\nval_data = tf.data.Dataset.from_generator(\n    lambda: data_generator(val_paths, val_labels),\n    output_signature=(\n        tf.TensorSpec(shape=(96, 96, 3), dtype=tf.float32),\n        tf.TensorSpec(shape=(), dtype=tf.float32)\n    )\n).batch(BATCH_SIZE, drop_remainder=True).cache().prefetch(tf.data.AUTOTUNE)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-19T09:32:33.753218Z","iopub.execute_input":"2025-05-19T09:32:33.753480Z","iopub.status.idle":"2025-05-19T09:32:33.820724Z","shell.execute_reply.started":"2025-05-19T09:32:33.753460Z","shell.execute_reply":"2025-05-19T09:32:33.820199Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras import layers, models, callbacks\n\nnew_optimizer = tf.keras.optimizers.Adam(learning_rate=1e-4)\n\nvgg_model_v3.compile(optimizer=new_optimizer, loss='binary_crossentropy', metrics=['accuracy'])\n\ncallbacks_list = [\n    callbacks.EarlyStopping(\n        monitor='val_loss',\n        patience=3,\n        restore_best_weights=True\n    ),\n    callbacks.ReduceLROnPlateau(\n        monitor='val_loss',\n        factor=0.2,\n        patience=3,\n        min_lr=1e-6\n    ),\n    # save best model\n    callbacks.ModelCheckpoint(\n        filepath='best_vgg_model_v3.keras',\n        monitor='val_accuracy',\n        save_best_only=True,\n        verbose=1\n    )\n]\n\nvgg_history_v4 = vgg_model_v3.fit(\n    train_data, \n    validation_data = val_data, \n    epochs=10,\n    steps_per_epoch=steps_per_epoch,\n    callbacks=callbacks_list\n)\nvgg_model_v3.evaluate(val_data)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-19T09:32:59.283476Z","iopub.execute_input":"2025-05-19T09:32:59.283747Z","iopub.status.idle":"2025-05-19T09:42:17.993461Z","shell.execute_reply.started":"2025-05-19T09:32:59.283726Z","shell.execute_reply":"2025-05-19T09:42:17.992837Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# Accuracy plot\nplt.figure(figsize=(12, 5))\n\nplt.subplot(1, 2, 1)\nplt.plot(vgg_history_v4.history['accuracy'], label='Train Accuracy')\nplt.plot(vgg_history_v4.history['val_accuracy'], label='Val Accuracy')\nplt.title('Model Accuracy')\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.legend()\n\n# Loss plot\nplt.subplot(1, 2, 2)\nplt.plot(vgg_history_v4.history['loss'], label='Train Loss')\nplt.plot(vgg_history_v4.history['val_loss'], label='Val Loss')\nplt.title('Model Loss')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-19T09:42:21.918656Z","iopub.execute_input":"2025-05-19T09:42:21.919207Z","iopub.status.idle":"2025-05-19T09:42:22.235017Z","shell.execute_reply.started":"2025-05-19T09:42:21.919183Z","shell.execute_reply":"2025-05-19T09:42:22.234040Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Here, we can see again that the val accuracy goes up as the train accuracy goes up, and the same is true for the loss. Both lines going in the same direction show that our final tuning worked neatly. We see many spikes, which show that it is pretty noisy and bounces around.","metadata":{}},{"cell_type":"code","source":"y_true_vgg, y_pred_vgg = get_y_true_and_y_pred(val_data, vgg_model, 0.5)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-19T09:42:30.529660Z","iopub.execute_input":"2025-05-19T09:42:30.529955Z","iopub.status.idle":"2025-05-19T09:42:44.377250Z","shell.execute_reply.started":"2025-05-19T09:42:30.529934Z","shell.execute_reply":"2025-05-19T09:42:44.376460Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix, ConfusionMatrixDisplay\n\n# Confusion Matrix\ncm = confusion_matrix(y_true_vgg, y_pred_vgg)\ndisp = ConfusionMatrixDisplay(confusion_matrix=cm)\ndisp.plot(cmap='Blues')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-19T09:42:44.378492Z","iopub.execute_input":"2025-05-19T09:42:44.378715Z","iopub.status.idle":"2025-05-19T09:42:44.533621Z","shell.execute_reply.started":"2025-05-19T09:42:44.378698Z","shell.execute_reply":"2025-05-19T09:42:44.532937Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We have reduced the false negatives, but there are still some, which means this would not necessarily be good for medical purposes.","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import roc_curve, auc\n\nfpr, tpr, _ = roc_curve(y_true_vgg, y_pred_probs_vgg)\nroc_auc = auc(fpr, tpr)\n\nplt.figure()\nplt.plot(fpr, tpr, label=f'ROC curve (AUC = {roc_auc:.2f})')\nplt.plot([0, 1], [0, 1], linestyle='--', color='gray')\nplt.xlabel('False Positive Rate')\nplt.ylabel('True Positive Rate (Recall)')\nplt.title('Receiver Operating Characteristic')\nplt.legend(loc='lower right')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-19T09:42:44.534321Z","iopub.execute_input":"2025-05-19T09:42:44.534569Z","iopub.status.idle":"2025-05-19T09:42:44.681121Z","shell.execute_reply.started":"2025-05-19T09:42:44.534540Z","shell.execute_reply":"2025-05-19T09:42:44.680358Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The Area under the curve is 0.97, which is really great! So, our model has improved by tuning from AUC 0.92 to AUC 0.97.\nNow, I want to look at the false negative rate for the final results, which is 1-Recall (true positive rate).","metadata":{}},{"cell_type":"markdown","source":"The code below (commented out) is used to submit to the Kaggle competition. I scored 0.7809.","metadata":{}},{"cell_type":"code","source":"#import os\n#import pandas as pd\n\n# Get all .tif filenames from the test folder\n#test_dir = \"/kaggle/input/histopathologic-cancer-detection/test/\"\n#test_filenames = sorted([f for f in os.listdir(test_dir) if f.endswith(\".tif\")])\n\n# Extract IDs by removing \".tif\"\n#test_ids = [f[:-4] for f in test_filenames]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-19T07:32:01.125022Z","iopub.execute_input":"2025-05-19T07:32:01.125284Z","iopub.status.idle":"2025-05-19T07:32:02.046175Z","shell.execute_reply.started":"2025-05-19T07:32:01.125264Z","shell.execute_reply":"2025-05-19T07:32:02.045436Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#import tensorflow as tf\n#import cv2\n#import numpy as np\n\n#def test_data_generator(paths):\n#    for path in paths:\n#        img = load_image_cv2(path)\n#        yield img\n\n# Create full paths\n#test_paths = [os.path.join(test_dir, fname) for fname in test_filenames]\n\n# Create tf.data.Dataset\n#test_dataset = tf.data.Dataset.from_generator(\n#    lambda: test_data_generator(test_paths),\n#    output_signature=tf.TensorSpec(shape=(96, 96, 3), dtype=tf.float32)\n#).batch(32).prefetch(tf.data.AUTOTUNE)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-19T07:32:03.418210Z","iopub.execute_input":"2025-05-19T07:32:03.418542Z","iopub.status.idle":"2025-05-19T07:32:03.515838Z","shell.execute_reply.started":"2025-05-19T07:32:03.418518Z","shell.execute_reply":"2025-05-19T07:32:03.515242Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#from tensorflow.keras.models import load_model\n\n#best_model = load_model('best_vgg_model_v3.keras')\n\n#pred_probs = best_model.predict(test_dataset, verbose=1)\n#pred_labels = (pred_probs > 0.5).astype(int).flatten()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-19T07:32:12.168575Z","iopub.execute_input":"2025-05-19T07:32:12.168845Z","execution_failed":"2025-05-19T07:34:26.619Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#submission_df = pd.DataFrame({\n#    'id': test_ids,\n#    'label': pred_labels\n#})\n\n#submission_df.to_csv(\"submission.csv\", index=False)","metadata":{"trusted":true,"execution":{"execution_failed":"2025-05-19T07:34:26.620Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Conclusion\n\nLet us load the models and calculate the accuracy and recall scores.","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import recall_score, accuracy_score\nfrom tensorflow.keras.models import load_model\n\nbasic_model = load_model('best_basic_model.keras')\ny_true_bm, y_pred_bm = get_y_true_and_y_pred(val_data, basic_model, 0.5)\n\nbm_rc = recall_score(y_true_bm, y_pred_bm)\nbm_ac = accuracy_score(y_true_bm, y_pred_bm)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T05:27:20.449774Z","iopub.execute_input":"2025-05-20T05:27:20.450519Z","iopub.status.idle":"2025-05-20T05:28:58.578190Z","shell.execute_reply.started":"2025-05-20T05:27:20.450492Z","shell.execute_reply":"2025-05-20T05:28:58.577629Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.metrics import recall_score, accuracy_score\nfrom tensorflow.keras.models import load_model\n\nvgg_model = load_model('best_vgg_model.keras')\ny_true_vgg, y_pred_vgg = get_y_true_and_y_pred(val_data, vgg_model, 0.5)\n\nvgg_rc = recall_score(y_true_vgg, y_pred_vgg)\nvgg_ac = accuracy_score(y_true_vgg, y_pred_vgg)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T05:29:16.861136Z","iopub.execute_input":"2025-05-20T05:29:16.861421Z","iopub.status.idle":"2025-05-20T05:29:32.106784Z","shell.execute_reply.started":"2025-05-20T05:29:16.861397Z","shell.execute_reply":"2025-05-20T05:29:32.106198Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.metrics import recall_score, accuracy_score\nfrom tensorflow.keras.models import load_model\n\nvgg_model_lr = load_model('best_vgg_model_v1.keras')\ny_true_vgg_lr, y_pred_vgg_lr = get_y_true_and_y_pred(val_data, vgg_model_lr, 0.5)\n\nvgg_lr_rc = recall_score(y_true_vgg_lr, y_pred_vgg_lr)\nvgg_lr_ac = accuracy_score(y_true_vgg_lr, y_pred_vgg_lr)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T05:32:25.669943Z","iopub.execute_input":"2025-05-20T05:32:25.670821Z","iopub.status.idle":"2025-05-20T05:32:39.653380Z","shell.execute_reply.started":"2025-05-20T05:32:25.670795Z","shell.execute_reply":"2025-05-20T05:32:39.652541Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.metrics import recall_score, accuracy_score\nfrom tensorflow.keras.models import load_model\n\nvgg_model_v2 = load_model('best_vgg_model_v2.keras')\ny_true_vgg_v2, y_pred_vgg_v2 = get_y_true_and_y_pred(val_data, vgg_model_v2, 0.5)\n\nvgg_v2_rc = recall_score(y_true_vgg_v2, y_pred_vgg_v2)\nvgg_v2_ac = accuracy_score(y_true_vgg_v2, y_pred_vgg_v2)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T05:32:39.654748Z","iopub.execute_input":"2025-05-20T05:32:39.655205Z","iopub.status.idle":"2025-05-20T05:32:53.858371Z","shell.execute_reply.started":"2025-05-20T05:32:39.655185Z","shell.execute_reply":"2025-05-20T05:32:53.857834Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.metrics import recall_score, accuracy_score\nfrom tensorflow.keras.models import load_model\n\nvgg_model_v3 = load_model('best_vgg_model_v3.keras')\ny_true_vgg_v3, y_pred_vgg_v3 = get_y_true_and_y_pred(val_data, vgg_model_v3, 0.5)\n\nvgg_v3_rc = recall_score(y_true_vgg_v3, y_pred_vgg_v3)\nvgg_v3_ac = accuracy_score(y_true_vgg_v3, y_pred_vgg_v3)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T05:32:55.963708Z","iopub.execute_input":"2025-05-20T05:32:55.964346Z","iopub.status.idle":"2025-05-20T05:33:09.963045Z","shell.execute_reply.started":"2025-05-20T05:32:55.964323Z","shell.execute_reply":"2025-05-20T05:33:09.962450Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We want to focus on the miss rate. As this is medical data, we don't want to (false) negatively identify patients.","metadata":{}},{"cell_type":"code","source":"results = {\n    'Model': ['Basic Model', 'VGG Model', 'VGG Model with 1e-4 LR', 'VGG Model v2', 'VGG Model v3'],\n    'Accuracy': [bm_ac, vgg_ac, vgg_lr_ac, vgg_v2_ac, vgg_v3_ac],\n    'Recall': [bm_rc, vgg_rc, vgg_lr_rc, vgg_v2_rc, vgg_v3_rc],\n    'Miss Rate': [1 - bm_rc, 1 - vgg_rc, 1 - vgg_lr_rc, 1 - vgg_v2_rc, 1 - vgg_v3_rc]\n}\n\ndf_result = pd.DataFrame(results).sort_values(by=['Miss Rate'], ascending=True)\ndf_result","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-20T05:44:59.878407Z","iopub.execute_input":"2025-05-20T05:44:59.879130Z","iopub.status.idle":"2025-05-20T05:44:59.889787Z","shell.execute_reply.started":"2025-05-20T05:44:59.879103Z","shell.execute_reply":"2025-05-20T05:44:59.889041Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"From the table of results above, we can clearly see that the VGGNet Model V3 (smaller learning rate, additional dropout layer, and L2 regularization) has really improved the model. The accuracy improved from 0.796 to 0.923. The recall (true positive rate) also improved from 0.772 to 0.896. Finally, the miss rate (or false negative rate) was 0.104. This means we will miss 1 out of every 10. Now, that does not sound great from a medical perspective.\n\n## Learning and Improvements\n\nWe saw that the model performed better each time we tuned it. This shows that there was a lot of overfitting on the training data, and we needed the additional regularization and dropout. But there is still more we can improve. We did not try different optimizers. We also needed to reduce the data size to make it runnable in Kaggle. \n\n* Increase dataset sample size.\n* Try different optimizers.\n* Add additional VGGNet layers.\n* Try different model types. We noticed early stopping. We may need to use Google's InceptionNet.\n* Longer epochs\n\n## Conclusion\n\nWhile this model received a reasonable score of 0.7809 from the actual Kaggle test data, from a medical perspective, I would not trust it. We can still make many improvements as we tune the models more going forward. Overall, it is a good start, but we still have a lot of work to do.","metadata":{}},{"cell_type":"markdown","source":"# References\n\nAitken, A. M. (2025). Titanic - ML from Disaster (Supervised Learning), 8. Retrieved 05/13/2025 from https://www.kaggle.com/code/alexandermaitken/titanic-ml-from-disaster-supervised-learning#Titanic---ML-from-Disaster\n\nCukierski, W. (2018). Histopathologic Cancer Detection, 1. Retrieved 04/10/2025 from https://www.kaggle.com/competitions/histopathologic-cancer-detection.","metadata":{}},{"cell_type":"markdown","source":"You can find a copy of the notebook in my Github: https://github.com/alexmaitken/csca5642-week3","metadata":{}}]}