{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":11848,"databundleVersionId":862157,"sourceType":"competition"}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# CSCA-5642: Kaggle Cancerous Cell Detection Project #\n#### Develop an algorithm that identifies metastatic cancer in small image patches from digital pathology scans. ####\n    \n* Author: Alexander Meau  \n* Email: alme9155@colorado.edu  \n* GitHub: [https://github.com/alme9155/csca-5642-week3/tree/main](https://github.com/alme9155/csca-5642-week3/tree/main)  \n","metadata":{}},{"cell_type":"markdown","source":"## I. Brief description of the problem and data ##\n\nThis project aims to tackle Kaggle's Cancer Detection challenge to identify cancerous cell images from non-cancerous using Convolutional Neural Network (CNN). The challenge is to create an algorithm to identify metastatic cancer in small image patches taken from larget digital pathlogy scans. The data of this Kaggle competition is a slightly modified version of the PatchCamelyon (PCam) benchmark dataset. \n### Dataset: ####\n* The training dataset contains about 220,025 image patches already labeled 1 as cancerous, 0 as non-cancerous in \"train_labels.csv\".\n* A positive label indicates that the center 32x32px region of a patch contains at least one pixel of tumor tissue.\n* Tumor tissue in the outer region of the patch does not influence the label. \n\n### Data Size and Dimension ####\n* Training dataset: 220,025 tiff images\n* Test dataset: 57,458 tiff images (~26% of training size)\n* Each image patch is 96x96 pixel of RGB color images in tiff formats.\n* Label CSV contains 2 columns: id, label (1 as cancerous, 0 as non-cancerous)\n\n### Competition Rules ###\n* Expected submission CSV files in the same format as \"train_labels.csv\" with two columns \"id, label\". (id: unique id from test set, label: 0, or 1)","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\ntif_files = []\n\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    print(f\"\\nos path: [{dirname}]\")\n    tif_file_count =0\n    for filename in filenames:\n        if filename.endswith('.tif'):\n            tif_file_count +=1\n            if tif_file_count < 3:\n                print(os.path.join(dirname, filename))\n            elif tif_file_count == 3:\n                print(\"...\")\n        else:\n            print(os.path.join(dirname, filename))  \n    if tif_file_count > 0:\n        print(f\"\\nTotal number of .tif files in {dirname}: {tif_file_count}\")\n        tif_file_count =0\n\n        \n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-06-30T01:26:25.654875Z","iopub.execute_input":"2025-06-30T01:26:25.655165Z","iopub.status.idle":"2025-06-30T01:32:24.337066Z","shell.execute_reply.started":"2025-06-30T01:26:25.655144Z","shell.execute_reply":"2025-06-30T01:32:24.336222Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## II. Exploratory Data Analysis (EDA) ##\nFile: \"train_label.csv\"\n- Inspect file columns, duplciates and null labels.\n- Pre-process data frame to add 'train_filepath' for ease access to locate the files\n- Pre-process data to convert 'label' as string\n- Drop any rows in the dataset where corresponding .tif file path is missing.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport os \n\ntrain_labels = pd.read_csv('/kaggle/input/histopathologic-cancer-detection/train_labels.csv')\nprint(\"shape:\",train_labels.shape)\nprint(\"-- Label head ---\")\nprint(\"################\")\nprint(train_labels.head())\nprint(\"-----------------\")\nprint(\"Missing values: \")\nprint(\"################\")\nprint(train_labels.isnull().sum())\nprint(\"-----------------\")\nprint(f\"Sum of duplicated labels: {train_labels.duplicated().sum()}.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T01:32:37.026299Z","iopub.execute_input":"2025-06-30T01:32:37.027031Z","iopub.status.idle":"2025-06-30T01:32:37.558366Z","shell.execute_reply.started":"2025-06-30T01:32:37.026992Z","shell.execute_reply":"2025-06-30T01:32:37.557458Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\n\nsns.countplot(x='label', data=train_labels)\nplt.xticks([0, 1], ['Non-Cancer (0)', 'Cancer (1)'])\nplt.title(\"Distribution of Cancer vs Non-Cancer Images\")\nplt.ylabel(\"Image Count\")\nplt.xlabel(\"Label\")\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T06:38:14.226161Z","iopub.execute_input":"2025-06-30T06:38:14.226881Z","iopub.status.idle":"2025-06-30T06:38:14.358122Z","shell.execute_reply.started":"2025-06-30T06:38:14.226858Z","shell.execute_reply":"2025-06-30T06:38:14.357371Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"### show sample image files\nimport cv2\nfrom PIL import Image\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\n\n# add 'train_filepath' to dataframe\ntrain_dir = Path(\"/kaggle/input/histopathologic-cancer-detection/train\")\ntrain_labels = pd.read_csv('/kaggle/input/histopathologic-cancer-detection/train_labels.csv')\ntrain_labels['train_filepath'] = train_labels['id'].apply(lambda x: str(train_dir / f'{x}.tif'))\n# # convert 'label' field to string.\n# train_labels['label'] = train_labels['label'].astype(str)\n\n# show sample file\ndef show_samples(label, df=train_labels, num_images=5):\n    subset = df[df['label'] == label].sample(num_images, random_state=42)\n    plt.figure(figsize=(12, 4))\n    for i, row in enumerate(subset.itertuples()):\n        img = cv2.imread(row.train_filepath)\n        img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n        center = (32, 32, 64, 64)  \n        cv2.rectangle(img, (center[0], center[1]), (center[2], center[3]), (255, 0, 0), 1)\n        plt.subplot(1, num_images, i+1)\n        plt.imshow(img)\n        plt.title(f'Label: {row.label}')\n        plt.axis('off')\n    plt.show()\n\n\nshow_samples(0)  # Non-cancerous\nshow_samples(1)  # Cancerous\n\n# check image size and channel\nimg_path = f\"/kaggle/input/histopathologic-cancer-detection/train/{train_labels['id'].iloc[0]}.tif\"\nimg = Image.open(img_path)\nprint(f\"image size: {img.size}\") \nprint(f\"image mode: {img.mode}\") ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T01:32:46.355813Z","iopub.execute_input":"2025-06-30T01:32:46.356431Z","iopub.status.idle":"2025-06-30T01:32:48.042370Z","shell.execute_reply.started":"2025-06-30T01:32:46.356400Z","shell.execute_reply":"2025-06-30T01:32:48.041808Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings('ignore')\n\nimport tensorflow as tf\nimport tensorflow_io as tfio\nprint(\"TensorFlow version:\", tf.__version__)\nprint(\"GPU available:\", tf.config.list_physical_devices('GPU'))\n\nimport os\nos.environ['TF_CPP_MIN_LOG_LEVEL'] = '3' ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T01:36:53.762841Z","iopub.execute_input":"2025-06-30T01:36:53.763430Z","iopub.status.idle":"2025-06-30T01:36:53.767924Z","shell.execute_reply.started":"2025-06-30T01:36:53.763410Z","shell.execute_reply":"2025-06-30T01:36:53.767218Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Pre-process the data ###\n\n- Preprocess and split the data 80/20 into training and validation dataset\n- Prepare load_image method to be used\n- Prepare tensor pipeline to load image, label pair in mini-batch of 64","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nimport cv2\n\n\n# Split the DataFrame into training and validation sets\ntrain_df, val_df = train_test_split(\n    train_labels,\n    test_size=0.2,\n    stratify=train_labels['label'],\n    random_state=42\n)\n\ndef load_image(path, label):\n    def _load_image_py(path_tensor):\n        path_str = path_tensor.numpy()\n        if isinstance(path_str, bytes):\n            path_str = path_str.decode(\"utf-8\")\n        img = cv2.imread(path_str)\n        if img is None:\n            raise ValueError(f\"Failed to load image at path: {path_str}\")\n        img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n        return img.astype(np.uint8)\n\n    image = tf.py_function(func=_load_image_py, inp=[path], Tout=tf.uint8)\n    image.set_shape([96, 96, 3])\n    image = tf.cast(image, tf.float32) / 255.0\n    label = tf.cast(label, tf.int32)  \n    return image, label\n\n# Create training and validation datasets\ntrain_dataset = tf.data.Dataset.from_tensor_slices((train_df['train_filepath'], train_df['label']))\ntrain_dataset = train_dataset.map(load_image, num_parallel_calls=tf.data.AUTOTUNE)\ntrain_dataset = train_dataset.shuffle(1000).batch(64).prefetch(tf.data.AUTOTUNE)\n\nvalidation_dataset = tf.data.Dataset.from_tensor_slices((val_df['train_filepath'], val_df['label']))\nvalidation_dataset = validation_dataset.map(load_image, num_parallel_calls=tf.data.AUTOTUNE)\nvalidation_dataset = validation_dataset.batch(64).prefetch(tf.data.AUTOTUNE)\n\nprint(\"Data Pre-processing complete..\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T01:37:01.393851Z","iopub.execute_input":"2025-06-30T01:37:01.394446Z","iopub.status.idle":"2025-06-30T01:37:01.603684Z","shell.execute_reply.started":"2025-06-30T01:37:01.394423Z","shell.execute_reply":"2025-06-30T01:37:01.603022Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## III. Model Architecture ##\n\n### Model Description ###\n\n- This model implements traditional CNN model with three layers architecture.\n- Model begins with 32 filters, and then this hyperparameter value doubled of each subsequent layer: 32, 64, 128 \n- Since each image size is 96 x 96 pixel of RGB channel, input shape = (96,96, 3)\n- To avoid overfitting, each convolution block is matched with pooling layer and batch-normalization.\n- First layer use 5x5 kernel to learn major feature and then 3x3 kernel following inspiration of VGG-16. \n\n#### Reference ###\n- Ref: https://www.geeksforgeeks.org/computer-vision/vgg-16-cnn-model\n- Ref: https://openaccess.thecvf.com/content_cvpr_2016/papers/He_Deep_Residual_Learning_CVPR_2016_paper.pdf\n\n<img src=\"https://developers.google.com/static/machine-learning/practica/image-classification/images/cnn_architecture.svg\" width=\"600\">","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Conv2D, BatchNormalization, Activation, MaxPooling2D, Flatten, Dense, Dropout\nfrom tensorflow.keras.metrics import AUC\n\n#input shape = 96x 96x 3 (width x height x RGB)\ncnn_model = Sequential([\n    # First CNN block \n    Conv2D(32, kernel_size=5, strides=1, padding='same', input_shape=(96, 96, 3)),\n    BatchNormalization(),\n    Activation('relu'),\n    MaxPooling2D(pool_size=2, strides=2), \n    \n    # Second CNN block\n    Conv2D(64, kernel_size=3, strides=1, padding='same'),\n    BatchNormalization(),\n    Activation('relu'),\n    MaxPooling2D(pool_size=2, strides=2), \n    \n    # Third CNN block\n    Conv2D(128, kernel_size=3, strides=1, padding='same'),\n    BatchNormalization(),\n    Activation('relu'),\n    MaxPooling2D(pool_size=2, strides=2), \n    \n    # Output layer\n    Flatten(),\n    Dropout(0.5),\n    Dense(256, activation='relu'),\n    Dropout(0.3),\n    Dense(1, activation='sigmoid') \n])\n\ncnn_model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy', AUC(name='auroc')])\ncnn_model.summary()\nthree_layer_history = cnn_model.fit(\n    train_dataset,\n    validation_data=validation_dataset,\n    epochs=10\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T04:18:57.158583Z","iopub.execute_input":"2025-06-30T04:18:57.159173Z","iopub.status.idle":"2025-06-30T04:59:49.562017Z","shell.execute_reply.started":"2025-06-30T04:18:57.159151Z","shell.execute_reply":"2025-06-30T04:59:49.561425Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nimport numpy as np\nfrom sklearn.metrics import confusion_matrix\n\nprint(\"\\nTraining Metrics:\")\nprint(\"---------------------------\")\nfor metric in ['loss', 'accuracy', 'auroc']:\n    print(f\"{metric.capitalize()}: {three_layer_history.history[metric][-1]:.4f}\")\n\nprint(\"\\nValidation Metrics:\")\nprint(\"---------------------------\")\nfor metric in ['val_loss', 'val_accuracy', 'val_auroc']:\n    print(f\"{metric.replace('val_', 'Validation ').capitalize()}: {three_layer_history.history[metric][-1]:.4f}\")\n\n\n\n# show confusion matrix\ny_true = []\ny_pred = []\nfor images, labels in validation_dataset:\n    preds = cnn_model.predict(images, verbose=0)\n    y_true.extend(labels.numpy())\n    y_pred.extend((preds > 0.5).astype(int).flatten())\ny_true = np.array(y_true)\ny_pred = np.array(y_pred)\n\ncm = confusion_matrix(y_true, y_pred)\nplt.figure(figsize=(6, 5))\nsns.heatmap(cm, annot=True, fmt='d', cmap='Blues', \n            xticklabels=['Negative', 'Positive'], \n            yticklabels=['Negative', 'Positive'])\n\nplt.title('Confusion Matrix (Validation DataSet)')\nplt.xlabel('Predicted')\nplt.ylabel('True')\nplt.show()\n\nprint(\"\\nConfusion Matrix Metrics:\")\nprint(f\"True Negatives (TN): {cm[0,0]}\")\nprint(f\"False Positives (FP): {cm[0,1]}\")\nprint(f\"False Negatives (FN): {cm[1,0]}\")\nprint(f\"True Positives (TP): {cm[1,1]}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T05:00:00.932434Z","iopub.execute_input":"2025-06-30T05:00:00.933278Z","iopub.status.idle":"2025-06-30T05:01:33.210904Z","shell.execute_reply.started":"2025-06-30T05:00:00.933246Z","shell.execute_reply":"2025-06-30T05:01:33.210166Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nimport numpy as np\nfrom sklearn.metrics import confusion_matrix\n\n\n# Training loss vs Validation Loss\nplt.figure(figsize=(15, 5))\nplt.subplot(1, 3, 1)\nplt.plot(three_layer_history.history['loss'], label='Training Loss')\nplt.plot(three_layer_history.history['val_loss'], label='Validation Loss')\nplt.title('Training vs Validation Loss')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\nplt.grid(True)\n\n# Training Accuracy vs Validation Accuracy\nplt.subplot(1, 3, 2)\nplt.plot(three_layer_history.history['accuracy'], label='Training Accuracy')\nplt.plot(three_layer_history.history['val_accuracy'], label='Validation Accuracy')\nplt.title('Training vs Validation Accuracy')\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.legend()\nplt.grid(True)\n\n# AUC \nplt.subplot(1, 3, 3)\nplt.plot(three_layer_history.history['auroc'], label='Training AUC')\nplt.plot(three_layer_history.history['val_auroc'], label='Validation AUC')\nplt.title('Training vs Validation AUC')\nplt.xlabel('Epoch')\nplt.ylabel('AUC')\nplt.legend()\nplt.grid(True)\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T05:01:40.031161Z","iopub.execute_input":"2025-06-30T05:01:40.031476Z","iopub.status.idle":"2025-06-30T05:01:40.577324Z","shell.execute_reply.started":"2025-06-30T05:01:40.031447Z","shell.execute_reply":"2025-06-30T05:01:40.576555Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## IV. Result Analysis Before Fine-Tuning ##\n\n- Overall, this CNN model performs moderately for the histopathologic cancer detection task before tuning, as it has a simple architecture and achieves 87% accuracy on the validation dataset. \n\n### Details ###\n- At the end, this simplistic CNN model achieves 94% accuracy in the training dataset, but drops accuracy to 87% in the validation set. The drop in accuracy suggested that the model might have overfit the training dataset.\n- The curve of training loss starts at 0.4 and decreases steadily to 0.1 around epoch 8. However, validation loss spikes sharply, and this suggest the model is memorizing the training data rather than generalizing it.\n- The accuracy curve follows similiar patterns as in the curve of the training loss. Even though there are spikes with the validation curve, the validation accuracy improved gradually.\n- Spikes are also observed in the AUC curve, but it did achieve the highest value of 0.97, suggesting the model is highly effective despite fluctuations.\n","metadata":{}},{"cell_type":"markdown","source":"## V. Explore Different CNN Model Architectures ##\n\nMy approach to explore different Convolutional Neural Network (CNN) architectures is as follows:\n\n- Compare CNN model performance of same layout with higher and lower convolution layers\n  - **2 convolutional layers**\n  - **3 convolutional layers**\n  - **5 convolutional layers**\n","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Conv2D, BatchNormalization, Activation, MaxPooling2D, Flatten, Dense, Dropout\nfrom tensorflow.keras.metrics import AUC\n\n#Two convolutional block architecture\ntwo_layer_cnn_model = Sequential([\n    Conv2D(32, kernel_size=5, strides=1, padding='same', input_shape=(96, 96, 3)),\n    BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n\n    Conv2D(64, kernel_size=3, strides=1, padding='same'),\n    BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n\n    # Output layer\n    Flatten(),Dropout(0.5),\n    Dense(256, activation='relu'),Dropout(0.3), Dense(1, activation='sigmoid') \n])\n\ntwo_layer_cnn_model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy', AUC(name='auroc')])\ntwo_layer_history = two_layer_cnn_model.fit(\n    train_dataset,\n    validation_data=validation_dataset,\n    epochs=10\n)\n\nprint(\"\\nTraining Metrics:\")\nprint(\"---------------------------\")\nfor metric in ['loss', 'accuracy', 'auroc']:\n    print(f\"{metric.capitalize()}: {two_layer_history.history[metric][-1]:.4f}\")\n\nprint(\"\\nValidation Metrics:\")\nprint(\"---------------------------\")\nfor metric in ['val_loss', 'val_accuracy', 'val_auroc']:\n    print(f\"{metric.replace('val_', 'Validation ').capitalize()}: {two_layer_history.history[metric][-1]:.4f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T03:31:33.298358Z","iopub.execute_input":"2025-06-30T03:31:33.298935Z","iopub.status.idle":"2025-06-30T04:11:18.149816Z","shell.execute_reply.started":"2025-06-30T03:31:33.298908Z","shell.execute_reply":"2025-06-30T04:11:18.149107Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings('ignore')\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Conv2D, BatchNormalization, Activation, MaxPooling2D, Flatten, Dense, Dropout\nfrom tensorflow.keras.metrics import AUC\n\n#input shape = 96x 96x 3 (width x height x RGB)\nfive_layer_cnn_model = Sequential([\n    Conv2D(32, kernel_size=5, strides=1, padding='same', input_shape=(96, 96, 3)),\n    BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n\n    Conv2D(64, kernel_size=3, strides=1, padding='same'),\n    BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n\n    Conv2D(128, kernel_size=3, strides=1, padding='same'),\n    BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n\n    Conv2D(256, kernel_size=3, strides=1, padding='same'),\n    BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n\n    Conv2D(512, kernel_size=3, strides=1, padding='same'),\n    BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n        \n    # Output layer\n    Flatten(),Dropout(0.5),\n    Dense(256, activation='relu'),Dropout(0.3), Dense(1, activation='sigmoid') \n])\n\nfive_layer_cnn_model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy', AUC(name='auroc')])\nfive_layer_history = five_layer_cnn_model.fit(\n    train_dataset,\n    validation_data=validation_dataset,\n    epochs=10\n)\n\nprint(\"\\nTraining Metrics:\")\nprint(\"---------------------------\")\nfor metric in ['loss', 'accuracy', 'auroc']:\n    print(f\"{metric.capitalize()}: {five_layer_history.history[metric][-1]:.4f}\")\n\nprint(\"\\nValidation Metrics:\")\nprint(\"---------------------------\")\nfor metric in ['val_loss', 'val_accuracy', 'val_auroc']:\n    print(f\"{metric.replace('val_', 'Validation ').capitalize()}: {five_layer_history.history[metric][-1]:.4f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T01:38:04.899838Z","iopub.execute_input":"2025-06-30T01:38:04.900372Z","iopub.status.idle":"2025-06-30T02:24:35.318828Z","shell.execute_reply.started":"2025-06-30T01:38:04.900350Z","shell.execute_reply":"2025-06-30T02:24:35.318197Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\nhistories = {\n    \"2-layer CNN\": two_layer_history,\n    \"3-layer CNN\": three_layer_history,\n    \"5-layer CNN\": five_layer_history\n}\n\nmetrics = ['accuracy', 'val_accuracy']\ntitles = {\n    'accuracy': 'Training Accuracy',\n    'val_accuracy': 'Validation Accuracy'\n}\n\nfor metric in ['accuracy', 'val_accuracy']:    \n    plt.figure(figsize=(6, 4))\n    for label, history in histories.items():\n        plt.plot(history.history[metric], label=label)\n    plt.title(f\"{titles[metric]} Comparison\")\n    plt.xlabel(\"Epoch\")\n    plt.ylabel(metric.split('_')[-1].capitalize())\n    plt.legend()\n    plt.grid(True)\n    plt.tight_layout()\n    plt.show()\n\nrecords = []\nfor name, hist in histories.items():\n    max_train_accuracy = max(hist.history['accuracy'])\n    max_val_accuracy = max(hist.history['val_accuracy'])\n    last_train_accuracy = hist.history['accuracy'][-1]\n    last_val_accuracy = hist.history['val_accuracy'][-1]\n    train_loss = hist.history['loss'][-1]\n    val_loss = hist.history['val_loss'][-1]\n    max_auroc = max(hist.history['val_auroc'])\n    last_auroc = hist.history['val_auroc'][-1]    \n    records.append({\n        \"Model\": name,\n        \"Max Training Accuracy\": round(max_train_accuracy, 4),\n        \"Max Validation Accuracy\": round(max_val_accuracy, 4),\n        \"Last Train Acc\": round(last_train_accuracy, 4),\n        \"Last Val Acc\": round(last_val_accuracy, 4),\n        \"Train Loss\": round(train_loss, 4),\n        \"Val Loss\": round(val_loss, 4),\n        \"Max Val AUC\": round(max_auroc, 4),\n        \"Last Val AUC\": round(last_auroc, 4)         \n    })\nmetrics_df = pd.DataFrame(records)\nmetrics_df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T08:56:58.026870Z","iopub.execute_input":"2025-06-30T08:56:58.027171Z","iopub.status.idle":"2025-06-30T08:56:58.422301Z","shell.execute_reply.started":"2025-06-30T08:56:58.027150Z","shell.execute_reply":"2025-06-30T08:56:58.421543Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## VI. Performance Analysis with Different Model Architectures ##\n\n- Deeper model achieve higher traing and validation accuracy\n  - 5 layer CNN reach 97% training accuracy, compared to 94% (3 layers), and 84% (2 layers)\n  - 5 layer CNN reach 93% training accuracy, compared to 88% (3 layers), and 86% (2 layers)\n  - Deeper network have greater capacity to learn more complex pattern\n\n- Deeper model may introduce higher risk of overfitting\n  - 5 layer model have relatively deeper gap between training and validation accuracy.\n  - Performance gap widens with depth, suggesting a higher risk of overfitting and early stopping is required (next performance tuning).","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nimport numpy as np\n\n# Training loss vs Validation Loss\nplt.figure(figsize=(15, 5))\nplt.subplot(1, 3, 1)\nplt.plot(five_layer_history.history['loss'], label='Training Loss')\nplt.plot(five_layer_history.history['val_loss'], label='Validation Loss')\nplt.title('Training vs Validation Loss \\n(Best Model Architecture)')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\nplt.grid(True)\n\n# Training Accuracy vs Validation Accuracy\nplt.subplot(1, 3, 2)\nplt.plot(five_layer_history.history['accuracy'], label='Training Accuracy')\nplt.plot(five_layer_history.history['val_accuracy'], label='Validation Accuracy')\nplt.title('Training vs Validation Accuracy \\n(Best Model Architecture)')\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.legend()\nplt.grid(True)\n\n# AUC \nplt.subplot(1, 3, 3)\nplt.plot(five_layer_history.history['auroc'], label='Training AUC')\nplt.plot(five_layer_history.history['val_auroc'], label='Validation AUC')\nplt.title('Training vs Validation AUC \\n(Best Model Architecture)')\nplt.xlabel('Epoch')\nplt.ylabel('AUC')\nplt.legend()\nplt.grid(True)\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T07:18:20.876005Z","iopub.execute_input":"2025-06-30T07:18:20.876795Z","iopub.status.idle":"2025-06-30T07:18:21.398217Z","shell.execute_reply.started":"2025-06-30T07:18:20.876767Z","shell.execute_reply":"2025-06-30T07:18:21.397528Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## VII. Hyper-parameter Fine-Tuning by reducing Number of Filters ##\n\n* **Validation accuracy lower than training accuracy** indicates **overfitting**, where the model performs well on training data but fails to generalize to unseen data.\n* Overfitting may result from too many filters that learn noise patterns from the training data.\n* Less number of filters encouraging to focus on essential patterns rather than memorizing training data.\n\n- Three set number of filters be used for performance comparison:\n  - (16, 32, 64, 128, 256), \n  - (24, 48, 96, 192, 384)\n  - (32, 64, 128, 256, 512) (best model architecture)\n","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Conv2D, BatchNormalization, Activation, MaxPooling2D, Flatten, Dense, Dropout\nfrom tensorflow.keras.metrics import AUC\n\n\nfilter_config = [(16, 32, 64, 128, 256), (24, 48, 96, 192, 384)]\n# add previous model result\nhistory_dict = {} \nhistory_dict[(32, 64, 128, 256, 512)] = five_layer_history\nbest_validation_accuracy = five_layer_history.history['val_accuracy'][-1]\nbest_filter_config = (32, 64, 128, 256, 512)\nbest_filter_history = five_layer_history\n\nfor idx, (f1, f2, f3, f4, f5) in enumerate(filter_config):\n    filter_label = (f1, f2, f3, f4, f5)\n    print(f\"\\nTraining Model with Filters: {filter_label}...\")\n    # best model architecture (5 layers)\n    model_filter_tuning = Sequential([\n        Conv2D(f1, kernel_size=5, strides=1, padding='same', input_shape=(96, 96, 3)),\n        BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n    \n        Conv2D(f2, kernel_size=3, strides=1, padding='same'),\n        BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n    \n        Conv2D(f3, kernel_size=3, strides=1, padding='same'),\n        BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n    \n        Conv2D(f4, kernel_size=3, strides=1, padding='same'),\n        BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n    \n        Conv2D(f5, kernel_size=3, strides=1, padding='same'),\n        BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n            \n        # Output layer\n        Flatten(),Dropout(0.5),\n        Dense(256, activation='relu'),Dropout(0.3), Dense(1, activation='sigmoid') \n    ])\n\n        \n    model_filter_tuning.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy', AUC(name='auroc')])\n    history_filter_tuning = model_filter_tuning.fit(train_dataset, validation_data=validation_dataset, epochs=10, verbose=1)\n    history_dict[filter_label] = history_filter_tuning\n    final_validation_accuracy =  history_filter_tuning.history['val_accuracy'][-1]\n    print(f\"Final Validation Accuracy for {filter_label}: {final_validation_accuracy:.4f}\")\n\n    # Update best if current model is better\n    if final_validation_accuracy > best_validation_accuracy:\n        best_validation_accuracy = final_validation_accuracy\n        best_filter_config = filter_label\n        best_filter_history = history_filter_tuning\n\nprint(\"Best Filter Configuration:\")\nprint(f\"Filters: {best_filter_config}\")\nprint(f\"Validation Accuracy: {best_validation_accuracy:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T07:18:35.466921Z","iopub.execute_input":"2025-06-30T07:18:35.467648Z","iopub.status.idle":"2025-06-30T08:40:23.251990Z","shell.execute_reply.started":"2025-06-30T07:18:35.467619Z","shell.execute_reply":"2025-06-30T08:40:23.251245Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Table of performance comparison\nrecords_filter = []\nfor name, hist in history_dict.items():\n    max_train_accuracy = max(hist.history['accuracy'])\n    max_val_accuracy = max(hist.history['val_accuracy'])\n    last_train_accuracy = hist.history['accuracy'][-1]\n    last_val_accuracy = hist.history['val_accuracy'][-1]\n    train_loss = hist.history['loss'][-1]\n    val_loss = hist.history['val_loss'][-1]\n    max_auroc = max(hist.history['val_auroc'])\n    last_auroc = hist.history['val_auroc'][-1]\n    \n    records_filter.append({\n        \"name\": name,\n        \"Max Train Accuracy\": round(max_train_accuracy, 4),\n        \"Max Validation Accuracy\": round(max_val_accuracy, 4),\n        \"Last Train Acc\": round(last_train_accuracy, 4),\n        \"Last Validation Acc\": round(last_val_accuracy, 4),\n        \"Train Loss\": round(train_loss, 4),\n        \"Validation Loss\": round(val_loss, 4),\n        \"Max Train AUC\": round(max_auroc, 4),\n        \"Last Validation AUC\": round(last_auroc, 4)        \n    })\nmetrics_df_filter = pd.DataFrame(records_filter)\nmetrics_df_filter","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T13:11:12.602900Z","iopub.execute_input":"2025-06-30T13:11:12.603182Z","iopub.status.idle":"2025-06-30T13:11:12.618168Z","shell.execute_reply.started":"2025-06-30T13:11:12.603160Z","shell.execute_reply":"2025-06-30T13:11:12.617417Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nimport numpy as np\nfrom sklearn.metrics import confusion_matrix\n\n\n# Training loss vs Validation Loss\nplt.figure(figsize=(15, 5))\nplt.subplot(1, 3, 1)\nplt.plot(best_filter_history.history['loss'], label='Training Loss')\nplt.plot(best_filter_history.history['val_loss'], label='Validation Loss')\nplt.title(f\"Training vs Validation Loss  \\n(Best Num of Filters {best_filter_config}\")\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\nplt.grid(True)\n\n# Training Accuracy vs Validation Accuracy\nplt.subplot(1, 3, 2)\nplt.plot(best_filter_history.history['accuracy'], label='Training Accuracy')\nplt.plot(best_filter_history.history['val_accuracy'], label='Validation Accuracy')\nplt.title(f\"Training vs Validation Accuracy \\n(Best Num of Filters {best_filter_config}\")\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.legend()\nplt.grid(True)\n\n# AUC \nplt.subplot(1, 3, 3)\nplt.plot(best_filter_history.history['auroc'], label='Training AUC')\nplt.plot(best_filter_history.history['val_auroc'], label='Validation AUC')\nplt.title(f\"Training vs Validation AUC \\n(Best Num of Filters {best_filter_config}\")\nplt.xlabel('Epoch')\nplt.ylabel('AUC')\nplt.legend()\nplt.grid(True)\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T08:47:04.914833Z","iopub.execute_input":"2025-06-30T08:47:04.915353Z","iopub.status.idle":"2025-06-30T08:47:05.424298Z","shell.execute_reply.started":"2025-06-30T08:47:04.915330Z","shell.execute_reply":"2025-06-30T08:47:05.423635Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## VIII. Result Analysis after Fine-Tuning \"Number of Filters\" Hyper-parameter ##\n\n### Result ###\n- Despite fluctuation, CNN Model with filter configuration (32, 64, 128, 256, 512) achieve the higest validation accuracy 93% and validation AUC (0.97)\n- CNN Model with configuration (24, 48, 96, 192, 384) has high training performance (0.96), but validation accuracy peaked at 0.92 (< 0.97)\n- CNN Model with configuration (16, 32, 64, 128, 256) has worst validation result with AUC drop to 0.84\n\n### Key Insights to why or why not the hyper-parameter fine-Tuning worked ###\nBackground:\n- All three filter configurations perform better on training data than the validation dataset.\n- This suggest model is overfitted to the training data and fails to generalize to unseen data.\n- Theoretically, a lower number of filters should reduce noise patterns from the training set.\n\nWhat worked:\n- The spike in the validation history suggests that the model has learned enough useful features to achieve a high AUC score, but still requires further regularization.\n- Even though fewer filter configurations (24, 48, 96, 192, 384) might have improved regularization, they do not learn enough complex features to exceed the original (32, 64, 128, 256, 512) filter configuration.\n- Filter configuration (16, 32, 64, 128, 256) performs the worst because it lacks the capacity to capture complex patterns.\n- Thus, for the best model submission, **filter configurations (32, 64, 128, 256, 512) remains the best choice.**\n","metadata":{}},{"cell_type":"markdown","source":"## IX. Hyper-parameter Fine-Tuning by reducing learning rates and early stopping ##\n\n* A higher learning rate allows for faster convergence but risks overlooking the minimum and causing fluctuations.\n* A lower learning rate should provide more stable convergence, but it also risks getting stuck in local minima.\n\n- Three set learning rates will be used to study effect of this hyper-parameter: \n  - Learning rate = 0.0005 \n  - Learning rate = 0.001 *(default value)*\n  - Learning rate = 0.005\n\n- A slower learning rate requires more epochs to reach minima. Thus, the epoch is doubled to 20 compared to the rest of the settings (epoch=10)","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\n\n# default lr = 0.001\noptimizer = tf.keras.optimizers.Adam()\ndefault_lr = optimizer.learning_rate.numpy()\nprint(f\"Default learning rate: {default_lr}\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Conv2D, BatchNormalization, Activation, MaxPooling2D, Flatten, Dense, Dropout\nfrom tensorflow.keras.metrics import AUC\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.callbacks import EarlyStopping\nfrom tensorflow.keras.callbacks import ModelCheckpoint\n\n# increase number of epoch to 50 as learning rate decreases\nlearning_rates_to_epoch = {\n    5e-4: 20,   # (lower learning rate: 0.0005)\n    5e-3: 10    # (higher learning rate=0.005)\n}\n\nhistory_dict_lr = {} \nhistory_dict_lr[default_lr] = five_layer_history\nbest_lr_validation_accuracy = five_layer_history.history['val_accuracy'][-1]\nbest_lr = default_lr\nbest_lr_history = five_layer_history\n\nfor lr, num_epochs in learning_rates_to_epoch.items():\n    print(f\"\\nTraining Model with Learning Rate: {lr}...\")\n    \n    model_lr_tuning = Sequential([\n        Conv2D(f1, kernel_size=5, strides=1, padding='same', input_shape=(96, 96, 3)),\n        BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n    \n        Conv2D(f2, kernel_size=3, strides=1, padding='same'),\n        BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n    \n        Conv2D(f3, kernel_size=3, strides=1, padding='same'),\n        BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n    \n        Conv2D(f4, kernel_size=3, strides=1, padding='same'),\n        BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n    \n        Conv2D(f5, kernel_size=3, strides=1, padding='same'),\n        BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n            \n        # Output layer\n        Flatten(),Dropout(0.5),\n        Dense(256, activation='relu'),Dropout(0.3), Dense(1, activation='sigmoid') \n    ])\n\n\n    model_lr_tuning.compile(optimizer=Adam(learning_rate=lr), loss='binary_crossentropy', metrics=['accuracy', AUC(name='auroc')])\n\n    # Early stopping and keep best result\n    early_stopping = EarlyStopping(monitor='val_auroc', patience=3, mode='max', restore_best_weights=True)\n    history_lr_turning = model_lr_tuning.fit(train_dataset, validation_data=validation_dataset,epochs=num_epochs, callbacks=[early_stopping])\n    history_dict_lr[lr] = history_lr_turning\n    final_lr_validation_accuracy =  history_lr_turning.history['val_accuracy'][-1]\n    print(f\"Final Validation Accuracy for {lr}: {final_lr_validation_accuracy:.4f}\")\n    \n    if final_lr_validation_accuracy > best_lr_validation_accuracy:\n        best_lr_validation_accuracy = final_lr_validation_accuracy\n        best_lr = lr\n        best_lr_history = history_lr_turning\n\nprint(\"Best Learing Rate Configuration:\")\nprint(f\"Learning Rate: {best_lr}\")\nprint(f\"Validation Accuracy: {best_lr_validation_accuracy:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T09:09:51.282668Z","iopub.execute_input":"2025-06-30T09:09:51.282954Z","iopub.status.idle":"2025-06-30T09:58:00.008793Z","shell.execute_reply.started":"2025-06-30T09:09:51.282935Z","shell.execute_reply":"2025-06-30T09:58:00.008114Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Table of performance comparison\nrecords_lr = []\nfor name, hist in history_dict_lr.items():\n    max_train_accuracy = max(hist.history['accuracy'])\n    max_val_accuracy = max(hist.history['val_accuracy'])\n    last_train_accuracy = hist.history['accuracy'][-1]\n    last_val_accuracy = hist.history['val_accuracy'][-1]\n    train_loss = hist.history['loss'][-1]\n    val_loss = hist.history['val_loss'][-1]\n    max_auroc = max(hist.history['val_auroc'])\n    last_auroc = hist.history['val_auroc'][-1]    \n    records_lr.append({\n        \"Learning Rate\": name,\n        \"Max Training Accuracy\": round(max_train_accuracy, 4),\n        \"Max Validation Accuracy\": round(max_val_accuracy, 4),\n        \"Last Train Acc\": round(last_train_accuracy, 4),\n        \"Last Val Acc\": round(last_val_accuracy, 4),\n        \"Train Loss\": round(train_loss, 4),\n        \"Val Loss\": round(val_loss, 4),\n        \"Max Val AUC\": round(max_auroc, 4),\n        \"Last Val AUC\": round(last_auroc, 4)                \n    })\nmetrics_df_lr = pd.DataFrame(records_lr)\nmetrics_df_lr","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T13:12:26.046525Z","iopub.execute_input":"2025-06-30T13:12:26.047027Z","iopub.status.idle":"2025-06-30T13:12:26.060383Z","shell.execute_reply.started":"2025-06-30T13:12:26.047003Z","shell.execute_reply":"2025-06-30T13:12:26.059713Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## X. Result Analysis after Fine-Tuning \"Learning Rate\" Hyper-parameter with Early Stopping ##\n\n### Result ###\n- Learning rate =0.001 (default) achieve the highest accuracy (0.93) and AUC (0.97)\n- Training and validation loss and both the lowest among all 3 options indicating overfitting is well-controlled.\n\n### Key Insights to why or why not the hyper-parameter fine-Tuning worked ###\n- When the learning rate is 0.001, early stopping is most likely in effect before overfitting starts.\n- When the learning rate is reduced (rate = 0.0005), slower convergence might have caused the model to underfit before early stopping, halting the training.\n- When the learning rate is increased (rate=0.005), faster convergence may cause instability without reaching the optimum.","metadata":{}},{"cell_type":"markdown","source":"## XI. Prepare final CNN model for submission ##\n\n- Model Architecture: 5 convoluntion block + output block\n- Input shape: (96, 96, 3)\n- Kernel size=5 for the 1st convolution block, and kernel size=3 for the rest.\n- Stride=1, padding=same,\n- Each convolution block include batchnormation max pooling size=2, strides=2\n- Filter configuration: 32, 64, 128, 256, 512\n- Learning rate: default (0.001) with Early stopping ","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Conv2D, BatchNormalization, Activation, MaxPooling2D, Flatten, Dense, Dropout\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.metrics import AUC\nfrom tensorflow.keras.callbacks import ModelCheckpoint, EarlyStopping\n\n\n#putting everything together for kaggle submission\n# Model Architecture: 5 layers\n# Learning rate: default (0.001)\n# Best filter configuration: 32, 64, 128, 256, 512\nmodel_best = Sequential([\n    Conv2D(32, kernel_size=5, strides=1, padding='same', input_shape=(96, 96, 3)),\n    BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n\n    Conv2D(64, kernel_size=3, strides=1, padding='same'),\n    BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n\n    Conv2D(128, kernel_size=3, strides=1, padding='same'),\n    BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n\n    Conv2D(256, kernel_size=3, strides=1, padding='same'),\n    BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n\n    Conv2D(512, kernel_size=3, strides=1, padding='same'),\n    BatchNormalization(), Activation('relu'), MaxPooling2D(pool_size=2, strides=2),\n        \n    # Output layer\n    Flatten(),Dropout(0.5),\n    Dense(256, activation='relu'),Dropout(0.3), Dense(1, activation='sigmoid') \n])\n\n# default Learning rate =0.001\nmodel_best.compile(optimizer=Adam(learning_rate=0.001), loss='binary_crossentropy', metrics=['accuracy', AUC(name='auroc')])\ncheckpoint = ModelCheckpoint('model_best.h5', monitor='val_auroc', save_best_only=True, mode='max')\n\nearly_stopping = EarlyStopping(monitor='val_auroc', patience=3, mode='max', restore_best_weights=True)\nhistory_best_model = model_best.fit(train_dataset, validation_data=validation_dataset,epochs=num_epochs, callbacks=[early_stopping, checkpoint])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T10:14:52.438828Z","iopub.execute_input":"2025-06-30T10:14:52.439511Z","iopub.status.idle":"2025-06-30T10:59:11.861083Z","shell.execute_reply.started":"2025-06-30T10:14:52.439489Z","shell.execute_reply":"2025-06-30T10:59:11.860497Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from tensorflow.keras.models import load_model\nimport pandas as pd\nimport os\n#model_loaded = load_model(\"model_best.h5\", compile=False)\n\n# Load test dataset and generate predictions\n\ntest_dir = Path(\"/kaggle/input/histopathologic-cancer-detection/test\")\n\ntest_df = pd.read_csv('/kaggle/input/histopathologic-cancer-detection/sample_submission.csv')  \ntest_df['test_filepath'] = test_df['id'].apply(lambda x: str(test_dir / f'{x}.tif'))\n\ntest_dataset = tf.data.Dataset.from_tensor_slices((test_df['test_filepath'], test_df['label']))\ntest_dataset = test_dataset.map(load_image, num_parallel_calls=tf.data.AUTOTUNE)\ntest_dataset = test_dataset.batch(64).prefetch(tf.data.AUTOTUNE)\n\n\npredictions = model_best.predict(test_dataset, verbose=1)\ntest_df['label'] = predictions.flatten()\n\n# Save submission file\nsubmission_path = 'submission.csv'\ntest_df[['id', 'label']].to_csv(submission_path, index=False)\nprint(f\"Submission file saved to {submission_path}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T11:32:06.195843Z","iopub.execute_input":"2025-06-30T11:32:06.196568Z","iopub.status.idle":"2025-06-30T11:33:02.828232Z","shell.execute_reply.started":"2025-06-30T11:32:06.196545Z","shell.execute_reply":"2025-06-30T11:33:02.827615Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission_path = '/kaggle/working/submission.csv'\ntest_df[['id', 'label']].to_csv(submission_path, index=False)\nprint(f\"Submission file saved to {submission_path}\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## XII. Conclusion ##\n\n### Best performance is achieved with the following configurations ###\n- Model Architecture: 5 convoluntion block + output block\n- Input shape: (96, 96, 3)\n- Kernel size=5 for the 1st convolution block, and kernel size=3 for the rest.\n- Stride=1, padding=same,\n- Each convolution block include batchnormation max pooling size=2, strides=2\n- Filter configuration: 32, 64, 128, 256, 512\n- Learning rate: default (0.001) with Early stopping\n\n\nDiscuss and interpret results as well as learnings and takeaways. What did and did not help improve the performance of your models? What improvements could you try in the future?","metadata":{}},{"cell_type":"markdown","source":"### Results ###\n- High precision and specificity mean the model is reliable with high numbers of TP and TN\n- Moderate false negative rate (1,881 missed cases) \n\n\n| Metric               | Formula                          | Value     |\n|----------------------|----------------------------------|-----------|\n| Accuracy             | (TP + TN) / (TP + TN + FP + FN) | 0.9348    |\n| Precision (PPV)      | TP / (TP + FP)                  | 0.9489    |\n| Recall (Sensitivity) | TP / (TP + FN)                  | 0.8946    |\n| Specificity          | TN / (TN + FP)                  | 0.9672    |\n| F1 Score             | 2 * (Precision * Recall) / (Precision + Recall) | 0.9209    |\n\nConfusion Matrix:\n|                  | Predicted Negative | Predicted Positive | Total   |\n|------------------|--------------------|---------------------|---------|\n| **Actual Negative** | TN = 25,322 (58.9%) | FP = 860 (2.0%)       | 26,182 (60.9%) |\n| **Actual Positive** | FN = 1,881 (4.4%)   | TP = 15,942 (37.1%)    | 17,823 (39.1%) |\n| **Total**          | 27,203 (63.2%)     | 16,802 (36.8%)         | 43,005 (100%) |\n","metadata":{}},{"cell_type":"markdown","source":"### What improve the performance of the model ###\n- Deeper model achieve higher traing and validation accuracy\n  - 5 layer CNN reach 97% training accuracy, compared to 94% (3 layers), and 84% (2 layers)\n  - 5 layer CNN reach 93% training accuracy, compared to 88% (3 layers), and 86% (2 layers)\n  - Deeper network have greater capacity to learn more complex pattern\n\n- Deeper model may introduce higher risk of overfitting\n  - 5 layer model have relatively deeper gap between training and validation accuracy.\n  - Performance gap widens with depth, suggesting a higher risk of overfitting and early stopping is required.\n\n- Increase filter depth helped feature extraction:\n  - The spike in the validation history suggests that the model has learned enough useful features to achieve a high AUC score, but still requires further regularization.\n  - Even though fewer filter configurations (24, 48, 96, 192, 384) might have improved regularization, they do not learn enough complex features to exceed the original (32, 64, 128, 256, 512) filter configuration.  -     - Filter configuration (16, 32, 64, 128, 256) performs the worst because it lacks the capacity to capture complex patterns.\n  - Thus, for the best model submission, **filter configurations (32, 64, 128, 256, 512) remains the best choice.**\n\n- Early stopping protects against overfitting but requires appropriate learning rates for optimal effect.\n  - Learning rate =0.001 (default) achieve the highest accuracy (0.93) and AUC (0.97)\n  - When the learning rate is reduced (rate = 0.0005), slower convergence might have caused the model to underfit before early stopping, halting the training.\n  - When the learning rate is increased (rate=0.005), faster convergence may cause instability without reaching the optimum.\n\n### What improvement could you try in the future? ###\n- More complex convolution model such as VGG-16/ VGG-32 by Oxford University [https://www.robots.ox.ac.uk/~vgg/research/very_deep/](https://www.robots.ox.ac.uk/~vgg/research/very_deep/)\n- Explore vision transformer(ViT) developed by Google Deep Brain Team [https://deepmind.google/discover/blog/rt-2-new-model-translates-vision-and-language-into-action/](https://deepmind.google/discover/blog/rt-2-new-model-translates-vision-and-language-into-action/)","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nimport numpy as np\nimport os\nfrom sklearn.metrics import confusion_matrix\n\ny_true_best_model = []\nfor _, labels in validation_dataset:\n    y_true_best_model.extend(labels.numpy())\ny_true_best_model = np.array(y_true_best_model)\ny_pred_prob_best_model = model_best.predict(validation_dataset).flatten()\ny_pred_best_model = (y_pred_prob_best_model >= 0.5).astype(int)\n\ncm = confusion_matrix(y_true_best_model, y_pred_best_model)\nplt.figure(figsize=(6, 5))\nsns.heatmap(cm, annot=True, fmt='d', cmap='Blues', \n            xticklabels=['Negative', 'Positive'], \n            yticklabels=['Negative', 'Positive'])\n\nplt.title('Confusion Matrix (Validation DataSet)')\nplt.xlabel('Predicted')\nplt.ylabel('True')\nplt.show()\n\nprint(\"\\nConfusion Matrix Metrics:\")\nprint(f\"True Negatives (TN): {cm[0,0]}\")\nprint(f\"False Positives (FP): {cm[0,1]}\")\nprint(f\"False Negatives (FN): {cm[1,0]}\")\nprint(f\"True Positives (TP): {cm[1,1]}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T11:46:51.669771Z","iopub.execute_input":"2025-06-30T11:46:51.670492Z","iopub.status.idle":"2025-06-30T11:48:21.106218Z","shell.execute_reply.started":"2025-06-30T11:46:51.670467Z","shell.execute_reply":"2025-06-30T11:48:21.105406Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nimport numpy as np\nimport os\nfrom sklearn.metrics import confusion_matrix\n\n\n# Training loss vs Validation Loss\nplt.figure(figsize=(15, 5))\nplt.subplot(1, 3, 1)\nplt.plot(history_best_model.history['loss'], label='Training Loss')\nplt.plot(history_best_model.history['val_loss'], label='Validation Loss')\nplt.title(f\"Training vs Validation Loss (Best Model)\")\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\nplt.grid(True)\n\n# Training Accuracy vs Validation Accuracy\nplt.subplot(1, 3, 2)\nplt.plot(history_best_model.history['accuracy'], label='Training Accuracy')\nplt.plot(history_best_model.history['val_accuracy'], label='Validation Accuracy')\nplt.title(f\"Training vs Validation Accuracy (Best Model)\")\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.legend()\nplt.grid(True)\n\n# AUC \nplt.subplot(1, 3, 3)\nplt.plot(history_best_model.history['auroc'], label='Training AUC')\nplt.plot(history_best_model.history['val_auroc'], label='Validation AUC')\nplt.title(f\"Training vs Validation AUC (Best Model)\")\nplt.xlabel('Epoch')\nplt.ylabel('AUC')\nplt.legend()\nplt.grid(True)\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-30T11:42:20.706057Z","iopub.execute_input":"2025-06-30T11:42:20.706345Z","iopub.status.idle":"2025-06-30T11:42:21.276388Z","shell.execute_reply.started":"2025-06-30T11:42:20.706325Z","shell.execute_reply":"2025-06-30T11:42:21.275644Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## XIII. Future Works and References ##\n- Experiment with VGG-16, VGG-32 by Oxford University [https://www.robots.ox.ac.uk/~vgg/research/very_deep/](https://www.robots.ox.ac.uk/~vgg/research/very_deep/)\n- Explore with Visual Transformer by Google Deep Brain [https://deepmind.google/discover/blog/rt-2-new-model-translates-vision-and-language-into-action/](https://deepmind.google/discover/blog/rt-2-new-model-translates-vision-and-language-into-action/)\n\n### References ###\n- Ref: Tensor flow tutorial [https://www.tensorflow.org/tutorials](https://www.tensorflow.org/tutorials) \n- Ref: TensorFlow Guide: [https://www.tensorflow.org/guide/keras/](https://www.tensorflow.org/guide/keras/)","metadata":{}}]}