{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":11848,"databundleVersionId":862157,"sourceType":"competition"}],"dockerImageVersionId":30646,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<a id=\"0\"></a> <br>\n # Table of Contents  \n\n- Part 0:  [Project Introduction](#1)\n- Part 1:  [Exploratory Data Analysis (EDA)](#2)   \n  - 1.1  [Import Files, assisgn source directories, and Load Data](#3)\n  - 1.2  [Use partial of training data for EDA](#4)\n  - 1.3  [Abnormal images in all data ](#5)\n- Part 2:  [Model Building and training](#6)\n  - 2.1 [Copy all normal images to directories, with train test split](#7)\n  - 2.2 [Baseline model](#8)\n  - 2.3 [Model Tuning 1 - Add more epochs to increase training AUC](#9)\n  - 2.4 [Model Tuning 2 - Add Model Complexity Attempt 1(Use Transfer Learning)](#10)\n- Part 3: [Test Results](#11)\n- Part 3: [Project Conclusion](#12)\n- [References](#13)\n","metadata":{}},{"cell_type":"markdown","source":"# <a id=\"1\"></a>\n<div style=\"text-align: center; background-color: #AFDDCE; font-size:100%; padding: 5px;border-radius:10px 10px;\">\n    <h1>Part 0: Project Introduction </h1>\n</div>","metadata":{}},{"cell_type":"markdown","source":"### Dataset Overview\n- The dataset originates from a Kaggle project focused on classifying small pathology images. While the information on data collection methods is not clear, the [dataset description](https://www.kaggle.com/competitions/histopathologic-cancer-detection/data) provides valuable insights and clarification on duplicates. \n- The dataset comprises separate directories for train and test images. Ground truth labels are available in the \"train_labels.csv\" file, where a positive label signifies the presence of tumor tissue in the center 32x32px region.\n- Although the exact dataset size is undisclosed (we'll discover in EDA), it is mentioned that duplicates were removed. The dataset maintains the same splits as the PCam benchmark.\n\n### Project Goals and focus\n- Project goal: Identify metastatic tissue in histopathologic scans of lymph node sections in the unseen (testing) dataset.\n- Beyond Prediction: Gain insights and identify potential bottlenecks/challenges into the application of CNNs for processing and analyzing medical images in large volume. \n\n\n### Model Training Approach:\n- Initial Base Run: A preliminary model was trained to establish a baseline.\n- Tuning: Refinements were made, focusing on epoch numbers as the model was still learning.\n- Transfer Learning: Utilized InceptionV3 for improved performance.\n\n### A few highlights:\n- EDA (Exploratory Data Analysis): Due to the dataset's size, a subset (1%) was used for efficient EDA. Noteworthy patterns and abnormal images were identified and excluded from training.\n- Data Handling: Numpy arrays were employed for faster EDA, but image copies were later used for model training.\n- Model Training Steps: A systematic approach involved an initial base run, tuning parameters, and implementing transfer learning with InceptionV3.\n","metadata":{}},{"cell_type":"markdown","source":"\n","metadata":{}},{"cell_type":"markdown","source":"\n","metadata":{}},{"cell_type":"markdown","source":"# <a id=\"2\"></a>\n<div style=\"text-align: center; background-color: #AFDDCE; font-size:100%; padding: 5px;border-radius:10px 10px;\">\n    <h1>Part 1: Exploratory Data Analysis (EDA)</h1>\n</div>","metadata":{}},{"cell_type":"markdown","source":"# <a id=\"3\"></a>\n\n<div style=\"text-align: center; background-color: #81badd; font-size:70%; padding: 2px;border-radius:10px;\">\n    <h2>Part 1.1: Import Files, assisgn source directories, and Load Data</h2>\n</div>","metadata":{}},{"cell_type":"code","source":"# Import Libraries\n# General\nimport numpy as np\nimport pandas as pd\nimport os\nimport shutil\nfrom glob import glob \nimport cv2\nimport gc \nimport time\nimport matplotlib.pyplot as plt\nimport matplotlib.patches as patches\nimport seaborn as sns\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import precision_recall_curve, f1_score, roc_curve,auc,recall_score, precision_score, accuracy_score, confusion_matrix,ConfusionMatrixDisplay\n\n \n%matplotlib inline\n\n\n# Tensorflow\nimport tensorflow as tf\nfrom tensorflow.keras.optimizers import Adam, RMSprop\nfrom tensorflow.keras.optimizers.schedules import ExponentialDecay\nfrom tensorflow.keras.metrics import AUC\nfrom tensorflow.keras.applications import InceptionV3, DenseNet121, NASNetMobile\nfrom tensorflow.keras import layers\nfrom tensorflow.keras import Model\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import BatchNormalization, Conv2D, Flatten, Dense, Dropout, MaxPooling2D\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tqdm import tqdm_notebook,trange\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T09:46:52.340981Z","iopub.execute_input":"2025-05-27T09:46:52.341698Z","iopub.status.idle":"2025-05-27T09:47:04.619959Z","shell.execute_reply.started":"2025-05-27T09:46:52.341671Z","shell.execute_reply":"2025-05-27T09:47:04.61925Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Assign directory \nbase_dir = '../input/histopathologic-cancer-detection'  #path\nlabel_path = os.path.join(base_dir, 'train_labels.csv')\ntrain_path = os.path.join(base_dir,'train')\ntest_path = os.path.join(base_dir,'test')\nsample_path = os.path.join(base_dir,'sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:47:13.545234Z","iopub.execute_input":"2025-05-27T09:47:13.546305Z","iopub.status.idle":"2025-05-27T09:47:13.550565Z","shell.execute_reply.started":"2025-05-27T09:47:13.546275Z","shell.execute_reply":"2025-05-27T09:47:13.549517Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Check numbers of files under train and test\nprint('samples in training data: ',len(os.listdir(train_path)))\nprint('samples in testingdata: ',len(os.listdir(test_path)))","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:47:16.360813Z","iopub.execute_input":"2025-05-27T09:47:16.361739Z","iopub.status.idle":"2025-05-27T09:47:25.083804Z","shell.execute_reply.started":"2025-05-27T09:47:16.361707Z","shell.execute_reply":"2025-05-27T09:47:25.082935Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# create a data frame to match training data image id and labels\ndf = pd.DataFrame({'path': glob(os.path.join(train_path,'*.tif'))}) # load the filenames\ndf['id'] = df.path.map(lambda x: x.split('/')[4].split(\".\")[0]) # get image id to merge with training labels\ntrain_labels = pd.read_csv(label_path)\ndf = df.merge(train_labels, on = \"id\") # merge labels and filepaths\ndf.head(2) ","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:47:27.140755Z","iopub.execute_input":"2025-05-27T09:47:27.141113Z","iopub.status.idle":"2025-05-27T09:47:32.057515Z","shell.execute_reply.started":"2025-05-27T09:47:27.141083Z","shell.execute_reply":"2025-05-27T09:47:32.056638Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Count the number of samples per class\nlabel_counts = df['label'].value_counts()\n\n# Print the counts\nprint(\"Class distribution:\")\nprint(label_counts)\n\n# Optional: Show as percentages\nprint(\"\\nPercentage distribution:\")\nprint(label_counts / len(df) * 100)\n\n# Plot the distribution\nplt.figure(figsize=(6,4))\nsns.barplot(x=label_counts.index, y=label_counts.values)\nplt.xlabel('Class Labels')\nplt.ylabel('Number of Samples')\nplt.title('Class Distribution')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T09:47:33.285784Z","iopub.execute_input":"2025-05-27T09:47:33.286586Z","iopub.status.idle":"2025-05-27T09:47:33.445109Z","shell.execute_reply.started":"2025-05-27T09:47:33.28656Z","shell.execute_reply":"2025-05-27T09:47:33.444284Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# check if df is created correctly\n# check image counts\nprint('df shape',df.shape)\n# check null values\nprint('nulls :',df.isnull().sum().sum())\n# check dup values\nprint(\"dups: \",df.duplicated().sum())","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:47:36.169245Z","iopub.execute_input":"2025-05-27T09:47:36.169891Z","iopub.status.idle":"2025-05-27T09:47:36.323898Z","shell.execute_reply.started":"2025-05-27T09:47:36.169859Z","shell.execute_reply":"2025-05-27T09:47:36.323175Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot the label count histogram\nprint('Positive labels in traing data: ', train_labels['label'].value_counts()[1] /  (train_labels['label'].value_counts()[0] + train_labels['label'].value_counts()[1])  )\nplt.figure(figsize=(5, 4))\nax = train_labels['label'].value_counts().sort_index().plot(kind='bar', color=['blue', 'orange'])\nplt.xticks([0, 1], labels=[f\"Negative N={train_labels['label'].value_counts()[0]}\", f\"Positive N={train_labels['label'].value_counts()[1]}\"], rotation=0)\nplt.title(f'Label Distributions in {train_labels.shape[0]} training data')\n","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:47:36.979166Z","iopub.execute_input":"2025-05-27T09:47:36.979889Z","iopub.status.idle":"2025-05-27T09:47:37.157578Z","shell.execute_reply.started":"2025-05-27T09:47:36.97986Z","shell.execute_reply":"2025-05-27T09:47:37.156933Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <a id=\"4\"></a>\n\n<div style=\"text-align: center; background-color: #81badd; font-size:70%; padding: 2px;border-radius:10px;\">\n    <h2>Part 1.2: Use partial of training data for EDA</h2>\n</div>","metadata":{}},{"cell_type":"code","source":"# train_labels.head()\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:47:40.433212Z","iopub.execute_input":"2025-05-27T09:47:40.433898Z","iopub.status.idle":"2025-05-27T09:47:40.442194Z","shell.execute_reply.started":"2025-05-27T09:47:40.433868Z","shell.execute_reply":"2025-05-27T09:47:40.441344Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### Here for partial EDA data, I convert images to numpy array with unit8 types. It's faster.\n* Note creating numpy array with huge data takes too much memory, which would cause problems when reach the stage of training deeper CNN. Need to copy files to directory for training later.","metadata":{}},{"cell_type":"code","source":"# source https://www.kaggle.com/code/gomezp/complete-beginner-s-guide-eda-keras-lb-0-93. \n# This method is fast but consumes too much memory. casue later I cannot train deep CNN \n# function to load data.\n\ndef load_data(pct, df):\n    \"\"\"This function loads percent of df and return image path/name and labels.\"\"\"\n    # Get the number of images to load\n    n = int(pct * df.shape[0])\n    \n    # Randomly select n indices from the dataframe\n    random_indices =  np.random.choice(df.shape[0], size=n, replace=False)\n    \n    # Allocate a numpy array for the images (n, 96x96px, 3 channels, values 0 - 255)\n    X = np.zeros([n, 96, 96, 3], dtype=np.uint8)\n    # Convert the labels to a numpy array\n    y = np.squeeze(df['label'].values)[random_indices]\n\n    # Read images using randomly selected indices\n    for i, idx in enumerate(random_indices):\n        X[i] = cv2.imread(df.loc[idx, 'path'])\n\n    return n, X, y\n","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:47:42.292746Z","iopub.execute_input":"2025-05-27T09:47:42.293314Z","iopub.status.idle":"2025-05-27T09:47:42.299152Z","shell.execute_reply.started":"2025-05-27T09:47:42.293288Z","shell.execute_reply":"2025-05-27T09:47:42.298266Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### Slow to load data each time. Just use 1% training data for EDA seems enough.","metadata":{}},{"cell_type":"code","source":"# Load partical data\nPCT = 0.01\nn,sample_X,sample_y = load_data(pct = PCT, df=df)","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:48:40.215798Z","iopub.execute_input":"2025-05-27T09:48:40.216117Z","iopub.status.idle":"2025-05-27T09:49:07.458548Z","shell.execute_reply.started":"2025-05-27T09:48:40.21609Z","shell.execute_reply":"2025-05-27T09:49:07.457859Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Take a look at image size and pixel info\nprint('Image shape: ',sample_X.shape)\nprint('Image maximum pixel value: ', sample_X.max())\nprint('Image minimum pixel value: ', sample_X.min())","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:49:15.944625Z","iopub.execute_input":"2025-05-27T09:49:15.945236Z","iopub.status.idle":"2025-05-27T09:49:15.96128Z","shell.execute_reply.started":"2025-05-27T09:49:15.94521Z","shell.execute_reply":"2025-05-27T09:49:15.960532Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# source: https://www.kaggle.com/code/qitvision/a-complete-ml-pipeline-fast-ai\n# read image with flipping to rgb for visualization purpose\ndef readImage(path):\n    # OpenCV reads the image in bgr format by default\n    bgr_img = cv2.imread(path)\n    b,g,r = cv2.split(bgr_img)\n    rgb_img = cv2.merge([r,g,b])\n    return rgb_img","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:49:23.029251Z","iopub.execute_input":"2025-05-27T09:49:23.030044Z","iopub.status.idle":"2025-05-27T09:49:23.034315Z","shell.execute_reply.started":"2025-05-27T09:49:23.030016Z","shell.execute_reply":"2025-05-27T09:49:23.033358Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_images(index, num_image_each):\n    '''\n    num_image_each needs to be an integer number >= 2\n    '''\n    fig, ax = plt.subplots(1, num_image_each, figsize=(20, 4))\n\n    for i, idx in enumerate(index):\n        path = os.path.join(train_path, train_labels.iloc[idx]['id'])\n        ax[i].imshow(readImage(path + '.tif'))\n\n    ax[0].set_ylabel('samples', size='large')\n    \n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:49:23.357161Z","iopub.execute_input":"2025-05-27T09:49:23.357588Z","iopub.status.idle":"2025-05-27T09:49:23.36299Z","shell.execute_reply.started":"2025-05-27T09:49:23.35756Z","shell.execute_reply":"2025-05-27T09:49:23.362112Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# plot some random negative and positive images  \n\n# number of positive or negative each\nnum_image_each = 8\n\n# Index of postive and negative samples\nindex_negative = np.random.choice(train_labels[train_labels['label']==0].index,num_image_each)\nindex_positive = np.random.choice(train_labels[train_labels['label']==1].index,num_image_each)\nprint(\"random negative index: \", index_negative)\nprint(\"random positive index: \", index_positive)\n\n# Plot some negatives images\nprint(\"\\nplot some negative images\")\nplot_images(index= index_negative, num_image_each=num_image_each)\n\n# Plot some positive images\nprint(\"\\nplot some positive images\")\nplot_images(index= index_positive, num_image_each=num_image_each)","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:49:47.041402Z","iopub.execute_input":"2025-05-27T09:49:47.042041Z","iopub.status.idle":"2025-05-27T09:49:48.909946Z","shell.execute_reply.started":"2025-05-27T09:49:47.042013Z","shell.execute_reply":"2025-05-27T09:49:48.909122Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### Context Knowledge:\nAccording to [Libre Pathology](https://librepathology.org/wiki/Lymph_node_metastasis), **these features might help us to identify tumor tissue** under microscopic:\n\n\n- +/-Cells with cytologic features of malignancy.\n  - Nuclear pleomorphism (variation in size, shape and staining).\n  - Nuclear atypia:\n    - Nuclear enlargement.\n    - Irregular nuclear membrane.\n    - Irregular chromatin pattern, esp. asymmetry.\n    - Large or irregular nucleolus.\n  - Abundant mitotic figures.\n\nWikipedia for [H&E stain ](https://en.wikipedia.org/wiki/H%26E_stain)","metadata":{}},{"cell_type":"markdown","source":"#### Draw some plots to see distributions of pixels in different color chanels for positive and negative examples","metadata":{}},{"cell_type":"code","source":"# seperate positive and negaive sample and get means\npos_samples = sample_X[sample_y == 1]\nneg_samples = sample_X[sample_y == 0]\n\n# means brightness of positive and negative samples\npos_sample_avgs=np.mean(pos_samples, axis=(1,2,3))\nneg_sample_avgs=np.mean(neg_samples, axis=(1,2,3))\n\n# max brightness of positive and negative samples\npos_sample_maxs=np.max(pos_samples, axis=(1,2,3))\nneg_sample_maxs=np.max(neg_samples, axis=(1,2,3))\n\n# min brightness of positive and negative samples\npos_sample_mins=np.min(pos_samples, axis=(1,2,3))\nneg_sample_mins=np.min(neg_samples, axis=(1,2,3))","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:49:53.68474Z","iopub.execute_input":"2025-05-27T09:49:53.685425Z","iopub.status.idle":"2025-05-27T09:49:53.75946Z","shell.execute_reply.started":"2025-05-27T09:49:53.685363Z","shell.execute_reply":"2025-05-27T09:49:53.758721Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# plot box plots for positive and negative sample chanel distributions\nchannels = [\"Red\", \"Green\", \"Blue\", \"RGB\"]\nsample_types = [\"Positive\", \"Negative\"]\nsamples = [pos_samples, neg_samples]\n\nfig, axs = plt.subplots(4, 2, sharey=True, figsize=(5, 8), dpi=150)\n\nfor i, channel in enumerate(channels):\n    for j, sample_type in enumerate(sample_types):\n        # Get the corresponding samples\n        current_samples = samples[j]\n\n        # Create filled box plot (vertically)\n        bp = axs[i, j].boxplot(current_samples[:, :, :, i % 3].flatten(), vert=True, patch_artist=True)\n        axs[i, j].set_title(f\"{sample_type} samples {channel}\")\n\n        # Set box color\n        colors = ['lightblue', 'lightcoral']\n        for patch, color in zip(bp['boxes'], colors):\n            patch.set_facecolor(color)\n\n        # Set y-axis label for the last column\n        if j == 1:\n            axs[i, j].set_ylabel(channel, rotation='horizontal', labelpad=35, fontsize=12)\n\n# Set common labels\nfor i in range(4):\n    axs[i, 0].set_ylabel(\"Channels\")\n    axs[i, 1].set_xlabel(\"Pixel value\")\n\n# Adjust layout\nfig.tight_layout()\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:49:54.707904Z","iopub.execute_input":"2025-05-27T09:49:54.70855Z","iopub.status.idle":"2025-05-27T09:49:56.960737Z","shell.execute_reply.started":"2025-05-27T09:49:54.708521Z","shell.execute_reply":"2025-05-27T09:49:56.959898Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# plot histograms for positive and negative sample chanel distributions\nnr_of_bins = 256\nchannels = [\"Red\", \"Green\", \"Blue\", \"RGB\"]\nsample_types = [\"Positive\", \"Negative\"]\nsamples = [pos_samples, neg_samples]\n\nfig, axs = plt.subplots(4, 2, figsize=(15, 8))\n\nfor i, channel in enumerate(channels):\n    for j, sample_type in enumerate(sample_types):\n        # Get the corresponding samples\n        current_samples = samples[j]\n\n        # Create histogram\n        axs[i, j].hist(current_samples[:, :, :, i % 3].flatten(), bins=nr_of_bins, density=True)\n        axs[i, j].set_title(f\"{sample_type} samples {channel}\")\n\n        # Set y-axis label for the last column\n        if j == 1:\n            axs[i, j].set_ylabel(channel, rotation='horizontal', labelpad=35, fontsize=12)\n\n# Set common labels\nfor i in range(4):\n    axs[i, 0].set_ylabel(\"Relative frequency\")\n    axs[i, 1].set_xlabel(\"Pixel value\")\n\n# Adjust layout\nfig.tight_layout()\n\n# Show the plot\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:50:01.039869Z","iopub.execute_input":"2025-05-27T09:50:01.040551Z","iopub.status.idle":"2025-05-27T09:50:05.453643Z","shell.execute_reply.started":"2025-05-27T09:50:01.04052Z","shell.execute_reply":"2025-05-27T09:50:05.452783Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot mean,min,max brightness histograms for positive and negative samples\nfig,axs = plt.subplots(3,2,figsize=(10, 10))\naxs[0,0].hist(pos_sample_avgs,bins= 60,density=True);\naxs[0,1].hist(neg_sample_avgs,bins= 60,density=True);\naxs[0,0].set_title(\"Mean brightness, positive samples\");\naxs[0,1].set_title(\"Mean brightness, negative samples\");\naxs[0,0].set_ylabel(\"Relative frequency\")\n\naxs[1,0].hist(pos_sample_mins,bins= 60,density=True);\naxs[1,1].hist(neg_sample_mins,bins= 60,density=True);\naxs[1,0].set_title(\"Min brightness, positive samples\");\naxs[1,1].set_title(\"Min brightness, negative samples\");\naxs[1,0].set_ylabel(\"Relative frequency\")\n\naxs[2,0].hist(pos_sample_maxs,bins= 60,density=True);\naxs[2,1].hist(neg_sample_maxs,bins= 60,density=True);\naxs[2,0].set_title(\"Max brightness, positive samples\");\naxs[2,1].set_title(\"Max brightness, negative samples\");\naxs[2,0].set_ylabel(\"Relative frequency\")\n","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:50:08.462651Z","iopub.execute_input":"2025-05-27T09:50:08.463287Z","iopub.status.idle":"2025-05-27T09:50:09.732225Z","shell.execute_reply.started":"2025-05-27T09:50:08.463259Z","shell.execute_reply":"2025-05-27T09:50:09.731448Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"##### Note the threshold below I found after image inspection. Since those codes are comment out, I just want to show some abnormal pictures here is samples.","metadata":{}},{"cell_type":"code","source":"# Calculate brightness for each image. \nbrightest_values = np.max(sample_X, axis=(1, 2, 3))\ndarkest_values = np.min(sample_X, axis=(1, 2, 3))\n\n# Find indices of images with abnormal brightness\n# Note the threshold I found after later abnormal image inspectaion on all images.\nabnormal_brightness_indices_sx = np.where(\n    (brightest_values < 75) | (darkest_values  >  180)\n)[0]\n\n\nprint( 'abnormal_indices in sample X: ',abnormal_brightness_indices_sx)\nprint( 'number of abnormal_indices in sample X: ',abnormal_brightness_indices_sx.shape[0])","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:50:12.862711Z","iopub.execute_input":"2025-05-27T09:50:12.86346Z","iopub.status.idle":"2025-05-27T09:50:12.882514Z","shell.execute_reply.started":"2025-05-27T09:50:12.86343Z","shell.execute_reply":"2025-05-27T09:50:12.88162Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# check brightest and darkest pixel of abnormal images in samples\nfor index in abnormal_brightness_indices_sx: \n    print(f'Image {index} brightest pixel value is {np.max(sample_X[index,:,:,:])}, darkest pixel value is {np.min(sample_X[index,:,:,:])}')\n   \n# Plot abnormal images in samples\nfig, ax = plt.subplots(1, len(abnormal_brightness_indices_sx), figsize=(2.5 * len(abnormal_brightness_indices_sx) , 2))\n\nfor i, idx in enumerate(abnormal_brightness_indices_sx):\n    ax[i].imshow(sample_X[idx,:,:,:])\n    ax[0].set_ylabel('abnormals in samples', size='large')\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:50:14.081127Z","iopub.execute_input":"2025-05-27T09:50:14.081703Z","iopub.status.idle":"2025-05-27T09:50:14.268359Z","shell.execute_reply.started":"2025-05-27T09:50:14.081675Z","shell.execute_reply":"2025-05-27T09:50:14.267593Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"background-color: #FFF2CC; padding: 10px;\">\n<span style=\"font-size: larger;\">\n\n### What I've found during EDA:\n    \n- There are **220,025 samples** in training data and **57,458** samples in testing data. \n- **Train/Test split: 80%/20%**.\n- There are **130,908 negative and 89,117 postitive labels.  Positive lable %: 41%**.\n- **Images has shape (96, 96, 3)**\n- **Pixal values range from 0 to 255**. \n- **Negative samples seems to have more brighter (200-250 higer value) pixels in red, green, blue, and RGB mixed chanels**.\n- **There are a few abnormal images**. Some image's smallest pixel values seems high, meaning they are too bright. Also, maybe a few images are too dark.\n    \n### Few Data preprocessing/clearning idea based on EDA:\n- For now, I don't plan to do further resampling to balance data. 40% positive samples using AUC seems okay. We are required to use AUC for this project. AUC isn't too sensitive to imbalanced dataset unless data is extremely imbalanced.\n- Normalize data using /.255. I will do this in image augmentation though.\n- Do some further inspectation on abnormal images. Then, exclude those abnormal images when copy all data to directories for training. \n    ","metadata":{}},{"cell_type":"markdown","source":"# <a id=\"5\"></a>\n\n<div style=\"text-align: center; background-color: #81badd; font-size:70%; padding: 2px;border-radius:10px;\">\n    <h2>Part 1.3 Abnormal images in all data </h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"### Abnormal brightness image inspection\n\n* Note I commented codes below out because those takes lots of RAM to include in final run.","metadata":{}},{"cell_type":"code","source":"# Load all data\n# st =time.time()\n# PCT = 1\n# n,X,y = load_data(pct = PCT, df=df)\n# print(time.time() - st)","metadata":{"execution":{"iopub.status.busy":"2025-05-27T07:57:14.004259Z","iopub.execute_input":"2025-05-27T07:57:14.004896Z","iopub.status.idle":"2025-05-27T07:57:14.008557Z","shell.execute_reply.started":"2025-05-27T07:57:14.004869Z","shell.execute_reply":"2025-05-27T07:57:14.007658Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Get and save abnormal indices\n# Calculate brightness for each image\n# brightest_values = np.max(X, axis=(1, 2, 3))\n# darkest_values = np.min(X, axis=(1, 2, 3))\n\n# # Find indices of images with abnormal brightness\n# abnormal_brightness_indices = np.where(\n#     (brightest_values < 75) | (darkest_values  >  180)\n# )[0]\n\n# print('Labels:', y[abnormal_brightness_indices])\n# print( 'number_of_abnormal_indices: ',abnormal_brightness_indices.shape[0])\n# np.save('/kaggle/working/abnormal_brightness_indices.npy', abnormal_brightness_indices)","metadata":{"execution":{"iopub.status.busy":"2025-05-27T07:57:16.369497Z","iopub.execute_input":"2025-05-27T07:57:16.370228Z","iopub.status.idle":"2025-05-27T07:57:16.374068Z","shell.execute_reply.started":"2025-05-27T07:57:16.370197Z","shell.execute_reply":"2025-05-27T07:57:16.373111Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# indices_string = f'abnormal_brightness_indices = {abnormal_brightness_indices.tolist()}'\n# print(indices_string)","metadata":{"execution":{"iopub.status.busy":"2025-05-27T07:57:16.72397Z","iopub.execute_input":"2025-05-27T07:57:16.724258Z","iopub.status.idle":"2025-05-27T07:57:16.72774Z","shell.execute_reply.started":"2025-05-27T07:57:16.724237Z","shell.execute_reply":"2025-05-27T07:57:16.726937Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# get abonormal indices (hard-coded to avoid re-run unecessary codes)\n\nabnormal_brightness_indices = [563, 1184, 1501, 2000, 2817, 3872, 4067, 4079, 4204, 4382, 4660, 5190, 5636, 6520, 6907, 7012, 7063, 7173, 8134, 8325, 8406, 8525, 9010, 9307, 9390, 9415, 9764, 9832, 10200, 10337, 11326, 11915, 12771, 12965, 13724, 13781, 13801, 13945, 14335, 14742, 16689, 16899, 17809, 18386, 18613, 18743, 18810, 18842, 19022, 19336, 19456, 19579, 20019, 20243, 20294, 20687, 21713, 21844, 22923, 23084, 23135, 23263, 23306, 23396, 23429, 23962, 25503, 25930, 25966, 26266, 26758, 26771, 27134, 27188, 27888, 28005, 29274, 30035, 30323, 30481, 30754, 31352, 32016, 32212, 32230, 33071, 33392, 33946, 34005, 34921, 35086, 35763, 35875, 36062, 36452, 36614, 36760, 36881, 37404, 37613, 37645, 37842, 38089, 38117, 38127, 38387, 38412, 38414, 39151, 39191, 40038, 40084, 40884, 41171, 41442, 41749, 41753, 42940, 43235, 43236, 43370, 43526, 43561, 43689, 43818, 44236, 45707, 46058, 47419, 48056, 48270, 49635, 49889, 50295, 50638, 51614, 52251, 53086, 53427, 53976, 54446, 54638, 54699, 54990, 55000, 55351, 55515, 56591, 56675, 57504, 57552, 57674, 58043, 58090, 58625, 58835, 59690, 59823, 61334, 61379, 61784, 61857, 62393, 62534, 62926, 62935, 63407, 63672, 65585, 66062, 66656, 66741, 66927, 67452, 67502, 67903, 68159, 68220, 68583, 68607, 68709, 69071, 71056, 72897, 72979, 73372, 73666, 73822, 74334, 74386, 74536, 74573, 74953, 75278, 75379, 75745, 75963, 76168, 76746, 76841, 77341, 78310, 78668, 78798, 79156, 79894, 80524, 81094, 81098, 81320, 81358, 81408, 82393, 82518, 82852, 82945, 83303, 83354, 83393, 83479, 83559, 83651, 84554, 84670, 84843, 85021, 85154, 85329, 85503, 86727, 87077, 87873, 87961, 88128, 88459, 88581, 89759, 89969, 90442, 90585, 91986, 92803, 93235, 93709, 94200, 94521, 95061, 95276, 95387, 95468, 95530, 95583, 95713, 95749, 96035, 96640, 96929, 98927, 99296, 99298, 99458, 99806, 99927, 100380, 100402, 100605, 100872, 101543, 102082, 102620, 102621, 102712, 102983, 104141, 104192, 104290, 104634, 104948, 105141, 105326, 106057, 106439, 106867, 106929, 107026, 107212, 107304, 107776, 107990, 108110, 108674, 108920, 109101, 109393, 110252, 110412, 111151, 112387, 112620, 112621, 113174, 113192, 113553, 114720, 114790, 114829, 115397, 115582, 116799, 117191, 117287, 117913, 118377, 118507, 118610, 118776, 118841, 119605, 119616, 119789, 120325, 120532, 121005, 121226, 121275, 121380, 121631, 122261, 122391, 122864, 123934, 124240, 124444, 124626, 125824, 126393, 126487, 126969, 127079, 127196, 127335, 127460, 128059, 128263, 128451, 128973, 129338, 129416, 129806, 130123, 130194, 130257, 130390, 130754, 131060, 131557, 131992, 132402, 132404, 132812, 133154, 133211, 133297, 134614, 134666, 135599, 135615, 135963, 136110, 136477, 137134, 137851, 137963, 138273, 138381, 138484, 138859, 139331, 139744, 139817, 139920, 140298, 140399, 141505, 142139, 142468, 142889, 143209, 143309, 143518, 144010, 144265, 144534, 145298, 145955, 146020, 146522, 147783, 147952, 148527, 148545, 149393, 150294, 150759, 151089, 152243, 152549, 152837, 153256, 153651, 154094, 154133, 154238, 154270, 154316, 154729, 155536, 155865, 156253, 156422, 156512, 157071, 157228, 157581, 158455, 158705, 158929, 159130, 159801, 160097, 160480, 160588, 161045, 161076, 161195, 161303, 161308, 161379, 161443, 162186, 162323, 162487, 162796, 162797, 163116, 163449, 163932, 164002, 164046, 164052, 164077, 164803, 164944, 165394, 166311, 166593, 167008, 167058, 167767, 167825, 167864, 167871, 168099, 168270, 168456, 168649, 168791, 168916, 168965, 169061, 169068, 169210, 169227, 169355, 169583, 169903, 169976, 169981, 171091, 171456, 171564, 171975, 172152, 172427, 172437, 172764, 172930, 173665, 174007, 176035, 176314, 176474, 176507, 176653, 177333, 177629, 177955, 178367, 178552, 179100, 179995, 180425, 180441, 180761, 181059, 181208, 181280, 181376, 181591, 181871, 181964, 182028, 182140, 182439, 182459, 182777, 182788, 182858, 183068, 183462, 183901, 185126, 185474, 185743, 185935, 186136, 186217, 186687, 187048, 187650, 188092, 188603, 188964, 188981, 189406, 189854, 190096, 190141, 190482, 190589, 190899, 190913, 191313, 191407, 191684, 191902, 191908, 192341, 192456, 192477, 192569, 193074, 193194, 193618, 193623, 193840, 193844, 193879, 194220, 194275, 195125, 196150, 196775, 197231, 197234, 198170, 198226, 198325, 198356, 199282, 199308, 199634, 200032, 200286, 201505, 201625, 201990, 202078, 202103, 202167, 203077, 203493, 203656, 203741, 204340, 204405, 204624, 205096, 205142, 205274, 205585, 205687, 205721, 206409, 207278, 207917, 208118, 208354, 208505, 209273, 209725, 210278, 210915, 211026, 211056, 211398, 211546, 211585, 212253, 212672, 212872, 212923, 213383, 213575, 213581, 213944, 214139, 214809, 214824, 215062, 215389, 216153, 216309, 216557, 216759, 217036, 217188, 217258, 217345, 217380, 217563, 217685, 217993, 218102, 218511, 219449, 219534, 219598]\n\n","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:50:31.057532Z","iopub.execute_input":"2025-05-27T09:50:31.057877Z","iopub.status.idle":"2025-05-27T09:50:31.075473Z","shell.execute_reply.started":"2025-05-27T09:50:31.05785Z","shell.execute_reply":"2025-05-27T09:50:31.074529Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#  Plot some abnormal images\n# fig, ax = plt.subplots(1, 10, figsize=(30, 4))\n\n# for i, idx in enumerate(abnormal_brightness_indices[11:21]):\n#     ax[i].imshow(X[idx, :, :, :])\n#     ax[0].set_ylabel('samples', size='large')\n# plt.show()","metadata":{"execution":{"iopub.status.busy":"2025-05-27T07:57:18.84446Z","iopub.execute_input":"2025-05-27T07:57:18.844738Z","iopub.status.idle":"2025-05-27T07:57:18.848534Z","shell.execute_reply.started":"2025-05-27T07:57:18.844716Z","shell.execute_reply":"2025-05-27T07:57:18.847593Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### During the inspectations above, I found that if the brightest pixel is less than 75 or the darkest pixel is greater than 180, the images are basically blank. We should remove them.\n\n","metadata":{}},{"cell_type":"markdown","source":"# <a id=\"6\"></a>\n<div style=\"text-align: center; background-color: #AFDDCE; font-size:100%; padding: 5px;border-radius:10px 10px;\">\n    <h1>Part 2: Model Building and training</h1>\n</div>","metadata":{}},{"cell_type":"markdown","source":"# <a id=\"7\"></a>\n\n<div style=\"text-align: center; background-color: #81badd; font-size:70%; padding: 2px;border-radius:10px;\">\n    <h2>Part 2.1: Copy all normal images to directories, with train test split</h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"#### Now let's load all (but only normal data) in. Copy files to directories takes a longer time (compared to copy to int8 numpy array)but can save more memory for later model training.\n\n","metadata":{}},{"cell_type":"code","source":"# drop abnormal images indices in df\ndf.drop(abnormal_brightness_indices, inplace=True)\nprint('df size after removing abnormal images: ', df.shape[0])","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:50:34.391905Z","iopub.execute_input":"2025-05-27T09:50:34.392487Z","iopub.status.idle":"2025-05-27T09:50:34.415494Z","shell.execute_reply.started":"2025-05-27T09:50:34.392457Z","shell.execute_reply":"2025-05-27T09:50:34.414602Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# train test split\ndf_train, df_val = train_test_split(df, test_size=0.10, random_state=100)","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:50:35.545919Z","iopub.execute_input":"2025-05-27T09:50:35.546482Z","iopub.status.idle":"2025-05-27T09:50:35.592015Z","shell.execute_reply.started":"2025-05-27T09:50:35.546452Z","shell.execute_reply":"2025-05-27T09:50:35.591359Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# reset index using id\ndf.set_index('id', inplace=True)\n","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:50:36.898477Z","iopub.execute_input":"2025-05-27T09:50:36.8991Z","iopub.status.idle":"2025-05-27T09:50:36.903967Z","shell.execute_reply.started":"2025-05-27T09:50:36.899072Z","shell.execute_reply":"2025-05-27T09:50:36.902924Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create directories for model augmentation\nwork_dir ='/kaggle/working/data'\ntrain_dir =os.path.join(work_dir,'train')\nval_dir=os.path.join(work_dir,'val')\n\nif os.path.exists(work_dir):# remove the entire directory and its contents if already exist\n    shutil.rmtree(work_dir)\n\nfor fold in [train_dir, val_dir]:\n    for subf in [\"0\", \"1\"]:\n        os.makedirs(os.path.join(fold, subf)) # this will create necessary parent directories as well :)       ","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:50:38.557232Z","iopub.execute_input":"2025-05-27T09:50:38.557898Z","iopub.status.idle":"2025-05-27T09:50:38.56325Z","shell.execute_reply.started":"2025-05-27T09:50:38.557867Z","shell.execute_reply":"2025-05-27T09:50:38.562366Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from concurrent.futures import ThreadPoolExecutor\nimport shutil, os\n\ndef copy_file(args):\n    src, dst = args\n    shutil.copy(src, dst)\n\n# Prepare file paths\ntrain_files = [\n    (os.path.join(train_path, id + '.tif'), os.path.join(train_dir, str(df.loc[id, 'label']), id + '.tif'))\n    for id in df_train['id'].values\n]\n\nval_files = [\n    (os.path.join(train_path, id + '.tif'), os.path.join(val_dir, str(df.loc[id, 'label']), id + '.tif'))\n    for id in df_val['id'].values\n]\n\n# Use multithreading\nwith ThreadPoolExecutor(max_workers=16) as executor:\n    executor.map(copy_file, train_files + val_files)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T09:50:40.843558Z","iopub.execute_input":"2025-05-27T09:50:40.844221Z","iopub.status.idle":"2025-05-27T09:53:22.988054Z","shell.execute_reply.started":"2025-05-27T09:50:40.844193Z","shell.execute_reply":"2025-05-27T09:53:22.987322Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# check image counts in Training and validation directories\n\n#Training\nprint('Positive images in training directory: ', len(os.listdir(os.path.join(train_dir, '1'))))\nprint('Negative images in training directory: ',len(os.listdir(os.path.join(train_dir, '0'))))\nprint('% Positive images in training directory: ',len(os.listdir(os.path.join(train_dir, '1'))) / (len(os.listdir(os.path.join(train_dir, '0'))) + len(os.listdir(os.path.join(train_dir, '1'))) ),'\\n')\n\n# Validation\nprint('Positive images in validation directory: ', len(os.listdir(os.path.join(val_dir, '1'))))\nprint('Negative images in validation directory: ',len(os.listdir(os.path.join(val_dir, '0'))))\nprint('% Positive images in validation directory: ',len(os.listdir(os.path.join(val_dir, '1'))) / (len(os.listdir(os.path.join(val_dir, '0'))) + len(os.listdir(os.path.join(val_dir, '1'))) ), '\\n')\n\n# Training vs. Validation\nprint('Images in Training directory: ', len(os.listdir(os.path.join(train_dir, '1'))) + len(os.listdir(os.path.join(train_dir, '0'))) )\nprint('Images in  validation directory: ', len(os.listdir(os.path.join(val_dir, '1'))) + len(os.listdir(os.path.join(val_dir, '0'))) )\nprint('% Training images: ', (len(os.listdir(os.path.join(train_dir, '1'))) + len(os.listdir(os.path.join(train_dir, '0')))) / (len(os.listdir(os.path.join(train_dir, '1'))) + len(os.listdir(os.path.join(train_dir, '0')))+len(os.listdir(os.path.join(val_dir, '1'))) + len(os.listdir(os.path.join(val_dir, '0'))) )    )\nprint('\\nTotal images: ',  (len(os.listdir(os.path.join(train_dir, '1'))) + len(os.listdir(os.path.join(train_dir, '0')))+len(os.listdir(os.path.join(val_dir, '1'))) + len(os.listdir(os.path.join(val_dir, '0'))) )    )\n","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:53:42.962308Z","iopub.execute_input":"2025-05-27T09:53:42.963184Z","iopub.status.idle":"2025-05-27T09:53:43.776209Z","shell.execute_reply.started":"2025-05-27T09:53:42.963155Z","shell.execute_reply":"2025-05-27T09:53:43.775172Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Collect garbage to free up memory\ndel pos_samples \ndel neg_samples \ndel pos_sample_avgs \ndel neg_sample_avgs \ndel pos_sample_mins\ndel neg_sample_mins \ndel pos_sample_maxs \ndel neg_sample_maxs\ndel df\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:53:49.381814Z","iopub.execute_input":"2025-05-27T09:53:49.382131Z","iopub.status.idle":"2025-05-27T09:53:49.635137Z","shell.execute_reply.started":"2025-05-27T09:53:49.382106Z","shell.execute_reply":"2025-05-27T09:53:49.633768Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <a id=\"8\"></a>\n\n<div style=\"text-align: center; background-color: #81badd; font-size:70%; padding: 2px;border-radius:10px;\">\n    <h2>Part 2.2: Baseline model</h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"### First I will just try to build a baseline model that I think might perform well. A few helpful knowledges. 👇\n","metadata":{}},{"cell_type":"markdown","source":"\n<div style=\"background-color: #FCFFE9; padding: 10px;\">\n    \n### Why CNN for images?📝\n  - **Problem caused by the big images data** (Even more and more pixels nowadays!)\n    - **Computation capacity**: consider the weights we need for flattening out all the pixels and use them as the first layer of a fully connected neural network!\n    - **Overfit**: Fully-connected neural network would have too many parameters. Lots of weights from empty spaces in the image aren't useful and they are redundant.\n  - **Translational Invariance in images**\n    - Fully connected layer would think an object shift to a different position is a different object beacause it learns the whole image.\n    - CNN uses small and different sized filters to learn/detect different primitive features (Ex: an ear, an eye, intensity of color, texture, shape, orientation)in an image. **What a filter learn in one position of an image can be re-used in other positions.** CNN slices those filters across the image and the **learning weights are shared** in different positions. This results in way less parameters/overfitting/waste of computational resources\n\n### Why multiple convolution layers? 📝\nThere is a **field of view (FOV)** effect of CNN. Here is an example:\n- 1st convolution filter of size 3 * 3 would look at the input image 3 * 3 at a time and assign it to 1 output pixal. That is, each size 1 pixal in the 1st convolution feature map gets it's information from an area of 3 * 3 in the original input image. 3 * 3 isn't big area, so the 1st filter can only detect local features like lines etc.\n- 2nd convolution filter of size 3 * 3 would look at the 1st convolution feature map image 3 * 3 at a time and assign it to 1 output pixal. That is, each size 1 pixal in the 2nd convolution feature map gets it's information from an area of size 3 * 3 in the 1st convolution feature map image. That is equivalent to look at an area of 9 * 9 from the original image ( (3 * 3) * (3 * 3)) based on layer 1. So, 2nd convolution layers can see more global features.\n- The FOV becomes larger when CNN gets more layers so CNN would see more global feature in later layers. For example:\n  - Layer 1 may detect features like edges and lines, color contrasts, textures, etc.\n  - Layer 2 may detect simple shapes and patterns formed by the combination of features from the first layer. like curves,combinations of edges, corners, and textures.\n  - Later layers may detect more global features like eyes, face, cars, etc. <br>\n \n(More details of FOV in [this paper](https://arxiv.org/abs/1311.2901))\n    \n ### Why prefer small filters? 📝\n\nNowadays small filters (like 3 * 3 )are more popular:\n\n- **Computation effeciency**: for instance, two 3 * 3 filters and one 5 * 5 filter both get you (N - 4) * (N -4) feature map. But the 3 * 3 filters only need 18 (=3 * 3 * 2) but 5 * 5 filter need 25 (5 * 5) parameters. \n- Deal with **Translational Invariance in images** better\n\n ### Why use 1 * 1 Filters? 📝\n    \n-  1 * 1 filters mimic applying a fully-connected layer to each pixals. This adds non-linearity to let layer learn more complex function.(Details in [this video](https://www.youtube.com/watch?v=c1RBQzKsDCk&list=PLpFsSf5Dm-pd5d3rjNtIXUHT-v7bdaEIe&index=115)). To illustrate: \n   - If we have a 6 * 6 * 32 (depth) input layer. At each slice step, a 1 * 1 filter would take 32 numbers (depth), multiply each 32 numbers by 32 weights, apply activation function (maybe Relu), then output 1 output units. This is like creating one neuron of a fully connected layer .\n   - So, each 1 * 1 filter would output a 6 * 6 * 32 input layer to 6 * 6 feature map. With more filters, the output would be 6 * 6 * number of filters. Thus, adding more filters is like creating more output neuron units.\n- 1 * 1 filters allow to shrink or add number of channels. Can be used for computation efficiency.  \n    \n### Why use Pooling layer? 📝\n\n    \n- Pooling layer in CNN uses subsampling method to **down-sample the spatial dimensions of the input feature maps**. \n- A **pooling filter of size p would usualy reduce the input image size by 1/p** (without considering the input image depth, which goes to 1 for each filter). Explaination:\n    - Output_image_size Output_image_size = (N + 2 * p - f)/s +1 (N: input image size; p: padding size; f: filter size; s: stride size)\n    - Pooling filter typically don't have padding. (p=0)\n    - The pooling window is usually applied to non-overlapping regions of the input. (So,usually s=f.)\n    - To illustrate, if we have a 10 * 10 image and a 2*2 pooling filter, the output size would be (10-2)/2+1 = 5. \n- Pooling filter sizes are **usually 2 * 2** or some **even size** (while **convolution filters are often odd size**), so usually it reduce half of the image size  \n- Two popular types:\n   - **Max pooling**: the maximum is taken within each region. For example, if the input image has a deepth of 3 and we have a 2*2 max pooling filter. The output pixal for each step would be the maximum of the 12 numbers (12=2 * 2 * 3). It emphasize the most dominant features in local regions and is used most often.\n   - **Average pooling**: the average is taken within each region. It provides a smoother down-sampling effect.\n\n\n    \n- Why using pooling layer? - It shrinks down the feature map size for a few reasons:\n  - **Computational efficiency**\n  - Helps to see more **global features**\n  - **Feature generalization**: retain most important information and ignore the details. The layer would be less sensitive to small shifts in the input.\n","metadata":{}},{"cell_type":"markdown","source":"### My Plan for training the model\n- Reduce Training Error:\n  - First focus on reducing training error, which means it's comprehensive and complex enough that it fits training data well. \n- Reduce Validation Error:\n  - Once satisfied with the training performance, introduce overfitting prevention techniques:\n    - Dropout\n    - L2 or L1 normalization\n    - Batch normalization (might just add it earlier to reduce training error)\n    - Data Augmentation\n- Attempt of Transfer Learning:\n  - After just a couple of attempts, if the model is still too simple, explore transfer learning with pre-trained models.\n  - Include structures for overfitting from the start\n\nNote: I've decided to withhold result submission until I'm thoroughly satisfied with my validation score (probably not until the last run😅). Given that the testing score is unlikely to significantly surpass the validation one, I'm prioritizing efficient use of resources. However, I will meticulously analyze each model's performance, conducting a detailed comparison of training and validation metrics. This analysis is crucial for guiding the next steps in training direction.\n","metadata":{}},{"cell_type":"markdown","source":"# Transfer Learning","metadata":{}},{"cell_type":"markdown","source":"<div style=\"background-color: #FCFFE9; padding: 10px;\">\n    \n## Transfer Learning\n\n#### Why transfer learning and how to use it?\n    \nIf dataset is very large, it can easily take more than a few weeks to train. It's good to use transfer learning in this case. A few ways to use it:\n- **As fixed feature extractor**: froze the weights in the convoluntion layers. Can add/train fully connected layers.\n- **Fine-tuning the CNN**: you data is big, similar to the transfer learning dataset but not exactly the same. You can 'de-forze' and train some CNN layers weights too.\n- **Use only part of CNN layer**: add a few of your own CNN layers also.\n  - A typical example for medical images.\n    - **Medical images** are harder to collect and even harder to label since they need to labeled by doctors.\n    - Trained by big ImageNet data, the first few layers can find some primitive features that can be used for medical images. We can use the earlier layers and add some later layers to train specifically for medical images.\n   \n    \n    \n#### When to use a pre-trained network?\n- **❌New dataset is small (ex: 20,000 with 100 categories) and similar to the original dataset (ex: 2,000,000 with 1000 categories)**\n    - Since the network is huge, you have froze the weight, which means you aren't training the model. Even not consider computation resource that you can actually re-train the model, the model is going to overfit since it's too big and more suitable for much bigger data.\n    \n- **✅New dataset is large and similar to the original dataset**\n    - can gain training time by using either fixed or fresh weights\n- **❌New dataset is small but very different from the original dataset**\n    - Overfit as in case in. Additionally, if very different, the transfer learning efficiency would be low.\n- **✅New dataset is large and very different from the original dataset**\n    - Can still use part of feature extractor of earlier layers. Still can save lots of training time.\n    \nVarious [models](https://keras.io/applications/) pre-trained on ImageNet.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"background-color: #FCFFE9; padding: 10px;\">\n    \n## InceptionV3 \n#### Insights from InceptionV3 Architecture:\n- Lower Layers:\n  - Basic convolutional Operations: such as Conv2d_1a_3x3, Conv2d_2a_3x3, and Conv2d_2b_3x3 \n  - Fundamental local features extractraction: such as edges, textures, colors, and gradients\n  - Simple Shapes and Contrast: In addition to low-level features, identify simple shapes and capturing contrast variations in the images.\n- Intermediate Layers:\n  - Mixed Operations:  blocks such as Mixed_3a, Mixed_4a, and Mixed_5a, combining different filter sizes and operations in parallel.\n  - Repetitive Structures capturing: such as repetitive structures within the images, recognizing patterns that repeat across the dataset.\n  - Color Variations and Relationships: further understand color variations and establish hierarchical relationships between different parts of objects in the images.\n  - Object Boundaries: delineating object boundaries, enhancing the model's ability to recognize distinct objects and their spatial relationships.\n- Deeper Inception Blocks:\n  - Mixed Operations: Blocks such as Mixed_5b, Mixed_6a, Mixed_7a\n  - Complex Shapes and Structures: improved understanding of complex shapes and intricate structures present in the images.\n  - Discrimination of Patterns: enhanced discrimination of complex patterns, distinguish fine-grained details and textures.\n\n#### My Adaptation Plan:\n- Feature Extraction:\n  - Utilize earlier blocks for low-level features \n- Intermediate Layers Modification:\n  - Introduce extra convolution layers for medical image nuances.\n  - Post 'mixed2' layer in InceptionV3, the feature map is roughly (9 * 9). Add padding to prevent rapid size reduction (retain spatial information)\n- Enhancements with 1x1 Filters:\n  - Incorporate 1x1 filters for improved learning.\n- Handle Overfitting:\n  - Batch normalization\n  - Dropout layers.\n- Image Augmentations for better generalization.","metadata":{}},{"cell_type":"markdown","source":"## Model 1 - transfer learning with Inception V3","metadata":{}},{"cell_type":"code","source":"# Build model 1\n\n# Load the InceptionV3 model\nincept_model = InceptionV3(weights='imagenet', include_top=False, input_shape=(96, 96, 3))\n\n# Freeze the weights of the layers.\nfor layer in incept_model.layers:\n    layer.trainable = False\n\n# get the last_layer of inception we want to use\nlast_layer = incept_model.get_layer('mixed2')\nprint('last layer output shape: ', last_layer.output_shape)\n\n# Create a new model using the functional API\nbase_model = Model(inputs=incept_model.input, outputs=last_layer.output)\n\n# Add layers to the model\nmymodel_3 = Sequential([\n    base_model,\n    \n    Conv2D(64, (3, 3), padding='same', activation='relu'), # keep the width and height\n    BatchNormalization(),\n    Conv2D(64, (1, 1),activation='relu'),\n    BatchNormalization(),\n    Dropout(0.2),\n    \n    Conv2D(128, (3, 3), activation='relu'),\n    BatchNormalization(),            \n    Conv2D(64, (1, 1), activation='relu'),\n    BatchNormalization(),\n    Dropout(0.2),\n\n    Flatten(),\n    Dense(128, activation='relu'),\n    BatchNormalization(),\n    Dropout(0.2),\n    \n    Dense(128, activation='relu'),\n    BatchNormalization(),\n    Dropout(0.2),\n    \n    Dense(1, activation='sigmoid')\n])\n\n# Print the model summary\nmymodel_3.build(input_shape=(None, 96, 96, 3))\nmymodel_3.summary()","metadata":{"execution":{"iopub.status.busy":"2025-05-27T08:00:48.184581Z","iopub.execute_input":"2025-05-27T08:00:48.18518Z","iopub.status.idle":"2025-05-27T08:00:55.693657Z","shell.execute_reply.started":"2025-05-27T08:00:48.185152Z","shell.execute_reply":"2025-05-27T08:00:55.69285Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from tensorflow.keras.optimizers.schedules import ExponentialDecay\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.metrics import AUC\n\n# Define the learning rate schedule (exponential decay)\ninitial_learning_rate = 0.001\ndecay_steps = 1000\ndecay_rate = 0.9 \nlr_schedule = ExponentialDecay(initial_learning_rate, decay_steps, decay_rate, staircase=True)\n\n# Compile the model\nmymodel_3.compile(\n    optimizer=Adam(learning_rate=lr_schedule),\n    loss='binary_crossentropy',\n    metrics=['accuracy', AUC()]  # Added 'accuracy' here\n)\n","metadata":{"execution":{"iopub.status.busy":"2025-05-27T08:02:09.989484Z","iopub.execute_input":"2025-05-27T08:02:09.98981Z","iopub.status.idle":"2025-05-27T08:02:10.94928Z","shell.execute_reply.started":"2025-05-27T08:02:09.989777Z","shell.execute_reply":"2025-05-27T08:02:10.948593Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Fit model 3\nBATCH_SIZE = 128\nEPOCHS = 5\n\n# data generator\ndatagen = ImageDataGenerator(\n    rescale=1./255,\n    rotation_range=20,\n    width_shift_range=0.2,\n    height_shift_range=0.2,\n    vertical_flip=True,\n    horizontal_flip=True,\n)\n\ndatagen_val = ImageDataGenerator(\n    rescale=1./255,\n)\n\ntrain_generator = datagen.flow_from_directory(directory=train_dir,\n                                            batch_size=BATCH_SIZE,\n                                            class_mode='binary',\n                                            target_size=(96, 96))\n\nval_generator = datagen_val.flow_from_directory(directory=val_dir, \n                                                batch_size=BATCH_SIZE, \n                                                class_mode='binary',\n                                               target_size=(96,96))\n\nsteps_per_epoch = len(train_generator) \nvalidation_steps = len(val_generator)\n\n# Fit model 3\nhistory_3 = mymodel_3.fit(\n    train_generator,\n    validation_data=val_generator,\n    epochs=EPOCHS,\n    steps_per_epoch=steps_per_epoch,\n    validation_steps=validation_steps,\n    verbose=1\n)","metadata":{"execution":{"iopub.status.busy":"2025-05-27T09:59:00.390741Z","iopub.execute_input":"2025-05-27T09:59:00.391061Z","iopub.status.idle":"2025-05-27T09:59:03.406283Z","shell.execute_reply.started":"2025-05-27T09:59:00.391036Z","shell.execute_reply":"2025-05-27T09:59:03.40543Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(history_3.history.keys())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T08:43:45.48086Z","iopub.execute_input":"2025-05-27T08:43:45.481501Z","iopub.status.idle":"2025-05-27T08:43:45.485774Z","shell.execute_reply.started":"2025-05-27T08:43:45.481474Z","shell.execute_reply":"2025-05-27T08:43:45.484816Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Extract final values from training history\nfinal_loss = history_3.history['loss'][-1]\nfinal_accuracy = history_3.history['accuracy'][-1]\nfinal_auc = history_3.history['auc'][-1]\n\nfinal_val_loss = history_3.history['val_loss'][-1]\nfinal_val_accuracy = history_3.history['val_accuracy'][-1]\nfinal_val_auc = history_3.history['val_auc'][-1]\n\n# Print final metrics\nprint(\"\\n=== Final Training & Validation Metrics ===\")\nprint(f\"Train Loss:          {final_loss:.4f}\")\nprint(f\"Train Accuracy:      {final_accuracy:.4f}\")\nprint(f\"Train AUC:           {final_auc:.4f}\")\nprint(f\"Validation Loss:     {final_val_loss:.4f}\")\nprint(f\"Validation Accuracy: {final_val_accuracy:.4f}\")\nprint(f\"Validation AUC:      {final_val_auc:.4f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T08:45:06.739289Z","iopub.execute_input":"2025-05-27T08:45:06.739984Z","iopub.status.idle":"2025-05-27T08:45:06.745714Z","shell.execute_reply.started":"2025-05-27T08:45:06.739956Z","shell.execute_reply":"2025-05-27T08:45:06.744841Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot metrics\nepochs = np.arange(1, len(history_3.history['loss']) + 1)\nplt.figure(figsize=(10, 4))\n\n# Accuracy plot\nplt.subplot(1, 2, 1)\nplt.plot(epochs, history_3.history['accuracy'], label='Train Accuracy')\nplt.plot(epochs, history_3.history['val_accuracy'], label='Val Accuracy')\nplt.xlabel('Epochs')\nplt.ylabel('Accuracy')\nplt.title('Accuracy Over Epochs')\nplt.legend()\nplt.grid(True)\n\n# Loss plot\nplt.subplot(1, 2, 2)\nplt.plot(epochs, history_3.history['loss'], label='Train Loss')\nplt.plot(epochs, history_3.history['val_loss'], label='Val Loss')\nplt.xlabel('Epochs')\nplt.ylabel('Loss')\nplt.title('Loss Over Epochs')\nplt.legend()\nplt.grid(True)\n\nplt.tight_layout()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2025-05-27T08:50:13.890734Z","iopub.execute_input":"2025-05-27T08:50:13.891474Z","iopub.status.idle":"2025-05-27T08:50:14.28542Z","shell.execute_reply.started":"2025-05-27T08:50:13.891446Z","shell.execute_reply":"2025-05-27T08:50:14.284587Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Model 3 Performance:\n- Achieved satisfactory training AUC of around 0.96 and a commendable validation AUC of 0.97 in the last epoch.\n- validation AUC plateaued after 8 epochs.\n\n### Next Steps:\n- Model evaluation, focusing on metrics such as precision, recall, F1 score, and confusion matrix to gain a holistic understanding of the model's performance. This approach ensures a balanced assessment of the model's predictive capabilities beyond AUC, providing insights into its precision, recall, and overall classification effectiveness.","metadata":{}},{"cell_type":"code","source":"# Predict validation set \nsamples_val = len(os.listdir(os.path.join(val_dir, '1'))) + len(os.listdir(os.path.join(val_dir, '0')))\nval_generator_for_pred = datagen_val.flow_from_directory(directory=val_dir, \n                                                batch_size=1, \n                                                class_mode='binary',\n                                               target_size=(96,96),\n                                                shuffle=False)\n\n\nval_predictions = mymodel_3.predict_generator(val_generator_for_pred,steps=samples_val, verbose=1)","metadata":{"execution":{"iopub.status.busy":"2025-05-27T08:46:44.387677Z","iopub.execute_input":"2025-05-27T08:46:44.388594Z","iopub.status.idle":"2025-05-27T08:48:22.650591Z","shell.execute_reply.started":"2025-05-27T08:46:44.388564Z","shell.execute_reply":"2025-05-27T08:48:22.649815Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Actual vs. predicted classes for validation set\n\n# actual\nval_labels = val_generator_for_pred.classes\n\n# convert probability to classes using 0.5 shreshold\nbinary_val_predictions = (val_predictions > 0.5).astype(int)\n","metadata":{"execution":{"iopub.status.busy":"2025-05-27T08:49:15.603978Z","iopub.execute_input":"2025-05-27T08:49:15.604292Z","iopub.status.idle":"2025-05-27T08:49:15.609303Z","shell.execute_reply.started":"2025-05-27T08:49:15.60427Z","shell.execute_reply":"2025-05-27T08:49:15.608367Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot ROC curve\n\n# Calculate AUC\nfpr, tpr, thresholds = roc_curve(val_labels, val_predictions)\nroc_auc = auc(fpr, tpr)\n\n# Plot \nplt.figure(figsize=(4, 4))\nplt.plot(fpr, tpr, color='orange', lw=2, label='ROC curve (AUC = {:.2f})'.format(roc_auc))\nplt.plot([0, 1], [0, 1], color='navy', lw=2, linestyle='--', label='Random Guessing')\nplt.xlabel('FPR')\nplt.ylabel('TPR')\nplt.title('ROC Curve')\nplt.legend(loc='lower right')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2025-05-27T08:49:16.783868Z","iopub.execute_input":"2025-05-27T08:49:16.784207Z","iopub.status.idle":"2025-05-27T08:49:16.945791Z","shell.execute_reply.started":"2025-05-27T08:49:16.784182Z","shell.execute_reply":"2025-05-27T08:49:16.944956Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Get recall, predicision, F1, accuracy scores for validation set\nrecall = recall_score(val_labels, binary_val_predictions)\nprecision = precision_score(val_labels, binary_val_predictions)\nF1 = f1_score(val_labels, binary_val_predictions)\nacc= accuracy_score(val_labels, binary_val_predictions)\n\nprint(\"Recall on validation set:\", recall)\nprint(\"Precision on validation set:\", precision)\nprint(\"F1 score n on validation set:\", F1)\nprint(\"Accuracy score n on validation set:\", acc)","metadata":{"execution":{"iopub.status.busy":"2025-05-27T08:49:20.413251Z","iopub.execute_input":"2025-05-27T08:49:20.413569Z","iopub.status.idle":"2025-05-27T08:49:20.438359Z","shell.execute_reply.started":"2025-05-27T08:49:20.413545Z","shell.execute_reply":"2025-05-27T08:49:20.437616Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot confusion matrix\ncm = confusion_matrix(val_labels, binary_val_predictions)\ndisp = ConfusionMatrixDisplay(confusion_matrix=cm, display_labels=['Non-cancer', 'Cancer'])\ndisp.plot(cmap='Blues', values_format='d')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2025-05-27T08:50:45.313554Z","iopub.execute_input":"2025-05-27T08:50:45.314425Z","iopub.status.idle":"2025-05-27T08:50:45.468689Z","shell.execute_reply.started":"2025-05-27T08:50:45.314394Z","shell.execute_reply":"2025-05-27T08:50:45.467877Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Xception Model - Model 2","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.applications import Xception\nfrom tensorflow.keras.models import Model, Sequential\nfrom tensorflow.keras.layers import Conv2D, BatchNormalization, Dropout, Dense, GlobalAveragePooling2D\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.callbacks import ReduceLROnPlateau, EarlyStopping\nfrom tensorflow.keras.metrics import AUC\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\n\n# Load Xception\nxception_model = Xception(weights='imagenet', include_top=False, input_shape=(96, 96, 3))\n\n# Unfreeze top 20 layers for fine-tuning\nfor layer in xception_model.layers[:-20]:\n    layer.trainable = False\nfor layer in xception_model.layers[-20:]:\n    layer.trainable = True\n\n# Define base model\nlast_layer = xception_model.get_layer('block14_sepconv2_act')\nbase_model = Model(inputs=xception_model.input, outputs=last_layer.output)\n\n# Define the complete model\nmymodel_3 = Sequential([\n    base_model,\n    \n    GlobalAveragePooling2D(),  # Replaces Flatten + reduces params\n\n    Dense(256, activation='relu', kernel_regularizer='l2'),\n    BatchNormalization(),\n    Dropout(0.4),\n    \n    Dense(128, activation='relu', kernel_regularizer='l2'),\n    BatchNormalization(),\n    Dropout(0.3),\n    \n    Dense(1, activation='sigmoid')\n])\n\n# Compile with dynamic LR adjustment\nmymodel_3.compile(\n    optimizer=Adam(learning_rate=1e-3),\n    loss='binary_crossentropy',\n    metrics=['accuracy',AUC()]\n)\n\n# Callbacks\ncallbacks = [\n    ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=1, verbose=1),\n    EarlyStopping(monitor='val_loss', patience=2, restore_best_weights=True, verbose=1)\n]\n\n# Fit model\nhistory_3 = mymodel_3.fit(\n    train_generator,\n    validation_data=val_generator,\n    epochs=5,\n    steps_per_epoch=steps_per_epoch,\n    validation_steps=validation_steps,\n    callbacks=callbacks,\n    verbose=1\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T10:01:31.171134Z","iopub.execute_input":"2025-05-27T10:01:31.171933Z","iopub.status.idle":"2025-05-27T10:44:06.806997Z","shell.execute_reply.started":"2025-05-27T10:01:31.171906Z","shell.execute_reply":"2025-05-27T10:44:06.806266Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(history_3.history.keys())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T10:44:52.590438Z","iopub.execute_input":"2025-05-27T10:44:52.590769Z","iopub.status.idle":"2025-05-27T10:44:52.595536Z","shell.execute_reply.started":"2025-05-27T10:44:52.590744Z","shell.execute_reply":"2025-05-27T10:44:52.594632Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Extract final values from training history\nfinal_loss = history_3.history['loss'][-1]\nfinal_accuracy = history_3.history['accuracy'][-1]\nfinal_auc = history_3.history['auc_2'][-1]\n\nfinal_val_loss = history_3.history['val_loss'][-1]\nfinal_val_accuracy = history_3.history['val_accuracy'][-1]\nfinal_val_auc = history_3.history['val_auc_2'][-1]\nfinal_lr = history_3.history['lr'][-1]\n\n# Print final metrics\nprint(\"\\n=== Final Training & Validation Metrics ===\")\nprint(f\"Train Loss:          {final_loss:.4f}\")\nprint(f\"Train Accuracy:      {final_accuracy:.4f}\")\nprint(f\"Train AUC:           {final_auc:.4f}\")\nprint(f\"Validation Loss:     {final_val_loss:.4f}\")\nprint(f\"Validation Accuracy: {final_val_accuracy:.4f}\")\nprint(f\"Validation AUC:      {final_val_auc:.4f}\")\nprint(f\"Final Learning Rate: {final_lr:.6f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T10:45:51.918274Z","iopub.execute_input":"2025-05-27T10:45:51.919037Z","iopub.status.idle":"2025-05-27T10:45:51.925067Z","shell.execute_reply.started":"2025-05-27T10:45:51.919011Z","shell.execute_reply":"2025-05-27T10:45:51.924243Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"epochs = np.arange(1, len(history_3.history['loss']) + 1)\nplt.figure(figsize=(15, 4))\n\n# Accuracy Plot\nplt.subplot(1, 3, 1)\nplt.plot(epochs, history_3.history['accuracy'], label='Train Accuracy')\nplt.plot(epochs, history_3.history['val_accuracy'], label='Val Accuracy')\nplt.xlabel('Epochs')\nplt.ylabel('Accuracy')\nplt.title('Accuracy Over Epochs')\nplt.legend()\nplt.grid(True)\n\n# AUC Plot\nplt.subplot(1, 3, 2)\nplt.plot(epochs, history_3.history['auc_2'], label='Train AUC')\nplt.plot(epochs, history_3.history['val_auc_2'], label='Val AUC')\nplt.xlabel('Epochs')\nplt.ylabel('AUC')\nplt.title('AUC Over Epochs')\nplt.legend()\nplt.grid(True)\n\n# Loss Plot\nplt.subplot(1, 3, 3)\nplt.plot(epochs, history_3.history['loss'], label='Train Loss')\nplt.plot(epochs, history_3.history['val_loss'], label='Val Loss')\nplt.xlabel('Epochs')\nplt.ylabel('Loss')\nplt.title('Loss Over Epochs')\nplt.legend()\nplt.grid(True)\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T10:46:26.091893Z","iopub.execute_input":"2025-05-27T10:46:26.092497Z","iopub.status.idle":"2025-05-27T10:46:26.685974Z","shell.execute_reply.started":"2025-05-27T10:46:26.092467Z","shell.execute_reply":"2025-05-27T10:46:26.684934Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Predict validation set \nsamples_val = len(os.listdir(os.path.join(val_dir, '1'))) + len(os.listdir(os.path.join(val_dir, '0')))\nval_generator_for_pred = datagen_val.flow_from_directory(directory=val_dir, \n                                                batch_size=1, \n                                                class_mode='binary',\n                                               target_size=(96,96),\n                                                shuffle=False)\n\n\nval_predictions = mymodel_3.predict_generator(val_generator_for_pred,steps=samples_val, verbose=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T10:47:04.643507Z","iopub.execute_input":"2025-05-27T10:47:04.644126Z","iopub.status.idle":"2025-05-27T10:49:35.396494Z","shell.execute_reply.started":"2025-05-27T10:47:04.644097Z","shell.execute_reply":"2025-05-27T10:49:35.395738Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Actual vs. predicted classes for validation set\n\n# actual\nval_labels = val_generator_for_pred.classes\n\n# convert probability to classes using 0.5 shreshold\nbinary_val_predictions = (val_predictions > 0.5).astype(int)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T10:49:39.765779Z","iopub.execute_input":"2025-05-27T10:49:39.766337Z","iopub.status.idle":"2025-05-27T10:49:39.770401Z","shell.execute_reply.started":"2025-05-27T10:49:39.76631Z","shell.execute_reply":"2025-05-27T10:49:39.769589Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot ROC curve\n\n# Calculate AUC\nfpr, tpr, thresholds = roc_curve(val_labels, val_predictions)\nroc_auc = auc(fpr, tpr)\n\n# Plot \nplt.figure(figsize=(4, 4))\nplt.plot(fpr, tpr, color='orange', lw=2, label='ROC curve (AUC = {:.2f})'.format(roc_auc))\nplt.plot([0, 1], [0, 1], color='navy', lw=2, linestyle='--', label='Random Guessing')\nplt.xlabel('FPR')\nplt.ylabel('TPR')\nplt.title('ROC Curve')\nplt.legend(loc='lower right')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T10:49:41.408224Z","iopub.execute_input":"2025-05-27T10:49:41.409027Z","iopub.status.idle":"2025-05-27T10:49:41.562675Z","shell.execute_reply.started":"2025-05-27T10:49:41.408999Z","shell.execute_reply":"2025-05-27T10:49:41.561814Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Get recall, predicision, F1, accuracy scores for validation set\nrecall = recall_score(val_labels, binary_val_predictions)\nprecision = precision_score(val_labels, binary_val_predictions)\nF1 = f1_score(val_labels, binary_val_predictions)\nacc= accuracy_score(val_labels, binary_val_predictions)\n\nprint(\"Recall on validation set:\", recall)\nprint(\"Precision on validation set:\", precision)\nprint(\"F1 score n on validation set:\", F1)\nprint(\"Accuracy score n on validation set:\", acc)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T10:49:44.106696Z","iopub.execute_input":"2025-05-27T10:49:44.107023Z","iopub.status.idle":"2025-05-27T10:49:44.130021Z","shell.execute_reply.started":"2025-05-27T10:49:44.106995Z","shell.execute_reply":"2025-05-27T10:49:44.129208Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot confusion matrix\ncm = confusion_matrix(val_labels, binary_val_predictions)\ndisp = ConfusionMatrixDisplay(confusion_matrix=cm, display_labels=['Non-cancer', 'Cancer'])\ndisp.plot(cmap='Blues', values_format='d')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T10:49:48.723001Z","iopub.execute_input":"2025-05-27T10:49:48.723654Z","iopub.status.idle":"2025-05-27T10:49:48.864638Z","shell.execute_reply.started":"2025-05-27T10:49:48.723627Z","shell.execute_reply":"2025-05-27T10:49:48.863754Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# DenseNet121 - Model 3","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.applications import DenseNet121\nfrom tensorflow.keras.models import Model, Sequential\nfrom tensorflow.keras.layers import Conv2D, BatchNormalization, Dropout, Dense, GlobalAveragePooling2D\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.callbacks import ReduceLROnPlateau, EarlyStopping\nfrom tensorflow.keras.metrics import AUC\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\n\n# Load DenseNet121 with ImageNet weights\ndensenet_model = DenseNet121(weights='imagenet', include_top=False, input_shape=(96, 96, 3))\n\n# Unfreeze the last 30 layers for fine-tuning\nfor layer in densenet_model.layers[:-30]:\n    layer.trainable = False\nfor layer in densenet_model.layers[-30:]:\n    layer.trainable = True\n\n# Define base model\nlast_layer = densenet_model.get_layer('conv5_block16_2_conv')  # Last dense block before global pooling\nbase_model = Model(inputs=densenet_model.input, outputs=last_layer.output)\n\n# Define the complete model\nmymodel_3 = Sequential([\n    base_model,\n    \n    GlobalAveragePooling2D(),  # Reduces parameters, acts like flattening\n\n    Dense(512, activation='relu', kernel_regularizer='l2'),\n    BatchNormalization(),\n    Dropout(0.5),\n    \n    Dense(256, activation='relu', kernel_regularizer='l2'),\n    BatchNormalization(),\n    Dropout(0.4),\n    \n    Dense(128, activation='relu', kernel_regularizer='l2'),\n    BatchNormalization(),\n    Dropout(0.3),\n    \n    Dense(1, activation='sigmoid')  # For binary classification\n])\n\n# Compile with dynamic LR adjustment\nmymodel_3.compile(\n    optimizer=Adam(learning_rate=1e-3),\n    loss='binary_crossentropy',\n    metrics=['accuracy',AUC()]\n)\n\n# Callbacks for learning rate adjustment and early stopping\ncallbacks = [\n    ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=2, verbose=1),\n    EarlyStopping(monitor='val_loss', patience=3, restore_best_weights=True, verbose=1)\n]\n\n# Define aggressive data augmentation strategies\ndatagen = ImageDataGenerator(\n    rescale=1./255,\n    rotation_range=30,\n    width_shift_range=0.2,\n    height_shift_range=0.2,\n    shear_range=0.2,\n    zoom_range=0.2,\n    horizontal_flip=True,\n    fill_mode='nearest'\n)\n\ndatagen_val = ImageDataGenerator(rescale=1./255)\n\n# Assuming `train_dir` and `val_dir` are your directories\ntrain_generator = datagen.flow_from_directory(\n    directory=train_dir,\n    batch_size=128,\n    class_mode='binary',\n    target_size=(96, 96)\n)\n\nval_generator = datagen_val.flow_from_directory(\n    directory=val_dir,\n    batch_size=128,\n    class_mode='binary',\n    target_size=(96, 96)\n)\n\nsteps_per_epoch = len(train_generator)\nvalidation_steps = len(val_generator)\n\n# Fit the model\nhistory_3 = mymodel_3.fit(\n    train_generator,\n    validation_data=val_generator,\n    epochs=5,\n    steps_per_epoch=steps_per_epoch,\n    validation_steps=validation_steps,\n    callbacks=callbacks,\n    verbose=1\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T10:49:58.583786Z","iopub.execute_input":"2025-05-27T10:49:58.584182Z","iopub.status.idle":"2025-05-27T11:33:26.78659Z","shell.execute_reply.started":"2025-05-27T10:49:58.584151Z","shell.execute_reply":"2025-05-27T11:33:26.785867Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(history_3.history.keys())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T11:34:32.639454Z","iopub.execute_input":"2025-05-27T11:34:32.640054Z","iopub.status.idle":"2025-05-27T11:34:32.644474Z","shell.execute_reply.started":"2025-05-27T11:34:32.640023Z","shell.execute_reply":"2025-05-27T11:34:32.643639Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Extract final values using correct keys\nfinal_loss = history_3.history['loss'][-1]\nfinal_accuracy = history_3.history['accuracy'][-1]\nfinal_auc = history_3.history['auc_3'][-1]\n\nfinal_val_loss = history_3.history['val_loss'][-1]\nfinal_val_accuracy = history_3.history['val_accuracy'][-1]\nfinal_val_auc = history_3.history['val_auc_3'][-1]\n\n# Print final metrics\nprint(\"\\n=== Final Training & Validation Metrics ===\")\nprint(f\"Train Loss:          {final_loss:.4f}\")\nprint(f\"Train Accuracy:      {final_accuracy:.4f}\")\nprint(f\"Train AUC:           {final_auc:.4f}\")\nprint(f\"Validation Loss:     {final_val_loss:.4f}\")\nprint(f\"Validation Accuracy: {final_val_accuracy:.4f}\")\nprint(f\"Validation AUC:      {final_val_auc:.4f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T11:35:13.749068Z","iopub.execute_input":"2025-05-27T11:35:13.74987Z","iopub.status.idle":"2025-05-27T11:35:13.756049Z","shell.execute_reply.started":"2025-05-27T11:35:13.749842Z","shell.execute_reply":"2025-05-27T11:35:13.755235Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import numpy as np\nimport matplotlib.pyplot as plt\n\nepochs = np.arange(1, len(history_3.history['loss']) + 1)\nplt.figure(figsize=(15, 4))\n\n# Accuracy Plot\nplt.subplot(1, 3, 1)\nplt.plot(epochs, history_3.history['accuracy'], label='Train Accuracy')\nplt.plot(epochs, history_3.history['val_accuracy'], label='Val Accuracy')\nplt.xlabel('Epochs')\nplt.ylabel('Accuracy')\nplt.title('Accuracy Over Epochs')\nplt.legend()\nplt.grid(True)\n\n# AUC Plot\nplt.subplot(1, 3, 2)\nplt.plot(epochs, history_3.history['auc_3'], label='Train AUC')\nplt.plot(epochs, history_3.history['val_auc_3'], label='Val AUC')\nplt.xlabel('Epochs')\nplt.ylabel('AUC')\nplt.title('AUC Over Epochs')\nplt.legend()\nplt.grid(True)\n\n# Loss Plot\nplt.subplot(1, 3, 3)\nplt.plot(epochs, history_3.history['loss'], label='Train Loss')\nplt.plot(epochs, history_3.history['val_loss'], label='Val Loss')\nplt.xlabel('Epochs')\nplt.ylabel('Loss')\nplt.title('Loss Over Epochs')\nplt.legend()\nplt.grid(True)\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T11:35:31.671059Z","iopub.execute_input":"2025-05-27T11:35:31.671444Z","iopub.status.idle":"2025-05-27T11:35:32.275512Z","shell.execute_reply.started":"2025-05-27T11:35:31.671415Z","shell.execute_reply":"2025-05-27T11:35:32.274642Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Predict validation set \nsamples_val = len(os.listdir(os.path.join(val_dir, '1'))) + len(os.listdir(os.path.join(val_dir, '0')))\nval_generator_for_pred = datagen_val.flow_from_directory(directory=val_dir, \n                                                batch_size=1, \n                                                class_mode='binary',\n                                               target_size=(96,96),\n                                                shuffle=False)\n\n\nval_predictions = mymodel_3.predict_generator(val_generator_for_pred,steps=samples_val, verbose=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T11:35:40.128484Z","iopub.execute_input":"2025-05-27T11:35:40.12907Z","iopub.status.idle":"2025-05-27T11:39:49.728416Z","shell.execute_reply.started":"2025-05-27T11:35:40.129038Z","shell.execute_reply":"2025-05-27T11:39:49.7276Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Actual vs. predicted classes for validation set\n\n# actual\nval_labels = val_generator_for_pred.classes\n\n# convert probability to classes using 0.5 shreshold\nbinary_val_predictions = (val_predictions > 0.5).astype(int)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T11:39:49.730145Z","iopub.execute_input":"2025-05-27T11:39:49.730449Z","iopub.status.idle":"2025-05-27T11:39:49.73487Z","shell.execute_reply.started":"2025-05-27T11:39:49.730412Z","shell.execute_reply":"2025-05-27T11:39:49.733941Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot ROC curve\n\n# Calculate AUC\nfpr, tpr, thresholds = roc_curve(val_labels, val_predictions)\nroc_auc = auc(fpr, tpr)\n\n# Plot \nplt.figure(figsize=(4, 4))\nplt.plot(fpr, tpr, color='orange', lw=2, label='ROC curve (AUC = {:.2f})'.format(roc_auc))\nplt.plot([0, 1], [0, 1], color='navy', lw=2, linestyle='--', label='Random Guessing')\nplt.xlabel('FPR')\nplt.ylabel('TPR')\nplt.title('ROC Curve')\nplt.legend(loc='lower right')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T11:39:51.298651Z","iopub.execute_input":"2025-05-27T11:39:51.299315Z","iopub.status.idle":"2025-05-27T11:39:51.458166Z","shell.execute_reply.started":"2025-05-27T11:39:51.29929Z","shell.execute_reply":"2025-05-27T11:39:51.4573Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Get recall, predicision, F1, accuracy scores for validation set\nrecall = recall_score(val_labels, binary_val_predictions)\nprecision = precision_score(val_labels, binary_val_predictions)\nF1 = f1_score(val_labels, binary_val_predictions)\nacc= accuracy_score(val_labels, binary_val_predictions)\n\nprint(\"Recall on validation set:\", recall)\nprint(\"Precision on validation set:\", precision)\nprint(\"F1 score n on validation set:\", F1)\nprint(\"Accuracy score n on validation set:\", acc)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T11:39:54.028745Z","iopub.execute_input":"2025-05-27T11:39:54.029405Z","iopub.status.idle":"2025-05-27T11:39:54.05217Z","shell.execute_reply.started":"2025-05-27T11:39:54.029364Z","shell.execute_reply":"2025-05-27T11:39:54.051301Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot confusion matrix\ncm = confusion_matrix(val_labels, binary_val_predictions)\ndisp = ConfusionMatrixDisplay(confusion_matrix=cm, display_labels=['Non-cancer', 'Cancer'])\ndisp.plot(cmap='Blues', values_format='d')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T11:39:55.085273Z","iopub.execute_input":"2025-05-27T11:39:55.085615Z","iopub.status.idle":"2025-05-27T11:39:55.234488Z","shell.execute_reply.started":"2025-05-27T11:39:55.085589Z","shell.execute_reply":"2025-05-27T11:39:55.233586Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# MobileNet","metadata":{}},{"cell_type":"code","source":"# Fit model \nBATCH_SIZE = 128\nEPOCHS = 5\n\n# data generator\ndatagen = ImageDataGenerator(\n    rescale=1./255,\n    rotation_range=20,\n    width_shift_range=0.2,\n    height_shift_range=0.2,\n    vertical_flip=True,\n    horizontal_flip=True,\n)\n\ndatagen_val = ImageDataGenerator(\n    rescale=1./255,\n)\n\ntrain_generator = datagen.flow_from_directory(directory=train_dir,\n                                            batch_size=BATCH_SIZE,\n                                            class_mode='binary',\n                                            target_size=(96, 96))\n\nval_generator = datagen_val.flow_from_directory(directory=val_dir, \n                                                batch_size=BATCH_SIZE, \n                                                class_mode='binary',\n                                               target_size=(96,96))\n\nsteps_per_epoch = len(train_generator) \nvalidation_steps = len(val_generator)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T11:40:01.395129Z","iopub.execute_input":"2025-05-27T11:40:01.396031Z","iopub.status.idle":"2025-05-27T11:40:04.33943Z","shell.execute_reply.started":"2025-05-27T11:40:01.396003Z","shell.execute_reply":"2025-05-27T11:40:04.338736Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from tensorflow.keras.applications import MobileNetV2\nfrom tensorflow.keras.models import Model, Sequential\nfrom tensorflow.keras.layers import GlobalAveragePooling2D, Dense, BatchNormalization, Dropout\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.callbacks import ReduceLROnPlateau, EarlyStopping\nfrom tensorflow.keras.metrics import AUC\n\n# Load MobileNetV2 with ImageNet weights\nmobilenet_model = MobileNetV2(weights='imagenet', include_top=False, input_shape=(96, 96, 3))\n\n# Freeze most layers and unfreeze the top few for fine-tuning\nfor layer in mobilenet_model.layers[:-20]:\n    layer.trainable = False\nfor layer in mobilenet_model.layers[-20:]:\n    layer.trainable = True\n\n# Define the base output\nlast_layer = mobilenet_model.output\n\n# Build the full model\nx = GlobalAveragePooling2D()(last_layer)\nx = Dense(256, activation='relu', kernel_regularizer='l2')(x)\nx = BatchNormalization()(x)\nx = Dropout(0.4)(x)\n\nx = Dense(128, activation='relu', kernel_regularizer='l2')(x)\nx = BatchNormalization()(x)\nx = Dropout(0.3)(x)\n\npredictions = Dense(1, activation='sigmoid')(x)\n\nmymodel = Model(inputs=mobilenet_model.input, outputs=predictions)\n\n# Compile the model\nmymodel_3.compile(\n    optimizer=Adam(learning_rate=1e-3),\n    loss='binary_crossentropy',\n    metrics=['accuracy',AUC()]\n)\n\n# Define callbacks\ncallbacks = [\n    ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=1, verbose=1),\n    EarlyStopping(monitor='val_loss', patience=2, restore_best_weights=True, verbose=1)\n]\n\n# Fit the model\nhistory_3 = mymodel_3.fit(\n    train_generator,\n    validation_data=val_generator,\n    epochs=5,\n    steps_per_epoch=steps_per_epoch,\n    validation_steps=validation_steps,\n    callbacks=callbacks,\n    verbose=1\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T11:40:39.863187Z","iopub.execute_input":"2025-05-27T11:40:39.86398Z","iopub.status.idle":"2025-05-27T12:23:56.521975Z","shell.execute_reply.started":"2025-05-27T11:40:39.863952Z","shell.execute_reply":"2025-05-27T12:23:56.521369Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(history_3.history.keys())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T12:24:02.330915Z","iopub.execute_input":"2025-05-27T12:24:02.331245Z","iopub.status.idle":"2025-05-27T12:24:02.335755Z","shell.execute_reply.started":"2025-05-27T12:24:02.331217Z","shell.execute_reply":"2025-05-27T12:24:02.334924Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"final_loss = history_3.history['loss'][-1]\nfinal_accuracy = history_3.history['accuracy'][-1]\nfinal_auc = history_3.history['auc_4'][-1]  # changed key\n\nfinal_val_loss = history_3.history['val_loss'][-1]\nfinal_val_accuracy = history_3.history['val_accuracy'][-1]\nfinal_val_auc = history_3.history['val_auc_4'][-1]  # changed key\n\n# Print final metrics\nprint(\"\\n=== Final Training & Validation Metrics ===\")\nprint(f\"Train Loss:          {final_loss:.4f}\")\nprint(f\"Train Accuracy:      {final_accuracy:.4f}\")\nprint(f\"Train AUC:           {final_auc:.4f}\")\nprint(f\"Validation Loss:     {final_val_loss:.4f}\")\nprint(f\"Validation Accuracy: {final_val_accuracy:.4f}\")\nprint(f\"Validation AUC:      {final_val_auc:.4f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T12:24:55.868156Z","iopub.execute_input":"2025-05-27T12:24:55.86849Z","iopub.status.idle":"2025-05-27T12:24:55.874746Z","shell.execute_reply.started":"2025-05-27T12:24:55.868464Z","shell.execute_reply":"2025-05-27T12:24:55.873935Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"epochs = np.arange(1, len(history_3.history['loss']) + 1)\nplt.figure(figsize=(15, 4))\n\n# Accuracy Plot\nplt.subplot(1, 3, 1)\nplt.plot(epochs, history_3.history['accuracy'], label='Train Accuracy')\nplt.plot(epochs, history_3.history['val_accuracy'], label='Val Accuracy')\nplt.xlabel('Epochs')\nplt.ylabel('Accuracy')\nplt.title('Accuracy Over Epochs')\nplt.legend()\nplt.grid(True)\n\n# AUC Plot\nplt.subplot(1, 3, 2)\nplt.plot(epochs, history_3.history['auc_4'], label='Train AUC')  # changed key\nplt.plot(epochs, history_3.history['val_auc_4'], label='Val AUC')  # changed key\nplt.xlabel('Epochs')\nplt.ylabel('AUC')\nplt.title('AUC Over Epochs')\nplt.legend()\nplt.grid(True)\n\n# Loss Plot\nplt.subplot(1, 3, 3)\nplt.plot(epochs, history_3.history['loss'], label='Train Loss')\nplt.plot(epochs, history_3.history['val_loss'], label='Val Loss')\nplt.xlabel('Epochs')\nplt.ylabel('Loss')\nplt.title('Loss Over Epochs')\nplt.legend()\nplt.grid(True)\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T12:25:08.791306Z","iopub.execute_input":"2025-05-27T12:25:08.791669Z","iopub.status.idle":"2025-05-27T12:25:09.401565Z","shell.execute_reply.started":"2025-05-27T12:25:08.791641Z","shell.execute_reply":"2025-05-27T12:25:09.400666Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Predict validation set \nsamples_val = len(os.listdir(os.path.join(val_dir, '1'))) + len(os.listdir(os.path.join(val_dir, '0')))\nval_generator_for_pred = datagen_val.flow_from_directory(directory=val_dir, \n                                                batch_size=1, \n                                                class_mode='binary',\n                                               target_size=(96,96),\n                                                shuffle=False)\n\n\nval_predictions = mymodel_3.predict_generator(val_generator_for_pred,steps=samples_val, verbose=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T12:25:16.02434Z","iopub.execute_input":"2025-05-27T12:25:16.02471Z","iopub.status.idle":"2025-05-27T12:29:20.618561Z","shell.execute_reply.started":"2025-05-27T12:25:16.024683Z","shell.execute_reply":"2025-05-27T12:29:20.617726Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Actual vs. predicted classes for validation set\n\n# actual\nval_labels = val_generator_for_pred.classes\n\n# convert probability to classes using 0.5 shreshold\nbinary_val_predictions = (val_predictions > 0.5).astype(int)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T12:29:24.158508Z","iopub.execute_input":"2025-05-27T12:29:24.158824Z","iopub.status.idle":"2025-05-27T12:29:24.163601Z","shell.execute_reply.started":"2025-05-27T12:29:24.1588Z","shell.execute_reply":"2025-05-27T12:29:24.162662Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot ROC curve\n\n# Calculate AUC\nfpr, tpr, thresholds = roc_curve(val_labels, binary_val_predictions)\nroc_auc = auc(fpr, tpr)\n\n# Plot \nplt.figure(figsize=(4, 4))\nplt.plot(fpr, tpr, color='orange', lw=2, label='ROC curve (AUC = {:.2f})'.format(roc_auc))\nplt.plot([0, 1], [0, 1], color='navy', lw=2, linestyle='--', label='Random Guessing')\nplt.xlabel('FPR')\nplt.ylabel('TPR')\nplt.title('ROC Curve')\nplt.legend(loc='lower right')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T12:29:30.635364Z","iopub.execute_input":"2025-05-27T12:29:30.636007Z","iopub.status.idle":"2025-05-27T12:29:30.789725Z","shell.execute_reply.started":"2025-05-27T12:29:30.635982Z","shell.execute_reply":"2025-05-27T12:29:30.788855Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Get recall, predicision, F1, accuracy scores for validation set\nrecall = recall_score(val_labels, binary_val_predictions)\nprecision = precision_score(val_labels, binary_val_predictions)\nF1 = f1_score(val_labels, binary_val_predictions)\nacc= accuracy_score(val_labels, binary_val_predictions)\n\nprint(\"Recall on validation set:\", recall)\nprint(\"Precision on validation set:\", precision)\nprint(\"F1 score n on validation set:\", F1)\nprint(\"Accuracy score n on validation set:\", acc)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T12:29:34.14133Z","iopub.execute_input":"2025-05-27T12:29:34.142104Z","iopub.status.idle":"2025-05-27T12:29:34.167161Z","shell.execute_reply.started":"2025-05-27T12:29:34.142071Z","shell.execute_reply":"2025-05-27T12:29:34.166276Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot confusion matrix\ncm = confusion_matrix(val_labels, binary_val_predictions)\ndisp = ConfusionMatrixDisplay(confusion_matrix=cm, display_labels=['Non-cancer', 'Cancer'])\ndisp.plot(cmap='Blues', values_format='d')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T12:29:36.211037Z","iopub.execute_input":"2025-05-27T12:29:36.211328Z","iopub.status.idle":"2025-05-27T12:29:36.357635Z","shell.execute_reply.started":"2025-05-27T12:29:36.211306Z","shell.execute_reply":"2025-05-27T12:29:36.35676Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Inception without  Data Augmentation","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.applications import InceptionV3\nfrom tensorflow.keras.models import Model, Sequential\nfrom tensorflow.keras.layers import Conv2D, Flatten, Dense, BatchNormalization, Dropout\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.callbacks import ReduceLROnPlateau, EarlyStopping\nfrom tensorflow.keras.metrics import AUC\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras.optimizers.schedules import ExponentialDecay\n\n# Load the InceptionV3 model\nincept_model = InceptionV3(weights='imagenet', include_top=False, input_shape=(96, 96, 3))\n\n# Freeze the weights of the layers.\nfor layer in incept_model.layers:\n    layer.trainable = False\n\n# get the last_layer of inception we want to use\nlast_layer = incept_model.get_layer('mixed2')\nprint('last layer output shape: ', last_layer.output_shape)\n\n# Create a new model using the functional API\nbase_model = Model(inputs=incept_model.input, outputs=last_layer.output)\n\n# Add layers to the model\nmymodel_3 = Sequential([\n    base_model,\n    \n    Conv2D(64, (3, 3), padding='same', activation='relu'),\n    BatchNormalization(),\n    Conv2D(64, (1, 1), activation='relu'),\n    BatchNormalization(),\n    Dropout(0.2),\n    \n    Conv2D(128, (3, 3), activation='relu'),\n    BatchNormalization(),            \n    Conv2D(64, (1, 1), activation='relu'),\n    BatchNormalization(),\n    Dropout(0.2),\n\n    Flatten(),\n    Dense(128, activation='relu'),\n    BatchNormalization(),\n    Dropout(0.2),\n    \n    Dense(128, activation='relu'),\n    BatchNormalization(),\n    Dropout(0.2),\n    \n    Dense(1, activation='sigmoid')\n])\n\n# Print the model summary\nmymodel_3.build(input_shape=(None, 96, 96, 3))\nmymodel_3.summary()\n\n# Define the learning rate schedule\ninitial_learning_rate = 0.001\ndecay_steps = 1000\ndecay_rate = 0.9 \nlr_schedule = ExponentialDecay(initial_learning_rate, decay_steps, decay_rate, staircase=True)\n\n# Compile the model\nmymodel_3.compile(\n    optimizer=Adam(learning_rate=lr_schedule),\n    loss='binary_crossentropy',\n    metrics=[AUC()]\n)\n\n# Training config\nBATCH_SIZE = 128\nEPOCHS = 5\n\n# ImageDataGenerators without augmentation\ndatagen = ImageDataGenerator(rescale=1./255)\ndatagen_val = ImageDataGenerator(rescale=1./255)\n\ntrain_generator = datagen.flow_from_directory(directory=train_dir,\n                                              batch_size=BATCH_SIZE,\n                                              class_mode='binary',\n                                              target_size=(96, 96))\n\nval_generator = datagen_val.flow_from_directory(directory=val_dir, \n                                                batch_size=BATCH_SIZE, \n                                                class_mode='binary',\n                                                target_size=(96, 96))\n\nsteps_per_epoch = len(train_generator)\nvalidation_steps = len(val_generator)\n\n# Fit model 3\nhistory_3 = mymodel_3.fit(\n    train_generator,\n    validation_data=val_generator,\n    epochs=EPOCHS,\n    steps_per_epoch=steps_per_epoch,\n    validation_steps=validation_steps,\n    verbose=1\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-04T15:33:05.329727Z","iopub.execute_input":"2025-05-04T15:33:05.330388Z","iopub.status.idle":"2025-05-04T15:45:25.305645Z","shell.execute_reply.started":"2025-05-04T15:33:05.330359Z","shell.execute_reply":"2025-05-04T15:45:25.304848Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(history_3.history.keys())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-04T15:45:31.543103Z","iopub.execute_input":"2025-05-04T15:45:31.543875Z","iopub.status.idle":"2025-05-04T15:45:31.548127Z","shell.execute_reply.started":"2025-05-04T15:45:31.543845Z","shell.execute_reply":"2025-05-04T15:45:31.547289Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot metrics\nepochs = np.arange(1, len(history_3.history['loss']) + 1)\nplt.figure(figsize=(10, 4))\n\n# AUC plot\nplt.subplot(1, 2, 1)\nplt.plot(epochs, history_3.history['auc_1'], label='Train AUC')\nplt.plot(epochs, history_3.history['val_auc_1'], label='Val AUC')\nplt.xlabel('Epochs')\nplt.ylabel('AUC')\nplt.title('AUC Over Epochs')\nplt.legend()\nplt.grid(True)\n\n# Loss plot\nplt.subplot(1, 2, 2)\nplt.plot(epochs, history_3.history['loss'], label='Train Loss')\nplt.plot(epochs, history_3.history['val_loss'], label='Val Loss')\nplt.xlabel('Epochs')\nplt.ylabel('Loss')\nplt.title('Loss Over Epochs')\nplt.legend()\nplt.grid(True)\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-04T15:45:55.501277Z","iopub.execute_input":"2025-05-04T15:45:55.501608Z","iopub.status.idle":"2025-05-04T15:45:55.897749Z","shell.execute_reply.started":"2025-05-04T15:45:55.501586Z","shell.execute_reply":"2025-05-04T15:45:55.896905Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Predict validation set \nsamples_val = len(os.listdir(os.path.join(val_dir, '1'))) + len(os.listdir(os.path.join(val_dir, '0')))\nval_generator_for_pred = datagen_val.flow_from_directory(directory=val_dir, \n                                                batch_size=1, \n                                                class_mode='binary',\n                                               target_size=(96,96),\n                                                shuffle=False)\n\n\nval_predictions = mymodel_3.predict_generator(val_generator_for_pred,steps=samples_val, verbose=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-04T15:45:58.682531Z","iopub.execute_input":"2025-05-04T15:45:58.68322Z","iopub.status.idle":"2025-05-04T15:47:34.019467Z","shell.execute_reply.started":"2025-05-04T15:45:58.683192Z","shell.execute_reply":"2025-05-04T15:47:34.018657Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Actual vs. predicted classes for validation set\n\n# actual\nval_labels = val_generator_for_pred.classes\n\n# convert probability to classes using 0.5 shreshold\nbinary_val_predictions = (val_predictions > 0.5).astype(int)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-04T15:47:39.756255Z","iopub.execute_input":"2025-05-04T15:47:39.756671Z","iopub.status.idle":"2025-05-04T15:47:39.761522Z","shell.execute_reply.started":"2025-05-04T15:47:39.756649Z","shell.execute_reply":"2025-05-04T15:47:39.760455Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot ROC curve\n\n# Calculate AUC\nfpr, tpr, thresholds = roc_curve(val_labels, binary_val_predictions)\nroc_auc = auc(fpr, tpr)\n\n# Plot \nplt.figure(figsize=(4, 4))\nplt.plot(fpr, tpr, color='orange', lw=2, label='ROC curve (AUC = {:.2f})'.format(roc_auc))\nplt.plot([0, 1], [0, 1], color='navy', lw=2, linestyle='--', label='Random Guessing')\nplt.xlabel('FPR')\nplt.ylabel('TPR')\nplt.title('ROC Curve')\nplt.legend(loc='lower right')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-04T15:47:42.19174Z","iopub.execute_input":"2025-05-04T15:47:42.192444Z","iopub.status.idle":"2025-05-04T15:47:42.350364Z","shell.execute_reply.started":"2025-05-04T15:47:42.192416Z","shell.execute_reply":"2025-05-04T15:47:42.349359Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Get recall, predicision, F1, accuracy scores for validation set\nrecall = recall_score(val_labels, binary_val_predictions)\nprecision = precision_score(val_labels, binary_val_predictions)\nF1 = f1_score(val_labels, binary_val_predictions)\nacc= accuracy_score(val_labels, binary_val_predictions)\n\nprint(\"Recall on validation set:\", recall)\nprint(\"Precision on validation set:\", precision)\nprint(\"F1 score n on validation set:\", F1)\nprint(\"Accuracy score n on validation set:\", acc)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-04T15:47:46.240677Z","iopub.execute_input":"2025-05-04T15:47:46.241501Z","iopub.status.idle":"2025-05-04T15:47:46.263646Z","shell.execute_reply.started":"2025-05-04T15:47:46.241471Z","shell.execute_reply":"2025-05-04T15:47:46.262801Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot confusion matrix\ncm = confusion_matrix(val_labels, binary_val_predictions)\ndisp = ConfusionMatrixDisplay(confusion_matrix=cm, display_labels=['Non-cancer', 'Cancer'])\ndisp.plot(cmap='Blues', values_format='d')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-04T15:48:31.171771Z","iopub.execute_input":"2025-05-04T15:48:31.172444Z","iopub.status.idle":"2025-05-04T15:48:31.319459Z","shell.execute_reply.started":"2025-05-04T15:48:31.172418Z","shell.execute_reply":"2025-05-04T15:48:31.318553Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# VGG16","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.applications import VGG16\nfrom tensorflow.keras.models import Model\nfrom tensorflow.keras.layers import GlobalAveragePooling2D, Dense, BatchNormalization, Dropout\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.callbacks import ReduceLROnPlateau, EarlyStopping\nfrom tensorflow.keras.metrics import AUC\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\n\n# Load VGG16 with ImageNet weights\nvgg_model = VGG16(weights='imagenet', include_top=False, input_shape=(96, 96, 3))\n\n# Freeze most layers and fine-tune last few\nfor layer in vgg_model.layers[:-10]:\n    layer.trainable = False\nfor layer in vgg_model.layers[-10:]:\n    layer.trainable = True\n\n# Define the output layer from base model\nx = vgg_model.output\nx = GlobalAveragePooling2D()(x)\n\n# Add custom layers\nx = Dense(256, activation='relu', kernel_regularizer='l2')(x)\nx = BatchNormalization()(x)\nx = Dropout(0.4)(x)\n\nx = Dense(128, activation='relu', kernel_regularizer='l2')(x)\nx = BatchNormalization()(x)\nx = Dropout(0.3)(x)\n\npredictions = Dense(1, activation='sigmoid')(x)\n\n# Build final model\nmymodel_3 = Model(inputs=vgg_model.input, outputs=predictions)\n\n# Compile\nmymodel_3.compile(\n    optimizer=Adam(learning_rate=1e-3),\n    loss='binary_crossentropy',\n    metrics=['accuracy',AUC()]\n)\n\n# Callbacks\ncallbacks = [\n    ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=1, verbose=1),\n    EarlyStopping(monitor='val_loss', patience=2, restore_best_weights=True, verbose=1)\n]\n\n# Train the model\nhistory_3 = mymodel_3.fit(\n    train_generator,\n    validation_data=val_generator,\n    epochs=5,\n    steps_per_epoch=steps_per_epoch,\n    validation_steps=validation_steps,\n    callbacks=callbacks,\n    verbose=1\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T12:30:08.681802Z","iopub.execute_input":"2025-05-27T12:30:08.682721Z","iopub.status.idle":"2025-05-27T13:12:03.148678Z","shell.execute_reply.started":"2025-05-27T12:30:08.68269Z","shell.execute_reply":"2025-05-27T13:12:03.147799Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(history_3.history.keys())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T13:12:06.745769Z","iopub.execute_input":"2025-05-27T13:12:06.746417Z","iopub.status.idle":"2025-05-27T13:12:06.750672Z","shell.execute_reply.started":"2025-05-27T13:12:06.746387Z","shell.execute_reply":"2025-05-27T13:12:06.749886Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Extract final values from training history\nfinal_loss = history_3.history['loss'][-1]\nfinal_accuracy = history_3.history['accuracy'][-1]\nfinal_auc = history_3.history['auc_6'][-1]\n\nfinal_val_loss = history_3.history['val_loss'][-1]\nfinal_val_accuracy = history_3.history['val_accuracy'][-1]\nfinal_val_auc = history_3.history['val_auc_6'][-1]\n\n# Print final metrics\nprint(\"\\n=== Final Training & Validation Metrics ===\")\nprint(f\"Train Loss:          {final_loss:.4f}\")\nprint(f\"Train Accuracy:      {final_accuracy:.4f}\")\nprint(f\"Train AUC:           {final_auc:.4f}\")\nprint(f\"Validation Loss:     {final_val_loss:.4f}\")\nprint(f\"Validation Accuracy: {final_val_accuracy:.4f}\")\nprint(f\"Validation AUC:      {final_val_auc:.4f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T13:13:11.905039Z","iopub.execute_input":"2025-05-27T13:13:11.905407Z","iopub.status.idle":"2025-05-27T13:13:11.911783Z","shell.execute_reply.started":"2025-05-27T13:13:11.905361Z","shell.execute_reply":"2025-05-27T13:13:11.910923Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"epochs = np.arange(1, len(history_3.history['loss']) + 1)\nplt.figure(figsize=(15, 4))\n\n# Accuracy Plotepochs = np.arange(1, len(history_3.history['loss']) + 1)\nplt.figure(figsize=(15, 4))\n\n# Accuracy Plot\nplt.subplot(1, 3, 1)\nplt.plot(epochs, history_3.history['accuracy'], label='Train Accuracy')\nplt.plot(epochs, history_3.history['val_accuracy'], label='Val Accuracy')\nplt.xlabel('Epochs')\nplt.ylabel('Accuracy')\nplt.title('Accuracy Over Epochs')\nplt.legend()\nplt.grid(True)\n\n# AUC Plot\nplt.subplot(1, 3, 2)\nplt.plot(epochs, history_3.history['auc_6'], label='Train AUC')\nplt.plot(epochs, history_3.history['val_auc_6'], label='Val AUC')\nplt.xlabel('Epochs')\nplt.ylabel('AUC')\nplt.title('AUC Over Epochs')\nplt.legend()\nplt.grid(True)\n\n# Loss Plot\nplt.subplot(1, 3, 3)\nplt.plot(epochs, history_3.history['loss'], label='Train Loss')\nplt.plot(epochs, history_3.history['val_loss'], label='Val Loss')\nplt.xlabel('Epochs')\nplt.ylabel('Loss')\nplt.title('Loss Over Epochs')\nplt.legend()\nplt.grid(True)\n\nplt.tight_layout()\nplt.show()\n\nplt.subplot(1, 3, 1)\nplt.plot(epochs, history_3.history['accuracy'], label='Train Accuracy')\nplt.plot(epochs, history_3.history['val_accuracy'], label='Val Accuracy')\nplt.xlabel('Epochs')\nplt.ylabel('Accuracy')\nplt.title('Accuracy Over Epochs')\nplt.legend()\nplt.grid(True)\n\n# AUC Plot\nplt.subplot(1, 3, 2)\nplt.plot(epochs, history_3.history['auc_6'], label='Train AUC')\nplt.plot(epochs, history_3.history['val_auc_6'], label='Val AUC')\nplt.xlabel('Epochs')\nplt.ylabel('AUC')\nplt.title('AUC Over Epochs')\nplt.legend()\nplt.grid(True)\n\n# Loss Plot\nplt.subplot(1, 3, 3)\nplt.plot(epochs, history_3.history['loss'], label='Train Loss')\nplt.plot(epochs, history_3.history['val_loss'], label='Val Loss')\nplt.xlabel('Epochs')\nplt.ylabel('Loss')\nplt.title('Loss Over Epochs')\nplt.legend()\nplt.grid(True)\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T13:14:14.001201Z","iopub.execute_input":"2025-05-27T13:14:14.001558Z","iopub.status.idle":"2025-05-27T13:14:15.043599Z","shell.execute_reply.started":"2025-05-27T13:14:14.001531Z","shell.execute_reply":"2025-05-27T13:14:15.042801Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Predict validation set \nsamples_val = len(os.listdir(os.path.join(val_dir, '1'))) + len(os.listdir(os.path.join(val_dir, '0')))\nval_generator_for_pred = datagen_val.flow_from_directory(directory=val_dir, \n                                                batch_size=1, \n                                                class_mode='binary',\n                                               target_size=(96,96),\n                                                shuffle=False)\n\n\nval_predictions = mymodel_3.predict_generator(val_generator_for_pred,steps=samples_val, verbose=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T13:14:18.172771Z","iopub.execute_input":"2025-05-27T13:14:18.173098Z","iopub.status.idle":"2025-05-27T13:15:38.255255Z","shell.execute_reply.started":"2025-05-27T13:14:18.173066Z","shell.execute_reply":"2025-05-27T13:15:38.254515Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Actual vs. predicted classes for validation set\n\n# actual\nval_labels = val_generator_for_pred.classes\n\n# convert probability to classes using 0.5 shreshold\nbinary_val_predictions = (val_predictions > 0.5).astype(int)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T13:15:47.873433Z","iopub.execute_input":"2025-05-27T13:15:47.873758Z","iopub.status.idle":"2025-05-27T13:15:47.878284Z","shell.execute_reply.started":"2025-05-27T13:15:47.873733Z","shell.execute_reply":"2025-05-27T13:15:47.877423Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot ROC curve\n\n# Calculate AUC\nfpr, tpr, thresholds = roc_curve(val_labels, binary_val_predictions)\nroc_auc = auc(fpr, tpr)\n\n# Plot \nplt.figure(figsize=(4, 4))\nplt.plot(fpr, tpr, color='orange', lw=2, label='ROC curve (AUC = {:.2f})'.format(roc_auc))\nplt.plot([0, 1], [0, 1], color='navy', lw=2, linestyle='--', label='Random Guessing')\nplt.xlabel('FPR')\nplt.ylabel('TPR')\nplt.title('ROC Curve')\nplt.legend(loc='lower right')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T13:15:59.095839Z","iopub.execute_input":"2025-05-27T13:15:59.09652Z","iopub.status.idle":"2025-05-27T13:15:59.245694Z","shell.execute_reply.started":"2025-05-27T13:15:59.096491Z","shell.execute_reply":"2025-05-27T13:15:59.244759Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Get recall, predicision, F1, accuracy scores for validation set\nrecall = recall_score(val_labels, binary_val_predictions)\nprecision = precision_score(val_labels, binary_val_predictions)\nF1 = f1_score(val_labels, binary_val_predictions)\nacc= accuracy_score(val_labels, binary_val_predictions)\n\nprint(\"Recall on validation set:\", recall)\nprint(\"Precision on validation set:\", precision)\nprint(\"F1 score n on validation set:\", F1)\nprint(\"Accuracy score n on validation set:\", acc)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T13:16:02.064646Z","iopub.execute_input":"2025-05-27T13:16:02.06545Z","iopub.status.idle":"2025-05-27T13:16:02.087077Z","shell.execute_reply.started":"2025-05-27T13:16:02.065411Z","shell.execute_reply":"2025-05-27T13:16:02.086345Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot confusion matrix\ncm = confusion_matrix(val_labels, binary_val_predictions)\ndisp = ConfusionMatrixDisplay(confusion_matrix=cm, display_labels=['Non-cancer', 'Cancer'])\ndisp.plot(cmap='Blues', values_format='d')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T13:16:13.181957Z","iopub.execute_input":"2025-05-27T13:16:13.182587Z","iopub.status.idle":"2025-05-27T13:16:13.321278Z","shell.execute_reply.started":"2025-05-27T13:16:13.182557Z","shell.execute_reply":"2025-05-27T13:16:13.320409Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# Updated model names\nmodels = ['InceptionV3', 'Xception', 'DenseNet121', 'MobileNet', 'VGG16']\n\n# Corresponding accuracy scores from validation\naccuracies = [\n    0.9215,  # InceptionV3\n    0.9031,  # Xception\n    0.8795,  # DenseNet121\n    0.9253,  # MobileNet\n    0.9077   # VGG16\n]\n\n# Plotting accuracy\nplt.figure(figsize=(10, 6))\nbars = plt.bar(models, accuracies, color='skyblue')\n\n# Add labels\nfor bar, acc in zip(bars, accuracies):\n    plt.text(bar.get_x() + bar.get_width() / 2, bar.get_height() + 0.005, f'{acc:.4f}', \n             ha='center', va='bottom', fontsize=10)\n\nplt.ylim(0.85, 0.95)\nplt.title('Model Validation Accuracy Comparison', fontsize=16)\nplt.ylabel('Accuracy', fontsize=12)\nplt.xlabel('Model', fontsize=12)\nplt.grid(axis='y', linestyle='--', alpha=0.7)\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T13:20:00.729214Z","iopub.execute_input":"2025-05-27T13:20:00.729561Z","iopub.status.idle":"2025-05-27T13:20:00.90502Z","shell.execute_reply.started":"2025-05-27T13:20:00.729537Z","shell.execute_reply":"2025-05-27T13:20:00.904152Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# Models\nmodels = ['InceptionV3', 'Xception', 'DenseNet121', 'MobileNet', 'VGG16']\n\n# Recall, Precision, and F1-score values from provided metrics\nrecall_scores = [0.8705, 0.9299, 0.9438, 0.9222, 0.9892]\nprecision_scores = [0.9315, 0.8465, 0.7967, 0.8968, 0.4316]\nf1_scores = [0.8999, 0.8862, 0.8640, 0.9093, 0.6010]\n\n# Plot Recall\nplt.figure(figsize=(8, 5))\nplt.bar(models, recall_scores, color='skyblue')\nplt.title('Recall on Validation Set')\nplt.ylabel('Recall')\nplt.ylim(0.8, 1.0)\nplt.xticks(rotation=45)\nplt.grid(axis='y')\nplt.tight_layout()\nplt.show()\n\n# Plot Precision\nplt.figure(figsize=(8, 5))\nplt.bar(models, precision_scores, color='salmon')\nplt.title('Precision on Validation Set')\nplt.ylabel('Precision')\nplt.ylim(0.4, 1.0)\nplt.xticks(rotation=45)\nplt.grid(axis='y')\nplt.tight_layout()\nplt.show()\n\n# Plot F1-Score\nplt.figure(figsize=(8, 5))\nplt.bar(models, f1_scores, color='mediumseagreen')\nplt.title('F1 Score on Validation Set')\nplt.ylabel('F1 Score')\nplt.ylim(0.5, 1.0)\nplt.xticks(rotation=45)\nplt.grid(axis='y')\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-05-27T13:20:14.297283Z","iopub.execute_input":"2025-05-27T13:20:14.29762Z","iopub.status.idle":"2025-05-27T13:20:14.794978Z","shell.execute_reply.started":"2025-05-27T13:20:14.297594Z","shell.execute_reply":"2025-05-27T13:20:14.794116Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Methodology used","metadata":{}},{"cell_type":"markdown","source":"\"\"\"  Model 1 – InceptionV3 + Custom Conv Head\nMethodology:\nBase Model: Pretrained InceptionV3 (ImageNet weights) with include_top=False.\n\nLayer Freezing: All layers are frozen (feature extractor only).\n\nCut-off Layer: Model output taken from mixed2 layer (intermediate layer).\n\nArchitecture:\n\nSeveral custom Conv2D → BatchNorm → Dropout layers.\n\nFollowed by Flatten, multiple Dense layers with BatchNorm, and dropout.\n\nEnds with a Dense(1, activation='sigmoid') for binary classification.\n\nNo fine-tuning of base model.\n\nPurpose: Leverage early features from InceptionV3 and add a deep CNN head for learning.\n\n✅ Model 2 – Xception + GlobalAveragePooling Head\nMethodology:\nBase Model: Xception pretrained on ImageNet, include_top=False.\n\nFine-Tuning: Top 20 layers of the model are unfrozen and trained.\n\nCut-off Layer: block14_sepconv2_act (final convolutional activation layer).\n\nArchitecture:\n\nGlobalAveragePooling2D instead of flattening.\n\nTwo Dense layers with L2 regularization, BatchNorm, and Dropout.\n\nFinal sigmoid output for binary classification.\n\nPurpose: Lighter model with fewer parameters using GAP, leveraging deeper pretrained features.\n\n✅ Model 3 – DenseNet121 + Deep Dense Head\nMethodology:\nBase Model: DenseNet121 pretrained with include_top=False.\n\nFine-Tuning: Last 30 layers unfrozen for fine-tuning.\n\nCut-off Layer: conv5_block16_2_conv (last convolution block).\n\nArchitecture:\n\nGlobalAveragePooling2D after base.\n\n3 Dense layers (512, 256, 128) with L2, BatchNorm, and Dropout.\n\nFinal sigmoid classification layer.\n\nData Augmentation: Extensive real-time augmentation used.\n\nPurpose: Use DenseNet’s deep features with an extensive classifier on top.\n\n✅ Model 4 – MobileNetV2 + Lightweight Dense Head\nMethodology:\nBase Model: MobileNetV2 with include_top=False, pretrained on ImageNet.\n\nFine-Tuning: Last 20 layers unfrozen.\n\nCut-off Layer: Output layer of base model.\n\nArchitecture:\n\nGlobalAveragePooling2D\n\nTwo Dense layers with L2, BatchNorm, Dropout.\n\nFinal sigmoid output layer.\n\nData Augmentation: Moderate augmentation with both horizontal and vertical flips.\n\nPurpose: Lightweight model for faster training and inference with efficient mobile architecture.\n\n✅ Model 5 – VGG16 + GAP & Dense Classifier\nMethodology:\nBase Model: VGG16, pretrained, include_top=False.\n\nFine-Tuning: Last 10 layers unfrozen for training.\n\nCut-off Layer: Last layer of base model (block5).\n\nArchitecture:\n\nGlobalAveragePooling2D\n\nTwo Dense layers with L2, BatchNorm, and Dropout.\n\nFinal sigmoid output layer.\n\nPurpose: Simple architecture with strong early-layer feature maps and classic classifier head.\"\"\"\n\n","metadata":{}},{"cell_type":"markdown","source":"# Differences in Methodology Across the Models\nBase Architecture Used:\n\nEach model uses a different pretrained CNN:\nInceptionV3, Xception, DenseNet121, MobileNetV2, and VGG16, all pretrained on ImageNet.\n\nLayer Freezing & Fine-Tuning Strategy:\n\nModel 1 (InceptionV3): All layers frozen (no fine-tuning).\n\nModels 2–5: Partial fine-tuning by unfreezing the top N layers:\n\nXception: last 20 layers\n\nDenseNet121: last 30 layers\n\nMobileNetV2: last 20 layers\n\nVGG16: last 10 layers\n\nOutput Layer from Base Model:\n\nModel 1 taps into an intermediate layer (mixed2) for feature extraction.\n\nModels 2–5 use the final output of the base CNN or near-final layers for better abstracted features.\n\nType of Classifier Head:\n\nModel 1: Uses multiple Conv2D + BatchNorm + Dropout layers before flattening and dense layers (custom CNN head).\n\nModels 2–5: Use GlobalAveragePooling2D followed by stacked Dense layers (simpler, parameter-efficient).\n\nClassifier Depth and Regularization:\n\nModel 1: Deeper custom head, no L2 regularization.\n\nModels 2–5: Shallower or moderately deep heads with L2 regularization on dense layers to combat overfitting.\n\nData Augmentation Usage:\n\nModel 3 (DenseNet121) and Model 4 (MobileNetV2): Use aggressive or moderate real-time image augmentation.\n\nOthers: Not explicitly shown or lighter augmentation assumed.\n\nModel Compilation & Callbacks:\n\nAll models use Adam optimizer, binary_crossentropy loss, and AUC as a metric.\n\nAll incorporate ReduceLROnPlateau and EarlyStopping, but patience levels vary slightly.\n\nThese differences reflect trade-offs in model complexity, training time, generalization ability, and resource efficiency, with deeper heads like in Model 1 favoring learning capacity and lighter ones (Model 4) favoring speed and deployment feasibility.\n\n\n\n\n\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"# 🎓 Thesis Title\n\"Evaluating the Performance of Deep Learning Models on Imbalanced Datasets for Novel Disease Detection\"\n\n💡 Importance of the Thesis\nDetecting new or rare diseases presents a major challenge in medical AI due to the imbalance in available data. Typically, there are far fewer positive cases (patients with the disease) than negative ones (healthy or other conditions). This class imbalance causes most traditional models to favor the majority class, leading to false security and missed diagnoses.\n\nThis thesis aims to analyze the effectiveness of five state-of-the-art deep learning models in handling such imbalanced medical data, to ensure early and accurate detection of new diseases.\n\n📏 Why Accuracy Isn’t Enough\nIn imbalanced datasets, accuracy can be misleading. A model predicting only the majority class can still score high on accuracy. For this reason, other metrics are far more important:\n\nMetric\tImportance in Imbalanced Datasets\nPrecision\tProportion of predicted positive cases that are actually correct. High precision reduces false alarms.\nRecall\tProportion of actual positive cases that are correctly identified. Crucial for catching real cases.\nF1 Score\tHarmonic mean of precision and recall. A balanced metric for imbalanced datasets.\nAUC\tMeasures overall ability to distinguish between classes at all thresholds. Higher is bett","metadata":{}},{"cell_type":"markdown","source":"| Model           | Accuracy (%) | Precision | Recall    | F1 Score  | AUC        | Validation Loss |\n| --------------- | ------------ | --------- | --------- | --------- | ---------- | --------------- |\n| **InceptionV3** | 92.15        | 0.931     | 0.870     | 0.899     | 0.9729     | 0.2006          |\n| **Xception**    | 90.31        | 0.847     | 0.930     | 0.886     | 0.9711     | 0.2406          |\n| **DenseNet121** | 87.95        | 0.797     | 0.944     | 0.864     | 0.9664     | 0.3218          |\n| **MobileNetV2** | **92.53**    | 0.897     | 0.922     | **0.909** | **0.9759** | 0.2089          |\n| **VGG16**       | 46.69        | 0.432     | **0.989** | 0.601     | 0.9626     | 0.2477          |\n","metadata":{}},{"cell_type":"markdown","source":"#  Best Performing Model: MobileNetV2\n✅ Reasons for Superior Performance:\nHighest F1 Score (0.909): Balanced performance for both precision and recall.\n\nHigh Recall (0.922): Captures a large number of actual disease cases.\n\nGood Precision (0.897): Ensures most flagged cases are truly positive.\n\nHighest AUC (0.9759): Best overall discrimination ability.\n\nLow Validation Loss (0.2089): Indicates better generalization and less overfitting.\n\n🧱 Architectural Advantage:\nDepthwise Separable Convolutions: Efficiently reduce parameters and overfitting.\n\nLightweight and Fast: MobileNet is designed to perform well with fewer computations.\n\nTransfer Learning: Fine-tuning of the top 20 layers allows adaptation to disease data without losing general features.\n\n⚠ Model Limitations – Why Others Didn’t Win\nInceptionV3: Excellent precision, but lower recall than MobileNet — may miss some actual disease cases.\n\nXception: High recall but compromises on precision — more false positives.\n\nDenseNet121: Slight overfitting visible in higher validation loss.\n\nVGG16: Extremely high recall but terrible precision — predicts nearly everything as disease, causing too many false alarms.\n\n🔬 Final Takeaway\nThis thesis highlights that MobileNetV2 is best suited for detecting new diseases from imbalanced datasets due to its architecture and balanced metric performance. The findings stress that in sensitive applications like healthcare, metrics like F1 Score, Recall, and AUC are more important than mere accuracy.\n\n","metadata":{}},{"cell_type":"markdown","source":"# Results","metadata":{}},{"cell_type":"markdown","source":"#Inception V3\n=== Final Training & Validation Metrics ===\nTrain Loss:          0.2252\nTrain Accuracy:      0.9101\nTrain AUC:           0.9658\nValidation Loss:     0.2006\nValidation Accuracy: 0.9215\nValidation AUC:      0.9729\n\nRecall on validation set: 0.8704785441473826\nPrecision on validation set: 0.9314821492967905\nF1 score n on validation set: 0.8999477382265839\nAccuracy score n on validation set: 0.9214640594375313\n\nXception Model\nTrain Loss:          0.2260\nTrain Accuracy:      0.9119\nTrain AUC:           0.9678\nValidation Loss:     0.2406\nValidation Accuracy: 0.9031\nValidation AUC:      0.9711\nFinal Learning Rate: 0.000500\n\nRecall on validation set: 0.9299033924960683\nPrecision on validation set: 0.8465078228857756\nF1 score n on validation set: 0.8862480595257214\nAccuracy score n on validation set: 0.9031405260039199(Prediction accuracy)\n\nDenseNet121\n=== Final Training & Validation Metrics ===\nTrain Loss:          0.2705\nTrain Accuracy:      0.8990\nTrain AUC:           0.9582\nValidation Loss:     0.3218\nValidation Accuracy: 0.8795\nValidation AUC:      0.9664\n\nRecall on validation set: 0.9438328465513368\nPrecision on validation set: 0.7967001706808269\nF1 score n on validation set: 0.8640477169888934\nAccuracy score n on validation set: 0.8794840238844067\n\nMobileNet\n=== Final Training & Validation Metrics ===\nTrain Loss:          0.2256\nTrain Accuracy:      0.9161\nTrain AUC:           0.9692\nValidation Loss:     0.2089\nValidation Accuracy: 0.9253\nValidation AUC:      0.9759\n\nRecall on validation set: 0.9221523253201528\nPrecision on validation set: 0.8967664409001529\nF1 score n on validation set: 0.9092822330527249\nAccuracy score n on validation set: 0.9253384383973745\n\nVGG16 Model\n=== Final Training & Validation Metrics ===\nTrain Loss:          0.2289\nTrain Accuracy:      0.9142\nTrain AUC:           0.9667\nValidation Loss:     0.2477\nValidation Accuracy: 0.9077\nValidation AUC:      0.9626\n\nRecall on validation set: 0.9892159065378566\nPrecision on validation set: 0.4315820427367183\nF1 score n on validation set: 0.6009690848290453\nAccuracy score n on validation set: 0.4669766169834541\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"# <a id=\"13\"></a>\n\n## References (really helpful notebooks🙏🏽)\nhttps://www.kaggle.com/code/jieshends2020/baseline-keras-cnn-roc-fast-10min-0-925-lb/edit<br>\nhttps://www.kaggle.com/code/akarshu121/cancer-detection-with-cnn-for-beginners<br>\nhttps://www.kaggle.com/code/qitvision/a-complete-ml-pipeline-fast-ai<br>\nhttps://www.kaggle.com/code/CVxTz/cnn-starter-nasnet-mobile-0-9709-lb<br>\nhttps://www.kaggle.com/code/suicaokhoailang/wip-densenet121-baseline-with-fastai<br>\nhttps://www.kaggle.com/code/vbookshelf/cnn-how-to-use-160-000-images-without-crashing<br>\n","metadata":{}}]}