{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":11848,"databundleVersionId":862157,"sourceType":"competition"}],"dockerImageVersionId":30646,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<a id=\"0\"></a> <br>\n # Table of Contents  \n\n- Part 0:  [Project Introduction](#1)\n- Part 1:  [Exploratory Data Analysis (EDA)](#2)   \n  - 1.1  [Import Files, assisgn source directories, and Load Data](#3)\n  - 1.2  [Use partial of training data for EDA](#4)\n  - 1.3  [Abnormal images in all data ](#5)\n- Part 2:  [Model Building and training](#6)\n  - 2.1 [Copy all normal images to directories, with train test split](#7)\n  - 2.2 [Baseline model](#8)\n  - 2.3 [Model Tuning 1 - Add more epochs to increase training AUC](#9)\n  - 2.4 [Model Tuning 2 - Add Model Complexity Attempt 1(Use Transfer Learning)](#10)\n- Part 3: [Test Results](#11)\n- Part 3: [Project Conclusion](#12)\n- [References](#13)\n","metadata":{}},{"cell_type":"markdown","source":"# <a id=\"1\"></a>\n<div style=\"text-align: center; background-color: #AFDDCE; font-size:100%; padding: 5px;border-radius:10px 10px;\">\n    <h1>Part 0: Project Introduction </h1>\n</div>","metadata":{}},{"cell_type":"markdown","source":"### Dataset Overview\n- The dataset originates from a Kaggle project focused on classifying small pathology images. While the information on data collection methods is not clear, the [dataset description](https://www.kaggle.com/competitions/histopathologic-cancer-detection/data) provides valuable insights and clarification on duplicates. \n- The dataset comprises separate directories for train and test images. Ground truth labels are available in the \"train_labels.csv\" file, where a positive label signifies the presence of tumor tissue in the center 32x32px region.\n- Although the exact dataset size is undisclosed (we'll discover in EDA), it is mentioned that duplicates were removed. The dataset maintains the same splits as the PCam benchmark.\n\n### Project Goals and focus\n- Project goal: Identify metastatic tissue in histopathologic scans of lymph node sections in the unseen (testing) dataset.\n- Beyond Prediction: Gain insights and identify potential bottlenecks/challenges into the application of CNNs for processing and analyzing medical images in large volume. \n\n\n### Model Training Approach:\n- Initial Base Run: A preliminary model was trained to establish a baseline.\n- Tuning: Refinements were made, focusing on epoch numbers as the model was still learning.\n- Transfer Learning: Utilized InceptionV3 for improved performance.\n\n### A few highlights:\n- EDA (Exploratory Data Analysis): Due to the dataset's size, a subset (1%) was used for efficient EDA. Noteworthy patterns and abnormal images were identified and excluded from training.\n- Data Handling: Numpy arrays were employed for faster EDA, but image copies were later used for model training.\n- Model Training Steps: A systematic approach involved an initial base run, tuning parameters, and implementing transfer learning with InceptionV3.\n","metadata":{}},{"cell_type":"markdown","source":"\n","metadata":{}},{"cell_type":"markdown","source":"\n","metadata":{}},{"cell_type":"markdown","source":"# <a id=\"2\"></a>\n<div style=\"text-align: center; background-color: #AFDDCE; font-size:100%; padding: 5px;border-radius:10px 10px;\">\n    <h1>Part 1: Exploratory Data Analysis (EDA)</h1>\n</div>","metadata":{}},{"cell_type":"markdown","source":"# <a id=\"3\"></a>\n\n<div style=\"text-align: center; background-color: #81badd; font-size:70%; padding: 2px;border-radius:10px;\">\n    <h2>Part 1.1: Import Files, assisgn source directories, and Load Data</h2>\n</div>","metadata":{}},{"cell_type":"code","source":"# Import Libraries\n# Generaal\nimport numpy as np\nimport pandas as pd\nimport os\nimport shutil\nfrom glob import glob \nimport cv2\nimport gc \nimport time\nimport matplotlib.pyplot as plt\nimport matplotlib.patches as patches\nimport seaborn as sns\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import precision_recall_curve, f1_score, roc_curve,auc,recall_score, precision_score, accuracy_score, confusion_matrix,ConfusionMatrixDisplay\n\n\n%matplotlib inline\n\n\n# Tensorflow\nimport tensorflow as tf\nfrom tensorflow.keras.optimizers import Adam, RMSprop\nfrom tensorflow.keras.optimizers.schedules import ExponentialDecay\nfrom tensorflow.keras.metrics import AUC\nfrom tensorflow.keras.applications import InceptionV3, DenseNet121, NASNetMobile\nfrom tensorflow.keras import layers\nfrom tensorflow.keras import Model\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import BatchNormalization, Conv2D, Flatten, Dense, Dropout, MaxPooling2D\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tqdm import tqdm_notebook,trange\n","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:53:21.304469Z","iopub.execute_input":"2024-02-05T21:53:21.305007Z","iopub.status.idle":"2024-02-05T21:53:39.773812Z","shell.execute_reply.started":"2024-02-05T21:53:21.304966Z","shell.execute_reply":"2024-02-05T21:53:39.772472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Assign directory \nbase_dir = '../input/histopathologic-cancer-detection'  #path\nlabel_path = os.path.join(base_dir, 'train_labels.csv')\ntrain_path = os.path.join(base_dir,'train')\ntest_path = os.path.join(base_dir,'test')\nsample_path = os.path.join(base_dir,'sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:53:39.776276Z","iopub.execute_input":"2024-02-05T21:53:39.777175Z","iopub.status.idle":"2024-02-05T21:53:39.784219Z","shell.execute_reply.started":"2024-02-05T21:53:39.777127Z","shell.execute_reply":"2024-02-05T21:53:39.782853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check numbers of files under train and test\nprint('samples in training data: ',len(os.listdir(train_path)))\nprint('samples in testingdata: ',len(os.listdir(test_path)))","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:53:39.786235Z","iopub.execute_input":"2024-02-05T21:53:39.786710Z","iopub.status.idle":"2024-02-05T21:53:43.844208Z","shell.execute_reply.started":"2024-02-05T21:53:39.786667Z","shell.execute_reply":"2024-02-05T21:53:43.841987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# create a data frame to match training data image id and labels\ndf = pd.DataFrame({'path': glob(os.path.join(train_path,'*.tif'))}) # load the filenames\ndf['id'] = df.path.map(lambda x: x.split('/')[4].split(\".\")[0]) # get image id to merge with training labels\ntrain_labels = pd.read_csv(label_path)\ndf = df.merge(train_labels, on = \"id\") # merge labels and filepaths\ndf.head(2) ","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:53:43.847019Z","iopub.execute_input":"2024-02-05T21:53:43.847467Z","iopub.status.idle":"2024-02-05T21:53:45.906080Z","shell.execute_reply.started":"2024-02-05T21:53:43.847438Z","shell.execute_reply":"2024-02-05T21:53:45.905231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check if df is created correctly\n# check image counts\nprint('df shape',df.shape)\n# check null values\nprint('nulls :',df.isnull().sum().sum())\n# check dup values\nprint(\"dups: \",df.duplicated().sum())","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:53:45.907392Z","iopub.execute_input":"2024-02-05T21:53:45.907776Z","iopub.status.idle":"2024-02-05T21:53:46.155117Z","shell.execute_reply.started":"2024-02-05T21:53:45.907730Z","shell.execute_reply":"2024-02-05T21:53:46.153848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot the label count histogram\nprint('Positive labels in traing data: ', train_labels['label'].value_counts()[1] /  (train_labels['label'].value_counts()[0] + train_labels['label'].value_counts()[1])  )\nplt.figure(figsize=(5, 4))\nax = train_labels['label'].value_counts().sort_index().plot(kind='bar', color=['blue', 'orange'])\nplt.xticks([0, 1], labels=[f\"Negative N={train_labels['label'].value_counts()[0]}\", f\"Positive N={train_labels['label'].value_counts()[1]}\"], rotation=0)\nplt.title(f'Label Distributions in {train_labels.shape[0]} training data')\n","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:53:46.156506Z","iopub.execute_input":"2024-02-05T21:53:46.156994Z","iopub.status.idle":"2024-02-05T21:53:46.479201Z","shell.execute_reply.started":"2024-02-05T21:53:46.156952Z","shell.execute_reply":"2024-02-05T21:53:46.477962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <a id=\"4\"></a>\n\n<div style=\"text-align: center; background-color: #81badd; font-size:70%; padding: 2px;border-radius:10px;\">\n    <h2>Part 1.2: Use partial of training data for EDA</h2>\n</div>","metadata":{}},{"cell_type":"code","source":"# train_labels.head()\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:53:46.480861Z","iopub.execute_input":"2024-02-05T21:53:46.481241Z","iopub.status.idle":"2024-02-05T21:53:46.494663Z","shell.execute_reply.started":"2024-02-05T21:53:46.481208Z","shell.execute_reply":"2024-02-05T21:53:46.493334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Here for partial EDA data, I convert images to numpy array with unit8 types. It's faster.\n* Note creating numpy array with huge data takes too much memory, which would cause problems when reach the stage of training deeper CNN. Need to copy files to directory for training later.","metadata":{}},{"cell_type":"code","source":"# source https://www.kaggle.com/code/gomezp/complete-beginner-s-guide-eda-keras-lb-0-93. \n# This method is fast but consumes too much memory. casue later I cannot train deep CNN \n# function to load data.\n\ndef load_data(pct, df):\n    \"\"\"This function loads percent of df and return image path/name and labels.\"\"\"\n    # Get the number of images to load\n    n = int(pct * df.shape[0])\n    \n    # Randomly select n indices from the dataframe\n    random_indices =  np.random.choice(df.shape[0], size=n, replace=False)\n    \n    # Allocate a numpy array for the images (n, 96x96px, 3 channels, values 0 - 255)\n    X = np.zeros([n, 96, 96, 3], dtype=np.uint8)\n    # Convert the labels to a numpy array\n    y = np.squeeze(df['label'].values)[random_indices]\n\n    # Read images using randomly selected indices\n    for i, idx in enumerate(random_indices):\n        X[i] = cv2.imread(df.loc[idx, 'path'])\n\n    return n, X, y\n","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:53:46.505741Z","iopub.execute_input":"2024-02-05T21:53:46.506057Z","iopub.status.idle":"2024-02-05T21:53:46.519081Z","shell.execute_reply.started":"2024-02-05T21:53:46.506020Z","shell.execute_reply":"2024-02-05T21:53:46.517933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Slow to load data each time. Just use 1% training data for EDA seems enough.","metadata":{}},{"cell_type":"code","source":"# Load partical data\nPCT = 0.01\nn,sample_X,sample_y = load_data(pct = PCT, df=df)","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:53:46.525188Z","iopub.execute_input":"2024-02-05T21:53:46.525587Z","iopub.status.idle":"2024-02-05T21:54:09.588470Z","shell.execute_reply.started":"2024-02-05T21:53:46.525556Z","shell.execute_reply":"2024-02-05T21:54:09.587356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Take a look at image size and pixel info\nprint('Image shape: ',sample_X.shape)\nprint('Image maximum pixel value: ', sample_X.max())\nprint('Image minimum pixel value: ', sample_X.min())","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:54:09.589834Z","iopub.execute_input":"2024-02-05T21:54:09.590374Z","iopub.status.idle":"2024-02-05T21:54:09.610857Z","shell.execute_reply.started":"2024-02-05T21:54:09.590341Z","shell.execute_reply":"2024-02-05T21:54:09.609889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# source: https://www.kaggle.com/code/qitvision/a-complete-ml-pipeline-fast-ai\n# read image with flipping to rgb for visualization purpose\ndef readImage(path):\n    # OpenCV reads the image in bgr format by default\n    bgr_img = cv2.imread(path)\n    b,g,r = cv2.split(bgr_img)\n    rgb_img = cv2.merge([r,g,b])\n    return rgb_img","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:54:09.612058Z","iopub.execute_input":"2024-02-05T21:54:09.613010Z","iopub.status.idle":"2024-02-05T21:54:09.619137Z","shell.execute_reply.started":"2024-02-05T21:54:09.612976Z","shell.execute_reply":"2024-02-05T21:54:09.617715Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_images(index, num_image_each):\n    '''\n    num_image_each needs to be an integer number >= 2\n    '''\n    fig, ax = plt.subplots(1, num_image_each, figsize=(20, 4))\n\n    for i, idx in enumerate(index):\n        path = os.path.join(train_path, train_labels.iloc[idx]['id'])\n        ax[i].imshow(readImage(path + '.tif'))\n\n    ax[0].set_ylabel('samples', size='large')\n    \n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:54:09.620635Z","iopub.execute_input":"2024-02-05T21:54:09.620989Z","iopub.status.idle":"2024-02-05T21:54:09.635037Z","shell.execute_reply.started":"2024-02-05T21:54:09.620951Z","shell.execute_reply":"2024-02-05T21:54:09.633769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot some random negative and positive images  \n\n# number of positive or negative each\nnum_image_each = 8\n\n# Index of postive and negative samples\nindex_negative = np.random.choice(train_labels[train_labels['label']==0].index,num_image_each)\nindex_positive = np.random.choice(train_labels[train_labels['label']==1].index,num_image_each)\nprint(\"random negative index: \", index_negative)\nprint(\"random positive index: \", index_positive)\n\n# Plot some negatives images\nprint(\"\\nplot some negative images\")\nplot_images(index= index_negative, num_image_each=num_image_each)\n\n# Plot some positive images\nprint(\"\\nplot some positive images\")\nplot_images(index= index_positive, num_image_each=num_image_each)","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:54:09.636575Z","iopub.execute_input":"2024-02-05T21:54:09.636914Z","iopub.status.idle":"2024-02-05T21:54:12.109911Z","shell.execute_reply.started":"2024-02-05T21:54:09.636888Z","shell.execute_reply":"2024-02-05T21:54:12.108599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Context Knowledge:\nAccording to [Libre Pathology](https://librepathology.org/wiki/Lymph_node_metastasis), **these features might help us to identify tumor tissue** under microscopic:\n\n\n- +/-Cells with cytologic features of malignancy.\n  - Nuclear pleomorphism (variation in size, shape and staining).\n  - Nuclear atypia:\n    - Nuclear enlargement.\n    - Irregular nuclear membrane.\n    - Irregular chromatin pattern, esp. asymmetry.\n    - Large or irregular nucleolus.\n  - Abundant mitotic figures.\n\nWikipedia for [H&E stain ](https://en.wikipedia.org/wiki/H%26E_stain)","metadata":{}},{"cell_type":"markdown","source":"#### Draw some plots to see distributions of pixels in different color chanels for positive and negative examples","metadata":{}},{"cell_type":"code","source":"# seperate positive and negaive sample and get means\npos_samples = sample_X[sample_y == 1]\nneg_samples = sample_X[sample_y == 0]\n\n# means brightness of positive and negative samples\npos_sample_avgs=np.mean(pos_samples, axis=(1,2,3))\nneg_sample_avgs=np.mean(neg_samples, axis=(1,2,3))\n\n# max brightness of positive and negative samples\npos_sample_maxs=np.max(pos_samples, axis=(1,2,3))\nneg_sample_maxs=np.max(neg_samples, axis=(1,2,3))\n\n# min brightness of positive and negative samples\npos_sample_mins=np.min(pos_samples, axis=(1,2,3))\nneg_sample_mins=np.min(neg_samples, axis=(1,2,3))","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:54:12.111493Z","iopub.execute_input":"2024-02-05T21:54:12.112238Z","iopub.status.idle":"2024-02-05T21:54:12.219703Z","shell.execute_reply.started":"2024-02-05T21:54:12.112204Z","shell.execute_reply":"2024-02-05T21:54:12.218368Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot box plots for positive and negative sample chanel distributions\nchannels = [\"Red\", \"Green\", \"Blue\", \"RGB\"]\nsample_types = [\"Positive\", \"Negative\"]\nsamples = [pos_samples, neg_samples]\n\nfig, axs = plt.subplots(4, 2, sharey=True, figsize=(5, 8), dpi=150)\n\nfor i, channel in enumerate(channels):\n    for j, sample_type in enumerate(sample_types):\n        # Get the corresponding samples\n        current_samples = samples[j]\n\n        # Create filled box plot (vertically)\n        bp = axs[i, j].boxplot(current_samples[:, :, :, i % 3].flatten(), vert=True, patch_artist=True)\n        axs[i, j].set_title(f\"{sample_type} samples {channel}\")\n\n        # Set box color\n        colors = ['lightblue', 'lightcoral']\n        for patch, color in zip(bp['boxes'], colors):\n            patch.set_facecolor(color)\n\n        # Set y-axis label for the last column\n        if j == 1:\n            axs[i, j].set_ylabel(channel, rotation='horizontal', labelpad=35, fontsize=12)\n\n# Set common labels\nfor i in range(4):\n    axs[i, 0].set_ylabel(\"Channels\")\n    axs[i, 1].set_xlabel(\"Pixel value\")\n\n# Adjust layout\nfig.tight_layout()\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:54:12.221589Z","iopub.execute_input":"2024-02-05T21:54:12.224309Z","iopub.status.idle":"2024-02-05T21:54:15.362581Z","shell.execute_reply.started":"2024-02-05T21:54:12.224264Z","shell.execute_reply":"2024-02-05T21:54:15.361507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot histograms for positive and negative sample chanel distributions\nnr_of_bins = 256\nchannels = [\"Red\", \"Green\", \"Blue\", \"RGB\"]\nsample_types = [\"Positive\", \"Negative\"]\nsamples = [pos_samples, neg_samples]\n\nfig, axs = plt.subplots(4, 2, figsize=(15, 8))\n\nfor i, channel in enumerate(channels):\n    for j, sample_type in enumerate(sample_types):\n        # Get the corresponding samples\n        current_samples = samples[j]\n\n        # Create histogram\n        axs[i, j].hist(current_samples[:, :, :, i % 3].flatten(), bins=nr_of_bins, density=True)\n        axs[i, j].set_title(f\"{sample_type} samples {channel}\")\n\n        # Set y-axis label for the last column\n        if j == 1:\n            axs[i, j].set_ylabel(channel, rotation='horizontal', labelpad=35, fontsize=12)\n\n# Set common labels\nfor i in range(4):\n    axs[i, 0].set_ylabel(\"Relative frequency\")\n    axs[i, 1].set_xlabel(\"Pixel value\")\n\n# Adjust layout\nfig.tight_layout()\n\n# Show the plot\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:54:15.363795Z","iopub.execute_input":"2024-02-05T21:54:15.364363Z","iopub.status.idle":"2024-02-05T21:54:22.090794Z","shell.execute_reply.started":"2024-02-05T21:54:15.364331Z","shell.execute_reply":"2024-02-05T21:54:22.089658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot mean,min,max brightness histograms for positive and negative samples\nfig,axs = plt.subplots(3,2,figsize=(10, 10))\naxs[0,0].hist(pos_sample_avgs,bins= 60,density=True);\naxs[0,1].hist(neg_sample_avgs,bins= 60,density=True);\naxs[0,0].set_title(\"Mean brightness, positive samples\");\naxs[0,1].set_title(\"Mean brightness, negative samples\");\naxs[0,0].set_ylabel(\"Relative frequency\")\n\naxs[1,0].hist(pos_sample_mins,bins= 60,density=True);\naxs[1,1].hist(neg_sample_mins,bins= 60,density=True);\naxs[1,0].set_title(\"Min brightness, positive samples\");\naxs[1,1].set_title(\"Min brightness, negative samples\");\naxs[1,0].set_ylabel(\"Relative frequency\")\n\naxs[2,0].hist(pos_sample_maxs,bins= 60,density=True);\naxs[2,1].hist(neg_sample_maxs,bins= 60,density=True);\naxs[2,0].set_title(\"Max brightness, positive samples\");\naxs[2,1].set_title(\"Max brightness, negative samples\");\naxs[2,0].set_ylabel(\"Relative frequency\")\n","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:54:22.092368Z","iopub.execute_input":"2024-02-05T21:54:22.092712Z","iopub.status.idle":"2024-02-05T21:54:24.430168Z","shell.execute_reply.started":"2024-02-05T21:54:22.092682Z","shell.execute_reply":"2024-02-05T21:54:24.428976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### Note the threshold below I found after image inspection. Since those codes are comment out, I just want to show some abnormal pictures here is samples.","metadata":{}},{"cell_type":"code","source":"# Calculate brightness for each image. \nbrightest_values = np.max(sample_X, axis=(1, 2, 3))\ndarkest_values = np.min(sample_X, axis=(1, 2, 3))\n\n# Find indices of images with abnormal brightness\n# Note the threshold I found after later abnormal image inspectaion on all images.\nabnormal_brightness_indices_sx = np.where(\n    (brightest_values < 75) | (darkest_values  >  180)\n)[0]\n\n\nprint( 'abnormal_indices in sample X: ',abnormal_brightness_indices_sx)\nprint( 'number of abnormal_indices in sample X: ',abnormal_brightness_indices_sx.shape[0])","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:54:24.432216Z","iopub.execute_input":"2024-02-05T21:54:24.432657Z","iopub.status.idle":"2024-02-05T21:54:24.456353Z","shell.execute_reply.started":"2024-02-05T21:54:24.432620Z","shell.execute_reply":"2024-02-05T21:54:24.454860Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check brightest and darkest pixel of abnormal images in samples\nfor index in abnormal_brightness_indices_sx: \n    print(f'Image {index} brightest pixel value is {np.max(sample_X[index,:,:,:])}, darkest pixel value is {np.min(sample_X[index,:,:,:])}')\n   \n# Plot abnormal images in samples\nfig, ax = plt.subplots(1, len(abnormal_brightness_indices_sx), figsize=(2.5 * len(abnormal_brightness_indices_sx) , 2))\n\nfor i, idx in enumerate(abnormal_brightness_indices_sx):\n    ax[i].imshow(sample_X[idx,:,:,:])\n    ax[0].set_ylabel('abnormals in samples', size='large')\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:54:24.458276Z","iopub.execute_input":"2024-02-05T21:54:24.458779Z","iopub.status.idle":"2024-02-05T21:54:25.452614Z","shell.execute_reply.started":"2024-02-05T21:54:24.458710Z","shell.execute_reply":"2024-02-05T21:54:25.451294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color: #FFF2CC; padding: 10px;\">\n<span style=\"font-size: larger;\">\n\n### What I've found during EDA:\n    \n- There are **220,025 samples** in training data and **57,458** samples in testing data. \n- **Train/Test split: 80%/20%**.\n- There are **130,908 negative and 89,117 postitive labels.  Positive lable %: 41%**.\n- **Images has shape (96, 96, 3)**\n- **Pixal values range from 0 to 255**. \n- **Negative samples seems to have more brighter (200-250 higer value) pixels in red, green, blue, and RGB mixed chanels**.\n- **There are a few abnormal images**. Some image's smallest pixel values seems high, meaning they are too bright. Also, maybe a few images are too dark.\n    \n### Few Data preprocessing/clearning idea based on EDA:\n- For now, I don't plan to do further resampling to balance data. 40% positive samples using AUC seems okay. We are required to use AUC for this project. AUC isn't too sensitive to imbalanced dataset unless data is extremely imbalanced.\n- Normalize data using /.255. I will do this in image augmentation though.\n- Do some further inspectation on abnormal images. Then, exclude those abnormal images when copy all data to directories for training. \n    ","metadata":{}},{"cell_type":"markdown","source":"# <a id=\"5\"></a>\n\n<div style=\"text-align: center; background-color: #81badd; font-size:70%; padding: 2px;border-radius:10px;\">\n    <h2>Part 1.3 Abnormal images in all data </h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"### Abnormal brightness image inspection\n\n* Note I commented codes below out because those takes lots of RAM to include in final run.","metadata":{}},{"cell_type":"code","source":"# Load all data\n# st =time.time()\n# PCT = 1\n# n,X,y = load_data(pct = PCT, df=df)\n# print(time.time() - st)","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:54:25.454401Z","iopub.execute_input":"2024-02-05T21:54:25.455174Z","iopub.status.idle":"2024-02-05T21:54:25.460469Z","shell.execute_reply.started":"2024-02-05T21:54:25.455127Z","shell.execute_reply":"2024-02-05T21:54:25.459272Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get and save abnormal indices\n# Calculate brightness for each image\n# brightest_values = np.max(X, axis=(1, 2, 3))\n# darkest_values = np.min(X, axis=(1, 2, 3))\n\n# # Find indices of images with abnormal brightness\n# abnormal_brightness_indices = np.where(\n#     (brightest_values < 75) | (darkest_values  >  180)\n# )[0]\n\n# print('Labels:', y[abnormal_brightness_indices])\n# print( 'number_of_abnormal_indices: ',abnormal_brightness_indices.shape[0])\n# np.save('/kaggle/working/abnormal_brightness_indices.npy', abnormal_brightness_indices)","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:54:25.462210Z","iopub.execute_input":"2024-02-05T21:54:25.462632Z","iopub.status.idle":"2024-02-05T21:54:25.474528Z","shell.execute_reply.started":"2024-02-05T21:54:25.462584Z","shell.execute_reply":"2024-02-05T21:54:25.473060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# indices_string = f'abnormal_brightness_indices = {abnormal_brightness_indices.tolist()}'\n# print(indices_string)","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:54:25.476246Z","iopub.execute_input":"2024-02-05T21:54:25.476710Z","iopub.status.idle":"2024-02-05T21:54:25.487328Z","shell.execute_reply.started":"2024-02-05T21:54:25.476666Z","shell.execute_reply":"2024-02-05T21:54:25.485906Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get abonormal indices (hard-coded to avoid re-run unecessary codes)\n\nabnormal_brightness_indices = [563, 1184, 1501, 2000, 2817, 3872, 4067, 4079, 4204, 4382, 4660, 5190, 5636, 6520, 6907, 7012, 7063, 7173, 8134, 8325, 8406, 8525, 9010, 9307, 9390, 9415, 9764, 9832, 10200, 10337, 11326, 11915, 12771, 12965, 13724, 13781, 13801, 13945, 14335, 14742, 16689, 16899, 17809, 18386, 18613, 18743, 18810, 18842, 19022, 19336, 19456, 19579, 20019, 20243, 20294, 20687, 21713, 21844, 22923, 23084, 23135, 23263, 23306, 23396, 23429, 23962, 25503, 25930, 25966, 26266, 26758, 26771, 27134, 27188, 27888, 28005, 29274, 30035, 30323, 30481, 30754, 31352, 32016, 32212, 32230, 33071, 33392, 33946, 34005, 34921, 35086, 35763, 35875, 36062, 36452, 36614, 36760, 36881, 37404, 37613, 37645, 37842, 38089, 38117, 38127, 38387, 38412, 38414, 39151, 39191, 40038, 40084, 40884, 41171, 41442, 41749, 41753, 42940, 43235, 43236, 43370, 43526, 43561, 43689, 43818, 44236, 45707, 46058, 47419, 48056, 48270, 49635, 49889, 50295, 50638, 51614, 52251, 53086, 53427, 53976, 54446, 54638, 54699, 54990, 55000, 55351, 55515, 56591, 56675, 57504, 57552, 57674, 58043, 58090, 58625, 58835, 59690, 59823, 61334, 61379, 61784, 61857, 62393, 62534, 62926, 62935, 63407, 63672, 65585, 66062, 66656, 66741, 66927, 67452, 67502, 67903, 68159, 68220, 68583, 68607, 68709, 69071, 71056, 72897, 72979, 73372, 73666, 73822, 74334, 74386, 74536, 74573, 74953, 75278, 75379, 75745, 75963, 76168, 76746, 76841, 77341, 78310, 78668, 78798, 79156, 79894, 80524, 81094, 81098, 81320, 81358, 81408, 82393, 82518, 82852, 82945, 83303, 83354, 83393, 83479, 83559, 83651, 84554, 84670, 84843, 85021, 85154, 85329, 85503, 86727, 87077, 87873, 87961, 88128, 88459, 88581, 89759, 89969, 90442, 90585, 91986, 92803, 93235, 93709, 94200, 94521, 95061, 95276, 95387, 95468, 95530, 95583, 95713, 95749, 96035, 96640, 96929, 98927, 99296, 99298, 99458, 99806, 99927, 100380, 100402, 100605, 100872, 101543, 102082, 102620, 102621, 102712, 102983, 104141, 104192, 104290, 104634, 104948, 105141, 105326, 106057, 106439, 106867, 106929, 107026, 107212, 107304, 107776, 107990, 108110, 108674, 108920, 109101, 109393, 110252, 110412, 111151, 112387, 112620, 112621, 113174, 113192, 113553, 114720, 114790, 114829, 115397, 115582, 116799, 117191, 117287, 117913, 118377, 118507, 118610, 118776, 118841, 119605, 119616, 119789, 120325, 120532, 121005, 121226, 121275, 121380, 121631, 122261, 122391, 122864, 123934, 124240, 124444, 124626, 125824, 126393, 126487, 126969, 127079, 127196, 127335, 127460, 128059, 128263, 128451, 128973, 129338, 129416, 129806, 130123, 130194, 130257, 130390, 130754, 131060, 131557, 131992, 132402, 132404, 132812, 133154, 133211, 133297, 134614, 134666, 135599, 135615, 135963, 136110, 136477, 137134, 137851, 137963, 138273, 138381, 138484, 138859, 139331, 139744, 139817, 139920, 140298, 140399, 141505, 142139, 142468, 142889, 143209, 143309, 143518, 144010, 144265, 144534, 145298, 145955, 146020, 146522, 147783, 147952, 148527, 148545, 149393, 150294, 150759, 151089, 152243, 152549, 152837, 153256, 153651, 154094, 154133, 154238, 154270, 154316, 154729, 155536, 155865, 156253, 156422, 156512, 157071, 157228, 157581, 158455, 158705, 158929, 159130, 159801, 160097, 160480, 160588, 161045, 161076, 161195, 161303, 161308, 161379, 161443, 162186, 162323, 162487, 162796, 162797, 163116, 163449, 163932, 164002, 164046, 164052, 164077, 164803, 164944, 165394, 166311, 166593, 167008, 167058, 167767, 167825, 167864, 167871, 168099, 168270, 168456, 168649, 168791, 168916, 168965, 169061, 169068, 169210, 169227, 169355, 169583, 169903, 169976, 169981, 171091, 171456, 171564, 171975, 172152, 172427, 172437, 172764, 172930, 173665, 174007, 176035, 176314, 176474, 176507, 176653, 177333, 177629, 177955, 178367, 178552, 179100, 179995, 180425, 180441, 180761, 181059, 181208, 181280, 181376, 181591, 181871, 181964, 182028, 182140, 182439, 182459, 182777, 182788, 182858, 183068, 183462, 183901, 185126, 185474, 185743, 185935, 186136, 186217, 186687, 187048, 187650, 188092, 188603, 188964, 188981, 189406, 189854, 190096, 190141, 190482, 190589, 190899, 190913, 191313, 191407, 191684, 191902, 191908, 192341, 192456, 192477, 192569, 193074, 193194, 193618, 193623, 193840, 193844, 193879, 194220, 194275, 195125, 196150, 196775, 197231, 197234, 198170, 198226, 198325, 198356, 199282, 199308, 199634, 200032, 200286, 201505, 201625, 201990, 202078, 202103, 202167, 203077, 203493, 203656, 203741, 204340, 204405, 204624, 205096, 205142, 205274, 205585, 205687, 205721, 206409, 207278, 207917, 208118, 208354, 208505, 209273, 209725, 210278, 210915, 211026, 211056, 211398, 211546, 211585, 212253, 212672, 212872, 212923, 213383, 213575, 213581, 213944, 214139, 214809, 214824, 215062, 215389, 216153, 216309, 216557, 216759, 217036, 217188, 217258, 217345, 217380, 217563, 217685, 217993, 218102, 218511, 219449, 219534, 219598]\n\n","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:54:25.489554Z","iopub.execute_input":"2024-02-05T21:54:25.490089Z","iopub.status.idle":"2024-02-05T21:54:25.532077Z","shell.execute_reply.started":"2024-02-05T21:54:25.490034Z","shell.execute_reply":"2024-02-05T21:54:25.530544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#  Plot some abnormal images\n# fig, ax = plt.subplots(1, 10, figsize=(30, 4))\n\n# for i, idx in enumerate(abnormal_brightness_indices[11:21]):\n#     ax[i].imshow(X[idx, :, :, :])\n#     ax[0].set_ylabel('samples', size='large')\n# plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:54:25.534078Z","iopub.execute_input":"2024-02-05T21:54:25.534534Z","iopub.status.idle":"2024-02-05T21:54:25.548692Z","shell.execute_reply.started":"2024-02-05T21:54:25.534490Z","shell.execute_reply":"2024-02-05T21:54:25.547276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### During the inspectations above, I found that if the brightest pixel is less than 75 or the darkest pixel is greater than 180, the images are basically blank. We should remove them.\n\n","metadata":{}},{"cell_type":"markdown","source":"# <a id=\"6\"></a>\n<div style=\"text-align: center; background-color: #AFDDCE; font-size:100%; padding: 5px;border-radius:10px 10px;\">\n    <h1>Part 2: Model Building and training</h1>\n</div>","metadata":{}},{"cell_type":"markdown","source":"# <a id=\"7\"></a>\n\n<div style=\"text-align: center; background-color: #81badd; font-size:70%; padding: 2px;border-radius:10px;\">\n    <h2>Part 2.1: Copy all normal images to directories, with train test split</h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"#### Now let's load all (but only normal data) in. Copy files to directories takes a longer time (compared to copy to int8 numpy array)but can save more memory for later model training.\n\n","metadata":{}},{"cell_type":"code","source":"# drop abnormal images indices in df\ndf.drop(abnormal_brightness_indices, inplace=True)\nprint('df size after removing abnormal images: ', df.shape[0])","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:54:25.550371Z","iopub.execute_input":"2024-02-05T21:54:25.550872Z","iopub.status.idle":"2024-02-05T21:54:25.586398Z","shell.execute_reply.started":"2024-02-05T21:54:25.550837Z","shell.execute_reply":"2024-02-05T21:54:25.585142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# train test split\ndf_train, df_val = train_test_split(df, test_size=0.10, random_state=100)","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:54:25.587826Z","iopub.execute_input":"2024-02-05T21:54:25.588177Z","iopub.status.idle":"2024-02-05T21:54:25.641682Z","shell.execute_reply.started":"2024-02-05T21:54:25.588148Z","shell.execute_reply":"2024-02-05T21:54:25.640462Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# reset index using id\ndf.set_index('id', inplace=True)\n","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:54:25.650541Z","iopub.execute_input":"2024-02-05T21:54:25.650980Z","iopub.status.idle":"2024-02-05T21:54:25.657348Z","shell.execute_reply.started":"2024-02-05T21:54:25.650943Z","shell.execute_reply":"2024-02-05T21:54:25.655985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create directories for model augmentation\nwork_dir ='/kaggle/working/data'\ntrain_dir =os.path.join(work_dir,'train')\nval_dir=os.path.join(work_dir,'val')\n\nif os.path.exists(work_dir):# remove the entire directory and its contents if already exist\n    shutil.rmtree(work_dir)\n\nfor fold in [train_dir, val_dir]:\n    for subf in [\"0\", \"1\"]:\n        os.makedirs(os.path.join(fold, subf)) # this will create necessary parent directories as well :)       ","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:54:25.658639Z","iopub.execute_input":"2024-02-05T21:54:25.659213Z","iopub.status.idle":"2024-02-05T21:54:25.672172Z","shell.execute_reply.started":"2024-02-05T21:54:25.659174Z","shell.execute_reply":"2024-02-05T21:54:25.670652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Copy all files to train and validation directies\n \n# copy train files \nfor id in df_train['id'].values:\n    fn =id +'.tif'\n    label = str(df.loc[id,'label']) \n    source_dir_fn=os.path.join(train_path, fn)\n    train_dir_fn=os.path.join(train_dir,label,fn)\n    shutil.copyfile(source_dir_fn,train_dir_fn )\n    \n# copy validation files\nfor id in df_val['id'].values:\n    fn =id +'.tif'\n    label = str(df.loc[id,'label']) \n    source_dir_fn =os.path.join(train_path, fn)\n    val_dir_fn =os.path.join(val_dir,label,fn)\n    shutil.copyfile(source_dir_fn,val_dir_fn )","metadata":{"execution":{"iopub.status.busy":"2024-02-05T21:54:25.673572Z","iopub.execute_input":"2024-02-05T21:54:25.673963Z","iopub.status.idle":"2024-02-05T22:29:18.297586Z","shell.execute_reply.started":"2024-02-05T21:54:25.673933Z","shell.execute_reply":"2024-02-05T22:29:18.293814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check image counts in Training and validation directories\n\n#Training\nprint('Positive images in training directory: ', len(os.listdir(os.path.join(train_dir, '1'))))\nprint('Negative images in training directory: ',len(os.listdir(os.path.join(train_dir, '0'))))\nprint('% Positive images in training directory: ',len(os.listdir(os.path.join(train_dir, '1'))) / (len(os.listdir(os.path.join(train_dir, '0'))) + len(os.listdir(os.path.join(train_dir, '1'))) ),'\\n')\n\n# Validation\nprint('Positive images in validation directory: ', len(os.listdir(os.path.join(val_dir, '1'))))\nprint('Negative images in validation directory: ',len(os.listdir(os.path.join(val_dir, '0'))))\nprint('% Positive images in validation directory: ',len(os.listdir(os.path.join(val_dir, '1'))) / (len(os.listdir(os.path.join(val_dir, '0'))) + len(os.listdir(os.path.join(val_dir, '1'))) ), '\\n')\n\n# Training vs. Validation\nprint('Images in Training directory: ', len(os.listdir(os.path.join(train_dir, '1'))) + len(os.listdir(os.path.join(train_dir, '0'))) )\nprint('Images in  validation directory: ', len(os.listdir(os.path.join(val_dir, '1'))) + len(os.listdir(os.path.join(val_dir, '0'))) )\nprint('% Training images: ', (len(os.listdir(os.path.join(train_dir, '1'))) + len(os.listdir(os.path.join(train_dir, '0')))) / (len(os.listdir(os.path.join(train_dir, '1'))) + len(os.listdir(os.path.join(train_dir, '0')))+len(os.listdir(os.path.join(val_dir, '1'))) + len(os.listdir(os.path.join(val_dir, '0'))) )    )\nprint('\\nTotal images: ',  (len(os.listdir(os.path.join(train_dir, '1'))) + len(os.listdir(os.path.join(train_dir, '0')))+len(os.listdir(os.path.join(val_dir, '1'))) + len(os.listdir(os.path.join(val_dir, '0'))) )    )\n","metadata":{"execution":{"iopub.status.busy":"2024-02-05T22:29:18.302609Z","iopub.execute_input":"2024-02-05T22:29:18.303504Z","iopub.status.idle":"2024-02-05T22:29:19.531557Z","shell.execute_reply.started":"2024-02-05T22:29:18.303427Z","shell.execute_reply":"2024-02-05T22:29:19.530144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Collect garbage to free up memory\ndel pos_samples \ndel neg_samples \ndel pos_sample_avgs \ndel neg_sample_avgs \ndel pos_sample_mins\ndel neg_sample_mins \ndel pos_sample_maxs \ndel neg_sample_maxs\ndel df\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2024-02-05T22:29:19.532950Z","iopub.execute_input":"2024-02-05T22:29:19.533346Z","iopub.status.idle":"2024-02-05T22:29:20.381910Z","shell.execute_reply.started":"2024-02-05T22:29:19.533313Z","shell.execute_reply":"2024-02-05T22:29:20.380541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <a id=\"8\"></a>\n\n<div style=\"text-align: center; background-color: #81badd; font-size:70%; padding: 2px;border-radius:10px;\">\n    <h2>Part 2.2: Baseline model</h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"### First I will just try to build a baseline model that I think might perform well. A few helpful knowledges. 👇\n","metadata":{}},{"cell_type":"markdown","source":"\n<div style=\"background-color: #FCFFE9; padding: 10px;\">\n    \n### Why CNN for images?📝\n  - **Problem caused by the big images data** (Even more and more pixels nowadays!)\n    - **Computation capacity**: consider the weights we need for flattening out all the pixels and use them as the first layer of a fully connected neural network!\n    - **Overfit**: Fully-connected neural network would have too many parameters. Lots of weights from empty spaces in the image aren't useful and they are redundant.\n  - **Translational Invariance in images**\n    - Fully connected layer would think an object shift to a different position is a different object beacause it learns the whole image.\n    - CNN uses small and different sized filters to learn/detect different primitive features (Ex: an ear, an eye, intensity of color, texture, shape, orientation)in an image. **What a filter learn in one position of an image can be re-used in other positions.** CNN slices those filters across the image and the **learning weights are shared** in different positions. This results in way less parameters/overfitting/waste of computational resources\n\n### Why multiple convolution layers? 📝\nThere is a **field of view (FOV)** effect of CNN. Here is an example:\n- 1st convolution filter of size 3 * 3 would look at the input image 3 * 3 at a time and assign it to 1 output pixal. That is, each size 1 pixal in the 1st convolution feature map gets it's information from an area of 3 * 3 in the original input image. 3 * 3 isn't big area, so the 1st filter can only detect local features like lines etc.\n- 2nd convolution filter of size 3 * 3 would look at the 1st convolution feature map image 3 * 3 at a time and assign it to 1 output pixal. That is, each size 1 pixal in the 2nd convolution feature map gets it's information from an area of size 3 * 3 in the 1st convolution feature map image. That is equivalent to look at an area of 9 * 9 from the original image ( (3 * 3) * (3 * 3)) based on layer 1. So, 2nd convolution layers can see more global features.\n- The FOV becomes larger when CNN gets more layers so CNN would see more global feature in later layers. For example:\n  - Layer 1 may detect features like edges and lines, color contrasts, textures, etc.\n  - Layer 2 may detect simple shapes and patterns formed by the combination of features from the first layer. like curves,combinations of edges, corners, and textures.\n  - Later layers may detect more global features like eyes, face, cars, etc. <br>\n \n(More details of FOV in [this paper](https://arxiv.org/abs/1311.2901))\n    \n ### Why prefer small filters? 📝\n\nNowadays small filters (like 3 * 3 )are more popular:\n\n- **Computation effeciency**: for instance, two 3 * 3 filters and one 5 * 5 filter both get you (N - 4) * (N -4) feature map. But the 3 * 3 filters only need 18 (=3 * 3 * 2) but 5 * 5 filter need 25 (5 * 5) parameters. \n- Deal with **Translational Invariance in images** better\n\n ### Why use 1 * 1 Filters? 📝\n    \n-  1 * 1 filters mimic applying a fully-connected layer to each pixals. This adds non-linearity to let layer learn more complex function.(Details in [this video](https://www.youtube.com/watch?v=c1RBQzKsDCk&list=PLpFsSf5Dm-pd5d3rjNtIXUHT-v7bdaEIe&index=115)). To illustrate: \n   - If we have a 6 * 6 * 32 (depth) input layer. At each slice step, a 1 * 1 filter would take 32 numbers (depth), multiply each 32 numbers by 32 weights, apply activation function (maybe Relu), then output 1 output units. This is like creating one neuron of a fully connected layer .\n   - So, each 1 * 1 filter would output a 6 * 6 * 32 input layer to 6 * 6 feature map. With more filters, the output would be 6 * 6 * number of filters. Thus, adding more filters is like creating more output neuron units.\n- 1 * 1 filters allow to shrink or add number of channels. Can be used for computation efficiency.  \n    \n### Why use Pooling layer? 📝\n\n    \n- Pooling layer in CNN uses subsampling method to **down-sample the spatial dimensions of the input feature maps**. \n- A **pooling filter of size p would usualy reduce the input image size by 1/p** (without considering the input image depth, which goes to 1 for each filter). Explaination:\n    - Output_image_size Output_image_size = (N + 2 * p - f)/s +1 (N: input image size; p: padding size; f: filter size; s: stride size)\n    - Pooling filter typically don't have padding. (p=0)\n    - The pooling window is usually applied to non-overlapping regions of the input. (So,usually s=f.)\n    - To illustrate, if we have a 10 * 10 image and a 2*2 pooling filter, the output size would be (10-2)/2+1 = 5. \n- Pooling filter sizes are **usually 2 * 2** or some **even size** (while **convolution filters are often odd size**), so usually it reduce half of the image size  \n- Two popular types:\n   - **Max pooling**: the maximum is taken within each region. For example, if the input image has a deepth of 3 and we have a 2*2 max pooling filter. The output pixal for each step would be the maximum of the 12 numbers (12=2 * 2 * 3). It emphasize the most dominant features in local regions and is used most often.\n   - **Average pooling**: the average is taken within each region. It provides a smoother down-sampling effect.\n\n\n    \n- Why using pooling layer? - It shrinks down the feature map size for a few reasons:\n  - **Computational efficiency**\n  - Helps to see more **global features**\n  - **Feature generalization**: retain most important information and ignore the details. The layer would be less sensitive to small shifts in the input.\n","metadata":{}},{"cell_type":"markdown","source":"### My Plan for training the model\n- Reduce Training Error:\n  - First focus on reducing training error, which means it's comprehensive and complex enough that it fits training data well. \n- Reduce Validation Error:\n  - Once satisfied with the training performance, introduce overfitting prevention techniques:\n    - Dropout\n    - L2 or L1 normalization\n    - Batch normalization (might just add it earlier to reduce training error)\n    - Data Augmentation\n- Attempt of Transfer Learning:\n  - After just a couple of attempts, if the model is still too simple, explore transfer learning with pre-trained models.\n  - Include structures for overfitting from the start\n\nNote: I've decided to withhold result submission until I'm thoroughly satisfied with my validation score (probably not until the last run😅). Given that the testing score is unlikely to significantly surpass the validation one, I'm prioritizing efficient use of resources. However, I will meticulously analyze each model's performance, conducting a detailed comparison of training and validation metrics. This analysis is crucial for guiding the next steps in training direction.\n","metadata":{}},{"cell_type":"markdown","source":"## Model 1 \n* note I comment out the code and left with train/validation AUC and loss plots. I don't want to run those too many times to save time. ","metadata":{}},{"cell_type":"code","source":"# Build model 1\n# Note for model 1 to model 3, I used copied numpy data. I later found this method consumes too much memory.\n# mymodel_1 = tf.keras.models.Sequential([\n#    Conv2D(32, (3, 3), activation='relu', input_shape=(96, 96, 3)),\n#    Conv2D(32, (3, 3), activation='relu'),\n#    MaxPooling2D((2, 2)),\n#    Conv2D(64, (3, 3), activation='relu'),\n#    Conv2D(64, (3, 3), activation='relu'),\n#    MaxPooling2D((2, 2)),\n#    Conv2D(128, (3, 3), activation='relu'),\n#    Conv2D(128, (3, 3), activation='relu'),\n#    MaxPooling2D((2, 2)),\n#    Conv2D(128, (1, 1), activation='relu'),\n#    Flatten(),\n#    Dense(128, activation='relu'),\n#    Dense(128, activation='relu'),\n#    Dense(1, activation='sigmoid')\n# ])\n\n# Model summary\n#mymodel_1.summary()","metadata":{"execution":{"iopub.status.busy":"2024-02-05T22:29:20.383589Z","iopub.execute_input":"2024-02-05T22:29:20.384082Z","iopub.status.idle":"2024-02-05T22:29:20.392729Z","shell.execute_reply.started":"2024-02-05T22:29:20.384040Z","shell.execute_reply":"2024-02-05T22:29:20.391564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# compile model 1\n\n# Define the learning rate schedule (exponential decay)\n# initial_learning_rate = 0.001\n# decay_steps = 1000 # total steps = (training size/batchsize) * number of epochs, 200,000/64*3= around 10000\n# decay_rate = 0.9\n# lr_schedule = ExponentialDecay(initial_learning_rate, decay_steps, decay_rate, staircase=True)\n\n# base_model.compile(\n#     optimizer=Adam(learning_rate= lr_schedule),\n#     loss='binary_crossentropy',\n#     metrics=[AUC()]\n# )\n","metadata":{"execution":{"iopub.status.busy":"2024-02-05T22:29:20.394375Z","iopub.execute_input":"2024-02-05T22:29:20.394748Z","iopub.status.idle":"2024-02-05T22:29:20.411411Z","shell.execute_reply.started":"2024-02-05T22:29:20.394719Z","shell.execute_reply":"2024-02-05T22:29:20.410166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Fit model 1\n# datagen = ImageDataGenerator(rescale=1.0 / 255.0)\n# train_generator = datagen.flow(X_train, y_train, batch_size=64)\n# val_generator = datagen.flow(X_val, y_val, batch_size=64)\n# mymodel_1.fit(train_generator, validation_data=val_generator, epochs=10, batch_size=64, verbose=2)\n","metadata":{"execution":{"iopub.status.busy":"2024-02-05T22:29:20.412795Z","iopub.execute_input":"2024-02-05T22:29:20.413208Z","iopub.status.idle":"2024-02-05T22:29:20.422837Z","shell.execute_reply.started":"2024-02-05T22:29:20.413176Z","shell.execute_reply":"2024-02-05T22:29:20.421967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot model 1 results\n\n# train, validation auc history from '.history' \n# harded so I only run it once. I can't run multiple tuning or I will be out of RAM/GPU :(\n\nmodel1_history_auc=     [0.7638, 0.8196, 0.8517, 0.8770, 0.9005]\nmodel1_history_auc_val= [0.8096, 0.8401, 0.8585, 0.8662, 0.8629]\n\nmodel1_history_loss=         [0.5863, 0.5084, 0.4687, 0.4301, 0.3895]\nmodel1_history_loss_val=     [0.5257, 0.4956, 0.4648, 0.4518, 0.4631 ]\n\n# history = mymodel_1.fit(X_train/255., y_train, validation_data=(X_val/255., y_val), epochs=5, batch_size=64, verbose=2)\n\n# Plot training and validation AUC\nplt.figure(figsize=(8, 3))\n\nplt.subplot(1, 2, 1)\n# plt.plot(history.history['auc'], label='Train AUC')\n# plt.plot(history.history['val_auc'], label='Validation AUC')\nplt.plot(range(1, 6),model1_history_auc, label='Train AUC')\nplt.plot(range(1, 6),model1_history_auc_val, label='Validation AUC')\nplt.title('Training and Validation AUC')\nplt.xlabel('Epoch')\nplt.ylabel('AUC')\nplt.legend()\n\n# Plot training and validation loss\nplt.subplot(1, 2, 2)\nplt.plot(range(1, 6),model1_history_loss, label='Train Loss')\nplt.plot(range(1, 6),model1_history_loss_val, label='Validation Loss')\nplt.title('Training and Validation Loss')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\n\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-02-05T22:29:20.423994Z","iopub.execute_input":"2024-02-05T22:29:20.424441Z","iopub.status.idle":"2024-02-05T22:29:20.922345Z","shell.execute_reply.started":"2024-02-05T22:29:20.424389Z","shell.execute_reply":"2024-02-05T22:29:20.921052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n### Model 1 Performance:\n- Achieved 0.9 training AUC after 5 epochs, potential for further improvement.\n- Validation AUC dipped after 4 epochs, indicating possible overfitting.\n\n### Tuning Plan:\n- Extend training to assess if the model's complexity is adequate.\n- Monitor validation AUC to ensure generalization performance.\n","metadata":{}},{"cell_type":"markdown","source":"# <a id=\"9\"></a>\n\n<div style=\"text-align: center; background-color: #81badd; font-size:70%; padding: 2px;border-radius:10px;\">\n    <h2>Part 2.3: Model Tuning 1 - Add more epochs to increase training AUC </h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"## Model 2 - add more epochs (train longer)\n* note I comment out the code and left with train/validation AUC and loss plots. I don't want to run those too many times to save time. ","metadata":{}},{"cell_type":"code","source":"# Build model 2\n# Note for model 1 to model 3, I used copied numpy data. I later found this method consumes too much memory.\nmymodel_2 = tf.keras.models.Sequential([\n    Conv2D(32, (3, 3), activation='relu', input_shape=(96, 96, 3)),\n    Conv2D(32, (3, 3), activation='relu'),\n    MaxPooling2D((2, 2)),\n    Conv2D(64, (3, 3), activation='relu'),\n    Conv2D(64, (3, 3), activation='relu'),\n    MaxPooling2D((2, 2)),\n    Conv2D(128, (3, 3), activation='relu'),\n    Conv2D(128, (3, 3), activation='relu'),\n    MaxPooling2D((2, 2)),\n    Conv2D(128, (1, 1), activation='relu'),\n    Flatten(),\n    Dense(128, activation='relu'),\n    Dense(128, activation='relu'),\n    Dense(1, activation='sigmoid')\n])\n\n# model summary\nmymodel_2.summary()","metadata":{"execution":{"iopub.status.busy":"2024-02-05T22:29:20.923871Z","iopub.execute_input":"2024-02-05T22:29:20.924211Z","iopub.status.idle":"2024-02-05T22:29:21.524508Z","shell.execute_reply.started":"2024-02-05T22:29:20.924181Z","shell.execute_reply":"2024-02-05T22:29:21.523113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# compile model 2\n\n# Define the learning rate schedule (exponential decay)\n# initial_learning_rate = 0.001\n# decay_steps = 1000 # total steps = (training size/batchsize) * number of epochs, 200,000/batch size*epochs\n# decay_rate = 0.8\n# lr_schedule = ExponentialDecay(initial_learning_rate, decay_steps, decay_rate, staircase=True)\n\n# mymodel_2.compile(\n#     optimizer=Adam(learning_rate= lr_schedule),\n#     loss='binary_crossentropy',\n#     metrics=[AUC()]\n# )\n","metadata":{"execution":{"iopub.status.busy":"2024-02-05T22:29:21.526523Z","iopub.execute_input":"2024-02-05T22:29:21.527467Z","iopub.status.idle":"2024-02-05T22:29:21.534696Z","shell.execute_reply.started":"2024-02-05T22:29:21.527417Z","shell.execute_reply":"2024-02-05T22:29:21.532851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# fit model 2\n# datagen = ImageDataGenerator(rescale=1.0 / 255.0)\n# train_generator = datagen.flow(X_train, y_train, batch_size=48)\n# val_generator = datagen.flow(X_val, y_val, batch_size=48)\n# mymodel_2.fit(train_generator, validation_data=val_generator, epochs=8, batch_size=48, verbose=2)\n","metadata":{"execution":{"iopub.status.busy":"2024-02-05T22:29:21.536482Z","iopub.execute_input":"2024-02-05T22:29:21.536997Z","iopub.status.idle":"2024-02-05T22:29:21.551644Z","shell.execute_reply.started":"2024-02-05T22:29:21.536965Z","shell.execute_reply":"2024-02-05T22:29:21.550645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot model 2 results\n\n# train, validation auc history from '.history' \n# hard coded so I only run it once. I can't run multiple tuning or I will be out of RAM/GPU :(\n\nmodel2_history_auc=     [0.7379, 0.8521, 0.8796, 0.8935, 0.8999, 0.9027, 0.9039, 0.9044]\nmodel2_history_auc_val= [0.8243, 0.8671, 0.8749, 0.8785, 0.8810, 0.8814, 0.8816, 0.8816]\n\nmodel2_history_loss=         [0.5842, 0.4674, 0.4263, 0.4030, 0.3915, 0.3864, 0.3841,0.3833]\nmodel2_history_loss_val=     [0.5049, 0.4509, 0.4335, 0.4282, 0.4241, 0.4239, 0.4233,0.4234 ]\n\n# history = mymodel_1.fit(X_train/255., y_train, validation_data=(X_val/255., y_val), epochs=5, batch_size=64, verbose=2)\n\n# Plot training and validation AUC\nplt.figure(figsize=(8, 3))\n\nplt.subplot(1, 2, 1)\n# plt.plot(history.history['auc'], label='Train AUC')\n# plt.plot(history.history['val_auc'], label='Validation AUC')\nplt.plot(range(1, 9), model2_history_auc, label='Train AUC')\nplt.plot(range(1, 9), model2_history_auc_val, label='Validation AUC')\nplt.title('Training and Validation AUC')\nplt.xlabel('Epoch')\nplt.ylabel('AUC')\nplt.legend()\n\n# Plot training and validation loss\nplt.subplot(1, 2, 2)\nplt.plot(range(1, 9),model2_history_loss, label='Train Loss')\nplt.plot(range(1, 9),model2_history_loss_val, label='Validation Loss')\nplt.title('Training and Validation Loss')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-05T22:29:21.552690Z","iopub.execute_input":"2024-02-05T22:29:21.553036Z","iopub.status.idle":"2024-02-05T22:29:22.009177Z","shell.execute_reply.started":"2024-02-05T22:29:21.553008Z","shell.execute_reply":"2024-02-05T22:29:22.007848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Model 2 Performance:\n- Model 2's training AUC is around 0.9, suggesting that extending training duration may not enhance performance. This indicates a potential need for increased model complexity.\n\n### Tuning Plan:\n- Address the complexity issue. Two approaches:\n  - Update Model Structure:\n    - Add more layers, filters, or neurons.\n    - Add batch normalization or other architectural enhancements.\n  - Transfer Learning:\n    - Using a pretrained model known for its complexity and broad training data. Note that success may vary based on data differences and training set size.\n    - For learning purpose, I will prioritize transfer learning as an initial step for learning purposes.\n    - If transfer learning falls short, revert to enhancing model complexity with additional layers and features.\n\n\nA few things about transfer learning 👇 and a great [artile](https://medium.com/starschema-blog/transfer-learning-the-dos-and-donts-165729d66625) to read","metadata":{}},{"cell_type":"markdown","source":"# <a id=\"10\"></a>\n\n<div style=\"text-align: center; background-color: #81badd; font-size:70%; padding: 2px;border-radius:10px;\">\n    <h2>Part 2.3: Model Tuning 2 - Add Model Complexity Attempt 1 (Use Transfer Learning) </h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"background-color: #FCFFE9; padding: 10px;\">\n    \n## Transfer Learning\n\n#### Why transfer learning and how to use it?\n    \nIf dataset is very large, it can easily take more than a few weeks to train. It's good to use transfer learning in this case. A few ways to use it:\n- **As fixed feature extractor**: froze the weights in the convoluntion layers. Can add/train fully connected layers.\n- **Fine-tuning the CNN**: you data is big, similar to the transfer learning dataset but not exactly the same. You can 'de-forze' and train some CNN layers weights too.\n- **Use only part of CNN layer**: add a few of your own CNN layers also.\n  - A typical example for medical images.\n    - **Medical images** are harder to collect and even harder to label since they need to labeled by doctors.\n    - Trained by big ImageNet data, the first few layers can find some primitive features that can be used for medical images. We can use the earlier layers and add some later layers to train specifically for medical images.\n   \n    \n    \n#### When to use a pre-trained network?\n- **❌New dataset is small (ex: 20,000 with 100 categories) and similar to the original dataset (ex: 2,000,000 with 1000 categories)**\n    - Since the network is huge, you have froze the weight, which means you aren't training the model. Even not consider computation resource that you can actually re-train the model, the model is going to overfit since it's too big and more suitable for much bigger data.\n    \n- **✅New dataset is large and similar to the original dataset**\n    - can gain training time by using either fixed or fresh weights\n- **❌New dataset is small but very different from the original dataset**\n    - Overfit as in case in. Additionally, if very different, the transfer learning efficiency would be low.\n- **✅New dataset is large and very different from the original dataset**\n    - Can still use part of feature extractor of earlier layers. Still can save lots of training time.\n    \nVarious [models](https://keras.io/applications/) pre-trained on ImageNet.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"background-color: #FCFFE9; padding: 10px;\">\n    \n## InceptionV3 \n#### Insights from InceptionV3 Architecture:\n- Lower Layers:\n  - Basic convolutional Operations: such as Conv2d_1a_3x3, Conv2d_2a_3x3, and Conv2d_2b_3x3 \n  - Fundamental local features extractraction: such as edges, textures, colors, and gradients\n  - Simple Shapes and Contrast: In addition to low-level features, identify simple shapes and capturing contrast variations in the images.\n- Intermediate Layers:\n  - Mixed Operations:  blocks such as Mixed_3a, Mixed_4a, and Mixed_5a, combining different filter sizes and operations in parallel.\n  - Repetitive Structures capturing: such as repetitive structures within the images, recognizing patterns that repeat across the dataset.\n  - Color Variations and Relationships: further understand color variations and establish hierarchical relationships between different parts of objects in the images.\n  - Object Boundaries: delineating object boundaries, enhancing the model's ability to recognize distinct objects and their spatial relationships.\n- Deeper Inception Blocks:\n  - Mixed Operations: Blocks such as Mixed_5b, Mixed_6a, Mixed_7a\n  - Complex Shapes and Structures: improved understanding of complex shapes and intricate structures present in the images.\n  - Discrimination of Patterns: enhanced discrimination of complex patterns, distinguish fine-grained details and textures.\n\n#### My Adaptation Plan:\n- Feature Extraction:\n  - Utilize earlier blocks for low-level features \n- Intermediate Layers Modification:\n  - Introduce extra convolution layers for medical image nuances.\n  - Post 'mixed2' layer in InceptionV3, the feature map is roughly (9 * 9). Add padding to prevent rapid size reduction (retain spatial information)\n- Enhancements with 1x1 Filters:\n  - Incorporate 1x1 filters for improved learning.\n- Handle Overfitting:\n  - Batch normalization\n  - Dropout layers.\n- Image Augmentations for better generalization.","metadata":{}},{"cell_type":"markdown","source":"## Model 3 - transfer learning with Inception V3","metadata":{}},{"cell_type":"code","source":"# Build model 3\n# Note for model 1 to model 3, I used copied numpy data. I later found this method consumes too much memory.\n\n# Load the InceptionV3 model\nincept_model = InceptionV3(weights='imagenet', include_top=False, input_shape=(96, 96, 3))\n\n# Freeze the weights of the layers.\nfor layer in incept_model.layers:\n    layer.trainable = False\n\n# get the last_layer of inception we want to use\nlast_layer = incept_model.get_layer('mixed2')\nprint('last layer output shape: ', last_layer.output_shape)\n\n# Create a new model using the functional API\nbase_model = Model(inputs=incept_model.input, outputs=last_layer.output)\n\n# Add layers to the model\nmymodel_3 = Sequential([\n    base_model,\n    \n    Conv2D(64, (3, 3), padding='same', activation='relu'), # keep the width and height\n    BatchNormalization(),\n    Conv2D(64, (1, 1),activation='relu'),\n    BatchNormalization(),\n    Dropout(0.2),\n    \n    Conv2D(128, (3, 3), activation='relu'),\n    BatchNormalization(),            \n    Conv2D(64, (1, 1), activation='relu'),\n    BatchNormalization(),\n    Dropout(0.2),\n\n    Flatten(),\n    Dense(128, activation='relu'),\n    BatchNormalization(),\n    Dropout(0.2),\n    \n    Dense(128, activation='relu'),\n    BatchNormalization(),\n    Dropout(0.2),\n    \n    Dense(1, activation='sigmoid')\n])\n\n# Print the model summary\nmymodel_3.build(input_shape=(None, 96, 96, 3))\nmymodel_3.summary()","metadata":{"execution":{"iopub.status.busy":"2024-02-05T22:29:22.011125Z","iopub.execute_input":"2024-02-05T22:29:22.011511Z","iopub.status.idle":"2024-02-05T22:29:29.937632Z","shell.execute_reply.started":"2024-02-05T22:29:22.011475Z","shell.execute_reply":"2024-02-05T22:29:29.936077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# compile model 3\n\n#Define the learning rate schedule (exponential decay)\ninitial_learning_rate = 0.001\ndecay_steps = 1000\ndecay_rate = 0.9 \nlr_schedule = ExponentialDecay(initial_learning_rate, decay_steps, decay_rate, staircase=True)\n\n# Compile the model\nmymodel_3.compile(\n    optimizer=Adam(learning_rate=lr_schedule),\n    loss='binary_crossentropy',\n    metrics=[AUC()]\n)","metadata":{"execution":{"iopub.status.busy":"2024-02-05T22:29:29.939580Z","iopub.execute_input":"2024-02-05T22:29:29.939985Z","iopub.status.idle":"2024-02-05T22:29:30.003582Z","shell.execute_reply.started":"2024-02-05T22:29:29.939948Z","shell.execute_reply":"2024-02-05T22:29:30.002428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Fit model 3\nBATCH_SIZE = 32\nEPOCHS = 10\n\n# data generator\ndatagen = ImageDataGenerator(\n    rescale=1./255,\n    rotation_range=20,\n    width_shift_range=0.2,\n    height_shift_range=0.2,\n    vertical_flip=True,\n    horizontal_flip=True,\n)\n\ndatagen_val = ImageDataGenerator(\n    rescale=1./255,\n)\n\ntrain_generator = datagen.flow_from_directory(directory=train_dir,\n                                            batch_size=BATCH_SIZE,\n                                            class_mode='binary',\n                                            target_size=(96, 96))\n\nval_generator = datagen_val.flow_from_directory(directory=val_dir, \n                                                batch_size=BATCH_SIZE, \n                                                class_mode='binary',\n                                               target_size=(96,96))\n\nsteps_per_epoch = len(train_generator) \nvalidation_steps = len(val_generator)\n\n# Fit model 3\nhistory_3 = mymodel_3.fit(\n    train_generator,\n    validation_data=val_generator,\n    epochs=EPOCHS,\n    steps_per_epoch=steps_per_epoch,\n    validation_steps=validation_steps,\n    verbose=1\n)","metadata":{"execution":{"iopub.status.busy":"2024-02-05T22:29:30.005488Z","iopub.execute_input":"2024-02-05T22:29:30.005963Z","iopub.status.idle":"2024-02-06T02:57:39.976310Z","shell.execute_reply.started":"2024-02-05T22:29:30.005921Z","shell.execute_reply":"2024-02-06T02:57:39.967553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot metrics\nepochs = np.arange(1, EPOCHS+1)\nplt.figure(figsize=(8, 3))\n\n# AUC\nplt.subplot(1, 2, 1)\nplt.plot(epochs, history_3.history['auc'], label='Train')\nplt.plot(epochs, history_3.history['val_auc'], label='Validation')\nplt.xlabel('Epochs')\nplt.ylabel('AUC')\nplt.legend()\nplt.title('AUC Over Epochs')\n\n# Loss\nplt.subplot(1, 2, 2)\nplt.plot(epochs, history_3.history['loss'], label='Train')\nplt.plot(epochs, history_3.history['val_loss'], label='Validation')\nplt.xlabel('Epochs')\nplt.ylabel('Loss')\nplt.legend()\nplt.title('Loss Over Epochs')\n\nplt.tight_layout()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:08:07.118520Z","iopub.execute_input":"2024-02-06T03:08:07.119013Z","iopub.status.idle":"2024-02-06T03:08:07.678826Z","shell.execute_reply.started":"2024-02-06T03:08:07.118975Z","shell.execute_reply":"2024-02-06T03:08:07.677475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Model 3 Performance:\n- Achieved satisfactory training AUC of around 0.96 and a commendable validation AUC of 0.97 in the last epoch.\n- validation AUC plateaued after 8 epochs.\n\n### Next Steps:\n- Model evaluation, focusing on metrics such as precision, recall, F1 score, and confusion matrix to gain a holistic understanding of the model's performance. This approach ensures a balanced assessment of the model's predictive capabilities beyond AUC, providing insights into its precision, recall, and overall classification effectiveness.","metadata":{}},{"cell_type":"code","source":"# Predict validation set \nsamples_val = len(os.listdir(os.path.join(val_dir, '1'))) + len(os.listdir(os.path.join(val_dir, '0')))\nval_generator_for_pred = datagen_val.flow_from_directory(directory=val_dir, \n                                                batch_size=1, \n                                                class_mode='binary',\n                                               target_size=(96,96),\n                                                shuffle=False)\n\n\nval_predictions = mymodel_3.predict_generator(val_generator_for_pred,steps=samples_val, verbose=1)\n","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:08:07.681489Z","iopub.execute_input":"2024-02-06T03:08:07.682002Z","iopub.status.idle":"2024-02-06T03:12:45.634359Z","shell.execute_reply.started":"2024-02-06T03:08:07.681953Z","shell.execute_reply":"2024-02-06T03:12:45.632480Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Actual vs. predicted classes for validation set\n\n# actual\nval_labels = val_generator_for_pred.classes\n\n# convert probability to classes using 0.5 shreshold\nbinary_val_predictions = (val_predictions > 0.5).astype(int)\n","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:12:45.636862Z","iopub.execute_input":"2024-02-06T03:12:45.637278Z","iopub.status.idle":"2024-02-06T03:12:45.645082Z","shell.execute_reply.started":"2024-02-06T03:12:45.637243Z","shell.execute_reply":"2024-02-06T03:12:45.643860Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot ROC curve\n\n# Calculate AUC\nfpr, tpr, thresholds = roc_curve(val_labels, val_predictions)\nroc_auc = auc(fpr, tpr)\n\n# Plot \nplt.figure(figsize=(4, 4))\nplt.plot(fpr, tpr, color='orange', lw=2, label='ROC curve (AUC = {:.2f})'.format(roc_auc))\nplt.plot([0, 1], [0, 1], color='navy', lw=2, linestyle='--', label='Random Guessing')\nplt.xlabel('FPR')\nplt.ylabel('TPR')\nplt.title('ROC Curve')\nplt.legend(loc='lower right')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:12:45.648578Z","iopub.execute_input":"2024-02-06T03:12:45.649111Z","iopub.status.idle":"2024-02-06T03:12:45.922113Z","shell.execute_reply.started":"2024-02-06T03:12:45.649069Z","shell.execute_reply":"2024-02-06T03:12:45.920710Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get recall, predicision, F1, accuracy scores for validation set\nrecall = recall_score(val_labels, binary_val_predictions)\nprecision = precision_score(val_labels, binary_val_predictions)\nF1 = f1_score(val_labels, binary_val_predictions)\nacc= accuracy_score(val_labels, binary_val_predictions)\n\nprint(\"Recall on validation set:\", recall)\nprint(\"Precision on validation set:\", precision)\nprint(\"F1 score n on validation set:\", F1)\nprint(\"Accuracy score n on validation set:\", acc)","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:12:45.924145Z","iopub.execute_input":"2024-02-06T03:12:45.925080Z","iopub.status.idle":"2024-02-06T03:12:45.966357Z","shell.execute_reply.started":"2024-02-06T03:12:45.925024Z","shell.execute_reply":"2024-02-06T03:12:45.965028Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot confusion matrix\ncm = confusion_matrix(val_labels, binary_val_predictions)\ndisp = ConfusionMatrixDisplay(confusion_matrix=cm, display_labels=['Non-cancer', 'Cancer'])\ndisp.plot(cmap='Blues', values_format='d')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:12:45.968510Z","iopub.execute_input":"2024-02-06T03:12:45.969154Z","iopub.status.idle":"2024-02-06T03:12:46.289572Z","shell.execute_reply.started":"2024-02-06T03:12:45.969115Z","shell.execute_reply":"2024-02-06T03:12:46.288476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Final Model Reflection:\n\n- Recall Concerns:\n  - Recall around 83% is suboptimal for cancer detection.\n  - 1350+ cancer images misclassified as non-cancer signals a notable false-negative rate.\n- Importance of Recall:\n  - High recall crucial in cancer detection.\n\n- Future Considerations:\n  - Boosting recall through class weight adjustment or model fine-tuning. Emphasizes the need to strike a balance between overall model performance and sensitivity in critical scenarios like cancer detection.","metadata":{}},{"cell_type":"markdown","source":"# <a id=\"11\"></a>\n<div style=\"text-align: center; background-color: #AFDDCE; font-size:100%; padding: 5px;border-radius:10px 10px;\">\n    <h1>Part 3: Test Results </h1>\n</div>","metadata":{}},{"cell_type":"code","source":"# Copy files to test_dir\n\n# Directories\ntest_source_path = os.path.join(base_dir,'test')\ntest_images = os.listdir(test_path)\ntest_dir=os.path.join(work_dir,'test')\ntest_dir_images=os.path.join(test_dir,'Pos and neg images') # put all images directly to work_dir won't work! Need to put images under a folder in test_dir \n\n# remove the entire directory and its contents if already exist\nif os.path.exists(test_dir):\n    shutil.rmtree(test_dir)\n    \n# Make directories   \nos.makedirs(test_dir)\nos.makedirs(test_dir_images)\n\n# Copy images\nfor fn in test_images:\n    source_fns = os.path.join(test_source_path, fn)\n    test_fns = os.path.join(test_dir_images, fn)\n    shutil.copyfile(source_fns,test_fns)\n\n","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:12:46.291227Z","iopub.execute_input":"2024-02-06T03:12:46.292305Z","iopub.status.idle":"2024-02-06T03:22:35.687067Z","shell.execute_reply.started":"2024-02-06T03:12:46.292249Z","shell.execute_reply":"2024-02-06T03:22:35.685629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check number of images under test dir\nprint(len(os.listdir(test_dir)))\nprint(len(os.listdir(test_source_path)))","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:22:35.688685Z","iopub.execute_input":"2024-02-06T03:22:35.689087Z","iopub.status.idle":"2024-02-06T03:22:35.739848Z","shell.execute_reply.started":"2024-02-06T03:22:35.689052Z","shell.execute_reply":"2024-02-06T03:22:35.738852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Test data generator\ndatagen_test = ImageDataGenerator(\n    rescale=1./255)\n\ntest_generator = datagen_test.flow_from_directory(directory=test_dir, \n                                                batch_size= BATCH_SIZE, \n                                                class_mode=None,\n                                               target_size=(96,96),\n                                                shuffle=False)\n\ntest_predictions = mymodel_3.predict(test_generator, verbose=1)","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:22:35.741227Z","iopub.execute_input":"2024-02-06T03:22:35.741804Z","iopub.status.idle":"2024-02-06T03:27:08.008894Z","shell.execute_reply.started":"2024-02-06T03:22:35.741746Z","shell.execute_reply":"2024-02-06T03:27:08.007439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create test df\ntest_df= pd.DataFrame({'id':  test_generator.filenames,'label': test_predictions.flatten()})\ntest_df.head(1)","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:27:08.013056Z","iopub.execute_input":"2024-02-06T03:27:08.013477Z","iopub.status.idle":"2024-02-06T03:27:08.065451Z","shell.execute_reply.started":"2024-02-06T03:27:08.013440Z","shell.execute_reply":"2024-02-06T03:27:08.064196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The id for submission should not have '.tif'","metadata":{}},{"cell_type":"code","source":"# Source https://www.kaggle.com/code/vbookshelf/cnn-how-to-use-160-000-images-without-crashing\n# Get rid of '.tif' in 'id' column\ndef extract_id(x):\n    \n    # split into a list\n    a = x.split('/')\n    # split into a list\n    b = a[1].split('.')\n    extracted_id = b[0]\n    \n    return extracted_id\n\ntest_df['id'] = test_df['id'].apply(extract_id)\n\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:27:08.067371Z","iopub.execute_input":"2024-02-06T03:27:08.067753Z","iopub.status.idle":"2024-02-06T03:27:08.154079Z","shell.execute_reply.started":"2024-02-06T03:27:08.067718Z","shell.execute_reply":"2024-02-06T03:27:08.152813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# submition\ntest_df.to_csv('CNN submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:27:08.155570Z","iopub.execute_input":"2024-02-06T03:27:08.156478Z","iopub.status.idle":"2024-02-06T03:27:08.434369Z","shell.execute_reply.started":"2024-02-06T03:27:08.156437Z","shell.execute_reply":"2024-02-06T03:27:08.433034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# remove files in test_dir to prevent error\nshutil.rmtree(test_dir)","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:27:08.435744Z","iopub.execute_input":"2024-02-06T03:27:08.436275Z","iopub.status.idle":"2024-02-06T03:27:11.108602Z","shell.execute_reply.started":"2024-02-06T03:27:08.436225Z","shell.execute_reply":"2024-02-06T03:27:11.107215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Testing AUC after submission🙌:\n- Score: 0.9334\n- Private score: 0.9361\n\n### Thoughts on testing AUC:\n  - There is a notable gap exists between testing AUC (around 0.93) and validation AUC (around 0.97). Possible Explanations and improvement tips:   \n    - Overfitting Issues: The discrepancy may indicate lingering overfitting concerns. Increasing augmentation and dropout might help.\n    - Validation Set Size: The validation set size at 10% might be limiting. Expanding it could provide a more representative evaluation of model generalization.\n    - AUC as Metric:  AUC, while a powerful metric, introduces complexities during model compilation, as it considers performance across all thresholds. In contrast, practical evaluation metrics like recall, precision, and the confusion matrix often rely on a fixed threshold, typically set at 0.5. Using another metric might help.\n\n \n","metadata":{}},{"cell_type":"markdown","source":"### Model Tuning Progress and Performance Comparison summary:\n\n- Model 1:\n\n  - Utilized 3 blocks of 2 convolutions followed by 1 max pooling structure.\n  - Achieved a training AUC of 0.90 and a validation AUC of 0.86.\n  - Training for 5 epochs revealed potential for further improvement.\n\n- Model 2:\n\n  - Retained Model 1 structure but extended training duration.\n  - Training AUC remained unchanged, while validation AUC increased to 0.88.\n  - Indicated that the initial model structure was too simplistic.\n- Model 3 (Transfer Learning):\n  - Employed Transfer Learning, significantly increasing model complexity and addressed overfitting problem at the same time.\n  - Demonstrated substantial improvement, elevating validation AUC from 0.88 to an impressive 0.97, and training AUC from 0.90 to 0.96.","metadata":{}},{"cell_type":"markdown","source":"# <a id=\"12\"></a>\n<div style=\"text-align: center; background-color: #AFDDCE; font-size:100%; padding: 5px;border-radius:10px 10px;\">\n    <h1>Part 3: Project Conclusion </h1>\n</div>","metadata":{}},{"cell_type":"markdown","source":"### Conclusion and take aways\nIn the course of this project, I gained valuable insights into leveraging Convolutional Neural Networks (CNNs) for medical image cancer detection. Several noteworthy takeaways:\n- Data Handling Challenges: Managing large datasets can be challenging. I explored various methods for efficient data loading, such as converting images to int8 numpy arrays. However, this approach proved memory-intensive, impacting model training.\n- Transfer Learning Impact: The adoption of transfer learning significantly elevated the model's complexity. Also, transfer learning model is usually trained on a more diverse and larger dataset,this infusion of external knowledge could empower the model to discern intricate patterns in medical images. Moreover, transfer Learning could address overfitting problems at the same time.\n- Considerations for Model Evaluation: AUC, while a powerful metric, introduces complexities during model compilation, as it considers performance across all thresholds. In contrast, practical evaluation metrics like recall, precision, and the confusion matrix often rely on a fixed threshold, typically set at 0.5.\n\n### Future Improvements:\n\n- Exploring Alternative Metrics: Experimenting with alternative metrics, such as accuracy, could provide a different perspective on model performance. However, achieving a balanced dataset would be crucial in this context.\n\n- Optimizing Region of Interest: Focusing on the central 32x32 region, where cancer detection is most relevant, may involve cropping images to enhance model efficiency and effectiveness.\n- Addressing testing AUC discrepancy:\n  - Experimenting with increased augmentation and dropout techniques to reduce overfitting.\n  - Consider expanding the validation set size for a more representative evaluation of model generalization.\n\n","metadata":{}},{"cell_type":"markdown","source":"# <a id=\"13\"></a>\n\n## References (really helpful notebooks🙏🏽)\nhttps://www.kaggle.com/code/jieshends2020/baseline-keras-cnn-roc-fast-10min-0-925-lb/edit<br>\nhttps://www.kaggle.com/code/akarshu121/cancer-detection-with-cnn-for-beginners<br>\nhttps://www.kaggle.com/code/qitvision/a-complete-ml-pipeline-fast-ai<br>\nhttps://www.kaggle.com/code/CVxTz/cnn-starter-nasnet-mobile-0-9709-lb<br>\nhttps://www.kaggle.com/code/suicaokhoailang/wip-densenet121-baseline-with-fastai<br>\nhttps://www.kaggle.com/code/vbookshelf/cnn-how-to-use-160-000-images-without-crashing<br>\n","metadata":{}}]}