{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Histopathologic Cancer Detection: Exploring architectural variations and hyperparameter impact on CNN","metadata":{}},{"cell_type":"markdown","source":"**Jay Manvirk (Ivan Loginov)**<br/>University of Colorado, Boulder<br/>jay.manvirk@gmail.com","metadata":{}},{"cell_type":"markdown","source":"# Table of Contents\n\n1. [Abstract](#chapter_1)\n2. [Introduction](#chapter_2)\n3. [Libraries and raw data](#chapter_3)\n    - 3.1 [Libraries](#chapter_3_1)\n    - 3.2 [Raw data](#chapter_3_2)\n4. [Exploratory data analysis](#chapter_4)\n    - 4.1 [Short datasets summary](#chapter_4_1)\n    - 4.2 [Number of records per label](#chapter_4_2)\n    - 4.3 [Images](#chapter_4_3)\n5. [Data preprocessing](#chapter_5)\n    - 5.1 [Sampling](#chapter_5_1)\n    - 5.2 [Train-test split](#chapter_5_2)\n    - 5.3 [Parallel preprocessing](#chapter_5_3)\n6. [Model architecture](#chapter_6)\n    - 6.1 [Base model](#chapter_6_1)\n    - 6.2 [Base + Additional layers](#chapter_6_2)\n    - 6.3 [Base + Wider layers](#chapter_6_3)\n    - 6.4 [Base + Max pooling](#chapter_6_4)\n    - 6.5 [Base + Dropout regularization](#chapter_6_5)\n7. [Model results](#chapter_7)\n    - 7.1 [Base model](#chapter_7_1)\n    - 7.2 [Base + Additional layers](#chapter_7_2)\n    - 7.3 [Base + Wider layers](#chapter_7_3)\n    - 7.4 [Base + Max pooling](#chapter_7_4)\n    - 7.5 [Base + Dropout regularization](#chapter_7_5)\n    - 7.6 [Table results comparison](#chapter_7_6)\n8. [Submission results](#chapter_8)\n9. [Conclusion](#chapter_9)\n10. [References](#chapter_10)","metadata":{}},{"cell_type":"markdown","source":"# 1. Abstract <a class=\"anchor\" id=\"chapter_1\"></a>","metadata":{}},{"cell_type":"markdown","source":"This study briefly investigates the impact of different architectures and hyperparameters on the performance of convolutional neural networks (CNNs) through a comparative analysis of five distinct models. The research focuses on understanding how variations in the model design influence the models' ability to classify pathology images. The dataset utilized in this study comprises a substantial collection of small pathology images, categorized into two classes: images containing metastatic cancer and images devoid of it. By altering model configurations such as network depth, layer sizes, max pooling and dropout layer inclusion, the study explores the models' accuracy, efficiency, and generalization capabilities.","metadata":{}},{"cell_type":"markdown","source":"# 2. Introduction <a class=\"anchor\" id=\"chapter_2\"></a>","metadata":{}},{"cell_type":"markdown","source":"The goal of this notebook is to understand different architectures and hyperparameter effect on model by comparing performance of the 5 simple CNNs:\n\n1. Base model, which is going to include convolutional, flatten and dense layers\n2. Base + Additional layers\n3. Base + Wider layers\n4. Base + Max pooling\n5. Base + Dropout regularization\n\nEvery specification presented here, except the first one, is a slight extension of the base model. With these examples we're aiming to descern the effects of the introduced designs on the model results. Concretely, we're going to measure models' runtime and train/test ROC AUC scores per epoch.\n\nWe will utilize a dataset consisting of small pathology images sourced from Kaggle for both training and submission purposes. The dataset comprises:\n* train set: 220 025 images\n* submission set: 57 458 images\n* image shape: (96, 96, 3)\n* image size: ~28 kB\n\nFor the sake of boosting the training speed of the aforementioned models we're going to incorporate Tensorflow parallelism and GPU accelerators.\n\nThe notebook is going to be divided into several sections in order to achieve our goal:\n* **Exploratory Data Analysis:**<br/>\nAs was mentioned in the preface of this dataset, the data has been already modified, so that it is duplicate-free. However, it's not enough to accept that data doesn't need any additional manipulation. Therefore we're going to look closer at the dataset internals for any additional cleaning or type conversion. Also we will examine the number of records per label to see whether the dataset needs sampling or not.\n* **Data preprocessing:**<br/>\nHere we will introduce 3 steps to improve future model trainings efficiency. Sampling, train-test split and parallel preprocessing.<br/>\nIn the Sampling section we will try to even number of records per label for better label representation during training. The train-test split section is presented only due to the reason, that we're not going to use cross validation. We will manually pass train and test sets directly to the models' fit function. And finally the first parallelism would be introduced in the parallel preprocessing section. It's going to be about how to utilize CPU cores efficiently to load images as fas as possible.\n* **Model architecture:**<br/>\nEvery model, except the base one, is going to be an updated version of the base model. In essence they are supposed to be simple and straightforward, so that it's easier to spot the impact of the different setups on the model. Another reason for the implemented simplicity is that this notebook doesn't take hours to run.\n* **Model results:**<br/>\nThis section is going to be filled with train/test scores per epoch charts, final table with the fitted models' runtime and the ROC AUC scores and results' explanation.","metadata":{}},{"cell_type":"markdown","source":"# 3. Libraries and data <a class=\"anchor\" id=\"chapter_3\"></a>","metadata":{}},{"cell_type":"markdown","source":"## 3.1 Libraries <a class=\"anchor\" id=\"chapter_3_1\"></a>","metadata":{}},{"cell_type":"code","source":"# basics\nimport os\nimport time\nimport numpy as np\n\n# EDA\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom PIL import Image\n\n# Data preprocessing\nimport tensorflow as tf\nimport tensorflow_io as tfio\nfrom sklearn.utils import resample\nfrom sklearn.model_selection import train_test_split\n\n# Convolutional neural network\nfrom keras.models import Sequential\nfrom tensorflow.keras import layers, models\n\n# helper functions\nfrom sklearn.metrics import roc_auc_score\nfrom tensorflow.keras.models import load_model","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3.2 Raw Data <a class=\"anchor\" id=\"chapter_3_2\"></a>","metadata":{}},{"cell_type":"code","source":"# Print list of files and directories in folder\ninput_dir = '/kaggle/input/histopathologic-cancer-detection'\nlist_l = [os.path.join(input_dir, x) for x in os.listdir(input_dir)]\nlist_l","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:06.121860Z","iopub.execute_input":"2023-10-12T09:27:06.122532Z","iopub.status.idle":"2023-10-12T09:27:06.131199Z","shell.execute_reply.started":"2023-10-12T09:27:06.122496Z","shell.execute_reply":"2023-10-12T09:27:06.130257Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set datasets and directory names\nsample_data = pd.read_csv(list_l[0])\ntrain_data = pd.read_csv(list_l[1])\ntrain_dir = list_l[3] + '/'\ntest_dir = list_l[2] + '/'","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:06.132873Z","iopub.execute_input":"2023-10-12T09:27:06.133223Z","iopub.status.idle":"2023-10-12T09:27:06.504340Z","shell.execute_reply.started":"2023-10-12T09:27:06.133193Z","shell.execute_reply":"2023-10-12T09:27:06.503421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Cleaning\ndel list_l","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:06.506710Z","iopub.execute_input":"2023-10-12T09:27:06.507123Z","iopub.status.idle":"2023-10-12T09:27:06.511197Z","shell.execute_reply.started":"2023-10-12T09:27:06.507093Z","shell.execute_reply":"2023-10-12T09:27:06.510239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Exploratory data analysis <a class=\"anchor\" id=\"chapter_4\"></a>","metadata":{}},{"cell_type":"markdown","source":"## 4.1 Short datasets summary <a class=\"anchor\" id=\"chapter_4_1\"></a>","metadata":{}},{"cell_type":"code","source":"def print_short_summary(name, data):\n    \"\"\"\n    Prints data head, shape and info.\n    Args:\n        name (str): name of dataset\n        data (dataframe): dataset in a pd.DataFrame format\n    \"\"\"\n    print(name)\n    print('\\n1. Data head:')\n    print(data.head())\n    print('\\n2. Data shape: {}'.format(data.shape))\n    print('\\n3. Data info:')\n    data.info()\n    \ndef print_number_files(dirpath):\n    print('{}: {} files'.format(dirpath, len(os.listdir(dirpath))))","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:06.512868Z","iopub.execute_input":"2023-10-12T09:27:06.513546Z","iopub.status.idle":"2023-10-12T09:27:06.523056Z","shell.execute_reply.started":"2023-10-12T09:27:06.513517Z","shell.execute_reply":"2023-10-12T09:27:06.522053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print_short_summary('Train data', train_data)","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:06.524934Z","iopub.execute_input":"2023-10-12T09:27:06.525599Z","iopub.status.idle":"2023-10-12T09:27:06.568155Z","shell.execute_reply.started":"2023-10-12T09:27:06.525533Z","shell.execute_reply":"2023-10-12T09:27:06.567371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print_short_summary('Sample data', sample_data)","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:06.569239Z","iopub.execute_input":"2023-10-12T09:27:06.570020Z","iopub.status.idle":"2023-10-12T09:27:06.586936Z","shell.execute_reply.started":"2023-10-12T09:27:06.569991Z","shell.execute_reply":"2023-10-12T09:27:06.585902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print_number_files(train_dir)","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:06.587959Z","iopub.execute_input":"2023-10-12T09:27:06.588449Z","iopub.status.idle":"2023-10-12T09:27:28.679565Z","shell.execute_reply.started":"2023-10-12T09:27:06.588422Z","shell.execute_reply":"2023-10-12T09:27:28.678537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print_number_files(test_dir)","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:28.680818Z","iopub.execute_input":"2023-10-12T09:27:28.681141Z","iopub.status.idle":"2023-10-12T09:27:34.518474Z","shell.execute_reply.started":"2023-10-12T09:27:28.681111Z","shell.execute_reply":"2023-10-12T09:27:34.517448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Cleaning\ndel print_short_summary, print_number_files","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:34.522182Z","iopub.execute_input":"2023-10-12T09:27:34.522632Z","iopub.status.idle":"2023-10-12T09:27:34.526899Z","shell.execute_reply.started":"2023-10-12T09:27:34.522608Z","shell.execute_reply":"2023-10-12T09:27:34.525964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.2 Number of records per label <a class=\"anchor\" id=\"chapter_4_2\"></a>","metadata":{}},{"cell_type":"markdown","source":"The number of records per label differs signinficantly. This might shift model's focus a bit more towards the majority class resulting in lower ROC AUC score. Additionaly it create a problem in the model scoring, if only one class is present in the selected batch. We will address these problems in the 5.1 Sampling section.","metadata":{}},{"cell_type":"code","source":"# Plot horizontal barplot of number of records per label\nplt.figure(figsize=(16, 9))\ntmp = train_data['label'].value_counts()\nsns.barplot(y=['No Cancer', 'Cancer'], x=tmp.values, orient='h')\nplt.xlabel('Number of records')\nplt.ylabel('Label')\nplt.title('Number of records per label')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:34.528122Z","iopub.execute_input":"2023-10-12T09:27:34.529064Z","iopub.status.idle":"2023-10-12T09:27:34.969786Z","shell.execute_reply.started":"2023-10-12T09:27:34.529032Z","shell.execute_reply":"2023-10-12T09:27:34.968891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Cleaning\ndel tmp","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:34.971082Z","iopub.execute_input":"2023-10-12T09:27:34.971989Z","iopub.status.idle":"2023-10-12T09:27:34.976270Z","shell.execute_reply.started":"2023-10-12T09:27:34.971957Z","shell.execute_reply":"2023-10-12T09:27:34.975114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.3 Images <a class=\"anchor\" id=\"chapter_4_3\"></a>","metadata":{}},{"cell_type":"code","source":"def get_images_to_plot(file_names):\n    \"\"\"\n    Returns list of images\n    Args:\n        file_names: list of filenames\n    Returns:\n        list of image objects\n    \"\"\"\n    return [Image.open(f) for f in file_names]\n\ndef get_image_label(dirname, data, labels, n = 5):\n    \"\"\"\n    Return dictionary with label-imagepath\n    Args:\n        dirname: name of the directory\n        data: dataset of file names\n        labels: list of labels\n        n (opt): number of images per label\n    Returns:\n        dict_img: dictionary with label-imagepath pairs\n    \"\"\"\n    dict_img = {}\n    for l in labels:\n        indexes = data['label'] == l\n        tmp = data[indexes][:n]\n        tmp = dirname + tmp['id'] + '.tif'\n        tmp = tmp.values\n        tmp = get_images_to_plot(tmp)\n        dict_img[l] = tmp\n        \n    return dict_img","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:34.977663Z","iopub.execute_input":"2023-10-12T09:27:34.977997Z","iopub.status.idle":"2023-10-12T09:27:34.988328Z","shell.execute_reply.started":"2023-10-12T09:27:34.977968Z","shell.execute_reply":"2023-10-12T09:27:34.987328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Print original image size\nimg_path = train_dir + train_data['id'][0] + '.tif'\nimg = Image.open(img_path)\nprint('Original image size: {}'.format(img.size))","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:34.989641Z","iopub.execute_input":"2023-10-12T09:27:34.990225Z","iopub.status.idle":"2023-10-12T09:27:35.066322Z","shell.execute_reply.started":"2023-10-12T09:27:34.990196Z","shell.execute_reply":"2023-10-12T09:27:35.065378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get 5 filenames per label\ndata = get_image_label(train_dir,train_data, [0,1])","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:35.067833Z","iopub.execute_input":"2023-10-12T09:27:35.068552Z","iopub.status.idle":"2023-10-12T09:27:35.167909Z","shell.execute_reply.started":"2023-10-12T09:27:35.068500Z","shell.execute_reply":"2023-10-12T09:27:35.166989Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Initialize subplots with 2 rows and 5 columns\nfig, axes = plt.subplots(nrows=2, ncols=5, figsize=(16, 9))\n\n# Loop through selected images and display in the respective rows\nlabels = ['No Cancer', 'Cancer']\nfor i in range(10):\n    row = i // 5\n    col = i % 5\n    axes[row, col].imshow(data[row][col])\n    axes[row, col].set_title(labels[row])\n    axes[row, col].axis('off')\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:35.169113Z","iopub.execute_input":"2023-10-12T09:27:35.169426Z","iopub.status.idle":"2023-10-12T09:27:36.239269Z","shell.execute_reply.started":"2023-10-12T09:27:35.169393Z","shell.execute_reply":"2023-10-12T09:27:36.238399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Cleaning\ndel get_images_to_plot, get_image_label, img_path, img\ndel data, fig, axes, labels, row, col","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:36.240192Z","iopub.execute_input":"2023-10-12T09:27:36.240499Z","iopub.status.idle":"2023-10-12T09:27:36.247088Z","shell.execute_reply.started":"2023-10-12T09:27:36.240471Z","shell.execute_reply":"2023-10-12T09:27:36.246060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5. Data preprocessing <a class=\"anchor\" id=\"chapter_5\"></a>","metadata":{}},{"cell_type":"markdown","source":"## 5.1 Sampling <a class=\"anchor\" id=\"chapter_5_1\"></a>","metadata":{}},{"cell_type":"markdown","source":"Downsampling or upsampling are about bias-variance trade-off. In the first one we reduce the number of records from the majority class. In the second one we replicate existing instances of minority class. In both cases we're trying to:\n* even number of labels to improve model performance, be it better generalization (upsample) or faster training (downsample)\n* lower the risk of error in the ROC AUC score calculation if only one label is present in the selected batch\n\nIn our case the choice falls in favor of downsampling. Since we're dealing with simple CNN comparison our aim here is not the top performance algorithm. Thus we need to tune not just the models but the training data as well in order to reduce computational time as much as possible. Of course, all of that has to be done without significant loss in the ROC AUC score.","metadata":{}},{"cell_type":"code","source":"# Majority class\nno_cancer = train_data[train_data['label'] == 0]\n# Minority class\ncancer = train_data[train_data['label'] == 1]\n\n# Downsample majority class to match minority class\nno_cancer_downsampled = resample(no_cancer,\n                              replace=False, \n                              n_samples=len(cancer),\n                              random_state=0)\n\nbalanced_train_data = pd.concat([no_cancer_downsampled, cancer])\n\n# Shuffle train data for training\nbalanced_train_data = balanced_train_data.sample(frac=1, random_state=0).reset_index(drop=True)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:36.249079Z","iopub.execute_input":"2023-10-12T09:27:36.249787Z","iopub.status.idle":"2023-10-12T09:27:36.354650Z","shell.execute_reply.started":"2023-10-12T09:27:36.249757Z","shell.execute_reply":"2023-10-12T09:27:36.353764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Cleaning\ndel no_cancer, cancer, no_cancer_downsampled","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:36.358760Z","iopub.execute_input":"2023-10-12T09:27:36.359516Z","iopub.status.idle":"2023-10-12T09:27:36.368520Z","shell.execute_reply.started":"2023-10-12T09:27:36.359482Z","shell.execute_reply":"2023-10-12T09:27:36.367143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5.2 Train-test split <a class=\"anchor\" id=\"chapter_5_2\"></a>","metadata":{}},{"cell_type":"markdown","source":"Since we're not going to rely on cross validation, we need to provide both train and test samples to CNN to measure overfitting during training. Thus we first split our train dataset into 75% and 25% of train and test data respectively based on the default values of train_test_split function.","metadata":{}},{"cell_type":"code","source":"# Get full path to image including extension\nimage_paths = train_dir + balanced_train_data['id'] + '.tif'\nimage_paths = image_paths.values\n\nlabels = balanced_train_data['label'].values\n\nX_train, X_test, y_train, y_test = train_test_split(image_paths\n                                                    , labels\n                                                    , test_size = 0.25\n                                                    , shuffle = True\n                                                    , random_state = 0)","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:36.372686Z","iopub.execute_input":"2023-10-12T09:27:36.374645Z","iopub.status.idle":"2023-10-12T09:27:36.502046Z","shell.execute_reply.started":"2023-10-12T09:27:36.374604Z","shell.execute_reply":"2023-10-12T09:27:36.501107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Cleaning\ndel image_paths, labels","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:36.506586Z","iopub.execute_input":"2023-10-12T09:27:36.508700Z","iopub.status.idle":"2023-10-12T09:27:36.516524Z","shell.execute_reply.started":"2023-10-12T09:27:36.508665Z","shell.execute_reply":"2023-10-12T09:27:36.515390Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5.3 Parallel preprocessing <a class=\"anchor\" id=\"chapter_5_3\"></a>","metadata":{}},{"cell_type":"markdown","source":"If we were to load images sequentially one-by-one, we would lose significant amount of time. Therefore it's more efficient to run this process in parallel in batches, so that most of the time would be spent on the actual training. With that in mind we're going to incorporate Tensorflow parallel processing functions to use optimal number of CPU cores provided in this notebook.\n\nAdditional important measures to increase training speed are the image resolution reduction and pixel scaling.<br/>Given our dataset of 96x96px images we can shrink them to 32x32px. It's an arbitrary reduction found across other kernels on Kaggle, but it helps in our case of speed-score trade-off.<br/>\nThe justification for pixel scaling lies in the idea of preventing gradients from becoming too large (exploding gradients) or too small (vanishing gradients) during backpropagation. Having this kind of constraint in the gradient values improve model performance overall.","metadata":{}},{"cell_type":"code","source":"def get_decoded_image(image_path, label=None):\n    \"\"\"\n    Load and preprocess images using TensorFlow I/O.\n    Decode image with 4 channels RGBA.\n    Resize image to 32x32px.\n    Scale pixels from 0 to 1.\n    Args:\n        image_path: path to TIFF image\n        label (optional): true label from train data\n    Returns:\n        (img, label): for train data\n        img: for test data\n    \"\"\"\n    img = tf.io.read_file(image_path)\n    img = tfio.experimental.image.decode_tiff(img)\n    img = tf.image.resize(img, [32, 32])\n    img = tf.cast(img, tf.float32) / 255.0\n    \n    return img if label is None else (img, label)\n\ndef get_prefetched_data(data, batch_size):\n    \"\"\"\n    Create a TensorFlow dataset from image paths and labels.\n    Execution in parallel.\n    Load, preprocess images and batch the data.\n    Prefetch batches to improve training performance.\n    Args:\n        data (tuple): image paths and corresponding labels\n        batch_size (int): number of samples per batch\n    Returns:\n        tf.data.Dataset: preprocessed and preloaded TensorFlow dataset for keras CNN\n    \"\"\"\n    # Autotune the degree of parallelism during training\n    AUTOTUNE = tf.data.experimental.AUTOTUNE\n    \n    # Create dataset from image paths and labels\n    dataset = tf.data.Dataset.from_tensor_slices(data)\n    \n    # Apply parallel processing to load and preprocess images\n    dataset = dataset.map(get_decoded_image, num_parallel_calls=AUTOTUNE)\n    dataset = dataset.batch(batch_size)\n    dataset = dataset.prefetch(buffer_size=AUTOTUNE)\n    \n    return dataset","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:36.520809Z","iopub.execute_input":"2023-10-12T09:27:36.523325Z","iopub.status.idle":"2023-10-12T09:27:36.537238Z","shell.execute_reply.started":"2023-10-12T09:27:36.523292Z","shell.execute_reply":"2023-10-12T09:27:36.536264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set 128 samples to be processed in each training step\nBATCH_SIZE = 128\n\ntrain_dataset = get_prefetched_data((X_train, y_train)\n                                    , BATCH_SIZE)\ntest_dataset = get_prefetched_data((X_test, y_test)\n                                   , BATCH_SIZE)","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:36.541065Z","iopub.execute_input":"2023-10-12T09:27:36.544035Z","iopub.status.idle":"2023-10-12T09:27:40.142234Z","shell.execute_reply.started":"2023-10-12T09:27:36.544003Z","shell.execute_reply":"2023-10-12T09:27:40.141250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del X_train, y_train, X_test, y_test","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:40.143488Z","iopub.execute_input":"2023-10-12T09:27:40.144031Z","iopub.status.idle":"2023-10-12T09:27:40.160784Z","shell.execute_reply.started":"2023-10-12T09:27:40.144000Z","shell.execute_reply":"2023-10-12T09:27:40.159872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 6. Model architecture <a class=\"anchor\" id=\"chapter_6\"></a>","metadata":{}},{"cell_type":"markdown","source":"Besides different architectures and hyperparameters that are going to be introduced in the following sections there are some that will be present in every model:\n1. **Activations:**<br/>\nReLU: introduces non-linearity into the model by replacing all negative values in the input with zero and supposedly computationally efficient<br/>\nSigmoid: within the range between 0 and 1 it represents the probability of belonging to the positive class, which is useful in our binary class problem\n2. **Loss function:**<br/>\nBinary crossentropy loss: forces the model to output probabilities close to 1 for cancer instances and close to 0 otherwise\n3. **Optimizer:**<br/>\nAdam optimizer: efficient and commonly used optimization algorithm\n3. **ROC AUC score:**<br/>\nThis metric is used in the competition based on this dataset. Therefore it's useful to employ it in the training sessions as well for better understanding of what to expect from the model submission results.\n4. **Batch size:**<br/>\nThe batch size of 128 is chosen arbitrary. In our case it increases the training speed while maintaining decent results of every model.\n5. **Number of epochs:**<br/>\nWe're going to train models on 5 epochs. This value is solely based on the idea of quick experimentation.\n\nOther parameters are based on either default values from tensorflow package or computed automatically like steps per epoch by tf.data module.","metadata":{}},{"cell_type":"markdown","source":"## 6.1 Base model <a class=\"anchor\" id=\"chapter_6_1\"></a>","metadata":{}},{"cell_type":"markdown","source":"Our base model is going to be simple CNN yet with several layers. All the parameters are based either on the default values of a model, common values among other works or dataset features.\nIt is going to be comprised of:\n\n1. **Convolutional Layer:**<br/>\nIt's said that Conv2D layer can capture spatial patterns in the input data. Ours is going to be with 32 filters of size (3, 3). The input shape (32, 32, 4) is used to address 32x32px images with 4 channels RGBA which were preprocessed by tfio.experimental.image.decode_tiff().\n3. **Flatten Layer:**<br/>\nThe Flatten layer transforms the 2D feature maps into a 1D vector, preparing the data for the subsequent dense layers.\n4. **Dense Layers:**<br/>\nThe first Dense layer with 32 units and the ReLU activation function enables the model to learn intricate relationships in the data. The final Dense layer with 1 unit and the sigmoid activation function is to output probability of having cancer in the image.","metadata":{}},{"cell_type":"code","source":"def get_model_base():\n    \"\"\"\n    Return base model architecture\n    \"\"\"\n    model = models.Sequential([\n        layers.Conv2D(32, (3, 3), activation='relu', input_shape=(32, 32, 4))\n\n        , layers.Flatten()\n\n        , layers.Dense(32, activation='relu')\n\n        , layers.Dense(1, activation='sigmoid')\n    ])\n    \n    return model","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:40.162135Z","iopub.execute_input":"2023-10-12T09:27:40.162485Z","iopub.status.idle":"2023-10-12T09:27:40.173009Z","shell.execute_reply.started":"2023-10-12T09:27:40.162434Z","shell.execute_reply":"2023-10-12T09:27:40.172120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 6.2 Base + Additional layers <a class=\"anchor\" id=\"chapter_6_2\"></a>","metadata":{}},{"cell_type":"markdown","source":"* **Additional layers:**<br/>\nThis is done to increase base model capacity for feature extraction, which eventually should lead to better accuracy in predictions. However, it opens another door to overfitting. We're going to add one convolutional layer and one dense layer with the same amout of units as in the previous ones.","metadata":{}},{"cell_type":"code","source":"def get_model_base_deep():\n    \"\"\"\n    Return deeper model architecture\n    \"\"\"\n    model = models.Sequential([\n        layers.Conv2D(32, (3, 3), activation='relu', input_shape=(32, 32, 4))\n        # Add new convolutional layer\n        , layers.Conv2D(32, (3, 3), activation='relu')\n\n        , layers.Flatten()\n\n        , layers.Dense(32, activation='relu')\n        # Add new dense layer\n        , layers.Dense(32, activation='relu')\n\n        , layers.Dense(1, activation='sigmoid')\n    ])\n    \n    return model","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:40.174187Z","iopub.execute_input":"2023-10-12T09:27:40.175282Z","iopub.status.idle":"2023-10-12T09:27:40.187342Z","shell.execute_reply.started":"2023-10-12T09:27:40.175251Z","shell.execute_reply":"2023-10-12T09:27:40.186453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 6.3 Base + Wider layers <a class=\"anchor\" id=\"chapter_6_3\"></a>","metadata":{}},{"cell_type":"markdown","source":"* **Wider layers:**<br/>\nWith more filters in the convolutional layer and more neurons in the dense layer we allow model to learn fine-grained features of the images. Generally that leads to better classification accuracy. In our case we double the amount of units in both Conv2d and Dense layers.","metadata":{}},{"cell_type":"code","source":"def get_model_base_wide():\n    \"\"\"\n    Return wider model architecture\n    \"\"\"\n    model_drop_bn = models.Sequential([\n        # Increase number of units from 32 to 64\n        layers.Conv2D(64, (3, 3), activation='relu', input_shape=(32, 32, 4))\n\n        , layers.Flatten()\n        \n        # Increase number of units from 32 to 64\n        , layers.Dense(64, activation='relu')\n \n        , layers.Dense(1, activation='sigmoid')\n    ])\n    \n    return model_drop_bn","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:40.188504Z","iopub.execute_input":"2023-10-12T09:27:40.189330Z","iopub.status.idle":"2023-10-12T09:27:40.198402Z","shell.execute_reply.started":"2023-10-12T09:27:40.189301Z","shell.execute_reply":"2023-10-12T09:27:40.197542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 6.4 Base + Max pooling <a class=\"anchor\" id=\"chapter_6_4\"></a>","metadata":{}},{"cell_type":"markdown","source":"* **Max pooling layer:**<br/>\nThe MaxPooling2D layer focuses on the most important features. Our setup is the window size of (2,2) and the strides of (2,2) which are the default values. With that we reduce the spatial dimensions of the feature map by half in both width and height, which can increase model's training speed.","metadata":{}},{"cell_type":"code","source":"def get_model_base_maxpool():\n    \"\"\"\n    Return maxpool model architecture\n    \"\"\"\n    model = models.Sequential([\n        layers.Conv2D(32, (3, 3), activation='relu', input_shape=(32, 32, 4))\n        # Add new layer of max pooling\n        , layers.MaxPooling2D((2, 2), strides = (2,2))\n\n        , layers.Flatten()\n\n        , layers.Dense(32, activation='relu')\n        \n        , layers.Dense(1, activation='sigmoid')\n    ])\n    \n    return model","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:40.202906Z","iopub.execute_input":"2023-10-12T09:27:40.203292Z","iopub.status.idle":"2023-10-12T09:27:40.209444Z","shell.execute_reply.started":"2023-10-12T09:27:40.203271Z","shell.execute_reply":"2023-10-12T09:27:40.208399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 6.5 Base + Dropout regularization <a class=\"anchor\" id=\"chapter_6_5\"></a>","metadata":{}},{"cell_type":"markdown","source":"* **Dropout layer:**<br/>\nTo tackle the problem of overfitting, we can add additional layer of Dropout Regularization. Dropout randomly sets a fraction of input units to 0 during training, which prevents model memorizing train data. We're going to set the dropout rate to the default value of 0.25, indicating that during training, 25% of the input units will be dropped out (set to 0).","metadata":{}},{"cell_type":"code","source":"def get_model_base_dropout():\n    \"\"\"\n    Return dropout model architecture\n    \"\"\"\n    model = models.Sequential([\n        layers.Conv2D(32, (3, 3), activation='relu', input_shape=(32, 32, 4))\n\n        , layers.Flatten()\n\n        , layers.Dense(32, activation='relu')\n        \n        # Add new dropout layer \n        , layers.Dropout(0.25)\n        \n        , layers.Dense(1, activation='sigmoid')\n    ])\n    \n    return model","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:40.210728Z","iopub.execute_input":"2023-10-12T09:27:40.211485Z","iopub.status.idle":"2023-10-12T09:27:40.218696Z","shell.execute_reply.started":"2023-10-12T09:27:40.211435Z","shell.execute_reply":"2023-10-12T09:27:40.217826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 7. Model results <a class=\"anchor\" id=\"chapter_7\"></a>","metadata":{}},{"cell_type":"code","source":"def get_compiled_model(func):\n    \"\"\"\n    Create model to be trained with a multi-GPU strategy.\n    Args:\n        func: function to get model architecture\n    Returns:\n        compiled_model: tensorflow model that performs data parallelism\n                            by copying all of the model's variables\n                            to each processor\n    \"\"\"\n    # Check if GPU is available\n    gpus = tf.config.experimental.list_physical_devices('GPU')\n    if gpus:\n        # Create a MirroredStrategy.\n        strategy = tf.distribute.MirroredStrategy()\n\n        print('Number of devices: {}'.format(strategy.num_replicas_in_sync))\n    else:\n        strategy = tf.distribute.OneDeviceStrategy(device=\"/cpu:0\")\n        print('No GPU available, falling back to CPU.')\n\n    with strategy.scope():\n        compiled_model = func()\n        compiled_model.compile(optimizer = tf.keras.optimizers.Adam()\n                              , loss = tf.keras.losses.BinaryCrossentropy()\n                              , metrics = [tf.keras.metrics.AUC()])\n\n    return compiled_model","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:40.219945Z","iopub.execute_input":"2023-10-12T09:27:40.220871Z","iopub.status.idle":"2023-10-12T09:27:40.232708Z","shell.execute_reply.started":"2023-10-12T09:27:40.220761Z","shell.execute_reply":"2023-10-12T09:27:40.231789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_model_scores(scores, model_name):\n    \"\"\"\n    Plot train and test ROC AUC scores of a model by epoch\n    \"\"\"\n    train_scores, test_scores = scores\n    epochs = range(1, len(train_scores) + 1)\n\n    # Plot train and test scores\n    plt.figure(figsize=(16, 9))\n    plt.plot(epochs, train_scores, label='Train score')\n    plt.plot(epochs, test_scores, label='Test score')\n    plt.title('Train and test ROC AUC scores of the {}'.format(model_name))\n    plt.xlabel('Epoch')\n    plt.ylabel('ROC AUC Score')\n    plt.legend()\n    plt.grid(True)\n    plt.show()\n\n    \ndef get_model_results(model_name, model):\n    \"\"\"\n    Return tuple of runtime, train and test scores.\n    Compile, fit and save model along the way.\n    Args:\n        model: fitted model\n    Returns:\n        (runtime, (train_scores, test_scores) )\n    \"\"\"\n    model = get_compiled_model(model)\n    \n    st = time.time()\n    model.fit(train_dataset, epochs=5, validation_data=test_dataset)\n    runtime = time.time() - st\n    \n    model.save('{}.h5'.format(model_name))\n    \n    train_scores = model.history.history['auc']\n    test_scores = model.history.history['val_auc']\n    \n    tf.keras.backend.clear_session()\n    \n    return (runtime, (train_scores, test_scores))","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:27:40.233948Z","iopub.execute_input":"2023-10-12T09:27:40.234562Z","iopub.status.idle":"2023-10-12T09:27:40.243787Z","shell.execute_reply.started":"2023-10-12T09:27:40.234531Z","shell.execute_reply":"2023-10-12T09:27:40.242834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 7.1 Base model <a class=\"anchor\" id=\"chapter_7_1\"></a>","metadata":{}},{"cell_type":"code","source":"# Get train and test scores of every epoch\nruntime_base, scores_base = get_model_results('model_base',get_model_base)","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:48:04.001636Z","iopub.execute_input":"2023-10-12T09:48:04.002351Z","iopub.status.idle":"2023-10-12T09:54:03.176127Z","shell.execute_reply.started":"2023-10-12T09:48:04.002310Z","shell.execute_reply":"2023-10-12T09:54:03.175209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot scores\nplot_model_scores(scores_base, 'base model')","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:54:05.397733Z","iopub.execute_input":"2023-10-12T09:54:05.398486Z","iopub.status.idle":"2023-10-12T09:54:05.712758Z","shell.execute_reply.started":"2023-10-12T09:54:05.398434Z","shell.execute_reply":"2023-10-12T09:54:05.711917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 7.2 Base + Additional layers <a class=\"anchor\" id=\"chapter_7_2\"></a>","metadata":{}},{"cell_type":"code","source":"# Get train and test scores of every epoch\nruntime_base_deep, scores_base_deep = get_model_results('model_base_deep'\n                                                          ,get_model_base_deep)","metadata":{"execution":{"iopub.status.busy":"2023-10-12T09:54:15.238048Z","iopub.execute_input":"2023-10-12T09:54:15.238383Z","iopub.status.idle":"2023-10-12T10:00:11.005984Z","shell.execute_reply.started":"2023-10-12T09:54:15.238355Z","shell.execute_reply":"2023-10-12T10:00:11.004851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot scores\nplot_model_scores(scores_base_deep, 'base + additional layers model')","metadata":{"execution":{"iopub.status.busy":"2023-10-12T10:00:15.507392Z","iopub.execute_input":"2023-10-12T10:00:15.508409Z","iopub.status.idle":"2023-10-12T10:00:15.827672Z","shell.execute_reply.started":"2023-10-12T10:00:15.508364Z","shell.execute_reply":"2023-10-12T10:00:15.826838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 7.3 Base + Wider layers <a class=\"anchor\" id=\"chapter_7_3\"></a>","metadata":{}},{"cell_type":"code","source":"# Get train and test scores of every epoch\nruntime_base_wide, scores_base_wide = get_model_results('model_base_wide'\n                                                          ,get_model_base_wide)","metadata":{"execution":{"iopub.status.busy":"2023-10-12T10:00:21.209695Z","iopub.execute_input":"2023-10-12T10:00:21.210023Z","iopub.status.idle":"2023-10-12T10:06:37.096211Z","shell.execute_reply.started":"2023-10-12T10:00:21.209997Z","shell.execute_reply":"2023-10-12T10:06:37.095210Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot scores\nplot_model_scores(scores_base_wide, 'base + wider layers model')","metadata":{"execution":{"iopub.status.busy":"2023-10-12T10:06:38.853271Z","iopub.execute_input":"2023-10-12T10:06:38.854404Z","iopub.status.idle":"2023-10-12T10:06:39.269558Z","shell.execute_reply.started":"2023-10-12T10:06:38.854361Z","shell.execute_reply":"2023-10-12T10:06:39.268636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 7.4 Base + Max pooling <a class=\"anchor\" id=\"chapter_7_4\"></a>","metadata":{}},{"cell_type":"code","source":"# Get train and test scores of every epoch\nruntime_base_maxpool, scores_base_maxpool = get_model_results('model_base_maxpool'\n                                                                ,get_model_base_maxpool)","metadata":{"execution":{"iopub.status.busy":"2023-10-12T10:06:51.805957Z","iopub.execute_input":"2023-10-12T10:06:51.806287Z","iopub.status.idle":"2023-10-12T10:12:58.924906Z","shell.execute_reply.started":"2023-10-12T10:06:51.806258Z","shell.execute_reply":"2023-10-12T10:12:58.923937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot scores\nplot_model_scores(scores_base_maxpool, 'base + max pooling model')","metadata":{"execution":{"iopub.status.busy":"2023-10-12T10:13:00.641001Z","iopub.execute_input":"2023-10-12T10:13:00.641352Z","iopub.status.idle":"2023-10-12T10:13:00.963474Z","shell.execute_reply.started":"2023-10-12T10:13:00.641321Z","shell.execute_reply":"2023-10-12T10:13:00.962642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 7.5 Base + Dropout regularization <a class=\"anchor\" id=\"chapter_7_5\"></a>","metadata":{}},{"cell_type":"code","source":"# Get train and test scores of every epoch\nruntime_base_dropout, scores_base_dropout = get_model_results('model_base_dropout'\n                                                                ,get_model_base_dropout)","metadata":{"execution":{"iopub.status.busy":"2023-10-12T10:13:08.483666Z","iopub.execute_input":"2023-10-12T10:13:08.483992Z","iopub.status.idle":"2023-10-12T10:18:52.519702Z","shell.execute_reply.started":"2023-10-12T10:13:08.483966Z","shell.execute_reply":"2023-10-12T10:18:52.518777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot scores\nplot_model_scores(scores_base_dropout, 'base + dropout model')","metadata":{"execution":{"iopub.status.busy":"2023-10-12T10:18:56.054558Z","iopub.execute_input":"2023-10-12T10:18:56.054886Z","iopub.status.idle":"2023-10-12T10:18:56.373637Z","shell.execute_reply.started":"2023-10-12T10:18:56.054858Z","shell.execute_reply":"2023-10-12T10:18:56.372782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 7.6 Table results comparison <a class=\"anchor\" id=\"chapter_7_6\"></a>","metadata":{}},{"cell_type":"markdown","source":"The winner among selected designs is the 'Base + Additional layers' model with the highest test ROC AUC score (0.898) and the second best runtime (355 seconds).<br/>\nHere are the possible explanations of the model results in comparison to the 'Base' model in a sorted descending order by test ROC AUC score:\n1. **Base + Additional layers:**<br/>\nAdding extra convolutional and dense layers allowed the network to learn more complex patterns in the data, leading to better generalization on the test dataset.\n2. **Base + Max pooling:**<br/>\nThe addition of max pooling layer is generally used to reduce model complexity, and in this case, it seems like this reduction helped in improving the Base model scores, however resulted in slower training speed which was unexpected.\n3. **Base + Wider layers:**<br/>\nUsually widening the layers captures more relevant features. But in our case it apparently captured noise in the train data more as the test score ROC AUC ended up lower than train one, showing signs of overfitting.\n4. **Base + Dropout:**<br/>\nThe dropout layer was introduced to mitigate overfitting by randomly dropping out connections during training. And while this approach reduced overfitting, in the end it led to a decrease in both training and test scores. This regularization reduced Base model's ability to learn the intricacies of the dataset.\n\nIt's also crucial to note the consistent gap in ROC AUC scores across all models, except for the Base. This persistent disparity, indicative of underfitting, reveals that these models, despite their relatively higher scores, struggle to grasp the intricate nuances within the training data.","metadata":{}},{"cell_type":"code","source":"# Print table results comparison\nresults = [('Base', runtime_base, scores_base)\n          ,('Base + Add. layers', runtime_base_deep, scores_base_deep)\n          ,('Base + Wider layers', runtime_base_wide, scores_base_wide)\n          ,('Base + Max pooling', runtime_base_maxpool, scores_base_maxpool)\n          ,('Base + Dropout', runtime_base_dropout, scores_base_dropout)]\ntable = []\nfor i in range(len(results)):\n    tmp = {\n            'model': results[i][0]\n            , 'runtime (sec)': results[i][1]\n            , 'train_roc_auc_score': results[i][2][0][-1]\n            , 'test_roc_auc_score': results[i][2][1][-1]\n        }\n    table.append(tmp)\n\n\npd.DataFrame(table).sort_values(by = ['test_roc_auc_score'\n                                      ,'runtime (sec)']\n                                , ascending = [False, True])","metadata":{"execution":{"iopub.status.busy":"2023-10-12T10:19:09.174092Z","iopub.execute_input":"2023-10-12T10:19:09.174423Z","iopub.status.idle":"2023-10-12T10:19:09.194533Z","shell.execute_reply.started":"2023-10-12T10:19:09.174397Z","shell.execute_reply":"2023-10-12T10:19:09.193344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Cleaning\ndel results, tmp, table, train_dataset, test_dataset, train_data, train_dir, test_dir\ndel get_model_results, plot_model_scores, get_compiled_model","metadata":{"execution":{"iopub.status.busy":"2023-10-11T18:01:28.163062Z","iopub.execute_input":"2023-10-11T18:01:28.164378Z","iopub.status.idle":"2023-10-11T18:01:28.184416Z","shell.execute_reply.started":"2023-10-11T18:01:28.164305Z","shell.execute_reply":"2023-10-11T18:01:28.183460Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 8. Submission results <a class=\"anchor\" id=\"chapter_8\"></a>","metadata":{}},{"cell_type":"markdown","source":"#### Public score: 0.8156\n\nThis the best ROC AUC score of our top performance model.","metadata":{}},{"cell_type":"code","source":"# Load top performed model\nmodel = load_model('model_base_deep.h5')","metadata":{"execution":{"iopub.status.busy":"2023-10-12T10:20:10.446842Z","iopub.execute_input":"2023-10-12T10:20:10.447306Z","iopub.status.idle":"2023-10-12T10:20:10.638599Z","shell.execute_reply.started":"2023-10-12T10:20:10.447266Z","shell.execute_reply":"2023-10-12T10:20:10.637533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create prefethed dataset of images to classify\nsubmis_data = test_dir + sample_data['id'] + '.tif'\nsubmis_data = submis_data.values\n\nsubmis_dataset = get_prefetched_data((submis_data)\n                                    , BATCH_SIZE)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-12T10:20:16.901929Z","iopub.execute_input":"2023-10-12T10:20:16.902592Z","iopub.status.idle":"2023-10-12T10:20:16.952274Z","shell.execute_reply.started":"2023-10-12T10:20:16.902559Z","shell.execute_reply":"2023-10-12T10:20:16.951404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set results\nresults = model.predict(submis_dataset)","metadata":{"execution":{"iopub.status.busy":"2023-10-12T10:20:21.966980Z","iopub.execute_input":"2023-10-12T10:20:21.967302Z","iopub.status.idle":"2023-10-12T10:24:43.976772Z","shell.execute_reply.started":"2023-10-12T10:20:21.967274Z","shell.execute_reply":"2023-10-12T10:24:43.975761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create table of ids and labels like sample_submission\nsample_data['label'] = np.ravel(np.round(results))","metadata":{"execution":{"iopub.status.busy":"2023-10-12T10:25:02.961961Z","iopub.execute_input":"2023-10-12T10:25:02.962293Z","iopub.status.idle":"2023-10-12T10:25:02.969248Z","shell.execute_reply.started":"2023-10-12T10:25:02.962265Z","shell.execute_reply":"2023-10-12T10:25:02.968366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Print submission table\nsample_data","metadata":{"execution":{"iopub.status.busy":"2023-10-12T10:25:05.941525Z","iopub.execute_input":"2023-10-12T10:25:05.942442Z","iopub.status.idle":"2023-10-12T10:25:05.955491Z","shell.execute_reply.started":"2023-10-12T10:25:05.942399Z","shell.execute_reply":"2023-10-12T10:25:05.954420Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Make submission\nsample_data.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-10-12T10:25:12.766374Z","iopub.execute_input":"2023-10-12T10:25:12.767024Z","iopub.status.idle":"2023-10-12T10:25:12.890542Z","shell.execute_reply.started":"2023-10-12T10:25:12.766994Z","shell.execute_reply":"2023-10-12T10:25:12.889668Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Cleaning\ndel submis_data, submis_dataset, sample_data\ndel get_decoded_image, get_prefetched_data","metadata":{"execution":{"iopub.status.busy":"2023-10-11T18:01:34.090668Z","iopub.execute_input":"2023-10-11T18:01:34.091120Z","iopub.status.idle":"2023-10-11T18:01:34.106293Z","shell.execute_reply.started":"2023-10-11T18:01:34.091084Z","shell.execute_reply":"2023-10-11T18:01:34.105045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 9. Conclusion <a class=\"anchor\" id=\"chapter_9\"></a>","metadata":{}},{"cell_type":"markdown","source":"In this study, we examined the impact of different architectures and hyperparameters on the performance of convolutional neural network (CNN) models. The comparative analysis was based on the problem of identifying metastatic cancer in small image patches taken from larger digital pathology scans.\n\nThe selected methods were:\n* Addition of convolutional and dense layers\n* Change in the width of layers\n* Max pooling introduction\n* Dropout regularization\n\nTo explore the effects of these configurations on the model we created 5 models: 1 base model and 4 updated. The updated models followed the base model architecture and were independent of each other.\n\nThe collected results showed that the 'Base + Additional layers' model emerged as the top performer. Additional layers enhanced pattern recognition, leading to better generalization.<br/>\nThe 'Base + Max pooling' model, while improving scores, slowed training unexpectedly.<br/>\n'Base + Wider layers' overfit the data, capturing noise.<br/>\nThe last model 'Base + Dropout' mitigated overfitting however reduced the model's ability to learn dataset intricacies.\n\nAlthough the increase in the model depth produced the best results, it also highlighted the need for significant enhancements. To supplement the methods discussed earlier, one can investigate the following strategies:\n* data augmentation: rotations, scaling, flipping, and other transformations\n* pretrained models: VGG, ResNet or Inception\n\nThese findings provide brief insights into the interplay between design and model performance. While they underscore the crucial role of thoughtful architecture and hyperparameter tuning, they do not definitively specify which options to exclude. Several other parameters, including learning rate, strides, batch normalization, and others, were not covered in this study. Therefore, additional research on their effects is necessary before establishing an extensive and complex network.","metadata":{}},{"cell_type":"markdown","source":"# 10. References <a class=\"anchor\" id=\"chapter_10\"></a>","metadata":{}},{"cell_type":"markdown","source":"* Will Cukierski. (2018). Histopathologic Cancer Detection. Kaggle. <br/>\nhttps://www.kaggle.com/competitions/histopathologic-cancer-detection/data\n* Better performance with the tf.data API<br/>\nhttps://www.tensorflow.org/guide/data_performance\n* Distributed training with Tensorflow<br/>\nhttps://www.tensorflow.org/guide/distributed_training\n* Tensorflow GPU<br/>\nhttps://www.tensorflow.org/guide/gpu\n* Karen Simonyan and Andrew Zisserman. (2014). Very Deep Convolutional Networks for Large-Scale Image Recognition<br/>\nhttps://arxiv.org/abs/1409.1556\n* Kaiming He, Xiangyu Zhang, Shaoqing Ren, (2015). Jian Sun. Deep Residual Learning for Image Recognition<br/>\nhttps://arxiv.org/abs/1512.03385\n* Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, Andrew Rabinovich. (2014). Going Deeper with Convolutions<br/>\nhttps://arxiv.org/abs/1409.4842","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}