{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# import libraries\n\nimport numpy as np \nimport pandas as pd \nimport random\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom sklearn.utils import shuffle\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import roc_curve, auc, roc_auc_score\nfrom sklearn.metrics import classification_report,confusion_matrix\n\nimport keras\nfrom keras import backend as K \nfrom keras.models import *\nfrom keras.layers import *\nfrom keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras.optimizers import Adam\nfrom keras.callbacks import ReduceLROnPlateau\nfrom tensorflow.keras.applications import ResNet50\nfrom keras.applications.vgg16 import VGG16\n\nimport tensorflow as tf\nimport os\nfrom skimage import io\n\n#import os\n#os.environ['TF_CPP_MIN_LOG_LEVEL'] = '2'","metadata":{"id":"f1ce21bf"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 1: Brief description of the problem and data","metadata":{"id":"f1f3d0d0"}},{"cell_type":"markdown","source":"### 1.1 Data","metadata":{"id":"5aa09f61"}},{"cell_type":"markdown","source":"In this dataset, we are provided with a large number of small pathology images to classify. Files are named with an image id. The train_labels.csv file provides the ground truth for the images in the train folder. We are predicting the labels for the images in the test folder. A positive label indicates that the center 32x32px region of a patch contains at least one pixel of tumor tissue. Tumor tissue in the outer region of the patch does not influence the label. ","metadata":{"id":"58014456"}},{"cell_type":"markdown","source":"The dataset was downloaded from https://www.kaggle.com/competitions/histopathologic-cancer-detection/data","metadata":{"id":"c2e349ab"}},{"cell_type":"markdown","source":"### 1.2 Project Topic","metadata":{"id":"3c8f2b05"}},{"cell_type":"markdown","source":"The goal of this project is to identify metastatic cancer in small image patches taken from larger digital pathology scans. To achieve this goal, I will:  \n+ Inspect, Visualize and Clean the data.  \n+ Build two CNN models.\n+ Run hyperparameter tuning, try different architectures for comparison. \n+ Analysis and Result.\n+ Conclusion.  \n","metadata":{"id":"33821533"}},{"cell_type":"markdown","source":"## Step 2: EDA - Inspect, Visualize and Clean the Data","metadata":{"id":"4f4c1016"}},{"cell_type":"markdown","source":"### 2.1 Inspect the data","metadata":{"id":"5cddd908"}},{"cell_type":"code","source":"#train_path = '../input/histopathologic-cancer-detection/train/'\n#test_path = '../input/histopathologic-cancer-detection/test/'\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_path = '../Week3/train/'\ntest_path = '../Week3/test/'","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# take a look at some first train_label rows\n#train_df = pd.read_csv('../input/histopathologic-cancer-detection/train_labels.csv')\ntrain_df = pd.read_csv('../Week3/train_labels.csv')\ntrain_df.head()\n","metadata":{"id":"7a8c49dc","outputId":"281291a7-ca65-4070-d1ba-5fb739a06fdf"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get a quick description of the data\ntrain_df.describe()\n","metadata":{"id":"950ae572","outputId":"9c559605-7ef7-475b-f709-17315e4a9666"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check null values in data\ntrain_df.isnull().sum()\n","metadata":{"id":"166b436f","outputId":"e79ecaea-66cf-415f-9210-c51dcbb2f1c3"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check for duplicate train_label\ntrain_df.duplicated(keep=False).sum()\n","metadata":{"id":"a1c6e782","outputId":"7af56a19-eb98-43b1-9ab7-c18421b5cd43"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# the structure of data also tells us the number of rows, columns and type of data\ntrain_df.info()\n","metadata":{"id":"b88b0b3a","outputId":"020c9ec2-7fdb-4325-ddb3-c64134b62b30"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get the number of entries in test set\n#test_df = pd.DataFrame({'id':os.listdir(test_path)})\ntest_df = pd.DataFrame({'id':os.listdir('../Week3/test')})\nprint(len(test_df))\n","metadata":{"id":"a96373df","outputId":"d94501e3-b3ea-447f-d3da-640a5d50b979"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# take a look at some first rows of test set\ntest_df.head()\n","metadata":{"id":"b1fe2eb2","outputId":"551a9347-3639-4345-f786-52517211d509"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check null values in test set\ntest_df.isnull().sum()\n","metadata":{"id":"161247a8","outputId":"07af946a-064f-42d7-ca0c-29b901980db6"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check for duplicate in test set\ntest_df.duplicated(keep=False).sum()\n","metadata":{"id":"309791e8","outputId":"019373f4-836a-410c-8078-5ff6dcb4014c"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the output above, we can summarize that:\n\n+ There are 220,025 entries and 2 columns in train_label data.\n+ There is no missing values.   \n+ There is no duplicated entries.  \n+ The id column is object and label column is integer with two values 0 (no_tumor_tissue) and 1 (has_tumor_tissue). \n+ There are 57,458 entries in test data.\n","metadata":{"id":"53706d76"}},{"cell_type":"markdown","source":"### 2.2 Visualize the data","metadata":{"id":"de1b6bdf"}},{"cell_type":"code","source":"# calculate the count of each label\ntrain_df['label'].value_counts()\n","metadata":{"id":"0efdc4ce","outputId":"7dbb80a8-b5f7-4562-f206-da98f2f2fa60"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# calculate the proportion of each label\ntrain_df['label'].value_counts()/len(train_df)*100\n","metadata":{"id":"f89c8ecf","outputId":"7b15ca7d-b000-4cb5-b8f7-c3892a454343"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot the count of each label\nfig, ax = plt.subplots(figsize=(6,6))\nsns.countplot(data=train_df, y='label', ax=ax).set(title='\\nFigure 1. The Count of Each Label\\n')\n\n# plot the proportion of each label\nlabels = train_df['label'].unique().tolist()\ncounts = train_df['label'].value_counts()\nsizes = [counts[v] for v in labels]\nfig1, ax1 = plt.subplots()\nax1.pie(sizes, labels=labels, autopct='%0.2f%%')\nax1.axis('equal')\nplt.title(\"\\nFigure 2. The Proportion of Each Label\\n\")\nplt.tight_layout()\nplt.show()\n","metadata":{"id":"72a6a66b","outputId":"d5473947-f922-4797-8916-13e12a9345ff"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Figure 1 shows the count of each label and figure 2 shows the proportions of each label. Looking at these two figures, we can see that in overall, the number of example for no_tumor_tissue label (0) was larger than the number of example for has_tumor_tissue label (1). And with the proportion of example for no_tumor_tissue label (0) was nearly 60% comparing to 40% of example for has_tumor_tissue (1), I think it will be better if we balance data by reducing the number of samples in label 0 since if one category was severely underrepresentated or, in contrast, overrepresentative in the train data, then it may cause our model to be biased and/or perform poorly on some or all of the test data.","metadata":{"id":"5a23fd48"}},{"cell_type":"code","source":"# Train image visualisations\ndef append_tif(string):\n    return string + \".tif\"\ntrain_df[\"id\"] = train_df[\"id\"].apply(append_tif)\ntrain_df[\"label\"] = train_df[\"label\"].astype(str)\n\nfig, axes = plt.subplots(5, 5, figsize=(15, 15))\nfor i, ax in enumerate(axes.flat):\n    file = str(train_path + train_df.id[i])\n    image = io.imread(file)\n    ax.imshow(image)\n    ax.set(xticks=[], yticks=[], xlabel = train_df.label[i])\nfig.suptitle('\\nFigure 3. Train image Visualizations') \nplt.show() \n","metadata":{"id":"c9930b3d","outputId":"ac67f27d-dceb-46ac-918c-c396ec4c2544"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Test image visualisations\nfig, axes = plt.subplots(5, 5, figsize=(15, 15))\nfor i, ax in enumerate(axes.flat):\n    file = str(test_path + test_df.id[i])\n    image = io.imread(file)\n    ax.imshow(image)\nfig.suptitle('\\nFigure 4. Test image Visualizations') \nplt.show()\n","metadata":{"id":"fdd3b5ff","outputId":"a34ff83d-0026-41fb-e421-864643a28f96"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.3 Clean the data / Data Preprocessing","metadata":{"id":"dcc6c9df"}},{"cell_type":"code","source":"# check the data format\nprint(K.image_data_format()) \n","metadata":{"id":"8796a0ad","outputId":"2f39164d-959e-44f6-f3f7-699ef2b6562b"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# set random state\nRANDOM_STATE = 42","metadata":{"id":"30a69d99"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# set batch size\nBATCH_SIZE = 10\n","metadata":{"id":"a1c474e0"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# set input shape\nimg_width, img_height = 64, 64\ninput_shape = (img_width, img_height, 3)\n","metadata":{"id":"38140e7c"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's balance the data and split train dataset to:\n+ train set: is the set used for training model\n+ validation set: is the set used during the model training to adjust the hyperparameters. (20%)\n","metadata":{"id":"yXo65BtXfG6t"}},{"cell_type":"code","source":"# balance the data\nSAMPLE = 80000\ntrain1 = train_df[train_df[\"label\"] == \"0\"].sample(SAMPLE, random_state=RANDOM_STATE)\ntrain2 = train_df[train_df[\"label\"] == \"1\"].sample(SAMPLE, random_state=RANDOM_STATE)\ntrain_dt = pd.concat([train1, train2], axis=0).reset_index(drop=True)\ntrain_dt[\"label\"].value_counts()\n","metadata":{"id":"ed36c3f4"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# split train dataset to train and validation set\ntrain_data, valid_data = train_test_split(train_dt,                                                      \n                                   random_state=RANDOM_STATE, \n                                   test_size=0.2, \n                                   shuffle=True, stratify=train_dt[\"label\"])\n\n# check value count in train and validation set\nprint(train_data[\"label\"].value_counts())\nprint(valid_data[\"label\"].value_counts())","metadata":{"id":"hM5u9C2psfxR"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Before we can proceed with building the model:  \n\nThe first step to working with neural networks is to normalize the dataset, otherwise, it could take a lot longer for the network to converge on a solution.\n\nThe usual way of normalizing a dataset is to scale the features, and this is done by subtracting the mean from each feature and dividing by the standard deviation. This will put the features on the same scale somewhere between 0 — 1.\n\nAs we are working with 32 x 32 NumPy arrays representing each image and each pixel in the array has an intensity somewhere between 1 — 255, a simpler way of getting all of these images on a scale between 0–1 is to divide each array by 255.","metadata":{"id":"94f372ef"}},{"cell_type":"code","source":"datagen = ImageDataGenerator(featurewise_center=False,  # set input mean to 0 over the dataset                           \n                             zoom_range = 0.2, # Randomly zoom image \n                             rotation_range = 30,  # randomly rotate images in the range (degrees, 0 to 180)\n                             width_shift_range=0.1,  # randomly shift images horizontally (fraction of total width)\n                             height_shift_range=0.1,  # randomly shift images vertically (fraction of total height)\n                             horizontal_flip = True,  # randomly flip images\n                             rescale=1./255)    # multiply the data by the value provided\n                            ","metadata":{"id":"bae32f43"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_generator = datagen.flow_from_dataframe(\n                            dataframe=train_data,\n                            directory=train_path,\n                            x_col=\"id\",\n                            y_col=\"label\",                            \n                            batch_size=BATCH_SIZE,                           \n                            seed=RANDOM_STATE,\n                            class_mode=\"binary\",\n                            target_size=(64,64))  ","metadata":{"id":"f9dab221","outputId":"f8b30f66-2459-48f0-e580-f040dc94fd28"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"validation_generator = datagen.flow_from_dataframe(\n                            dataframe=valid_data,\n                            directory=train_path,\n                            x_col=\"id\",\n                            y_col=\"label\",\n                            batch_size=BATCH_SIZE,\n                            seed=RANDOM_STATE,\n                            class_mode=\"binary\",\n                            target_size=(64,64))","metadata":{"id":"8c628f89","outputId":"631247f4-238f-4f0a-b23f-062618fc2807"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 3: Describe Model Architecture","metadata":{"id":"87a69072"}},{"cell_type":"markdown","source":"### Model 1:\nModel is comprised of:  \n\nA simple CNN model with 3 Convolutional layers followed by max-pooling layers. A dropout layer is added at the final convolutional layer to avoid overfitting. BatchNormalization normalize the activation of the previous layer at each batch. Sigmoid is used as the activation function for the final layer of the binary classifier. Use binary-entropy loss function for our binary-class classification problem. For simplicity, use accuracy as our evaluation metrics to evaluate the model during training and testing.\n  + optimization: Adam\n  + learning rate: 0.0001\n  + hidden layer activations: relu\n  + final layer dropout: 0.4\n  + final layer activation: sigmoid because of the binary classification\n","metadata":{"id":"1f03058e"}},{"cell_type":"code","source":"model = Sequential()\n# first convolutional layer\nmodel.add(Conv2D(32, (3, 3), input_shape=input_shape))\nmodel.add(BatchNormalization())\nmodel.add(Activation('relu'))\nmodel.add(MaxPooling2D(pool_size=(2, 2)))\n\n# second convolutional layer\nmodel.add(Conv2D(64, (3, 3)))\nmodel.add(BatchNormalization())\nmodel.add(Activation('relu'))\nmodel.add(MaxPooling2D(pool_size=(2, 2)))\n  \n# third convolutional layer\nmodel.add(Conv2D(128, (3, 3)))\nmodel.add(BatchNormalization())\nmodel.add(Activation('relu'))\nmodel.add(MaxPooling2D(pool_size=(2, 2)))\n\nmodel.add(Flatten())\nmodel.add(Dense(256))\nmodel.add(BatchNormalization())\nmodel.add(Activation('relu'))\nmodel.add(Dropout(0.4))\n\n# Out layer\nmodel.add(Dense(1))\nmodel.add(Activation('sigmoid'))\n\nmodel.summary()","metadata":{"id":"9c6bd4c7","outputId":"9d5eb6dd-ed07-46ee-978f-8c3ec1343d2a"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's compile the model now using Adam as our optimizer and binary crossentropy as the loss function. We are using a lower learning rate of 0.0001 for a smoother curve.","metadata":{"id":"cd1c8123"}},{"cell_type":"code","source":"# compile the model\nopt = Adam(learning_rate=0.0001)\nmodel.compile(loss='binary_crossentropy',\n              optimizer=opt,\n              metrics=['accuracy'])\n","metadata":{"id":"57bb6174"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# let’s train our model for 20 epochs\nSTEP_SIZE_TRAIN=train_generator.n//train_generator.batch_size\nSTEP_SIZE_VALID=validation_generator.n//validation_generator.batch_size\nhistory = model.fit(train_generator,\n                    epochs = 20 , \n                    steps_per_epoch=STEP_SIZE_TRAIN,\n                    validation_data = validation_generator,\n                    validation_steps=STEP_SIZE_VALID)\n            ","metadata":{"id":"dd2acf68","outputId":"68e39b5d-85bb-4416-f791-f10ddbc9fa3a"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.save(\"../Week3/my_model1\")\n","metadata":{"id":"0h-kHUxDrIi0","outputId":"b6fa9ac3-8e5a-4e1b-b6ed-2f1c1cb6dd6a"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"acc = history.history['accuracy']\nval_acc = history.history['val_accuracy']\nloss = history.history['loss']\nval_loss = history.history['val_loss']\n\nepochs_range = range(20)\n\nplt.figure(figsize=(15, 15))\nplt.subplot(2, 2, 1)\nplt.plot(epochs_range, acc, label='Training Accuracy')\nplt.plot(epochs_range, val_acc, label='Validation Accuracy')\nplt.legend(loc='lower right')\nplt.title('Training and Validation Accuracy')\n\nplt.subplot(2, 2, 2)\nplt.plot(epochs_range, loss, label='Training Loss')\nplt.plot(epochs_range, val_loss, label='Validation Loss')\nplt.legend(loc='upper right')\nplt.title('Training and Validation Loss')\nplt.show()","metadata":{"id":"31ed8113","outputId":"785aa44e-e843-461d-af87-b6450329c874"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Model 2","metadata":{}},{"cell_type":"markdown","source":"Next, let's use Earlystopping \nto avoid overfitting by terminating the process early. Since the goal of a training is to minimize the loss. With this, we can set up the metric as:  \n    + monitor : val_loss, value being monitored.  \n    + mode: min, training will stop when the quantity monitored has stopped decreasing.   \n    + patience: 3, number of epochs with no improvement after which training will be stopped.  \n    \nMoreover, let's use Reduce learning rate when a metric has stopped improving. Models often benefit from reducing the learning rate by a factor of 2-10 once learning stagnates. This callback monitors a quantity and if no improvement is seen for a 'patience' number of epochs, the learning rate is reduced.  \n    + factor: factor by which the learning rate will be reduced (new_learning_rate = learning_rate * factor).  \n    + min_lr: lower bound on the learning rate.  \n\nA model.fit() training loop will check at end of every epoch whether the loss is no longer decreasing, considering the min_delta and patience if applicable. Once it's found no longer decreasing, model.stop_training is marked True and the training terminates.","metadata":{}},{"cell_type":"code","source":"# define an Earlystopping\ncheckpoint_filepath = '../Week3/checkpoint'\nmp= tf.keras.callbacks.ModelCheckpoint(filepath=checkpoint_filepath, save_weights_only=True,\n                               verbose=1, save_best_only=True)\nes= tf.keras.callbacks.EarlyStopping(monitor='val_loss', mode='min', patience=3, verbose=1)\nreduce_lr = ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=3, \n                                   verbose=1, mode='min', min_lr=0.00001)\ncallback=[es, mp, reduce_lr]\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_model = Sequential()\n# first convolutional layer\nnew_model.add(Conv2D(32, (3, 3), input_shape=input_shape))\nnew_model.add(BatchNormalization())\nnew_model.add(Activation('relu'))\nnew_model.add(MaxPooling2D(pool_size=(2, 2)))\n\n# second convolutional layer\nnew_model.add(Conv2D(64, (3, 3)))\nnew_model.add(BatchNormalization())\nnew_model.add(Activation('relu'))\nnew_model.add(MaxPooling2D(pool_size=(2, 2)))\n  \n# third convolutional layer\nnew_model.add(Conv2D(128, (3, 3)))\nnew_model.add(BatchNormalization())\nnew_model.add(Activation('relu'))\nnew_model.add(MaxPooling2D(pool_size=(2, 2)))\n\nnew_model.add(Flatten())\nnew_model.add(Dense(256))\nnew_model.add(BatchNormalization())\nnew_model.add(Activation('relu'))\nnew_model.add(Dropout(0.4))\n\n# Out layer\nnew_model.add(Dense(1))\nnew_model.add(Activation('sigmoid'))\n\nnew_model.summary()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# compile new model\nnew_model.compile(loss='binary_crossentropy',\n              optimizer=Adam(learning_rate=0.0001),\n              metrics=['accuracy'])\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# let’s train our model for 20 epochs\nnew_history = new_model.fit(train_generator,\n                    epochs = 20 , \n                    steps_per_epoch=STEP_SIZE_TRAIN,\n                    validation_data = validation_generator,\n                    validation_steps=STEP_SIZE_VALID,\n                    callbacks=callback)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_model.save(\"../Week3/my_model2\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_acc = new_history.history['accuracy']\nnew_val_acc = new_history.history['val_accuracy']\nnew_loss = new_history.history['loss']\nnew_val_loss = new_history.history['val_loss']\nepochs_range = range(7)\n\nplt.figure(figsize=(15, 15))\nplt.subplot(2, 2, 1)\nplt.plot(epochs_range, new_acc, label='Training Accuracy')\nplt.plot(epochs_range, new_val_acc, label='Validation Accuracy')\nplt.legend(loc='lower right')\nplt.title('Training and Validation Accuracy')\n\nplt.subplot(2, 2, 2)\nplt.plot(epochs_range, new_loss, label='Training Loss')\nplt.plot(epochs_range, new_val_loss, label='Validation Loss')\nplt.legend(loc='upper right')\nplt.title('Training and Validation Loss')\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 4: Results and Analysis","metadata":{"id":"32db4b42"}},{"cell_type":"code","source":"# check what index keras has internally assigned to each label\nprint(validation_generator.class_indices)\n","metadata":{"id":"yBb5flkfwp2E"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the true labels\ny_true = validation_generator.classes\n","metadata":{"id":"A_g2g7b7wp7-"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since our dataset is not heavily imbalanced then I would like to use ROC AUC as an evaluation metric for binary classification problem. The Receiver Operator Characteristic (ROC) is a probability curve that plots the TPR against FPR at various threshold values and essentially separates the ‘signal’ from the ‘noise.’ In other words, it shows the performance of a classification model at all classification thresholds. The Area Under the Curve (AUC) is the measure of the ability of a binary classifier to distinguish between classes and is used as a summary of the ROC curve.\n\nThe higher the AUC, the better the model’s performance at distinguishing between the positive and negative classes.","metadata":{"id":"zfd8FnW0SwHI"}},{"cell_type":"markdown","source":"### Model 1","metadata":{}},{"cell_type":"code","source":"val_loss1, val_acc1 = model.evaluate(validation_generator)\nprint('val_loss_model1:', val_loss1)\nprint('val_acc_model1:', val_acc1)\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# predict validation dataset\npredictions1 = model.predict(validation_generator, verbose=1)\npredictions1\n","metadata":{"id":"tB7geNTkwolA"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# calculate auc_score\nfpr1, tpr1, thresholds1 = roc_curve(y_true, predictions1, pos_label=1)\nauc_score1 = auc(fpr1, tpr1)\nauc_score1","metadata":{"id":"rkI5BmHEVbXt"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot([0,1], [0,1], linestyle='--', color='blue')\nplt.plot(fpr1, tpr1, label='area = {:.2f}'.format(auc_score1))\n# axis labels\nplt.xlabel('False Positive Rate')\nplt.ylabel('True Positive Rate')\n# show the legend\nplt.legend()\n# show the plot\nplt.show()\n","metadata":{"id":"V9noIh-nZP_9"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can print out the classification report to see the precision and accuracy.","metadata":{"id":"OADz5A7LyaBZ"}},{"cell_type":"code","source":"# Get the prediction binary\ny_pred1 = np.where(predictions1 > 0.5, 1, 0)\n","metadata":{"id":"VczaKo7AXy6n"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# print out the classification report\nprint(classification_report(y_true, y_pred1, target_names = ['no_tumor_tissue (Class 0)','has_tumor_tissue (Class 1)']))\n","metadata":{"id":"M2JegTNBybmm"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# print out the confusion matrix\ncm1 = confusion_matrix(y_true, y_pred1)\nsns.heatmap(cm1, annot=True, fmt=\".0f\")\n","metadata":{"id":"A_8BO-4fyrB5"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Model 2","metadata":{}},{"cell_type":"code","source":"# the best epoch will be used.\nnew_model.load_weights('../Week3/checkpoint')\nval_loss2, val_acc2 = new_model.evaluate(validation_generator)\nprint('val_loss_model2:', val_loss2)\nprint('val_acc_model2:', val_acc2)\n","metadata":{"id":"vClkxbOCRvHT"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# predict validation dataset\npredictions2 = new_model.predict(validation_generator)\npredictions2\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# calculate auc_score\nfpr2, tpr2, thresholds2 = roc_curve(y_true, predictions2, pos_label=1)\nauc_score2 = auc(fpr2, tpr2)\nauc_score2\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot([0,1], [0,1], linestyle='--', color='blue')\nplt.plot(fpr2, tpr2, label='area = {:.2f}'.format(auc_score2))\n# axis labels\nplt.xlabel('False Positive Rate')\nplt.ylabel('True Positive Rate')\n# show the legend\nplt.legend()\n# show the plot\nplt.show()\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the prediction binary\ny_pred2 = np.where(predictions2 > 0.5, 1, 0)\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# print out the classification report\nprint(classification_report(y_true, y_pred2, target_names = ['no_tumor_tissue (Class 0)','has_tumor_tissue (Class 1)']))\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# print out the confusion matrix\ncm2 = confusion_matrix(y_true, y_pred2)\nsns.heatmap(cm2, annot=True, fmt=\".0f\")\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Predict test data and print out the submission","metadata":{"id":"ZobWffgJE5Ob"}},{"cell_type":"code","source":"test_datagen = ImageDataGenerator(rescale=1./255)","metadata":{"id":"79ca8b1b"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_generator = test_datagen.flow_from_dataframe(\n                            dataframe=test_df,\n                            directory=test_path,\n                            x_col=\"id\",\n                            y_col=None,\n                            batch_size=BATCH_SIZE,\n                            shuffle=False,\n                            seed=RANDOM_STATE,\n                            class_mode=None,\n                            target_size=(64,64))","metadata":{"id":"3495c35a"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# predict validation dataset\nt_predictions = new_model.predict(test_generator, verbose=1)\nt_predictions\n","metadata":{"id":"d06857a9"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the new prediction binary\ntest_pred = np.where(t_predictions > 0.5, 1, 0)\n","metadata":{"id":"SAe3-Wkh5IH4"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# create submission dataframe\ntest_predictions = np.transpose(test_pred)[0]\nsubmission = pd.DataFrame()\nsubmission['id'] = test_df['id'].apply(lambda x: x.split('.')[0])\nsubmission['label'] = test_predictions\nsubmission.head()\n","metadata":{"id":"f064e9c4"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# view test prediction counts\nsubmission['label'].value_counts()","metadata":{"id":"7d29d43a"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot the count of each label\nfig, ax = plt.subplots(figsize=(6,6))\nsns.countplot(data=submission, y='label', ax=ax).set(title='\\nFigure 5. The Count of Each Label\\n')\n\n# plot the proportion of each label\nlabels = submission['label'].unique().tolist()\ncounts = submission['label'].value_counts()\nsizes = [counts[v] for v in labels]\nfig1, ax1 = plt.subplots()\nax1.pie(sizes, labels=labels, autopct='%0.2f%%')\nax1.axis('equal')\nplt.title(\"\\nFigure 6. The Proportion of Each Label\\n\")\nplt.tight_layout()\nplt.show()","metadata":{"id":"1bMCLjlfV3xY"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# convert to csv to submit to competition\n#submission.to_csv('submission.csv', index=False)\nsubmission.to_csv('../Week3/submission.csv', index=False)\n","metadata":{"id":"asbRyg2wXGLT"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 5: Conclusion","metadata":{"id":"5b773b0b"}},{"cell_type":"code","source":"compare_table = pd.DataFrame({\"Model\": [\"Model1\", \"Model2\"],\n                        \"val_acc\": [round(val_acc1, 3), round(val_acc2, 3)],\n                        \"val_loss\": [round(val_loss1, 3), round(val_loss2, 3)],\n                        \"AUC\": [round(auc_score1, 2), round(auc_score2, 2)]})\ncompare_table","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Model 1 has the higher validation accuracy and lower validation loss compare to model2, However AUC score of two models are just the same although it took more time to run model1 than model2 because model2 used Earlystopping and Reduce Learning Rate to optimize the model. I think these two models might be overfitting, so besides these two models, I tried building some models with different learning rate and different values of dense, drop out. For example, when I chose a learning rate like 0.00001, I observed that the model just ran and ended up with an early stop at epoch 4 because of the learning rate was too small, so it was stuck at epoch 4. But due to the limitation of time and memory, I could just build these simple CNN models and get AUC of 0.5. Hence, I believe that there are many ways could improve the result such as run this model by increasing the number of epochs or trying to test with many different parameters might get better results. ","metadata":{}},{"cell_type":"code","source":"","metadata":{"id":"f31ce78a"},"execution_count":null,"outputs":[]}]}