{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Histopathologic Cancer Detection Resnet 50 Model\n\nThe goal of this project is to train a Convolutional Neural Network, specifically Resnet with 50 layers, in order to identify if the center of the image of a body tissue collected from a pathology scan contains cancerous tissues or not. This is a binary classification problem, meaning that each training example is labeled as either being cancerous(1) or not cancerous(0). \n\nCancerous in this case means that at least one pixel of the image contains tumorous tissue.\n\n\nThe dataset is based on PatchCamelyon (PCam) benchmark dataset (the original PCam dataset contains duplicate images due to its probabilistic sampling, however, the version presented on Kaggle does not contain duplicates). The source of this data is this Kaggle competition: https://www.kaggle.com/competitions/histopathologic-cancer-detection\n\nThe dataset consists of 220025 training images that are all labeled with if they are cancerous or not. Additionally, there are 57458 unlabeled test images. The goal of this project for the built neural network to predict reasonably well the labels on this test set. The test data set does not come with test labels, so the test performance will be gotten by making submissions to the kaggle competition.\n\nIt is to be noted that while we can make submissions the public leaderboard for this competition has been freezed. We can still submit to get the AUC which is the performance metric for this project but we will not get into public leaderboard.","metadata":{}},{"cell_type":"markdown","source":"## 1. Packages","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nimport numpy as np\nimport tensorflow.keras.layers as tfl\nimport matplotlib.pyplot as plt\nimport pandas as pd\nfrom tensorflow.keras.models import Model\nfrom tensorflow.keras.initializers import random_uniform, glorot_uniform\n\nimport os\nimport shutil\nimport json\n\nfrom PIL import Image","metadata":{"execution":{"iopub.status.busy":"2023-08-20T07:50:59.469749Z","iopub.execute_input":"2023-08-20T07:50:59.470144Z","iopub.status.idle":"2023-08-20T07:51:08.787848Z","shell.execute_reply.started":"2023-08-20T07:50:59.470114Z","shell.execute_reply":"2023-08-20T07:51:08.786865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Creating a function to streamline the Train data set   \ndef train_img_path(id_str):\n    return os.path.join(r\"/kaggle/input/histopathologic-cancer-detection/train\", f\"{id_str}.tif\")","metadata":{"execution":{"iopub.status.busy":"2023-08-20T07:51:18.382414Z","iopub.execute_input":"2023-08-20T07:51:18.382801Z","iopub.status.idle":"2023-08-20T07:51:18.387952Z","shell.execute_reply.started":"2023-08-20T07:51:18.382771Z","shell.execute_reply":"2023-08-20T07:51:18.386994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2. Exploratory Data Analysis","metadata":{}},{"cell_type":"markdown","source":"In this section we will load the data and do cleaning if needed and then perform basic EDA.\n\nHowever it is to be noted that this dataset has already been heavily tuned and cleaned by Kaggle. There is no need for data cleaning but we can still perform basic EDA.","metadata":{}},{"cell_type":"code","source":"example_path = \"/kaggle/input/histopathologic-cancer-detection/train/f38a6374c348f90b587e046aac6079959adf3835.tif\"\nexample_img = Image.open(example_path)\nexample_array = np.array(example_img)\nprint(f\"Image Shape = {example_array.shape}\")\nplt.imshow(example_img)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-20T07:51:19.535618Z","iopub.execute_input":"2023-08-20T07:51:19.536056Z","iopub.status.idle":"2023-08-20T07:51:19.895520Z","shell.execute_reply.started":"2023-08-20T07:51:19.536019Z","shell.execute_reply":"2023-08-20T07:51:19.894330Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have the images with size of (96,96,3) that is height is 96, width is 96, and we have 3 channels which are basic RGB channels. All the images has same shape in both train and test set.\n\nNow we will create a dataframe to store id, label, and filename which will be used in training generator for fitting the data.","metadata":{}},{"cell_type":"code","source":"train_labels_df = pd.read_csv('/kaggle/input/histopathologic-cancer-detection/train_labels.csv')\ntrain_labels_df[\"filename\"] = train_labels_df[\"id\"].apply(train_img_path)\ntrain_labels_df[\"label\"] = train_labels_df[\"label\"].astype(str)\ntrain_labels_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-20T07:51:22.736993Z","iopub.execute_input":"2023-08-20T07:51:22.737835Z","iopub.status.idle":"2023-08-20T07:51:24.198947Z","shell.execute_reply.started":"2023-08-20T07:51:22.737786Z","shell.execute_reply":"2023-08-20T07:51:24.197482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels_df.shape","metadata":{"execution":{"iopub.status.busy":"2023-08-20T07:51:24.264730Z","iopub.execute_input":"2023-08-20T07:51:24.265115Z","iopub.status.idle":"2023-08-20T07:51:24.273420Z","shell.execute_reply.started":"2023-08-20T07:51:24.265085Z","shell.execute_reply":"2023-08-20T07:51:24.272130Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"set(train_labels_df['label'])","metadata":{"execution":{"iopub.status.busy":"2023-08-20T07:51:27.531477Z","iopub.execute_input":"2023-08-20T07:51:27.532230Z","iopub.status.idle":"2023-08-20T07:51:27.575239Z","shell.execute_reply.started":"2023-08-20T07:51:27.532191Z","shell.execute_reply":"2023-08-20T07:51:27.573993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have 220,025 images in the train data set with 2 unique labels. 0 for not cancerous and 1 for cancerous tissues.","metadata":{}},{"cell_type":"code","source":"train_labels_df['label'].value_counts(normalize = True)","metadata":{"execution":{"iopub.status.busy":"2023-08-20T07:51:29.314792Z","iopub.execute_input":"2023-08-20T07:51:29.315923Z","iopub.status.idle":"2023-08-20T07:51:29.366552Z","shell.execute_reply.started":"2023-08-20T07:51:29.315864Z","shell.execute_reply":"2023-08-20T07:51:29.365593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Around 40.5% of the train data set are cancerous tissues and 59.5% are tissues without cancerous tumor. Dataset is slightly unbalanced, but it is not to a limit that we may need to fine tune the data set.","metadata":{}},{"cell_type":"code","source":"sample_data = np.empty((100, 96, 96, 3), dtype=np.uint8)\nsample_labels = np.empty(100, dtype=np.int8)\nfor i in range(len(train_labels_df))[:100]:\n    img_path = train_img_path(train_labels_df['id'][i])\n    img = Image.open(img_path)\n    sample_data[i] = np.array(img)\n    sample_labels[i] = train_labels_df['label'][i]","metadata":{"execution":{"iopub.status.busy":"2023-08-20T07:51:31.191998Z","iopub.execute_input":"2023-08-20T07:51:31.192389Z","iopub.status.idle":"2023-08-20T07:51:31.962152Z","shell.execute_reply.started":"2023-08-20T07:51:31.192357Z","shell.execute_reply":"2023-08-20T07:51:31.960689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Non-Cancerous Images\")\n\nselected_images = np.random.choice(sample_data[sample_labels == 0].shape[0], 12, replace=False)\ngrid_size = int(np.ceil(np.sqrt(12)))\n\nfig, axs = plt.subplots(grid_size, grid_size, figsize=(5, 5))\n\nfor i, ax in enumerate(axs.flatten()):\n    if i < 12:\n        ax.imshow(sample_data[sample_labels == 0][selected_images[i]])\n        ax.axis('off') \n    else:\n        fig.delaxes(ax) \n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-20T07:51:38.143324Z","iopub.execute_input":"2023-08-20T07:51:38.144538Z","iopub.status.idle":"2023-08-20T07:51:38.840066Z","shell.execute_reply.started":"2023-08-20T07:51:38.144496Z","shell.execute_reply":"2023-08-20T07:51:38.836620Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Cancerous Images\")\n\nselected_images = np.random.choice(sample_data[sample_labels == 1].shape[0], 12, replace=False)\ngrid_size = int(np.ceil(np.sqrt(12)))\n\nfig, axs = plt.subplots(grid_size, grid_size, figsize=(5, 5))\n\nfor i, ax in enumerate(axs.flatten()):\n    if i < 12:\n        ax.imshow(sample_data[sample_labels == 1][selected_images[i]])\n        ax.axis('off') \n    else:\n        fig.delaxes(ax) \n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-20T07:52:01.993072Z","iopub.execute_input":"2023-08-20T07:52:01.994215Z","iopub.status.idle":"2023-08-20T07:52:02.680300Z","shell.execute_reply.started":"2023-08-20T07:52:01.994176Z","shell.execute_reply":"2023-08-20T07:52:02.679371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above random images of the cancerous and non canceorus tissues we can see that tumor is not recognizable to an untrained human, for general populance both are not that different.","metadata":{}},{"cell_type":"markdown","source":"## 3. Model Designing","metadata":{}},{"cell_type":"markdown","source":"In this part we will perform various steps required to properly create the Resnet 50 model. We will load the data from the disk, specify train and validation data generators while creating test generator for the final submission.\n\nWe will then create a modified Resnet 50 suited for this task.","metadata":{}},{"cell_type":"code","source":"test_path = \"/kaggle/input/histopathologic-cancer-detection/test\"\ntest_ids = [filename[:-4] for filename in os.listdir(test_path)]\ntest_filenames = [os.path.join(test_path, filename) for filename in os.listdir(test_path)]\ntest_df = pd.DataFrame()\ntest_df[\"id\"] = test_ids\ntest_df[\"filename\"] = test_filenames","metadata":{"execution":{"iopub.status.busy":"2023-08-20T07:51:46.159570Z","iopub.execute_input":"2023-08-20T07:51:46.160300Z","iopub.status.idle":"2023-08-20T07:51:47.779296Z","shell.execute_reply.started":"2023-08-20T07:51:46.160260Z","shell.execute_reply":"2023-08-20T07:51:47.778101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"datagen = tf.keras.preprocessing.image.ImageDataGenerator(rescale = 1/255, validation_split = 0.2)","metadata":{"execution":{"iopub.status.busy":"2023-08-20T04:25:25.910666Z","iopub.execute_input":"2023-08-20T04:25:25.911422Z","iopub.status.idle":"2023-08-20T04:25:25.917238Z","shell.execute_reply.started":"2023-08-20T04:25:25.911386Z","shell.execute_reply":"2023-08-20T04:25:25.915676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_generator = datagen.flow_from_dataframe(\n    shuffle = True,\n    dataframe = train_labels_df,\n    x_col = \"filename\",\n    y_col = \"label\",\n    target_size = (96, 96),\n    color_mode = \"rgb\",\n    batch_size = 32,\n    class_mode = \"binary\",\n    subset = \"training\",\n    validate_filenames = False,\n    seed = 10\n)\n\nvalidation_generator = datagen.flow_from_dataframe(\n    shuffle = True,\n    dataframe=train_labels_df,\n    x_col = \"filename\",\n    y_col = \"label\",\n    target_size=(96, 96),\n    color_mode = \"rgb\",\n    batch_size = 32,\n    class_mode = \"binary\",\n    subset = \"validation\",\n    validate_filenames = False,\n    seed = 10\n)","metadata":{"execution":{"iopub.status.busy":"2023-08-20T04:25:31.588286Z","iopub.execute_input":"2023-08-20T04:25:31.588673Z","iopub.status.idle":"2023-08-20T04:25:33.283073Z","shell.execute_reply.started":"2023-08-20T04:25:31.588640Z","shell.execute_reply":"2023-08-20T04:25:33.281927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_generator = datagen.flow_from_dataframe(\n    dataframe = test_df,\n    x_col = \"filename\",\n    y_col = None,\n    target_size = (96, 96),\n    color_mode = \"rgb\",\n    batch_size = 64,\n    shuffle = False,\n    class_mode = None,\n    validate_filenames = False,\n    seed = 10\n)","metadata":{"execution":{"iopub.status.busy":"2023-08-20T07:01:08.127192Z","iopub.execute_input":"2023-08-20T07:01:08.128156Z","iopub.status.idle":"2023-08-20T07:01:08.256983Z","shell.execute_reply.started":"2023-08-20T07:01:08.128120Z","shell.execute_reply":"2023-08-20T07:01:08.255875Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_steps = 176020//32  # 8000 images for training\nval_steps = 44005//32  ","metadata":{"execution":{"iopub.status.busy":"2023-08-20T04:25:48.040623Z","iopub.execute_input":"2023-08-20T04:25:48.041351Z","iopub.status.idle":"2023-08-20T04:25:48.046361Z","shell.execute_reply.started":"2023-08-20T04:25:48.041315Z","shell.execute_reply":"2023-08-20T04:25:48.045073Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we will create Resnet50 model.\n\nIt is to be noted that we can actually use transfer learning and import the model and its trained parameters based on imagenet from the tensorflow API. But for this project we will create a modified model.\n\nThere will be two types of blocks, identity blocks where input and output dimension remains the same, and convolutional blocks where input and output dimensions are allowed to be changed. In both blocks we will use strong skip connections.\n\nAfter convolutional we will use fully collected layers with output being a dense Sigmoid unit.\n\nIn the end we will have a Resnet model with 50 layers.","metadata":{}},{"cell_type":"code","source":"def identity_block(X, f, filters, training=True, initializer = random_uniform):\n    \"\"\"\n    \n    Arguments:\n    X -- input tensor of shape (m, n_H_prev, n_W_prev, n_C_prev)\n    f -- integer, specifying the shape of the middle CONV's window for the main path\n    filters -- python list of integers, defining the number of filters in the CONV layers of the main path\n    training -- True: Behave in training mode\n                False: Behave in inference mode\n    initializer -- to set up the initial weights of a layer. Equals to random uniform initializer\n    \n    Returns:\n    X -- output of the identity block, tensor of shape (m, n_H, n_W, n_C)\n    \"\"\"\n    \n    # Filters\n    F1, F2, F3 = filters\n    \n    # Save the input value. You'll need this later to add back to the main path. \n    X_shortcut = X\n    \n    # First component of main path\n    X = tfl.Conv2D(filters = F1, kernel_size = 1, strides = (1,1), padding = 'valid', kernel_initializer = initializer(seed=0))(X)\n    X = tfl.BatchNormalization(axis = 3)(X, training = training) # Default axis\n    X = tfl.Activation('relu')(X)\n    \n    ## Set the padding = 'same'\n    X = tfl.Conv2D(filters = F2, kernel_size = f, strides = 1, padding = 'same', kernel_initializer = initializer(seed=0))(X)\n    X = tfl.BatchNormalization(axis = 3)(X, training = training)\n    X = tfl.Activation('relu')(X)\n\n\n    ## Set the padding = 'valid'\n    X = tfl.Conv2D(filters = F3, kernel_size = 1, strides = 1, padding = 'valid', kernel_initializer = initializer(seed = 0))(X)\n    X = tfl.BatchNormalization(axis = 3)(X, training = training) \n    \n    ## Final step: Add shortcut value to main path, and pass it through a RELU activation (≈2 lines)\n    X = tfl.Add()([X, X_shortcut])\n    X = tfl.Activation('relu')(X)\n\n    return X\n","metadata":{"execution":{"iopub.status.busy":"2023-08-20T04:25:52.778745Z","iopub.execute_input":"2023-08-20T04:25:52.779148Z","iopub.status.idle":"2023-08-20T04:25:52.790921Z","shell.execute_reply.started":"2023-08-20T04:25:52.779116Z","shell.execute_reply":"2023-08-20T04:25:52.789753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def convolutional_block(X, f, filters, s = 2, training=True, initializer = glorot_uniform):\n    \"\"\"\n    Implementation of the convolutional block\n    \n    Arguments:\n    X -- input tensor of shape (m, n_H_prev, n_W_prev, n_C_prev)\n    f -- integer, specifying the shape of the middle CONV's window for the main path\n    filters -- python list of integers, defining the number of filters in the CONV layers of the main path\n    s -- Integer, specifying the stride to be used\n    training -- True: Behave in training mode\n                False: Behave in inference mode\n    initializer -- to set up the initial weights of a layer. Equals to Glorot uniform initializer, \n                   also called Xavier uniform initializer.\n    \n    Returns:\n    X -- output of the convolutional block, tensor of shape (m, n_H, n_W, n_C)\n    \"\"\"\n    \n    # Filters\n    F1, F2, F3 = filters\n    \n    # Save the input value\n    X_shortcut = X\n\n    X = tfl.Conv2D(filters = F1, kernel_size = 1, strides = (s, s), padding='valid', kernel_initializer = initializer(seed=0))(X)\n    X = tfl.BatchNormalization(axis = 3)(X, training=training)\n    X = tfl.Activation('relu')(X)\n    \n    X = tfl.Conv2D(filters = F2, kernel_size = (f,f), strides = 1, padding='same', kernel_initializer = initializer(seed=0))(X) \n    X = tfl.BatchNormalization(axis = 3)(X, training=training)\n    X = tfl.Activation('relu')(X) \n\n    X = tfl.Conv2D(filters = F3, kernel_size = 1, strides = 1, padding='valid', kernel_initializer = initializer(seed=0))(X)\n    X = tfl.BatchNormalization(axis = 3)(X, training=training) \n\n    X_shortcut = tfl.Conv2D(filters = F3, kernel_size = (1,1), strides = (s, s), padding='valid', kernel_initializer = initializer(seed=0))(X_shortcut)\n    X_shortcut = tfl.BatchNormalization(axis = 3)(X_shortcut, training=training)\n    \n    X = tfl.Add()([X, X_shortcut])\n    X = tfl.Activation('relu')(X)\n    \n    return X","metadata":{"execution":{"iopub.status.busy":"2023-08-20T04:26:06.045616Z","iopub.execute_input":"2023-08-20T04:26:06.046543Z","iopub.status.idle":"2023-08-20T04:26:06.058447Z","shell.execute_reply.started":"2023-08-20T04:26:06.046504Z","shell.execute_reply":"2023-08-20T04:26:06.057260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def ResNet50(input_shape = (96, 96, 3)):\n    \"\"\"\n    Stage-wise implementation of the architecture of the popular ResNet50:\n    CONV2D -> BATCHNORM -> RELU -> MAXPOOL -> CONVBLOCK -> IDBLOCK*2 -> CONVBLOCK -> IDBLOCK*3\n    -> CONVBLOCK -> IDBLOCK*5 -> CONVBLOCK -> IDBLOCK*2 -> AVGPOOL -> FLATTEN -> DENSE \n\n    Arguments:\n    input_shape -- shape of the images of the dataset\n\n    Returns:\n    model -- a Model() instance in Keras\n    \"\"\"\n    \n    # Define the input as a tensor with shape input_shape\n    X_input = tfl.Input(input_shape)\n\n    \n    # Zero-Padding\n    X = tfl.ZeroPadding2D((3, 3))(X_input)\n    \n    # Stage 1\n    X = tfl.Conv2D(96, (7, 7), strides = (2, 2), kernel_initializer = glorot_uniform(seed=0))(X)\n    X = tfl.BatchNormalization(axis = 3)(X)\n    X = tfl.Activation('relu')(X)\n    X = tfl.MaxPooling2D((3, 3), strides=(2, 2))(X)\n\n    # Stage 2\n    X = convolutional_block(X, f = 3, filters = [64, 64, 256], s = 1)\n    X = identity_block(X, 3, [64, 64, 256])\n    X = identity_block(X, 3, [64, 64, 256])\n\n    ### START CODE HERE\n    \n    # Use the instructions above in order to implement all of the Stages below\n    # Make sure you don't miss adding any required parameter\n    \n    ## Stage 3 (≈4 lines)\n    # `convolutional_block` with correct values of `f`, `filters` and `s` for this stage\n    X = convolutional_block(X, f = 3, filters = [128, 128, 512], s = 2)\n    \n    # the 3 `identity_block` with correct values of `f` and `filters` for this stage\n    X = identity_block(X, 3, [128, 128, 512])\n    X = identity_block(X, 3, [128, 128, 512])\n    X = identity_block(X, 3, [128, 128, 512])\n\n    # Stage 4 (≈6 lines)\n    # add `convolutional_block` with correct values of `f`, `filters` and `s` for this stage\n    X = convolutional_block(X, f = 3, filters = [256, 256, 1024], s = 2)\n    \n    # the 5 `identity_block` with correct values of `f` and `filters` for this stage\n    X = identity_block(X, 3, [256, 256, 1024])\n    X = identity_block(X, 3, [256, 256, 1024])\n    X = identity_block(X, 3, [256, 256, 1024])\n    X = identity_block(X, 3,[256, 256, 1024])\n    X = identity_block(X, 3, [256, 256, 1024])\n\n    # Stage 5 (≈3 lines)\n    # add `convolutional_block` with correct values of `f`, `filters` and `s` for this stage\n    X = convolutional_block(X, f = 3, filters = [512, 512, 2048], s = 2)\n    \n    # the 2 `identity_block` with correct values of `f` and `filters` for this stage\n    X = identity_block(X, 3, [512, 512, 2048])\n    X = identity_block(X, 3, [512, 512, 2048])\n\n    # AVGPOOL (≈1 line). Use \"X = AveragePooling2D()(X)\"\n    X = tfl.AveragePooling2D(pool_size = (2,2))(X)\n    \n\n    # output layer\n    X = tfl.Flatten()(X)\n    X = tfl.Dense(1, kernel_initializer = glorot_uniform(seed=0))(X)\n    \n    \n    # Create model\n    model = Model(inputs = X_input, outputs = X)\n\n    return model","metadata":{"execution":{"iopub.status.busy":"2023-08-20T04:26:46.577550Z","iopub.execute_input":"2023-08-20T04:26:46.577971Z","iopub.status.idle":"2023-08-20T04:26:46.599102Z","shell.execute_reply.started":"2023-08-20T04:26:46.577938Z","shell.execute_reply":"2023-08-20T04:26:46.598019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = ResNet50(input_shape = (96, 96, 3))\nprint(model.summary())","metadata":{"execution":{"iopub.status.busy":"2023-08-20T04:26:51.189298Z","iopub.execute_input":"2023-08-20T04:26:51.189699Z","iopub.status.idle":"2023-08-20T04:26:56.404074Z","shell.execute_reply.started":"2023-08-20T04:26:51.189667Z","shell.execute_reply":"2023-08-20T04:26:56.403140Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Above is the model summary. We have 23,551,681 trainable parameters. If we use more deeper network number of parameters will increase with possible increase in Accuracy or AUC. But for this model we will use resnet 50 only.","metadata":{}},{"cell_type":"markdown","source":"## 4. Model Deployment","metadata":{}},{"cell_type":"markdown","source":"In this part we will use the model created and fit the model with Adam optimizer and loss as binary cross entropy. We have used from logits = True for better accuracy. If reader wants to avoid it they can modify the last output layer in Resnet50 above and make the Activation function as sigmoid instead of linear.","metadata":{}},{"cell_type":"code","source":"model.compile(optimizer = tf.keras.optimizers.Adam(learning_rate = 0.001), loss=tf.keras.losses.BinaryCrossentropy(from_logits = True), metrics=['accuracy'])","metadata":{"execution":{"iopub.status.busy":"2023-08-20T04:27:31.466903Z","iopub.execute_input":"2023-08-20T04:27:31.467675Z","iopub.status.idle":"2023-08-20T04:27:31.499373Z","shell.execute_reply.started":"2023-08-20T04:27:31.467637Z","shell.execute_reply":"2023-08-20T04:27:31.498214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = model.fit(\n    train_generator,\n    steps_per_epoch = train_steps,\n    validation_data = validation_generator,\n    validation_steps = val_steps,\n    epochs = 10\n)","metadata":{"execution":{"iopub.status.busy":"2023-08-20T04:45:28.527641Z","iopub.execute_input":"2023-08-20T04:45:28.528083Z","iopub.status.idle":"2023-08-20T06:29:06.803239Z","shell.execute_reply.started":"2023-08-20T04:45:28.528049Z","shell.execute_reply":"2023-08-20T06:29:06.801766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It takes around 2 hours to train the model with 10 epochs and Kaggle P100 GPU. We can increase the number of epochs but only Training loss decrease for some time after 10 epochs with little change in validation loss and accuracy. In fact we can see from above that after 6th Epoch Validation Accruacy has not changed much.\n\nSo, it will not be efficient to train with more epochs but readers can do it if they have resources and time for this.","metadata":{}},{"cell_type":"markdown","source":"## 5. Model Evaluation","metadata":{}},{"cell_type":"markdown","source":"Model evaluation is limited here as we don't have test dataset labels. We will still perform basic evaluation and see why we cant do indepth model evaluation here.","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, roc_auc_score, roc_curve\n\n\nval_predictions = tf.nn.sigmoid(model.predict(validation_generator)).numpy()\nval_pred_classes = (val_predictions > 0.5).astype(int).flatten() #Threshold is assumed to be 0.5 for this cell\n\n# True labels\ntrue_labels = validation_generator.classes\n\n# Ensure the lengths match\nval_pred_classes = val_pred_classes[:len(true_labels)]\n\n# Calculate metrics\naccuracy = accuracy_score(true_labels, val_pred_classes)\nprecision = precision_score(true_labels, val_pred_classes)\nrecall = recall_score(true_labels, val_pred_classes)\nf1 = f1_score(true_labels, val_pred_classes)\nroc_auc = roc_auc_score(true_labels, val_predictions)\n\nprint(f\"Accuracy: {accuracy:.4f}\")\nprint(f\"Precision: {precision:.4f}\")\nprint(f\"Recall: {recall:.4f}\")\nprint(f\"F1 Score: {f1:.4f}\")\nprint(f\"ROC-AUC: {roc_auc:.4f}\")","metadata":{"execution":{"iopub.status.busy":"2023-08-20T06:49:56.087072Z","iopub.execute_input":"2023-08-20T06:49:56.087468Z","iopub.status.idle":"2023-08-20T06:51:17.609124Z","shell.execute_reply.started":"2023-08-20T06:49:56.087437Z","shell.execute_reply":"2023-08-20T06:51:17.607950Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Above we have used a threshold of 0.5 which is just assumed and is totally not correct. There was no decision boundary given with the dataset so we cant use above metrics to test the performance of the model.\n\nIn order to test the model performance we need to create a submission and get the AUC we will obtain from such submission.","metadata":{}},{"cell_type":"code","source":"best_loss = history['val_loss'][i_min]\nbest_accuracy = history['val_accuracy'][i_min]\nbest_auc = history['val_auc'][i_min]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"acc = [0.] + history.history['accuracy']\nval_acc = [0.] + history.history['val_accuracy']\n\nloss = history.history['loss']\nval_loss = history.history['val_loss']\n\nplt.figure(figsize=(8, 8))\nplt.subplot(2, 1, 1)\nplt.plot(acc, label='Training Accuracy')\nplt.plot(val_acc, label='Validation Accuracy')\nplt.legend(loc='lower right')\nplt.ylabel('Accuracy')\nplt.ylim([min(plt.ylim()),1])\nplt.title('Training and Validation Accuracy')\n\nplt.subplot(2, 1, 2)\nplt.plot(loss, label='Training Loss')\nplt.plot(val_loss, label='Validation Loss')\nplt.legend(loc='upper right')\nplt.ylabel('Cross Entropy')\nplt.ylim([0,1.0])\nplt.title('Training and Validation Loss')\nplt.xlabel('epoch')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-20T06:54:51.791075Z","iopub.execute_input":"2023-08-20T06:54:51.791521Z","iopub.status.idle":"2023-08-20T06:54:52.340886Z","shell.execute_reply.started":"2023-08-20T06:54:51.791488Z","shell.execute_reply":"2023-08-20T06:54:52.339712Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see from above images after few early epochs we cant see notable changes in Training and Valdiation metrics. While Training Binary Cross Entroy loss has a downward trend and its possible to get further low error Validation loss has more flat curve.\n\nAnd for the accuracy both Training and Validation Accuracy has reached a satisfactory level beyond 0.94 after 6th epoch. There is little change after that. Also, since both Accuracies don't have big gaps we can assure that there is little to no Model Variance.\n\nWe will submit the file and get the AUC from final submission in next step.","metadata":{}},{"cell_type":"markdown","source":"## 6. Submission","metadata":{}},{"cell_type":"code","source":"test_probs = model.predict(test_generator)","metadata":{"execution":{"iopub.status.busy":"2023-08-20T07:27:06.432003Z","iopub.execute_input":"2023-08-20T07:27:06.432493Z","iopub.status.idle":"2023-08-20T07:29:02.614954Z","shell.execute_reply.started":"2023-08-20T07:27:06.432454Z","shell.execute_reply":"2023-08-20T07:29:02.613816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_labels = np.round(test_probs).astype(int).flatten()\nout_df = pd.DataFrame()\nout_df[\"id\"] = test_ids\nout_df[\"label\"] = test_labels\nout_df.to_csv(os.path.join('/kaggle/working/', \"test_labels.csv\"), index=False)","metadata":{"execution":{"iopub.status.busy":"2023-08-20T07:17:00.484961Z","iopub.execute_input":"2023-08-20T07:17:00.485726Z","iopub.status.idle":"2023-08-20T07:17:00.729918Z","shell.execute_reply.started":"2023-08-20T07:17:00.485680Z","shell.execute_reply":"2023-08-20T07:17:00.728805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import shutil\nsubmission_file = r\"/kaggle/working/test_labels.csv\"\nshutil.copy(submission_file, \"submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-08-20T07:19:16.963592Z","iopub.execute_input":"2023-08-20T07:19:16.964630Z","iopub.status.idle":"2023-08-20T07:19:16.975651Z","shell.execute_reply.started":"2023-08-20T07:19:16.964589Z","shell.execute_reply":"2023-08-20T07:19:16.974243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Conclusion","metadata":{}},{"cell_type":"markdown","source":"In the final submission we got the public area under the ROC curve or AUC score of 0.9463 and private score of 0.9167. This is acceptable result as we used only one deep Resnet Model of 50 layers. \n\nWe can improve the score by creating even deeper neural network or by even using transfer learning and import Mobilenet or Inception-Resnet architectures or even more deeper architecture beyond 150 hidedn layer.\n\nWe can also improve the model accuracy by using data augmentation and increasing the size of training data set.\n\nOverall with Resnet50 we got Test AUC of 0.9463 in Histopathological cancer detection data set. Traiing loss was 0.0989 with accuracy of 0.9627 and validation loss was 0.1408 with val_accuracy of 0.9476. We can strive to improve the model by using more deeper architecutes.","metadata":{}}]}