{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduciton","metadata":{}},{"cell_type":"markdown","source":"The goal of this project is to train a neural network in order to identify if an image of body tissue collected from a pathology scan is cancerous or not. This is a binary classification problem, meaning that each training example is labeled as either being cancerous(1) or not cancerous(0). Cancerous in this case means that at least one pixel of the image contains tumorous tissue. \n\nThe source of this data is this Kaggle competition: https://www.kaggle.com/competitions/histopathologic-cancer-detection\n\nThe dataset consists of 220025 training images that are all labeled with if they are cancerous or not. Additionally, there are 57458 unlabeled test images. The goal of this project for the built neural network to predict reasonably well the labels on this test set. The dataset does not come with test labels, so the test performance will be gotten by making submissions to the kaggle competition. ","metadata":{}},{"cell_type":"markdown","source":"# Import Libraries and Define Helper Functions","metadata":{}},{"cell_type":"code","source":"import os\nimport shutil\nimport json\n\nimport numpy as np\nimport pandas as pd\n\nfrom PIL import Image\nimport matplotlib.pyplot as plt\n\nimport tensorflow as tf\nfrom tensorflow.keras.callbacks import ModelCheckpoint, EarlyStopping\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Dense, Flatten, Conv2D, MaxPooling2D, Dropout\nfrom tensorflow.keras.utils import plot_model\nfrom tensorflow.keras.optimizers.legacy import Adam","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:10:46.586319Z","start_time":"2023-07-21T03:10:42.891477Z"},"execution":{"iopub.status.busy":"2023-08-06T19:09:23.119131Z","iopub.execute_input":"2023-08-06T19:09:23.120389Z","iopub.status.idle":"2023-08-06T19:09:23.128217Z","shell.execute_reply.started":"2023-08-06T19:09:23.120320Z","shell.execute_reply":"2023-08-06T19:09:23.127024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def display_random_images(image_data_array, N):\n    \"\"\"Displays N random images from the given image_data_array\"\"\"\n    # Calculate the grid size\n    grid_size = int(np.ceil(np.sqrt(N)))\n\n    # Randomly select N images\n    selected_images = np.random.choice(image_data_array.shape[0], N, replace=False)\n\n    fig, axs = plt.subplots(grid_size, grid_size, figsize=(10, 10))\n\n    for i, ax in enumerate(axs.flatten()):\n        if i < N:\n            ax.imshow(image_data_array[selected_images[i]])\n            ax.axis('off')  # Turn off axis for each subplot\n        else:\n            fig.delaxes(ax)  # Remove excess subplots\n\n    plt.tight_layout()\n    plt.show()\n    \ndef make_train_path(id_str):\n    return os.path.join(r\"/kaggle/input/histopathologic-cancer-detection/train\", f\"{id_str}.tif\")","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:10:46.602337Z","start_time":"2023-07-21T03:10:46.588321Z"},"execution":{"iopub.status.busy":"2023-08-06T19:09:25.731094Z","iopub.execute_input":"2023-08-06T19:09:25.731518Z","iopub.status.idle":"2023-08-06T19:09:25.741682Z","shell.execute_reply.started":"2023-08-06T19:09:25.731483Z","shell.execute_reply":"2023-08-06T19:09:25.740199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exploratory Data Analysis","metadata":{}},{"cell_type":"markdown","source":"Start by having a look at what one of these images look like and what structure it has. ","metadata":{}},{"cell_type":"code","source":"example_path = \"/kaggle/input/histopathologic-cancer-detection/train/0000f8a4da4c286eee5cf1b0d2ab82f979989f7b.tif\"\nexample_img = Image.open(example_path)\nexample_array = np.array(example_img)\nprint(f\"Image Shape = {example_array.shape}\")\nexample_img","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:10:46.698426Z","start_time":"2023-07-21T03:10:46.604340Z"},"execution":{"iopub.status.busy":"2023-08-06T19:09:26.171432Z","iopub.execute_input":"2023-08-06T19:09:26.171828Z","iopub.status.idle":"2023-08-06T19:09:26.531381Z","shell.execute_reply.started":"2023-08-06T19:09:26.171797Z","shell.execute_reply":"2023-08-06T19:09:26.530245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The image is visually a collection of cells from a tissue sample. Its resolution 96x96 pixels and it is stored as an RGB image, meaning there are separate channels for red, green, and blue.\n\nThe training labels are stored in `train_labels.csv`, which allows the label of each image to be looked up using the id of the image, which comes from the filename of each image file. \n\nThis labels file will be loaded into a data frame to confirm its structure. Additionally, the filename of each training image will be added, which will be useful for loading in the data in bulk later on. ","metadata":{}},{"cell_type":"code","source":"train_labels_path = r\"/kaggle/input/histopathologic-cancer-detection/train_labels.csv\"\ntrain_labels_df = pd.read_csv(train_labels_path)\ntrain_labels_df[\"filename\"] = train_labels_df[\"id\"].apply(make_train_path)\ntrain_labels_df[\"label\"] = train_labels_df[\"label\"].astype(str)\ntrain_labels_df.head()","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:10:47.702340Z","start_time":"2023-07-21T03:10:46.701429Z"},"execution":{"iopub.status.busy":"2023-08-06T19:09:26.533086Z","iopub.execute_input":"2023-08-06T19:09:26.533450Z","iopub.status.idle":"2023-08-06T19:09:27.758026Z","shell.execute_reply.started":"2023-08-06T19:09:26.533420Z","shell.execute_reply":"2023-08-06T19:09:27.757119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Check to see how balanced the dataset is by visualizing what percentage of the training dataset is cancerous(1) and what percentage is non-cancerous(0).","metadata":{}},{"cell_type":"code","source":"unique_labels, counts = np.unique(train_labels_df.label.values, return_counts=True)\nplt.pie(counts/np.sum(counts), labels=unique_labels, autopct='%1.1f%%')\nplt.show()","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:10:47.942550Z","start_time":"2023-07-21T03:10:47.704344Z"},"execution":{"iopub.status.busy":"2023-08-06T19:09:27.759799Z","iopub.execute_input":"2023-08-06T19:09:27.760352Z","iopub.status.idle":"2023-08-06T19:09:28.308056Z","shell.execute_reply.started":"2023-08-06T19:09:27.760309Z","shell.execute_reply":"2023-08-06T19:09:28.306394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Cancerous examples only make up 40.5% of the dataset while non-cancerous examples make up the remaining 59.5%. While this is slightly imbalanced, it is not imbalanced enough to perform any operations to balance the dataset, at least not on the first pass.","metadata":{}},{"cell_type":"markdown","source":"Next, a small subset of the training images will be loaded in (200 images) and will be used to visualize each of the classes individually to see if there are any visually obvious features of each class.","metadata":{}},{"cell_type":"code","source":"sample_data = np.empty((200, 96, 96, 3), dtype=np.uint8)\nsample_labels = np.empty(200, dtype=np.int8)\nfor i in range(len(train_labels_df))[:200]:\n    img_path = make_train_path(train_labels_df.id.values[i])\n    img = Image.open(img_path)\n    sample_data[i] = np.array(img)\n    sample_labels[i] = train_labels_df.label.values[i]","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:10:48.709252Z","start_time":"2023-07-21T03:10:47.944553Z"},"execution":{"iopub.status.busy":"2023-08-06T19:09:28.310447Z","iopub.execute_input":"2023-08-06T19:09:28.311405Z","iopub.status.idle":"2023-08-06T19:09:29.542111Z","shell.execute_reply.started":"2023-08-06T19:09:28.311345Z","shell.execute_reply":"2023-08-06T19:09:29.541002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Display 64 random images that are non-cancerous and 64 random images that are cancerous. ","metadata":{}},{"cell_type":"code","source":"print(\"Non-Cancerous Images\")\ndisplay_random_images(sample_data[sample_labels == 0], 64)","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:10:50.336732Z","start_time":"2023-07-21T03:10:48.710253Z"},"execution":{"iopub.status.busy":"2023-08-06T19:09:29.544414Z","iopub.execute_input":"2023-08-06T19:09:29.545721Z","iopub.status.idle":"2023-08-06T19:09:32.258029Z","shell.execute_reply.started":"2023-08-06T19:09:29.545639Z","shell.execute_reply":"2023-08-06T19:09:32.254047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Cancerous Images\")\ndisplay_random_images(sample_data[sample_labels == 1], 64)","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:10:51.998243Z","start_time":"2023-07-21T03:10:50.338734Z"},"execution":{"iopub.status.busy":"2023-08-06T19:09:35.173570Z","iopub.execute_input":"2023-08-06T19:09:35.174004Z","iopub.status.idle":"2023-08-06T19:09:38.118611Z","shell.execute_reply.started":"2023-08-06T19:09:35.173967Z","shell.execute_reply":"2023-08-06T19:09:38.117166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Both the cancerous and non-cancerous images visually seem to contain similar features. Both sets of images are very diverse and contain features such as small collections of purple cells, big white blobs, and dark patches. Without proper medical training, it is difficult to come up with any features visually that distinguish a cancerous image from a non-cancerous one. Hopefully a neural network can do better. ","metadata":{}},{"cell_type":"markdown","source":"# Data Generators","metadata":{}},{"cell_type":"markdown","source":"The training set for this task is very large, consisting of 220025 96x96x3 images, which works out to be 5.72 GB of data. It would be impractical to load in and store all of this data at once. Therefore a technique called batch loading will be used which loads images into memory in batches as they are needed for training and then replaces them with the next batch. This way, only one batch of images are in memory at a given time, which significantly reduces the memory necessary for training.\n\nStart by defining an `ImageDataGenerator` object. It will preprocess the images by rescaling them so each pixel is between 0 and 1 and also perform an 80/20 train-validation split. ","metadata":{}},{"cell_type":"code","source":"datagen = ImageDataGenerator(rescale=1./255, validation_split=0.2)","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:10:52.014258Z","start_time":"2023-07-21T03:10:52.000245Z"},"execution":{"iopub.status.busy":"2023-08-06T19:09:38.120506Z","iopub.execute_input":"2023-08-06T19:09:38.120883Z","iopub.status.idle":"2023-08-06T19:09:38.125831Z","shell.execute_reply.started":"2023-08-06T19:09:38.120853Z","shell.execute_reply":"2023-08-06T19:09:38.124765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now define the train and validation generators, the will read in images based on the filepaths specified in `train_labels_df` and store their labels based on the labels in that dataframe. A batch size of 32 images will be used.  ","metadata":{}},{"cell_type":"code","source":"train_generator = datagen.flow_from_dataframe(\n    dataframe=train_labels_df,\n    x_col=\"filename\",\n    y_col=\"label\",\n    target_size=(96, 96),\n    color_mode=\"rgb\",\n    batch_size=32,\n    class_mode=\"binary\",\n    subset=\"training\",\n    validate_filenames=False\n)\n\nvalidation_generator = datagen.flow_from_dataframe(\n    dataframe=train_labels_df,\n    x_col=\"filename\",\n    y_col=\"label\",\n    target_size=(96, 96),\n    color_mode=\"rgb\",\n    batch_size=32,\n    class_mode=\"binary\",\n    subset=\"validation\",\n    validate_filenames=False\n)","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:11:20.540480Z","start_time":"2023-07-21T03:10:52.019262Z"},"execution":{"iopub.status.busy":"2023-08-06T19:09:38.578167Z","iopub.execute_input":"2023-08-06T19:09:38.579145Z","iopub.status.idle":"2023-08-06T19:09:40.168362Z","shell.execute_reply.started":"2023-08-06T19:09:38.579106Z","shell.execute_reply":"2023-08-06T19:09:40.167134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A generator will also be made for the test data. First a similar dataframe `test_df` must be made which is in the same format as `train_labels_df`, just without a column for labels. ","metadata":{}},{"cell_type":"code","source":"test_path = \"/kaggle/input/histopathologic-cancer-detection/test\"\ntest_ids = [filename[:-4] for filename in os.listdir(test_path)]\ntest_filenames = [os.path.join(test_path, filename) for filename in os.listdir(test_path)]\ntest_df = pd.DataFrame()\ntest_df[\"id\"] = test_ids\ntest_df[\"filename\"] = test_filenames\n\ntest_generator = datagen.flow_from_dataframe(\n    dataframe=test_df,\n    x_col=\"filename\",\n    y_col=None,\n    target_size=(96, 96),\n    color_mode=\"rgb\",\n    batch_size=64,\n    shuffle=False,\n    class_mode=None,\n    validate_filenames=False\n)","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:11:24.751834Z","start_time":"2023-07-21T03:11:20.542481Z"},"execution":{"iopub.status.busy":"2023-08-06T19:09:40.170204Z","iopub.execute_input":"2023-08-06T19:09:40.170642Z","iopub.status.idle":"2023-08-06T19:09:45.596932Z","shell.execute_reply.started":"2023-08-06T19:09:40.170611Z","shell.execute_reply":"2023-08-06T19:09:45.595658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model Training Funciton","metadata":{}},{"cell_type":"markdown","source":"Next define some functions for training neural networks. Several neural networks will be trained so defining these functions significantly cuts down on shared code. \n\nThe first function train_model is responsible for performing the model training. It takes in an untrained model and trains it based on the training data, also using validation data to monitor the model's performance and using early stopping to prevent overfitting. This early stopping occurs if the validation loss has not improved for 10 epochs. The function saves the training history and dumps both the trained model and the history to a file when it is done training, allowing the model to be loaded back in if necessary. If load_from_file is true, the model and history will be loaded in from file rather than retrained.\n\nOnce a model is either trained or loaded, if plot is true, it will create a plot visualizing the training and validation loss, accuracy, and AUC over the training epochs. If run_test is true, it will run the second function make_output_csv, which processes the test data and outputs the results in an appropriate format to be submitted to the Kaggle competition for evaluation.\n\nFinally, a summary of the epoch with the lowest validation loss (which is the epoch whose weights are saved in the model after early stopping) is printed and that summary information is also returned. ","metadata":{}},{"cell_type":"code","source":"def train_model(model, train_generator, validation_generator, test_generator, folder_path, load_from_file, plot=True, run_test=True):\n    np.random.seed(42)\n    tf.random.set_seed(42)\n    \n    checkpoint_path = os.path.join(folder_path, \"model.ckpt\")\n    history_path = os.path.join(folder_path, \"history.json\")\n    \n    os.makedirs(folder_path, exist_ok=True)\n    \n    if load_from_file:\n        print(\"Loading Model from File\")\n        model.load_weights(checkpoint_path)\n        with open(history_path) as f:\n            history = json.load(f)\n            \n    else:\n        shutil.rmtree(folder_path)\n        checkpoint = ModelCheckpoint(checkpoint_path, save_weights_only=True, verbose=1, save_best_only=True)\n        early_stopping = EarlyStopping(monitor=\"val_loss\", patience=10, verbose=1)\n\n\n        history = model.fit(train_generator,\n                            steps_per_epoch=len(train_generator),\n                            validation_data=validation_generator,\n                            validation_steps=len(validation_generator),\n                            epochs=100,\n                            callbacks=[early_stopping, checkpoint]).history\n\n        with open(history_path, \"w+\") as f:\n            json.dump(history, f)\n            \n    if plot:\n        epochs = np.arange(1, len(history[\"val_loss\"]) + 1, 1)\n        fig, axs = plt.subplots(1, 3, figsize=(15, 5))\n\n        axs[0].plot(epochs, history[\"loss\"], label=\"train\")\n        axs[0].plot(epochs, history[\"val_loss\"], label=\"validation\")\n        axs[0].set_title('Model loss')\n        axs[0].set_ylabel('Loss')\n        axs[0].set_xlabel('Epoch')\n        axs[0].legend()\n\n        axs[1].plot(epochs, history[\"accuracy\"], label=\"train\")\n        axs[1].plot(epochs, history[\"val_accuracy\"], label=\"validation\")\n        axs[1].set_title('Model accuracy')\n        axs[1].set_ylabel('Accuracy')\n        axs[1].set_xlabel('Epoch')\n        axs[1].legend()\n\n        axs[2].plot(epochs, history[\"auc\"], label=\"train\")\n        axs[2].plot(epochs, history[\"val_auc\"], label=\"validation\")\n        axs[2].set_title('Model AUC')\n        axs[2].set_ylabel('AUC')\n        axs[2].set_xlabel('Epoch')\n        axs[2].legend()\n\n        plt.tight_layout()\n        plt.show()\n        \n    if run_test:\n        make_output_csv(model, test_generator, folder_path)\n        \n    i_min = np.argmin(history[\"val_loss\"])\n    best_epoch = i_min+1\n    best_loss = history['val_loss'][i_min]\n    best_accuracy = history['val_accuracy'][i_min]\n    best_auc = history['val_auc'][i_min]\n    \n    \n    return best_epoch, best_loss, best_accuracy, best_auc\n\ndef make_output_csv(model, test_generator, folder_path):\n    \"\"\"Runs the test data in test_generator aganist a trained model and outputs the results to test_labels.csv in the specified folder\"\"\"\n    test_probs = model.predict(test_generator)\n    test_labels = np.round(test_probs).astype(int).flatten()\n    out_df = pd.DataFrame()\n    out_df[\"id\"] = test_ids\n    out_df[\"label\"] = test_labels\n    out_df.to_csv(os.path.join(folder_path, \"test_labels.csv\"), index=False)","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:11:24.783863Z","start_time":"2023-07-21T03:11:24.754836Z"},"execution":{"iopub.status.busy":"2023-08-06T19:09:45.599256Z","iopub.execute_input":"2023-08-06T19:09:45.599668Z","iopub.status.idle":"2023-08-06T19:09:45.655998Z","shell.execute_reply.started":"2023-08-06T19:09:45.599638Z","shell.execute_reply":"2023-08-06T19:09:45.654916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Fully Connected","metadata":{}},{"cell_type":"markdown","source":"The first model that will be trained for this problem is a simple, dense fully connected neural network. It will have one hidden layer with 128 nodes that uses a ReLU activation function and a single output node that uses a sigmoid activation function.\n\nAs this is a binary classification problem, binary cross entropy will be used as the loss function and the ADAM optimizer will be used. In addition to the loss, the metrics of accuracy and AUC will be monitored. All future models will also be compiled in the same way. ","metadata":{}},{"cell_type":"code","source":"dense_model = Sequential()\ndense_model.add(Flatten(input_shape=(96, 96, 3)))\ndense_model.add(Dense(128, activation=\"relu\"))\ndense_model.add(Dense(1, activation=\"sigmoid\"))\n\ndense_model.compile(loss='binary_crossentropy', optimizer=Adam(), metrics=['accuracy', tf.keras.metrics.AUC(name='auc')])\ndense_model.summary()","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:11:27.480837Z","start_time":"2023-07-21T03:11:24.786866Z"},"execution":{"iopub.status.busy":"2023-08-06T19:23:21.152922Z","iopub.execute_input":"2023-08-06T19:23:21.153410Z","iopub.status.idle":"2023-08-06T19:23:21.255189Z","shell.execute_reply.started":"2023-08-06T19:23:21.153371Z","shell.execute_reply":"2023-08-06T19:23:21.254095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dense_train_res = train_model(dense_model, train_generator, validation_generator, test_generator, \"/kaggle/input/cancerDetection/dense_model\", load_from_file=True, run_test=False)\nprint(f\"Best Epoch: {dense_train_res[0]}\")\nprint(f\"Best Model Loss: {dense_train_res[1]}\")\nprint(f\"Best Model Accuracy: {dense_train_res[2]}\")\nprint(f\"Best Model AUC: {dense_train_res[3]}\")","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:11:28.208494Z","start_time":"2023-07-21T03:11:27.482840Z"},"execution":{"iopub.status.busy":"2023-08-06T19:23:24.183522Z","iopub.execute_input":"2023-08-06T19:23:24.183934Z","iopub.status.idle":"2023-08-06T19:23:25.757643Z","shell.execute_reply.started":"2023-08-06T19:23:24.183900Z","shell.execute_reply":"2023-08-06T19:23:25.756767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This optimization ran for 34 epochs, being ended by early stopping because the validation loss stopped improving after epoch 24, where it had a validation accuracy of 71.5% and an AUC of 0.775. Based on how the train and validation loss and AUC start to diverge from each other around epoch 10, this is when it started to over fit to the training data. \n\nThe test AUC from Kaggle is **0.66**, which is a fair amount lower than the validation AUC of 0.775. This is further evidence that the model started to over fit earlier on and at the final model does not generalize well to the test data. ","metadata":{}},{"cell_type":"markdown","source":"A `results_df` will be defined to store each model's validation loss, accuracy, and AUC as well as the test AUC. This will help compare all models later on. ","metadata":{}},{"cell_type":"code","source":"results_df = pd.DataFrame(columns=[\"Model Name\", \"Validation Loss\", \"Validation Accuracy\", \"Validation AUC\", \"Test AUC\"])\nresults_df.loc[len(results_df.index)] = [\"Dense\", dense_train_res[1], dense_train_res[2], dense_train_res[3], 0.66]","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:11:28.224509Z","start_time":"2023-07-21T03:11:28.210496Z"},"execution":{"iopub.status.busy":"2023-08-06T19:23:38.920243Z","iopub.execute_input":"2023-08-06T19:23:38.920754Z","iopub.status.idle":"2023-08-06T19:23:38.930053Z","shell.execute_reply.started":"2023-08-06T19:23:38.920708Z","shell.execute_reply":"2023-08-06T19:23:38.928688Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Convolutional","metadata":{}},{"cell_type":"markdown","source":"Because the goal is to detect certain features within these images that may indicate that they contain tumorous tissue, a more logical architecture for the neural network than a dense network is a convolutional neural network. A convolutional neural network is able to pick up on features, regardless of where they are in the image making them well suited for this task\n\nTo start with try a simple convolutional neural network with two convolutional layers that each have 32 3x3 convolutional kernels and use a ReLU activation function. Each of these convolutional layers is followed by a 2x2 max pooling layer. Finally, the output of the second max pooling layer is flattened and fed into the same dense architecture used before, with a hidden layer of 128 nodes and a final single output node. ","metadata":{}},{"cell_type":"code","source":"conv_model = Sequential()\nconv_model.add(Conv2D(32, (3, 3), activation='relu', input_shape=(96, 96, 3)))\nconv_model.add(MaxPooling2D((2, 2)))\nconv_model.add(Conv2D(32, (3, 3), activation='relu'))\nconv_model.add(MaxPooling2D((2, 2)))\nconv_model.add(Flatten())\nconv_model.add(Dense(128, activation=\"relu\"))\nconv_model.add(Dense(1, activation=\"sigmoid\"))\n\nconv_model.compile(loss='binary_crossentropy', optimizer=Adam(), metrics=['accuracy', tf.keras.metrics.AUC(name='auc')])\n\nconv_model.summary()","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:11:28.320598Z","start_time":"2023-07-21T03:11:28.226511Z"},"execution":{"iopub.status.busy":"2023-08-06T19:24:18.363926Z","iopub.execute_input":"2023-08-06T19:24:18.364313Z","iopub.status.idle":"2023-08-06T19:24:18.504501Z","shell.execute_reply.started":"2023-08-06T19:24:18.364281Z","shell.execute_reply":"2023-08-06T19:24:18.503117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This convolutional architecture notably has 1.6 million fewer trainable parameters than the previous dense architecture, despite being deeper. ","metadata":{}},{"cell_type":"code","source":"conv_train_res = train_model(conv_model, train_generator, validation_generator, test_generator, \"/kaggle/input/cancerDetection/conv_model\", load_from_file=True, run_test=False)\nprint(f\"Best Epoch: {conv_train_res[0]}\")\nprint(f\"Best Model Loss: {conv_train_res[1]}\")\nprint(f\"Best Model Accuracy: {conv_train_res[2]}\")\nprint(f\"Best Model AUC: {conv_train_res[3]}\")","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:11:29.058267Z","start_time":"2023-07-21T03:11:28.322599Z"},"execution":{"iopub.status.busy":"2023-08-06T19:24:22.237732Z","iopub.execute_input":"2023-08-06T19:24:22.238148Z","iopub.status.idle":"2023-08-06T19:24:24.083061Z","shell.execute_reply.started":"2023-08-06T19:24:22.238116Z","shell.execute_reply":"2023-08-06T19:24:24.081901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This optimization ran for 16 epochs, being ended by early stopping because the validation loss stopped improving after epoch 6, where it had a validation accuracy of 84.9% and an AUC of 0.92. Based on how the train and validation loss and AUC start to diverge from each other around epoch 6, there is very severe over fitting that started happening very early on and only got worse with every epoch. Despite this early over fitting, this convolutional model is 13.5% more accurate than the previous dense model, demonstrating the effectiveness of the convolutional neural network for this task. \n\nThe test AUC from Kaggle is **0.7566**, which is a fair amount lower than the validation AUC of 0.849. This is further evidence that the model was starting to over fit even in epoch 6.   \n","metadata":{"ExecuteTime":{"end_time":"2023-07-17T17:04:59.423221Z","start_time":"2023-07-17T17:04:59.412212Z"}}},{"cell_type":"code","source":"results_df.loc[len(results_df.index)] = [\"Convolutional\", conv_train_res[1], conv_train_res[2], conv_train_res[3], 0.7566]","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:11:29.074284Z","start_time":"2023-07-21T03:11:29.060269Z"},"execution":{"iopub.status.busy":"2023-08-06T19:24:26.348919Z","iopub.execute_input":"2023-08-06T19:24:26.349296Z","iopub.status.idle":"2023-08-06T19:24:26.360873Z","shell.execute_reply.started":"2023-08-06T19:24:26.349265Z","shell.execute_reply":"2023-08-06T19:24:26.359719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Convolutional With Dropout","metadata":{}},{"cell_type":"markdown","source":"In order to try and remedy the early overfitting in the previous model, add some regularization in the form of a dropout layer before the second convolutional layer, before the dense hidden layer, and before the output. These drop out layers randomly set 50% of their inputs to zero, which helps prevent overfitting by creating a more robust model with more meaningful weights. Except for the dropout layers, the architecture is the same as the previous network. ","metadata":{}},{"cell_type":"code","source":"conv_dropout_model = Sequential()\nconv_dropout_model.add(Conv2D(32, (3, 3), activation='relu', input_shape=(96, 96, 3)))\nconv_dropout_model.add(MaxPooling2D((2, 2)))\nconv_dropout_model.add(Dropout(0.5))\nconv_dropout_model.add(Conv2D(32, (3, 3), activation='relu'))\nconv_dropout_model.add(MaxPooling2D((2, 2)))\nconv_dropout_model.add(Dropout(0.5))\nconv_dropout_model.add(Flatten())\nconv_dropout_model.add(Dense(128, activation=\"relu\"))\nconv_dropout_model.add(Dropout(0.5))\nconv_dropout_model.add(Dense(1, activation=\"sigmoid\"))\n\nconv_dropout_model.compile(loss='binary_crossentropy', optimizer=Adam(), metrics=['accuracy', tf.keras.metrics.AUC(name='auc')])\n\nconv_dropout_model.summary()","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:11:29.169371Z","start_time":"2023-07-21T03:11:29.076287Z"},"execution":{"iopub.status.busy":"2023-08-06T19:24:54.981990Z","iopub.execute_input":"2023-08-06T19:24:54.982431Z","iopub.status.idle":"2023-08-06T19:24:55.146271Z","shell.execute_reply.started":"2023-08-06T19:24:54.982394Z","shell.execute_reply":"2023-08-06T19:24:55.144945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"conv_dropout_train_res = train_model(conv_dropout_model, train_generator, validation_generator, test_generator, \"/kaggle/input/cancerDetection/conv_dropout_model\", load_from_file=True, run_test=False)\nprint(f\"Best Epoch: {conv_dropout_train_res[0]}\")\nprint(f\"Best Model Loss: {conv_dropout_train_res[1]}\")\nprint(f\"Best Model Accuracy: {conv_dropout_train_res[2]}\")\nprint(f\"Best Model AUC: {conv_dropout_train_res[3]}\")","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:11:29.803948Z","start_time":"2023-07-21T03:11:29.171373Z"},"execution":{"iopub.status.busy":"2023-08-06T19:25:05.966654Z","iopub.execute_input":"2023-08-06T19:25:05.967162Z","iopub.status.idle":"2023-08-06T19:25:07.356128Z","shell.execute_reply.started":"2023-08-06T19:25:05.967118Z","shell.execute_reply":"2023-08-06T19:25:07.354822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This optimization ran for 56 epochs, being ended by early stopping because the validation loss stopped improving after epoch 46, where it had a validation accuracy of 87.6% and an AUC of 0.949. This model certainly did not overfit to the training data as early as the previous model but a new issue has appeared where the validation loss, accuracy, and AUC are extremely variable from epoch to epoch. Despite this issue, there is still an improvement of 1.8% in the validation accuracy from the previous model, showing that adding the dropout layers did help. \n\nThe test AUC from Kaggle is **0.8001**, which is a fair amount lower than the validation AUC of 0.949, showing that there is still some over fitting going on.","metadata":{"ExecuteTime":{"end_time":"2023-07-17T18:15:11.683418Z","start_time":"2023-07-17T18:15:11.668406Z"}}},{"cell_type":"code","source":"results_df.loc[len(results_df.index)] = [\"Convolutional Dropout Batch 32\", conv_dropout_train_res[1], conv_dropout_train_res[2], conv_dropout_train_res[3], 0.8001]","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:11:29.819963Z","start_time":"2023-07-21T03:11:29.805950Z"},"execution":{"iopub.status.busy":"2023-08-06T19:25:09.965495Z","iopub.execute_input":"2023-08-06T19:25:09.965882Z","iopub.status.idle":"2023-08-06T19:25:09.975678Z","shell.execute_reply.started":"2023-08-06T19:25:09.965845Z","shell.execute_reply":"2023-08-06T19:25:09.974301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Tune Batch Size","metadata":{}},{"cell_type":"markdown","source":"In batch optimization like what is being used here, the gradient of the loss function with respect to each tunable weight is estimated at every step using only a subset of the total training data. So far 32 images has been the batch sized used to train all of the above models. However if this batch size is not large enough to be a statistically significant sample of the total training dataset, then these estimated gradients will be inaccurate. If the estimated gradients are inaccurate then there are two significant consequences. First, the model will both not optimize as much as it could every epoch and second, the model will not generalize as well to the validation or test data. \n\nIn this case there are over 200k training images, so it is likely that the batch size of 32 has not been enough to get a statistically significant sample of the training data to estimate gradients with. This could explain the highly variable validation accuracy in the previous model, with the accuracy on the validation set at each epoch being tied to how well or poorly the gradient was estimated from the small batches.\n\nTherefore the next step is to increase the batch size in order to better estimate the gradients at each step and hopefully produce less variable validation accuracy and a more robust model. To start with the batch size will be doubled to 64 and the same model as the previous one will be trained with this larger batch size. ","metadata":{}},{"cell_type":"code","source":"train_generator_64 = datagen.flow_from_dataframe(\n    dataframe=train_labels_df,\n    x_col=\"filename\",\n    y_col=\"label\",\n    target_size=(96, 96),\n    color_mode=\"rgb\",\n    batch_size=64,\n    class_mode=\"binary\",\n    subset=\"training\",\n    validate_filenames=False\n)\n\nvalidation_generator_64 = datagen.flow_from_dataframe(\n    dataframe=train_labels_df,\n    x_col=\"filename\",\n    y_col=\"label\",\n    target_size=(96, 96),\n    color_mode=\"rgb\",\n    batch_size=64,\n    class_mode=\"binary\",\n    subset=\"validation\",\n    validate_filenames=False\n)","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:11:59.386129Z","start_time":"2023-07-21T03:11:29.821966Z"},"execution":{"iopub.status.busy":"2023-08-06T19:25:56.222453Z","iopub.execute_input":"2023-08-06T19:25:56.223113Z","iopub.status.idle":"2023-08-06T19:25:57.745894Z","shell.execute_reply.started":"2023-08-06T19:25:56.223076Z","shell.execute_reply":"2023-08-06T19:25:57.744620Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"conv_dropout_model_64 = Sequential()\nconv_dropout_model_64.add(Conv2D(32, (3, 3), activation='relu', input_shape=(96, 96, 3)))\nconv_dropout_model_64.add(MaxPooling2D((2, 2)))\nconv_dropout_model_64.add(Dropout(0.5))\nconv_dropout_model_64.add(Conv2D(32, (3, 3), activation='relu'))\nconv_dropout_model_64.add(MaxPooling2D((2, 2)))\nconv_dropout_model_64.add(Dropout(0.5))\nconv_dropout_model_64.add(Flatten())\nconv_dropout_model_64.add(Dense(128, activation=\"relu\"))\nconv_dropout_model_64.add(Dropout(0.5))\nconv_dropout_model_64.add(Dense(1, activation=\"sigmoid\"))\n\nconv_dropout_model_64.compile(loss='binary_crossentropy', optimizer=Adam(), metrics=['accuracy', tf.keras.metrics.AUC(name='auc')])\n\nconv_dropout_model_64.summary()","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:11:59.498231Z","start_time":"2023-07-21T03:11:59.388131Z"},"execution":{"iopub.status.busy":"2023-08-06T19:26:00.928420Z","iopub.execute_input":"2023-08-06T19:26:00.928881Z","iopub.status.idle":"2023-08-06T19:26:01.099759Z","shell.execute_reply.started":"2023-08-06T19:26:00.928847Z","shell.execute_reply":"2023-08-06T19:26:01.098634Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"conv_dropout_train_res_64 = train_model(conv_dropout_model_64, train_generator_64, validation_generator_64, test_generator, \"/kaggle/input/cancerDetection/conv_dropout_64_model\", load_from_file=True, run_test=False)\nprint(f\"Best Epoch: {conv_dropout_train_res_64[0]}\")\nprint(f\"Best Model Loss: {conv_dropout_train_res_64[1]}\")\nprint(f\"Best Model Accuracy: {conv_dropout_train_res_64[2]}\")\nprint(f\"Best Model AUC: {conv_dropout_train_res_64[3]}\")","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:12:00.106768Z","start_time":"2023-07-21T03:11:59.500232Z"},"execution":{"iopub.status.busy":"2023-08-06T19:26:10.850767Z","iopub.execute_input":"2023-08-06T19:26:10.851152Z","iopub.status.idle":"2023-08-06T19:26:12.208148Z","shell.execute_reply.started":"2023-08-06T19:26:10.851122Z","shell.execute_reply":"2023-08-06T19:26:12.207070Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This optimization ran for 33 epochs, being ended by early stopping because the validation loss stopped improving after epoch 23, where it had a validation accuracy of 87.5% and an AUC of 0.949. These are very similar results to the previous batch size of 32 and the validation accuracy is still highly variable. Despite this, there is now a more noticeable downward trend in the validation loss despite the variability, meaning that the larger batch size is starting to help. \n\nThe test AUC from Kaggle is **0.8076**, which is a fair amount lower than the validation AUC of 0.949. It is however a little bit higher than the test AUC from when the batch size was 32, showing that this model generalizes to the test data a bit better. ","metadata":{}},{"cell_type":"code","source":"results_df.loc[len(results_df.index)] = [\"Convolutional Dropout Batch 64\", conv_dropout_train_res_64[1], conv_dropout_train_res_64[2], conv_dropout_train_res_64[3], 0.8076]","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:12:00.122221Z","start_time":"2023-07-21T03:12:00.108770Z"},"execution":{"iopub.status.busy":"2023-08-06T19:26:17.550170Z","iopub.execute_input":"2023-08-06T19:26:17.550618Z","iopub.status.idle":"2023-08-06T19:26:17.560268Z","shell.execute_reply.started":"2023-08-06T19:26:17.550582Z","shell.execute_reply":"2023-08-06T19:26:17.558604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, double the batch size again from 64 to 128 and repeat the same process. ","metadata":{}},{"cell_type":"code","source":"train_generator_128 = datagen.flow_from_dataframe(\n    dataframe=train_labels_df,\n    x_col=\"filename\",\n    y_col=\"label\",\n    target_size=(96, 96),\n    color_mode=\"rgb\",\n    batch_size=128,\n    class_mode=\"binary\",\n    subset=\"training\",\n    validate_filenames=False\n)\n\nvalidation_generator_128 = datagen.flow_from_dataframe(\n    dataframe=train_labels_df,\n    x_col=\"filename\",\n    y_col=\"label\",\n    target_size=(96, 96),\n    color_mode=\"rgb\",\n    batch_size=128,\n    class_mode=\"binary\",\n    subset=\"validation\",\n    validate_filenames=False\n)","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:12:28.679199Z","start_time":"2023-07-21T03:12:00.123223Z"},"execution":{"iopub.status.busy":"2023-08-06T19:26:33.418845Z","iopub.execute_input":"2023-08-06T19:26:33.419282Z","iopub.status.idle":"2023-08-06T19:26:34.928990Z","shell.execute_reply.started":"2023-08-06T19:26:33.419245Z","shell.execute_reply":"2023-08-06T19:26:34.927650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"conv_dropout_model_128 = Sequential()\nconv_dropout_model_128.add(Conv2D(32, (3, 3), activation='relu', input_shape=(96, 96, 3)))\nconv_dropout_model_128.add(MaxPooling2D((2, 2)))\nconv_dropout_model_128.add(Dropout(0.5))\nconv_dropout_model_128.add(Conv2D(32, (3, 3), activation='relu'))\nconv_dropout_model_128.add(MaxPooling2D((2, 2)))\nconv_dropout_model_128.add(Dropout(0.5))\nconv_dropout_model_128.add(Flatten())\nconv_dropout_model_128.add(Dense(128, activation=\"relu\"))\nconv_dropout_model_128.add(Dropout(0.5))\nconv_dropout_model_128.add(Dense(1, activation=\"sigmoid\"))\n\nconv_dropout_model_128.compile(loss='binary_crossentropy', optimizer=Adam(), metrics=['accuracy', tf.keras.metrics.AUC(name='auc')])\n\nconv_dropout_model_128.summary()","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:12:28.790300Z","start_time":"2023-07-21T03:12:28.681200Z"},"execution":{"iopub.status.busy":"2023-08-06T19:26:42.149491Z","iopub.execute_input":"2023-08-06T19:26:42.149928Z","iopub.status.idle":"2023-08-06T19:26:42.325857Z","shell.execute_reply.started":"2023-08-06T19:26:42.149891Z","shell.execute_reply":"2023-08-06T19:26:42.324953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"conv_dropout_train_res_128 = train_model(conv_dropout_model_128, train_generator_128, validation_generator_128, test_generator, \"/kaggle/input/cancerDetection/conv_dropout_128_model\", load_from_file=True, run_test=False)\nprint(f\"Best Epoch: {conv_dropout_train_res_128[0]}\")\nprint(f\"Best Model Loss: {conv_dropout_train_res_128[1]}\")\nprint(f\"Best Model Accuracy: {conv_dropout_train_res_128[2]}\")\nprint(f\"Best Model AUC: {conv_dropout_train_res_128[3]}\")","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:12:29.398854Z","start_time":"2023-07-21T03:12:28.795304Z"},"execution":{"iopub.status.busy":"2023-08-06T19:26:56.365251Z","iopub.execute_input":"2023-08-06T19:26:56.366493Z","iopub.status.idle":"2023-08-06T19:26:57.737955Z","shell.execute_reply.started":"2023-08-06T19:26:56.366441Z","shell.execute_reply":"2023-08-06T19:26:57.736620Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This optimization ran for 53 epochs, being ended by early stopping because the validation loss stopped improving after epoch 43, where it had a validation accuracy of 89% and an AUC of 0.954. These are slightly better results to the previous batch size of 64 and the validation accuracy is now much less variable highly variable, following the training results much more closely. This is starting to show that increasing the batch size is helping the model optimize and generalize \n\nThe test AUC from Kaggle is **0.7979**, which is a fair amount lower than the validation AUC of 0.954. Despite this model visually fitting more closely to the training data, this is a slight decrease in test performance compared to the batch size of 64.","metadata":{}},{"cell_type":"code","source":"results_df.loc[len(results_df.index)] = [\"Convolutional Dropout Batch 128\", conv_dropout_train_res_128[1], conv_dropout_train_res_128[2], conv_dropout_train_res_128[3], 0.7979]","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:12:29.414856Z","start_time":"2023-07-21T03:12:29.400855Z"},"execution":{"iopub.status.busy":"2023-08-06T19:27:04.219623Z","iopub.execute_input":"2023-08-06T19:27:04.220696Z","iopub.status.idle":"2023-08-06T19:27:04.229068Z","shell.execute_reply.started":"2023-08-06T19:27:04.220658Z","shell.execute_reply":"2023-08-06T19:27:04.227806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally, the batch size will be doubled one more time from 128 to 256","metadata":{}},{"cell_type":"code","source":"train_generator_256 = datagen.flow_from_dataframe(\n    dataframe=train_labels_df,\n    x_col=\"filename\",\n    y_col=\"label\",\n    target_size=(96, 96),\n    color_mode=\"rgb\",\n    batch_size=256,\n    class_mode=\"binary\",\n    subset=\"training\",\n    validate_filenames=False\n)\n\nvalidation_generator_256 = datagen.flow_from_dataframe(\n    dataframe=train_labels_df,\n    x_col=\"filename\",\n    y_col=\"label\",\n    target_size=(96, 96),\n    color_mode=\"rgb\",\n    batch_size=256,\n    class_mode=\"binary\",\n    subset=\"validation\",\n    validate_filenames=False\n)","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:12:58.002838Z","start_time":"2023-07-21T03:12:29.416859Z"},"execution":{"iopub.status.busy":"2023-08-06T19:28:12.730628Z","iopub.execute_input":"2023-08-06T19:28:12.731027Z","iopub.status.idle":"2023-08-06T19:28:14.256269Z","shell.execute_reply.started":"2023-08-06T19:28:12.730996Z","shell.execute_reply":"2023-08-06T19:28:14.254766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"conv_dropout_model_256 = Sequential()\nconv_dropout_model_256.add(Conv2D(32, (3, 3), activation='relu', input_shape=(96, 96, 3)))\nconv_dropout_model_256.add(MaxPooling2D((2, 2)))\nconv_dropout_model_256.add(Dropout(0.5))\nconv_dropout_model_256.add(Conv2D(32, (3, 3), activation='relu'))\nconv_dropout_model_256.add(MaxPooling2D((2, 2)))\nconv_dropout_model_256.add(Dropout(0.5))\nconv_dropout_model_256.add(Flatten())\nconv_dropout_model_256.add(Dense(128, activation=\"relu\"))\nconv_dropout_model_256.add(Dropout(0.5))\nconv_dropout_model_256.add(Dense(1, activation=\"sigmoid\"))\n\nconv_dropout_model_256.compile(loss='binary_crossentropy', optimizer=Adam(), metrics=['accuracy', tf.keras.metrics.AUC(name='auc')])\n\nconv_dropout_model_256.summary()","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:12:58.098926Z","start_time":"2023-07-21T03:12:58.004841Z"},"execution":{"iopub.status.busy":"2023-08-06T19:28:15.810160Z","iopub.execute_input":"2023-08-06T19:28:15.810623Z","iopub.status.idle":"2023-08-06T19:28:15.976568Z","shell.execute_reply.started":"2023-08-06T19:28:15.810586Z","shell.execute_reply":"2023-08-06T19:28:15.975079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"conv_dropout_train_res_256 = train_model(conv_dropout_model_256, train_generator_256, validation_generator_256, test_generator, \"/kaggle/input/cancerDetection/conv_dropout_256_model\", load_from_file=True, run_test=False)\nprint(f\"Best Epoch: {conv_dropout_train_res_256[0]}\")\nprint(f\"Best Model Loss: {conv_dropout_train_res_256[1]}\")\nprint(f\"Best Model Accuracy: {conv_dropout_train_res_256[2]}\")\nprint(f\"Best Model AUC: {conv_dropout_train_res_256[3]}\")","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:12:58.851611Z","start_time":"2023-07-21T03:12:58.100928Z"},"execution":{"iopub.status.busy":"2023-08-06T19:28:18.963628Z","iopub.execute_input":"2023-08-06T19:28:18.964059Z","iopub.status.idle":"2023-08-06T19:28:20.320652Z","shell.execute_reply.started":"2023-08-06T19:28:18.964022Z","shell.execute_reply":"2023-08-06T19:28:20.319398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This optimization ran for 85 epochs, being ended by early stopping because the validation loss stopped improving after epoch 75, where it had a validation accuracy of 90.2% and an AUC of 0.964. Similar to the batch size of 128, this batch size of 256 leads to less variable validation metrics, especially AUC, which matches the train AUC very closely at later epochs. There did appear to be some instability between epochs 10 and 20, with the validation accuracy sharply falling before starting to increase again. It is also notable that this optimization ran for many more epochs than any other model before early stopping stopped it, indicating that the optimization was much more stable and gradual than previous models. \n\nThe test AUC from Kaggle is **0.8246**, which is the best test AUC achieved. This demonstrates that increasing the batch size led to even better performance.","metadata":{}},{"cell_type":"code","source":"results_df.loc[len(results_df.index)] = [\"Convolutional Dropout Batch 256\", conv_dropout_train_res_256[1], conv_dropout_train_res_256[2], conv_dropout_train_res_256[3], 0.8246]","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:12:58.867632Z","start_time":"2023-07-21T03:12:58.852611Z"},"execution":{"iopub.status.busy":"2023-08-06T19:28:24.431771Z","iopub.execute_input":"2023-08-06T19:28:24.432152Z","iopub.status.idle":"2023-08-06T19:28:24.440942Z","shell.execute_reply.started":"2023-08-06T19:28:24.432122Z","shell.execute_reply":"2023-08-06T19:28:24.439877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Results","metadata":{}},{"cell_type":"markdown","source":"Below is the final results table comparing each of the created models and summarizing their validation loss, accuracy, and auc along with the test AUC","metadata":{}},{"cell_type":"code","source":"results_df","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:12:58.883611Z","start_time":"2023-07-21T03:12:58.868634Z"},"execution":{"iopub.status.busy":"2023-08-06T19:28:33.250103Z","iopub.execute_input":"2023-08-06T19:28:33.250513Z","iopub.status.idle":"2023-08-06T19:28:33.266319Z","shell.execute_reply.started":"2023-08-06T19:28:33.250481Z","shell.execute_reply":"2023-08-06T19:28:33.265014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To better visualize all of these metrics and how they compare for each model, create some bar charts","metadata":{}},{"cell_type":"code","source":"fig, axs = plt.subplots(2, 2, figsize=(15,10))\n\n# Validation Loss bar chart\naxs[0, 0].bar(results_df['Model Name'], results_df['Validation Loss'], color ='maroon')\naxs[0, 0].set_title('Validation Loss')\naxs[0, 0].set_xticks(range(len(results_df['Model Name'])))\naxs[0, 0].set_xticklabels(results_df['Model Name'], rotation=45, ha='right')\n\n# Validation Accuracy bar chart\naxs[0, 1].bar(results_df['Model Name'], results_df['Validation Accuracy'], color ='blue')\naxs[0, 1].set_title('Validation Accuracy')\naxs[0, 1].set_xticks(range(len(results_df['Model Name'])))\naxs[0, 1].set_xticklabels(results_df['Model Name'], rotation=45, ha='right')\naxs[0, 1].set_ylim(0.6, 1)\n\n# Validation AUC bar chart\naxs[1, 0].bar(results_df['Model Name'], results_df['Validation AUC'], color ='green')\naxs[1, 0].set_title('Validation AUC')\naxs[1, 0].set_xticks(range(len(results_df['Model Name'])))\naxs[1, 0].set_xticklabels(results_df['Model Name'], rotation=45, ha='right')\naxs[1, 0].set_ylim(0.7, 1)\n\n# Test AUC bar chart\naxs[1, 1].bar(results_df['Model Name'], results_df['Test AUC'], color ='purple')\naxs[1, 1].set_title('Test AUC')\naxs[1, 1].set_xticks(range(len(results_df['Model Name'])))\naxs[1, 1].set_xticklabels(results_df['Model Name'], rotation=45, ha='right')\naxs[1, 1].set_ylim(0.6, 1)\n\nplt.tight_layout()\nplt.show()","metadata":{"ExecuteTime":{"end_time":"2023-07-21T03:12:59.491165Z","start_time":"2023-07-21T03:12:58.884614Z"},"execution":{"iopub.status.busy":"2023-08-06T19:28:39.521982Z","iopub.execute_input":"2023-08-06T19:28:39.522463Z","iopub.status.idle":"2023-08-06T19:28:40.680387Z","shell.execute_reply.started":"2023-08-06T19:28:39.522425Z","shell.execute_reply":"2023-08-06T19:28:40.679270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Observation about each model based on these results:\n* The dense model was by far the worst performing model across all metrics. This makes sense as a dense, fully connected neural network is not well suited for this task, but does form a good baseline to compare other models to. \n* The first convolutional model with no dropout performed significantly better than the dense model, but still worse than all of the later convolutional models accross all metrics\n* Adding dropout to the convolutional model improved it accross all metrics. \n* When tuning the batch size hyperparameter, it was not the case that the model improved across all metrics for each new, larger batch size. Batch size 64 had a slightly higher loss and AUC than batch size 32.\n* Across all metrics the best model was the convolutional model with dropout and a batch size of 256","metadata":{}},{"cell_type":"markdown","source":"# Conclusion","metadata":{}},{"cell_type":"markdown","source":"In this project, several neural network models were created in order to classify 96x96 RGB images of tissue as either cancerous or non-cancerous. First a dense, single hidden layer network was trained that did not perform very well. Next, a simple convolutional model was trained which performed much better than the dense model but was found to over fit easily to the training data. \n\nTo address this over fitting, the third model added dropout layers to regularize the weights of the model and prevent over fitting, this exposed the issue of highly variable validation metrics. In order to correct the highly variable validation metrics, the hyperparameter of batch size was tuned with values of 32, 64, 128, and 256 being tried. The best performing model in the end was the convolutional model with dropout trained with a batch size of 256, which achieved a test AUC of **0.8246**.\n\nThere are several ways that this model could be extended and improved even further.\n* The best performing convolutional models still only had two convolutional layers and one fully connected hidden layer. It is likely that by making this neural network even deeper, with more convolutional and max pooling layers, that even more complex features could be found that lead to even better model performance.\n* Only one hyperparameter, the batch size, was tuned in order to correct the issue of highly variable validation metrics. There are many more hyperpareters in the model however such as the learning rate, the kernel sizes, the number of hidden layers, or the pooling layer sizes to name a few. Further hyperparameter tuning could be performed on some of these to further improve the model. ","metadata":{}},{"cell_type":"markdown","source":"# Make Submission","metadata":{}},{"cell_type":"code","source":"import shutil\nsubmission_file = r\"/kaggle/input/cancerDetection/conv_dropout_256_model/test_labels.csv\"\nshutil.copy(submission_file, \"submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-08-06T19:35:22.180854Z","iopub.execute_input":"2023-08-06T19:35:22.181673Z","iopub.status.idle":"2023-08-06T19:35:22.194414Z","shell.execute_reply.started":"2023-08-06T19:35:22.181631Z","shell.execute_reply":"2023-08-06T19:35:22.193023Z"},"trusted":true},"execution_count":null,"outputs":[]}]}