{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":11848,"databundleVersionId":862157,"sourceType":"competition"}],"dockerImageVersionId":31012,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# CNN Cancer Detection Mini-Project\n\n#### DTSA 5511 Introduction to Deep Learning, Week 3\n\n&nbsp;\n\n<b>Instructions</b>: <i>This Kaggle competition is a binary image classification problem where you will identify metastatic cancer in small image patches taken from larger digital pathology scans.\n\nThe project has 125 total points. The instructions summarize the criteria you will use to guide your submission and review others' submissions.</i>","metadata":{}},{"cell_type":"markdown","source":"First, I'll import the libraries I need for this project.","metadata":{}},{"cell_type":"code","source":"import os\nimport time\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport pandas as pd\nimport numpy as np\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import confusion_matrix, classification_report\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator, load_img, img_to_array\nfrom tensorflow.keras.optimizers.schedules import ExponentialDecay\nfrom tensorflow.keras.applications import VGG16\nfrom tensorflow.keras import Input, layers, models, optimizers\nfrom keras import optimizers","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T03:07:05.679742Z","iopub.execute_input":"2025-04-19T03:07:05.680053Z","iopub.status.idle":"2025-04-19T03:07:05.686229Z","shell.execute_reply.started":"2025-04-19T03:07:05.680032Z","shell.execute_reply":"2025-04-19T03:07:05.685251Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Then, I'll import the data needed for the project. I'm limiting the dataset to just 5,000 images, since training on the full dataset took a very long time to complete.","metadata":{}},{"cell_type":"code","source":"# This is our folder with all of the data\ndata_path = \"/kaggle/input/histopathologic-cancer-detection\"\nprint(\"Files in dataset folder:\", os.listdir(data_path), '\\n')\n\n# The 'train' and 'test' folders have images to be classified by the model\ntrain_path = os.path.join(data_path, \"train\")\nprint(\"Sample images:\", os.listdir(train_path)[:5], '\\n')\n\n# train_labels.csv provides the labels for training images (1=cancerous, 0=non-cancerous)\nlabels = pd.read_csv(os.path.join(data_path, \"train_labels.csv\"))[:5000]\nprint(labels.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T01:56:12.842863Z","iopub.execute_input":"2025-04-19T01:56:12.843430Z","iopub.status.idle":"2025-04-19T01:56:15.372908Z","shell.execute_reply.started":"2025-04-19T01:56:12.843369Z","shell.execute_reply":"2025-04-19T01:56:15.371945Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Brief description of the problem and data (5 pts)\n\n<b>Instructions</b>: <i>Briefly describe the challenge problem and NLP. Describe the size, dimension, structure, etc., of the data.</i>","metadata":{}},{"cell_type":"markdown","source":"In this project, I am faced with the challenge problem of classifying images of tumor tissue as either cancerous or non-cancerous. This is a special case where computer vision techniques, like convolutional neural netowrks, could be well-equipped to solve this problem.\n\nFirst, let's look at a randomly-chosen image and it's label to get a sense for this dataset.","metadata":{}},{"cell_type":"code","source":"random_index = np.random.randint(0, len(labels))\n\nimage_id = labels.iloc[random_index]['id']\nlabel = labels.iloc[random_index]['label']\n\nimg_path = f'{data_path}/train/{image_id}.tif'\nimage = load_img(img_path)\n\nplt.imshow(image)\nplt.axis('off')\nplt.title(f\"Label: {label}\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T01:56:15.373995Z","iopub.execute_input":"2025-04-19T01:56:15.374346Z","iopub.status.idle":"2025-04-19T01:56:15.512802Z","shell.execute_reply.started":"2025-04-19T01:56:15.374301Z","shell.execute_reply":"2025-04-19T01:56:15.511490Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let's take a look at the size and shape of the data.","metadata":{}},{"cell_type":"code","source":"num_image_files = len([f for f in os.listdir(f'{data_path}/train/') if f.endswith('.tif')])\nprint(f\"Number of image files: {num_image_files}\")\n\nprint(f\"Labels dataset shape: {labels.shape}\")  # Expecting 5,000 rows for the training set","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T01:56:15.513652Z","iopub.execute_input":"2025-04-19T01:56:15.513937Z","iopub.status.idle":"2025-04-19T01:56:17.650667Z","shell.execute_reply.started":"2025-04-19T01:56:15.513915Z","shell.execute_reply":"2025-04-19T01:56:17.649731Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"first_image_file = [f for f in os.listdir(f'{data_path}/train/') if f.endswith('.tif')][0]\n\nimg_path = os.path.join(f'{data_path}/train/', first_image_file)\nimage = load_img(img_path)\n\nimage_array = img_to_array(image)\nprint(f\"Image shape: {image_array.shape}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T01:56:17.651639Z","iopub.execute_input":"2025-04-19T01:56:17.651956Z","iopub.status.idle":"2025-04-19T01:56:19.694856Z","shell.execute_reply.started":"2025-04-19T01:56:17.651934Z","shell.execute_reply":"2025-04-19T01:56:19.693966Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We looks like we have 220,025 images, but I'll only use 5,000 for training, since I don't have a GPU and training on the full 220,025 images was taking too long to complete. The images are 96 pixels by 96 pixels, with three channels (presumably R, G, B). The `labels` dataframe has two columns, one for the image ID, and the other with the label (1 or 0).\n\nIt's also worth noting the criteria for an image to be designated as 'cancerous' in this dataset: \"A positive label indicates that the center 32x32px region of a patch contains at least one pixel of tumor tissue. Tumor tissue in the outer region of the patch does not influence the label. This outer region is provided to enable fully-convolutional models that do not use zero-padding, to ensure consistent behavior when applied to a whole-slide image.\"","metadata":{}},{"cell_type":"markdown","source":"### Exploratory Data Analysis (EDA) — Inspect, Visualize and Clean the Data (15 pts)\n\n<b>Instructions</b>: <i>Show a few visualizations like histograms. Describe any data cleaning procedures. Based on your EDA, what is your plan of analysis?</i>","metadata":{}},{"cell_type":"markdown","source":"How balanced is our dataset? Do we have a majority of images that are cancerous or non-cancerous?","metadata":{}},{"cell_type":"code","source":"labels['label'].value_counts().sort_index().plot(kind='bar')\nplt.title('Distribution of Labels (Cancer Detection)')\nplt.xlabel('Label (0 = No Cancer, 1 = Cancer)')\nplt.ylabel('Number of Images')\nplt.xticks([0, 1])\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T01:56:19.695807Z","iopub.execute_input":"2025-04-19T01:56:19.696044Z","iopub.status.idle":"2025-04-19T01:56:19.878215Z","shell.execute_reply.started":"2025-04-19T01:56:19.696025Z","shell.execute_reply":"2025-04-19T01:56:19.877239Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(labels.describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T01:56:19.879221Z","iopub.execute_input":"2025-04-19T01:56:19.879909Z","iopub.status.idle":"2025-04-19T01:56:19.890599Z","shell.execute_reply.started":"2025-04-19T01:56:19.879885Z","shell.execute_reply":"2025-04-19T01:56:19.889700Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It looks like the dataset has more non-cancerous images (roughly 60% of the data) than cancerous images (roughly 40% of the data). However, it does not appear to be way out of balance. ","metadata":{}},{"cell_type":"markdown","source":"Are there any missing labels we need to handle?","metadata":{}},{"cell_type":"code","source":"labeled_ids = set(labels['id'])\nimage_ids_in_folder = set(f.replace('.tif', '') for f in os.listdir(f'{data_path}/train/') if f.endswith('.tif'))\nunlabeled_images = image_ids_in_folder - labeled_ids\nprint(f\"Number of image files without a label: {len(unlabeled_images)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T01:56:19.894016Z","iopub.execute_input":"2025-04-19T01:56:19.894553Z","iopub.status.idle":"2025-04-19T01:56:22.044501Z","shell.execute_reply.started":"2025-04-19T01:56:19.894528Z","shell.execute_reply":"2025-04-19T01:56:22.043332Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Great, I verified that all images in the training set have a label.","metadata":{}},{"cell_type":"markdown","source":"What are the ranges of values in the pixels of the images? We'll take a random image and plot a histogram of it's values in each channel.","metadata":{}},{"cell_type":"code","source":"random_index = np.random.randint(0, len(labels))\nimage_path = os.path.join(f'{data_path}/train/', os.listdir(f'{data_path}/train/')[random_index])\nimg = load_img(image_path)\nimg_array = img_to_array(img).astype(np.uint8)\n\nred_channel = img_array[:, :, 0].flatten()\ngreen_channel = img_array[:, :, 1].flatten()\nblue_channel = img_array[:, :, 2].flatten()\n\n# Plot histograms\nplt.hist(red_channel, bins=256, color='red', alpha=0.5)\nplt.hist(green_channel, bins=256, color='green', alpha=0.5)\nplt.hist(blue_channel, bins=256, color='blue', alpha=0.5)\nplt.title('Pixel Value Distribution from a Random Image')\nplt.xlabel('Pixel Intensity (0-255)')\nplt.ylabel('Frequency')\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T01:56:22.045332Z","iopub.execute_input":"2025-04-19T01:56:22.045637Z","iopub.status.idle":"2025-04-19T01:56:25.154548Z","shell.execute_reply.started":"2025-04-19T01:56:22.045614Z","shell.execute_reply":"2025-04-19T01:56:25.153576Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Looks like the values span from 0 to 255, which is typical for RGB images. However, the model is going to work better if we limit the values to a smaller range, while maintaining the important relationships within the data. I'll make sure to normalize our image data to the range [0, 1] when we train the model.","metadata":{}},{"cell_type":"markdown","source":"Lastly, Keras wants the labels to be strings, so I converted the label column in the dataframe to string.","metadata":{}},{"cell_type":"code","source":"labels['label'] = labels['label'].astype(str)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T01:56:25.155449Z","iopub.execute_input":"2025-04-19T01:56:25.155724Z","iopub.status.idle":"2025-04-19T01:56:25.162554Z","shell.execute_reply.started":"2025-04-19T01:56:25.155703Z","shell.execute_reply":"2025-04-19T01:56:25.161271Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Model Architecture (25 pts)\n\n<b>Instructions</b>: <i>Describe your model architecture and reasoning for why you believe that specific architecture would be suitable for this problem. Compare multiple architectures and tune hyperparameters.</i>","metadata":{}},{"cell_type":"markdown","source":"Now that we've explored the data, the data need to be split into training and validation sets for the model.","metadata":{}},{"cell_type":"code","source":"labels['id'] = labels['id'].astype(str) + '.tif'  # Add .tif to the end of each image ID in the dataframe\ntrain, val = train_test_split(labels, test_size=0.2, stratify=labels['label'], random_state=42)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T01:56:25.163584Z","iopub.execute_input":"2025-04-19T01:56:25.163937Z","iopub.status.idle":"2025-04-19T01:56:25.199879Z","shell.execute_reply.started":"2025-04-19T01:56:25.163913Z","shell.execute_reply":"2025-04-19T01:56:25.198543Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The data will be fed into the CNN using a generator. Loading all of the images we're going to train with into RAM all at once is overwhelming and slow. The generator avoids this issue by iteratively providing batches of images to the CNN. The generator also scales values from [0, 255] to [0, 1].","metadata":{}},{"cell_type":"code","source":"# This is the generator that will provide batches of images to the model\n# It also handles the normalization step with the rescale parameter\nimggen = ImageDataGenerator(rescale=1./255)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T01:56:25.201146Z","iopub.execute_input":"2025-04-19T01:56:25.201437Z","iopub.status.idle":"2025-04-19T01:56:25.207837Z","shell.execute_reply.started":"2025-04-19T01:56:25.201404Z","shell.execute_reply":"2025-04-19T01:56:25.206862Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Generator connected to the training set\ntrain_generator = imggen.flow_from_dataframe(\n    train,\n    directory=f'{data_path}/train/',\n    x_col='id',\n    y_col='label',\n    batch_size=16,\n    target_size=(96, 96),\n    color_mode='rgb',\n    class_mode='binary',\n    shuffle=False\n)\n\n# Generator connected to the validation set\nval_generator = imggen.flow_from_dataframe(\n    val,\n    directory=f'{data_path}/train/',\n    x_col='id',\n    y_col='label',\n    batch_size=16,\n    target_size=(96, 96),\n    color_mode='rgb',\n    class_mode='binary',\n    shuffle=False\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T01:56:25.208946Z","iopub.execute_input":"2025-04-19T01:56:25.209348Z","iopub.status.idle":"2025-04-19T01:56:28.962901Z","shell.execute_reply.started":"2025-04-19T01:56:25.209326Z","shell.execute_reply":"2025-04-19T01:56:28.961832Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"As a baseline, I'll train a simple, feed-forward artificial neural network and see how it performs.","metadata":{}},{"cell_type":"code","source":"simple_ann = models.Sequential([\n    Input(shape=(96, 96, 3)),\n    layers.Flatten(),\n    layers.Dense(512, activation='relu'),\n    layers.Dense(256, activation='relu'),\n    layers.Dense(1, activation='sigmoid')\n])\n\nsimple_ann.compile(\n    optimizer=optimizers.RMSprop(),\n    loss='binary_crossentropy',\n    metrics=['accuracy']\n)\n\nsimple_ann_start = time.time()\nsimple_ann_results = simple_ann.fit(\n    train_generator,\n    validation_data=val_generator,\n    steps_per_epoch=20,\n    epochs=5,\n)\nsimple_ann_end = time.time()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T01:56:28.964184Z","iopub.execute_input":"2025-04-19T01:56:28.964620Z","iopub.status.idle":"2025-04-19T01:57:11.749933Z","shell.execute_reply.started":"2025-04-19T01:56:28.964594Z","shell.execute_reply":"2025-04-19T01:57:11.748959Z"},"scrolled":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"As you can see, the simple, feed-forward neural network didn't perform too well. Since this is a binary classification problem, an accuracy of 0.5 would indicate a model has no skill. This neural network does better than that, with a 0.60 accuracy on the validation dataset. However, we see no indication that the model improves from one epoch to the next.\n \nIn this week's lectures, we learned about convolutional layers that can be added to neural networks. Convolutional layers apply filters that slide across the input image. They help detect small patterns in the image, like edges, textures, or shapes. These detected features are then passed on to deeper layers, which allow the model to learn increasingly complex representations. I'll try including this in my model architecture:","metadata":{}},{"cell_type":"code","source":"simple_cnn = models.Sequential([\n    layers.Input(shape=(96, 96, 3)),\n    layers.Conv2D(32, (3, 3), activation='relu'),\n    layers.MaxPooling2D(pool_size=(2, 2)),\n    layers.Conv2D(64, (3, 3), activation='relu'),\n    layers.MaxPooling2D(pool_size=(2, 2)),\n    layers.Flatten(),\n    layers.Dense(128, activation='relu'),\n    layers.Dense(1, activation='sigmoid')\n])\n\nsimple_cnn.compile(\n    optimizer=optimizers.RMSprop(),\n    loss='binary_crossentropy',\n    metrics=['accuracy']\n)\n\nsimple_cnn.fit(\n    train_generator,\n    validation_data=val_generator,\n    steps_per_epoch=10,\n    epochs=20,\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T01:57:11.751172Z","iopub.execute_input":"2025-04-19T01:57:11.751869Z","iopub.status.idle":"2025-04-19T01:58:34.931402Z","shell.execute_reply.started":"2025-04-19T01:57:11.751834Z","shell.execute_reply":"2025-04-19T01:58:34.930310Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Ok, we're seeing some progress! I introduced two convolutional layers, followed by two max pooling layers, plus a single dense layer at the end before providing the output. The validation accuracy is getting better. It's above 0.60 for most epochs. However, it bounces around quite a bit from epoch to epoch. This is typically a sign that the learning rate is too high. I'll lower the learning rate and see if that helps.","metadata":{}},{"cell_type":"code","source":"simple_cnn = models.Sequential([\n    layers.Input(shape=(96, 96, 3)),\n    layers.Conv2D(32, (3, 3), activation='relu'),\n    layers.MaxPooling2D(pool_size=(2, 2)),\n    layers.Conv2D(64, (3, 3), activation='relu'),\n    layers.MaxPooling2D(pool_size=(2, 2)),\n    layers.Flatten(),\n    layers.Dense(128, activation='relu'),\n    layers.Dense(1, activation='sigmoid')\n])\n\nsimple_cnn.compile(\n    optimizer=optimizers.RMSprop(learning_rate=0.00001),\n    loss='binary_crossentropy',\n    metrics=['accuracy']\n)\n\nsimple_cnn_start = time.time()\nsimple_cnn_results = simple_cnn.fit(\n    train_generator,\n    validation_data=val_generator,\n    steps_per_epoch=40,\n    epochs=75,\n)\nsimple_cnn_end = time.time()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T01:58:34.932531Z","iopub.execute_input":"2025-04-19T01:58:34.932943Z","iopub.status.idle":"2025-04-19T02:08:21.432078Z","shell.execute_reply.started":"2025-04-19T01:58:34.932919Z","shell.execute_reply":"2025-04-19T02:08:21.430892Z"},"scrolled":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The default learning rate in Keras is 0.001. I made it 0.00001. The validation accuracy doesn't swing as wildy between each epoch, but it is still not smooth either. Also, it's important to consider that a slower learning rate requires more epochs to converge. It would be helpful to plot the accuracy from each epoch to understand whether the model is converging.","metadata":{}},{"cell_type":"code","source":"plt.plot(simple_cnn_results.history['accuracy'], label='Training Accuracy')\nplt.plot(simple_cnn_results.history['val_accuracy'], label='Validation Accuracy')\nplt.xlabel('Epochs')\nplt.ylabel('Accuracy')\nplt.title('Training and Validation Accuracy')\nplt.legend()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T02:08:21.433234Z","iopub.execute_input":"2025-04-19T02:08:21.433625Z","iopub.status.idle":"2025-04-19T02:08:21.652686Z","shell.execute_reply.started":"2025-04-19T02:08:21.433592Z","shell.execute_reply":"2025-04-19T02:08:21.651739Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It looks like the model is starting to level off around 50 to 70 epochs. Since I have limited computing resources, and especially since I don't have a GPU, it's easier for me to work with models that train faster. Training for 50 epochs or more can take several minutes.","metadata":{}},{"cell_type":"markdown","source":"I also want to try batch normalization. This is a technique that normalizes the outputs of layers before activations, which helps activations from getting too large or too small. This can stabilize learning, allowing for higher learning rates and fewer epochs. Let's see how it performs:","metadata":{}},{"cell_type":"code","source":"batch_cnn = models.Sequential([\n    layers.Input(shape=(96, 96, 3)),\n    \n    layers.Conv2D(32, (3, 3)),\n    layers.BatchNormalization(),\n    layers.Activation('relu'),\n    layers.MaxPooling2D(pool_size=(2, 2)),\n    \n    layers.Conv2D(64, (3, 3)),\n    layers.BatchNormalization(),\n    layers.Activation('relu'),\n    layers.MaxPooling2D(pool_size=(2, 2)),\n    \n    layers.Flatten(),\n    \n    layers.Dense(128),\n    layers.BatchNormalization(),\n    layers.Activation('relu'),\n    \n    layers.Dense(1, activation='sigmoid')\n])\n\nbatch_cnn.compile(\n    optimizer=optimizers.RMSprop(learning_rate=0.00001),\n    loss='binary_crossentropy',\n    metrics=['accuracy']\n)\n\nbatch_cnn_results = batch_cnn.fit(\n    train_generator,\n    validation_data=val_generator,\n    steps_per_epoch=40,\n    epochs=75,\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T02:21:19.461971Z","iopub.execute_input":"2025-04-19T02:21:19.462259Z","iopub.status.idle":"2025-04-19T02:34:25.481724Z","shell.execute_reply.started":"2025-04-19T02:21:19.462239Z","shell.execute_reply":"2025-04-19T02:34:25.478839Z"},"scrolled":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.plot(batch_cnn_results.history['accuracy'], label='Training Accuracy')\nplt.plot(batch_cnn_results.history['val_accuracy'], label='Validation Accuracy')\nplt.xlabel('Epochs')\nplt.ylabel('Accuracy')\nplt.title('Training and Validation Accuracy')\nplt.legend()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T02:34:25.483634Z","iopub.execute_input":"2025-04-19T02:34:25.483997Z","iopub.status.idle":"2025-04-19T02:34:25.713246Z","shell.execute_reply.started":"2025-04-19T02:34:25.483947Z","shell.execute_reply":"2025-04-19T02:34:25.711938Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Batch normalization certainly helped this model reach convergence faster, after just 15 or 20 epochs. I'll limit the model to 20 epochs so that we avoid overfitting.","metadata":{}},{"cell_type":"code","source":"batch_cnn = models.Sequential([\n    layers.Input(shape=(96, 96, 3)),\n    \n    layers.Conv2D(32, (3, 3)),\n    layers.BatchNormalization(),\n    layers.Activation('relu'),\n    layers.MaxPooling2D(pool_size=(2, 2)),\n    \n    layers.Conv2D(64, (3, 3)),\n    layers.BatchNormalization(),\n    layers.Activation('relu'),\n    layers.MaxPooling2D(pool_size=(2, 2)),\n    \n    layers.Flatten(),\n    \n    layers.Dense(128),\n    layers.BatchNormalization(),\n    layers.Activation('relu'),\n    \n    layers.Dense(1, activation='sigmoid')\n])\n\nbatch_cnn.compile(\n    optimizer=optimizers.RMSprop(learning_rate=0.00001),\n    loss='binary_crossentropy',\n    metrics=['accuracy']\n)\n\nbatch_cnn_start = time.time()\nbatch_cnn_results = batch_cnn.fit(\n    train_generator,\n    validation_data=val_generator,\n    steps_per_epoch=40,\n    epochs=20,\n)\nbatch_cnn_end = time.time()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T02:34:25.715345Z","iopub.execute_input":"2025-04-19T02:34:25.715943Z","iopub.status.idle":"2025-04-19T02:37:53.197459Z","shell.execute_reply.started":"2025-04-19T02:34:25.715895Z","shell.execute_reply":"2025-04-19T02:37:53.196529Z"},"scrolled":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The model ended up with a validation accuracy of 0.65, which is much better than the simple neural network, which had a validation accuracy of 0.60. I was also able to achieve this in 20 epochs, which is a manageable number of epochs for my computer.","metadata":{}},{"cell_type":"markdown","source":"I'm also curious whether a pretrained model can perform better than this. We talked about VGG and ResNet in the lectures. Let's see how they perform on this task:","metadata":{}},{"cell_type":"code","source":"vgg = VGG16(\n    weights='imagenet', \n    include_top=False, \n    input_shape=(96, 96, 3)\n)\nvgg.trainable = False\n\nvgg = models.Sequential([\n    vgg,\n    layers.Flatten(),\n    layers.Dense(128),\n    layers.BatchNormalization(),\n    layers.Activation('relu'),\n    layers.Dense(1, activation='sigmoid')\n])\n\nvgg.compile(\n    optimizer=optimizers.RMSprop(learning_rate=0.00001),\n    loss='binary_crossentropy',\n    metrics=['accuracy']\n)\n\nvgg_start = time.time()\nvgg_history = vgg.fit(\n    train_generator,\n    validation_data=val_generator,\n    epochs=10,\n    steps_per_epoch=40,\n    validation_steps=20\n)\nvgg_end = time.time()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T02:37:53.199521Z","iopub.execute_input":"2025-04-19T02:37:53.199834Z","iopub.status.idle":"2025-04-19T02:45:43.933491Z","shell.execute_reply.started":"2025-04-19T02:37:53.199811Z","shell.execute_reply":"2025-04-19T02:45:43.932318Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Training the VGG model was even slower than the models I built myself, but that is to be expected with a pretrained model. VGG and other pretrained models take advantage of transfer learning. These models were trained on a large dataset, and then can be adapted for other tasks like this one by reusing it's learned features. However, this technique requires a large amount of data to train with in order to see effective results.\n\nThe validation accuracy in this VGG model I attempted to use above is very erratic, but also generally between 0.7 and 0.9 for validation accuracy. These results provide promise for folks who can train with more data and use better processors, like GPUs. In my case, this pretrained model ended with a validation accuracy of 0.73, which is slightly better than my \"homemade\" convolutional neural network.","metadata":{}},{"cell_type":"markdown","source":"### Results and Analysis (35 pts)\n\n<b>Instructions</b>: <i>Run hyperparameter tuning, try different architectures for comparison, apply techniques to improve training or performance, and discuss what helped.\n\nIncludes results with tables and figures. There is an analysis of why or why not something worked well, troubleshooting, and a hyperparameter optimization procedure summary.</i>","metadata":{}},{"cell_type":"markdown","source":"I already went through iterations of model architecture and tuning hyperparamters, like learning rates, in the preivous section, so I'll focus more on a discussion of the accuracy and performance of the models I trained above.\n\nHere's a summary of the accuracy and training time for each model:","metadata":{}},{"cell_type":"code","source":"model_summary = pd.DataFrame()\nmodel_summary['Model'] = ['Simple Artificial NN', 'Convolutional NN', 'Convolutional NN w/ Batch Norm.', 'VGG Pre-trained NN']\nmodel_summary['Accuracy'] = [simple_ann_results.history['val_accuracy'][-1],\n                             simple_cnn_results.history['val_accuracy'][-1],\n                             batch_cnn_results.history['val_accuracy'][-1],\n                             vgg_history.history['val_accuracy'][-1]]\nmodel_summary['Training Time (sec)'] = [simple_ann_end - simple_ann_start,\n                                        simple_cnn_end - simple_cnn_start,\n                                        batch_cnn_end - batch_cnn_start,\n                                        vgg_end - vgg_start]\nmodel_summary['Training Batches'] = [5, 75, 20, 10]\nmodel_summary['Training Time per Batch (sec)'] = model_summary['Training Time (sec)'] / model_summary['Training Batches']\nmodel_summary","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T02:45:43.934531Z","iopub.execute_input":"2025-04-19T02:45:43.934822Z","iopub.status.idle":"2025-04-19T02:45:43.961708Z","shell.execute_reply.started":"2025-04-19T02:45:43.934800Z","shell.execute_reply":"2025-04-19T02:45:43.960563Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The artificial neural network only implemented dense, feedforward layers, and provided a baseline for the rest of the models I trained. It's accuracy was the poorest, since we know that dense layers can't learn image features well. That's why I then created a convolutional neural network, which performed much better. However, I struggled to get the training to be more stable. I tried a convolutional neural network with batch normalization before each activation. This allowed the model to converge much sooner, which was great. However, it didn't do much to stabilize the training process. I also tried a pretrained model, which had a performance in between my two homemade CNNs, but was also unstable during training, and probably needed more data and computing resources to improve upon.\n\nI tweaked hyperparameters throughout the process, mostly focusing on learning rate and number of epochs until I found a good balance, but I also attempted some work with momentum, decay, and dropout. The balancing act between learning rate and epochs was key to getting the model to converge, but without overfitting.\n\nLet's take a closer look at the performance of the first convolutional neural network I created. I made predictions on the validation set and then put them into a confusion matrix.","metadata":{}},{"cell_type":"code","source":"y_true = val_generator.classes\ny_pred_prob = simple_cnn.predict(val_generator)\ny_pred = (y_pred_prob > 0.5).astype(int).flatten()\ncm = confusion_matrix(y_true, y_pred)\nprint(classification_report(y_true, y_pred))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T03:16:03.180107Z","iopub.execute_input":"2025-04-19T03:16:03.180562Z","iopub.status.idle":"2025-04-19T03:16:05.982279Z","shell.execute_reply.started":"2025-04-19T03:16:03.180534Z","shell.execute_reply":"2025-04-19T03:16:05.981267Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sns.heatmap(cm, annot=True, fmt='d', cmap='Greens')\nplt.xlabel('Predicted')\nplt.ylabel('Actual')\nplt.title('Simple CNN Confusion Matrix')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T03:06:04.173542Z","iopub.execute_input":"2025-04-19T03:06:04.173874Z","iopub.status.idle":"2025-04-19T03:06:04.372594Z","shell.execute_reply.started":"2025-04-19T03:06:04.173850Z","shell.execute_reply":"2025-04-19T03:06:04.371591Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This was my best model, which a validation accuracy of 0.77. We know that the dataset is unbalanced, so there's more class 0 than class 1 samples in the dataset.  The precision and recall are generally around 0.77 as well, so the accuracy measure is a good indicator of the overall skill of the model (it's not the case here that one class is being predicted extremley well, and the other extremeley poorly, which is a situation where accuracy can be a misleading metric).\n\nI'll move forward using the first convolutional neural network and use it to make predictions for the Kaggle challenge.","metadata":{}},{"cell_type":"code","source":"# Get the names of the images in the test dataset\ntest_image_dir = f'{data_path}/test/'\ntest_df = pd.DataFrame({'id': os.listdir(test_image_dir)})","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T02:45:43.962877Z","iopub.execute_input":"2025-04-19T02:45:43.963227Z","iopub.status.idle":"2025-04-19T02:45:44.697436Z","shell.execute_reply.started":"2025-04-19T02:45:43.963197Z","shell.execute_reply":"2025-04-19T02:45:44.696198Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_datagen = ImageDataGenerator(rescale=1./255)\n\n# This generator will serve test images to the model to make predictions\ntest_generator = test_datagen.flow_from_dataframe(\n    dataframe=test_df,\n    directory=f'{data_path}/test/',\n    x_col='id',\n    y_col=None,\n    target_size=(96, 96),\n    color_mode='rgb',\n    batch_size=16,\n    class_mode=None,\n    shuffle=False\n)\n\n# Get predictions from test images\npredictions = simple_cnn.predict(test_generator)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T02:45:44.698557Z","iopub.execute_input":"2025-04-19T02:45:44.698814Z","iopub.status.idle":"2025-04-19T02:54:10.772189Z","shell.execute_reply.started":"2025-04-19T02:45:44.698795Z","shell.execute_reply":"2025-04-19T02:54:10.771113Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Bring predictions together which \npredicted_classes = (predictions > 0.5).astype(\"int32\")\ntest_df['id'] = test_df['id'].str[:-4]\ntest_df['label'] = predicted_classes\nprint(test_df.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T02:58:02.058035Z","iopub.execute_input":"2025-04-19T02:58:02.058455Z","iopub.status.idle":"2025-04-19T02:58:02.082485Z","shell.execute_reply.started":"2025-04-19T02:58:02.058419Z","shell.execute_reply":"2025-04-19T02:58:02.081509Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Output CSV for submission to competition\ntest_df.to_csv('submission.csv', index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-19T02:58:55.153260Z","iopub.execute_input":"2025-04-19T02:58:55.153697Z","iopub.status.idle":"2025-04-19T02:58:55.275365Z","shell.execute_reply.started":"2025-04-19T02:58:55.153674Z","shell.execute_reply":"2025-04-19T02:58:55.274512Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Conclusion (15 pts)\n\n<b>Instructions</b>: <i>Discuss and interpret results as well as learnings and takeaways. What did and did not help improve the performance of your models? What improvements could you try in the future?</i>","metadata":{}},{"cell_type":"markdown","source":"Overall, the results showed me the power of convolutional neural networks (CNN). All of the CNNs I trained outperformed the basic artificial neural network, since convolutional layers can pick out patterns from images that a typical dense layer cannot. \n\nHere's what helped improve the performance of my models (and some lessons learned):\n* Including convolutional layers to help with the task of training on image input\n* Batch normalization allowed my neural network to converge faster, so I could prototype more quickly to make further improvements\n* A lower learning rate is key to finding a minimum without bouncing around and overshooting the minimum of the loss function\n* More epochs might be needed if the model is not converging, especially if the learning rate is slow\n\nHere's what didn't help improve the performance of my models (and some lessons learned):\n* I was limited to a CPU, so training many epochs across many models took a while and slowed down my progress\n* I limited my training set to 5,000 images so that I could train faster, probably at the expense of some model skill and stability\n* I couldn't get the momentum parameter on the optomizer to make any meaningful difference in model performance\n\nHere's what I'd like to try in the future:\n* I'd like to work with a GPU in the future so I can train models faster and more easily prototype improvements\n* I'd like to find an example project where momentum could be more helpful, because I did not see it improve performance when I tried it here\n* I'd also like to experiment with data augmentation, which allows you to duplicate images in the training set, except rotated or resized, in order to improve model performance","metadata":{}}]}