{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":11848,"databundleVersionId":862157,"sourceType":"competition"}],"dockerImageVersionId":30616,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\n\nThis project centers on the application of Convolutional Neural Networks (CNN) to Histopathologic Cancer Detection, a critical task in medical diagnostics. Histopathologic analysis involves examining tissue samples under a microscope to identify signs of disease, such as cancer. The use of CNN in this domain aims to automate and enhance the accuracy of detecting cancer cells in histopathology images.\n\nThe challenge at hand is to develop a CNN model capable of discerning subtle patterns and features in histopathologic slides that are indicative of cancer. This is a nuanced task, given the complexity of tissue structures and the subtle variations that distinguish benign from malignant cells.\n\nThe dataset (Cukierski, 2018) employed for this endeavor is sourced from a Kaggle competition, specifically tailored for Histopathologic Cancer Detection. It comprises thousands of annotated high-resolution images of lymph node sections. Each image has been meticulously labeled, marking the presence of metastatic tissue – a crucial indicator of cancer.","metadata":{}},{"cell_type":"markdown","source":"# Explaratory Data Analysis and Cleanup\n\nIn this section, we'll try to explore the data as well as prepare it to be used for the models.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom PIL import Image\nimport tensorflow as tf\nimport os\n\nBASE_DIR = '/kaggle/input/histopathologic-cancer-detection/'\n\ntrain_df = pd.read_csv(BASE_DIR + 'train_labels.csv')\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-13T04:23:57.149837Z","iopub.execute_input":"2023-12-13T04:23:57.150115Z","iopub.status.idle":"2023-12-13T04:24:10.499191Z","shell.execute_reply.started":"2023-12-13T04:23:57.150085Z","shell.execute_reply":"2023-12-13T04:24:10.498064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2023-12-13T07:48:12.999125Z","iopub.execute_input":"2023-12-13T07:48:12.999534Z","iopub.status.idle":"2023-12-13T07:48:13.044209Z","shell.execute_reply.started":"2023-12-13T07:48:12.999502Z","shell.execute_reply":"2023-12-13T07:48:13.043320Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The training CSV file contains a total of **220,025** sets of IDs and labels.\n\nThe IDs here correspond to the name of the image file in the adjacent folder, while the labels indicate whether the image with given ID is of a cancerous cell or not.","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize=(30, 6))\ntrain_imgs = os.listdir(BASE_DIR+\"train\")\nfor idx, img in enumerate(np.random.choice(train_imgs, 20)):\n    ax = fig.add_subplot(2, 20//2, idx+1, xticks=[], yticks=[])\n    im = Image.open(BASE_DIR+\"train/\" + img)\n    plt.imshow(im)\n    label = train_df.loc[train_df['id'] == img.split('.')[0], 'label'].values[0]\n    ax.set_title(f'Label: {label}')","metadata":{"execution":{"iopub.status.busy":"2023-12-13T04:24:10.500941Z","iopub.execute_input":"2023-12-13T04:24:10.501264Z","iopub.status.idle":"2023-12-13T04:24:15.810436Z","shell.execute_reply.started":"2023-12-13T04:24:10.501238Z","shell.execute_reply":"2023-12-13T04:24:15.809410Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see some random images here from the training set. It is very difficult, especially for someone like me without a medical background, to tell which ones are cancerous and which ones aren't without the labels.","metadata":{}},{"cell_type":"code","source":"np.array(im).shape","metadata":{"execution":{"iopub.status.busy":"2023-12-13T04:24:15.811856Z","iopub.execute_input":"2023-12-13T04:24:15.812378Z","iopub.status.idle":"2023-12-13T04:24:15.818027Z","shell.execute_reply.started":"2023-12-13T04:24:15.812348Z","shell.execute_reply":"2023-12-13T04:24:15.817131Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['label'].value_counts().plot(kind='pie', legend=True, autopct='%1.1f%%')","metadata":{"execution":{"iopub.status.busy":"2023-12-13T04:24:15.820323Z","iopub.execute_input":"2023-12-13T04:24:15.820583Z","iopub.status.idle":"2023-12-13T04:24:15.988827Z","shell.execute_reply.started":"2023-12-13T04:24:15.820557Z","shell.execute_reply":"2023-12-13T04:24:15.987922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It appears like the data provided for training purposes contains 60% values with label 0 and 40% with 1.","metadata":{}},{"cell_type":"markdown","source":"Since this project is going to use keras, we will need to prepare a keras dataset. Keras utils method `image_dataset_from_directory` does not support tif files at the time of this writing. As such, the files need to be converted to something supported by the method.","metadata":{}},{"cell_type":"code","source":"def convert_image(image_file, label=0, with_label=True, subset='train'):\n    image_name = image_file.split('.')[0]\n    \n    if with_label:\n#         label = train_df[train_df['id'] == image_name]['label'].iloc[0]\n        output_dir = f'png/{subset}/{label}'\n    else:\n        output_dir = f'png/{subset}'\n\n    os.makedirs(output_dir, exist_ok=True)\n\n    output_path = f'{output_dir}/{image_name}.png'\n    if not os.path.exists(output_path):\n        with Image.open(BASE_DIR + f'{subset}/{image_file}') as tiff_img:\n            png = tiff_img.convert(\"RGB\")\n            png.save(output_path)","metadata":{"execution":{"iopub.status.busy":"2023-12-13T04:24:15.990560Z","iopub.execute_input":"2023-12-13T04:24:15.991213Z","iopub.status.idle":"2023-12-13T04:24:15.998941Z","shell.execute_reply.started":"2023-12-13T04:24:15.991178Z","shell.execute_reply":"2023-12-13T04:24:15.997945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import multiprocessing\nfrom tqdm import tqdm\n\ndef convert_image_wrapper(args):\n    return convert_image(*args)\n\ndef process_images(image_files, labels=[], with_label=True, subset='train'):\n    num_processes = multiprocessing.cpu_count()\n    \n    with multiprocessing.Pool(processes=num_processes) as pool:\n        tasks = [(filename, labels[index] if with_label else 0, with_label, subset) for index, filename in enumerate(image_files)]\n        for _ in tqdm(pool.imap_unordered(convert_image_wrapper, tasks), total=len(tasks)):\n            pass","metadata":{"execution":{"iopub.status.busy":"2023-12-13T04:24:16.000377Z","iopub.execute_input":"2023-12-13T04:24:16.000717Z","iopub.status.idle":"2023-12-13T04:24:16.015402Z","shell.execute_reply.started":"2023-12-13T04:24:16.000686Z","shell.execute_reply":"2023-12-13T04:24:16.014537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This process takes some time (10+ minutes)\nprocess_images((train_df['id']+'.tif').values.tolist(), labels=train_df['label'].values.tolist())","metadata":{"execution":{"iopub.status.busy":"2023-12-13T04:24:16.016997Z","iopub.execute_input":"2023-12-13T04:24:16.017397Z","iopub.status.idle":"2023-12-13T04:33:32.085099Z","shell.execute_reply.started":"2023-12-13T04:24:16.017367Z","shell.execute_reply":"2023-12-13T04:33:32.084001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"process_images(os.listdir(BASE_DIR+\"test\"), with_label=False, subset='test')","metadata":{"execution":{"iopub.status.busy":"2023-12-13T04:33:32.086536Z","iopub.execute_input":"2023-12-13T04:33:32.086873Z","iopub.status.idle":"2023-12-13T04:36:19.571041Z","shell.execute_reply.started":"2023-12-13T04:33:32.086840Z","shell.execute_reply":"2023-12-13T04:36:19.569938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we should have PNG equivalent of the training data separated into different folders based on their classes. We should be able to load the datasets.","metadata":{}},{"cell_type":"code","source":"train_dataset = tf.keras.utils.image_dataset_from_directory('/kaggle/working/png/train', \n                                                            label_mode='binary',\n                                                            image_size=(96,96), \n                                                            seed=42,\n                                                            validation_split=0.2,\n                                                            subset='training',\n                                                            batch_size=128)\n\nval_dataset = tf.keras.utils.image_dataset_from_directory('/kaggle/working/png/train', \n                                                            label_mode='binary',\n                                                            image_size=(96,96), \n                                                            seed=42,\n                                                            validation_split=0.2,\n                                                            subset='validation',\n                                                            batch_size=128)\n","metadata":{"execution":{"iopub.status.busy":"2023-12-13T04:36:19.572567Z","iopub.execute_input":"2023-12-13T04:36:19.572911Z","iopub.status.idle":"2023-12-13T04:36:47.082145Z","shell.execute_reply.started":"2023-12-13T04:36:19.572878Z","shell.execute_reply":"2023-12-13T04:36:47.081191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_dataset = tf.keras.utils.image_dataset_from_directory('/kaggle/working/png/test',\n                                                            label_mode=None,\n                                                            image_size=(96,96),\n                                                            shuffle=False,\n                                                            batch_size=1)","metadata":{"execution":{"iopub.status.busy":"2023-12-13T04:36:47.084704Z","iopub.execute_input":"2023-12-13T04:36:47.085029Z","iopub.status.idle":"2023-12-13T04:36:49.762169Z","shell.execute_reply.started":"2023-12-13T04:36:47.085002Z","shell.execute_reply":"2023-12-13T04:36:49.761098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Developing the Model\n\nSince we have the datasets ready, we can continue with model development.\n\nI decided to first trying to build a completely custom model. The model consists of three convolutional layers for feature extraction followed by three dense layers for classification. \n\nThe convolutional layers have max pooling in between them. \n\nThe dense layers utilize the ReLU activation function, except for the output layer which uses sigmoid.","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras import layers, models\n\nmodel = tf.keras.Sequential()\nmodel.add(layers.Rescaling(1./255, input_shape=(96,96,3)))\nmodel.add(layers.Conv2D(32, (2, 2), strides=(2,2), activation='relu'))\nmodel.add(layers.MaxPooling2D((2, 2)))\nmodel.add(layers.Conv2D(64, (2, 2), strides=(2,2), activation='relu'))\nmodel.add(layers.MaxPooling2D((2, 2)))\nmodel.add(layers.Conv2D(128, (2, 2), strides=(2,2), activation='relu'))\nmodel.add(layers.Flatten())\nmodel.add(layers.Dense(64, activation='relu'))\nmodel.add(layers.Dense(32, activation='relu'))\nmodel.add(layers.Dense(1, activation='sigmoid'))\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2023-12-13T06:18:05.149448Z","iopub.execute_input":"2023-12-13T06:18:05.150134Z","iopub.status.idle":"2023-12-13T06:18:05.268830Z","shell.execute_reply.started":"2023-12-13T06:18:05.150101Z","shell.execute_reply":"2023-12-13T06:18:05.267907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.compile(optimizer='adam',\n              loss=tf.keras.losses.BinaryCrossentropy(),\n              metrics=['accuracy'])","metadata":{"execution":{"iopub.status.busy":"2023-12-13T06:42:07.335818Z","iopub.execute_input":"2023-12-13T06:42:07.336213Z","iopub.status.idle":"2023-12-13T06:42:07.355357Z","shell.execute_reply.started":"2023-12-13T06:42:07.336181Z","shell.execute_reply":"2023-12-13T06:42:07.354307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = model.fit(train_dataset, epochs=10, validation_data=val_dataset)","metadata":{"execution":{"iopub.status.busy":"2023-12-13T06:42:07.761705Z","iopub.execute_input":"2023-12-13T06:42:07.762024Z","iopub.status.idle":"2023-12-13T06:53:57.547777Z","shell.execute_reply.started":"2023-12-13T06:42:07.761991Z","shell.execute_reply":"2023-12-13T06:53:57.546949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(history.history['accuracy'], label='accuracy')\nplt.plot(history.history['val_accuracy'], label = 'val_accuracy')\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.ylim([0.5, 1])\nplt.title('Train vs Validation Accuracy Per Epoch')\nplt.legend(loc='lower right')","metadata":{"execution":{"iopub.status.busy":"2023-12-13T06:54:08.185362Z","iopub.execute_input":"2023-12-13T06:54:08.185782Z","iopub.status.idle":"2023-12-13T06:54:08.509628Z","shell.execute_reply.started":"2023-12-13T06:54:08.185726Z","shell.execute_reply":"2023-12-13T06:54:08.508684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(history.history['loss'], label='loss')\nplt.plot(history.history['val_loss'], label = 'val_loss')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\nplt.title('Train vs Validation Loss Per Epoch')","metadata":{"execution":{"iopub.status.busy":"2023-12-13T06:54:15.741916Z","iopub.execute_input":"2023-12-13T06:54:15.742376Z","iopub.status.idle":"2023-12-13T06:54:16.661011Z","shell.execute_reply.started":"2023-12-13T06:54:15.742345Z","shell.execute_reply":"2023-12-13T06:54:16.660106Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With __86%__ accuracy in the validation set, the model performed reasonably well. Based on the graphs, we can see that the model might be overfitting after the fourth epoch, as validation loss and accuracy both started to decrease from that point. ","metadata":{}},{"cell_type":"markdown","source":"Now lets create a submission entry for the competition. This requires a CSV file with ID and labels based on the provided test data.","metadata":{}},{"cell_type":"code","source":"test_imgs = os.listdir(\"/kaggle/working/png/test\")\nmodel1_pred_df = pd.DataFrame(columns=['id', 'label'])\ntest_imgs=sorted(test_imgs)","metadata":{"execution":{"iopub.status.busy":"2023-12-13T06:54:54.623420Z","iopub.execute_input":"2023-12-13T06:54:54.623802Z","iopub.status.idle":"2023-12-13T06:54:54.692387Z","shell.execute_reply.started":"2023-12-13T06:54:54.623772Z","shell.execute_reply":"2023-12-13T06:54:54.691443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = model.predict(test_dataset)","metadata":{"execution":{"iopub.status.busy":"2023-12-13T06:54:55.309009Z","iopub.execute_input":"2023-12-13T06:54:55.309838Z","iopub.status.idle":"2023-12-13T06:56:58.638987Z","shell.execute_reply.started":"2023-12-13T06:54:55.309802Z","shell.execute_reply":"2023-12-13T06:56:58.638057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model1_pred_df['id'] = [filename.split('.')[0] for filename in test_imgs]\nmodel1_pred_df['label'] = np.round(predictions.flatten()).astype('int')\nmodel1_pred_df","metadata":{"execution":{"iopub.status.busy":"2023-12-13T06:56:58.649029Z","iopub.execute_input":"2023-12-13T06:56:58.649351Z","iopub.status.idle":"2023-12-13T06:56:58.695646Z","shell.execute_reply.started":"2023-12-13T06:56:58.649323Z","shell.execute_reply":"2023-12-13T06:56:58.694780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model1_pred_df.to_csv('custom_model_predictions.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-12-13T06:56:58.696910Z","iopub.execute_input":"2023-12-13T06:56:58.697264Z","iopub.status.idle":"2023-12-13T06:56:58.898882Z","shell.execute_reply.started":"2023-12-13T06:56:58.697229Z","shell.execute_reply":"2023-12-13T06:56:58.898111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After submitting the model to the competition, it was observed that it achieved a public score of **0.8239**.","metadata":{}},{"cell_type":"markdown","source":"Then I ventured to developing a model with ImageNet as the base model. Using ImageNet's weights, the dense layers were applied similar to the custom model.","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.applications import VGG16\nbase_model = tf.keras.applications.MobileNetV2(input_shape=(96,96,3),\n                                               include_top=False,\n                                               weights='imagenet')\nbase_model.trainable = False\nbase_model.summary()","metadata":{"execution":{"iopub.status.busy":"2023-12-13T06:18:35.820679Z","iopub.execute_input":"2023-12-13T06:18:35.821426Z","iopub.status.idle":"2023-12-13T06:18:37.560609Z","shell.execute_reply.started":"2023-12-13T06:18:35.821387Z","shell.execute_reply":"2023-12-13T06:18:37.559667Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"global_average_layer = tf.keras.layers.GlobalAveragePooling2D()\nprediction_layer = tf.keras.layers.Dense(1, activation='sigmoid')\n\ninputs = tf.keras.Input(shape=(96, 96, 3))\nx = base_model(inputs, training=False)\nx = global_average_layer(x)\nx = tf.keras.layers.Dropout(0.2)(x)\nx = layers.Dense(64, activation='relu')(x)\nx = layers.Dense(32, activation='relu')(x)\noutputs = prediction_layer(x)\nimagenet_model = tf.keras.Model(inputs, outputs)\nimagenet_model.summary()","metadata":{"execution":{"iopub.status.busy":"2023-12-13T07:03:23.638123Z","iopub.execute_input":"2023-12-13T07:03:23.638860Z","iopub.status.idle":"2023-12-13T07:03:24.193135Z","shell.execute_reply.started":"2023-12-13T07:03:23.638821Z","shell.execute_reply":"2023-12-13T07:03:24.192305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"base_learning_rate = 0.001\nimagenet_model.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=base_learning_rate),\n              loss=tf.keras.losses.BinaryCrossentropy(),\n              metrics=['accuracy'])","metadata":{"execution":{"iopub.status.busy":"2023-12-13T07:05:00.022909Z","iopub.execute_input":"2023-12-13T07:05:00.023274Z","iopub.status.idle":"2023-12-13T07:05:00.040192Z","shell.execute_reply.started":"2023-12-13T07:05:00.023246Z","shell.execute_reply":"2023-12-13T07:05:00.039151Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"imagenet_history = imagenet_model.fit(train_dataset, epochs=10, validation_data=val_dataset)","metadata":{"execution":{"iopub.status.busy":"2023-12-13T07:05:01.582510Z","iopub.execute_input":"2023-12-13T07:05:01.582890Z","iopub.status.idle":"2023-12-13T07:17:51.695418Z","shell.execute_reply.started":"2023-12-13T07:05:01.582859Z","shell.execute_reply":"2023-12-13T07:17:51.694593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(imagenet_history.history['accuracy'], label='accuracy')\nplt.plot(imagenet_history.history['val_accuracy'], label = 'val_accuracy')\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.ylim([0.5, 1])\nplt.title('Imagenet Train vs Validation Accuracy Per Epoch')\nplt.legend(loc='lower right')","metadata":{"execution":{"iopub.status.busy":"2023-12-13T07:19:04.170444Z","iopub.execute_input":"2023-12-13T07:19:04.171367Z","iopub.status.idle":"2023-12-13T07:19:04.486218Z","shell.execute_reply.started":"2023-12-13T07:19:04.171329Z","shell.execute_reply":"2023-12-13T07:19:04.485338Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(imagenet_history.history['loss'], label='loss')\nplt.plot(imagenet_history.history['val_loss'], label = 'val_loss')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\nplt.title('Imagenet Train vs Validation Loss Per Epoch')","metadata":{"execution":{"iopub.status.busy":"2023-12-13T07:19:08.330390Z","iopub.execute_input":"2023-12-13T07:19:08.331100Z","iopub.status.idle":"2023-12-13T07:19:08.607191Z","shell.execute_reply.started":"2023-12-13T07:19:08.331064Z","shell.execute_reply":"2023-12-13T07:19:08.606305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Interestingly, the accuracy for both validation and training data remained almost constant throughout the training. The loss for both steadily went down (apart from at the fifth epoch for validation where something wild happened).\n\nLet's try predicting test results with this model.","metadata":{}},{"cell_type":"code","source":"model2_pred_df = pd.DataFrame(columns=['id', 'label'])\nimagenet_predictions = imagenet_model.predict(test_dataset)\n","metadata":{"execution":{"iopub.status.busy":"2023-12-13T07:19:16.056895Z","iopub.execute_input":"2023-12-13T07:19:16.057257Z","iopub.status.idle":"2023-12-13T07:24:54.895692Z","shell.execute_reply.started":"2023-12-13T07:19:16.057227Z","shell.execute_reply":"2023-12-13T07:24:54.894579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model2_pred_df['id'] = [filename.split('.')[0] for filename in test_imgs]\nmodel2_pred_df['label'] = np.round(imagenet_predictions.flatten()).astype('int')\nmodel2_pred_df","metadata":{"execution":{"iopub.status.busy":"2023-12-13T07:27:37.248492Z","iopub.execute_input":"2023-12-13T07:27:37.249041Z","iopub.status.idle":"2023-12-13T07:27:37.294762Z","shell.execute_reply.started":"2023-12-13T07:27:37.249003Z","shell.execute_reply":"2023-12-13T07:27:37.293630Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model2_pred_df.to_csv('imagenet_predictions.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-12-13T07:27:41.504172Z","iopub.execute_input":"2023-12-13T07:27:41.504588Z","iopub.status.idle":"2023-12-13T07:27:41.707406Z","shell.execute_reply.started":"2023-12-13T07:27:41.504557Z","shell.execute_reply":"2023-12-13T07:27:41.706464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Upon submitting, the results from ImageNet model to the competition, a public score of **0.7855** was achieved.","metadata":{}},{"cell_type":"markdown","source":"# Results and Analysis\n\nThe project explored two distinct CNN models for histopathologic cancer detection: a custom-built model and an ImageNet-based model. The custom model, comprising three convolutional layers with max pooling and three dense layers, achieved a validation accuracy of 86% and demonstrated potential overfitting beyond the fourth epoch. This was evident as both validation loss and accuracy began to decrease. The model scored 0.8239 in the Kaggle competition. \n\nIn contrast, the ImageNet-based model maintained stable performance throughout training, with an unexpected fluctuation in the fifth epoch. Despite its stability, it scored slightly lower in the competition at 0.7855.\n\n| Model Type | Validation Accuracy | Stability | Competition Score | Remarks |\n| -----------|---------------------|-----------|-------------------|---------|\n| Custom Model | 86% | Low | 0.8239 | Overfitting after 4th epoch |\n| ImageNet-Based | 80% | High | 0.7855 | Consistent performance |\n\n\n## Troubleshooting and Future Considerations\n\nFor the custom model, implementing techniques such as dropout, early stopping, or adding regularization could help mitigate overfitting. For the ImageNet-based model, exploring different architectures or fine-tuning additional layers might improve its ability to capture more complex patterns in the data.\n\nIn summary, both models provide foundational insights for histopathologic cancer detection, yet each requires specific adjustments to fully optimize their performance. Future work should focus on balancing model complexity with the ability to generalize, along with exploring systematic hyperparameter optimization strategies.","metadata":{}},{"cell_type":"markdown","source":"# Conclusion\n\nIn this project, we explored CNNs for histopathologic cancer detection using two models: a custom-built model and an ImageNet-based model. The custom model showed high accuracy (86%) but tended to overfit after the fourth epoch. It scored 0.8239 in the Kaggle competition, indicating its effectiveness despite overfitting issues. The ImageNet-based model, while stable, achieved a lower score of 0.7855, suggesting that it might not fully capture the dataset's complexity.\n\nNo hyperparameter optimization was performed, which could be a direction for future improvement. Both models have their strengths - the custom model in accuracy and the ImageNet model in stability. This project highlights the potential of CNNs in medical imaging but also the need for balanced model design and hyperparameter tuning to enhance performance and generalization. Future efforts could focus on combining the best aspects of both models and systematic hyperparameter optimization for better cancer detection capabilities.","metadata":{}},{"cell_type":"markdown","source":"# References\n\n- Cukierski, W.: *Histopathologic Cancer Detection*, Kaggle, https://kaggle.com/competitions/histopathologic-cancer-detection, 2018.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}