{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":11848,"databundleVersionId":862157,"sourceType":"competition"}],"dockerImageVersionId":30588,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# ***Histopathic Cancer Detection***\n## DSCI 598\n Jeffery Boczkaja |\n Sara Bronson |\n Shantel Johnson","metadata":{}},{"cell_type":"markdown","source":"## Overview: Developing an Algorithm for Cancer Detection in Digital Pathology\n\nThe primary goal of this project is to develop a sophisticated algorithm capable of identifying metastatic cancer from small image patches which are derived from larger digital pathology scans. This complex task is approached as a binary classification problem where the aim is to accurately categorize each image patch into one of two distinct classes: those exhibiting signs of cancer and those that do not.\n\nTo achieve our objective we will utilize Convolutional Neural Networks (CNNs), a class of deep neural networks highly effective in analyzing visual imagery. CNNs are particularly well-suited for this task due to their ability to detect patterns and features in images, making them ideal for identifying intricate and subtle variations in tissue samples that are indicative of cancerous cells.\n\nThrough this approach the project aims to harness the power of advanced machine learning techniques to aid in the early detection of cancer, potentially leading to better patient outcomes. The successful development of this algorithm could represent a significant step forward in the application of artificial intelligence in medical imaging and diagnostics.","metadata":{}},{"cell_type":"markdown","source":"# Section 1: Importing Packages\n\nWe initialize the Python environment by importing necesarry libraries and packages that will be used throughout the notebook. We will use KerasTuner to help find hyper parameters. We also set the random seeds for reproducibility. ","metadata":{}},{"cell_type":"code","source":"#import packages\nimport os\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport matplotlib.image as mpimg\nfrom sklearn.model_selection import train_test_split\nimport pickle\nimport tensorflow as tf\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import *\nfrom kerastuner.tuners import RandomSearch\nfrom kerastuner.engine.hyperparameters import HyperParameters\n\n# Set the seed for reproducibility\nnp.random.seed(66)\ntf.random.set_seed(66)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-12-17T05:10:46.918648Z","iopub.execute_input":"2023-12-17T05:10:46.918901Z","iopub.status.idle":"2023-12-17T05:10:59.810126Z","shell.execute_reply.started":"2023-12-17T05:10:46.918877Z","shell.execute_reply":"2023-12-17T05:10:59.809158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Section 2: Data Loading and Initial Setup\n\nIn this section we load the image files into the notebook. This involves importing the labeled data which will be used for training the model.","metadata":{}},{"cell_type":"code","source":"#train and test folder\nprint('Number of images in train set',len(os.listdir('../input/histopathologic-cancer-detection/train')))\nprint('Number of images in test set',len(os.listdir('../input/histopathologic-cancer-detection/test')))","metadata":{"execution":{"iopub.status.busy":"2023-12-17T05:10:59.812055Z","iopub.execute_input":"2023-12-17T05:10:59.812569Z","iopub.status.idle":"2023-12-17T05:11:03.246835Z","shell.execute_reply.started":"2023-12-17T05:10:59.812541Z","shell.execute_reply":"2023-12-17T05:11:03.245809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the training data into a DataFrame. \n# Print the shape of the resulting DataFrame.\n\nhcd = pd.read_csv('/kaggle/input/histopathologic-cancer-detection/train_labels.csv')\nprint(hcd.shape)","metadata":{"execution":{"iopub.status.busy":"2023-12-17T05:11:03.248065Z","iopub.execute_input":"2023-12-17T05:11:03.248367Z","iopub.status.idle":"2023-12-17T05:11:03.560451Z","shell.execute_reply.started":"2023-12-17T05:11:03.248341Z","shell.execute_reply":"2023-12-17T05:11:03.559473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the first few rows of the dataframe.\nhcd.head() ","metadata":{"execution":{"iopub.status.busy":"2023-12-17T05:11:03.563387Z","iopub.execute_input":"2023-12-17T05:11:03.563804Z","iopub.status.idle":"2023-12-17T05:11:03.580369Z","shell.execute_reply.started":"2023-12-17T05:11:03.563768Z","shell.execute_reply":"2023-12-17T05:11:03.579300Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#label distrobution\n(hcd.label.value_counts() / len(hcd)).to_frame()","metadata":{"execution":{"iopub.status.busy":"2023-12-17T05:11:03.581363Z","iopub.execute_input":"2023-12-17T05:11:03.581659Z","iopub.status.idle":"2023-12-17T05:11:03.602463Z","shell.execute_reply.started":"2023-12-17T05:11:03.581635Z","shell.execute_reply":"2023-12-17T05:11:03.601425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Adding a variable for the image directory\nimg_dir = '/kaggle/input/histopathologic-cancer-detection/train'\n","metadata":{"execution":{"iopub.status.busy":"2023-12-17T05:11:03.603861Z","iopub.execute_input":"2023-12-17T05:11:03.604170Z","iopub.status.idle":"2023-12-17T05:11:03.608397Z","shell.execute_reply.started":"2023-12-17T05:11:03.604145Z","shell.execute_reply":"2023-12-17T05:11:03.607459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Section 3: Data Visualization and Exploration\n\nIn this section, we delve into a visual exploration of our dataset. We examine the structure and composition of the data by inspecting the DataFrame, focusing particularly on the distribution of labels to help understand the balance between the two different classes. Next we enhance our comprehension of the dataset by visually inspecting a selection of random images. This step provides insights into the nature of the images we are dealing with. It also aids in better understanding how cancerous and non-cancerous samples may differ visually.","metadata":{}},{"cell_type":"code","source":"sample = hcd.sample(n=9).reset_index()\n\nplt.figure(figsize=(3,3))\n\nfor i, row in sample.iterrows():\n\n    img = mpimg.imread(f'{img_dir}/{row.id}.tif')    \n    label = row.label\n\n    plt.subplot(3,3,i+1)\n    plt.imshow(img)\n    plt.text(0, -5, f'Class {label}', color='k')\n        \n    plt.axis('off')\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-17T05:11:03.609929Z","iopub.execute_input":"2023-12-17T05:11:03.610284Z","iopub.status.idle":"2023-12-17T05:11:04.213066Z","shell.execute_reply.started":"2023-12-17T05:11:03.610250Z","shell.execute_reply":"2023-12-17T05:11:04.211541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#using data generators \ntrain_df, valid_df = train_test_split(hcd, test_size=0.2, random_state=39, stratify=hcd.label)\n\nprint(train_df.shape)\nprint(valid_df.shape)","metadata":{"execution":{"iopub.status.busy":"2023-12-17T05:11:04.214698Z","iopub.execute_input":"2023-12-17T05:11:04.215222Z","iopub.status.idle":"2023-12-17T05:11:04.333582Z","shell.execute_reply.started":"2023-12-17T05:11:04.215174Z","shell.execute_reply":"2023-12-17T05:11:04.332709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Section 4: Data Preprocessing and Preparation for Modeling\n\nIn this section we prepare the data and images for input into our model. We begin by employing the train_test_split method to divide our dataset into training and validation subsets, ensuring a thorough evaluation of the model's performance. Additionally, we append the appropriate file extension to the id column, linking each data entry to its corresponding image file. Lastly, we utilize Data Generators plus scalers for efficient and effective image data augmentation and normalization. These steps are critical for optimizing the performance of our convolutional neural network. They facilitate the handling of image data and enhance the model's ability to generalize from the training data.","metadata":{}},{"cell_type":"code","source":"train_df['id'] = train_df['id'] + '.tif'\nvalid_df['id'] = valid_df['id'] + '.tif'\n","metadata":{"execution":{"iopub.status.busy":"2023-12-17T05:11:04.334943Z","iopub.execute_input":"2023-12-17T05:11:04.335223Z","iopub.status.idle":"2023-12-17T05:11:04.393804Z","shell.execute_reply.started":"2023-12-17T05:11:04.335199Z","shell.execute_reply":"2023-12-17T05:11:04.392802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Creating Data Generators for CNN\ntrain_datagen = ImageDataGenerator(\n    rescale=1/255,\n    rotation_range=40,\n    width_shift_range=0.2,\n    height_shift_range=0.2,\n    shear_range=0.2,\n    zoom_range=0.2,\n    horizontal_flip=True,\n    fill_mode='nearest'\n)\n\nvalid_datagen = ImageDataGenerator(rescale=1/255)\n\ntrain_df['label'] = train_df['label'].astype(str)\nvalid_df['label'] = valid_df['label'].astype(str)\n\ntrain_generator = train_datagen.flow_from_dataframe(\n    dataframe=train_df,\n    directory=img_dir,\n    x_col='id',\n    y_col='label',\n    target_size=(96, 96),\n    batch_size=32,\n    class_mode='binary'\n)\n\nvalidation_generator = valid_datagen.flow_from_dataframe(\n    dataframe=valid_df,\n    directory=img_dir,\n    x_col='id',\n    y_col='label',\n    target_size=(96, 96),\n    batch_size=32,\n    class_mode='binary'\n)","metadata":{"execution":{"iopub.status.busy":"2023-12-17T05:11:04.395100Z","iopub.execute_input":"2023-12-17T05:11:04.395456Z","iopub.status.idle":"2023-12-17T05:17:59.062177Z","shell.execute_reply.started":"2023-12-17T05:11:04.395425Z","shell.execute_reply":"2023-12-17T05:17:59.061406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"TR_STEPS = len(train_generator)\nVA_STEPS = len(validation_generator)\n\nprint('Number of batches in the training set:',TR_STEPS)\nprint('Number of batches in the validation set:',VA_STEPS)","metadata":{"execution":{"iopub.status.busy":"2023-12-17T05:17:59.064932Z","iopub.execute_input":"2023-12-17T05:17:59.065220Z","iopub.status.idle":"2023-12-17T05:17:59.070063Z","shell.execute_reply.started":"2023-12-17T05:17:59.065197Z","shell.execute_reply":"2023-12-17T05:17:59.069232Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Section 5: Model Architecture Definition and Hyperparameter Optimization\n\nIn this section we focus on constructing the neural network model. The model is built using Sequential, incorporating several convolutional layers. For optimization we utilize the Adam optimizer.\n\nWe undertake an exploration of different learning rates, testing the model with various optimizer values to determine the most effective rate for training. The process of identifying the best hyperparameters is carried out using a hyperparameter tuning approach. We employ this technique to search through various configurations, ultimately identifying the optimal set of hyperparameters. These optimal parameters are then applied to fine-tune our model using a tuner, ensuring that the model is as effective as possible in learning from our data and making accurate predictions.","metadata":{}},{"cell_type":"code","source":"def build_model(hp):\n    model = Sequential()\n\n    # Convolutional layers\n    model.add(Conv2D(hp.Int('conv1_units', min_value=32, max_value=128, step=32), (3, 3), activation='relu', input_shape=(96, 96, 3)))\n    model.add(MaxPooling2D((2, 2)))\n\n    model.add(Conv2D(hp.Int('conv2_units', min_value=32, max_value=128, step=32), (3, 3), activation='relu'))\n    model.add(MaxPooling2D((2, 2)))\n\n    model.add(Conv2D(hp.Int('conv3_units', min_value=32, max_value=128, step=32), (3, 3), activation='relu'))\n    model.add(MaxPooling2D((2, 2)))\n\n    model.add(Flatten())\n\n    # Dense layers\n    model.add(Dense(hp.Int('dense_units', min_value=64, max_value=256, step=64), activation='relu'))\n\n    # Dropout layer\n    model.add(Dropout(hp.Float('dropout_rate', min_value=0.2, max_value=0.5, step=0.1)))\n\n    # Output layer\n    model.add(Dense(1, activation='sigmoid'))\n\n    # Compile the model\n    opt = tf.keras.optimizers.Adam(hp.Choice('learning_rate', values=[1e-2, 1e-3, 1e-4]))\n    model.compile(loss='binary_crossentropy', optimizer=opt, metrics=['accuracy', tf.keras.metrics.AUC()])\n    return model","metadata":{"execution":{"iopub.status.busy":"2023-12-17T05:17:59.071153Z","iopub.execute_input":"2023-12-17T05:17:59.071409Z","iopub.status.idle":"2023-12-17T05:17:59.091670Z","shell.execute_reply.started":"2023-12-17T05:17:59.071387Z","shell.execute_reply":"2023-12-17T05:17:59.090822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tuner = RandomSearch(\n    build_model,\n    objective='val_accuracy',\n    max_trials=3,  \n    directory='HCD_tuner_dir',  \n    project_name='HCD_SB_6'\n)\n\ntuner.search(train_generator, epochs=5, validation_data=validation_generator, validation_steps=VA_STEPS)\n\n# Get the best hyperparameters\nbest_hps = tuner.get_best_hyperparameters(num_trials=1)[0]","metadata":{"execution":{"iopub.status.busy":"2023-12-17T05:17:59.092766Z","iopub.execute_input":"2023-12-17T05:17:59.093098Z","iopub.status.idle":"2023-12-17T08:50:57.213794Z","shell.execute_reply.started":"2023-12-17T05:17:59.093067Z","shell.execute_reply":"2023-12-17T08:50:57.212796Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_model = tuner.hypermodel.build(best_hps)","metadata":{"execution":{"iopub.status.busy":"2023-12-17T08:50:57.215204Z","iopub.execute_input":"2023-12-17T08:50:57.215788Z","iopub.status.idle":"2023-12-17T08:50:57.321735Z","shell.execute_reply.started":"2023-12-17T08:50:57.215748Z","shell.execute_reply":"2023-12-17T08:50:57.320812Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Section 6: Model Training and Performance Visualization\n\nIn this section, we integrate the previously identified optimal hyperparameters into our model and begin training. The model is trained over a total of 5 epochs. After this is complete we focus on evaluating the model's performance. We achieve this by plotting the training and validation accuracy, as well as the loss metrics, on line graphs. These visual representations provide an intuitive understanding of how the model has improved and stabilized over each epoch, and they are crucial for assessing the effectiveness of our training process and the overall reliability of the model in classifying the images.","metadata":{}},{"cell_type":"code","source":"h1 = final_model.fit(\n    train_generator,\n    steps_per_epoch=TR_STEPS,\n    epochs=10,\n    validation_data=validation_generator,\n    validation_steps=VA_STEPS,\n    verbose=1\n)\n","metadata":{"execution":{"iopub.status.busy":"2023-12-17T08:50:57.322941Z","iopub.execute_input":"2023-12-17T08:50:57.323299Z","iopub.status.idle":"2023-12-17T11:01:42.130129Z","shell.execute_reply.started":"2023-12-17T08:50:57.323266Z","shell.execute_reply":"2023-12-17T11:01:42.129340Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = h1.history\nepoch_range =range(1, len(history['loss'])+1)\n\nplt.figure(figsize=[14,4])\nplt.subplot(1,2,1)\nplt.plot(epoch_range, history['loss'], label='Training')\nplt.plot(epoch_range, history['val_loss'], label='Validation')\nplt.xlabel('Epoch'); plt.ylabel('Loss'); plt.title('Loss')\nplt.legend()\nplt.subplot(1,2,2)\nplt.plot(epoch_range, history['accuracy'], label='Training')\nplt.plot(epoch_range, history['val_accuracy'], label='Validation')\nplt.xlabel('Epoch'); plt.ylabel('Accuracy'); plt.title('Accuracy')\nplt.legend()\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-17T11:01:42.131523Z","iopub.execute_input":"2023-12-17T11:01:42.131814Z","iopub.status.idle":"2023-12-17T11:01:42.724637Z","shell.execute_reply.started":"2023-12-17T11:01:42.131789Z","shell.execute_reply":"2023-12-17T11:01:42.723649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#  Section 7: Saving the Trained Model\n\nHere we save the fully trained model for use when creating our test predictions. ","metadata":{"execution":{"iopub.status.busy":"2023-12-10T05:25:37.056294Z","iopub.execute_input":"2023-12-10T05:25:37.057253Z","iopub.status.idle":"2023-12-10T05:25:37.070816Z","shell.execute_reply.started":"2023-12-10T05:25:37.057211Z","shell.execute_reply":"2023-12-10T05:25:37.069493Z"}}},{"cell_type":"code","source":"# Save the model\nfinal_model.save('HCD_SB_6_Model.h5')\n\n# Save the history\nwith open('HCD_SB_6_Model_history.pkl', 'wb') as file:\n    pickle.dump(h1.history, file)","metadata":{"execution":{"iopub.status.busy":"2023-12-17T11:01:42.725810Z","iopub.execute_input":"2023-12-17T11:01:42.726103Z","iopub.status.idle":"2023-12-17T11:01:42.781444Z","shell.execute_reply.started":"2023-12-17T11:01:42.726078Z","shell.execute_reply":"2023-12-17T11:01:42.780465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Section 8: Documentation of Alternative Approaches\n\nThis markdown cell has the list of notebooks we tried and didn't produce the best results. These represent the trial and error process that has led to the creation of this model.\n\nhttps://www.kaggle.com/code/sarabronson/seab-take-1-hcd  \nhttps://www.kaggle.com/code/sarabronson/seab-take-1-hcdx  \nhttps://www.kaggle.com/code/sarabronson/seab-take-1-hcd-bf3be1  \nhttps://www.kaggle.com/code/sarabronson/hcd-tf  \nhttps://www.kaggle.com/code/sarabronson/hcd-pytorch-take1  \nhttps://www.kaggle.com/code/sarabronson/hcd-pytorch-take2  \nhttps://www.kaggle.com/code/sarabronson/sb-hcd-kt-sub  \nhttps://www.kaggle.com/code/shayjohnson/hcd-notebook  \nhttps://www.kaggle.com/code/shayjohnson/hcd-test-3  \nhttps://www.kaggle.com/code/shayjohnson/hcd-test-submission  \nhttps://www.kaggle.com/code/jefferyboczkaja/test-notebook  \nhttps://www.kaggle.com/code/jefferyboczkaja/test-submission  \nhttps://www.kaggle.com/code/jefferyboczkaja/hcd-pytorch-submission","metadata":{}},{"cell_type":"markdown","source":"# Section 9: Links to Submission and Summary Notebooks.\n\nThis is the link to the summary notebook. \n\nhttps://www.kaggle.com/code/sarabronson/hcd-summary-notebook\n\nThis is the link to the submission notebook.\n\nhttps://www.kaggle.com/code/sarabronson/sb-hcd-kt-sub","metadata":{"execution":{"iopub.status.busy":"2023-12-10T05:26:54.829929Z","iopub.execute_input":"2023-12-10T05:26:54.83026Z","iopub.status.idle":"2023-12-10T05:26:54.836441Z","shell.execute_reply.started":"2023-12-10T05:26:54.830233Z","shell.execute_reply":"2023-12-10T05:26:54.835205Z"}}}]}