{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## **CNN Cancer Detection**\n\n### **Overview**\n\nIn today's world we are fortunate enough to have advanced technology which allows the medical field to provide better patient care. Cancer is a highly researched area because there are so many people who suffer and/or die of cancerous disease. This study aims to improve cancer detection in lymph nodes by using computer vision machine learning technqiues. We will examine the data given in this competition to get a better understanding of it. Then, we will run multiple convolutional neural network models with the intent to be able to classify cancerous (1) and non-cancerous cells (0). With improved and faster cancer detection, patients will be able to recieve life-saving treatments faster. This the first step in those peoples cancer survival story. This notebook covers the thought process as to how to create simple CNN models. We will create two models, one without hyperparameter tuning, and one with tuning. Finally, we will suggest ways to improve the models in future studies.\n\nLayout for this notebook was given in assignment brief and is as follows: \n\n\n1. Brief Description of the Problem and Data\n2. Exploratory Data Analysis (EDA) - Inspect, Visualize and Clean the Data\n3. Describe Model Architecture \n4. Results and Analysis","metadata":{"execution":{"iopub.status.busy":"2022-12-24T17:33:11.729583Z","iopub.execute_input":"2022-12-24T17:33:11.730151Z","iopub.status.idle":"2022-12-24T17:33:11.756335Z","shell.execute_reply.started":"2022-12-24T17:33:11.730063Z","shell.execute_reply":"2022-12-24T17:33:11.754504Z"}}},{"cell_type":"code","source":"# Import libraries\n\n# General libraries \nimport numpy as np \nimport pandas as pd \nimport os\nimport random\nfrom sklearn.utils import shuffle\nimport shutil\n\n# Visualizations\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.express as px\nimport matplotlib.patches as patches\n\n# Work with images\nfrom skimage.transform import rotate\nfrom skimage import io\nimport cv2 as cv\n\n# Model Development\nfrom sklearn.model_selection import train_test_split\nimport tensorflow as tf\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras.layers import RandomFlip, RandomZoom, RandomRotation\nfrom tensorflow.keras.layers import Conv2D, MaxPooling2D, AveragePooling2D\nfrom tensorflow.keras.layers import Dense, Flatten, Dropout\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import BatchNormalization\nfrom tensorflow.keras.optimizers import Adam\n\nimport warnings\nwarnings.simplefilter(\"ignore\", category=DeprecationWarning)","metadata":{"execution":{"iopub.status.busy":"2022-12-25T20:47:02.063896Z","iopub.execute_input":"2022-12-25T20:47:02.064863Z","iopub.status.idle":"2022-12-25T20:47:11.299085Z","shell.execute_reply.started":"2022-12-25T20:47:02.064748Z","shell.execute_reply":"2022-12-25T20:47:11.298077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get files\ntest_path = '../input/histopathologic-cancer-detection/test/'\ntrain_path = '../input/histopathologic-cancer-detection/train/'\nsample_submission = pd.read_csv('../input/histopathologic-cancer-detection/sample_submission.csv')\ntrain_data = pd.read_csv('../input/histopathologic-cancer-detection/train_labels.csv')","metadata":{"execution":{"iopub.status.busy":"2022-12-25T20:47:11.300926Z","iopub.execute_input":"2022-12-25T20:47:11.302433Z","iopub.status.idle":"2022-12-25T20:47:11.699199Z","shell.execute_reply.started":"2022-12-25T20:47:11.302364Z","shell.execute_reply":"2022-12-25T20:47:11.698218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1. Brief Description of the Problem and Data\n\n* The dataset contains the histopathological Images, each image is 96px * 96px with 3 channels. \n* We have two datasets, a training and testing set already split for us.\n* The training set contains 220,025 unique images and the test set contains about 57,500.\n* To use these images in a machine learning model, we are also given an identifying dataframe with two columns: 'id' which is the unique image ID correpsonding to the training directory, and 'label' which tells us the classification category.\n* Each label is either a 0 or 1, depending whether the image is non-cancerous (0) or cancerous (1). \n* In the competition description, we find that if at least one pixel of an image is identified as cancerous then the whole image is therefore marked with a 1, otherwise it is 0. It is important to note that we do not have any missing values in this data which will make preprocessing more efficient.\n","metadata":{}},{"cell_type":"code","source":"# declare constants for reproduciblity\nRANDOM_STATE = 49","metadata":{"execution":{"iopub.status.busy":"2022-12-25T20:47:11.700547Z","iopub.execute_input":"2022-12-25T20:47:11.700915Z","iopub.status.idle":"2022-12-25T20:47:11.706041Z","shell.execute_reply.started":"2022-12-25T20:47:11.700878Z","shell.execute_reply":"2022-12-25T20:47:11.704965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# have a look at the format of the data\ntrain_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-25T20:47:11.709042Z","iopub.execute_input":"2022-12-25T20:47:11.709544Z","iopub.status.idle":"2022-12-25T20:47:11.728207Z","shell.execute_reply.started":"2022-12-25T20:47:11.709508Z","shell.execute_reply":"2022-12-25T20:47:11.727159Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# So, how many images are there in each of the folder in the training dataset?\n\nprint(len(os.listdir('../input/histopathologic-cancer-detection/train')))\nprint(len(os.listdir('../input/histopathologic-cancer-detection/test')))","metadata":{"execution":{"iopub.status.busy":"2022-12-25T20:47:11.729713Z","iopub.execute_input":"2022-12-25T20:47:11.730140Z","iopub.status.idle":"2022-12-25T20:47:15.352319Z","shell.execute_reply.started":"2022-12-25T20:47:11.730105Z","shell.execute_reply":"2022-12-25T20:47:15.351073Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# take a look at the data further\ntrain_data.describe()","metadata":{"execution":{"iopub.status.busy":"2022-12-25T20:47:15.353647Z","iopub.execute_input":"2022-12-25T20:47:15.354334Z","iopub.status.idle":"2022-12-25T20:47:15.386265Z","shell.execute_reply.started":"2022-12-25T20:47:15.354296Z","shell.execute_reply":"2022-12-25T20:47:15.385420Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check information, data types, and for missing data\ntrain_data.info()","metadata":{"execution":{"iopub.status.busy":"2022-12-25T20:47:15.388343Z","iopub.execute_input":"2022-12-25T20:47:15.388704Z","iopub.status.idle":"2022-12-25T20:47:15.414373Z","shell.execute_reply.started":"2022-12-25T20:47:15.388669Z","shell.execute_reply":"2022-12-25T20:47:15.413525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. (EDA) - Visualize and Clean the Data\n\nFirst, we will visualize the data. Then we will clean/preprocess the data. \n\n* We can see in the histogram and pie chart below that we have 59.5% of the labels are 0 (non-cancerous images) and 40.5% are labeled 1 (cancerous images). \n* We were told in the competition description that the data is 50/50 split between cancerous and non-cancerous images however, from what we are finding here, we have a split which is closer to 40/60. This means that our data is unbalanced, however it is not severely unbalanced either (compared to a split such as 30/70 or even 10/90). \n* We also have thousands of images to train with. For this reason, we can assume we will be able to create a sufficiently performing model which identifies cancerous images. ","metadata":{}},{"cell_type":"code","source":"# create histogram\n\nprint(pd.DataFrame(data={'Label Counts': train_data['label'].value_counts()}))\nsns.countplot(x=train_data['label'], palette='colorblind').set(title='Label Counts Histogram');","metadata":{"execution":{"iopub.status.busy":"2022-12-25T20:47:15.416518Z","iopub.execute_input":"2022-12-25T20:47:15.416855Z","iopub.status.idle":"2022-12-25T20:47:15.639479Z","shell.execute_reply.started":"2022-12-25T20:47:15.416820Z","shell.execute_reply":"2022-12-25T20:47:15.638777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#create pie chart\n\nfig = px.pie(train_data, \n             values = train_data['label'].value_counts().values, \n             names = train_data['label'].unique())\nfig.update_layout(\n    title={\n        'text': \"Label Percentage Pie Chart\",\n        'y':.99,\n        'x':0.5,\n        'xanchor': 'center',\n        'yanchor': 'top'})\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-25T20:47:15.640650Z","iopub.execute_input":"2022-12-25T20:47:15.641590Z","iopub.status.idle":"2022-12-25T20:47:16.587140Z","shell.execute_reply.started":"2022-12-25T20:47:15.641552Z","shell.execute_reply":"2022-12-25T20:47:16.586240Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Below we can see some example images from the training data.We were told that each image has the (potentially) cancerous cells centered in each 32x32 pixel image, so we have drawn a box around this area as a focal point. ","metadata":{}},{"cell_type":"code","source":"# Visualize a few images\n\nfig, ax = plt.subplots(5, 5, figsize=(15, 15))\nfor i, axis in enumerate(ax.flat):\n    file = str(train_path + train_data.id[i] + '.tif')\n    image = io.imread(file)\n    axis.imshow(image)\n    box = patches.Rectangle((32,32),32,32, linewidth=2, edgecolor='r',facecolor='none', linestyle='-')\n    axis.add_patch(box)\n    axis.set(xticks=[], yticks=[], xlabel = train_data.label[i]);\n    #cv2.waitKey(0)","metadata":{"execution":{"iopub.status.busy":"2022-12-25T20:47:16.591332Z","iopub.execute_input":"2022-12-25T20:47:16.591646Z","iopub.status.idle":"2022-12-25T20:47:17.996588Z","shell.execute_reply.started":"2022-12-25T20:47:16.591619Z","shell.execute_reply":"2022-12-25T20:47:17.995436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. Describe Model Architecture \n\n**Model:**\n1. Normalize images pre-training (image/255)\n2. Output layer activation (sigmoid)\n3. Optimization (Adam)\n4. Learning rate (0.0001) \n5. Hidden layer activations (ReLU)\n6. Dropout (0.3)","metadata":{}},{"cell_type":"code","source":"# set model constants\n\nBATCH_SIZE = 256","metadata":{"execution":{"iopub.status.busy":"2022-12-25T20:47:17.997545Z","iopub.execute_input":"2022-12-25T20:47:17.997862Z","iopub.status.idle":"2022-12-25T20:47:18.002539Z","shell.execute_reply.started":"2022-12-25T20:47:17.997832Z","shell.execute_reply":"2022-12-25T20:47:18.001677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# prepare data for training\ndef append_tif(string):\n    return string+\".tif\"\n\ntrain_data[\"id\"] = train_data[\"id\"].apply(append_tif)\ntrain_data['label'] = train_data['label'].astype(str)\n\n# randomly shuffle training data\ntrain_data = shuffle(train_data, random_state=RANDOM_STATE)","metadata":{"execution":{"iopub.status.busy":"2022-12-25T20:47:18.005316Z","iopub.execute_input":"2022-12-25T20:47:18.006348Z","iopub.status.idle":"2022-12-25T20:47:18.239639Z","shell.execute_reply.started":"2022-12-25T20:47:18.006297Z","shell.execute_reply":"2022-12-25T20:47:18.238657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# modify training data by normalizing it and split data into training and validation sets\n\ndatagen = ImageDataGenerator(rescale=1./255.,\n                            validation_split=0.15)","metadata":{"execution":{"iopub.status.busy":"2022-12-25T20:47:18.241354Z","iopub.execute_input":"2022-12-25T20:47:18.241746Z","iopub.status.idle":"2022-12-25T20:47:18.246987Z","shell.execute_reply.started":"2022-12-25T20:47:18.241709Z","shell.execute_reply":"2022-12-25T20:47:18.246069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# generate training data\ntrain_generator = datagen.flow_from_dataframe(\n    dataframe=train_data,\n    directory=train_path,\n    x_col=\"id\",\n    y_col=\"label\",\n    subset=\"training\",\n    batch_size=BATCH_SIZE,\n    seed=RANDOM_STATE,\n    class_mode=\"binary\",\n    target_size=(64,64))        # original image = (96, 96) ","metadata":{"execution":{"iopub.status.busy":"2022-12-25T20:47:18.248227Z","iopub.execute_input":"2022-12-25T20:47:18.249025Z","iopub.status.idle":"2022-12-25T20:51:24.653438Z","shell.execute_reply.started":"2022-12-25T20:47:18.248988Z","shell.execute_reply":"2022-12-25T20:51:24.652428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# generate validation data\nvalid_generator = datagen.flow_from_dataframe(\n    dataframe=train_data,\n    directory=train_path,\n    x_col=\"id\",\n    y_col=\"label\",\n    subset=\"validation\",\n    batch_size=BATCH_SIZE,\n    seed=RANDOM_STATE,\n    class_mode=\"binary\",\n    target_size=(64,64))       # original image = (96, 96) ","metadata":{"execution":{"iopub.status.busy":"2022-12-25T20:51:24.655259Z","iopub.execute_input":"2022-12-25T20:51:24.655644Z","iopub.status.idle":"2022-12-25T20:53:06.061478Z","shell.execute_reply.started":"2022-12-25T20:51:24.655608Z","shell.execute_reply":"2022-12-25T20:53:06.060441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Setup GPU accelerator - configure Strategy. Assume TPU...if not set default for GPU/CPU\ntpu = None\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.TPUStrategy(tpu)\nexcept ValueError:\n    strategy = tf.distribute.get_strategy()","metadata":{"execution":{"iopub.status.busy":"2022-12-25T20:53:06.062843Z","iopub.execute_input":"2022-12-25T20:53:06.063682Z","iopub.status.idle":"2022-12-25T20:53:06.077447Z","shell.execute_reply.started":"2022-12-25T20:53:06.063644Z","shell.execute_reply":"2022-12-25T20:53:06.076412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The model that I have choosen for this problem has been taken from <a href = 'https://www.kaggle.com/fmarazzi/baseline-keras-cnn-roc-fast-10min-0-925-lb'>Baseline Keras CNN</a>","metadata":{}},{"cell_type":"code","source":"# set ROC AUC as metric\n\nROC_1 = tf.keras.metrics.AUC()\n\n# use GPU\nwith strategy.scope():kernel_size = (3,3)\npool_size= (2,2)\nfirst_filters = 32\nsecond_filters = 64\nthird_filters = 128\n\ndropout_conv = 0.3\ndropout_dense = 0.3\n\n\nmodel = Sequential()\nmodel.add(Conv2D(first_filters, kernel_size, activation = 'relu', input_shape = (64, 64, 3))) # original image = (96, 96, 3) \nmodel.add(Conv2D(first_filters, kernel_size, activation = 'relu'))\nmodel.add(Conv2D(first_filters, kernel_size, activation = 'relu'))\nmodel.add(MaxPooling2D(pool_size = pool_size)) \nmodel.add(Dropout(dropout_conv))\n\nmodel.add(Conv2D(second_filters, kernel_size, activation ='relu'))\nmodel.add(Conv2D(second_filters, kernel_size, activation ='relu'))\nmodel.add(Conv2D(second_filters, kernel_size, activation ='relu'))\nmodel.add(MaxPooling2D(pool_size = pool_size))\nmodel.add(Dropout(dropout_conv))\n\nmodel.add(Conv2D(third_filters, kernel_size, activation ='relu'))\nmodel.add(Conv2D(third_filters, kernel_size, activation ='relu'))\nmodel.add(Conv2D(third_filters, kernel_size, activation ='relu'))\nmodel.add(MaxPooling2D(pool_size = pool_size))\nmodel.add(Dropout(dropout_conv))\n\nmodel.add(Flatten())\nmodel.add(Dense(256, activation = \"relu\"))\nmodel.add(Dropout(dropout_dense))\nmodel.add(Dense(1, activation='sigmoid'))\n\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2022-12-25T21:02:53.728787Z","iopub.execute_input":"2022-12-25T21:02:53.729155Z","iopub.status.idle":"2022-12-25T21:02:53.889527Z","shell.execute_reply.started":"2022-12-25T21:02:53.729122Z","shell.execute_reply":"2022-12-25T21:02:53.888254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#compile\nadam_optimizer = Adam(learning_rate=0.0001)\nmodel.compile(loss='binary_crossentropy', metrics=['accuracy', ROC_1], optimizer=adam_optimizer)","metadata":{"execution":{"iopub.status.busy":"2022-12-25T21:02:56.554634Z","iopub.execute_input":"2022-12-25T21:02:56.555322Z","iopub.status.idle":"2022-12-25T21:02:56.565808Z","shell.execute_reply.started":"2022-12-25T21:02:56.555287Z","shell.execute_reply":"2022-12-25T21:02:56.564364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"EPOCHS = 20\n\n# train the model\nhistory_model = model.fit(\n                        train_generator,\n                        epochs = EPOCHS,\n                        validation_data = valid_generator)","metadata":{"execution":{"iopub.status.busy":"2022-12-25T21:02:58.236370Z","iopub.execute_input":"2022-12-25T21:02:58.237111Z","iopub.status.idle":"2022-12-25T22:51:34.518292Z","shell.execute_reply.started":"2022-12-25T21:02:58.237072Z","shell.execute_reply":"2022-12-25T22:51:34.517315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get the metric names so we can use evaulate_generator\nmodel.metrics_names","metadata":{"execution":{"iopub.status.busy":"2022-12-25T22:51:34.520352Z","iopub.execute_input":"2022-12-25T22:51:34.521121Z","iopub.status.idle":"2022-12-25T22:51:34.529308Z","shell.execute_reply.started":"2022-12-25T22:51:34.521082Z","shell.execute_reply":"2022-12-25T22:51:34.528375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot model accuracy per epoch \nplt.plot(history_model.history['accuracy'])\nplt.plot(history_model.history['val_accuracy'])\nplt.title('Model One Accuracy per Epoch')\nplt.ylabel('accuracy')\nplt.xlabel('epoch')\nplt.legend(['train', 'validate'], loc='upper left')\nplt.show();","metadata":{"execution":{"iopub.status.busy":"2022-12-25T22:51:34.530899Z","iopub.execute_input":"2022-12-25T22:51:34.531280Z","iopub.status.idle":"2022-12-25T22:51:34.750688Z","shell.execute_reply.started":"2022-12-25T22:51:34.531245Z","shell.execute_reply":"2022-12-25T22:51:34.749843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot model loss per epoch\nplt.plot(history_model.history['loss'])\nplt.plot(history_model.history['val_loss'])\nplt.title('Model One Loss per Epoch')\nplt.ylabel('loss')\nplt.xlabel('epoch')\nplt.legend(['train', 'validate'], loc='upper left')\nplt.show();","metadata":{"execution":{"iopub.status.busy":"2022-12-25T22:51:34.753026Z","iopub.execute_input":"2022-12-25T22:51:34.753293Z","iopub.status.idle":"2022-12-25T22:51:34.959200Z","shell.execute_reply.started":"2022-12-25T22:51:34.753267Z","shell.execute_reply":"2022-12-25T22:51:34.958291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot model ROC per epoch\nplt.plot(history_model.history['auc_1'])\nplt.plot(history_model.history['val_auc_1'])\nplt.title('Model One AUC ROC per Epoch')\nplt.ylabel('ROC')\nplt.xlabel('epoch')\nplt.legend(['train', 'validate'], loc='upper left')\nplt.show();","metadata":{"execution":{"iopub.status.busy":"2022-12-25T22:51:54.394969Z","iopub.execute_input":"2022-12-25T22:51:54.395333Z","iopub.status.idle":"2022-12-25T22:51:54.608369Z","shell.execute_reply.started":"2022-12-25T22:51:54.395300Z","shell.execute_reply":"2022-12-25T22:51:54.607445Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Test final model against the test set**\n\nNow that we have a trained model, we can test it on the unseen test data images. We must also normalize the test data like we did with the training data. ","metadata":{}},{"cell_type":"code","source":"#double check what you're aiming the submission data set to look like\nsample_submission.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-25T22:51:55.084029Z","iopub.execute_input":"2022-12-25T22:51:55.084392Z","iopub.status.idle":"2022-12-25T22:51:55.093754Z","shell.execute_reply.started":"2022-12-25T22:51:55.084348Z","shell.execute_reply":"2022-12-25T22:51:55.092811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#create a dataframe to run the predictions\ntest_df = pd.DataFrame({'id':os.listdir(test_path)})\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-25T22:51:55.284785Z","iopub.execute_input":"2022-12-25T22:51:55.285621Z","iopub.status.idle":"2022-12-25T22:51:56.064059Z","shell.execute_reply.started":"2022-12-25T22:51:55.285580Z","shell.execute_reply":"2022-12-25T22:51:56.063173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# prepare test data (in same way as train data)\ndatagen_test = ImageDataGenerator(rescale=1./255.)\n\ntest_generator = datagen_test.flow_from_dataframe(\n    dataframe=test_df,\n    directory=test_path,\n    x_col='id', \n    y_col=None,\n    target_size=(64,64),         # original image = (96, 96) \n    batch_size=1,\n    shuffle=False,\n    class_mode=None)","metadata":{"execution":{"iopub.status.busy":"2022-12-25T23:09:04.460977Z","iopub.execute_input":"2022-12-25T23:09:04.461331Z","iopub.status.idle":"2022-12-25T23:09:31.105216Z","shell.execute_reply.started":"2022-12-25T23:09:04.461299Z","shell.execute_reply":"2022-12-25T23:09:31.104134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#run model to find predictions\n\n# predictions = model.predict(test_generator, verbose=1)\npredictions = model.predict(test_generator, verbose=1)","metadata":{"execution":{"iopub.status.busy":"2022-12-25T23:09:31.108576Z","iopub.execute_input":"2022-12-25T23:09:31.109594Z","iopub.status.idle":"2022-12-25T23:12:22.700877Z","shell.execute_reply.started":"2022-12-25T23:09:31.109555Z","shell.execute_reply":"2022-12-25T23:12:22.699912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Are the number of predictions correct?\n# Should be 57458.\n\nlen(predictions)","metadata":{"execution":{"iopub.status.busy":"2022-12-25T23:12:22.702489Z","iopub.execute_input":"2022-12-25T23:12:22.702868Z","iopub.status.idle":"2022-12-25T23:12:22.711715Z","shell.execute_reply.started":"2022-12-25T23:12:22.702832Z","shell.execute_reply":"2022-12-25T23:12:22.710667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#create submission dataframe\npredictions = np.transpose(predictions)[0]\nsubmission_df = pd.DataFrame()\nsubmission_df['id'] = test_df['id'].apply(lambda x: x.split('.')[0])\nsubmission_df['label'] = list(map(lambda x: 0 if x < 0.5 else 1, predictions))\nsubmission_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-25T23:12:22.714223Z","iopub.execute_input":"2022-12-25T23:12:22.714676Z","iopub.status.idle":"2022-12-25T23:12:22.840759Z","shell.execute_reply.started":"2022-12-25T23:12:22.714641Z","shell.execute_reply":"2022-12-25T23:12:22.839714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#view test prediction counts\nsubmission_df['label'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-12-25T23:12:22.842267Z","iopub.execute_input":"2022-12-25T23:12:22.842708Z","iopub.status.idle":"2022-12-25T23:12:22.850949Z","shell.execute_reply.started":"2022-12-25T23:12:22.842672Z","shell.execute_reply":"2022-12-25T23:12:22.849883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#plot test predictions\nsns.countplot(data=submission_df, x='label').set(title='Predicted Labels for Test Set');","metadata":{"execution":{"iopub.status.busy":"2022-12-25T23:12:22.852364Z","iopub.execute_input":"2022-12-25T23:12:22.852784Z","iopub.status.idle":"2022-12-25T23:12:23.033160Z","shell.execute_reply.started":"2022-12-25T23:12:22.852741Z","shell.execute_reply":"2022-12-25T23:12:23.032195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#convert to csv to submit to competition\nsubmission_df.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-12-25T23:12:23.034595Z","iopub.execute_input":"2022-12-25T23:12:23.034946Z","iopub.status.idle":"2022-12-25T23:12:23.116124Z","shell.execute_reply.started":"2022-12-25T23:12:23.034909Z","shell.execute_reply":"2022-12-25T23:12:23.115260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Results and Analysis\n\nWe can see from the above plots and diagrams for each model how well they performed with the training (and validation) sets. We see that model is a complex model with regards to the ROC metric. We see in the Model that the accuracy and loss do not steady, nor does the ROC in the model. This could pertain to the fact that we trained with very few epochs (20) and a simple CNN model with so many pictures may need more \"time\" to train to converge. \n\nAfter submitting the trained model separately on the test set, we can see (below) how the model performed.","metadata":{}},{"cell_type":"markdown","source":"## **I hope that you found my notebook useful.**","metadata":{}}]}