{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Histopathologic Cancer Detection with CNN\n\n**Overview**\n\nThis project focuses on creating an algorithm to identify metastatic cancer in small image patches taken from larger digital pathology scans. We will examine the data given in this competition to get a better understanding of it. Then, we will run multiple convolutional neural network models with the intent to be able to classify cancerous(1) and non-cancerous cells(0). With improved and faster cancer detection, patients will be able to recieve life-saving treatments faster. We will create two models, one without hyperparameter tuning and one with tuning. Finally, we will suggest ways to improve the models in future studies. Let's begin!","metadata":{}},{"cell_type":"code","source":"#Library imports\nimport os\nimport pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline\nfrom sklearn.utils import shuffle\nimport shutil\n\n#Image processing\nfrom skimage.transform import rotate\nfrom skimage import io\nimport cv2 as cv\n\n#Model deployment\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras.layers import RandomFlip, RandomZoom, RandomRotation\nfrom tensorflow.keras.layers import Conv2D, MaxPooling2D, AveragePooling2D\nfrom tensorflow.keras.layers import Dense, Dropout, Flatten, Activation\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import BatchNormalization\nfrom tensorflow.keras.optimizers import Adam","metadata":{"execution":{"iopub.status.busy":"2022-11-23T15:15:32.679937Z","iopub.execute_input":"2022-11-23T15:15:32.681049Z","iopub.status.idle":"2022-11-23T15:15:32.694298Z","shell.execute_reply.started":"2022-11-23T15:15:32.681008Z","shell.execute_reply":"2022-11-23T15:15:32.693148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Directories\nos.listdir('../input/histopathologic-cancer-detection/')","metadata":{"execution":{"iopub.status.busy":"2022-11-23T11:02:55.992761Z","iopub.execute_input":"2022-11-23T11:02:55.993039Z","iopub.status.idle":"2022-11-23T11:02:56.002197Z","shell.execute_reply.started":"2022-11-23T11:02:55.993014Z","shell.execute_reply":"2022-11-23T11:02:56.001161Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Number of images in training and test set \nprint(len(os.listdir('../input/histopathologic-cancer-detection/train')))\nprint(len(os.listdir('../input/histopathologic-cancer-detection/test')))","metadata":{"execution":{"iopub.status.busy":"2022-11-23T11:02:56.027940Z","iopub.execute_input":"2022-11-23T11:02:56.028521Z","iopub.status.idle":"2022-11-23T11:02:56.218862Z","shell.execute_reply.started":"2022-11-23T11:02:56.028493Z","shell.execute_reply":"2022-11-23T11:02:56.217767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = pd.read_csv('../input/histopathologic-cancer-detection/train_labels.csv')\nsample_submission = pd.read_csv('../input/histopathologic-cancer-detection/sample_submission.csv')\ntest_path = '../input/histopathologic-cancer-detection/test/'\ntrain_path = '../input/histopathologic-cancer-detection/train/'\ntest_data = pd.DataFrame({'id':os.listdir(test_path)})\n\n# removing this training image because it caused a training error previously\ntrain_data[train_data['id'] != 'dd6dfed324f9fcb6f93f46f32fc800f2ec196be2']\n\n# removing this training image because it's not visible\ntrain_data[train_data['id'] != '9369c7278ec8bcc6c880d99194de09fc2bd4efbe']\n\nprint(train_data.shape)","metadata":{"execution":{"iopub.status.busy":"2022-11-23T17:39:11.242439Z","iopub.execute_input":"2022-11-23T17:39:11.243362Z","iopub.status.idle":"2022-11-23T17:39:12.911736Z","shell.execute_reply.started":"2022-11-23T17:39:11.243322Z","shell.execute_reply":"2022-11-23T17:39:12.910621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# declare constants for reproduciblity\nRANDOM_STATE = 101","metadata":{"execution":{"iopub.status.busy":"2022-11-23T11:02:56.587247Z","iopub.execute_input":"2022-11-23T11:02:56.587901Z","iopub.status.idle":"2022-11-23T11:02:56.592456Z","shell.execute_reply.started":"2022-11-23T11:02:56.587862Z","shell.execute_reply":"2022-11-23T11:02:56.591426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exploratory Data Analysis","metadata":{}},{"cell_type":"code","source":"train_data.sample(10)","metadata":{"execution":{"iopub.status.busy":"2022-11-23T11:02:56.596229Z","iopub.execute_input":"2022-11-23T11:02:56.596561Z","iopub.status.idle":"2022-11-23T11:02:56.622388Z","shell.execute_reply.started":"2022-11-23T11:02:56.596532Z","shell.execute_reply":"2022-11-23T11:02:56.621165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.sample(10)","metadata":{"execution":{"iopub.status.busy":"2022-11-23T11:02:56.625203Z","iopub.execute_input":"2022-11-23T11:02:56.625643Z","iopub.status.idle":"2022-11-23T11:02:56.637977Z","shell.execute_reply.started":"2022-11-23T11:02:56.625604Z","shell.execute_reply":"2022-11-23T11:02:56.636659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Visualizations\n\nFrom the training data visualizations, we see there is approximately a 60/40 split between **cancerous(1)** and **non-cancerous(0)** tissue respectively which is a fairly balanced dataset for predictions. Seeing that we are working with thousands of images, we should be able to produce a relatively accurate model.","metadata":{}},{"cell_type":"code","source":"print(train_data['label'].value_counts())\nsns.countplot(x=train_data['label'], palette='colorblind').set(title='Cancer Label Counts');","metadata":{"execution":{"iopub.status.busy":"2022-11-23T11:02:56.640190Z","iopub.execute_input":"2022-11-23T11:02:56.640702Z","iopub.status.idle":"2022-11-23T11:02:56.904228Z","shell.execute_reply.started":"2022-11-23T11:02:56.640663Z","shell.execute_reply":"2022-11-23T11:02:56.903424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set(style='whitegrid')\npie_chart=pd.DataFrame(train_data['label'].replace(0,'Non-cancerous tissue').replace(1,'Cancerous tissue').value_counts())\npie_chart.reset_index(inplace=True)\npie_chart.plot(kind='pie', title='Category Images',y = 'label', \n             autopct='%1.1f%%', shadow=False, labels=pie_chart['index'], legend = False, fontsize=14, figsize=(18,8))","metadata":{"execution":{"iopub.status.busy":"2022-11-23T11:02:56.907533Z","iopub.execute_input":"2022-11-23T11:02:56.907838Z","iopub.status.idle":"2022-11-23T11:02:57.134795Z","shell.execute_reply.started":"2022-11-23T11:02:56.907811Z","shell.execute_reply":"2022-11-23T11:02:57.133253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Below are some sample images from the training data. To the layman, eyeballing these images would not distinguish whether they are cancerous(1) or non-cancerous(0). The correct label for each image is shown below it.","metadata":{}},{"cell_type":"code","source":"#Image Visualisations\nfig, ax = plt.subplots(4, 4, figsize=(15, 15))\nfor i, axis in enumerate(ax.flat):\n    file = str(train_path + train_data.id[i] + '.tif')\n    image = io.imread(file)\n    axis.imshow(image)\n    axis.set(xticks=[], yticks=[], xlabel = train_data.label[i]);","metadata":{"execution":{"iopub.status.busy":"2022-11-23T11:02:57.137248Z","iopub.execute_input":"2022-11-23T11:02:57.138587Z","iopub.status.idle":"2022-11-23T11:02:58.584829Z","shell.execute_reply.started":"2022-11-23T11:02:57.138538Z","shell.execute_reply":"2022-11-23T11:02:58.583659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Data Preprocessing**\n\nThere were no missing data in the training set.However, there were two images that were causing issues in the dataset.Those were taken out.\n\nWhat we will do with the image data is **shuffle** the data so that the model doesn't learn based on the image ordering/pattern of input, which could potentially have consequences in the model training. We will also split the data into training and validation set to improve model development. During training we will also normalize the pixels by dividing by 255.0, which should help data processing and model training.","metadata":{}},{"cell_type":"markdown","source":"# Model Architecture\n\nWe will be using the Keras library to run a convolutional neural network (CNN). The first model will be run without hyperparameters and the second will be run with it.\n\nOur CNN model will have a network such that there are two convolutional layers then a MaxPool layer, and we repeat this n number of times. Specifically, we will create a fairly simple model with two (n=2) of these clusters. In other words, our model will be input --> Conv2D --> Conv2D --> MaxPool --> Conv2D --> Conv2D --> MaxPool --> Flatten --> Output with sigmoid activation.\n\nFirst model:\n\nNormalize images pre-training (image/255)\nOutput layer activation (sigmoid)\nSecond model contains all the first model parameters, but we also add:\n\nDropout (0.1)\nBatch Normalization\nOptimization (Adam)\nLearning rate (0.0001)\nHidden layer activations (ReLU)\nBefore we start with the models, let's describe what each of the parameters will do to the model. In both models we will use 1 and 2, numbers 3 to 5 will be used in the second and final models.\n\nNormalize images : this will take the pixels and divide each pixel by 255 to normalize the data and have values between 0-1.\nOutput layer activation : we will use a sigmoid activation function on the output layer since we are working with binary data\nDropout : we will set dropout at 0.1 which will randonly select some weights and set them to equal 0 which regularizes the model because it is using a smaller number of weights for each training run.\nOptimization : we will use adaptive moment estimation (Adam) for optimizing the model which essentially mimics momentum for gradient adn gradient-squared.\nLearning rate: we will set our learning rate to 0.0001 which will assist in the gradient descent such that as the model learns, the speed of learning decreases so that it is less likely to overstep the (hopefully) global minimum.\nHidden layer activations: we will use rectified linear regression (ReLU) as our hidden layer activation function which will help the model to converge better, prevent saturation, and provide less need for computation power.\nIn addition, we will use fairly large batch sizes, set at 256 to help reduce variance. Finally we wil train our two models with 10 epochs. We will use accuracy and the ROC-AUC curve to measure model performance, as well as binary cross-entropy as our loss function.\n\n","metadata":{}},{"cell_type":"code","source":"# set model constants\nBATCH_SIZE = 256","metadata":{"execution":{"iopub.status.busy":"2022-11-23T11:02:58.585866Z","iopub.execute_input":"2022-11-23T11:02:58.586749Z","iopub.status.idle":"2022-11-23T11:02:58.590817Z","shell.execute_reply.started":"2022-11-23T11:02:58.586711Z","shell.execute_reply":"2022-11-23T11:02:58.590033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# prepare data for training\ndef append_tif(string):\n    return string+\".tif\"\n\ntrain_data[\"id\"] = train_data[\"id\"].apply(append_tif)\ntrain_data['label'] = train_data['label'].astype(str)\n\n# randomly shuffle training data\ntrain_data = shuffle(train_data, random_state=RANDOM_STATE)","metadata":{"execution":{"iopub.status.busy":"2022-11-23T11:02:58.594683Z","iopub.execute_input":"2022-11-23T11:02:58.595778Z","iopub.status.idle":"2022-11-23T11:02:58.839979Z","shell.execute_reply.started":"2022-11-23T11:02:58.595743Z","shell.execute_reply":"2022-11-23T11:02:58.838933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# modify training data by normalizing it \n# and split data into training and validation sets\ndatagen = ImageDataGenerator(rescale=1./255.,\n                            validation_split=0.15)","metadata":{"execution":{"iopub.status.busy":"2022-11-23T11:02:58.841475Z","iopub.execute_input":"2022-11-23T11:02:58.841848Z","iopub.status.idle":"2022-11-23T11:02:58.847429Z","shell.execute_reply.started":"2022-11-23T11:02:58.841803Z","shell.execute_reply":"2022-11-23T11:02:58.846165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# generate training data\ntrain_generator = datagen.flow_from_dataframe(\n    dataframe=train_data,\n    directory=train_path,\n    x_col=\"id\",\n    y_col=\"label\",\n    subset=\"training\",\n    batch_size=BATCH_SIZE,\n    seed=RANDOM_STATE,\n    class_mode=\"binary\",\n    target_size=(64,64))  ","metadata":{"execution":{"iopub.status.busy":"2022-11-23T11:02:58.849284Z","iopub.execute_input":"2022-11-23T11:02:58.849662Z","iopub.status.idle":"2022-11-23T11:11:00.927986Z","shell.execute_reply.started":"2022-11-23T11:02:58.849618Z","shell.execute_reply":"2022-11-23T11:11:00.926912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# generate validation data\nvalid_generator = datagen.flow_from_dataframe(\n    dataframe=train_data,\n    directory=train_path,\n    x_col=\"id\",\n    y_col=\"label\",\n    subset=\"validation\",\n    batch_size=BATCH_SIZE,\n    seed=RANDOM_STATE,\n    class_mode=\"binary\",\n    target_size=(64,64))  ","metadata":{"execution":{"iopub.status.busy":"2022-11-23T11:11:00.929539Z","iopub.execute_input":"2022-11-23T11:11:00.929966Z","iopub.status.idle":"2022-11-23T11:14:54.749235Z","shell.execute_reply.started":"2022-11-23T11:11:00.929904Z","shell.execute_reply":"2022-11-23T11:14:54.748128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tpu = None\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.TPUStrategy(tpu)\nexcept ValueError:\n    strategy = tf.distribute.get_strategy()","metadata":{"execution":{"iopub.status.busy":"2022-11-23T11:14:54.750828Z","iopub.execute_input":"2022-11-23T11:14:54.751464Z","iopub.status.idle":"2022-11-23T11:14:54.787945Z","shell.execute_reply.started":"2022-11-23T11:14:54.751424Z","shell.execute_reply":"2022-11-23T11:14:54.787052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Model build 1**","metadata":{}},{"cell_type":"code","source":"# set ROC AUC as metric\nROC_1 = tf.keras.metrics.AUC()\n\n# use GPU\nwith strategy.scope():\n    \n    #model design\n    model_one = Sequential()\n    \n    model_one.add(Conv2D(filters=16, kernel_size=(3,3)))\n    model_one.add(Conv2D(filters=16, kernel_size=(3,3)))\n    model_one.add(MaxPooling2D(pool_size=(2,2)))\n\n    model_one.add(Conv2D(filters=32, kernel_size=(3,3)))\n    model_one.add(Conv2D(filters=32, kernel_size=(3,3)))\n    model_one.add(AveragePooling2D(pool_size=(2,2)))\n\n    model_one.add(Flatten())\n    model_one.add(Dense(1, activation='sigmoid'))\n    \n    #model input size\n    model_one.build(input_shape=(BATCH_SIZE, 64, 64, 3))        # original image = (96, 96, 3) \n    \n    #compiling\n    model_one.compile(loss='binary_crossentropy', metrics=['accuracy', ROC_1])\n    \nmodel_one.summary()","metadata":{"execution":{"iopub.status.busy":"2022-11-23T11:20:21.173132Z","iopub.execute_input":"2022-11-23T11:20:21.173568Z","iopub.status.idle":"2022-11-23T11:20:21.240963Z","shell.execute_reply.started":"2022-11-23T11:20:21.173521Z","shell.execute_reply":"2022-11-23T11:20:21.239701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"EPOCHS = 10\n\n# train the model\nhistory_model_one = model_one.fit_generator(\n                        train_generator,\n                        epochs = EPOCHS,\n                        validation_data = valid_generator)","metadata":{"execution":{"iopub.status.busy":"2022-11-23T13:27:42.742201Z","iopub.execute_input":"2022-11-23T13:27:42.743206Z","iopub.status.idle":"2022-11-23T14:51:53.232857Z","shell.execute_reply.started":"2022-11-23T13:27:42.743163Z","shell.execute_reply":"2022-11-23T14:51:53.231778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot model accuracy per epoch \nplt.plot(history_model_one.history['accuracy'])\nplt.plot(history_model_one.history['val_accuracy'])\nplt.title('Model One Accuracy per Epoch')\nplt.ylabel('accuracy')\nplt.xlabel('epoch')\nplt.legend(['train', 'validate'], loc='upper left')\nplt.show();\n\n# plot model loss per epoch\nplt.plot(history_model_one.history['loss'])\nplt.plot(history_model_one.history['val_loss'])\nplt.title('Model One Loss per Epoch')\nplt.ylabel('loss')\nplt.xlabel('epoch')\nplt.legend(['train', 'validate'], loc='upper left')\nplt.show();\n\n# plot model ROC per epoch\nplt.plot(history_model_one.history['auc'])\nplt.plot(history_model_one.history['val_auc'])\nplt.title('Model One AUC ROC per Epoch')\nplt.ylabel('ROC')\nplt.xlabel('epoch')\nplt.legend(['train', 'validate'], loc='upper left')\nplt.show();","metadata":{"execution":{"iopub.status.busy":"2022-11-23T15:17:58.962223Z","iopub.execute_input":"2022-11-23T15:17:58.962588Z","iopub.status.idle":"2022-11-23T15:17:59.673092Z","shell.execute_reply.started":"2022-11-23T15:17:58.962557Z","shell.execute_reply":"2022-11-23T15:17:59.672060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Model build 2**","metadata":{}},{"cell_type":"code","source":"ROC_2 = tf.keras.metrics.AUC()\n\nwith strategy.scope():\n    \n    #create model\n    model_two = Sequential()\n    \n    model_two.add(Conv2D(filters=16, kernel_size=(3,3), activation='relu', ))\n    model_two.add(Conv2D(filters=16, kernel_size=(3,3), activation='relu'))\n    model_two.add(MaxPooling2D(pool_size=(2,2)))\n    model_two.add(Dropout(0.1))\n    \n    model_two.add(BatchNormalization())\n    model_two.add(Conv2D(filters=32, kernel_size=(3,3), activation='relu'))\n    model_two.add(Conv2D(filters=32, kernel_size=(3,3), activation='relu'))\n    model_two.add(AveragePooling2D(pool_size=(2,2)))\n    model_two.add(Dropout(0.1))\n    \n    model_two.add(BatchNormalization())\n    model_two.add(Conv2D(filters=32, kernel_size=(3,3), activation='relu'))\n    model_two.add(Flatten())\n    model_two.add(Dense(1, activation='sigmoid'))\n    \n    #build model by input size\n    model_two.build(input_shape=(BATCH_SIZE, 64, 64, 3))       # original image = (96, 96, 3) \n    \n    #compile\n    adam_optimizer = Adam(learning_rate=0.0001)\n    model_two.compile(loss='binary_crossentropy', metrics=['accuracy', ROC_2], optimizer=adam_optimizer)\n\n#quick look at model\nmodel_two.summary()","metadata":{"execution":{"iopub.status.busy":"2022-11-23T15:17:59.675122Z","iopub.execute_input":"2022-11-23T15:17:59.676004Z","iopub.status.idle":"2022-11-23T15:17:59.779827Z","shell.execute_reply.started":"2022-11-23T15:17:59.675962Z","shell.execute_reply":"2022-11-23T15:17:59.778869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"EPOCHS = 10\n\n# train model\nhistory_model_two = model_two.fit_generator(\n                        train_generator,\n                        epochs = EPOCHS,\n                        validation_data = valid_generator)","metadata":{"execution":{"iopub.status.busy":"2022-11-23T15:17:59.781210Z","iopub.execute_input":"2022-11-23T15:17:59.781556Z","iopub.status.idle":"2022-11-23T17:05:10.373567Z","shell.execute_reply.started":"2022-11-23T15:17:59.781521Z","shell.execute_reply":"2022-11-23T17:05:10.370888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# graph loss\nplt.plot(history_model_two.history['accuracy'])\nplt.plot(history_model_two.history['val_accuracy'])\nplt.title('Model Two Accuracy')\nplt.ylabel('accuracy')\nplt.xlabel('epoch')\nplt.legend(['train', 'validate'], loc='upper left')\nplt.show();\n\nplt.plot(history_model_two.history['loss'])\nplt.plot(history_model_two.history['val_loss'])\nplt.title('Model Two Loss')\nplt.ylabel('loss')\nplt.xlabel('epoch')\nplt.legend(['train', 'validate'], loc='upper left')\nplt.show();\n\n# plot model ROC per epoch\nplt.plot(history_model_two.history['auc_1'])\nplt.plot(history_model_two.history['val_auc_1'])\nplt.title('Model Two AUC ROC per Epoch')\nplt.ylabel('ROC')\nplt.xlabel('epoch')\nplt.legend(['train', 'validate'], loc='upper left')\nplt.show();\n","metadata":{"execution":{"iopub.status.busy":"2022-11-23T17:05:41.810840Z","iopub.execute_input":"2022-11-23T17:05:41.811271Z","iopub.status.idle":"2022-11-23T17:05:42.529097Z","shell.execute_reply.started":"2022-11-23T17:05:41.811237Z","shell.execute_reply":"2022-11-23T17:05:42.528151Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# preparing test data\ndatagen_test = ImageDataGenerator(rescale=1./255.)\n\ntest_generator = datagen_test.flow_from_dataframe(\n    dataframe=test_data,\n    directory=test_path,\n    x_col='id', \n    y_col=None,\n    target_size=(64,64),         \n    batch_size=1,\n    shuffle=False,\n    class_mode=None)","metadata":{"execution":{"iopub.status.busy":"2022-11-23T17:05:42.531034Z","iopub.execute_input":"2022-11-23T17:05:42.531372Z","iopub.status.idle":"2022-11-23T17:09:39.785007Z","shell.execute_reply.started":"2022-11-23T17:05:42.531344Z","shell.execute_reply":"2022-11-23T17:09:39.783926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#creating predictions\n\npredictions = model_two.predict(test_generator, verbose=1)","metadata":{"execution":{"iopub.status.busy":"2022-11-23T17:09:39.786467Z","iopub.execute_input":"2022-11-23T17:09:39.787476Z","iopub.status.idle":"2022-11-23T17:16:27.044816Z","shell.execute_reply.started":"2022-11-23T17:09:39.787436Z","shell.execute_reply":"2022-11-23T17:16:27.043764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#create submission dataframe\npredictions = np.transpose(predictions)[0]\nsubmission_df = pd.DataFrame()\nsubmission_df['id'] = test_data['id'].apply(lambda x: x.split('.')[0])\nsubmission_df['label'] = list(map(lambda x: 0 if x < 0.5 else 1, predictions))\nsubmission_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-11-23T17:16:27.051036Z","iopub.execute_input":"2022-11-23T17:16:27.054187Z","iopub.status.idle":"2022-11-23T17:16:27.218523Z","shell.execute_reply.started":"2022-11-23T17:16:27.054148Z","shell.execute_reply":"2022-11-23T17:16:27.217555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df['label'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-11-23T18:13:34.093602Z","iopub.execute_input":"2022-11-23T18:13:34.093985Z","iopub.status.idle":"2022-11-23T18:13:34.103894Z","shell.execute_reply.started":"2022-11-23T18:13:34.093954Z","shell.execute_reply":"2022-11-23T18:13:34.102616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Submission label percentage\nsns.set(style='whitegrid')\npie_chart=pd.DataFrame(submission_df['label'].replace(0,'Non-cancerous tissue').replace(1,'Cancerous tissue').value_counts())\npie_chart.reset_index(inplace=True)\npie_chart.plot(kind='pie', title='Submission Image Classification',y = 'label', \n             autopct='%1.1f%%', shadow=False, labels=pie_chart['index'], legend = False, fontsize=14, figsize=(18,8))","metadata":{"execution":{"iopub.status.busy":"2022-11-23T18:11:57.763489Z","iopub.execute_input":"2022-11-23T18:11:57.763872Z","iopub.status.idle":"2022-11-23T18:11:57.913760Z","shell.execute_reply.started":"2022-11-23T18:11:57.763842Z","shell.execute_reply":"2022-11-23T18:11:57.912427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Competition Submission\nsubmission_df.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-11-23T19:33:58.567322Z","iopub.execute_input":"2022-11-23T19:33:58.567701Z","iopub.status.idle":"2022-11-23T19:33:58.696428Z","shell.execute_reply.started":"2022-11-23T19:33:58.567669Z","shell.execute_reply":"2022-11-23T19:33:58.695090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Conclusion and Results \n\n**Results**\nFrom the above plots and diagrams for each model how well they performed with the training and validation sets. We see that model two performed slightly better than the model one and this was attributed to the hyperparameter tuning. \n\n**Conclusion**\nOverall the model performed fairly accurate predictions. Due to time and memory constraints, we could only deploy a simple CNN network.\nSome ways to make the models better could be a variety of factors. A valid suggestion would be to train the model with augmented images. In the preprocessing step, we only normalized the images. Instead, you could normalize, flip, zoom in/out, stretch, rotate, etc. with the images so that the model learns more instances of how images can be shaped.Using more epochs could allow the model to learn better. However, it is important that we don't overfit the data by training it too long. Since this project is for demonstration purposes, we did not use more epochs either. It is difficult to say which modification to the notebook would return much better results without actually doing it and running the models however, we are confident that these are good steps to try in another notebook.\n","metadata":{}}]}