{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 1. Problem Statement\n\n> ### Task - The problem is mainly a BINARY IMAGE CLASSIFICATION PROBLEM. The Problem focuses on identifying the presence of metastases from a 96 * 96 digital histopathology images\n\n> ### Metric Evaluation - Submissions are evaluated on area under the ROC curve between the predicted probability and the observed target. \n\n<img src='https://i.stack.imgur.com/kqxaJ.png' style=\"width:500px;height:300px;\">\n\n\n# 2. Analysis of the problem Statment\n\n> ## What Exactly the problem statment conveys to us?\n> ### 1. The problem deals with the Binary Classification of the Image that has a shape of 96px * 96px. It involves identifying the metastases from the 96px * 96px digital histapathology images.\n\n> ### 2. One key challenge is that the metastases can be as small as single cells in a large area of tissue.\n","metadata":{"papermill":{"duration":0.037832,"end_time":"2020-09-22T03:02:53.610792","exception":false,"start_time":"2020-09-22T03:02:53.572960","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"### About the Domain: \nObviously, I do not know much about Biology,I made some notes about the following terminologies :\n* Histopathology\n* Lymphocytes\n* Lymph Nodes\n\n### So, let us see some of the biological terminologies involved \n\n### 1. Histopathology - Histopathology is the diagnosis and study of diseases of the tissues, and involves examining tissues and/or cells under a microscope. Histopathologists are responsible for making tissue diagnoses and helping clinicians manage a patient's care.\n\n\n### 2. Lymphocytes - Lymphocytes are white blood cells that are also one of the body's main types of immune cells. They are made in the bone marrow and found in the blood and lymph tissue. The immune system is a complex network of cells known as immune cells that include lymphocytes.\n\n\n### 3. Lymph Nodes- Lymph nodes are small lumps of tissue that contain white blood cells, which fight infection. They filter lymph fluid, which is composed of fluid and waste products from your body tissues. Lymph nodes also help activate your immune system if you have an infection.\n\n\n","metadata":{"papermill":{"duration":0.038134,"end_time":"2020-09-22T03:02:53.765466","exception":false,"start_time":"2020-09-22T03:02:53.727332","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# 2.  Data Understanding\n\n* The dataset contains the histopathological Images, each image is 96px * 96px. \n\n* A positive label indicates that the center 32x32px region of a patch contains at least one pixel of tumor tissue. Tumor tissue in the outer region of the patch does not influence the label. This outer region is provided to enable fully-convolutional models that do not use zero-padding, to ensure consistent behavior when applied to a whole-slide image.\n\n* Kaggle says that :\n                    'The original PCam dataset contains duplicate images due to its probabilistic sampling,\n                    however, the version presented on Kaggle does not contain duplicates. We have otherwise \n                    maintained the same data and splits as the PCam benchmark.'\n                   \n* Also, one of the hing is that the problem states that the training Data contains **50/50** Images of both the labels i.e. the training contains equal proportion of both the labels, however on analysis it was found to be nearly equal to **60/40**, which we will consider while we design the model\n\n\n\n\n* ### **IS DATA RELEVANT TO THE PROBLEM ?**\n> This dataset is a combination of two independent datasets collected in Radboud University Medical Center (Nijmegen, the Netherlands), and the University Medical Center Utrecht (Utrecht, the Netherlands). The slides are produced by routine clinical practices and a trained pathologist would examine similar images for identifying metastases.\n","metadata":{"papermill":{"duration":0.039215,"end_time":"2020-09-22T03:02:53.920005","exception":false,"start_time":"2020-09-22T03:02:53.880790","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# 3. Designing the Model","metadata":{"papermill":{"duration":0.038351,"end_time":"2020-09-22T03:02:53.997149","exception":false,"start_time":"2020-09-22T03:02:53.958798","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Importing  Libraries\nfrom numpy.random import seed\nseed(101)\n\nimport pandas as pd\nimport numpy as np\n\n\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras.layers import Conv2D, MaxPooling2D\nfrom tensorflow.keras.layers import Dense, Dropout, Flatten, Activation\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.callbacks import EarlyStopping, ReduceLROnPlateau, ModelCheckpoint\nfrom tensorflow.keras.optimizers import Adam\n\nimport os\nimport cv2\n\nfrom sklearn.utils import shuffle\nfrom sklearn.metrics import confusion_matrix\nfrom sklearn.model_selection import train_test_split\nimport itertools\nimport shutil\nimport matplotlib.pyplot as plt\n%matplotlib inline\ntf.random.set_seed(101)","metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","papermill":{"duration":6.552565,"end_time":"2020-09-22T03:03:00.589774","exception":false,"start_time":"2020-09-22T03:02:54.037209","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:34:26.195365Z","iopub.execute_input":"2023-03-17T14:34:26.195817Z","iopub.status.idle":"2023-03-17T14:34:34.929125Z","shell.execute_reply.started":"2023-03-17T14:34:26.195772Z","shell.execute_reply":"2023-03-17T14:34:34.928036Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Setting Some Pre-Requisites\nIMAGE_SIZE=96\nIMAGE_CHANNELS=3\nSAMPLE_SIZE=80000         # We will be training 80,000 samples from each label","metadata":{"papermill":{"duration":0.04744,"end_time":"2020-09-22T03:03:00.677681","exception":false,"start_time":"2020-09-22T03:03:00.630241","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:34:34.931425Z","iopub.execute_input":"2023-03-17T14:34:34.932199Z","iopub.status.idle":"2023-03-17T14:34:34.938079Z","shell.execute_reply.started":"2023-03-17T14:34:34.932160Z","shell.execute_reply":"2023-03-17T14:34:34.936120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# So, what are the files which are available?\n\nos.listdir('../input/histopathologic-cancer-detection')","metadata":{"papermill":{"duration":0.050925,"end_time":"2020-09-22T03:03:00.768717","exception":false,"start_time":"2020-09-22T03:03:00.717792","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:34:34.939468Z","iopub.execute_input":"2023-03-17T14:34:34.940497Z","iopub.status.idle":"2023-03-17T14:34:34.954047Z","shell.execute_reply.started":"2023-03-17T14:34:34.940460Z","shell.execute_reply":"2023-03-17T14:34:34.953020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# So, how many images are there in each of the folder in the training dataset?\n\nprint(len(os.listdir('../input/histopathologic-cancer-detection/train')))\nprint(len(os.listdir('../input/histopathologic-cancer-detection/test')))","metadata":{"papermill":{"duration":5.191953,"end_time":"2020-09-22T03:03:06.000818","exception":false,"start_time":"2020-09-22T03:03:00.808865","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:34:34.958031Z","iopub.execute_input":"2023-03-17T14:34:34.958516Z","iopub.status.idle":"2023-03-17T14:34:38.848156Z","shell.execute_reply.started":"2023-03-17T14:34:34.958488Z","shell.execute_reply":"2023-03-17T14:34:38.847040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Creating a dataframe of all the training images\n\ndf_data = pd.read_csv('../input/histopathologic-cancer-detection/train_labels.csv')\n\n# removing this image because it caused a training error previously\ndf_data[df_data['id'] != 'dd6dfed324f9fcb6f93f46f32fc800f2ec196be2']\n\n# removing this image because it's black\ndf_data[df_data['id'] != '9369c7278ec8bcc6c880d99194de09fc2bd4efbe']\n\n\nprint(df_data.shape)","metadata":{"papermill":{"duration":0.393054,"end_time":"2020-09-22T03:03:06.437063","exception":false,"start_time":"2020-09-22T03:03:06.044009","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:34:38.849805Z","iopub.execute_input":"2023-03-17T14:34:38.850465Z","iopub.status.idle":"2023-03-17T14:34:39.286458Z","shell.execute_reply.started":"2023-03-17T14:34:38.850425Z","shell.execute_reply":"2023-03-17T14:34:39.284607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_data['label'].value_counts()","metadata":{"papermill":{"duration":0.059222,"end_time":"2020-09-22T03:03:06.539680","exception":false,"start_time":"2020-09-22T03:03:06.480458","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:34:39.288181Z","iopub.execute_input":"2023-03-17T14:34:39.288532Z","iopub.status.idle":"2023-03-17T14:34:39.302296Z","shell.execute_reply.started":"2023-03-17T14:34:39.288492Z","shell.execute_reply":"2023-03-17T14:34:39.301040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# source: https://www.kaggle.com/gpreda/honey-bee-subspecies-classification\n\ndef draw_category_images(col_name,figure_cols, df, IMAGE_PATH):\n    \n    \"\"\"\n    Give a column in a dataframe,\n    this function takes a sample of each class and displays that\n    sample on one row. The sample size is the same as figure_cols which\n    is the number of columns in the figure.\n    Because this function takes a random sample, each time the function is run it\n    displays different images.\n    \"\"\"\n    \n\n    categories = (df.groupby([col_name])[col_name].nunique()).index\n    f, ax = plt.subplots(nrows=len(categories),ncols=figure_cols, \n                         figsize=(4*figure_cols,4*len(categories))) # adjust size here\n    # draw a number of images for each location\n    for i, cat in enumerate(categories):\n        sample = df[df[col_name]==cat].sample(figure_cols) # figure_cols is also the sample size\n        for j in range(0,figure_cols):\n            file=IMAGE_PATH + sample.iloc[j]['id'] + '.tif'\n            im=cv2.imread(file)\n            ax[i, j].imshow(im, resample=True, cmap='gray')\n            ax[i, j].set_title(cat, fontsize=16)  \n    plt.tight_layout()\n    plt.show()","metadata":{"papermill":{"duration":0.057307,"end_time":"2020-09-22T03:03:06.640679","exception":false,"start_time":"2020-09-22T03:03:06.583372","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:34:39.303742Z","iopub.execute_input":"2023-03-17T14:34:39.304142Z","iopub.status.idle":"2023-03-17T14:34:39.315180Z","shell.execute_reply.started":"2023-03-17T14:34:39.304104Z","shell.execute_reply":"2023-03-17T14:34:39.313802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"IMAGE_PATH = '../input/histopathologic-cancer-detection/train/' \n\ndraw_category_images('label',4, df_data, IMAGE_PATH)","metadata":{"papermill":{"duration":2.160484,"end_time":"2020-09-22T03:03:08.843893","exception":false,"start_time":"2020-09-22T03:03:06.683409","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:34:39.317025Z","iopub.execute_input":"2023-03-17T14:34:39.317497Z","iopub.status.idle":"2023-03-17T14:34:40.536441Z","shell.execute_reply.started":"2023-03-17T14:34:39.317457Z","shell.execute_reply":"2023-03-17T14:34:40.535525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create the Train and Validation Sets\n\ndf_0=df_data[df_data['label']==0].sample(SAMPLE_SIZE,random_state=101)\ndf_1=df_data[df_data['label']==1].sample(SAMPLE_SIZE,random_state=101)\n\n# concat the dataframes\ndf_data = pd.concat([df_0, df_1], axis=0).reset_index(drop=True)\n# shuffle\ndf_data = shuffle(df_data)\n\ndf_data['label'].value_counts()","metadata":{"papermill":{"duration":0.226444,"end_time":"2020-09-22T03:03:09.195097","exception":false,"start_time":"2020-09-22T03:03:08.968653","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:34:40.537429Z","iopub.execute_input":"2023-03-17T14:34:40.537779Z","iopub.status.idle":"2023-03-17T14:34:40.602967Z","shell.execute_reply.started":"2023-03-17T14:34:40.537742Z","shell.execute_reply":"2023-03-17T14:34:40.601883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Now, for the train-test split\n\n# stratify=y creates a balanced validation set.\ny = df_data['label']\n\ndf_train, df_val = train_test_split(df_data, test_size=0.10, random_state=101, stratify=y)\n\nprint(df_train.shape)\nprint(df_val.shape)","metadata":{"papermill":{"duration":0.267586,"end_time":"2020-09-22T03:03:09.562763","exception":false,"start_time":"2020-09-22T03:03:09.295177","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:34:40.608417Z","iopub.execute_input":"2023-03-17T14:34:40.609044Z","iopub.status.idle":"2023-03-17T14:34:40.675993Z","shell.execute_reply.started":"2023-03-17T14:34:40.609004Z","shell.execute_reply":"2023-03-17T14:34:40.674826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a new directory so that we will be using the ImageDataGenerator\nbase_dir='base_dir'\nos.mkdir(base_dir)\n\n# now we create 2 folders inside 'base_dir':\n\n# train_dir\n    # a_no_tumor_tissue\n    # b_has_tumor_tissue\n\n# val_dir\n    # a_no_tumor_tissue\n    # b_has_tumor_tissue\n\n\n\n# create a path to 'base_dir' to which we will join the names of the new folders\n# train_dir\ntrain_dir = os.path.join(base_dir, 'train_dir')\nos.mkdir(train_dir)\n\n# val_dir\nval_dir = os.path.join(base_dir, 'val_dir')\nos.mkdir(val_dir)\n\n\n\n# [CREATE FOLDERS INSIDE THE TRAIN AND VALIDATION FOLDERS]\n# Inside each folder we create seperate folders for each class\n\n# create new folders inside train_dir\nno_tumor_tissue = os.path.join(train_dir, 'a_no_tumor_tissue')\nos.mkdir(no_tumor_tissue)\nhas_tumor_tissue = os.path.join(train_dir, 'b_has_tumor_tissue')\nos.mkdir(has_tumor_tissue)\n\n\n# create new folders inside val_dir\nno_tumor_tissue = os.path.join(val_dir, 'a_no_tumor_tissue')\nos.mkdir(no_tumor_tissue)\nhas_tumor_tissue = os.path.join(val_dir, 'b_has_tumor_tissue')\nos.mkdir(has_tumor_tissue)","metadata":{"papermill":{"duration":0.136193,"end_time":"2020-09-22T03:03:09.805179","exception":false,"start_time":"2020-09-22T03:03:09.668986","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:34:40.677827Z","iopub.execute_input":"2023-03-17T14:34:40.678225Z","iopub.status.idle":"2023-03-17T14:34:40.687322Z","shell.execute_reply.started":"2023-03-17T14:34:40.678185Z","shell.execute_reply":"2023-03-17T14:34:40.686270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check that the folders have been created\nos.listdir('base_dir/train_dir')","metadata":{"papermill":{"duration":0.077164,"end_time":"2020-09-22T03:03:09.950867","exception":false,"start_time":"2020-09-22T03:03:09.873703","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:34:40.689016Z","iopub.execute_input":"2023-03-17T14:34:40.689374Z","iopub.status.idle":"2023-03-17T14:34:40.701927Z","shell.execute_reply.started":"2023-03-17T14:34:40.689339Z","shell.execute_reply":"2023-03-17T14:34:40.700766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set the id as the index in df_data\ndf_data.set_index('id', inplace=True)","metadata":{"papermill":{"duration":0.088484,"end_time":"2020-09-22T03:03:10.106868","exception":false,"start_time":"2020-09-22T03:03:10.018384","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:34:40.703784Z","iopub.execute_input":"2023-03-17T14:34:40.704289Z","iopub.status.idle":"2023-03-17T14:34:40.711012Z","shell.execute_reply.started":"2023-03-17T14:34:40.704253Z","shell.execute_reply":"2023-03-17T14:34:40.710118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get a list of train and val images\ntrain_list = list(df_train['id'])\nval_list = list(df_val['id'])\n\n\n\n# Transfer the train images\n\nfor image in train_list:\n    \n    # the id in the csv file does not have the .tif extension therefore we add it here\n    fname = image + '.tif'\n    # get the label for a certain image\n    target = df_data.loc[image,'label']\n    \n    # these must match the folder names\n    if target == 0:\n        label = 'a_no_tumor_tissue'\n    if target == 1:\n        label = 'b_has_tumor_tissue'\n    \n    # source path to image\n    src = os.path.join('../input/histopathologic-cancer-detection/train', fname)\n    # destination path to image\n    dst = os.path.join(train_dir, label, fname)\n    # copy the image from the source to the destination\n    shutil.copyfile(src, dst)\n\n\n# Transfer the val images\n\nfor image in val_list:\n    \n    # the id in the csv file does not have the .tif extension therefore we add it here\n    fname = image + '.tif'\n    # get the label for a certain image\n    target = df_data.loc[image,'label']\n    \n    # these must match the folder names\n    if target == 0:\n        label = 'a_no_tumor_tissue'\n    if target == 1:\n        label = 'b_has_tumor_tissue'\n    \n\n    # source path to image\n    src = os.path.join('../input/histopathologic-cancer-detection/train', fname)\n    # destination path to image\n    dst = os.path.join(val_dir, label, fname)\n    # copy the image from the source to the destination\n    shutil.copyfile(src, dst)","metadata":{"papermill":{"duration":356.828724,"end_time":"2020-09-22T03:09:07.003601","exception":false,"start_time":"2020-09-22T03:03:10.174877","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:34:40.712369Z","iopub.execute_input":"2023-03-17T14:34:40.712874Z","iopub.status.idle":"2023-03-17T14:52:39.014399Z","shell.execute_reply.started":"2023-03-17T14:34:40.712837Z","shell.execute_reply":"2023-03-17T14:52:39.013325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check how many train images we have in each folder\n\nprint(len(os.listdir('base_dir/train_dir/a_no_tumor_tissue')))\nprint(len(os.listdir('base_dir/train_dir/b_has_tumor_tissue')))","metadata":{"papermill":{"duration":0.191751,"end_time":"2020-09-22T03:09:07.268978","exception":false,"start_time":"2020-09-22T03:09:07.077227","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:52:39.019134Z","iopub.execute_input":"2023-03-17T14:52:39.021781Z","iopub.status.idle":"2023-03-17T14:52:39.157129Z","shell.execute_reply.started":"2023-03-17T14:52:39.021738Z","shell.execute_reply":"2023-03-17T14:52:39.156227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check how many val images we have in each folder\n\nprint(len(os.listdir('base_dir/val_dir/a_no_tumor_tissue')))\nprint(len(os.listdir('base_dir/val_dir/b_has_tumor_tissue')))","metadata":{"papermill":{"duration":0.089046,"end_time":"2020-09-22T03:09:07.423606","exception":false,"start_time":"2020-09-22T03:09:07.334560","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:52:39.161281Z","iopub.execute_input":"2023-03-17T14:52:39.163572Z","iopub.status.idle":"2023-03-17T14:52:39.187221Z","shell.execute_reply.started":"2023-03-17T14:52:39.163533Z","shell.execute_reply":"2023-03-17T14:52:39.186347Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set up the generators\ntrain_path = 'base_dir/train_dir'\nvalid_path = 'base_dir/val_dir'\ntest_path = '../input/histopathologic-cancer-detection/test'\n\nnum_train_samples = len(df_train)\nnum_val_samples = len(df_val)\ntrain_batch_size = 10\nval_batch_size = 10\n\n\ntrain_steps = np.ceil(num_train_samples / train_batch_size)\nval_steps = np.ceil(num_val_samples / val_batch_size)","metadata":{"papermill":{"duration":0.08,"end_time":"2020-09-22T03:09:07.570221","exception":false,"start_time":"2020-09-22T03:09:07.490221","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:52:39.191564Z","iopub.execute_input":"2023-03-17T14:52:39.193832Z","iopub.status.idle":"2023-03-17T14:52:39.201857Z","shell.execute_reply.started":"2023-03-17T14:52:39.193795Z","shell.execute_reply":"2023-03-17T14:52:39.200804Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"datagen = ImageDataGenerator(rescale=1.0/255)\n\ntrain_gen = datagen.flow_from_directory(train_path,\n                                        target_size=(IMAGE_SIZE,IMAGE_SIZE),\n                                        batch_size=train_batch_size,\n                                        class_mode='categorical')\n\nval_gen = datagen.flow_from_directory(valid_path,\n                                        target_size=(IMAGE_SIZE,IMAGE_SIZE),\n                                        batch_size=val_batch_size,\n                                        class_mode='categorical')\n\n# Note: shuffle=False causes the test dataset to not be shuffled\ntest_gen = datagen.flow_from_directory(valid_path,\n                                        target_size=(IMAGE_SIZE,IMAGE_SIZE),\n                                        batch_size=1,\n                                        class_mode='categorical',\n                                        shuffle=False)","metadata":{"papermill":{"duration":11.181689,"end_time":"2020-09-22T03:09:18.817994","exception":false,"start_time":"2020-09-22T03:09:07.636305","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:52:39.203256Z","iopub.execute_input":"2023-03-17T14:52:39.203821Z","iopub.status.idle":"2023-03-17T14:52:47.049880Z","shell.execute_reply.started":"2023-03-17T14:52:39.203788Z","shell.execute_reply":"2023-03-17T14:52:47.048881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The model that I have choosen for this problem has been taken from <a href = 'https://www.kaggle.com/fmarazzi/baseline-keras-cnn-roc-fast-10min-0-925-lb'>Baseline Keras CNN</a>","metadata":{"papermill":{"duration":0.067085,"end_time":"2020-09-22T03:09:18.953557","exception":false,"start_time":"2020-09-22T03:09:18.886472","status":"completed"},"tags":[]}},{"cell_type":"code","source":"kernel_size = (3,3)\npool_size= (2,2)\nfirst_filters = 32\nsecond_filters = 64\nthird_filters = 128\n\ndropout_conv = 0.3\ndropout_dense = 0.3\n\n\nmodel = Sequential()\nmodel.add(Conv2D(first_filters, kernel_size, activation = 'relu', input_shape = (96, 96, 3)))\nmodel.add(Conv2D(first_filters, kernel_size, activation = 'relu'))\nmodel.add(Conv2D(first_filters, kernel_size, activation = 'relu'))\nmodel.add(MaxPooling2D(pool_size = pool_size)) \nmodel.add(Dropout(dropout_conv))\n\nmodel.add(Conv2D(second_filters, kernel_size, activation ='relu'))\nmodel.add(Conv2D(second_filters, kernel_size, activation ='relu'))\nmodel.add(Conv2D(second_filters, kernel_size, activation ='relu'))\nmodel.add(MaxPooling2D(pool_size = pool_size))\nmodel.add(Dropout(dropout_conv))\n\nmodel.add(Conv2D(third_filters, kernel_size, activation ='relu'))\nmodel.add(Conv2D(third_filters, kernel_size, activation ='relu'))\nmodel.add(Conv2D(third_filters, kernel_size, activation ='relu'))\nmodel.add(MaxPooling2D(pool_size = pool_size))\nmodel.add(Dropout(dropout_conv))\n\nmodel.add(Flatten())\nmodel.add(Dense(256, activation = \"relu\"))\nmodel.add(Dropout(dropout_dense))\nmodel.add(Dense(2, activation = \"softmax\"))\n\nmodel.summary()","metadata":{"papermill":{"duration":3.354194,"end_time":"2020-09-22T03:09:22.377049","exception":false,"start_time":"2020-09-22T03:09:19.022855","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:52:47.051303Z","iopub.execute_input":"2023-03-17T14:52:47.051766Z","iopub.status.idle":"2023-03-17T14:52:50.362708Z","shell.execute_reply.started":"2023-03-17T14:52:47.051727Z","shell.execute_reply":"2023-03-17T14:52:50.361875Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.compile(Adam(lr=0.0001), loss='binary_crossentropy', \n              metrics=['accuracy'])","metadata":{"papermill":{"duration":0.091086,"end_time":"2020-09-22T03:09:22.538813","exception":false,"start_time":"2020-09-22T03:09:22.447727","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:52:50.363742Z","iopub.execute_input":"2023-03-17T14:52:50.364100Z","iopub.status.idle":"2023-03-17T14:52:50.394787Z","shell.execute_reply.started":"2023-03-17T14:52:50.364062Z","shell.execute_reply":"2023-03-17T14:52:50.393818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the labels that are associated with each index\nprint(val_gen.class_indices)","metadata":{"papermill":{"duration":0.081487,"end_time":"2020-09-22T03:09:22.688346","exception":false,"start_time":"2020-09-22T03:09:22.606859","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:52:50.396246Z","iopub.execute_input":"2023-03-17T14:52:50.396854Z","iopub.status.idle":"2023-03-17T14:52:50.402190Z","shell.execute_reply.started":"2023-03-17T14:52:50.396818Z","shell.execute_reply":"2023-03-17T14:52:50.401129Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"filepath = \"model.h5\"\ncheckpoint = ModelCheckpoint(filepath, monitor='val_acc', verbose=1, \n                             save_best_only=True, mode='max')\n\nreduce_lr = ReduceLROnPlateau(monitor='val_acc', factor=0.5, patience=2, \n                                   verbose=1, mode='max', min_lr=0.00001)\n                              \n                              \ncallbacks_list = [checkpoint, reduce_lr]\n\nhistory = model.fit_generator(train_gen, steps_per_epoch=train_steps, \n                    validation_data=val_gen,\n                    validation_steps=val_steps,\n                    epochs=20, verbose=1,\n                   callbacks=callbacks_list)","metadata":{"papermill":{"duration":4258.409207,"end_time":"2020-09-22T04:20:21.202678","exception":false,"start_time":"2020-09-22T03:09:22.793471","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T14:52:50.403790Z","iopub.execute_input":"2023-03-17T14:52:50.404539Z","iopub.status.idle":"2023-03-17T16:11:43.744274Z","shell.execute_reply.started":"2023-03-17T14:52:50.404500Z","shell.execute_reply":"2023-03-17T16:11:43.743304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get the metric names so we can use evaulate_generator\nmodel.metrics_names","metadata":{"papermill":{"duration":25.534679,"end_time":"2020-09-22T04:21:12.861223","exception":false,"start_time":"2020-09-22T04:20:47.326544","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:11:43.746152Z","iopub.execute_input":"2023-03-17T16:11:43.746497Z","iopub.status.idle":"2023-03-17T16:11:43.756010Z","shell.execute_reply.started":"2023-03-17T16:11:43.746458Z","shell.execute_reply":"2023-03-17T16:11:43.754938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Here the best epoch will be used.\n\n\n\nval_loss, val_acc = \\\nmodel.evaluate_generator(test_gen, \n                        steps=len(df_val))\n\nprint('val_loss:', val_loss)\nprint('val_acc:', val_acc)","metadata":{"papermill":{"duration":64.43058,"end_time":"2020-09-22T04:22:43.906452","exception":false,"start_time":"2020-09-22T04:21:39.475872","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:11:43.757604Z","iopub.execute_input":"2023-03-17T16:11:43.758134Z","iopub.status.idle":"2023-03-17T16:13:05.734053Z","shell.execute_reply.started":"2023-03-17T16:11:43.758075Z","shell.execute_reply":"2023-03-17T16:13:05.732908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# display the loss and accuracy curves\n\nimport matplotlib.pyplot as plt\n\nacc = history.history['accuracy']\nval_acc = history.history['val_accuracy']\nloss = history.history['loss']\nval_loss = history.history['val_loss']\n\nepochs = range(1, len(acc) + 1)\n\nplt.plot(epochs, loss, 'bo', label='Training loss')\nplt.plot(epochs, val_loss, 'b', label='Validation loss')\nplt.title('Training and validation loss')\nplt.legend()\nplt.figure()\n\nplt.plot(epochs, acc, 'bo', label='Training acc')\nplt.plot(epochs, val_acc, 'b', label='Validation acc')\nplt.title('Training and validation accuracy')\nplt.legend()\nplt.figure()","metadata":{"papermill":{"duration":27.07921,"end_time":"2020-09-22T04:23:36.420132","exception":false,"start_time":"2020-09-22T04:23:09.340922","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:13:05.735608Z","iopub.execute_input":"2023-03-17T16:13:05.736181Z","iopub.status.idle":"2023-03-17T16:13:06.209917Z","shell.execute_reply.started":"2023-03-17T16:13:05.736140Z","shell.execute_reply":"2023-03-17T16:13:06.208905Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Validation and Analysis \n\n* ### Metrics\n* ### Prediction and Activation Visualizations\n* ### ROC and AUC","metadata":{"papermill":{"duration":26.017599,"end_time":"2020-09-22T04:24:28.207927","exception":false,"start_time":"2020-09-22T04:24:02.190328","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# make a prediction\npredictions = model.predict_generator(test_gen, steps=len(df_val), verbose=1)","metadata":{"papermill":{"duration":66.038298,"end_time":"2020-09-22T04:25:59.348357","exception":false,"start_time":"2020-09-22T04:24:53.310059","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:13:06.211285Z","iopub.execute_input":"2023-03-17T16:13:06.211903Z","iopub.status.idle":"2023-03-17T16:14:12.469412Z","shell.execute_reply.started":"2023-03-17T16:13:06.211861Z","shell.execute_reply":"2023-03-17T16:14:12.468332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions.shape","metadata":{"papermill":{"duration":26.701849,"end_time":"2020-09-22T04:26:54.046103","exception":false,"start_time":"2020-09-22T04:26:27.344254","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:14:12.473848Z","iopub.execute_input":"2023-03-17T16:14:12.474726Z","iopub.status.idle":"2023-03-17T16:14:12.482972Z","shell.execute_reply.started":"2023-03-17T16:14:12.474686Z","shell.execute_reply":"2023-03-17T16:14:12.481664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This is how to check what index keras has internally assigned to each class. \ntest_gen.class_indices","metadata":{"papermill":{"duration":26.898984,"end_time":"2020-09-22T04:27:47.518037","exception":false,"start_time":"2020-09-22T04:27:20.619053","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:14:12.490542Z","iopub.execute_input":"2023-03-17T16:14:12.490828Z","iopub.status.idle":"2023-03-17T16:14:12.497613Z","shell.execute_reply.started":"2023-03-17T16:14:12.490801Z","shell.execute_reply":"2023-03-17T16:14:12.496450Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Put the predictions into a dataframe.\n# The columns need to be ordered to match the output of the previous cell\n\ndf_preds = pd.DataFrame(predictions, columns=['no_tumor_tissue', 'has_tumor_tissue'])\n\ndf_preds.head()\n","metadata":{"papermill":{"duration":27.846086,"end_time":"2020-09-22T04:28:41.543386","exception":false,"start_time":"2020-09-22T04:28:13.697300","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:14:12.499357Z","iopub.execute_input":"2023-03-17T16:14:12.499986Z","iopub.status.idle":"2023-03-17T16:14:12.516697Z","shell.execute_reply.started":"2023-03-17T16:14:12.499924Z","shell.execute_reply":"2023-03-17T16:14:12.515716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the true labels\ny_true = test_gen.classes\n\n# Get the predicted labels as probabilities\ny_pred = df_preds['has_tumor_tissue']","metadata":{"papermill":{"duration":28.1778,"end_time":"2020-09-22T04:29:36.184008","exception":false,"start_time":"2020-09-22T04:29:08.006208","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:14:12.518319Z","iopub.execute_input":"2023-03-17T16:14:12.518919Z","iopub.status.idle":"2023-03-17T16:14:12.524812Z","shell.execute_reply.started":"2023-03-17T16:14:12.518883Z","shell.execute_reply":"2023-03-17T16:14:12.523646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import roc_auc_score\n\nroc_auc_score(y_true, y_pred)","metadata":{"papermill":{"duration":28.236966,"end_time":"2020-09-22T04:30:31.405742","exception":false,"start_time":"2020-09-22T04:30:03.168776","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:14:12.526728Z","iopub.execute_input":"2023-03-17T16:14:12.527101Z","iopub.status.idle":"2023-03-17T16:14:12.549383Z","shell.execute_reply.started":"2023-03-17T16:14:12.527066Z","shell.execute_reply":"2023-03-17T16:14:12.548450Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the labels of the test images.\n\ntest_labels = test_gen.classes\ntest_labels.shape","metadata":{"papermill":{"duration":28.191591,"end_time":"2020-09-22T04:31:26.910379","exception":false,"start_time":"2020-09-22T04:30:58.718788","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:14:12.550866Z","iopub.execute_input":"2023-03-17T16:14:12.551205Z","iopub.status.idle":"2023-03-17T16:14:12.557764Z","shell.execute_reply.started":"2023-03-17T16:14:12.551169Z","shell.execute_reply":"2023-03-17T16:14:12.556645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# argmax returns the index of the max value in a row\ncm = confusion_matrix(test_labels, predictions.argmax(axis=1))\n# Print the label associated with each class\ntest_gen.class_indices","metadata":{"papermill":{"duration":27.804324,"end_time":"2020-09-22T04:32:22.284748","exception":false,"start_time":"2020-09-22T04:31:54.480424","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:14:12.559456Z","iopub.execute_input":"2023-03-17T16:14:12.560374Z","iopub.status.idle":"2023-03-17T16:14:12.573007Z","shell.execute_reply.started":"2023-03-17T16:14:12.560336Z","shell.execute_reply":"2023-03-17T16:14:12.572034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import plot_confusion_matrix","metadata":{"papermill":{"duration":27.123071,"end_time":"2020-09-22T04:33:17.473019","exception":false,"start_time":"2020-09-22T04:32:50.349948","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:14:12.574281Z","iopub.execute_input":"2023-03-17T16:14:12.575316Z","iopub.status.idle":"2023-03-17T16:14:12.579913Z","shell.execute_reply.started":"2023-03-17T16:14:12.575280Z","shell.execute_reply":"2023-03-17T16:14:12.578887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Delete base_dir and it's sub folders to free up disk space.\n\nshutil.rmtree('base_dir')\n#[CREATE A TEST FOLDER DIRECTORY STRUCTURE]\n\n# We will be feeding test images from a folder into predict_generator().\n# Keras requires that the path should point to a folder containing images and not\n# to the images themselves. That is why we are creating a folder (test_images) \n# inside another folder (test_dir).\n\n# test_dir\n    # test_images\n\n# create test_dir\ntest_dir = 'test_dir'\nos.mkdir(test_dir)\n    \n# create test_images inside test_dir\ntest_images = os.path.join(test_dir, 'test_images')\nos.mkdir(test_images)\n# check that the directory we created exists\nos.listdir('test_dir')","metadata":{"papermill":{"duration":36.457038,"end_time":"2020-09-22T04:34:22.370355","exception":false,"start_time":"2020-09-22T04:33:45.913317","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:14:12.581378Z","iopub.execute_input":"2023-03-17T16:14:12.582544Z","iopub.status.idle":"2023-03-17T16:14:18.632101Z","shell.execute_reply.started":"2023-03-17T16:14:12.582505Z","shell.execute_reply":"2023-03-17T16:14:18.629917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Transfer the test images into image_dir\n\ntest_list = os.listdir('../input/histopathologic-cancer-detection/test')\n\nfor image in test_list:\n    \n    fname = image\n    \n    # source path to image\n    src = os.path.join('../input/histopathologic-cancer-detection/test', fname)\n    # destination path to image\n    dst = os.path.join(test_images, fname)\n    # copy the image from the source to the destination\n    shutil.copyfile(src, dst)\n# check that the images are now in the test_images\n# Should now be 57458 images in the test_images folder\n\nlen(os.listdir('test_dir/test_images'))","metadata":{"papermill":{"duration":148.722944,"end_time":"2020-09-22T04:37:17.956588","exception":false,"start_time":"2020-09-22T04:34:49.233644","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:14:18.633647Z","iopub.execute_input":"2023-03-17T16:14:18.634313Z","iopub.status.idle":"2023-03-17T16:23:58.590162Z","shell.execute_reply.started":"2023-03-17T16:14:18.634270Z","shell.execute_reply":"2023-03-17T16:23:58.589170Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_path ='test_dir'\n\n\n# Here we change the path to point to the test_images folder.\n\ntest_gen = datagen.flow_from_directory(test_path,\n                                        target_size=(IMAGE_SIZE,IMAGE_SIZE),\n                                        batch_size=1,\n                                        class_mode='categorical',\n                                        shuffle=False)","metadata":{"papermill":{"duration":29.436224,"end_time":"2020-09-22T04:38:16.091981","exception":false,"start_time":"2020-09-22T04:37:46.655757","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:23:58.591522Z","iopub.execute_input":"2023-03-17T16:23:58.591890Z","iopub.status.idle":"2023-03-17T16:24:00.039011Z","shell.execute_reply.started":"2023-03-17T16:23:58.591851Z","shell.execute_reply":"2023-03-17T16:24:00.037884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_test_images = 57458\n\n\n\npredictions = model.predict_generator(test_gen, steps=num_test_images, verbose=1)","metadata":{"papermill":{"duration":175.190036,"end_time":"2020-09-22T04:41:40.002246","exception":false,"start_time":"2020-09-22T04:38:44.812210","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:24:00.040301Z","iopub.execute_input":"2023-03-17T16:24:00.040677Z","iopub.status.idle":"2023-03-17T16:28:00.363519Z","shell.execute_reply.started":"2023-03-17T16:24:00.040639Z","shell.execute_reply":"2023-03-17T16:28:00.362403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Are the number of predictions correct?\n# Should be 57458.\n\nlen(predictions)","metadata":{"papermill":{"duration":29.556273,"end_time":"2020-09-22T04:42:37.878382","exception":false,"start_time":"2020-09-22T04:42:08.322109","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:28:00.378805Z","iopub.execute_input":"2023-03-17T16:28:00.379144Z","iopub.status.idle":"2023-03-17T16:28:00.387331Z","shell.execute_reply.started":"2023-03-17T16:28:00.379108Z","shell.execute_reply":"2023-03-17T16:28:00.386058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Put the predictions into a dataframe\n\ndf_preds = pd.DataFrame(predictions, columns=['no_tumor_tissue', 'has_tumor_tissue'])\n\ndf_preds.head()","metadata":{"papermill":{"duration":29.419676,"end_time":"2020-09-22T04:43:35.680065","exception":false,"start_time":"2020-09-22T04:43:06.260389","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:28:00.390088Z","iopub.execute_input":"2023-03-17T16:28:00.390787Z","iopub.status.idle":"2023-03-17T16:28:00.409286Z","shell.execute_reply.started":"2023-03-17T16:28:00.390743Z","shell.execute_reply":"2023-03-17T16:28:00.408112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This outputs the file names in the sequence in which \n# the generator processed the test images.\ntest_filenames = test_gen.filenames\n\n# add the filenames to the dataframe\ndf_preds['file_names'] = test_filenames\n\ndf_preds.head()","metadata":{"papermill":{"duration":29.588672,"end_time":"2020-09-22T04:44:34.169396","exception":false,"start_time":"2020-09-22T04:44:04.580724","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:28:00.411887Z","iopub.execute_input":"2023-03-17T16:28:00.412327Z","iopub.status.idle":"2023-03-17T16:28:00.430031Z","shell.execute_reply.started":"2023-03-17T16:28:00.412277Z","shell.execute_reply":"2023-03-17T16:28:00.429014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create an id column\n\n# A file name now has this format: \n# test_images/00006537328c33e284c973d7b39d340809f7271b.tif\n\n# This function will extract the id:\n# 00006537328c33e284c973d7b39d340809f7271b\n\n\ndef extract_id(x):\n    \n    # split into a list\n    a = x.split('/')\n    # split into a list\n    b = a[1].split('.')\n    extracted_id = b[0]\n    \n    return extracted_id\n\ndf_preds['id'] = df_preds['file_names'].apply(extract_id)\n\ndf_preds.head()","metadata":{"papermill":{"duration":29.757677,"end_time":"2020-09-22T04:45:32.139076","exception":false,"start_time":"2020-09-22T04:45:02.381399","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:28:00.431807Z","iopub.execute_input":"2023-03-17T16:28:00.432375Z","iopub.status.idle":"2023-03-17T16:28:00.483553Z","shell.execute_reply.started":"2023-03-17T16:28:00.432339Z","shell.execute_reply":"2023-03-17T16:28:00.482641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the predicted labels.\n# We were asked to predict a probability that the image has tumor tissue\ny_pred = df_preds['has_tumor_tissue']\n\n# get the id column\nimage_id = df_preds['id']","metadata":{"papermill":{"duration":29.551461,"end_time":"2020-09-22T04:46:29.694107","exception":false,"start_time":"2020-09-22T04:46:00.142646","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:28:00.485140Z","iopub.execute_input":"2023-03-17T16:28:00.485503Z","iopub.status.idle":"2023-03-17T16:28:00.490363Z","shell.execute_reply.started":"2023-03-17T16:28:00.485469Z","shell.execute_reply":"2023-03-17T16:28:00.489281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Confusion Matrix","metadata":{"papermill":{"duration":29.547336,"end_time":"2020-09-22T04:47:27.755152","exception":false,"start_time":"2020-09-22T04:46:58.207816","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# 5. Submission","metadata":{"papermill":{"duration":29.392735,"end_time":"2020-09-22T04:48:25.679221","exception":false,"start_time":"2020-09-22T04:47:56.286486","status":"completed"},"tags":[]}},{"cell_type":"code","source":"submission = pd.DataFrame({'id':image_id, \n                           'label':y_pred, \n                          }).set_index('id')\n\nsubmission.to_csv('patch_preds.csv', columns=['label']) \nsubmission.head()","metadata":{"papermill":{"duration":29.943072,"end_time":"2020-09-22T04:49:23.991342","exception":false,"start_time":"2020-09-22T04:48:54.048270","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:28:00.492044Z","iopub.execute_input":"2023-03-17T16:28:00.492751Z","iopub.status.idle":"2023-03-17T16:28:00.605245Z","shell.execute_reply.started":"2023-03-17T16:28:00.492704Z","shell.execute_reply":"2023-03-17T16:28:00.604262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Delete the test_dir directory we created to prevent a Kaggle error.\n# Kaggle allows a max of 500 files to be saved.\n\nshutil.rmtree('test_dir')","metadata":{"papermill":{"duration":30.889758,"end_time":"2020-09-22T04:50:23.418911","exception":false,"start_time":"2020-09-22T04:49:52.529153","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:28:00.607720Z","iopub.execute_input":"2023-03-17T16:28:00.608677Z","iopub.status.idle":"2023-03-17T16:28:02.785304Z","shell.execute_reply.started":"2023-03-17T16:28:00.608635Z","shell.execute_reply":"2023-03-17T16:28:02.784097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Confusion Matrix","metadata":{"papermill":{"duration":29.208911,"end_time":"2020-09-22T04:51:21.311273","exception":false,"start_time":"2020-09-22T04:50:52.102362","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt     \n\nax= plt.subplot()\nsns.heatmap(cm, annot=True, ax = ax); #annot=True to annotate cells\n\n# labels, title and ticks\nax.set_xlabel('Predicted labels');ax.set_ylabel('True labels'); \nax.set_title('Confusion Matrix'); ","metadata":{"papermill":{"duration":29.079485,"end_time":"2020-09-22T04:52:19.411350","exception":false,"start_time":"2020-09-22T04:51:50.331865","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-17T16:28:02.786738Z","iopub.execute_input":"2023-03-17T16:28:02.787411Z","iopub.status.idle":"2023-03-17T16:28:03.311875Z","shell.execute_reply.started":"2023-03-17T16:28:02.787368Z","shell.execute_reply":"2023-03-17T16:28:03.310929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Conclusion\nFrom the confusion matrix, we can conclude that the model correctly predicts true positives, with no false negatives. ","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}