{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Histopathologic cancer detection project using CNN\n\n<img src= \"https://storage.googleapis.com/kaggle-media/competitions/playground/Microscope\" style='width: 500px;'>","metadata":{}},{"cell_type":"markdown","source":"## Problem Statement\n\nThis is a project about creating & training a CNN model, **to predict labels of medical images.**\n\nAccording to the competition main page, the dataset is a slightly modified version of the PCam dataset, and a positive label value for each image indicates that the center 32x32px region of a patch contains at least one pixel of tumor tissue. Details of the dataset can be found here. https://github.com/basveeling/pcam\n\nInside the input directory, I can see one .csv file containing pairs of each image's id and the true label, and there are two folders that contain the actual images of **.tif format.**\n\nI will use **Keras** library to create / import some models, and optimize & evaluate the model parameters to finalize the submission output.\n\n## Environment Setup & Basic EDA","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport seaborn as sns\n\nimport matplotlib.pyplot as plt\nfrom tifffile import imread\n\nimport keras,os\nfrom keras.models import Sequential\nfrom keras.layers import Dense, Conv2D, MaxPool2D , Flatten, BatchNormalization, Activation\nfrom keras.preprocessing.image import ImageDataGenerator\nimport numpy as np\nfrom keras.optimizers import Adam","metadata":{"execution":{"iopub.status.busy":"2023-03-16T11:53:17.541888Z","iopub.execute_input":"2023-03-16T11:53:17.542418Z","iopub.status.idle":"2023-03-16T11:53:17.551277Z","shell.execute_reply.started":"2023-03-16T11:53:17.542330Z","shell.execute_reply":"2023-03-16T11:53:17.550189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dir_images_train = '../input/histopathologic-cancer-detection/train/'\nlabels_train = pd.read_csv('../input/histopathologic-cancer-detection/train_labels.csv')\nimages_train = labels_train\nimages_train['label'] = images_train['label'].astype(str)\nimages_train['id'] = labels_train['id'] + '.tif'\n\ndir_images_test = '../input/histopathologic-cancer-detection/test/'\n\nprint(labels_train.head(), '\\n')\nprint(images_train.head(), '\\n')\nprint(labels_train.info())","metadata":{"execution":{"iopub.status.busy":"2023-03-16T11:53:17.554117Z","iopub.execute_input":"2023-03-16T11:53:17.555090Z","iopub.status.idle":"2023-03-16T11:53:18.183144Z","shell.execute_reply.started":"2023-03-16T11:53:17.554951Z","shell.execute_reply":"2023-03-16T11:53:18.182092Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Examining the dimension of the data\n\nI first initialized some variables to the directory of the data we want, because generator functions that I define may use this value later.\n\nBased on the output, there are **220025 instances** of medical images labeled to either 0 or 1 in the train set, and there seems to be **no null entries.**\n\nThe test set only contains the images and not the labels, so I'll have to divide the training set for validation later.","metadata":{}},{"cell_type":"code","source":"sns.displot(data=labels_train, x='label')\n\nlabel_count = labels_train['label'].value_counts()\n\nprint(label_count[0] / label_count[1])","metadata":{"execution":{"iopub.status.busy":"2023-03-16T11:53:18.185000Z","iopub.execute_input":"2023-03-16T11:53:18.185432Z","iopub.status.idle":"2023-03-16T11:53:18.690926Z","shell.execute_reply.started":"2023-03-16T11:53:18.185391Z","shell.execute_reply":"2023-03-16T11:53:18.689959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Examining the distribution plot of the data\n\nIt can be seen that there are **more negative cases than positive cases.** The ratio of negative cases to positive cases is about **1.46.**\n\nThere's no much room for additional analysis here, let's just take a look at how these images look like.","metadata":{}},{"cell_type":"code","source":"rows, cols = 6, 6\n\nfig, axes = plt.subplots(rows, cols, figsize=(6, 6))\n\nfor i in range(6 * 6):\n    image = imread(dir_images_train + images_train['id'][i])\n\n    row, col = i // cols, i % cols\n\n    axes[row, col].imshow(image)\n    axes[row, col].axis('off')\n\nplt.subplots_adjust(wspace = 0.1, hspace = 0.3)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-16T11:53:18.692633Z","iopub.execute_input":"2023-03-16T11:53:18.693067Z","iopub.status.idle":"2023-03-16T11:53:19.673642Z","shell.execute_reply.started":"2023-03-16T11:53:18.693029Z","shell.execute_reply":"2023-03-16T11:53:19.672651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Examining the actual image\n\nLooks cool, but I can't really tell anything right out from the image since I have zero domain knowledge. \n\nBut I can see that the distribution of shapes inside the image is quite dense, so I think there shouldn't be any strides applied.\n\n## Defining model architecture\n\nI am first going to try out building a **simple VGG-style model,**, then try a modified version of the first model with **batch normalization** applied. I will compare and use the better one for submission.\n\nThe first model will be as follows:\n\n**Input -> Conv2D -> Conv2D -> MaxPool2D -> Conv2D -> Conv2D -> MaxPool2D -> Flatten -> Dense -> Dense -> Output**\n\nI will use **ReLU** as the activation function for the hidden layers, and a mix of **ReLU and sigmoid** for the dense layers before output.\n\nThere will be an additional dense layer because the dimension of the data is really large and each image is detailed, and I think the model needs it.\n\nThere will be no strides, because the image is quite detailed.\n\n## Building the first model","metadata":{}},{"cell_type":"code","source":"model = Sequential()\n\nmodel.add(Conv2D(filters=16, kernel_size=(3,3), activation='relu'))\nmodel.add(Conv2D(filters=16, kernel_size=(3,3), activation='relu'))\nmodel.add(MaxPool2D(pool_size=(2,2)))\n\nmodel.add(Conv2D(filters=32, kernel_size=(3,3), activation='relu'))\nmodel.add(Conv2D(filters=32, kernel_size=(3,3), activation='relu'))\nmodel.add(MaxPool2D(pool_size=(2,2)))\n\nmodel.add(Flatten())\nmodel.add(Dense(units=256, activation='relu'))\nmodel.add(Dense(units=1, activation='sigmoid'))\n\nbatch_size = 256\n\nmodel.build(input_shape=(batch_size, 64, 64, 3))\n\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2023-03-16T11:53:19.676675Z","iopub.execute_input":"2023-03-16T11:53:19.677565Z","iopub.status.idle":"2023-03-16T11:53:19.793259Z","shell.execute_reply.started":"2023-03-16T11:53:19.677524Z","shell.execute_reply":"2023-03-16T11:53:19.792571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Examining summary of the first model\n\nMost of the parameters came from the first dense model. I'll keep it because I think that the problem needs more than just a single unit dense layer.","metadata":{}},{"cell_type":"code","source":"os.system('pip install visualkeras')\nimport visualkeras\nvisualkeras.layered_view(model, legend=True)","metadata":{"execution":{"iopub.status.busy":"2023-03-16T11:53:19.794465Z","iopub.execute_input":"2023-03-16T11:53:19.794886Z","iopub.status.idle":"2023-03-16T11:53:29.434509Z","shell.execute_reply.started":"2023-03-16T11:53:19.794839Z","shell.execute_reply":"2023-03-16T11:53:29.433448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Examining visualization of the first model\n\nThe visualization matches the summary, and it resembles a reduced VGG model.\n\n## Preparing the training & validation data","metadata":{}},{"cell_type":"code","source":"generator = ImageDataGenerator(rescale=1./255, validation_split=0.25)\n\ndata_train = generator.flow_from_dataframe(\n    dataframe = images_train,\n    x_col='id', # filenames\n    y_col='label', # labels\n    directory=dir_images_train,\n    subset='training',\n    class_mode='binary',\n    batch_size=batch_size,\n    target_size=(64, 64))\n\ndata_validate=generator.flow_from_dataframe(\n    dataframe=images_train,\n    x_col='id', # filenames\n    y_col='label', # labels\n    directory=dir_images_train,\n    subset=\"validation\",\n    class_mode='binary',\n    batch_size=batch_size,\n    target_size=(64, 64))","metadata":{"execution":{"iopub.status.busy":"2023-03-16T11:53:29.435907Z","iopub.execute_input":"2023-03-16T11:53:29.437861Z","iopub.status.idle":"2023-03-16T12:00:52.452429Z","shell.execute_reply.started":"2023-03-16T11:53:29.437823Z","shell.execute_reply":"2023-03-16T12:00:52.451335Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Compiling & training the first model","metadata":{}},{"cell_type":"code","source":"opt = Adam(learning_rate=0.0001)\n\nmodel.compile(optimizer=opt, loss='binary_crossentropy', metrics=['accuracy'])","metadata":{"execution":{"iopub.status.busy":"2023-03-16T12:00:52.453922Z","iopub.execute_input":"2023-03-16T12:00:52.454410Z","iopub.status.idle":"2023-03-16T12:00:52.467309Z","shell.execute_reply.started":"2023-03-16T12:00:52.454352Z","shell.execute_reply":"2023-03-16T12:00:52.466205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"hist = model.fit(data_train, validation_data=data_validate, epochs=10)","metadata":{"execution":{"iopub.status.busy":"2023-03-16T12:00:52.469026Z","iopub.execute_input":"2023-03-16T12:00:52.470115Z","iopub.status.idle":"2023-03-16T12:59:17.972090Z","shell.execute_reply.started":"2023-03-16T12:00:52.470075Z","shell.execute_reply":"2023-03-16T12:59:17.971004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Evaluating the first model","metadata":{}},{"cell_type":"code","source":"plt.plot(hist.history[\"accuracy\"])\nplt.plot(hist.history['val_accuracy'])\nplt.plot(hist.history['loss'])\nplt.plot(hist.history['val_loss'])\n\nplt.title(\"Model Evaluation\")\nplt.ylabel(\"Accuracy\")\nplt.xlabel(\"Epoch\")\nplt.legend([\"Accuracy\",\"Validation Accuracy\",\"Loss\",\"Validation Loss\"])\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-16T12:59:17.974397Z","iopub.execute_input":"2023-03-16T12:59:17.974810Z","iopub.status.idle":"2023-03-16T12:59:18.214734Z","shell.execute_reply.started":"2023-03-16T12:59:17.974768Z","shell.execute_reply":"2023-03-16T12:59:18.213347Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Examining accuracy plot of of the first model\n\nThe slope is very consistent, and the accuracy converges with the validation accuracy, which is a good sign.\n\nThe model looks like a good fit overall, and the final training accuracy was **accuracy: 0.8638**, and validation accuracy was **val_accuracy: 0.8657**.\n\nThere are no significant signs of overfitting or underfitting visible in the plots.\n\n## Building the second model\n\nThis time, I will add some batch normalization layers **between the convolution and the activation function.**","metadata":{}},{"cell_type":"code","source":"model_bn = Sequential()\n\nmodel_bn.add(Conv2D(filters=16, kernel_size=(3,3)))\nmodel_bn.add(BatchNormalization())\nmodel_bn.add(Activation('relu'))\nmodel_bn.add(Conv2D(filters=16, kernel_size=(3,3)))\nmodel_bn.add(BatchNormalization())\nmodel_bn.add(Activation('relu'))\nmodel_bn.add(MaxPool2D(pool_size=(2,2)))\n\nmodel_bn.add(Conv2D(filters=32, kernel_size=(3,3)))\nmodel_bn.add(BatchNormalization())\nmodel_bn.add(Activation('relu'))\nmodel_bn.add(Conv2D(filters=32, kernel_size=(3,3)))\nmodel_bn.add(BatchNormalization())\nmodel_bn.add(Activation('relu'))\nmodel_bn.add(MaxPool2D(pool_size=(2,2)))\n\nmodel_bn.add(Flatten())\nmodel_bn.add(Dense(units=256, activation='relu'))\nmodel_bn.add(Dense(units=1, activation='sigmoid'))\n\nmodel_bn.build(input_shape=(batch_size, 64, 64, 3))\n\nmodel_bn.summary()","metadata":{"execution":{"iopub.status.busy":"2023-03-16T13:17:18.310607Z","iopub.execute_input":"2023-03-16T13:17:18.311542Z","iopub.status.idle":"2023-03-16T13:17:34.943734Z","shell.execute_reply.started":"2023-03-16T13:17:18.311491Z","shell.execute_reply":"2023-03-16T13:17:34.943039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Examining summary of the second model\n\nI can see that batch normalization layers are now added between the convolution layers and the activation functions, and the number of parameters increased a bit due to it.","metadata":{}},{"cell_type":"code","source":"visualkeras.layered_view(model_bn, legend=True)","metadata":{"execution":{"iopub.status.busy":"2023-03-16T12:59:18.413452Z","iopub.execute_input":"2023-03-16T12:59:18.413832Z","iopub.status.idle":"2023-03-16T12:59:18.481865Z","shell.execute_reply.started":"2023-03-16T12:59:18.413793Z","shell.execute_reply":"2023-03-16T12:59:18.480683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Examining visualization of the second model\n\nThe visualization matches the summary, and now it looks a bit different from than a traditional VGG model, having additional layers for batch normalization between hidden layers.\n\n## Compiling & training the second model","metadata":{}},{"cell_type":"code","source":"opt_bn = Adam(learning_rate=0.0001)\n\nmodel_bn.compile(optimizer=opt_bn, loss='binary_crossentropy', metrics=['accuracy'])","metadata":{"execution":{"iopub.status.busy":"2023-03-16T13:17:46.596660Z","iopub.execute_input":"2023-03-16T13:17:46.597154Z","iopub.status.idle":"2023-03-16T13:17:46.614023Z","shell.execute_reply.started":"2023-03-16T13:17:46.597111Z","shell.execute_reply":"2023-03-16T13:17:46.612462Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"hist_bn = model_bn.fit(data_train, validation_data=data_validate, epochs=10)","metadata":{"execution":{"iopub.status.busy":"2023-03-16T13:18:11.906507Z","iopub.execute_input":"2023-03-16T13:18:11.907519Z","iopub.status.idle":"2023-03-16T14:07:24.024569Z","shell.execute_reply.started":"2023-03-16T13:18:11.907476Z","shell.execute_reply":"2023-03-16T14:07:24.023426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Evaluating the second model","metadata":{}},{"cell_type":"code","source":"plt.plot(hist_bn.history[\"accuracy\"])\nplt.plot(hist_bn.history['val_accuracy'])\nplt.plot(hist_bn.history['loss'])\nplt.plot(hist_bn.history['val_loss'])\n\nplt.title(\"Model Evaluation\")\nplt.ylabel(\"Accuracy\")\nplt.xlabel(\"Epoch\")\nplt.legend([\"Accuracy\",\"Validation Accuracy\",\"Loss\",\"Validation Loss\"])\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-16T14:08:54.662897Z","iopub.execute_input":"2023-03-16T14:08:54.663347Z","iopub.status.idle":"2023-03-16T14:08:54.947918Z","shell.execute_reply.started":"2023-03-16T14:08:54.663303Z","shell.execute_reply":"2023-03-16T14:08:54.946751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Examining accuracy plot of of the second  model\n\nThe slope is quite consistent for the accuracy, and the accuracy achieved numbers higher than 0.9, but the slope for the validation has became more inconsistent. \n\nThe difference in average slope and the final absolute value of the difference itself is not that significant, so I can say that there's no sign of overfitting / underfitting of the data, but since the validation lines oscillated a bit, it can be inferred that the effects of the batch normalization could have biased the model's fit more towards the training data.\n\nTo sum up, the second model also looks like a good fit overall, and the final training accuracy was **accuracy: 0.9049**, and validation accuracy was **val_accuracy: 0.8889**\n\nBased on the plot and the values, I can conclude that **the second model with batch normalization performs much better** than the first one.\n\n## Conclusion\n\nA simple **VGG-style model** that consists of multiple combinations of **a pair of convolution layer and a max-pooling layer** was shown to work well for this specific type of classification problem, and applying **batch normalization** between each convolution layer and activation function not only improved the **accuracy** of the first epoch significantly, it improved the final results and even decreased the **loss** even further.\n\nBut with the batch normalization, the plots for the validation accuracy & loss became more inconsistent, implying that the additional layers could have made the model biased towards the training data.\n\nI think **adding additional dense layer** just before the last dense layer helped achieveing these results, by letting the model to learn more details from such a detailed image. I limited the number of dense layers to two due to limited computational resources that I had, but may try more layers with larger number of units in the future, with more epochs also.\n\n## Submission","metadata":{}},{"cell_type":"code","source":"images_test = pd.DataFrame({'id':os.listdir(dir_images_test)})\n\ngenerator_test = ImageDataGenerator(rescale=1./255)\n\ndata_test = generator_test.flow_from_dataframe(\n    dataframe = images_test,\n    x_col='id', # filenames\n    directory=dir_images_test,\n    class_mode=None,\n    batch_size=1,\n    target_size=(64, 64),\n    shuffle=False)","metadata":{"execution":{"iopub.status.busy":"2023-03-16T14:17:03.808045Z","iopub.execute_input":"2023-03-16T14:17:03.808441Z","iopub.status.idle":"2023-03-16T14:18:16.770230Z","shell.execute_reply.started":"2023-03-16T14:17:03.808399Z","shell.execute_reply":"2023-03-16T14:18:16.769133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = model_bn.predict(data_test, verbose=1)","metadata":{"execution":{"iopub.status.busy":"2023-03-16T14:18:20.114089Z","iopub.execute_input":"2023-03-16T14:18:20.114497Z","iopub.status.idle":"2023-03-16T14:24:42.088122Z","shell.execute_reply.started":"2023-03-16T14:18:20.114460Z","shell.execute_reply":"2023-03-16T14:24:42.087045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(predictions)","metadata":{"execution":{"iopub.status.busy":"2023-03-16T14:24:44.732530Z","iopub.execute_input":"2023-03-16T14:24:44.732930Z","iopub.status.idle":"2023-03-16T14:24:44.745301Z","shell.execute_reply.started":"2023-03-16T14:24:44.732889Z","shell.execute_reply":"2023-03-16T14:24:44.744009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred = np.transpose(predictions)[0]\n\nprint(pred)\n\nsubmission_df = pd.DataFrame()\nsubmission_df['id'] = images_test['id'].apply(lambda x: x.split('.')[0])\nsubmission_df['label'] = list(map(lambda x: 0 if x < 0.5 else 1, pred))\n\nprint(submission_df.head())\n\nsubmission_df['label'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-03-16T14:25:55.435683Z","iopub.execute_input":"2023-03-16T14:25:55.436054Z","iopub.status.idle":"2023-03-16T14:25:55.608011Z","shell.execute_reply.started":"2023-03-16T14:25:55.436021Z","shell.execute_reply":"2023-03-16T14:25:55.606909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-03-16T14:26:06.082294Z","iopub.execute_input":"2023-03-16T14:26:06.083015Z","iopub.status.idle":"2023-03-16T14:26:06.167033Z","shell.execute_reply.started":"2023-03-16T14:26:06.082977Z","shell.execute_reply":"2023-03-16T14:26:06.165953Z"},"trusted":true},"execution_count":null,"outputs":[]}]}