{"cells":[{"metadata":{},"cell_type":"markdown","source":"# **Histopathologic Cancer Detection**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Cancer Detection project, By BEKKAR Abdellatif\n# Aim is to familiarity with the use of CNNs with tensorflow for image classification","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**By BEKKAR Abdellatif**","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**About the Data Set**\n\nhe data for this kernel is a slightly modified version of the PatchCamelyon (PCam) benchmark dataset. The original PCam dataset contains duplicate images due to its probabilistic sampling, however, the version presented on Kaggle does not contain duplicates.\n\nThe PatchCamelyon benchmark is a new and challenging image classification dataset. It consists of 327.680 color images (96 x 96px) extracted from histopathologic scans of lymph node sections. Each image is annoted with a binary label indicating presence of metastatic tissue. PCam provides a new benchmark for machine learning models: bigger than CIFAR10, smaller than imagenet, trainable on a single GPU.\n\nPCam packs the clinically-relevant task of metastasis detection into a straight-forward binary image classification task, akin to CIFAR-10 and MNIST. Models can easily be trained on a single GPU in a couple hours, and achieve competitive scores in the Camelyon16 tasks of tumor detection and whole-slide image diagnosis. Furthermore, the balance between task-difficulty and tractability makes it a prime suspect for fundamental machine learning research on topics as active learning, model uncertainty, and explainability.\n\n**The images are labeled as 0 or 1, where 0 = No Tumor Tissue and 1 = Has Tumor Tissue(s)**","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nimport matplotlib.pyplot as plt\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.\n\nfrom time import time\nimport seaborn as sns\nimport plotly.graph_objects as go\n\nfrom sklearn.model_selection import train_test_split\nfrom keras.preprocessing.image import ImageDataGenerator\n\n\n# import the necessary packages\nfrom keras.models import Sequential\nfrom keras.layers.convolutional import Conv2D\nfrom keras.layers.convolutional import MaxPooling2D\nfrom keras.layers.core import Activation\nfrom keras.layers.core import Flatten\nfrom keras.layers.core import Dropout\nfrom keras.layers.core import Dense\nfrom keras.optimizers import Adam\nfrom keras import backend as K\nfrom keras.callbacks import EarlyStopping, ReduceLROnPlateau","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Exploratory Data Analysis","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"***Total Samples Available***","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#Total Samples Available\nprint('Train Images = ',len(os.listdir('../input/histopathologic-cancer-detection/train')))\nprint('Test Images = ',len(os.listdir('../input/histopathologic-cancer-detection/test')))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"***Create a DataFrame of all Train Image Labels***","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df = pd.read_csv('../input/histopathologic-cancer-detection/train_labels.csv',dtype=str)\nprint(df.head())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Number of image : ', len(df))\nimg = plt.imread(\"../input/histopathologic-cancer-detection/train/\"+df.iloc[0]['id']+'.tif')\nprint('Images shape', img.shape)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"***Visualize some Train Images***","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"for i in range(5):\n    img = plt.imread(\"../input/histopathologic-cancer-detection/train/\"+df.iloc[i]['id']+'.tif')\n    print(df.iloc[i]['label'])\n    plt.imshow(img)\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"***See the distribution of Train Labels***","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Descriptive Analytics for given Dataset\n\nprint(df.label.value_counts())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize = (6,6)) \nax = sns.countplot(df.label).set_title('Label Counts', fontsize = 18)\nplt.annotate(df.label.value_counts()[0],\n            xy = (0,df.label.value_counts()[0] + 2000),\n            va = 'bottom',\n            ha = 'center',\n            fontsize = 12)\nplt.annotate(df.label.value_counts()[1],\n            xy = (1,df.label.value_counts()[1] + 2000),\n            va = 'bottom',\n            ha = 'center',\n            fontsize = 12)\nplt.ylim(0,150000)\nplt.ylabel('Count', fontsize = 16)\nplt.xlabel('Labels', fontsize = 16)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"labels = [\"No Cancer - 0\", \"Cancer - 1\"]\nvalues = df.label.value_counts()\n\nd = go.Figure(data=[go.Pie(labels=labels, values=values, hole=.5, marker_colors=[\"rgb(0, 76, 153)\",\"rgb(255, 158, 60)\"])])\nd.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Here the Label-1 is 59,5% and Label-0 is 40,5% of the whole train images. There is a little imbalance here which we can rectify to get better performance.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Feature Engineering\n\nSplit into Train and Validation Sets with Keras ImageGenerator","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#add .tif to ids in the dataframe to use flow_from_dataframe\ndf[\"id\"]=df[\"id\"].apply(lambda x : x +\".tif\")\ndf.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_path = '../input/histopathologic-cancer-detection/train'\nvalid_path = '../input/histopathologic-cancer-detection/train'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_datagen = ImageDataGenerator(validation_split=0.20,\n                          rescale=1/255.0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_generator=train_datagen.flow_from_dataframe(\n    dataframe=df,\n    directory=train_path,\n    x_col=\"id\",\n    y_col=\"label\",\n    subset=\"training\",\n    batch_size=64,\n    shuffle=True,\n    class_mode=\"binary\",\n    target_size=(96,96))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"valid_generator=train_datagen.flow_from_dataframe(\n    dataframe=df,\n    directory=valid_path,\n    x_col=\"id\",\n    y_col=\"label\",\n    subset=\"validation\",\n    batch_size=64,\n    shuffle=True,\n    class_mode=\"binary\",\n    target_size=(96,96))\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Define the model","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Model structure (optimizer: Adam):**\n* In\n* Conv2D(32)*3 -> Dropout (0.3) -> MaxPool2D (3)\n* Conv2D(64)*3 -> Dropout (0.3) -> MaxPool2D (3)\n* Conv2D(128)*3 -> Dropout (0.3) -> MaxPool2D (3)\n* Flatten\n* Dense (128)\n* Dropout\n* Out\n","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"model = Sequential()\n\nmodel.add(Conv2D(filters = 32, kernel_size = (3,3), padding = 'same', activation = 'relu', input_shape = (96, 96, 3)))\nmodel.add(Conv2D(filters = 32, kernel_size = (3,3), padding = 'same', activation = 'relu'))\nmodel.add(Conv2D(filters = 32, kernel_size = (3,3), padding = 'same', activation = 'relu'))\nmodel.add(Dropout(0.2))\nmodel.add(MaxPooling2D(pool_size=(3,3)))\n\nmodel.add(Conv2D(filters = 64, kernel_size = (3,3), padding = 'same', activation = 'relu'))\nmodel.add(Conv2D(filters = 64, kernel_size = (3,3), padding = 'same', activation = 'relu'))\nmodel.add(Conv2D(filters = 64, kernel_size = (3,3), padding = 'same', activation = 'relu'))\nmodel.add(Dropout(0.3))\nmodel.add(MaxPooling2D(pool_size=(3,3)))\n\nmodel.add(Conv2D(filters = 128, kernel_size = (3,3), padding = 'same', activation = 'relu'))\nmodel.add(Conv2D(filters = 128, kernel_size = (3,3), padding = 'same', activation = 'relu'))\nmodel.add(Conv2D(filters = 128, kernel_size = (3,3), padding = 'same', activation = 'relu'))\nmodel.add(Dropout(0.5))\nmodel.add(MaxPooling2D(pool_size=(3,3)))\n\nmodel.add(Flatten())\nmodel.add(Dense(128, activation='relu'))\nmodel.add(Dropout(0.5))\nmodel.add(Dense(1, activation='sigmoid'))\nmodel.summary()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model.compile(optimizer='adam' , loss='binary_crossentropy', metrics=['accuracy'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"STEP_SIZE_TRAIN=train_generator.n//train_generator.batch_size\nSTEP_SIZE_VALID=valid_generator.n//valid_generator.batch_size\nEPOCHS=20","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\nearlystopper = EarlyStopping(monitor='val_accuracy', patience=3, verbose=1, restore_best_weights=True)\nreducel = ReduceLROnPlateau(monitor='val_accuracy', patience=2, verbose=1, factor=0.1)\n\n\nhistory = model.fit_generator(generator=train_generator, \n                    steps_per_epoch=STEP_SIZE_TRAIN, \n                    validation_data=valid_generator,\n                    validation_steps=STEP_SIZE_VALID,\n                    epochs=EPOCHS,\n                   callbacks=[reducel, earlystopper])\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# CNN Model Evaluation","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def show_final_history(history):\n    fig, ax = plt.subplots(1, 2, figsize=(15,5))\n    ax[0].set_title('loss')\n    ax[0].plot(history.epoch, history.history[\"loss\"], label=\"Train loss\")\n    ax[0].plot(history.epoch, history.history[\"val_loss\"], label=\"Validation loss\")\n    ax[1].set_title('accuracy')\n    ax[1].plot(history.epoch, history.history[\"accuracy\"], label=\"Train acc\")\n    ax[1].plot(history.epoch, history.history[\"val_accuracy\"], label=\"Validation acc\")\n    ax[0].legend()\n    ax[1].legend()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"show_final_history(history)\nprint(\"Validation Accuracy: \" + str(history.history['val_accuracy'][-1:]))","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}