{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# CNN Histopathologic Cancer Detection ","metadata":{}},{"cell_type":"markdown","source":"## Overview\nThis notebook is a practice of utilizing the TensorFlow and Keras to build a Convolutional Neural Network (CNN) for predicting cancer detection from digital pathology images. The practice is based on a Kaggle competition, and the data can be obtained from the competition website at https://www.kaggle.com/competitions/histopathologic-cancer-detection/data.\n\nThis notebook can also be found at https://github.com/Lorby04/msds/tree/main/dl/week3","metadata":{}},{"cell_type":"markdown","source":"1. Data preparing\n\nDownloading data from the source, extract the compressed files to local disk.\nThe original image is in tiff format which is not supported by tensorflow. Although there's other packages can handle the format, but it requires mandy dependecies which makes the image loading tasks being complex. To simplify the procedure, finally I converted the tiff image to jpeg format with the package Pillow. \nAfter the images are prepared, I'm planning to use the Keras API \"image_dataset_from_directory()\" to create tensorflow dataset from the directory. The API requires the directories to be organized hierarchically with subdirectory to be the name of the class_name or label. So, the hierarchical directory is created along with the downloading/converting/saving procedure as the implemention in the below cells.","metadata":{}},{"cell_type":"code","source":"import pathlib\nimport os\nimport pandas as pd\nfrom PIL import Image","metadata":{"id":"nvbPWI6RmAtO"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Subdir'0' mapps to the label 0 and subdir '1' mapps to the label 1","metadata":{}},{"cell_type":"code","source":"train_csv = \"./histopathologic-cancer-detection/train_labels.csv\"\n\norigin_train_dir = \"./histopathologic-cancer-detection/train\"\norigin_test_dir = \"./histopathologic-cancer-detection/test\"\n\ndata_dir = \"./hpcd\"\n\ntrain_dir = data_dir + \"/train\"\ntest_dir = data_dir + \"/test\"\n\ndir_true = train_dir + \"/1\"\ndir_false = train_dir + \"/0\"\n\norigin_train_path = pathlib.Path(origin_train_dir).with_suffix('')\norigin_test_path = pathlib.Path(origin_test_dir).with_suffix('')\n\ntrain_path = pathlib.Path(train_dir).with_suffix('')\ntest_path = pathlib.Path(test_dir).with_suffix('')\n\ntrain_df = None","metadata":{"id":"4yXxxTHLm-hu"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def download_data():\n  data_dir = 'histopathologic-cancer-detection' \n  dataset_url =  \"https://www.kaggle.com/competitions/histopathologic-cancer-detection/data\"\n  cmd = \"pip install opendatasets\"\n  os.system(cmd)\n  import opendatasets as od\n  od.download(dataset_url)","metadata":{"id":"ejJXtlphCGbB"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Convert the image from tif to jpg\n#Move train data to subclass directory\ndef new_dir(directory):\n  cmd = \"mkdir \" + directory\n  os.system(cmd)\n\ndef organize_data():\n  new_dir(data_dir)\n  new_dir(train_dir)\n  new_dir(test_dir)\n  new_dir(dir_true)\n  new_dir(dir_false)\n\n  train_df = pd.read_csv(train_csv)\n\n  #convert the tif image in train to jpg and save to new directory\n  for row in train_df.itertuples(index = False):\n    tif = origin_train_dir+ \"/\" +row.id + \".tif\"\n    jpg = \"/\" + row.id + \".jpg\" \n    if row.label == 0:\n      jpg = dir_false + jpg\n    else:\n      jpg = dir_true + jpg\n    \n    img = Image.open(tif)\n    img.save(jpg)\n\n  #convert the tif image in test to jpg and save to new directory\n    test_images = list(origin_test_path.glob('*.tif'))\n    for f in test_images:\n      img = Image.open(f)\n      f_name = f.name.replace(\".tif\", \".jpg\")\n      nf = test_dir + \"/\" + f_name\n      img.save(nf)\n\n  ","metadata":{"id":"jgAq5haVk6Ux"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Only perform the data preparing for the first time","metadata":{}},{"cell_type":"code","source":"if not os.path.exists(data_dir):\n  download_data()\n  organize_data()\n","metadata":{"id":"52URkUcO5o7B","outputId":"50b6849a-ae50-4bc9-a9d0-9531b51113d1"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#!pip uninstall PIL\n#!python3 -m pip install --upgrade pip\n#!python3 -m pip install --upgrade Pillow","metadata":{"id":"m82Yri1j73t7"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Exploratory Data Analysis (EDA)\nThe original data includes two sets, one for traning, the other one for testing. The preparing procedure put the training data to the train_dir/{class_name}, the testing data is put to the test_dir.\nIn the following sections, we will have initial analysis and visualization of the data.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport os\n\nimport tensorflow as tf\nimport tensorflow_datasets as tfds\n","metadata":{"id":"MNyeXEpe63PH"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The number of images for tranning is 220,025, and also verify the name is as expected.","metadata":{}},{"cell_type":"code","source":"images = list(train_path.glob('*/*.jpg'))\nimage_count = len(images)\nprint(\"Number of images:\",image_count)\nprint(\"First image:\", images[0])","metadata":{"id":"3cB1Usex4q_n","outputId":"8c0887dc-b383-4559-c611-12c81844cbdc"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Cross verify the data by checking the tain_csv file provided along with the data.\nThe csv file includes the file id (filename excluding the extention) and the label.","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(train_csv)\ntrain_df.info()\ntrain_df.head()","metadata":{"id":"4vIf90g6LTmd"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Have a glance of the statistics information of the training data from the csv file. From the data, 60% of the training images are not detected with cancer (marked as \"Negative\" in the pie chart), and 40% are detected with cancer (marked as \"Positive\")","metadata":{}},{"cell_type":"code","source":"train_df['label'].value_counts()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n%matplotlib inline\n\nlabels = ['Negative', 'Positive']\nplt.pie(train_df['label'].value_counts(),autopct='%1.1f%%',labels=labels)\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Cross check the raw data with the files in the traning directory. The data matches with the csv file.","metadata":{}},{"cell_type":"code","source":"images_true = list((pathlib.Path(dir_true).with_suffix('')).glob('*.jpg'))\nimage_count_true = len(images_true)\nprint(\"Number of images with cancer being detected:\",image_count_true)\n\nimages_false = list((pathlib.Path(dir_false).with_suffix('')).glob('*.jpg'))\nimage_count_false = len(images_false)\nprint(\"Number of images with cancer not being detected:\",image_count_false)","metadata":{"id":"faWo8MNkolDX"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#from sklearn.model_selection import train_test_split\n#X_train, X_verify, y_train, y_verify = train_test_split(X, y, random_state=13, train_size=0.8)\n#print(len(X_train), len(X_verify), len(y_train), len(y_verify))","metadata":{"id":"Wm27AFnVLVoP"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Take some of the training images to be shown on the screen, the first set are positive images and the second set are negative images. ","metadata":{}},{"cell_type":"code","source":"import PIL\nimport PIL.TiffImagePlugin\nfrom PIL import Image","metadata":{"id":"kuG26ZJXqTu9"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig_width = 5\nfig_height = 5\nfig,ax = plt.subplots(fig_height,fig_width)\n\nprint(\"Sample images with cancer being detected:\")\nfor i in range(fig_width * fig_height):\n    with open(images_true[i], \"rb\") as f:\n        img = Image.open(f)\n        ax[i%fig_height][i//fig_width].imshow(img)","metadata":{"id":"INBJ2VGGqUv_","outputId":"9d6ed49c-0583-4a2a-de34-ea35f00280ff"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig_width = 5\nfig_height = 5\nfig,ax = plt.subplots(fig_height,fig_width)\n        \nprint(\"Sample images with cancer not being detected:\")\nfor i in range(fig_width * fig_height):\n    with open(images_false[i], \"rb\") as f:\n        img = Image.open(f)\n        ax[i%fig_height][i//fig_width].imshow(img)","metadata":{"id":"LOtaJvJIqdK5","outputId":"78caebaf-9f1e-44ee-a1a0-301d98befd30"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To process the mimages, we need to know the size attibute of the images, it's (96,96) as shown below.","metadata":{}},{"cell_type":"code","source":"print(Image.open(images[0]).size)","metadata":{"id":"mxsj0JgVRXU_","outputId":"c0cd9ee2-d1b8-429b-9341-25a4452c703e"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Generate then load the training image dataset with image_dataset_from_directory API from the directory. Split 20% from the dataset to be used for validation.","metadata":{}},{"cell_type":"markdown","source":"Confirm all training files are in place again.","metadata":{}},{"cell_type":"code","source":"print(\"image_count:\",len(images), image_count)\nlist_ds = tf.data.Dataset.list_files(str(train_path/'*/*'), shuffle=True)\nlist_ds = list_ds.shuffle(image_count, reshuffle_each_iteration=True)","metadata":{"id":"hb9UTOS8czpm"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for f in list_ds.take(5):\n    print(f.numpy())","metadata":{"id":"8bBoATXrqorx"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels = train_df['label'].unique()\nprint(labels)","metadata":{"id":"yNHeojdoqrRA"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img_size = Image.open(images[0]).size\nclass_names = ['0', '1']\nbatch_size = 32\norig_train_ds = tf.keras.utils.image_dataset_from_directory(\n  train_dir,\n  class_names = class_names,\n  validation_split=0.2,\n  subset=\"training\",\n  seed=123,\n  image_size=img_size,\n  batch_size=batch_size)\n","metadata":{"id":"pDY7UtSMqxTN"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"orig_val_ds = tf.keras.utils.image_dataset_from_directory(\n  train_dir,\n  class_names = class_names,\n  validation_split=0.2,\n  subset=\"validation\",\n  seed=123,\n  image_size=img_size,\n  batch_size=batch_size)","metadata":{"id":"70ZmCIzFwWyA"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Check the type of the dataset and the information of elements in the dataset.","metadata":{}},{"cell_type":"code","source":"print(orig_train_ds, \"\\n\", orig_val_ds)","metadata":{"id":"9EHRrwPT4mWB","outputId":"81168628-268c-4bba-f88f-fbf222b9d814"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It indicates that each batch includes 32 images, among which, each image is 96*96 colored with RGB","metadata":{}},{"cell_type":"code","source":"for image_batch, labels_batch in orig_train_ds:\n  print(image_batch.shape)\n  print(labels_batch.shape)\n  break","metadata":{"id":"B_uPx0F5Hg_Y"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## NN Model Architecture \nThe high level architecture of the training network are referring to the tensorflow training material. On top of the referrd architecture, in this practice, several different modles with changing the hyper parameters such as different activation function and different loss funtion are tried to find the best one.\nTo have better performance, preloaded mecahnism is included.","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras import layers","metadata":{"id":"ktmNH_UsEtJB"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Preload the dataset to cache to accelarate the training.","metadata":{}},{"cell_type":"code","source":"AUTOTUNE = tf.data.AUTOTUNE\n\ntrain_ds = orig_train_ds.cache().prefetch(buffer_size=AUTOTUNE)\nval_ds = orig_val_ds.cache().prefetch(buffer_size=AUTOTUNE)","metadata":{"id":"eKPZVYVYH7nj"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Try the model with different hyper parameters to get the best one, to reduce the time, the epochs for each parameters set is set to 5, it may not nbe very accurate.\nGPU is enabled.","metadata":{}},{"cell_type":"code","source":"hyper_params = [\n  {\n    'num_classes':2,\n    'activation':'relu',\n    'loss':tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True)\n  },\n  {\n    'num_classes':2,\n    'activation':'sigmoid',\n    'loss':tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True)\n  },\n  {\n    'num_classes':2,\n    'activation':'softmax',\n    'loss':tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True)\n  },\n  {\n    'num_classes':2,\n    'activation':'relu',\n    'loss':tf.keras.losses.MeanSquaredLogarithmicError()\n  },\n  {\n    'num_classes':2,\n    'activation':'sigmoid',\n    'loss':tf.keras.losses.MeanSquaredLogarithmicError()\n  },\n  {\n    'num_classes':2,\n    'activation':'softmax',\n    'loss':tf.keras.losses.MeanSquaredLogarithmicError()\n  }\n]","metadata":{"id":"FrmuBFx9Ifwb"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras.optimizers.legacy import Adam\nopt = Adam()\n\nbest_model = None\nbest_param = None\nbest_result = 0\nfor p in hyper_params:\n    print(\"Model params:\",p)\n    model = tf.keras.Sequential([\n        tf.keras.layers.Rescaling(1./255, input_shape=(img_size[0],img_size[1],3)),\n        tf.keras.layers.RandomFlip(\"horizontal_and_vertical\"),\n        #tf.keras.layers.RandomRotation(0.2),\n        tf.keras.layers.Conv2D(64, 3, activation=p['activation']),\n        tf.keras.layers.MaxPooling2D(),\n        tf.keras.layers.Conv2D(64, 3, activation=p['activation']),\n        tf.keras.layers.MaxPooling2D(),\n        tf.keras.layers.Conv2D(64, 3, activation=p['activation']),\n        tf.keras.layers.MaxPooling2D(),\n        tf.keras.layers.Flatten(),\n        tf.keras.layers.Dense(128, activation=p['activation']),\n        tf.keras.layers.Dense(p['num_classes'])\n    ])\n\n    model.compile(\n      opt,\n      loss=p['loss'],\n      metrics=['accuracy'])\n    \n    model.fit(\n      train_ds,\n      validation_data=val_ds,\n      epochs=5\n    )\n    \n    metrics = model.get_metrics_result()\n    acc = metrics['accuracy'].numpy() #- metrics['loss'].numpy()\n    if acc > best_result:\n        best_model = model\n        best_param = p\n        best_result = acc","metadata":{"id":"IiH8poPZIG_G"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Summary of the best model and the corresponding hyper-parameters.","metadata":{}},{"cell_type":"code","source":"model = best_model\nmodel.summary()\nmodel.get_metrics_result()\n","metadata":{"id":"ZSIAXLQXIaBI"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Parameter:\", best_param)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To get better result, retrain the best model with more epochs(15) and early-stop callbacks.  The early-stop callbacks will cause the training to be stopped when the val_accuracy decreases in two continous rounds. With the early-stop mechanism, it can avoid significant overfitting problem.\nTo prevent overfitting further, the L2 regularity and drop-out were added initially, but the final result was very bad, so the two groups are removed from the final model.\n\nThe model includes the following layers:\n1. Preprocessing (Rescale) and augument (Flip) layer\n2. Output layer\n3. 3 hidden CNN layer with 64*3 nodes per layer and 3 hidden pooling layer\n4. Optimization algorithm: Adam\n5. Activation function:relu\n6. Other parameters: default value","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras import regularizers\np = best_param\nepochs = 15\nmodel = tf.keras.Sequential([\n    tf.keras.layers.Rescaling(1./255, input_shape=(img_size[0],img_size[1],3)),\n    tf.keras.layers.RandomFlip(\"horizontal_and_vertical\"),\n    #tf.keras.layers.RandomRotation(0.2),\n    tf.keras.layers.Conv2D(64, 3, activation=p['activation']),\n    tf.keras.layers.MaxPooling2D(),\n    tf.keras.layers.Conv2D(64, 3, activation=p['activation']),\n    tf.keras.layers.MaxPooling2D(),\n    tf.keras.layers.Conv2D(64, 3, activation=p['activation']),\n    tf.keras.layers.MaxPooling2D(),\n    \n\n    tf.keras.layers.Flatten(),\n\n    #Add L2 regulation and dropout layer to avoid overfitting\n    #layers.Dense(128, kernel_regularizer=regularizers.l2(0.0001),\n    #             activation='elu'),\n    \n    #layers.Dense(128, kernel_regularizer=regularizers.l2(0.0001),\n    #             activation='elu'),\n    #layers.Dropout(0.5),\n    #layers.Dense(128, kernel_regularizer=regularizers.l2(0.0001),\n    #             activation='elu'),\n    #layers.Dropout(0.5),\n    #layers.Dense(128, kernel_regularizer=regularizers.l2(0.0001),\n    #             activation='elu'),\n    #layers.Dropout(0.5),\n    \n    \n    tf.keras.layers.Dense(128, activation=p['activation']),\n    #tf.keras.layers.Dropout(0.3),\n    tf.keras.layers.Dense(p['num_classes'])\n])\n\nmodel.compile(\n    opt,\n    loss=p['loss'],\n    metrics=['accuracy']\n)\n\nmodel.fit(\n    train_ds,\n    validation_data=val_ds,\n    epochs=epochs,\n    callbacks=[tf.keras.callbacks.EarlyStopping(monitor='val_accuracy', patience=2)]\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The training is stopped earlier when epoch=5. Check the history information of the training.","metadata":{}},{"cell_type":"code","source":"history = model.history.history\nprint(history.keys())\nhistory","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#The plot function referrs to the code from https://machinelearningmastery.com/display-deep-learning-model-training-history-in-keras/\ndef plot_accuracy(history):\n    plt.plot(history['accuracy'],'*-')\n    plt.plot(history['val_accuracy'],\"x-\")\n    plt.title('model accuracy')\n    plt.ylabel('accuracy')\n    plt.xlabel('epoch')\n    plt.legend(['train', 'validation'], loc='upper left')\n    plt.grid()\n    plt.show()\n# summarize history for loss\ndef plot_loss(history):\n    plt.plot(history['loss'],'*-')\n    plt.plot(history['val_loss'],'x-')\n    plt.title('model loss')\n    plt.ylabel('loss')\n    plt.xlabel('epoch')\n    plt.legend(['train', 'validation'], loc='upper left')\n    plt.grid()\n    plt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_accuracy(history)\nplot_loss(history)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Refer to the tutorial, to make the probability more obvious, a softmax layer is added as the latest layer of the model.","metadata":{}},{"cell_type":"markdown","source":"## Test\nBefore testing, a softmax layer is added to the model to make the probability more obvious.\nThe test images are from the previously prepared test directory.\nTo make the result clearer, a pandas dataframe is created for the record.","metadata":{}},{"cell_type":"markdown","source":"Add softmax layer","metadata":{}},{"cell_type":"code","source":"model.add(tf.keras.layers.Softmax())","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Load data and create the dataframe respectively.","metadata":{}},{"cell_type":"code","source":"test_files = np.array(os.listdir(test_dir))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df = pd.DataFrame(test_files, columns=['id'])\nprint(test_df)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_datagen = tf.keras.preprocessing.image.ImageDataGenerator() #rescale is embedded in the model, no pre-rescaling\ntest_ds = test_datagen.flow_from_dataframe(\n    test_df,\n    test_dir,\n    class_mode=None,\n    shuffle= False,\n    x_col = 'id',\n    y_col = None,\n    target_size = img_size\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Predict. The result of the prediction is the probability of each class_name per record. The probability is mapped to the class_names then put to the dataframe at the end. ","metadata":{}},{"cell_type":"code","source":"pred = model.predict(test_ds)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(pred.shape)\nprint(pred)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df = test_df.applymap(lambda x: os.path.splitext(x)[0])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred = pd.DataFrame(pred)\ntest_df['label'] = pred.apply(np.argmax, axis=1)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.to_csv('submission.csv',index=False)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df['label'].value_counts()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels = ['Negative', 'Positive']\nplt.pie(test_df['label'].value_counts(),autopct='%1.1f%%',labels=labels)\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.pie(train_df['label'].value_counts(),autopct='%1.1f%%',labels=labels)\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Conclusion and Analysis\nFrom the statistics information of training data, the percentage of negative and positive are 60% and 40% respectively. The statistics of the predicted testing data indicates the percentage of negative and positive are 66% and 34% respectively.\nAs a rough estimation, there's 6% gap between the distribution of the prediction and the real distribution, which means the accuracy of the prediction(testing) is less than 90% (1-6/60), considering there are also some false positive and false negative, the accuracy of the prediction could be around 75% to 80%.\nThere are some improvements can be applied to get better result.\n1. Use tiff format directly.\n2. Fine tune with trying more different hyper-parameters \n3. Add rotation and other more image augumentation (it does not work in my environment however)\n4. Enhance the network","metadata":{}}]}