{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Intro\nWelcome to the [Human Protein Atlas - Single Cell Classification](https://www.kaggle.com/c/hpa-single-cell-image-classification).\n\n![](https://storage.googleapis.com/kaggle-competitions/kaggle/23823/logos/header.png)\n\nFor a TPU tutorial of this compedition we recommend this [notebook](https://www.kaggle.com/drcapa/human-protein-atlas-tpu-tutorial/).\n\n<span style=\"color: royalblue;\">Please vote the notebook up if it helps you. Feel free to leave a comment above the notebook. Thank you. </span>"},{"metadata":{},"cell_type":"markdown","source":"# Libraries"},{"metadata":{"trusted":true},"cell_type":"code","source":"import os\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport matplotlib\nimport cv2\n\nfrom sklearn.model_selection import train_test_split\n\nimport tensorflow as tf\nfrom keras.utils import to_categorical, Sequence\nfrom keras.models import Sequential\nfrom keras.layers import Dense, Dropout, Flatten, Conv2D, MaxPool2D, BatchNormalization\nfrom keras.optimizers import RMSprop,Adam\nfrom keras.applications import ResNet50","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Path"},{"metadata":{"trusted":true},"cell_type":"code","source":"path = '/kaggle/input/hpa-single-cell-image-classification/'\nos.listdir(path)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Load Data"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_data = pd.read_csv(path+'train.csv')\nsamp_subm = pd.read_csv(path+'sample_submission.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Parameter"},{"metadata":{"trusted":true},"cell_type":"code","source":"img_size = 64\nimg_channel = 3\nnum_classes = 19\nbatch_size = 64","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Overview"},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Number train samples:', len(train_data.index))\nprint('Number train images:', len(os.listdir(path+'train')))\nprint('Number submission samples:', len(samp_subm.index))\nprint('Number submission images:', len(os.listdir(path+'test')))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_data.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Load Images\nAll samples consist of four files - blue, green, red, and yellow. Colors are \n* red for [microtubule channels](https://en.wikipedia.org/wiki/Microtubule)\n* blue for nuclei channels\n* yellow for [Endoplasmic Reticulum (ER)](https://en.wikipedia.org/wiki/Endoplasmic_reticulum) channels\n* green for protein"},{"metadata":{"trusted":true},"cell_type":"code","source":"colors_dict = {'red': 'microtubule', 'blue': 'nuclei', 'yellow': 'Endoplasmic Reticulum', 'green': 'protein'}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Load the first image of the train dataset:"},{"metadata":{"trusted":true},"cell_type":"code","source":"image_id = train_data.loc[0, 'ID']\nimage_file = cv2.imread(path+'train/'+image_id+'_blue.png')\nimage_file.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Show the 4 images for the first sample of the train dataset:"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"image_id = train_data.loc[0, 'ID']\nfig, axs = plt.subplots(1, 4, figsize=(20, 5))\nfig.subplots_adjust(hspace = .2, wspace=.1)\naxs = axs.ravel()\ncolors = list(colors_dict.keys())\nfor i in range(len(colors)):\n    filename = ''.join([image_id, '_', colors[i], '.png'])\n    image_file = cv2.imread(path+'train/'+filename)\n    axs[i].imshow(image_file)\n    axs[i].set_title(colors_dict[colors[i]])\n    axs[i].set_xticklabels([])\n    axs[i].set_yticklabels([])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Encoding Labels\nThis is a multilabel classification. The labels are separeted by | in the train dataset."},{"metadata":{"trusted":true},"cell_type":"code","source":"# original label\nprint('input :', train_data.loc[0, 'Label'])\n# label as list\nprint('step 1:', train_data.loc[0, 'Label'].split('|'))\n# label as list on integers\nprint('step 2:', list(map(int, train_data.loc[0, 'Label'].split('|'))))\n# label to binary class matrix\nlabel = to_categorical(list(map(int, train_data.loc[0, 'Label'].split('|'))), num_classes=19)\nprint('step 3:', label)\n# sum the labels\nlabel = label.sum(axis=0)\nprint('step 4:', label)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Data Generator\nWe define a data generator to load the data on demand."},{"metadata":{"trusted":true},"cell_type":"code","source":"class DataGenerator(Sequence):\n    def __init__(self, path, list_IDs, labels, batch_size, img_size, img_channel):\n        self.path = path\n        self.list_IDs = list_IDs\n        self.labels = labels\n        self.batch_size = batch_size\n        self.img_size = img_size\n        self.img_channel = img_channel\n        self.indexes = np.arange(len(self.list_IDs))\n        \n    def __len__(self):\n        len_ = int(len(self.list_IDs)/self.batch_size)\n        if len_*self.batch_size < len(self.list_IDs):\n            len_ += 1\n        return len_\n    \n    \n    def __getitem__(self, index):\n        indexes = self.indexes[index*self.batch_size:(index+1)*self.batch_size]\n        list_IDs_temp = [self.list_IDs[k] for k in indexes]\n        X, y = self.__data_generation(list_IDs_temp)\n        return X, y\n\n    \n    def __data_generation(self, list_IDs_temp):\n        X = np.zeros((self.batch_size, self.img_size, self.img_size, self.img_channel))\n        y = np.zeros((self.batch_size, num_classes), dtype=int)\n        for i, ID in enumerate(list_IDs_temp):\n            data_file = cv2.imread(self.path+ID+'_blue.png')\n            img = cv2.resize(data_file, (self.img_size, self.img_size))\n            X[i, ] = img/255.\n            # Prepare label\n            label = self.labels[i]\n            label = label.split('|')\n            label = list(map(int, label))\n            label = to_categorical(label, num_classes=num_classes)\n            label = label.sum(axis=0)\n            y[i, ] = label\n        return X, y","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Define Model"},{"metadata":{"trusted":true},"cell_type":"code","source":"weights='../input/models/resnet50_weights_tf_dim_ordering_tf_kernels_notop.h5'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"metrics = [tf.keras.metrics.AUC(name='auc', multi_label=True)]\nlearning_rate = 1e-3","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"conv_base = ResNet50(include_top=False,\n                     weights=weights,\n                     input_shape=(img_size, img_size, img_channel))\nconv_base.trainable = True","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model = Sequential()\nmodel.add(conv_base)\nmodel.add(Flatten())\nmodel.add(Dense(64, activation='relu'))\nmodel.add(BatchNormalization())\nmodel.add(Dropout(0.1))\nmodel.add(Dense(num_classes, activation='sigmoid'))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model.compile(optimizer=Adam(lr=learning_rate), loss=\"binary_crossentropy\", metrics=metrics)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model.summary()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Train Model"},{"metadata":{"trusted":true},"cell_type":"code","source":"epochs = 3","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_IDs, val_IDs, y_train, y_val = train_test_split(train_data['ID'], train_data['Label'], test_size=0.33, random_state=2021)\ntrain_IDs.index=range(len(train_IDs))\ny_train.index=range(len(train_IDs))\nval_IDs.index=range(len(val_IDs))\ny_val.index=range(len(val_IDs))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_generator = DataGenerator(path+'train/', train_IDs, y_train, batch_size, img_size, img_channel)\nval_generator = DataGenerator(path+'train/', val_IDs, y_val, batch_size, img_size, img_channel)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"history = model.fit_generator(generator=train_generator,\n                              validation_data=val_generator,\n                              epochs = epochs)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Analyse Training"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"fig, axs = plt.subplots(1, 2, figsize=(20, 6))\nfig.subplots_adjust(hspace = .2, wspace=.2)\naxs = axs.ravel()\nloss = history.history['loss']\nloss_val = history.history['val_loss']\nepochs = range(1, len(loss)+1)\naxs[0].plot(epochs, loss, 'bo', label='loss_train')\naxs[0].plot(epochs, loss_val, 'ro', label='loss_val')\naxs[0].set_title('Value of the loss function')\naxs[0].set_xlabel('epochs')\naxs[0].set_ylabel('value of the loss function')\naxs[0].legend()\naxs[0].grid()\nacc = history.history['auc']\nacc_val = history.history['val_auc']\naxs[1].plot(epochs, acc, 'bo', label='accuracy_train')\naxs[1].plot(epochs, acc_val, 'ro', label='accuracy_val')\naxs[1].set_title('Accuracy')\naxs[1].set_xlabel('Epochs')\naxs[1].set_ylabel('Value of accuracy')\naxs[1].legend()\naxs[1].grid()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Write Output"},{"metadata":{"trusted":true},"cell_type":"code","source":"output = samp_subm.copy()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"output.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"output.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Next Steps\n* Extend the data generator for all colors (blue, red, yellow, green). Currently onle blue is used.\n* Predict test data."}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}