{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Intro\nWelcome to the [RANZCR CLiP - Catheter and Line Position Challenge](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/data).\n![](https://storage.googleapis.com/kaggle-competitions/kaggle/23870/logos/header.png)\n\nIn this competition, we will detect the presence and position of catheters and lines on chest x-rays.\n\nThere are 11 tagets to predict:\n* ETT - Abnormal - endotracheal tube placement abnormal\n* ETT - Borderline - endotracheal tube placement borderline abnormal\n* ETT - Normal - endotracheal tube placement normal\n* NGT - Abnormal - nasogastric tube placement abnormal\n* NGT - Borderline - nasogastric tube placement borderline abnormal\n* NGT - Incompletely Imaged - nasogastric tube placement inconclusive due to imaging\n* NGT - Normal - nasogastric tube placement borderline normal\n* CVC - Abnormal - central venous catheter placement abnormal\n* CVC - Borderline - central venous catheter placement borderline abnormal\n* CVC - Normal - central venous catheter placement normal\n* Swan Ganz Catheter Present\n\n<span style=\"color: royalblue;\">Please vote the notebook up if it helps you. Thank you. </span>"},{"metadata":{},"cell_type":"markdown","source":"# Libraries"},{"metadata":{"trusted":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport os\nimport matplotlib.pyplot as plt\nimport cv2","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from keras.utils import to_categorical, Sequence\nfrom keras.models import Sequential\nfrom keras.layers import Dense, Dropout, Flatten, Conv2D, MaxPool2D, Activation, BatchNormalization,GlobalAveragePooling2D\nfrom keras.optimizers import RMSprop,Adam\nfrom keras.applications import ResNet50, MobileNet\nfrom tensorflow.keras.applications import EfficientNetB3\nimport tensorflow as tf","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import warnings\nwarnings.filterwarnings(\"ignore\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Path"},{"metadata":{"trusted":true},"cell_type":"code","source":"path = '/kaggle/input/ranzcr-clip-catheter-line-classification/'\nos.listdir(path)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":false},"cell_type":"markdown","source":"# Load Data"},{"metadata":{"trusted":true},"cell_type":"code","source":"train = pd.read_csv(path+'train.csv')\ntrain_anno = pd.read_csv(path+'train_annotations.csv')\nsamp_subm = pd.read_csv(path+'sample_submission.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Functions\nWe define some helper functions."},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_example(data, type_):\n    fig, axs = plt.subplots(1, 5, figsize=(25, 12))\n    fig.subplots_adjust(hspace = .2, wspace=.2)\n    axs = axs.ravel()\n    temp = data[data[type_]==1]\n    for i in range(5):\n        idx = temp.index[i]\n        image_id = temp.loc[idx, 'StudyInstanceUID']\n        image_file = cv2.imread(''.join([path, 'train/', image_id, '.jpg']))\n        image_file = cv2.cvtColor(image_file, cv2.COLOR_BGR2RGB)\n        axs[i].imshow(image_file)\n        axs[i].set_title(type_)\n        axs[i].set_xticklabels([])\n        axs[i].set_yticklabels([])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Overview"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print('number of train samples:', len(train.index))\nprint('numebr of test samples:', len(samp_subm.index))\nprint('number of train images:', len(os.listdir(path+'train/')))\nprint('number of test images:', len(os.listdir(path+'test/')))\nprint('number of unique patient ids:', len(train['PatientID'].unique()))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There are less patients as samples. So for some patients there is more than one image."},{"metadata":{},"cell_type":"markdown","source":"# Parameters"},{"metadata":{"trusted":true},"cell_type":"code","source":"image_size = 222\nimage_channel = 3\nnum_classes = 11\nlabels = train[train.columns[1:-1]].columns.tolist()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# EDA"},{"metadata":{},"cell_type":"markdown","source":"The distribution of the labels is unblanced."},{"metadata":{"trusted":true},"cell_type":"code","source":"train[train.columns[1:-1]].sum()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.bar(x=labels, height=np.sum(train[labels], axis=0))\nplt.grid()\nplt.xticks(rotation=90)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## ETT - Abnormal"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_example(train, 'ETT - Abnormal')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## ETT - Borderline"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_example(train, 'ETT - Borderline')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## ETT - Normal"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_example(train, 'ETT - Normal')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## NGT - Abnormal"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_example(train, 'NGT - Abnormal')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## NGT - Borderline"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_example(train, 'NGT - Borderline')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## NGT - Incompletely Imaged"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_example(train, 'NGT - Incompletely Imaged')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## NGT - Normal"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_example(train, 'NGT - Normal')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## CVC - Abnormal"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_example(train, 'CVC - Abnormal')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## CVC - Borderline"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_example(train, 'CVC - Borderline')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## CVC - Normal"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_example(train, 'CVC - Normal')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Swan Ganz Catheter Present"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_example(train, 'Swan Ganz Catheter Present')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Class Weights"},{"metadata":{"trusted":true},"cell_type":"code","source":"class_weight = dict(zip(range(num_classes), train[labels].sum().values/len(train.index)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Data Generator\nWe define a data generator to laod the data on demand."},{"metadata":{"trusted":true},"cell_type":"code","source":"class DataGenerator(Sequence):\n    def __init__(self, path, data, batch_size,\n                 image_size, image_channel, num_classes):\n        self.path = path\n        self.list_IDs = data['StudyInstanceUID']\n        self.labels = data[data.columns[1:12]]\n        self.batch_size = batch_size\n        self.image_size = image_size\n        self.image_channel = image_channel\n        self.num_classes = num_classes\n        self.indexes = np.arange(len(self.list_IDs))\n        \n    def __len__(self):\n        l = int(len(self.list_IDs)/self.batch_size)\n        if l*self.batch_size < len(self.list_IDs):\n            l += 1\n        return l\n        \n    \n    def __getitem__(self, index):\n        indexes = self.indexes[index*self.batch_size:(index+1)*self.batch_size]\n        list_IDs_temp = [self.list_IDs[k] for k in indexes]\n        X, y = self.__data_generation(list_IDs_temp)\n        return X, y\n\n    \n    def __data_generation(self, list_IDs_temp):\n        X = np.zeros((self.batch_size, self.image_size, self.image_size, self.image_channel))\n        y = np.zeros((self.batch_size, self.num_classes), dtype=int)\n        for i, ID in enumerate(list_IDs_temp):\n            data_file = cv2.imread(''.join([self.path, ID, '.jpg']))\n            image = cv2.resize(data_file, (self.image_size, self.image_size))\n            X[i, ] = image/255.\n            y[i, ] = self.labels.iloc[i]\n        return X, y","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Define Model"},{"metadata":{},"cell_type":"markdown","source":"## MobileNet"},{"metadata":{"trusted":true},"cell_type":"code","source":"#weights='../input/models/mobilenet_1_0_224_tf_no_top.h5'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#conv_base = MobileNet(include_top=False,\n#                     weights=weights,\n#                     input_shape=(image_size, image_size, image_channel))\n#conv_base.trainable = True","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#model = Sequential()\n#model.add(conv_base)\n#model.add(GlobalAveragePooling2D())\n#model.add(Dense(1024, activation='relu'))\n#model.add(Dense(1024, activation='relu'))\n#model.add(Dense(512, activation='relu'))\n#model.add(Dense(num_classes, activation='sigmoid'))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## EfficientNet"},{"metadata":{"trusted":true},"cell_type":"code","source":"weights = '../input/models/efficientnetb3_notop.h5'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"conv_base = EfficientNetB3(include_top=False,\n                          weights=weights,\n                          input_shape=(image_size, image_size, image_channel))\nconv_base.trainable = True","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model = Sequential()\nmodel.add(conv_base)\nmodel.add(GlobalAveragePooling2D())\nmodel.add(Dense(num_classes, activation='sigmoid'))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"metrics = [tf.keras.metrics.AUC(name='auc', multi_label=True)]\nmodel.compile(optimizer=Adam(), loss='binary_crossentropy', metrics=metrics)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model.summary()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Train Model"},{"metadata":{"trusted":true},"cell_type":"code","source":"epochs = 5\nbatch_size = 64","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_gen = DataGenerator(path+'train/', train, batch_size, image_size, image_channel, num_classes)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"history = model.fit_generator(generator=train_gen,\n                              epochs = epochs,\n                              workers=4)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Predict Test Data"},{"metadata":{"trusted":true},"cell_type":"code","source":"test_gen = DataGenerator(path+'test/', samp_subm, batch_size, image_size, image_channel, num_classes)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"predict = model.predict_generator(test_gen, verbose=1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Write Output"},{"metadata":{"trusted":true},"cell_type":"code","source":"output = pd.DataFrame(predict, columns = labels)\noutput.insert(0, 'StudyInstanceUID', samp_subm['StudyInstanceUID'])\noutput.dropna(inplace=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"output.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.bar(x=labels, height=np.sum(output[labels], axis=0))\nplt.grid()\nplt.xticks(rotation=90)\nplt.show()","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}