{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### What is the competition?\n​\nThis competition is about identifying and localizing COVID-19 abnormalities on chest radiographs. We need to categorize a given chest radiograph image into 4 different classes namely\n​\n* Negative for Pneumonia\n* Typical Appearance \n* Indeterminate Appearance\n* Atypical Appearance\n​\nThe scope of the competition also includes predicting a bounding box that describes the abnormalities. However, for the scope of this notebook, we will focus on the below items:\n1. Understanding the dataset by visualizing some of the DICOM images\n2. Process the DICOM image data for a simple binary classification CNN Model\n3. Export DICOM images as jpg files in separate train and test folders. \n\nWe will use the jpg files and this notebook as the basis for the Neural Network, for which we will create a separate notebook. \n\n## Understanding the Dataset: \nRef: (https://www.dicomstandard.org)\n### What is DICOM?\n​\nDICOM® — [Digital Imaging and Communications in Medicine](https://www.dicomstandard.org) — is the ISO recognized international standard for medical images and related information. It defines the formats for medical images that can be exchanged with the data and quality necessary for clinical use.\n​","metadata":{}},{"cell_type":"markdown","source":"Lets Import the necessary libraries","metadata":{}},{"cell_type":"code","source":"!conda install gdcm -c conda-forge -y","metadata":{"execution":{"iopub.status.busy":"2021-07-23T22:22:17.007637Z","iopub.execute_input":"2021-07-23T22:22:17.007998Z","iopub.status.idle":"2021-07-23T22:23:19.045064Z","shell.execute_reply.started":"2021-07-23T22:22:17.007964Z","shell.execute_reply":"2021-07-23T22:23:19.043937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\nimport numpy as np\nimport pandas as pd\nfrom pandas import DataFrame\nfrom tensorflow import keras\nfrom keras import backend as K\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras import layers\nfrom tensorflow.keras.layers import Activation, Dense, Flatten, BatchNormalization, Conv2D, MaxPooling2D, Dropout\nfrom tensorflow.keras.layers import Input\nfrom tensorflow.keras.optimizers import Adam, SGD, RMSprop\nfrom tensorflow.keras.metrics import categorical_crossentropy\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom sklearn.metrics import confusion_matrix\nimport plotly.io as pio\nimport itertools\nimport os\nimport shutil\nimport random\nimport glob\nfrom tqdm.notebook import tqdm\nimport matplotlib.pyplot as plt\nimport cv2\nimport plotly_express as px\n\nimport os\n\nfrom PIL import Image\nimport pandas as pd\nfrom tqdm.auto import tqdm\nimport pydicom\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\n\n\n\n#from pydicom import dcmread\n\nfrom sklearn.metrics import roc_curve, auc, roc_auc_score, accuracy_score, average_precision_score, classification_report, confusion_matrix, plot_confusion_matrix\nfrom sklearn.model_selection import train_test_split\n\nimport warnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)\n%matplotlib inline","metadata":{"execution":{"iopub.status.busy":"2021-07-23T22:23:40.073075Z","iopub.execute_input":"2021-07-23T22:23:40.073391Z","iopub.status.idle":"2021-07-23T22:23:40.091776Z","shell.execute_reply.started":"2021-07-23T22:23:40.073357Z","shell.execute_reply":"2021-07-23T22:23:40.090871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\nfrom colorama import Fore, Back, Style\n\ny_ = Fore.YELLOW\nr_ = Fore.RED\ng_ = Fore.GREEN\nb_ = Fore.BLUE\nm_ = Fore.MAGENTA\nc_ = Fore.CYAN\nres = Style.RESET_ALL\n","metadata":{"execution":{"iopub.status.busy":"2021-07-23T21:52:45.156657Z","iopub.execute_input":"2021-07-23T21:52:45.156970Z","iopub.status.idle":"2021-07-23T21:52:45.161589Z","shell.execute_reply.started":"2021-07-23T21:52:45.156941Z","shell.execute_reply":"2021-07-23T21:52:45.160551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Assigning path to csv datasets and create dataframes for the image and study level data. We will further visualize the data as Dataframes.","metadata":{}},{"cell_type":"code","source":"PATH = '/kaggle/input/datafiles/'\nsubmission = pd.read_csv('/kaggle/input/datafiles/sample_submission.csv', index_col=None)\nimage_df = pd.read_csv('/kaggle/input/datafiles/train_image_level.csv', index_col=None)\nstudy_df = pd.read_csv('/kaggle/input/datafiles/train_study_level.csv', index_col=None)\npd.set_option('display.max_columns', None)  \npd.set_option('display.max_colwidth', None)\nprint(f\"{y_}Train image level csv shape : {image_df.shape}{res}\\n{g_}Train study level csv shape : {study_df.shape}{res}\")","metadata":{"execution":{"iopub.status.busy":"2021-07-23T21:55:04.163740Z","iopub.execute_input":"2021-07-23T21:55:04.164114Z","iopub.status.idle":"2021-07-23T21:55:04.204943Z","shell.execute_reply.started":"2021-07-23T21:55:04.164079Z","shell.execute_reply":"2021-07-23T21:55:04.204049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Lets see the study level data\nstudy_df","metadata":{"execution":{"iopub.status.busy":"2021-07-23T21:55:35.168746Z","iopub.execute_input":"2021-07-23T21:55:35.169111Z","iopub.status.idle":"2021-07-23T21:55:35.182400Z","shell.execute_reply.started":"2021-07-23T21:55:35.169071Z","shell.execute_reply":"2021-07-23T21:55:35.181504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As expected, we have a classification dataset with multiple classes to work with. Lets see what the image level data looks like. ","metadata":{}},{"cell_type":"code","source":"image_df","metadata":{"execution":{"iopub.status.busy":"2021-07-23T21:57:21.030041Z","iopub.execute_input":"2021-07-23T21:57:21.030369Z","iopub.status.idle":"2021-07-23T21:57:21.044757Z","shell.execute_reply.started":"2021-07-23T21:57:21.030338Z","shell.execute_reply":"2021-07-23T21:57:21.042980Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The Image level Data is mostly has information related to the bounding boxes and labels associated with the conditions. Even though we decided to restrict ourselves from the Localization part of this competition, we will at least visualize how this image data looks like. ","metadata":{}},{"cell_type":"code","source":"#Now lets check how our Image dataset is distributed across the 4 labels. \n#Creating a pivot table\nstudy_grp = pd.melt(study_df, id_vars=list(study_df.columns)[:1], value_vars=list(study_df.columns)[1:],\n             var_name='label', value_name='value')\nstudy_grp = study_grp.loc[study_grp['value']!=0]\n\n#Assigning Random selection of Colors to each label.\ncolors = {'Typical Appearance' : '#DCD427',\n'Negative for Pneumonia' : '#0092CC',\n'Indeterminate Appearance' : '#CC3333',\n          'Atypical Appearance' : '#E6E6E6'\n         }\nstudy_grp = study_grp.groupby('label').sum().sort_values('value',ascending=False).reset_index()\nstudy_grp['color'] = study_grp['label'].apply(lambda x: colors[x])\nstudy_grp","metadata":{"execution":{"iopub.status.busy":"2021-07-23T22:03:18.106634Z","iopub.execute_input":"2021-07-23T22:03:18.106963Z","iopub.status.idle":"2021-07-23T22:03:18.141329Z","shell.execute_reply.started":"2021-07-23T22:03:18.106933Z","shell.execute_reply":"2021-07-23T22:03:18.140483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Credit: https://www.kaggle.com/rajsengo \n#Creating a function to visualize Data for each label. \ndef plot_study_label(df):\n    pio.templates.default = \"plotly_dark\"\n    fig = px.bar(df, x='label', y='value',\n             hover_data=['label', 'value'], color='label',\n             #labels={column: label},\n             color_discrete_map=colors,\n             text='value')\n    fig.update_layout(xaxis={'categoryorder':'array', 'categoryarray': df['label'],\n                             'title' : None, \n                             'showgrid':False},\n                      yaxis={'showgrid':False,\n                            'title' : 'Count'},\n                      showlegend=False,\n                     title = 'Study samples in train data')\n    fig.update_traces(textfont_size=16)\n    fig.show()","metadata":{"execution":{"iopub.status.busy":"2021-07-23T22:05:29.398188Z","iopub.execute_input":"2021-07-23T22:05:29.398534Z","iopub.status.idle":"2021-07-23T22:05:29.404925Z","shell.execute_reply.started":"2021-07-23T22:05:29.398503Z","shell.execute_reply":"2021-07-23T22:05:29.403960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_study_label(study_grp)","metadata":{"execution":{"iopub.status.busy":"2021-07-23T22:02:16.653302Z","iopub.execute_input":"2021-07-23T22:02:16.653648Z","iopub.status.idle":"2021-07-23T22:02:17.077109Z","shell.execute_reply.started":"2021-07-23T22:02:16.653621Z","shell.execute_reply":"2021-07-23T22:02:17.076178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## VISUALIZE AND PROCESS DICOM IMAGE\n### Overview (Ref: https://www.dicomstandard.org)\nDICOM - Digital Imaging and Communication in Medicine is a standard format for encoding and transmitting medical images. This format stores image metadata like patient information, image acquistion parameters, image size, pixel size, etc along with the actual image. The image metadata is stored in the DICOM header. The image pixel data may be compressed using various techniques like JPEG, lossless JPEG, run length encoding (RLE), etc. Let's load a sample image to look at its header (metadata) and the actual pixel_array.","metadata":{}},{"cell_type":"code","source":"import pydicom\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\n\ndef read_xray(path, voi_lut = True, fix_monochrome = True):\n    # Original from: https://www.kaggle.com/raddar/convert-dicom-to-np-array-the-correct-way\n    dicom = pydicom.read_file(path)\n    \n    # VOI LUT (if available by DICOM device) is used to transform raw DICOM data to \n    # \"human-friendly\" view\n    if voi_lut:\n        data = apply_voi_lut(dicom.pixel_array, dicom)\n    else:\n        data = dicom.pixel_array\n               \n    # depending on this value, X-ray may look inverted - fix that:\n    if fix_monochrome and dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = np.amax(data) - data\n        \n    data = data - np.min(data)\n    data = data / np.max(data)\n    data = (data * 255).astype(np.uint8)\n        \n    return data\n\ndef resize(array, size, keep_ratio=False, resample=Image.LANCZOS):\n    # Original from: https://www.kaggle.com/xhlulu/vinbigdata-process-and-resize-to-image\n    im = Image.fromarray(array)\n    \n    if keep_ratio:\n        im.thumbnail((size, size), resample)\n    else:\n        im = im.resize((size, size), resample)\n    \n    return im","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv('../input/siim-covid19-detection/train_image_level.csv')\npath = '../input/siim-covid19-detection/train/ae3e63d94c13/288554eb6182/e00f9fe0cce5.dcm'\ndicom = pydicom.read_file(path)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_id = []\ndim0 = []\ndim1 = []\nsplits = []\n\nfor split in ['test', 'train']:\n    save_dir = f'/kaggle/tmp/{split}/'\n\n    os.makedirs(save_dir, exist_ok=True)\n    \n    for dirname, _, filenames in tqdm(os.walk(f'../input/siim-covid19-detection/{split}')):\n        for file in filenames:\n            # set keep_ratio=True to have original aspect ratio\n            xray = read_xray(os.path.join(dirname, file))\n            im = resize(xray, size=256)  \n            im.save(os.path.join(save_dir, file.replace('dcm', 'jpg')))\n\n            image_id.append(file.replace('.dcm', ''))\n            dim0.append(xray.shape[0])\n            dim1.append(xray.shape[1])\n            splits.append(split)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Merging study_level and image_level\n# rename id column in study_level to StudyInstanceUID\nstudy_df.rename(columns = {'id':'StudyInstanceUID'}, inplace = True)\n\n# remove _study from StudyInstanceUID\nstudy_df['StudyInstanceUID'] = study_df['StudyInstanceUID'].str.replace('_study', '')\n\n# merge\ndf_train = pd.merge(image_df, study_df, on='StudyInstanceUID')\n\n# remove _image from id column\ndf_train['id'] = df_train['id'].str.replace('_image', '')\n\n# rename id column as imageID\ndf_train.rename(columns = {'id':'imageID'}, inplace = True)\n\n# renaming target columns\ndf_train.rename(columns = {'Negative for Pneumonia':'negative'}, inplace = True)\ndf_train.rename(columns = {'Typical Appearance':'typical'}, inplace = True)\ndf_train.rename(columns = {'Indeterminate Appearance':'indeterminate'}, inplace = True)\ndf_train.rename(columns = {'Atypical Appearance':'atypical'}, inplace = True)\n\n# Create a new target column\ncategories = ['negative','typical','indeterminate','atypical']\ndf = df_train[categories]\ndf_train[\"target\"] = pd.Series(df.columns[np.where(df!=0)[1]])\ndf_train.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#TRAIN_DIR = f'{PATH}'\n#paths = []\n\n#for instance_id in tqdm(df_train['StudyInstanceUID']):\n    #paths.append(glob.glob(os.path.join(TRAIN_DIR, instance_id +\"/*/*\"))[0])\n\n#df_train['path'] = paths\n#df_train[:5]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_train['imageID Length'] = df_train['imageID'].apply(len)\n#df_train['imageID Length'].unique()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_tr = df_train[['path', 'negative']]\n#df_tr","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#df_training, df_valid = train_test_split(df_tr, test_size=0.20, random_state=13, stratify=df['negative'])\n#len(df_training), len(df_valid)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_jpg_directory = '../input/train-jpg-images'\ntest_jpg_directory = '../input/testjpgfiles'\n\ndef getImagePaths(path):\n    image_names = []\n    for dirname, _, filenames in os.walk(path):\n        for filename in filenames:\n            fullpath = os.path.join(dirname, filename)\n            image_names.append(fullpath)\n    return image_names\n\ntrain_images_path = getImagePaths(train_jpg_directory)\ntest_images_path = getImagePaths(test_jpg_directory)\n\nprint(f\"{y_}Number of train images: {g_} {len(train_images_path)}\\n\")\nprint(f\"{y_}Number of test images: {g_} {len(test_images_path)}\\n\")\n\ndef getShape(data, images_paths):\n    shape = cv2.imread(images_paths[0]).shape\n    for image_path in images_paths:\n        image_shape=cv2.imread(image_path).shape\n        if (image_shape!=shape):\n            return data +\" - Different image shape\"\n        else:\n            return data +\" - Same image shape \" + str(shape)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"getShape('train',train_images_path)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_images = DataFrame(train_images_path,columns=['train_images_path'])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_images['imageID'] = df_train_images['train_images_path'].str.slice(26,38)\ndf_train_images['Image_Name'] = df_train_images['train_images_path'].str.slice(26,42)\ndf_train_images","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_final = pd.merge(df_train,df_train_images, on='imageID')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_final","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_final = df_train_final.drop(['boxes', 'label', 'StudyInstanceUID', 'typical', 'indeterminate', 'atypical', 'target', 'imageID'], axis=1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_final","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_final['negative'] = df_train_final['negative'].astype('str')\n\nBATCH_SIZE = 16\nIMG_SIZE = (150, 150)\ntrain_datagen = ImageDataGenerator(\n    rescale=1./255,\n    #preprocessing_function=preprocess_input,\n    rotation_range=20,\n    width_shift_range=0.2,\n    height_shift_range=0.2,\n    brightness_range=[0.6, 1.3],\n    shear_range=0.3,\n    zoom_range=[0.8, 1.0],\n    horizontal_flip=True,\n    vertical_flip=True,\n    fill_mode='constant'\n)\n\ntest_datagen = ImageDataGenerator(\n    rescale=1./255,\n#     preprocessing_function=preprocess_input,\n)\n\ntrain_generator = train_datagen.flow_from_dataframe(\n                                        dataframe=df_train_final,\n                                        directory='../input/train-jpg-images/',\n                                        x_col='Image_Name',\n                                        y_col='negative',\n                                        class_mode='binary',\n                                        target_size=IMG_SIZE,\n                                        batch_size=BATCH_SIZE)\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split data into train-test data sets\n\nX = df_train_final.loc[:,'Image_Name']\ny = df_train_final.loc[:,'negative']\n\n# Split\ntrain_x, val_x, train_y, val_y = train_test_split(X, y, \n                                                  test_size = 0.1, \n                                                  random_state = 27, \n                                                  stratify=y)\n\n# Train df\ndf_train = pd.DataFrame(columns=['Image_Name','negative'])\ndf_train['Image_Name'] = train_x\ndf_train['negative'] = train_y\n\n# Test df\ndf_test= pd.DataFrame(columns=['image_name','negative'])\ndf_test['Image_Name'] = val_x\ndf_test['negative'] = val_y\n\ndf_train.reset_index(drop=True, inplace=True)\ndf_test.reset_index(drop=True, inplace=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"home_path = r'../input/train-jpg-images/'\n\n# Images\ntrain_images = df_train.loc[:,'Image_Name']\ntrain_labels = df_train.loc[:,'negative']\n\ntest_images = df_test.loc[:,'Image_Name']\ntest_labels = df_test.loc[:,'negative']\n\n# Train images\nx_train = []\nfor i in train_images:\n    image = home_path + i\n    img = cv2.imread(image)\n    x_train.append(img)\n\n# Train labels\ny_train=keras.utils.to_categorical(train_labels)\n\n# Test images\nx_test = []\nfor i in test_images:\n    image = home_path+i\n    img = cv2.imread(image)\n    x_test.append(img)\n\n# Test labels\ny_test=keras.utils.to_categorical(test_labels)\n\n# Normalize images\nx_train = np.array(x_train, dtype=\"float\") / 255.0\nx_test = np.array(x_test, dtype=\"float\") / 255.0","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#x_train = x_train[:500]\n#y_train = y_train[:500]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_train.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# specify image dimensions to be used\nimg_width, img_height = 150, 150\n\n# specify first channels to represent color channels \nif K.image_data_format() == 'channels_first':\n    input_shape = (1, img_width, img_height)\nelse:\n    input_shape = (img_width, img_height, 3)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"input_shape\n# Reshaping list of images\n#x_train = x_train.reshape(list(x_train.shape) + [1]) \n#x_test = x_test.reshape(list(x_test.shape) + [1])\n#y_train = y_train.reshape(list(y_train.shape) + [1])\n#y_test = y_test.reshape(list(y_test.shape) + [1])\n                                                        ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"y_train shape:\", x_train.shape)\nprint(y_train.shape[0], \"train samples\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_train[0]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Model architechture\nmodel = Sequential()\n\nmodel.add(Conv2D(32, (3, 3), input_shape=(256,256,3), activation='relu'))\nmodel.add(MaxPooling2D(pool_size=(2, 2)))\n\nmodel.add(Conv2D(32, (3, 3), activation='relu'))\nmodel.add(MaxPooling2D(pool_size=(2, 2)))\n\n#model.add(Conv2D(64, (3, 3), activation='relu'))\n#model.add(MaxPooling2D(pool_size=(2, 2)))\n\nmodel.add(Flatten())\n#model.add(Dense(64, activation='relu'))\n#model.add(Dropout(0.24))\n#model.add(Dense(2,activation='softmax'))\n\nmodel.summary()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.input","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Compile\noptim = Adam(learning_rate=0.001)\nmodel.compile(loss='binary_crossentropy',\n              optimizer=optim,\n              metrics=['accuracy'])\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Fit\nhistory = model.fit(x_train, y_train, epochs = 5, batch_size = 2, verbose = 1,\n                    validation_data = (x_test, y_test))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}