{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style='color:red;font-weight:500;'>SIIM-FISABIO-RSNA COVID-19 Detection: An Extended EDA </h1>\n\nIn this competition, we are provided with <span style='color:blue;font-weight:500;'>DICOM images </span> of chest X-ray radiographs, and we are asked to identify and localize COVID-19 abnormalities. This is important because typical diagnosis of COVID-19 requires molecular testing (polymerase chain reaction) requires several hours, while chest radiographs can be obtained in minutes, but it is hard to distinguish between COVID-19 pneumonia and other other viral and bacterial pneumonias. Therefore, in this competition, be hope to develop AI that that eventually help radiologists diagnose the millions of COVID-19 patients more confidently and quickly.\n\nI'll provide a quick and simple EDA to help you get started with this very interesting competition!\n","metadata":{}},{"cell_type":"markdown","source":"<span style='color:green;font-size:20px;font-weight:500;'>This Notebook is forked from anoder EDA notebook. But in this one , I am going to explain all the fact about this competition that confused me from beginning. </span>","metadata":{}},{"cell_type":"markdown","source":"<span style='color:crimson;font-size:20px;font-weight:500;'><b> First thing before Getting Started : </b> The data might look huge, but there is only a small number of images. We will determine the exact number of images in later part of the notebook. The huge size is mainly due to `DICOM` format. If we extract only the images in PNG format, even in high quality the dataset comes down to 3-4 GB only.  </span> ","metadata":{}},{"cell_type":"markdown","source":"<span style='color:blue;font-size:15px;font-weight:500;'>Some kind Kagglers have converted the dataset to JPG or PNG format. Though I will show in this notebook how to convert from `.dcm` to .`.jpg'/'.png`, but the converted datasets will also be linked. </span> ","metadata":{}},{"cell_type":"markdown","source":"\n<span style='color:blue;font-size:18px;font-weight:500;'>DICOM : </span> `It is the standard format for the communication and management of medical imaging information and related data.`\n\n","metadata":{}},{"cell_type":"code","source":"! conda install -c conda-forge gdcm -y","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-05-23T12:23:29.012267Z","iopub.execute_input":"2021-05-23T12:23:29.012717Z","iopub.status.idle":"2021-05-23T12:23:29.893279Z","shell.execute_reply.started":"2021-05-23T12:23:29.01263Z","shell.execute_reply":"2021-05-23T12:23:29.892296Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# import numpy as np # linear algebra\n# import pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n# import os\n# import pydicom\n# import glob\n# from tqdm.notebook import tqdm\n# from pydicom.pixel_data_handlers.util import apply_voi_lut\n# import matplotlib.pyplot as plt\n# from skimage import exposure\n# import cv2\n# import warnings\n# from fastai.vision.all import *\n# from fastai.medical.imaging import *\n# warnings.filterwarnings('ignore')\n\nimport os\nimport pydicom\nimport glob\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt \n","metadata":{"execution":{"iopub.status.busy":"2021-05-23T15:08:27.091641Z","iopub.execute_input":"2021-05-23T15:08:27.092064Z","iopub.status.idle":"2021-05-23T15:08:27.336534Z","shell.execute_reply.started":"2021-05-23T15:08:27.091978Z","shell.execute_reply":"2021-05-23T15:08:27.335262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style='color:green;'>A look at the provided data </h1>","metadata":{}},{"cell_type":"code","source":"dataset_path = '../input/siim-covid19-detection/'\n\nfor path in glob.glob(dataset_path  + '*'):\n    print(path)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2021-05-23T15:08:27.338015Z","iopub.execute_input":"2021-05-23T15:08:27.338315Z","iopub.status.idle":"2021-05-23T15:08:27.344998Z","shell.execute_reply.started":"2021-05-23T15:08:27.338282Z","shell.execute_reply":"2021-05-23T15:08:27.343796Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style='color:blue;font-size:18px;font-weight:500;'> We can see that we have:</span>\n\n* `train_study_level.csv` - the train study-level metadata, with one row for each study, including correct labels.\n* `train_image_level.csv` - the train image-level metadata, with one row for each image, including both correct labels and any bounding boxes in a dictionary format. Some images in both test and train have multiple bounding boxes.\n* `sample_submission.csv` - a sample submission file containing all image- and study-level IDs.\n* `train` folder - comprises 6,334 chest scans in DICOM format, stored in paths with the form `study`/`series`/`image`\n* `test` folder - The hidden test dataset is of roughly the same scale as the training dataset.\n","metadata":{}},{"cell_type":"markdown","source":"<span style='color:crimson;font-size:18px;font-weight:500;'> Train folder analysis: </span>\n\n\n`We see that there are some folders in train folder. Each of them have atleast one subfolder in them, and each of the subfolders have at least one dcm file, ","metadata":{}},{"cell_type":"code","source":"train_images = '../input/siim-covid19-detection/train/'\n\nn_folders = len(glob.glob(train_images  + '*'))\nn_subfolders = len(glob.glob(train_images  + '*/*'))\nn_images = len(glob.glob(train_images  + '*/*/*.dcm'))\n\nprint(f'There are {n_subfolders} subfolders in {n_folders} folders.')\nprint(f'There are altogether {n_images} images.')","metadata":{"execution":{"iopub.status.busy":"2021-05-23T15:08:55.034156Z","iopub.execute_input":"2021-05-23T15:08:55.034560Z","iopub.status.idle":"2021-05-23T15:09:22.017014Z","shell.execute_reply.started":"2021-05-23T15:08:55.034522Z","shell.execute_reply":"2021-05-23T15:09:22.015906Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"`So , some folders have more than one subfolders and some subfolders have more than one image.`","metadata":{}},{"cell_type":"code","source":"folders = glob.glob(train_images  + '*')\n\nsubfolder_dict = {}\n\nfor folder in folders:\n    n = len(glob.glob(folder + '/*'))\n    if n in subfolder_dict.keys():\n        subfolder_dict[n] += 1\n    else:\n        subfolder_dict[n] = 1\n        \nfor k, v in subfolder_dict.items():\n    print(f'There is {v} subfolders with {k} images.')\n\n","metadata":{"execution":{"iopub.status.busy":"2021-05-23T15:09:22.018892Z","iopub.execute_input":"2021-05-23T15:09:22.019281Z","iopub.status.idle":"2021-05-23T15:09:24.590290Z","shell.execute_reply.started":"2021-05-23T15:09:22.019236Z","shell.execute_reply":"2021-05-23T15:09:24.589475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"subfolders = glob.glob(train_images  + '*/*')\n\nimage_dict = {}\n\nfor subfolder in subfolders:\n    n = len(glob.glob(subfolder + '/*'))\n    if (n!=1):\n        print (subfolder.split('/')[-1])\n        for image in glob.glob(subfolder + '/*'):\n            print ('-->' + image.split('/')[-1])\n\n    if n in image_dict.keys():\n        image_dict[n] += 1\n    else:\n        image_dict[n] = 1\n        \nfor k, v in image_dict.items():\n    print(f'There is {v} subfolders with {k} images.')\n","metadata":{"execution":{"iopub.status.busy":"2021-05-23T14:48:13.115208Z","iopub.execute_input":"2021-05-23T14:48:13.115587Z","iopub.status.idle":"2021-05-23T14:48:18.410012Z","shell.execute_reply.started":"2021-05-23T14:48:13.115540Z","shell.execute_reply":"2021-05-23T14:48:18.408702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style='color:crimson;font-size:18px;font-weight:500;'> Major thing to remember is: </span>\n`There are 6331 subfolders in 6054 folders.\nThere are altogether 6334 images. Lets keep those numbers around.`","metadata":{}},{"cell_type":"markdown","source":"<span style='color:green;font-size:18px;font-weight:500;'>What is Study Level? What is Image Level? </span> <br><br> `Wait for a sec...We will get the answer  just after analyzing the CSV files.`\n\n","metadata":{}},{"cell_type":"markdown","source":"<span style='color:red;font-size:18px;font-weight:500;'>train_study_level.csv :</span>","metadata":{}},{"cell_type":"code","source":"train_study_df = pd.read_csv(dataset_path +'/train_study_level.csv')\n\ntrain_study_df","metadata":{"execution":{"iopub.status.busy":"2021-05-23T15:10:56.782102Z","iopub.execute_input":"2021-05-23T15:10:56.782493Z","iopub.status.idle":"2021-05-23T15:10:56.810380Z","shell.execute_reply.started":"2021-05-23T15:10:56.782460Z","shell.execute_reply":"2021-05-23T15:10:56.809271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So we have `6054` rows in this. All the rows are classified in following classes:\n* `Negative for Pneumonia`\n* `Typical Appearance`\n* `Intermediate Appearance`\n* `Atypical Appearance`\n\n <b>So the `Study Level`  looks something like a classification task</b>. We will check later, if this is single label classification or multilabel. \n\nDo you remeber those numbers ?? <br> \nYes, The `train_study_level.csv` targets these folder number.<b> Each Folder refers to a study.</b>\n\n<span style='color:blue;font-size:16px;font-weight:500;'>Now, Checking frequency of classes...</span>\n","metadata":{}},{"cell_type":"code","source":"study_classes = train_study_df.drop(columns = ['id']).columns.tolist()\nplt.figure(figsize = (10,5))\nplt.bar([1,2,3,4], train_study_df[study_classes].values.sum(axis=0))\nplt.xticks([1,2,3,4],study_classes)\nplt.ylabel('Frequency')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-05-23T15:18:09.304757Z","iopub.execute_input":"2021-05-23T15:18:09.305329Z","iopub.status.idle":"2021-05-23T15:18:09.467451Z","shell.execute_reply.started":"2021-05-23T15:18:09.305281Z","shell.execute_reply":"2021-05-23T15:18:09.466473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This seems okay. Because, naturally number of `Typycal Appearance` will be highest, because that is the general case. Then There should be number of `Negative for Pneumonia` , which would be less than the normal case, as those X-rays are for pataints, not random people.There will be a moderate number of cases that the doctors will fail to identify. Those are `Indeteminate Appearence`. And the lowest number willbe of wired `Atypical Cases`.\n\n<span style='color:blue;font-size:16px;font-weight:500;'>Lets check if any row has multiples labels</span>","metadata":{}},{"cell_type":"code","source":"train_study_df[study_classes].sum(axis = 1).value_counts()","metadata":{"execution":{"iopub.status.busy":"2021-05-23T15:22:42.627327Z","iopub.execute_input":"2021-05-23T15:22:42.627692Z","iopub.status.idle":"2021-05-23T15:22:42.638740Z","shell.execute_reply.started":"2021-05-23T15:22:42.627657Z","shell.execute_reply":"2021-05-23T15:22:42.638009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that each row has sum 1 after summing them along the row. So, `each of the 6054 studies have only one label each.`","metadata":{}},{"cell_type":"markdown","source":"<span style='color:red;font-size:18px;font-weight:500;'>train_study_level.csv :</span>","metadata":{}},{"cell_type":"code","source":"train_image_df = pd.read_csv(dataset_path +'/train_image_level.csv')\n\ntrain_image_df","metadata":{"execution":{"iopub.status.busy":"2021-05-23T15:25:24.182937Z","iopub.execute_input":"2021-05-23T15:25:24.183326Z","iopub.status.idle":"2021-05-23T15:25:24.232406Z","shell.execute_reply.started":"2021-05-23T15:25:24.183295Z","shell.execute_reply":"2021-05-23T15:25:24.231623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Firstly, Bboxex are goven here. That means it has something to do with localizing some portion of image. <b>This is an `Objet Detection Task`</b><br>\nDo you recognize the row number?? <br>\nYap! That is the total image number.\n\n<b>Thus `Image Level` means nothing but prediction for each image.</b>\n\nWe have our bounding box labels provided in the `label` column. The format is as follows:\n\n`[class ID] [confidence score] [bounding box]`\n\n* class ID - either `opacity` or `none`\n* confidence score - confidence from your neural network model. If none, the confidence is `1`.\n* bounding box - typical `xmin ymin xmax ymax` format. If class ID is none, the bounding box is `1 0 0 1 1`.\n\nThe bounding boxes are also provided in easily readable dictionary format in column `boxes`, and the study that each image is a part of is provided in`StudyInstanceUID`.\n\n<span style='color:blue;font-size:16px;font-weight:500;'>Now we will create a new column, splitting the label to more understandable list format.</span>","metadata":{}},{"cell_type":"code","source":"train_image_df['split_label'] = train_image_df.label.apply(lambda x: [x.split()[offs:offs+6] for offs in range(0, len(x.split()), 6)])\n\ntrain_image_df['split_label']","metadata":{"execution":{"iopub.status.busy":"2021-05-23T15:49:38.392184Z","iopub.execute_input":"2021-05-23T15:49:38.392591Z","iopub.status.idle":"2021-05-23T15:49:38.429457Z","shell.execute_reply.started":"2021-05-23T15:49:38.392559Z","shell.execute_reply":"2021-05-23T15:49:38.428440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Okay, If the label is None, there is inly one detection. But if not `None` , there might be a number of detections.\n\n<span style='color:blue;font-size:16px;font-weight:500;'>Now I will create one more column to estimate how many bounding boxes are there for each image.</span>\n","metadata":{}},{"cell_type":"code","source":"train_image_df['label_len'] = train_image_df['split_label'].apply(lambda x: 0 if (x[0][0] == 'none') else len(x))\ntrain_image_df['label_len'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2021-05-23T16:02:16.426977Z","iopub.execute_input":"2021-05-23T16:02:16.427298Z","iopub.status.idle":"2021-05-23T16:02:16.439609Z","shell.execute_reply.started":"2021-05-23T16:02:16.427270Z","shell.execute_reply":"2021-05-23T16:02:16.438609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, It looks like , there is 2040 images with no detection. 3113 images with 2 detections (understadibly in two lobes of lung).\n\n<span style='color:blue;font-size:16px;font-weight:500;'>Let's quick look at the distribution of opacity vs none:</span>","metadata":{}},{"cell_type":"code","source":"classes_freq = []\nfor i in range(len(train_image_df)):\n    for j in train_image_df.iloc[i].split_label: classes_freq.append(j[0])\nplt.hist(classes_freq)\nplt.ylabel('Frequency')","metadata":{"execution":{"iopub.status.busy":"2021-05-23T15:50:16.188697Z","iopub.execute_input":"2021-05-23T15:50:16.189085Z","iopub.status.idle":"2021-05-23T15:50:16.907316Z","shell.execute_reply.started":"2021-05-23T15:50:16.189054Z","shell.execute_reply":"2021-05-23T15:50:16.906218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\n","metadata":{}},{"cell_type":"code","source":"bbox_areas = []\nfor i in range(len(train_image_df)):\n    for j in train_image_df.iloc[i].split_label:\n        bbox_areas.append((float(j[4])-float(j[2]))*(float(j[5])*float(j[3])))\nplt.hist(bbox_areas)\nplt.ylabel('Frequency')","metadata":{"execution":{"iopub.status.busy":"2021-05-23T15:57:29.388954Z","iopub.execute_input":"2021-05-23T15:57:29.389400Z","iopub.status.idle":"2021-05-23T15:57:30.135333Z","shell.execute_reply.started":"2021-05-23T15:57:29.389364Z","shell.execute_reply":"2021-05-23T15:57:30.134217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style='color:crimson;font-size:20px;font-weight:500;'><b> If you like the EDA approach or the notebook, a lot more thing will be added. Thank you. </b>  </span> ","metadata":{}},{"cell_type":"markdown","source":"That's it for now!\n\n**Please upvote if you found this helpful!**","metadata":{}}]}