{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **SIIM-COVID19 : DATA EXPLORATION**","metadata":{}},{"cell_type":"code","source":"from IPython.display import Image\nImage(\"../input/headercovid/header.png\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-05-26T05:00:41.938057Z","iopub.execute_input":"2021-05-26T05:00:41.938673Z","iopub.status.idle":"2021-05-26T05:00:41.967502Z","shell.execute_reply.started":"2021-05-26T05:00:41.938550Z","shell.execute_reply":"2021-05-26T05:00:41.966495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a href='https://fr.freepik.com/photos/fond'>Fond photo créé par kjpargeter - fr.freepik.com</a>","metadata":{}},{"cell_type":"markdown","source":"### *1. DATA PROCESSING*","metadata":{}},{"cell_type":"markdown","source":"Let's import some libraries we will use in this notebook ...","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport os\nimport matplotlib.pyplot as plt\nimport matplotlib\nimport pydicom as dicom\nimport cv2\nimport ast\nimport warnings\nfrom collections import Counter\nimport seaborn as sns\nwarnings.filterwarnings('ignore')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2021-05-26T05:00:45.358442Z","iopub.execute_input":"2021-05-26T05:00:45.358841Z","iopub.status.idle":"2021-05-26T05:00:45.742433Z","shell.execute_reply.started":"2021-05-26T05:00:45.358808Z","shell.execute_reply":"2021-05-26T05:00:45.741663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"First of all, we define the path variable and check files and folders :","metadata":{}},{"cell_type":"code","source":"path = '/kaggle/input/siim-covid19-detection/'\nos.listdir(path)","metadata":{"execution":{"iopub.status.busy":"2021-05-26T05:00:48.990411Z","iopub.execute_input":"2021-05-26T05:00:48.990942Z","iopub.status.idle":"2021-05-26T05:00:48.998466Z","shell.execute_reply.started":"2021-05-26T05:00:48.990899Z","shell.execute_reply":"2021-05-26T05:00:48.997810Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It's time to import data in two differents dataframes :","metadata":{}},{"cell_type":"code","source":"idf = pd.read_csv(path + 'train_image_level.csv')\nsdf = pd.read_csv(path + 'train_study_level.csv')","metadata":{"execution":{"iopub.status.busy":"2021-05-26T05:00:51.215477Z","iopub.execute_input":"2021-05-26T05:00:51.215871Z","iopub.status.idle":"2021-05-26T05:00:51.256179Z","shell.execute_reply.started":"2021-05-26T05:00:51.215836Z","shell.execute_reply":"2021-05-26T05:00:51.255123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It's always better to check first few lines of each dataframes ...\nWe can already guess that there is a join between the 2 dataframes on \"StudyInstanceUID\" from image and \"id\" from study. We also note the suffix \"_study\" on the latter.","metadata":{}},{"cell_type":"code","source":"idf.head(3)","metadata":{"execution":{"iopub.status.busy":"2021-05-26T05:00:53.815509Z","iopub.execute_input":"2021-05-26T05:00:53.815999Z","iopub.status.idle":"2021-05-26T05:00:53.829147Z","shell.execute_reply.started":"2021-05-26T05:00:53.815965Z","shell.execute_reply":"2021-05-26T05:00:53.827946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sdf.head(3)","metadata":{"execution":{"iopub.status.busy":"2021-05-26T05:00:56.285475Z","iopub.execute_input":"2021-05-26T05:00:56.285883Z","iopub.status.idle":"2021-05-26T05:00:56.296423Z","shell.execute_reply.started":"2021-05-26T05:00:56.285850Z","shell.execute_reply":"2021-05-26T05:00:56.295441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's display some stats :","metadata":{}},{"cell_type":"code","source":"print(\"Nb of rows in study level file :\", sdf.shape[0])\nprint(\"Nb of rows in image level file :\", idf.shape[0])\n\nprint(\"Nb of StudyInstanceUID in image level file :\", \n        len(idf['StudyInstanceUID'].unique()))\n\nprint(\"Nb of null StudyInstanceUID in image level file :\",\n        len(idf[idf['StudyInstanceUID'].isna()]))\n\nprint(\"Nb of StudyInstanceUID in study level file :\", \n        len(sdf['id'].unique()))","metadata":{"execution":{"iopub.status.busy":"2021-05-26T05:00:58.595593Z","iopub.execute_input":"2021-05-26T05:00:58.596094Z","iopub.status.idle":"2021-05-26T05:00:58.609576Z","shell.execute_reply.started":"2021-05-26T05:00:58.596059Z","shell.execute_reply":"2021-05-26T05:00:58.608243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are less studyId than imageId, but the number of studies Id is the same in both files. Is that mean there are several imagesId for one study ? \nLet's check this out !","metadata":{}},{"cell_type":"markdown","source":"We first merge the two dataframes :","metadata":{}},{"cell_type":"code","source":"mdf = pd.merge(idf, sdf, \n               left_on=idf['StudyInstanceUID'], \n               right_on=sdf['id'].map(lambda x: x[:-6]))\nmdf.head(3)","metadata":{"execution":{"iopub.status.busy":"2021-05-26T05:01:01.925720Z","iopub.execute_input":"2021-05-26T05:01:01.926077Z","iopub.status.idle":"2021-05-26T05:01:01.950572Z","shell.execute_reply.started":"2021-05-26T05:01:01.926048Z","shell.execute_reply":"2021-05-26T05:01:01.949455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Then, in order to check the result of the join that we just realised, we verify the presence of null values in right part of merged dataframe :","metadata":{}},{"cell_type":"code","source":"print(\"Study id without image correspondence : \", mdf[mdf['id_y'].isna()].shape[0])","metadata":{"execution":{"iopub.status.busy":"2021-05-26T05:01:05.105514Z","iopub.execute_input":"2021-05-26T05:01:05.105925Z","iopub.status.idle":"2021-05-26T05:01:05.114293Z","shell.execute_reply.started":"2021-05-26T05:01:05.105893Z","shell.execute_reply":"2021-05-26T05:01:05.113399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Everything seems to be ok, we do some cleaning in the dataframe resulting from the join.","metadata":{}},{"cell_type":"code","source":"mdf = mdf.drop(['key_0', 'id_y'], axis = 1)\nmdf = mdf.rename(columns={'id_x': 'imageID',\n                          'Negative for Pneumonia': 'negative',\n                          'Typical Appearance': 'typical',\n                          'Indeterminate Appearance': 'indeterminate',\n                          'Atypical Appearance': 'atypical'})\nmdf.head(3)","metadata":{"execution":{"iopub.status.busy":"2021-05-26T05:01:16.981378Z","iopub.execute_input":"2021-05-26T05:01:16.981811Z","iopub.status.idle":"2021-05-26T05:01:16.997708Z","shell.execute_reply.started":"2021-05-26T05:01:16.981775Z","shell.execute_reply":"2021-05-26T05:01:16.996728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And finally we calculate the number of studyID assigned to one image, then two, etc ...","metadata":{}},{"cell_type":"code","source":"cnt = Counter(mdf.groupby('StudyInstanceUID')['imageID'].nunique())\nutotal = 0\nbtotal = 0\nfor key in cnt:\n    print('Nb of studies associated to {} images : {}'.format(key, cnt[key]))\n    utotal += cnt[key]\n    btotal += key * cnt[key]\nprint(\"\\nTotal of uniques studies : \", utotal)\nprint(\"Total of studies : \", btotal)\n","metadata":{"execution":{"iopub.status.busy":"2021-05-26T05:01:20.691406Z","iopub.execute_input":"2021-05-26T05:01:20.691819Z","iopub.status.idle":"2021-05-26T05:01:20.709816Z","shell.execute_reply.started":"2021-05-26T05:01:20.691787Z","shell.execute_reply":"2021-05-26T05:01:20.708185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most studies are associated with a single image, but there are a few cases where the same study can be associated with 2 or more images. We will see this in more detail in the exploratory analysis.","metadata":{}},{"cell_type":"markdown","source":"### *2. EXPLORATORY DATA ANALYSIS*","metadata":{}},{"cell_type":"markdown","source":"To understand what we are talking about let's take an image and then compare the values of label and boxes fields.","metadata":{}},{"cell_type":"code","source":"print(\"LABEL : \", mdf[mdf['imageID']=='000a312787f2_image'].label[0])\nprint(\"\\nBOXES : \",mdf[mdf['imageID']=='000a312787f2_image'].boxes[0])","metadata":{"execution":{"iopub.status.busy":"2021-05-26T05:01:24.655669Z","iopub.execute_input":"2021-05-26T05:01:24.656020Z","iopub.status.idle":"2021-05-26T05:01:24.665555Z","shell.execute_reply.started":"2021-05-26T05:01:24.655991Z","shell.execute_reply":"2021-05-26T05:01:24.664454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see several things :\n- more than one box can be associated to an image, \n- boxes's coordinates are listed into label,except that in place of (x y w h) format, coordinates indicated in label are formated as (x1 y1 x2 y2).\n- finally, diagnostic is not present in label, so we have to consider that opacity term means for us to search for the corresponding class.","metadata":{}},{"cell_type":"markdown","source":"Last check to perform is to verify if it can be more than one diagnostic per image :","metadata":{}},{"cell_type":"code","source":"print('Nb of images with more than one diagnostic : ', \n        len(mdf[mdf['negative']+mdf['typical']+mdf['indeterminate']+mdf['atypical']>1]))","metadata":{"execution":{"iopub.status.busy":"2021-05-26T05:01:29.095434Z","iopub.execute_input":"2021-05-26T05:01:29.095777Z","iopub.status.idle":"2021-05-26T05:01:29.110735Z","shell.execute_reply.started":"2021-05-26T05:01:29.095748Z","shell.execute_reply":"2021-05-26T05:01:29.109253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems that we have all we need to begin to plot some graphics.\n<br>Let's add some informations : the diagnosis as categorical variable and the number of spot (understand boxes).","metadata":{}},{"cell_type":"code","source":"mdf['nbSpot'] = mdf['label'].apply(lambda lab: lab.count('opacity'))\nmdf['diagnosis'] = mdf[['negative','typical','indeterminate','atypical']].idxmax(axis=1)","metadata":{"execution":{"iopub.status.busy":"2021-05-26T05:01:32.745347Z","iopub.execute_input":"2021-05-26T05:01:32.745758Z","iopub.status.idle":"2021-05-26T05:01:32.764859Z","shell.execute_reply.started":"2021-05-26T05:01:32.745722Z","shell.execute_reply":"2021-05-26T05:01:32.763797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Graphs serie 1 : Diagnosis repartition**","metadata":{}},{"cell_type":"code","source":"diagnosisColors = ['b','r','g','y']\n\nplt.figure(figsize = (14, 6))\n\nplt.suptitle('Diagnosis repartition')\nplt.subplot(1, 2, 1)\n\nsns.set();\nax=sns.countplot(x = mdf['diagnosis'].sort_values(), palette = diagnosisColors)\n\nfor p in ax.patches:\n        ax.annotate('{}'.format(p.get_height()), \n                    (p.get_x()+0.2, p.get_height()+50))\n\nplt.subplot(1, 2, 2)\n\nmdf['diagnosis'].value_counts(normalize=True).sort_index().plot(kind='pie', \n                                                 colors = diagnosisColors, \n                                                 explode = [0.025, 0.025, 0.025, 0.2],\n                                                 autopct = lambda x : str(round(x, 2)) + '%')\n\nplt.show()\n\n","metadata":{"execution":{"iopub.status.busy":"2021-05-26T05:05:28.090705Z","iopub.execute_input":"2021-05-26T05:05:28.091055Z","iopub.status.idle":"2021-05-26T05:05:28.383900Z","shell.execute_reply.started":"2021-05-26T05:05:28.091026Z","shell.execute_reply":"2021-05-26T05:05:28.382821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can already see that the \"typical\" diagnosis represents almost half of the dataset. Conversely, the \"atypical\" diagnosis represents only 7.63%.","metadata":{}},{"cell_type":"code","source":"def plotDiagnosis(ax, text, color):\n    ax=sns.countplot(x = mdf[mdf[text]==1].nbSpot, color=color)\n\n    for p in ax.patches:\n        ax.annotate('{}'.format(p.get_height()), \n                    (p.get_x()+0.2, p.get_height()+3))\n    plt.ylabel('')\n    plt.xlabel('nb ' + text + ' spots per image')\n    \nfig = plt.figure(figsize=(15,4))\n\nplt.subplot(1, 3, 1)\nplotDiagnosis(ax, 'atypical', diagnosisColors[0])\n\nplt.subplot(1, 3, 2)\nplotDiagnosis(ax, 'indeterminate', diagnosisColors[1])\n\nplt.subplot(1, 3, 3)\nplotDiagnosis(ax, 'typical', diagnosisColors[3])\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-05-26T05:12:19.346412Z","iopub.execute_input":"2021-05-26T05:12:19.346884Z","iopub.status.idle":"2021-05-26T05:12:19.852196Z","shell.execute_reply.started":"2021-05-26T05:12:19.346846Z","shell.execute_reply":"2021-05-26T05:12:19.851014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\"atypical\" and \"indeterminate\" diagnosis follow more or less the same pattern: a majority of images have 1 single identified spot, then come the images with 2 spots.\nFor the \"typical\" diagnosis, on the other hand, the vast majority of images have 2 spots.","metadata":{}}]}