{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<br>\n<h1 style = \"font-size:60px; font-family:Garamond ; font-weight : normal; background-color: #f6f5f5 ; color : #fe346e; text-align: center; border-radius: 200px 200px;\"> SIIM COVID-19 Detection: Complete EDA   <br> Exploratory Data Analysis 🧐 & Modeling</h1>\n<br>","metadata":{}},{"cell_type":"markdown","source":"# Objective \nIn this competition, we are identifying and localizing COVID-19 abnormalities on chest radiographs. This is an object detection and classification problem.\n\nFor each test image, you will be predicting a bounding box and class for all findings. If you predict that there are no findings, you should create a prediction of \"none 1 0 0 1 1\" (\"none\" is the class ID for no finding, and this provides a one-pixel bounding box with a confidence of 1.0).\n\nFurther, for each test study, you should make a determination within the following labels:\n\n'Negative for Pneumonia' 'Typical Appearance' 'Indeterminate Appearance' 'Atypical Appearance'\n\nTo make a prediction of one of the above labels, create a prediction string similar to the \"none\" class above: e.g. atypical 1 0 0 1 1\n\nPlease see the Evaluation page for more details about formatting predictions.\n\nThe images are in DICOM format, which means they contain additional data that might be useful for visualizing and classifying.\n\n# Dataset information\n\nThe train dataset comprises 6,334 chest scans in DICOM format, which were de-identified to protect patient privacy. All images were labeled by a panel of experienced radiologists for the presence of opacities as well as overall appearance.\n\nNote that all images are stored in paths with the form study/series/image. The study ID here relates directly to the study-level predictions, and the image ID is the ID used for image-level predictions.\n\nThe hidden test dataset is of roughly the same scale as the training dataset.\n\n### Files\n* train_study_level.csv - the train study-level metadata, with one row for each study, including correct labels.\n* train_image_level.csv - the train image-level metadata, with one row for each image, including both correct labels and any bounding boxes in a dictionary format. Some images in both test and train have multiple bounding boxes.\n* sample_submission.csv - a sample submission file containing all image- and study-level IDs.\n\n### Columns\n#### train_study_level.csv\n\n* id - unique study identifier\n* Negative for Pneumonia - 1 if the study is negative for pneumonia, 0 otherwise\n* Typical Appearance - 1 if the study has this appearance, 0 otherwise\n* Indeterminate Appearance  - 1 if the study has this appearance, 0 otherwise\n* Atypical Appearance  - 1 if the study has this appearance, 0 otherwise\n\n#### train_image_level.csv\n\n* id - unique image identifier\n* boxes - bounding boxes in easily-readable dictionary format\n* label - the correct prediction label for the provided bounding boxes","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"!pip install -q sweetviz\n!pip install -q klib","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n\nimport matplotlib.pyplot as plt\nimport matplotlib\nimport seaborn as sns\nimport pydicom as dicom\nimport cv2\nimport ast\n\n\nimport plotly.express as px\nimport plotly.graph_objects as go\nimport plotly.figure_factory as ff\nfrom plotly.subplots import make_subplots\nfrom plotly.offline import init_notebook_mode, iplot\ninit_notebook_mode(connected=True)\n\n\nimport sweetviz\nimport klib\n\n\nimport tensorflow as tf\n\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"AUTO = tf.data.experimental.AUTOTUNE\n\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()  \n    print('Running on TPU ', tpu.master())\nexcept ValueError:\n    tpu = None\n\nif tpu:\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.experimental.TPUStrategy(tpu)\nelse:\n    strategy = tf.distribute.get_strategy()\n\nprint(\"REPLICAS: \", strategy.num_replicas_in_sync)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"PATH='/kaggle/input/siim-covid19-detection/'\nTRAIN_PATH = \"../input/siim-covid19-detection/train\"\n#TEST_PATH = GCS_DS_PATH + \"/test\"\nTEST_PATH = \"../input/siim-covid19-detection/test\"\nTRAIN_FILES = tf.io.gfile.glob(TRAIN_PATH+\"/*/*/*.dcm\")\nTEST_FILES = tf.io.gfile.glob(TEST_PATH+\"/*/*/*.dcm\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_img=pd.read_csv(PATH+'train_image_level.csv')\ndf_train_study=pd.read_csv(PATH+'train_study_level.csv')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"classes_dict = {\n    0 : \"Negative for Pneumonia\",\n    1  : \"Typical Appearance\",\n    2  : \"Indeterminate Appearance\",\n    3  : \"Atypical Appearance\"\n}","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#getting filepath from study_id or image_id\ndef get_path(file_id,main_path,id_type):\n    name = file_id.split(\"_\")[0]\n    if id_type == \"study\":\n        path = tf.io.gfile.glob(main_path+f\"/{name}/*/*.dcm\")[0]\n    else:\n        path = tf.io.gfile.glob(main_path+f\"/*/*/{name}.dcm\")[0]\n    return path","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"df_train_img.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Total Images in directory for model training :- ',len(df_train_img['id'].unique()))\nprint('Total Images which does not have Pneumonia   :- ',df_train_img[df_train_img['boxes'].isnull()].shape[0])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_img.loc[0, 'StudyInstanceUID']","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_img","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path_train = PATH+'/train/'+df_train_img.loc[0, 'StudyInstanceUID']+'/'+'81456c9c5423'+'/'\nimg_id = df_train_img.loc[0, 'id'].replace('_image', '.dcm')\ndata_file = dicom.dcmread(path_train+img_id)\nimg = data_file.pixel_array","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## What is there in DICOM metadata file","metadata":{}},{"cell_type":"markdown","source":"Digital Imaging and Communications in Medicine (DICOM) is the standard for the communication and management of medical imaging information and related data.DICOM is most commonly used for storing and transmitting medical images enabling the integration of medical imaging devices such as scanners, servers, workstations, printers, network hardware, and picture archiving and communication systems (PACS) from multiple manufacturers. It has been widely adopted by hospitals and is making inroads into smaller applications like dentists' and doctors' offices.\n\nDICOM files can be exchanged between two entities that are capable of receiving image and patient data in DICOM format. The different devices come with DICOM Conformance Statements which state which DICOM classes they support. The standard includes a file format definition and a network communications protocol that uses TCP/IP to communicate between systems.\n\nThe National Electrical Manufacturers Association (NEMA) holds the copyright to the published standard which was developed by the DICOM Standards Committee, whose members are also partly members of NEMA.It is also known as NEMA standard PS3, and as ISO standard 12052:2017 \"Health informatics -- Digital imaging and communication in medicine (DICOM) including workflow and data management\".","metadata":{}},{"cell_type":"code","source":"data_file","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Shape of the Image","metadata":{}},{"cell_type":"code","source":"print('Image shape:', img.shape)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"boxes = ast.literal_eval(df_train_img.loc[0, 'boxes'])\nboxes","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 1, figsize=(20, 4))\n\nfor box in boxes:\n    p = matplotlib.patches.Rectangle((box['x'], box['y']), box['width'], box['height'],\n                                     ec='r', fc='none', lw=2.)\n    ax.add_patch(p)\nax.imshow(img, cmap='gray')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Let see more samples","metadata":{}},{"cell_type":"code","source":"fig, axs = plt.subplots(3, 3, figsize=(20, 20))\nfig.subplots_adjust(hspace = .1, wspace=.1)\naxs = axs.ravel()\n\nfor row in range(9):\n    study = df_train_img.loc[row, 'StudyInstanceUID']\n    path_in = PATH+'train/'+study+'/'\n    folder = os.listdir(path_in)\n    path_file = path_in+folder[0]\n    filename = os.listdir(path_file)[0]\n    file_id = filename.split('.')[0]\n    \n    data_file = dicom.dcmread(path_file+'/'+file_id+'.dcm')\n    img = data_file.pixel_array\n    if (df_train_img.loc[row, 'boxes']!=df_train_img.loc[row, 'boxes']) == False:\n        boxes = ast.literal_eval(df_train_img.loc[row, 'boxes'])\n    \n        for box in boxes:\n            p = matplotlib.patches.Rectangle((box['x'], box['y']), box['width'], box['height'],\n                                     ec='r', fc='none', lw=2.)\n            axs[row].add_patch(p)\n    axs[row].imshow(img, cmap='gray')\n    axs[row].set_title(df_train_img.loc[row, 'label'].split(' ')[0])\n    axs[row].set_xticklabels([])\n    axs[row].set_yticklabels([])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def split_label(s):\n    return s.split(' ')[0]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_img['label_name'] = df_train_img['label'].apply(split_label)\ndf_train_img['label_name'].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_study","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"klib.missingval_plot(df_train_study)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Distribution of Pneumonia Symptoms","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(2,2,figsize=(20,16))\nsns.kdeplot(df_train_study[\"Negative for Pneumonia\"], shade=True,ax=ax[0,0],color=\"#ffb4a2\")\nax[0,0].set_title(\"Negative for Pneumonia Distribution\",font=\"Serif\", fontsize=15)\nsns.kdeplot(df_train_study[\"Typical Appearance\"], shade=True,ax=ax[0,1],color=\"#e5989b\")\nax[0,1].set_title(\"Typical Appearance Distribution\",font=\"Serif\", fontsize=15)\nsns.kdeplot(df_train_study[\"Indeterminate Appearance\"], shade=True,ax=ax[1,0],color=\"#b5838d\")\nax[1,0].set_title(\"Indeterminate Appearance Distribution\",font=\"Serif\", fontsize=15)\nsns.kdeplot(df_train_study[\"Atypical Appearance\"], shade=True,ax=ax[1,1],color=\"#6d6875\")\nax[1,1].set_title(\"Atypical Appearance Distribution\",font=\"Serif\", fontsize=15)\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_img.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_study['StudyInstanceUID'] = df_train_study['id'].apply(lambda x: x.replace('_study', ''))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_img = df_train_img.merge(df_train_study[['Negative for Pneumonia', 'Typical Appearance','Indeterminate Appearance', 'Atypical Appearance','StudyInstanceUID']], on='StudyInstanceUID')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_img","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Parallel categories plot of targets","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(35,20))\nfig = px.parallel_categories(df_train_img[['Negative for Pneumonia', 'Typical Appearance',\n       'Indeterminate Appearance', 'Atypical Appearance']], color=\"Negative for Pneumonia\", color_continuous_scale=\"sunset\",\\\n                             title=\"Parallel categories plot of targets\")\nfig","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Observation :- \n    Negative for Pneumonia have share of other symptomes ","metadata":{}},{"cell_type":"code","source":"df_train = df_train_img.copy()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#converting into one-hot label\ndf_train[\"one_hot\"] = df_train.apply(lambda x : np.array([x[\"Negative for Pneumonia\"],\n                                                        x[\"Typical Appearance\"],\n                                                        x[\"Indeterminate Appearance\"],\n                                                        x[\"Atypical Appearance\"]]),axis=1)\n\ndf_train = df_train.drop([\"Negative for Pneumonia\",\"Typical Appearance\",\"Indeterminate Appearance\",\"Atypical Appearance\"],axis=1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train[\"label_id\"] = df_train[\"one_hot\"].map(lambda x : classes_dict[np.argmax(x)])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### .... Inprogress ","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}}]}