{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"!conda install -c conda-forge gdcm -y\n\nimport numpy as np \nimport pandas as pd \nimport os\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport matplotlib.patches as patches\n%matplotlib inline\nimport glob\nfrom pydicom import dcmread\nfrom pydicom.data import get_testdata_file\n\nfrom tqdm import tqdm\n\nimport ast\n\n!pip install hvplot\nimport hvplot.pandas \n\n!pip install pylibjpeg pylibjpeg-libjpeg pydicom python-gdcm\nimport gdcm\nimport pylibjpeg","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-06-20T13:09:03.324552Z","iopub.execute_input":"2021-06-20T13:09:03.325071Z","iopub.status.idle":"2021-06-20T13:09:49.280193Z","shell.execute_reply.started":"2021-06-20T13:09:03.325025Z","shell.execute_reply":"2021-06-20T13:09:49.278911Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#        *****     INTRODUCTION  ******\n\nCOVID - 19 is a mild infection resulting in inflammation and fluid in lungs **. This disease is similar to other pneumonias, which makes it difficult to diagnose. The idea of the project is to **detect and localize COVID-19** in order to help doctors to provide a quick and confident diagnosis. This will allow them to get the right treatment before severe effects of the virus.\n\nThe reason for using chest X-rays is that they are very easy to take and obtained in minutes. In addition, we can locate the disease and get a better idea of the patient's condition.\n\n\n    ****    OBJECTIVES   ***** \n\nIn this notebook, we will **Visualise and Understand ** the data from a dataset X-rays. The objective is to get an overview of the data and to see if some information can be gathered. Feel free to leave a comment if you have any questions, I would be happy to answer them :)\n\n> Link for the data : [SIIM-FISABIO-RSNA COVID-19 Detection](https://www.kaggle.com/c/siim-covid19-detection)","metadata":{"execution":{"iopub.status.busy":"2021-06-20T08:42:13.03224Z","iopub.execute_input":"2021-06-20T08:42:13.032608Z","iopub.status.idle":"2021-06-20T08:42:13.04181Z","shell.execute_reply.started":"2021-06-20T08:42:13.032575Z","shell.execute_reply":"2021-06-20T08:42:13.040778Z"}}},{"cell_type":"code","source":"# Read the data\n\ndf_train_images = pd.read_csv('../input/siim-covid19-detection/train_image_level.csv')\ndf_train_study = pd.read_csv('../input/siim-covid19-detection/train_study_level.csv')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-20T09:45:37.674022Z","iopub.execute_input":"2021-06-20T09:45:37.674312Z","iopub.status.idle":"2021-06-20T09:45:37.719955Z","shell.execute_reply.started":"2021-06-20T09:45:37.67428Z","shell.execute_reply":"2021-06-20T09:45:37.719086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# In this dataset, we have two important files :\n\n- *train_study_level.csv* : We will find all the study information as a corresponding label for the image.\n- *train_image_level.csv* : We will fing image information and the associated bounding boxes to locate the disease.\n\nThe train dataset comprimises **6,334 chest scans in DICOM format**, which were de-identified to protect patient privacy. All images were labeled by a panel of experienced radiologists for the presence of opacities as well as overall appearance.","metadata":{}},{"cell_type":"markdown","source":"# Analysis of the studies\n\nIn this section we will focus mainly on the analysis of the studies. We will try to understand the different categories and their distribution.","metadata":{}},{"cell_type":"code","source":"labels = ['Negative for Pneumonia', 'Typical Appearance', 'Indeterminate Appearance', 'Atypical Appearance']\n\nprint(\"Number of study : \", len(df_train_study))\n\ndf_train_study.head()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-20T09:45:37.72141Z","iopub.execute_input":"2021-06-20T09:45:37.721839Z","iopub.status.idle":"2021-06-20T09:45:37.73655Z","shell.execute_reply.started":"2021-06-20T09:45:37.721797Z","shell.execute_reply":"2021-06-20T09:45:37.735439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, in this file, we have 6054 rows. However, we have 6344 chest X-rays. So a study can group several images of a patient. One question that can be asked is why several images for one study. Are there different shots that have been taken? Was it taken in a different time frame? \n\nRegarding the diagnosis, we will have 4 labels :\n\n1. **Negative for Pneumonia**: No lung opacities.\n\n2. **Typical Appearance**: Multifocal bilateral, peripheral opacities with rounded morphology, lower lung–predominant distribution\n\n3. **Indeterminate Appearance**: Absence of typical findings AND unilateral, central or upper lung predominant distribution\n\n4. **Atypical Appearance**: Pneumothorax, pleural effusion, pulmonary edema, lobar consolidation, solitary lung nodule or mass, diffuse tiny nodules, cavity.\n\n\n> https://www.kaggle.com/c/siim-covid19-detection/discussion/240250\n\n\n**Pulmonary opacification represents the result of a decrease in the ratio of gas to soft tissue** (blood, lung parenchyma and stroma) in the lung. When reviewing an area of increased attenuation (opacification) on a chest radiograph or CT it is vital to determine where the opacification is. The patterns can broadly be divided into airspace opacification, lines and dots.\n\n> https://radiopaedia.org/articles/pulmonary-opacification","metadata":{}},{"cell_type":"markdown","source":"## See some sample images from the different categories","metadata":{}},{"cell_type":"code","source":"NUMBER_OF_SAMPLE = 5\n\ndef read_image_from_study(study_id):\n    study_name = study_id.split('_')[0]\n    file = glob.glob(\"../input/siim-covid19-detection/train/\" + study_name + \"/*/*.dcm\")\n    ds = dcmread(file[0])\n    return ds.pixel_array\n\ndef show_sample_data_from_study(sample_images, NB_SAMPLE = 5):\n    fig, axes = plt.subplots(nrows=1, ncols=NB_SAMPLE, figsize=(NB_SAMPLE * 4, 4))\n    i = 0\n    for index, row in sample_images.iterrows():\n        img = read_image_from_study(row['id'])\n        axes[i].imshow(img, cmap=plt.cm.gray, aspect='auto')\n        axes[i].axis('off')\n        i += 1\n    fig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-20T09:45:37.738597Z","iopub.execute_input":"2021-06-20T09:45:37.739116Z","iopub.status.idle":"2021-06-20T09:45:37.75055Z","shell.execute_reply.started":"2021-06-20T09:45:37.738988Z","shell.execute_reply":"2021-06-20T09:45:37.749373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Negative for pneumonia","metadata":{}},{"cell_type":"code","source":"sample_negative_pneumonia = df_train_study[df_train_study['Negative for Pneumonia'] == 1].sample(n=NUMBER_OF_SAMPLE, random_state=42)\nshow_sample_data_from_study(sample_negative_pneumonia, NUMBER_OF_SAMPLE)","metadata":{"execution":{"iopub.status.busy":"2021-06-20T09:45:37.752193Z","iopub.execute_input":"2021-06-20T09:45:37.752469Z","iopub.status.idle":"2021-06-20T09:45:42.266245Z","shell.execute_reply.started":"2021-06-20T09:45:37.752441Z","shell.execute_reply":"2021-06-20T09:45:42.265181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Typical appearance","metadata":{}},{"cell_type":"code","source":"sample_negative_pneumonia = df_train_study[df_train_study['Typical Appearance'] == 1].sample(n=NUMBER_OF_SAMPLE, random_state=42)\nshow_sample_data_from_study(sample_negative_pneumonia, NUMBER_OF_SAMPLE)","metadata":{"execution":{"iopub.status.busy":"2021-06-20T09:45:42.267718Z","iopub.execute_input":"2021-06-20T09:45:42.268288Z","iopub.status.idle":"2021-06-20T09:45:46.392214Z","shell.execute_reply.started":"2021-06-20T09:45:42.268239Z","shell.execute_reply":"2021-06-20T09:45:46.391136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Indeterminate appearance","metadata":{}},{"cell_type":"code","source":"sample_negative_pneumonia = df_train_study[df_train_study['Indeterminate Appearance'] == 1].sample(n=NUMBER_OF_SAMPLE, random_state=42)\nshow_sample_data_from_study(sample_negative_pneumonia, NUMBER_OF_SAMPLE)","metadata":{"execution":{"iopub.status.busy":"2021-06-20T09:45:46.393704Z","iopub.execute_input":"2021-06-20T09:45:46.394271Z","iopub.status.idle":"2021-06-20T09:45:50.469019Z","shell.execute_reply.started":"2021-06-20T09:45:46.394214Z","shell.execute_reply":"2021-06-20T09:45:50.467853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Atypical appearance","metadata":{}},{"cell_type":"code","source":"sample_negative_pneumonia = df_train_study[df_train_study['Atypical Appearance'] == 1].sample(n=NUMBER_OF_SAMPLE, random_state=42)\nshow_sample_data_from_study(sample_negative_pneumonia, NUMBER_OF_SAMPLE)","metadata":{"execution":{"iopub.status.busy":"2021-06-20T09:45:50.470468Z","iopub.execute_input":"2021-06-20T09:45:50.471245Z","iopub.status.idle":"2021-06-20T09:45:54.337627Z","shell.execute_reply.started":"2021-06-20T09:45:50.471194Z","shell.execute_reply":"2021-06-20T09:45:54.336745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"At a first glance, it's really complicated to really see the opacities. It's a difficult task.\nIf we focus a little bit, for the *Typical appearance*, we could see some opacities.\nWith a bigger picture, we can maybe have a better view.\n\nNote : I think that could be interesting to use images enhancement in order to help the visualisation of x-rays images. [I found on github a library called *X-Ray Images Enhancement* that could be interesting](https://github.com/asalmada/x-ray-images-enhancement). I will try it in another kernel. If someone already applies this kind of techniques or used that library, feel free to share your experience with us :)","metadata":{}},{"cell_type":"markdown","source":"## Distribution of the different categories","metadata":{}},{"cell_type":"code","source":"# Count for each labels the number of occurence\nstudy_case = [df_train_study[label].value_counts()[1] for label in labels]\n\nplt.figure(figsize=(15, 6))\nplt.bar(labels, study_case)\nplt.title('Distribution of the different categories')\nplt.show()\n\nplt.figure(figsize=(8, 8))\nplt.pie(study_case, labels=labels, autopct='%1.1f%%')\nplt.title('Proportion of the different categories')\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-20T09:45:54.339295Z","iopub.execute_input":"2021-06-20T09:45:54.340014Z","iopub.status.idle":"2021-06-20T09:45:54.607692Z","shell.execute_reply.started":"2021-06-20T09:45:54.339939Z","shell.execute_reply":"2021-06-20T09:45:54.606509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def count_column(x):\n    return x.sum()\n    \ndf_train_count = df_train_study[labels].apply(count_column, axis=1)\nprint(\"Number of multiple categories ?\", df_train_count[df_train_count != 1].sum())","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-20T09:45:54.60942Z","iopub.execute_input":"2021-06-20T09:45:54.609877Z","iopub.status.idle":"2021-06-20T09:45:54.945077Z","shell.execute_reply.started":"2021-06-20T09:45:54.609829Z","shell.execute_reply":"2021-06-20T09:45:54.943923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For the study case, we don't have multiple categories. Which means that **each categories are distinct**. \nRegarding the distribution of the data, **we have unbalanced categories**. As we can see in our diagram, we have 47% for *Typical Appearance*. And, regarding the *Atypical Appearance*, we have 7,8%. \n\nSo, when the preprocessing of our training dataset, we should take in consideration that we're dealing with unbalance data in order to avoid important prediction on unique label.","metadata":{}},{"cell_type":"markdown","source":"# Analysis of the images","metadata":{}},{"cell_type":"code","source":"print(\"Number of images : \", len(df_train_images))\ndf_train_images.head()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-20T09:45:54.946424Z","iopub.execute_input":"2021-06-20T09:45:54.946742Z","iopub.status.idle":"2021-06-20T09:45:54.960467Z","shell.execute_reply.started":"2021-06-20T09:45:54.946711Z","shell.execute_reply":"2021-06-20T09:45:54.959354Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#  EXPERIMENTAL SETUP *****\n","metadata":{}},{"cell_type":"markdown","source":"To recall, we had seen in the previous part that we have multiple images for a given study. It could be interesting to visualize those data in order to understand why we have multiple images.","metadata":{}},{"cell_type":"code","source":"print(\"Number of duplicate images :\", df_train_images.id.duplicated().sum())\nprint(\"Number of duplicate study :\", df_train_images.StudyInstanceUID.duplicated().sum())","metadata":{"execution":{"iopub.status.busy":"2021-06-20T09:45:54.961919Z","iopub.execute_input":"2021-06-20T09:45:54.962297Z","iopub.status.idle":"2021-06-20T09:45:54.977363Z","shell.execute_reply.started":"2021-06-20T09:45:54.962264Z","shell.execute_reply":"2021-06-20T09:45:54.975998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Count the number of duplicated images\nunique_study_duplicate = df_train_images[df_train_images.StudyInstanceUID.duplicated()].StudyInstanceUID.unique()\nprint(\"Some duplicated id : \", ' ; '.join(unique_study_duplicate[:10]))\n\nimages_with_duplicate_study = df_train_images[df_train_images.StudyInstanceUID.isin(unique_study_duplicate)]\nprint(\"Number of image concernd with duplication :\", len(images_with_duplicate_study))","metadata":{"execution":{"iopub.status.busy":"2021-06-20T09:45:54.979112Z","iopub.execute_input":"2021-06-20T09:45:54.979536Z","iopub.status.idle":"2021-06-20T09:45:54.995594Z","shell.execute_reply.started":"2021-06-20T09:45:54.979492Z","shell.execute_reply":"2021-06-20T09:45:54.994516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Visualization of some studies","metadata":{}},{"cell_type":"code","source":"def read_image_from_image(study_name, image_id):\n    image_name = image_id.split('_')[0]\n    file = glob.glob(\"../input/siim-covid19-detection/train/\" + study_name + \"/*/\" + image_name + \".dcm\")\n    ds = dcmread(file[0])\n    return ds.pixel_array\n\ndef show_sample_duplicate(samples):\n    nb_show_sample = min(5, len(samples))\n    fig, axes = plt.subplots(nrows=1, ncols=nb_show_sample, figsize=(nb_show_sample * 4, 4))\n    i = 0\n    for index, row in samples.iterrows():\n        img = read_image_from_image(row['StudyInstanceUID'], row['id'])\n        axes[i].imshow(img, cmap=plt.cm.gray, aspect='auto')\n        axes[i].axis('off')\n        i += 1\n        if i == 5:\n            break\n        \n    fig.suptitle(samples.StudyInstanceUID.unique()[0], fontsize=20)\n    fig.show()\n    \n\n# Get some sample from duplicate study\nnp.random.seed(42)\nduplicated_study_sample = np.random.choice(unique_study_duplicate, 5)\n\n# See the different values\nfor sample_study_name in duplicated_study_sample:\n    sample_duplicate_image = df_train_images[df_train_images.StudyInstanceUID == sample_study_name]\n    \n    show_sample_duplicate(sample_duplicate_image)    ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-20T09:45:54.997051Z","iopub.execute_input":"2021-06-20T09:45:54.997354Z","iopub.status.idle":"2021-06-20T09:46:06.375533Z","shell.execute_reply.started":"2021-06-20T09:45:54.997323Z","shell.execute_reply":"2021-06-20T09:46:06.374383Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems that for the images present for a given study, there are duplicated but with a different quality. If we took the first, the second and the last, we clearly have different brightness. Nevertheless the fourth seems to be the same. Finally, the third one is 4 different images with different brightness and cropping.","metadata":{}},{"cell_type":"markdown","source":"  ####   THE PARTICULAR DETAILS ARE BELOW ###\n\nDuring my research, I found a particular ID, which have duplicated images.","metadata":{}},{"cell_type":"code","source":"df_train_images[df_train_images.StudyInstanceUID == \"0fd2db233deb\"]","metadata":{"execution":{"iopub.status.busy":"2021-06-20T09:46:06.376976Z","iopub.execute_input":"2021-06-20T09:46:06.377284Z","iopub.status.idle":"2021-06-20T09:46:06.39285Z","shell.execute_reply.started":"2021-06-20T09:46:06.377254Z","shell.execute_reply":"2021-06-20T09:46:06.391829Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_study[df_train_study.id == \"0fd2db233deb_study\"]","metadata":{"execution":{"iopub.status.busy":"2021-06-20T09:46:06.39441Z","iopub.execute_input":"2021-06-20T09:46:06.394759Z","iopub.status.idle":"2021-06-20T09:46:06.412801Z","shell.execute_reply.started":"2021-06-20T09:46:06.394721Z","shell.execute_reply":"2021-06-20T09:46:06.411507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show_sample_duplicate(df_train_images[df_train_images.StudyInstanceUID == \"0fd2db233deb\"])","metadata":{"execution":{"iopub.status.busy":"2021-06-20T09:46:06.414442Z","iopub.execute_input":"2021-06-20T09:46:06.414793Z","iopub.status.idle":"2021-06-20T09:46:10.817401Z","shell.execute_reply.started":"2021-06-20T09:46:06.414761Z","shell.execute_reply":"2021-06-20T09:46:10.816128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Regarding the 0fd2db233deb ID, we have duplicated images. Moreover, regarding the image information, we have a box information for a unique rows. \n\nSo, in our dataset, we have **duplicate images**. Some are different (with different brigthness, different cropping, different angle) and some are the same. They represent 512 images of our dataset. It's represent about 8% (512 * 100 / 6334) of our dataset. ","metadata":{}},{"cell_type":"code","source":"# Rename the 'StudyInstanceUID' column\ndf_train_study['StudyInstanceUID'] = df_train_study['id'].apply(lambda x : x.replace('_study', ''))\n\n# Get the duplicated study\ndf_study_from_duplicate = df_train_study[df_train_study['StudyInstanceUID'].isin(images_with_duplicate_study['StudyInstanceUID'].unique())]\n\n# Get the duplicated images\ndf_image_from_duplicate = df_train_images[df_train_images.StudyInstanceUID.isin(unique_study_duplicate)]\n\n\n# Count for each category the number of duplicated study\nduplicate_study_case = [df_study_from_duplicate[label].value_counts()[1] for label in labels]\ntotal_study_case = [df_train_study[label].value_counts()[1] for label in labels]\n\n# Get the percentage for each category\nratio_duplicate = [x / y for x, y in zip(duplicate_study_case, total_study_case)] \n\nprint(\"Ratio total duplicated image : \", len(df_image_from_duplicate) / len(df_train_images))\nprint(\"Ratio total duplicated study : \", sum(duplicate_study_case) / sum(total_study_case))\n\nprint()\nprint(\"Percentage of duplicated study for each category :\")\nprint() \n\nfor i in range(len(ratio_duplicate)):\n    print(labels[i], \" : \", ratio_duplicate[i])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-20T09:46:10.819099Z","iopub.execute_input":"2021-06-20T09:46:10.819512Z","iopub.status.idle":"2021-06-20T09:46:10.850114Z","shell.execute_reply.started":"2021-06-20T09:46:10.819468Z","shell.execute_reply":"2021-06-20T09:46:10.848723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(8, 8))\nplt.pie(ratio_duplicate, labels=labels, autopct='%1.1f%%', normalize=True)\nplt.suptitle(\"Distribution of duplications for each category\", fontsize=20)\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-20T09:46:10.851691Z","iopub.execute_input":"2021-06-20T09:46:10.852017Z","iopub.status.idle":"2021-06-20T09:46:10.983544Z","shell.execute_reply.started":"2021-06-20T09:46:10.851985Z","shell.execute_reply":"2021-06-20T09:46:10.98244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this section, we saw the different duplicated study and images. Those could be x-ray images that could be retaken, maybe duplicated from copy/past or even images that have been analyze multiple times. Radiography analisys is a complex task, and some errors are possible even for the most brillant doctor. So we should keep in mind that maybe we could have error in our dataset.\n\nNevertheless, concerning the application of the duplicate images, multiple possibilites are available. \n- The simplest solution is to decide to ignore these files. As this is a small percentage of our dataset, this might be feasible. With this, we could avoid the duplication problem.\n\n- The other possibility is that we could decide to get some of the data. I mean not all the data have to be throws away. Some of them are duplicate files. For them, it would be good if we could analyse the group of images and keep only the best ones, with all the metadata information collected from the others. We could do the same for other similar images, those with a different brightness and cropping.","metadata":{}},{"cell_type":"markdown","source":"## See some image with their boxes","metadata":{}},{"cell_type":"code","source":"sample_images_with_boxes = df_train_images[df_train_images.boxes.notna()].sample(n=10, random_state=42)\nsample_images_with_boxes.head()","metadata":{"execution":{"iopub.status.busy":"2021-06-20T09:46:10.984872Z","iopub.execute_input":"2021-06-20T09:46:10.985174Z","iopub.status.idle":"2021-06-20T09:46:11.002263Z","shell.execute_reply.started":"2021-06-20T09:46:10.985142Z","shell.execute_reply":"2021-06-20T09:46:11.001529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axes = plt.subplots(nrows=2, ncols=5, figsize=(5 * 4, 4 * 2))\n\ni = 0\n# Iterate through the sample\nfor index, row in sample_images_with_boxes.iterrows():\n    # Read and show image\n    img = read_image_from_image(row['StudyInstanceUID'], row['id'])\n    axes[i // NUMBER_OF_SAMPLE, i % NUMBER_OF_SAMPLE].imshow(img, cmap=plt.cm.gray, aspect='auto')\n    \n    # The boxes are saved as str, we need to translate them to array of dict\n    array_boxes = ast.literal_eval(row.boxes) \n    \n    # Now, show the boxes\n    for box in array_boxes:\n        rect = patches.Rectangle((box['x'], box['y']),\n                                 box['width'], \n                                 box['height'], \n                                 edgecolor='r', \n                                 facecolor=\"none\")\n        \n        axes[i // NUMBER_OF_SAMPLE, i % NUMBER_OF_SAMPLE].add_patch(rect)\n    \n    # Remove axis information\n    axes[i // NUMBER_OF_SAMPLE, i % NUMBER_OF_SAMPLE].axis('off')\n    i += 1","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-20T09:46:11.003185Z","iopub.execute_input":"2021-06-20T09:46:11.003458Z","iopub.status.idle":"2021-06-20T09:46:18.321618Z","shell.execute_reply.started":"2021-06-20T09:46:11.003422Z","shell.execute_reply":"2021-06-20T09:46:18.320555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With these images, we can first see that the images in this sample are really different. If we take the second one, I can barely see the content (and I have to turn the brightness of my screen to the maximum!). The fourth one is also interesting because the image has been rotated and cropped. In this sample we can really see the different image contrasts we have.\n\n\nRegarding the boxes, on this sample, we see that we usually have two boxes and are places on the left and on the right. Moreover, they are mainly between the inferior and the middle lobe. However, this is a sample, we cannot make generalization on this small amount of data.\n\n\n<img src=\"https://cdn.lecturio.com/assets/Lobes-and-fissures-of-the-lungs-1200x570.jpg\" width=\"800\" />\n\nCredit : https://www.lecturio.com/concepts/lungs/ - Image by Lecturio.","metadata":{}},{"cell_type":"markdown","source":"### ALGORITHM OR SIMULATIONS ###","metadata":{}},{"cell_type":"code","source":"sample_images_with_boxes = df_train_images[df_train_images.boxes.notna()]\nbox_size = pd.DataFrame()\n\nfor boxes in sample_images_with_boxes.boxes:\n    array_boxes = ast.literal_eval(boxes) \n    for box in array_boxes:\n        box_size = box_size.append(box, ignore_index=True)\n\nbox_size.head()","metadata":{"execution":{"iopub.status.busy":"2021-06-20T09:46:18.32327Z","iopub.execute_input":"2021-06-20T09:46:18.32368Z","iopub.status.idle":"2021-06-20T09:46:33.286775Z","shell.execute_reply.started":"2021-06-20T09:46:18.323618Z","shell.execute_reply":"2021-06-20T09:46:33.286033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"box_size.describe()","metadata":{"execution":{"iopub.status.busy":"2021-06-20T09:46:33.287886Z","iopub.execute_input":"2021-06-20T09:46:33.28836Z","iopub.status.idle":"2021-06-20T09:46:33.314087Z","shell.execute_reply.started":"2021-06-20T09:46:33.288307Z","shell.execute_reply":"2021-06-20T09:46:33.313303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show image size\nsizes = box_size.groupby(['height', 'width']).size().reset_index().rename(columns={0 : 'count'})\nsizes.hvplot.scatter(\n    x='height', \n    y='width', \n    size='count',\n    title='Box size distribution',\n    xlim=(0,3141), ylim=(0,1920), \n    grid=True, \n    height=500, width=1000).options(scaling_factor=0.1, line_alpha=1, fill_alpha=0)","metadata":{"execution":{"iopub.status.busy":"2021-06-20T09:46:33.315315Z","iopub.execute_input":"2021-06-20T09:46:33.315769Z","iopub.status.idle":"2021-06-20T09:46:33.455739Z","shell.execute_reply.started":"2021-06-20T09:46:33.315722Z","shell.execute_reply":"2021-06-20T09:46:33.455033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Regarding the size of the boxes, they seem to have similar shapes, but very variable sizes. ","metadata":{}},{"cell_type":"markdown","source":"#### Annotation label\n\nFor each box, we have an *opacity* tag. The question we might ask is whether we have any other tags.","metadata":{}},{"cell_type":"code","source":"o = []\nfor label in df_train_images.label.values:\n    a = label.split(' ')\n    o.append(a[0])\n    \npd.Series(o).value_counts()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-20T09:46:33.456785Z","iopub.execute_input":"2021-06-20T09:46:33.457213Z","iopub.status.idle":"2021-06-20T09:46:33.472528Z","shell.execute_reply.started":"2021-06-20T09:46:33.457169Z","shell.execute_reply":"2021-06-20T09:46:33.471861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here, we know we have only two tags for boxes : *none* or *opacity*","metadata":{}},{"cell_type":"markdown","source":"# DICOM image metadata analysis\n\nThe x-ray images are stored using a DICOM (*Digital Imaging and Communication in Medicine*)is the standard for digital files created during medical imaging examinations. It also covers the specifications concerning their archiving and their transmission over a network (particularly important aspects in the medical field). Independent of technologies (scanner, MRI, etc.) and manufacturers, it allows standardised access to medical imaging results. In addition to the digital images from medical examinations, DICOM files also carry a lot of textual information about the patient (marital status, age, weight, etc.), the examination carried out (region explored, imaging technique used, etc.), the acquisition date, the practitioner, etc.\n\n> https://sti-biotechnologies-pedagogie.web.ac-grenoble.fr/content/fichiers-dicom-format-dcm-en-imagerie-medicale\n\nBy analysing these files, we might be able to find interesting points that we could exploit.","metadata":{}},{"cell_type":"code","source":"# Merge the two dataframe\ndf_merged_data = df_train_study.merge(df_train_images, on=\"StudyInstanceUID\")","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-06-20T09:46:33.479269Z","iopub.execute_input":"2021-06-20T09:46:33.479678Z","iopub.status.idle":"2021-06-20T09:46:33.497244Z","shell.execute_reply.started":"2021-06-20T09:46:33.479619Z","shell.execute_reply":"2021-06-20T09:46:33.49649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualize some DICOM image","metadata":{}},{"cell_type":"code","source":"path = '../input/siim-covid19-detection/train/00086460a852/9e8302230c91/65761e66de9f.dcm'\nds = dcmread(path)\nprint(ds)\nplt.imshow(ds.pixel_array, cmap=plt.cm.gray)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-06-20T09:46:33.508016Z","iopub.execute_input":"2021-06-20T09:46:33.508372Z","iopub.status.idle":"2021-06-20T09:46:34.330178Z","shell.execute_reply.started":"2021-06-20T09:46:33.508277Z","shell.execute_reply":"2021-06-20T09:46:34.329208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path = '../input/siim-covid19-detection/train/057c02a959f1/6de2191aa170/ba463980acdb.dcm'\nds = dcmread(path)\nprint(ds)\nplt.imshow(ds.pixel_array, cmap=plt.cm.gray)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-06-20T09:46:34.331541Z","iopub.execute_input":"2021-06-20T09:46:34.331856Z","iopub.status.idle":"2021-06-20T09:46:35.215446Z","shell.execute_reply.started":"2021-06-20T09:46:34.331826Z","shell.execute_reply":"2021-06-20T09:46:35.214431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Get meta-information from train images","metadata":{}},{"cell_type":"code","source":"def dcm2metadata(sample):\n    metadata = {}\n    for key in sample.keys():\n        if key.group < 50:\n            item = sample.get(key)\n        if hasattr(item, 'description') and hasattr(item, 'value'):\n            metadata[item.description()] = str(item.value)\n    return metadata\n\nTRAIN_PATH = \"../input/siim-covid19-detection/train\"\ntrain_images_path = glob.glob(TRAIN_PATH + \"/*/*/*.dcm\")\nimage_metadata = pd.DataFrame()\n\n\nfor image in tqdm(train_images_path):    \n    # Read only the metadata here\n    ds = dcmread(image, stop_before_pixels=True)\n    info = dcm2metadata(ds)\n    image_metadata = image_metadata.append(info, ignore_index=True)\n        \nimage_metadata.head()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-20T09:46:35.216756Z","iopub.execute_input":"2021-06-20T09:46:35.217042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_metadata.columns","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Based on the metadata, I decide to focus only on the following columns :\n\n- Patient ID\n- Patient's Sex \n- Modality\n- Body Part Examined\n- Image type\n- Columns\n- Rows","metadata":{}},{"cell_type":"markdown","source":"### Patient ID","metadata":{}},{"cell_type":"code","source":"print(\"Number of unique patient : \", len(image_metadata[\"Patient ID\"].unique()))","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Patient's Sex","metadata":{}},{"cell_type":"code","source":"nb_male = len(image_metadata[image_metadata[\"Patient's Sex\"] == 'M'])\nnb_female = len(image_metadata[image_metadata[\"Patient's Sex\"] == 'F'])\n\nplt.figure(figsize=(6,6))\nplt.title(\"Gender distribution\")\nplt.pie([nb_male, nb_female], labels=['Male', 'Female'], autopct='%1.1f%%', colors=['b', 'r'])\nplt.show()","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Modality","metadata":{}},{"cell_type":"code","source":"image_metadata[\"Modality\"].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Modality information :\n- DX : Digital Radiography\n- CR : Computed Radiography\n\n> References : https://www.dicomlibrary.com/dicom/modality/\n\nTL;DR\n\nThese are two methods of achieving x-ray images. Thus, DX offers superior throughput compared to CR.\n\nMore information :\n> Computed radiography (CR) cassettes use photo-stimulated luminescence screens to capture the X-ray image, instead of traditional X-ray film. The CR cassette goes into a reader to convert the data into a digital image. Digital radiography (DR) systems use active matrix flat panels consisting of a detection layer deposited over an active matrix array of thin film transistors and photodiodes. With DR the image is converted to digital data in real-time and is available for review within seconds.\n\n> While both CR and DR have a wider dose range and can be post processed to eliminate mistakes and avoid repeat examinations, DR has some significant advantages over CR. DR improves workflow by producing higher quality images instantaneously while providing two to three times more dose efficiency than CR.\n\n> The good and bad of CR is that it enables digital imaging with the traditional workflow of X-ray film. With CR, like film, no synchronization to the generator is required, which had been a requirement for DR imaging. However, recent advances in DR panels are improving their flexibility, portability, and affordability.\n\nCited : Rick Colbeth - June 6, 2016 - https://www.vareximaging.com/computed-radiography-cr-and-digital-radiography-dr-which-should-you-choose","metadata":{}},{"cell_type":"markdown","source":"### Body Part Examined","metadata":{}},{"cell_type":"code","source":"image_metadata[\"Body Part Examined\"].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Regarding the different words, I think we can regroup the words that seems to be *Thorax* (*TORAX*, *T?RAX*, *THORAX*, *2- TORAX*, *TÒRAX*).\n\n- The empty category is a bit odd. Maybe doctors forgot to assign the body part when he took the x-ray?\n\n- *PORT CHEST* : referrence the upper chest where we can found a portal system, a small medical appliance, use to inject drugs or use to collect blood sample\n> https://en.wikipedia.org/wiki/Port_(medical)\n\n- *Pecho* : is spanish term for chest. We can group the *Pecho* and *PECHO* category.\n\n- *SKULL* : I don't really know what it means. On the sample bellow, it seems the same as chest radiography.\n\n- *ABDOMEN* : seems to be larger image in height. On the sample, we can see that we have the chest bust also the abdomen parti visible.","metadata":{}},{"cell_type":"code","source":"def get_sample_body_part(sample, title):\n    fig, axes = plt.subplots(nrows=1, ncols=3, figsize=(3 * 4, 4))\n    fig.suptitle(title)\n    i = 0\n    for study_id in sample['Study Instance UID'].values:\n        path = glob.glob('../input/siim-covid19-detection/train/' + study_id + '/*/*.dcm')\n        ds = dcmread(path[0])\n        axes[i].imshow(ds.pixel_array, cmap=plt.cm.gray, aspect='auto')\n        axes[i].axis('off')\n        i+=1    \n\nsample_port_chest = image_metadata[image_metadata[\"Body Part Examined\"] == \"PORT CHEST\"].sample(n=3, random_state=42)\nget_sample_body_part(sample_port_chest, 'Radiography Port Chest')","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_port_chest = image_metadata[image_metadata[\"Body Part Examined\"] == \"\"].sample(n=3, random_state=42)\nget_sample_body_part(sample_port_chest, 'Radiography Empty')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_port_chest = image_metadata[image_metadata[\"Body Part Examined\"] == \"SKULL\"].sample(n=3, random_state=42)\nget_sample_body_part(sample_port_chest, 'Radiography Skull')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_port_chest = image_metadata[image_metadata[\"Body Part Examined\"] == \"Pecho\"].sample(n=3, random_state=42)\nget_sample_body_part(sample_port_chest, 'Radiography Pecho')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_port_chest = image_metadata[image_metadata[\"Body Part Examined\"] == \"ABDOMEN\"].sample(n=3, random_state=42)\nget_sample_body_part(sample_port_chest, 'Radiography ABDOMEN')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*** DISCUSSIONS ***","metadata":{}},{"cell_type":"code","source":"image_metadata[\"Image Type\"].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Pixel Data Characteristics\n    - is the image an ORIGINAL Image; an image whose pixel values are based on original or source data\n    - is the image a DERIVED Image; an image whose pixel values have been derived in some manner from the pixel value of one or more other images\n\n\n\n- Patient Examination Characteristics\n    - is the image a PRIMARY Image; an image created as a direct result of the patient examination\n    - is the image a SECONDARY Image; an image created after the initial patient examination\n\n\n\n- Modality Specific Characteristics\n\n- Implementation specific identifiers; other implementation specific identifiers shall be documented in an implementation's conformance statement.\n\n\n\n> https://dicom.innolitics.com/ciods/ct-image/general-image/00080008","metadata":{}},{"cell_type":"markdown","source":"### Image size analysis","metadata":{}},{"cell_type":"code","source":"# Convert dtype\nimage_metadata.Columns = np.array(image_metadata.Columns, dtype=int)\nimage_metadata.Rows = np.array(image_metadata.Rows, dtype=int)\n\n# Show image size\nsizes = image_metadata.groupby(['Columns', 'Rows']).size().reset_index().rename(columns={0 : 'count'})\nsizes.hvplot.scatter(\n    x='Columns', \n    y='Rows', \n    size='count',\n    title='Image size distribution',\n    xlim=(0,5000), ylim=(0,5000), \n    grid=True, \n    height=500, width=1000).options(scaling_factor=0.1, line_alpha=1, fill_alpha=0)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-06-20T13:10:11.436661Z","iopub.execute_input":"2021-06-20T13:10:11.436956Z","iopub.status.idle":"2021-06-20T13:10:11.533996Z","shell.execute_reply.started":"2021-06-20T13:10:11.436929Z","shell.execute_reply":"2021-06-20T13:10:11.532193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see on the top diagram, the size of the images seems to follow a linear line starting from 0. It seems that we have globally square sized images. Moreover, we have a high concentration of images with a size between 2000 and 3000 pixels.","metadata":{}},{"cell_type":"markdown","source":"*** CONCLUSION ***\n\nFINALLY , i conclude that , In this  kaggle notebook I have seen a lot of information with that i summarised the main major ideas:\n\n- We have unbalanced data.\n- A study can contain several images. Those images can be duplicated.\n- The brightness of the images changes a lot.\n- According to the metadata, the set of images corresponds well to the location of the chest.\n- The image appear to be square and its size is concentrated between 2000 and 3000 pixels.\n- The images were taken equally between CR and DX. \n\n\nat lastly i loved do this interesting topic finding covid19 chest x rays bu this machine learning method , hope every were excited to do this !!","metadata":{}},{"cell_type":"markdown","source":"*** REEFERENCES ***\n1) INTERNET\n2) WWW.KAGGLE.COM\n3) SOCIAL MEDIA \n4) TEXT BOOKS\n5) AND A GOOD GUIDANCE FROM MY INTERNSHIP TEAM LEADER.","metadata":{}},{"cell_type":"code","source":"            ***  RESULTS ***\n    finally , this chest x ray can reveal many things from our body includings below\n    : the condition of our lungs .\n    : x rays can detect cancer.\n    : infection of air collecting in the lungs(\"pneumothorax\").\n    : chest x ray shows heart problem changes in the lung .  \n        ","metadata":{},"execution_count":null,"outputs":[]}]}