{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# RSNA pneumonia detection  (EDA)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"## Contents\n\n* **1. Introduction**\n* **2. Data preparation**\n    * 2.1 Load relevant libraries\n    * 2.2 Data reading\n* **3. Data exploration**\n    * 3.1 Missing value exploration\n    * 3.2 Combine train and class data\n    * 3.3 Explore DICOM data\n    * 3.4 Combine image data and class and labels data\n    * 3.5 Add meta information from DICOM data\n    * 3.6 Show bounding box distribution\n    * 3.7 Save the processed result\n","metadata":{"execution":{"iopub.status.busy":"2023-02-04T00:21:09.466227Z","iopub.execute_input":"2023-02-04T00:21:09.466807Z","iopub.status.idle":"2023-02-04T00:21:09.477139Z","shell.execute_reply.started":"2023-02-04T00:21:09.466739Z","shell.execute_reply":"2023-02-04T00:21:09.475499Z"}}},{"cell_type":"markdown","source":"## 1. Introduction\n\nThe objective of this work is to explore the medical images show casing lungs that may or may not be healthy. The unhealthy lungs are said to be infected with RSNA Pnemonia. Therefore, the medical images given will be analysed visually as well as spatially. \n\nIn this kernel we shall work with stage 2 data.","metadata":{}},{"cell_type":"markdown","source":"## 2. Data preparation\n### 2.1 Load relevant libraries","metadata":{"execution":{"iopub.status.busy":"2023-02-04T00:24:02.559040Z","iopub.execute_input":"2023-02-04T00:24:02.559567Z","iopub.status.idle":"2023-02-04T00:24:02.566557Z","shell.execute_reply.started":"2023-02-04T00:24:02.559526Z","shell.execute_reply":"2023-02-04T00:24:02.564595Z"}}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom tqdm.notebook import tqdm\nfrom glob import glob\nfrom tqdm.notebook import tqdm_notebook\nfrom matplotlib.patches import Rectangle\nimport pydicom as cdm","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:51.252810Z","iopub.execute_input":"2023-03-21T05:32:51.253272Z","iopub.status.idle":"2023-03-21T05:32:51.260537Z","shell.execute_reply.started":"2023-03-21T05:32:51.253236Z","shell.execute_reply":"2023-03-21T05:32:51.259242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.2 Data reading","metadata":{}},{"cell_type":"markdown","source":"We are going to read the data for files containing stage 2 detailed class information and stage 2 train labels information.","metadata":{}},{"cell_type":"code","source":"# Loading of the data\ndet_class_info = pd.read_csv(\"../input/rsna-pneumonia-detection-challenge/stage_2_detailed_class_info.csv\")\ntrain_labels = pd.read_csv(\"../input/rsna-pneumonia-detection-challenge/stage_2_train_labels.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:51.271735Z","iopub.execute_input":"2023-03-21T05:32:51.273290Z","iopub.status.idle":"2023-03-21T05:32:51.408443Z","shell.execute_reply.started":"2023-03-21T05:32:51.273245Z","shell.execute_reply":"2023-03-21T05:32:51.407369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Checking the sizes of the datasets\nprint(f\"Detailed class information: {det_class_info.shape}\")\nprint(f\"Train labels information: {train_labels.shape}\")","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:51.410765Z","iopub.execute_input":"2023-03-21T05:32:51.411125Z","iopub.status.idle":"2023-03-21T05:32:51.418198Z","shell.execute_reply.started":"2023-03-21T05:32:51.411093Z","shell.execute_reply":"2023-03-21T05:32:51.416874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that the data sets have the same dimention of rows.","metadata":{}},{"cell_type":"code","source":"# Check the first five variables of the detailed class information\ndet_class_info.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:51.656816Z","iopub.execute_input":"2023-03-21T05:32:51.657276Z","iopub.status.idle":"2023-03-21T05:32:51.668231Z","shell.execute_reply.started":"2023-03-21T05:32:51.657235Z","shell.execute_reply":"2023-03-21T05:32:51.667090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Checking the first five variables of the train labels dataframe\ntrain_labels.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:51.670527Z","iopub.execute_input":"2023-03-21T05:32:51.670954Z","iopub.status.idle":"2023-03-21T05:32:51.692466Z","shell.execute_reply.started":"2023-03-21T05:32:51.670917Z","shell.execute_reply":"2023-03-21T05:32:51.691590Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- The detailed class information dataset shows the patient ID and class  data that depicts wether or not a person has normal lungs or not.\n\n- train label data set shows the patient ID, the window (x and y, width and height) showing evidence of pneumonia and the target variables indication binary information of 0's and 1's. \n\n- Basically, you can see that item number in the ***det_class_info*** has Lung Opacity and this items in the ***train_labels*** datafram has a target value of zero. Thus, these two data sets are interestingly linked.","metadata":{"execution":{"iopub.status.busy":"2023-02-04T01:21:31.900202Z","iopub.execute_input":"2023-02-04T01:21:31.900731Z","iopub.status.idle":"2023-02-04T01:21:31.924195Z","shell.execute_reply.started":"2023-02-04T01:21:31.900690Z","shell.execute_reply":"2023-02-04T01:21:31.922416Z"}}},{"cell_type":"markdown","source":"## 3. Data exploration","metadata":{}},{"cell_type":"markdown","source":"### 3.1 Missing value exploration","metadata":{}},{"cell_type":"code","source":"# checking for missing values in the in both train labels and class_info data sets.\ndef missing_items(data):\n    total = data.isnull().sum().sort_values(ascending = False)\n    percent = (data.isnull().sum()/data.isnull().count()*100).sort_values(ascending = False)\n    return np.transpose(pd.concat([total, percent], axis = 1, keys = ['Total', 'Percent']))\n# Checking for missing values in train labels dataframe\nmissing_items(train_labels)","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:51.693733Z","iopub.execute_input":"2023-03-21T05:32:51.694245Z","iopub.status.idle":"2023-03-21T05:32:51.725276Z","shell.execute_reply.started":"2023-03-21T05:32:51.694211Z","shell.execute_reply":"2023-03-21T05:32:51.724130Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"68.39 % of the data is missing in (x, y, width, height, patientId, Target) in the train labels dataframe. This missing data represents the ***Target 0***  which contains data items of (***not lung opacity***) condition.","metadata":{}},{"cell_type":"code","source":"# Checking for missing values in detailed class dataframe\nmissing_items(det_class_info)","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:51.727478Z","iopub.execute_input":"2023-03-21T05:32:51.727971Z","iopub.status.idle":"2023-03-21T05:32:51.756231Z","shell.execute_reply.started":"2023-03-21T05:32:51.727923Z","shell.execute_reply":"2023-03-21T05:32:51.754922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the detailed class information dataset, it can be seen that 0 % of the data is missing.","metadata":{}},{"cell_type":"code","source":"# Plot the distribution of the class detailed information\nf, ax = plt.subplots(1, 1, figsize = (15, 5))\ntotal = float(len(det_class_info))\ndet_class_info.groupby('class').size().plot.bar()\nfor p in ax.patches:\n    height = p.get_height()   \n    ax.text(p.get_x() + p.get_width()/2., \n           height + 3, f'{int(100 * height/total)} %', ha = 'center', bbox=dict(facecolor='red', alpha=0.5))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:51.758604Z","iopub.execute_input":"2023-03-21T05:32:51.759004Z","iopub.status.idle":"2023-03-21T05:32:52.012858Z","shell.execute_reply.started":"2023-03-21T05:32:51.758968Z","shell.execute_reply":"2023-03-21T05:32:52.011653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that the majority of the data is made up of ***No Lung Opacity / Not Normal*** consiting of 39 % of the data while the least among data classes is made up of ***Normal** classes comprising of 29 % of total data.","metadata":{"execution":{"iopub.status.busy":"2023-03-18T02:18:49.803372Z","iopub.execute_input":"2023-03-18T02:18:49.803742Z","iopub.status.idle":"2023-03-18T02:18:49.809746Z","shell.execute_reply.started":"2023-03-18T02:18:49.803710Z","shell.execute_reply":"2023-03-18T02:18:49.808868Z"}}},{"cell_type":"code","source":"# Checking the actual count of the data items in the detailed class dataset\ndef feature_distribution(data, feature):\n    # Label count\n    l_counts = data[feature].value_counts() # Label_counts\n    \n    # Total number of samples\n    t_samples = len(data) # total samples\n    \n    # Count the number of samples in each class\n    print(f'Feature: {feature}')\n    for i in range(len(l_counts)):\n        lbl = l_counts.index[i] # Label\n        cnt = l_counts.values[i] # Count\n        print(f'{lbl:<30s}: {cnt}')\n\nfeature_distribution(det_class_info, 'class')","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:52.014269Z","iopub.execute_input":"2023-03-21T05:32:52.014690Z","iopub.status.idle":"2023-03-21T05:32:52.026428Z","shell.execute_reply.started":"2023-03-21T05:32:52.014656Z","shell.execute_reply":"2023-03-21T05:32:52.024714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It can be seen that the data item ***No Lung Opacity / Not Normal*** lungs has the a total number of 11821 which is 39 % of the data while ***Lung Opacity*** dataset has 9555 items comprising of 31 % of the total data and lastly, ***Normal*** lungs items are 8851 in total consiting of 29 % of the dataset.","metadata":{}},{"cell_type":"code","source":"# Total number of patient cases\nprint('Patient Cases:',det_class_info['patientId'].value_counts().shape[0])","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:52.028346Z","iopub.execute_input":"2023-03-21T05:32:52.028739Z","iopub.status.idle":"2023-03-21T05:32:52.057543Z","shell.execute_reply.started":"2023-03-21T05:32:52.028682Z","shell.execute_reply":"2023-03-21T05:32:52.056178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The detailed class dataframe contains 26684 patient cases.","metadata":{}},{"cell_type":"markdown","source":"### 3.2 Combine train and class data","metadata":{}},{"cell_type":"markdown","source":"***Merging Method***","metadata":{}},{"cell_type":"code","source":"# Merging the two datasets using the patient ID\ncombined_df = train_labels.merge(det_class_info, left_on='patientId', right_on='patientId', how='inner')","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:52.061901Z","iopub.execute_input":"2023-03-21T05:32:52.062446Z","iopub.status.idle":"2023-03-21T05:32:52.098352Z","shell.execute_reply.started":"2023-03-21T05:32:52.062397Z","shell.execute_reply":"2023-03-21T05:32:52.097079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# show the row dimension of the combined data and the one of the actual data sets to check\n# wether the combination has not created duplicates.\nprint('Combined_data row dimension:', combined_df.shape[0])\nprint('Actual data row dimension:', det_class_info.shape[0])","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:52.099847Z","iopub.execute_input":"2023-03-21T05:32:52.100180Z","iopub.status.idle":"2023-03-21T05:32:52.106729Z","shell.execute_reply.started":"2023-03-21T05:32:52.100150Z","shell.execute_reply":"2023-03-21T05:32:52.105328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After merging, we can see that our data became more than the actual dimension of one of the two data sets. Meaning that this merging method has created duplicates.","metadata":{}},{"cell_type":"markdown","source":"***Concatenation method***","metadata":{}},{"cell_type":"code","source":"# We tried to combine the two data sets again but using cconcatenation method.\ndetailed = det_class_info.drop('patientId', 1)\n#combined_df = pd.concat([train_labels, det_class_info.drop('patientId', 1)], 1)\ncombined_df = pd.concat([train_labels, detailed], 1)\nprint(combined_df.shape[0], 'Combined data')","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:52.108929Z","iopub.execute_input":"2023-03-21T05:32:52.109306Z","iopub.status.idle":"2023-03-21T05:32:52.130095Z","shell.execute_reply.started":"2023-03-21T05:32:52.109263Z","shell.execute_reply":"2023-03-21T05:32:52.128692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that the row dimension of the combined data set is equal to that of the actual data sets. Therefore, cconcatenation method had worked just fine.","metadata":{}},{"cell_type":"code","source":"# Checking the first 5 sample in the combined datasets\ncombined_df.sample(5)","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:52.132340Z","iopub.execute_input":"2023-03-21T05:32:52.132802Z","iopub.status.idle":"2023-03-21T05:32:52.153804Z","shell.execute_reply.started":"2023-03-21T05:32:52.132755Z","shell.execute_reply":"2023-03-21T05:32:52.152415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Explore the distribution of boxes***","metadata":{}},{"cell_type":"code","source":"# Extracting the boxes from the data\nboxes = combined_df.groupby('patientId').size().reset_index(name='boxes')\ncombined_box_df = pd.merge(combined_df, boxes, on = 'patientId')\ncombined_box_df.head(6)","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:52.157758Z","iopub.execute_input":"2023-03-21T05:32:52.158117Z","iopub.status.idle":"2023-03-21T05:32:52.312137Z","shell.execute_reply.started":"2023-03-21T05:32:52.158086Z","shell.execute_reply":"2023-03-21T05:32:52.310229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Showing the distribution of boxes with respect to the number of patients\nboxes.groupby('boxes').size().reset_index(name = 'patients')","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:52.313984Z","iopub.execute_input":"2023-03-21T05:32:52.314547Z","iopub.status.idle":"2023-03-21T05:32:52.332985Z","shell.execute_reply.started":"2023-03-21T05:32:52.314494Z","shell.execute_reply":"2023-03-21T05:32:52.331178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Showing the relationship between class and target variables***","metadata":{}},{"cell_type":"code","source":"# Relationship between class and taget\ncombined_box_df.groupby(['class', 'Target']).size().reset_index(name = 'NO. of Patients')","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:52.334647Z","iopub.execute_input":"2023-03-21T05:32:52.335781Z","iopub.status.idle":"2023-03-21T05:32:52.360367Z","shell.execute_reply.started":"2023-03-21T05:32:52.335743Z","shell.execute_reply":"2023-03-21T05:32:52.358764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Showing the relationship between class and target graphically***","metadata":{}},{"cell_type":"code","source":"# Plotting the distribution of each class with respect to the target variable.\nfig, ax = plt.subplots(nrows = 1, figsize = (10, 8))\ndata = combined_box_df.groupby('Target')['class'].value_counts()\ndf = pd.DataFrame(data = {'X_ray exams' : data.values}, index = data.index).reset_index()\nsns.barplot(ax = ax, x = 'Target', y = 'X_ray exams', hue = 'class', data = df)\nplt.title('X_ray examination of classe of lungs and the target variable')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:52.362297Z","iopub.execute_input":"2023-03-21T05:32:52.363140Z","iopub.status.idle":"2023-03-21T05:32:52.665423Z","shell.execute_reply.started":"2023-03-21T05:32:52.363070Z","shell.execute_reply":"2023-03-21T05:32:52.664259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Target value of 0 depicts that no pathology was detected and that they are either of normal class or not-normal/no-lung-opancity while target values of 1 depicts that a pathalogy was detected.","metadata":{}},{"cell_type":"code","source":"# Plot the x, y, width and height dimensions for data sets with Lung Opactity.\ntarget1 = combined_box_df[combined_box_df['Target']==1]\nplt.figure()\nfig, ax = plt.subplots(2, 2, figsize = (12,12))\nsns.histplot(target1['x'], kde = True, bins = 50, color = 'blue', ax=ax[0,0])\nsns.histplot(target1['y'], kde=True, bins=50, color='red', ax=ax[0,1])\nsns.histplot(target1['width'], kde=True, bins=50, color='green', ax=ax[1,0])\nsns.histplot(target1['height'], kde=True, bins=50, color='cyan', ax=ax[1,1])\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:52.669683Z","iopub.execute_input":"2023-03-21T05:32:52.670929Z","iopub.status.idle":"2023-03-21T05:32:54.123304Z","shell.execute_reply.started":"2023-03-21T05:32:52.670888Z","shell.execute_reply":"2023-03-21T05:32:54.122326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Checking the medical images***","metadata":{}},{"cell_type":"code","source":"import os\nimage_sample_path = os.listdir('../input/rsna-pneumonia-detection-challenge/stage_2_train_images')[:5]\nprint(image_sample_path)","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:54.124735Z","iopub.execute_input":"2023-03-21T05:32:54.125798Z","iopub.status.idle":"2023-03-21T05:32:54.154609Z","shell.execute_reply.started":"2023-03-21T05:32:54.125744Z","shell.execute_reply":"2023-03-21T05:32:54.153298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load train and test images\nimage_train_path = os.listdir('../input/rsna-pneumonia-detection-challenge/stage_2_train_images')\n#image_train_path2 = os.PathLike('../input/rsna-pneumonia-detection-challenge/stage_2_train_images')\nimage_test_path = os.listdir('../input/rsna-pneumonia-detection-challenge/stage_2_test_images')\n\nprint('train set image count:', len(image_train_path), '\\ntest set image count:', len(image_test_path))","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:54.156192Z","iopub.execute_input":"2023-03-21T05:32:54.156713Z","iopub.status.idle":"2023-03-21T05:32:54.178236Z","shell.execute_reply.started":"2023-03-21T05:32:54.156664Z","shell.execute_reply":"2023-03-21T05:32:54.176755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check for duplicates in the train data set\nprint('Unique patientId in combined box df: ', combined_box_df['patientId'].nunique())","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:54.179972Z","iopub.execute_input":"2023-03-21T05:32:54.180605Z","iopub.status.idle":"2023-03-21T05:32:54.196531Z","shell.execute_reply.started":"2023-03-21T05:32:54.180555Z","shell.execute_reply":"2023-03-21T05:32:54.194952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The number of unique samples in the patientId is exactly equal to the number of samples in the image train data set. SInce the combined_box_df contains of 30227, it means that the data contains duplicates of 3543.","metadata":{}},{"cell_type":"markdown","source":"## 3.3 Exploring DICOM meta data","metadata":{}},{"cell_type":"code","source":"import pydicom as dcm\nsamplePatientID = list(combined_box_df[:3].T.to_dict().values())[0]['patientId']\nsamplePatientID = samplePatientID+'.dcm'\ndicom_file_path = os.path.join(\"../input/rsna-pneumonia-detection-challenge/stage_2_train_images/\",samplePatientID)\ndicom_file_dataset=dcm.read_file(dicom_file_path)\ndicom_file_dataset","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:54.197987Z","iopub.execute_input":"2023-03-21T05:32:54.199145Z","iopub.status.idle":"2023-03-21T05:32:54.235774Z","shell.execute_reply.started":"2023-03-21T05:32:54.199105Z","shell.execute_reply":"2023-03-21T05:32:54.234394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"plotting DICON images where target=1","metadata":{}},{"cell_type":"code","source":"def show_dicom_images(data):\n    img_data = list(data.T.to_dict().values())\n    f, ax = plt.subplots(2, 2, figsize = (15, 10))\n    for i, data_row in enumerate(img_data):\n        patientImage = data_row['patientId']+'.dcm'\n        image_path = os.path.join(\"../input/rsna-pneumonia-detection-challenge/stage_2_train_images/\",patientImage)\n        data_row_image_data = dcm.read_file(image_path)\n        modality = data_row_image_data.Modality\n        age = data_row_image_data.PatientAge\n        sex = data_row_image_data.PatientSex\n        data_row_img = dcm.dcmread(image_path)\n        ax[i//2, i%2].imshow(data_row_img.pixel_array, cmap=plt.cm.bone)\n        ax[i//2, i%2].axis('off')\n        ax[i//2, i%2].set_title(f\"'ID: {data_row['patientId']}\\nModality: {modality} Age: {age} Sex: {sex} Target: {data_row['Target']}\\nClass: {data_row['class']}\\nWindow: {data_row['x']}: {data_row['y']}: {data_row['width']}: {data_row['height']}'\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:54.238250Z","iopub.execute_input":"2023-03-21T05:32:54.238874Z","iopub.status.idle":"2023-03-21T05:32:54.251136Z","shell.execute_reply.started":"2023-03-21T05:32:54.238825Z","shell.execute_reply":"2023-03-21T05:32:54.249650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show_dicom_images(combined_box_df[combined_box_df['Target']==1].sample(4))","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:54.252438Z","iopub.execute_input":"2023-03-21T05:32:54.253695Z","iopub.status.idle":"2023-03-21T05:32:55.076301Z","shell.execute_reply.started":"2023-03-21T05:32:54.253636Z","shell.execute_reply":"2023-03-21T05:32:55.075397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def show_dicom_images_with_boxes(data):\n    img_data = list((data.T.to_dict().values()))\n    f, ax = plt.subplots(2,2, figsize = (15, 10))\n    for i, data_row in enumerate(img_data):\n        patientImage = data_row['patientId']+'.dcm'\n        imagePath = os.path.join(\"../input/rsna-pneumonia-detection-challenge/stage_2_train_images/\", patientImage)\n        data_row_img_data = dcm.read_file(imagePath)\n        modality = data_row_img_data.Modality\n        age = data_row_img_data.PatientAge\n        sex = data_row_img_data.PatientSex\n        data_row_img = dcm.dcmread(imagePath)\n        ax[i//2, i%2].imshow(data_row_img.pixel_array, cmap = plt.cm.bone)\n        ax[i//2, i%2].axis('off')\n        ax[i//2, i%2].set_title(f\"ID: {data_row['patientId']}\\n Modality: {modality} Age: {age} Sex: {sex} Target: {data_row['Target']}\\nClass: {data_row['class']}\")\n        rows = combined_box_df[combined_box_df['patientId']==data_row['patientId']]\n        box_data = list(rows.T.to_dict().values())\n        for j, row in enumerate(box_data):\n            ax[i//2, i%2].add_patch(Rectangle(xy=(row['x'], row['y']), width=row['width'], height=row['height'], color='yellow', alpha=0.1))\n            \n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:55.077738Z","iopub.execute_input":"2023-03-21T05:32:55.078787Z","iopub.status.idle":"2023-03-21T05:32:55.091061Z","shell.execute_reply.started":"2023-03-21T05:32:55.078747Z","shell.execute_reply":"2023-03-21T05:32:55.089941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show_dicom_images_with_boxes(combined_box_df[combined_box_df['Target']==1].sample(4))","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:55.092752Z","iopub.execute_input":"2023-03-21T05:32:55.093235Z","iopub.status.idle":"2023-03-21T05:32:55.956689Z","shell.execute_reply.started":"2023-03-21T05:32:55.093183Z","shell.execute_reply":"2023-03-21T05:32:55.955413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As indicated by the boxes in the figures above, we can see that some lungs in ***Target 1*** are problematic as they are infected with lung opancity.","metadata":{}},{"cell_type":"code","source":"# Plot lung data with target value 0\nshow_dicom_images(combined_box_df[combined_box_df['Target']==0].sample(4))","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:55.958649Z","iopub.execute_input":"2023-03-21T05:32:55.958996Z","iopub.status.idle":"2023-03-21T05:32:56.783496Z","shell.execute_reply.started":"2023-03-21T05:32:55.958964Z","shell.execute_reply":"2023-03-21T05:32:56.781760Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the images shown about for ***Target 0***, we can see that the boxes depicting unhealthy lungs or lungs with opancity are missing. However, from these images alone we cant tell whether the lungs all thes lungs are completely health.","metadata":{}},{"cell_type":"markdown","source":"***Showing boxes distribution so as to have a clear depction of unhealthy lungs***","metadata":{}},{"cell_type":"markdown","source":"## 3.4 Combine image data and class and labels data\n\nShowing them as distributions can give a higher to identify the unhealthy lungs.","metadata":{}},{"cell_type":"code","source":"# Combine image data with labels and class data (combined dataframe)\n\nimage_train_path = ('../input/rsna-pneumonia-detection-challenge/stage_2_train_images')\n\nimage_df = pd.DataFrame({'path': glob(os.path.join(image_train_path, '*.dcm'))})\nimage_df['patientId'] = image_df['path'].map(lambda x: os.path.splitext(os.path.basename(x))[0])\n\nimage_boxes_df = pd.merge(combined_box_df, image_df, on = 'patientId', how = 'left').sort_values('patientId')\nprint(image_boxes_df.shape[0], 'image bounding boxes')\nimage_boxes_df.sample(5)","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:56.785438Z","iopub.execute_input":"2023-03-21T05:32:56.786215Z","iopub.status.idle":"2023-03-21T05:32:57.044446Z","shell.execute_reply.started":"2023-03-21T05:32:56.786174Z","shell.execute_reply":"2023-03-21T05:32:57.042846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3.5 Add meta data","metadata":{}},{"cell_type":"code","source":"# We are going to parse the DICOM meta information and then add it to the combined box data set.\nPATH=\"../input/rsna-pneumonia-detection-challenge/\"\n#\"../input/rsna-pneumonia-detection-challenge/stage_2_train_images/\"\nvars = ['Modality', 'PatientAge', 'PatientSex', 'BodyPartExamined', 'ViewPosition', 'ConversionType',\n       'Rows', 'Columns', 'PixelSpacing']\n\ndef process_dicom_data(data_df, data_path):\n    for var in vars:\n        data_df[var] = None\n    image_names = os.listdir(PATH + data_path)\n    for i, img_name in tqdm_notebook(enumerate(image_names)): \n        imagePath = os.path.join(PATH, data_path, img_name)\n        data_row_img_data = dcm.read_file(imagePath)\n        idx = (data_df['patientId']==data_row_img_data.PatientID)\n        data_df.loc[idx, 'Modality'] = data_row_img_data.Modality\n        data_df.loc[idx, 'PatientAge'] = pd.to_numeric(data_row_img_data.PatientAge)\n        data_df.loc[idx, 'PatientSex'] = data_row_img_data.PatientSex\n        data_df.loc[idx, 'BodyPartExamined'] = data_row_img_data.BodyPartExamined\n        data_df.loc[idx, 'ViewPosition'] = data_row_img_data.ViewPosition\n        data_df.loc[idx, 'ConversionType'] = data_row_img_data.ConversionType\n        data_df.loc[idx, 'Rows'] = data_row_img_data.Rows\n        data_df.loc[idx, 'Columns'] = data_row_img_data.Columns\n        data_df.loc[idx, 'PixelSpacing'] = str.format(\"{:4.3f}\", data_row_img_data.PixelSpacing[0])","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:57.049635Z","iopub.execute_input":"2023-03-21T05:32:57.050016Z","iopub.status.idle":"2023-03-21T05:32:57.061063Z","shell.execute_reply.started":"2023-03-21T05:32:57.049981Z","shell.execute_reply":"2023-03-21T05:32:57.059658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"process_dicom_data(combined_box_df, 'stage_2_train_images/')","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:32:57.063004Z","iopub.execute_input":"2023-03-21T05:32:57.063408Z","iopub.status.idle":"2023-03-21T05:42:46.868098Z","shell.execute_reply.started":"2023-03-21T05:32:57.063370Z","shell.execute_reply":"2023-03-21T05:42:46.866646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a sample dataframe covering different cases of the data\nsample_df = image_boxes_df.groupby(['Target', 'class', 'boxes']).\\\n    apply(lambda x: x[x['patientId']==x.sample(1)['patientId'].values[0]]).\\\n    reset_index(drop = True)\nsample_df","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:42:46.870758Z","iopub.execute_input":"2023-03-21T05:42:46.871270Z","iopub.status.idle":"2023-03-21T05:42:46.940563Z","shell.execute_reply.started":"2023-03-21T05:42:46.871219Z","shell.execute_reply":"2023-03-21T05:42:46.938855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3.6 Show bounding box distribution","metadata":{}},{"cell_type":"code","source":"# Sampling of window boxes where lung opacity was detected and plotting\nN = 9555\narea = (30 * np.random.rand(N))**2\n \nbbox = image_boxes_df.query('Target == 1')\nbbox.plot.scatter(x = 'x', y = 'y',  s= area, c = 'red', alpha = 0.05)","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:42:46.942463Z","iopub.execute_input":"2023-03-21T05:42:46.942989Z","iopub.status.idle":"2023-03-21T05:42:47.518677Z","shell.execute_reply.started":"2023-03-21T05:42:46.942938Z","shell.execute_reply":"2023-03-21T05:42:47.517387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plotting centers of lung opancity rectangles by taking a sample of 3000\nfig, ax = plt.subplots(1, 1, figsize = (10, 10))\ntarget_sample = image_boxes_df.sample(5000)\ntarget_sample['xc'] = target_sample['x'] + target_sample['width'] / 2\ntarget_sample['yc'] = target_sample['y'] + target_sample['height'] / 2\nplt.title('Centers of lung opacity show in brown over yellow rectangles\\nSample size: 3000')\ntarget_sample.plot.scatter(x='xc', y='yc', xlim=(0, 1024), ylim=(0, 1024),\n                          ax=ax, alpha=0.8, marker=\".\", color='brown')\nfor i, crt_sample in target_sample.iterrows():\n    ax.add_patch(Rectangle(xy=(crt_sample['x'], crt_sample['y']), width=crt_sample['width'],\n                          height=crt_sample['height'], alpha=3.5e-3, color='red'))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:42:47.520559Z","iopub.execute_input":"2023-03-21T05:42:47.521394Z","iopub.status.idle":"2023-03-21T05:42:57.638963Z","shell.execute_reply.started":"2023-03-21T05:42:47.521343Z","shell.execute_reply":"2023-03-21T05:42:57.637353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that the disribution of windows is highly concentrated around the middle.","metadata":{}},{"cell_type":"markdown","source":"***Showing the boxes as segments***\n\nThis helps us to easily capture a propability that can aid us to see regions that are likely to have lung opacity.","metadata":{}},{"cell_type":"code","source":"# Show the boxes\nX, Y = 1024, 1024\nxx, yy = np.meshgrid(np.linspace(0, X, Y), np.linspace(0, X, Y),\n                    indexing = 'xy')\nimage_prob = np.zeros_like(xx)\nfor _, c_row in bbox.sample(5000).iterrows():\n    c_mask = (xx >= c_row['x']) & (xx <= (c_row['x'] + c_row['width']))\n    c_mask &= (yy >= c_row['y']) & (yy <= c_row['y'] + c_row['height'])\n    image_prob += c_mask\nfig, ax1 = plt.subplots(1,1, figsize = (10, 10))\nax1.imshow(image_prob, cmap = 'hot')","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:42:57.641035Z","iopub.execute_input":"2023-03-21T05:42:57.641454Z","iopub.status.idle":"2023-03-21T05:43:29.457527Z","shell.execute_reply.started":"2023-03-21T05:42:57.641415Z","shell.execute_reply":"2023-03-21T05:43:29.456173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Overlying the probability on some images***","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(2, 3, figsize = (20, 15))\nfor p_ax, (p_path, p_rows) in zip(ax.flatten(),\n                    sample_df.groupby(['path'])):\n    p_img_arr = dcm.read_file(p_path).pixel_array\n    \n    p_img = plt.cm.gray(p_img_arr)\n    p_img += 0.25*plt.cm.hot(image_prob/image_prob.max())\n    p_img = np.clip(p_img, 0, 1)\n    p_ax.imshow(p_img)\n    \n\n    p_ax.set_title(f\"{p_rows.iloc[0]['class']}\")\n    for i, (_, p_row) in enumerate(p_rows.dropna().iterrows()):     \n        p_ax.plot(p_row['x'], p_row['y'], 's', label = f\"{p_row['class']}\")\n        p_ax.add_patch(Rectangle(xy=(p_row['x'], p_row['y']),\n                                width = p_row['width'],\n                                height = p_row['height'], \n                                alpha = 0.5,\n                                fill=False))\n        if i==0: p_ax.legend()\nfig.savefig('figures.png')","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:43:29.458932Z","iopub.execute_input":"2023-03-21T05:43:29.459450Z","iopub.status.idle":"2023-03-21T05:43:33.154532Z","shell.execute_reply.started":"2023-03-21T05:43:29.459411Z","shell.execute_reply":"2023-03-21T05:43:33.153300Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To tell the difference as to whether there is lung opacity or not is still quite challenging even after laying the probability distributions, however, can see that lungs that are normal has a lower intersity of the prob distrubution while those with lungs Opacity show a higher intensity of prob distributions, the same can be observed in those lungs that have no lung opacity but are not normal.","metadata":{}},{"cell_type":"markdown","source":"## 3.7 Saving the processed results with DICOM data for training purposes.","metadata":{}},{"cell_type":"code","source":"image_boxes_df.to_csv('image_boxes_df.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2023-03-21T05:50:41.983805Z","iopub.execute_input":"2023-03-21T05:50:41.984343Z","iopub.status.idle":"2023-03-21T05:50:42.198764Z","shell.execute_reply.started":"2023-03-21T05:50:41.984285Z","shell.execute_reply":"2023-03-21T05:50:42.197412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## References\n\n[1] Kevin Mader, Lung Opacity Overview, https://www.kaggle.com/kmader/lung-opacity-overview\n[2] Gabriel Preda, RSNA Pneumonia Detection EDA, https://www.kaggle.com/code/gpreda/rsna-pneumonia-detection-eda\n","metadata":{}}]}