{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Exploratory Data Analysis of RSNA Pneumonia Dataset\n\n#### Data fields\n- patientId - A patientId. Each patientId corresponds to a unique image.\n- x - the upper-left x coordinate of the bounding box.\n- y - the upper-left y coordinate of the bounding box.\n- width - the width of the bounding box.\n- height - the height of the bounding box.\n- Target - the binary Target, indicating whether this sample has evidence of pneumonia.","metadata":{}},{"cell_type":"code","source":"import random\n\nimport numpy as np\nimport pandas as pd\n\nimport matplotlib.pyplot as plt\nfrom matplotlib.pyplot import figure\nfigure(figsize=(15, 12), dpi=120)\nimport seaborn as sns\nsns.set(style='whitegrid') #set seaborn plotting aesthetics\n%matplotlib inline\n\nimport os\nimport csv\nfrom pathlib import Path\nimport pydicom\nfrom glob import glob\nfrom matplotlib.patches import Rectangle","metadata":{"ExecuteTime":{"end_time":"2022-10-21T04:12:00.322921Z","start_time":"2022-10-21T04:11:49.513059Z"},"executionInfo":{"elapsed":3117,"status":"ok","timestamp":1666403325880,"user":{"displayName":"Rahul S","userId":"05974193569914813309"},"user_tz":-330},"id":"bRthd52meusb","execution":{"iopub.status.busy":"2022-11-20T23:18:05.203612Z","iopub.execute_input":"2022-11-20T23:18:05.204594Z","iopub.status.idle":"2022-11-20T23:18:07.262531Z","shell.execute_reply.started":"2022-11-20T23:18:05.204486Z","shell.execute_reply":"2022-11-20T23:18:07.261396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Reading the data\n\ntrain_labels = pd.read_csv('../input/rsna-pneumonia-detection-challenge/stage_2_train_labels.csv')\ntrain_class = pd.read_csv('../input/rsna-pneumonia-detection-challenge/stage_2_detailed_class_info.csv')\n\ntrain_path = os.listdir(Path('../input/rsna-pneumonia-detection-challenge/stage_2_train_images'))\ntest_path = os.listdir(Path('../input/rsna-pneumonia-detection-challenge/stage_2_test_images'))\n\nprint('Number of Duplicated records in train_labels file:', train_labels.duplicated().sum())\nprint('------------------------------------------------------------------------')\nprint(train_labels.info())\nprint('------------------------------------------------------------------------')\nprint(train_labels.head(15))","metadata":{"executionInfo":{"elapsed":1037,"status":"ok","timestamp":1666403358520,"user":{"displayName":"Rahul S","userId":"05974193569914813309"},"user_tz":-330},"id":"K8MGboFYr93i","outputId":"852662b1-43df-4e37-e3a6-50b547693b0b","execution":{"iopub.status.busy":"2022-11-20T23:20:57.418443Z","iopub.execute_input":"2022-11-20T23:20:57.418811Z","iopub.status.idle":"2022-11-20T23:20:59.271470Z","shell.execute_reply.started":"2022-11-20T23:20:57.418779Z","shell.execute_reply":"2022-11-20T23:20:59.270321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Each record in the train_labels table contains \n  1. a patientId (one unique value per patient)\n  2. corresponding abnormality bounding box defined by the upper-left hand corner (x, y) coordinate \n  3. a target (either 0 or 1 for absence or presence of pneumonia, respectively) \n- There are many NaN values in four columns. But they seem to follow a trend. \n  - For each record in which we have any of value in the tuple (x,y,width,height) as NeN, the other 3 will also be NaN.\n   - This seems plausible. x and y are values of a tuple. Together they stand for a location, and width and height also form a pair. All four, together, define the bound of the abnormality in on the lung, called as opacity. \n   - All will be NaN together, or none will.\n  - If the values in the tuple (x,y,width,height) is NaN, then the value in the column 'Target' is definitely 0, that is, the patient is not pneumonic.\n   - Target, as per data dictionary, stands for whether the patient is pneuomonic or not. And in this study, we are looking for  the same through an analysis of images. ","metadata":{}},{"cell_type":"code","source":"print(\"Confirming the connection between pnuemonia and presence of 'abrnomality' on the lung. The number of unique Target values when the feature has value NaN:\", len(train_labels[train_labels.x.isna()]['Target'].unique()))","metadata":{"execution":{"iopub.status.busy":"2022-11-20T23:21:41.057310Z","iopub.execute_input":"2022-11-20T23:21:41.057685Z","iopub.status.idle":"2022-11-20T23:21:41.072433Z","shell.execute_reply.started":"2022-11-20T23:21:41.057636Z","shell.execute_reply":"2022-11-20T23:21:41.070645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Number of unqiue values turn out to be 1, which imply all patients for whom the record shows absence of 'abnormality', will be non-pneumonic.\n- Continuing, thus the missing values in the dataset have a reason. There are empty non-existent values conditioned on Target being 1 or zero. In other words, the Target column is dependent on the tuple of these four features. We don't need it. We won't remove it though. This is out prediction variable. Instead, we will match the values in this table withe images and perform out analysis. \n- Let's check if the dataset has duplicate values. And drop them if there are.","metadata":{}},{"cell_type":"code","source":"print('Number of Duplicated records in train_labels file:', train_class.duplicated().sum())\nprint('------------------------------------------------------------------------')\nprint('Dropping duplicate records from train_class')\ntrain_class.drop_duplicates(inplace = True)\nprint('Checking the number of Duplicated records in train_class file:', train_class.duplicated().sum())\n\nprint('------------------------------------------------------------')\nprint('Understanding the train_class file')\nprint(train_class.info())\nprint('------------------------------------------------------------------------')\nprint(train_class.head(15))","metadata":{"execution":{"iopub.status.busy":"2022-11-20T23:22:15.468646Z","iopub.execute_input":"2022-11-20T23:22:15.469020Z","iopub.status.idle":"2022-11-20T23:22:15.510496Z","shell.execute_reply.started":"2022-11-20T23:22:15.468989Z","shell.execute_reply":"2022-11-20T23:22:15.509489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Each record in the train_class table contains\n    1. a patientId\n    2. a class \n- There are no missing values, as is clear from the info() table.\n- From the same table,  it's clear that the number of records in the two tables: train_class and train_labels is same: 30227","metadata":{"id":"vppKyxJnsadr"}},{"cell_type":"code","source":"# checking for unique patientId values. \nprint('Number of unique patientId values in train_class: ', train_class['patientId'].nunique())\nprint('Number of unique patientId values in train_label: ', train_labels.patientId.nunique())\nprint('Distribution of the classes:')\ntrain_class['class'].value_counts().plot(kind='pie', autopct='%1.1f%%')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-11-20T23:22:20.544868Z","iopub.execute_input":"2022-11-20T23:22:20.545314Z","iopub.status.idle":"2022-11-20T23:22:20.793621Z","shell.execute_reply.started":"2022-11-20T23:22:20.545281Z","shell.execute_reply":"2022-11-20T23:22:20.791985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Number of patientIds are same in both the train_labels dataset and train_class.","metadata":{}},{"cell_type":"markdown","source":"### Meaning of labels\n\nLet's now understand what these labels mean.\n- 'Normal' indicates that the lung x-ray is normal, and so is the lung. \n- 'Lung Opacity' confirms that the the lung has opacity and is indicative of pneumonia.\n- 'No Lung Opacity / Not Normal' indicates that while pneumonia was determined not to be present, there was nonetheless some type of abnormality on the image and oftentimes this finding may mimic the appearance of true pneumonia. \n\nThe third label is going to be challenging. These images have some abnormality, but they are not pneumomic. Our machine learning algorithm should be able to read this into the images, and not get fooled. In other words, it should be able to correctly classify the abnormal images into 'with pnuemonia' and 'without pneumonia'.\n\nAlso, there's a relatively uniform split between the three classes, with nearly 2/3rd of the data comprising of no pneumonia (either completely normal or no lung opacity / not normal). Compared to most medical imaging datasets, where the prevalence of disease is quite low, this dataset has been significantly enriched with pathology.\n\nLet's analyze the distribution of the Target values in train_labels. ","metadata":{}},{"cell_type":"code","source":"print('Distribution of Target values','\\n',train_labels['Target'].value_counts())\nsns.countplot(x=train_labels['Target'])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-11-20T23:22:27.477128Z","iopub.execute_input":"2022-11-20T23:22:27.477478Z","iopub.status.idle":"2022-11-20T23:22:27.670837Z","shell.execute_reply.started":"2022-11-20T23:22:27.477449Z","shell.execute_reply":"2022-11-20T23:22:27.669935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 9555 records with Target = 1 and same number of records with class = Lung Opacity. We don't need to analyze the data with code to know that they corressspond to the same patientId. If the x-ray of a patient has region(s) of lung opacity, then they are pneumonic.","metadata":{}},{"cell_type":"markdown","source":"### merging train_labels & train_class into train_meta","metadata":{}},{"cell_type":"code","source":"train_meta = pd.merge(train_labels, train_class)\nprint('Checking info and a few random samples of newly created train_meta to assert everything is alright:')\nprint(train_meta.info())\nprint(train_meta.sample(15))","metadata":{"execution":{"iopub.status.busy":"2022-11-20T23:22:32.145627Z","iopub.execute_input":"2022-11-20T23:22:32.146021Z","iopub.status.idle":"2022-11-20T23:22:32.198211Z","shell.execute_reply.started":"2022-11-20T23:22:32.145988Z","shell.execute_reply":"2022-11-20T23:22:32.196970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A visual inspection confirms our ideas about 'Target' and 'class' and their relationship. Let us group the data according to them, and see if our assumptions were right with a bivariate countplot","metadata":{}},{"cell_type":"code","source":"sns.countplot(x=train_meta['Target'], hue=train_meta['class'])\nplt.title('Countplot of class wrt Target')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-11-20T23:22:36.846545Z","iopub.execute_input":"2022-11-20T23:22:36.846915Z","iopub.status.idle":"2022-11-20T23:22:37.088150Z","shell.execute_reply.started":"2022-11-20T23:22:36.846869Z","shell.execute_reply":"2022-11-20T23:22:37.086858Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There is no incongruence. If class = Lung Opacity, then Target = 1. And if Target = 1, class is necessarily Lung Opacity. In other words, all the 'abnormalities' with class = No Lung Opacity/Not Normal come under the category of Target=0.","metadata":{}},{"cell_type":"code","source":"print('Number of images in train set is:', len(train_path))\nprint('Number of patients in csv file as per their Id is:', train_meta.patientId.nunique())\nprint('Number of images in test set is:', len(test_path))","metadata":{"execution":{"iopub.status.busy":"2022-11-20T23:22:42.121361Z","iopub.execute_input":"2022-11-20T23:22:42.121716Z","iopub.status.idle":"2022-11-20T23:22:42.136739Z","shell.execute_reply.started":"2022-11-20T23:22:42.121679Z","shell.execute_reply":"2022-11-20T23:22:42.135630Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So far we have been doing good. Their are no inconsistencies. Let us check the first image in the train set as per its index value in the corresponding meta table.","metadata":{}},{"cell_type":"markdown","source":"# DICOM","metadata":{}},{"cell_type":"code","source":"first_img = train_meta.patientId[0]\nimg1_file = '../input/rsna-pneumonia-detection-challenge/stage_2_train_images/%s.dcm' %first_img\nimg1_meta = pydicom.read_file(img1_file)\n\nimg1_px = img1_meta.pixel_array\nprint('The size of the image is: ', img1_px.shape, '\\n')\nprint('The pixel values of the image as a numpy array are:\\n', img1_px, '\\n')\nprint('The image looks like:\\n', plt.imshow(img1_px))\nprint('The meta information saved in DICOM file is as follows:')\nprint(img1_meta)","metadata":{"execution":{"iopub.status.busy":"2022-11-20T23:22:45.259652Z","iopub.execute_input":"2022-11-20T23:22:45.260037Z","iopub.status.idle":"2022-11-20T23:22:45.684870Z","shell.execute_reply.started":"2022-11-20T23:22:45.260004Z","shell.execute_reply":"2022-11-20T23:22:45.683933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most of the standard headers containing patient identifable information have been anonymized (removed) so we are left with a relatively sparse set of metadata. The primary field we will be accessing is the underlying pixel data.\n\nA few inferences that can be made from the above meta data contained in the dicom\n- A lot of infomation from the dicom has been removed. For example, we don't have the patient's name which is the same as their id. Also, there is no physician name and study id, nor is there the DOB of the patient. \n - There are 1024 rows and columns in the image- which should be the standard size of all the images. We will check. Else we will have to resize all the images in one size before feeding them in a Deep Neural Network.\n - It is monochrome i.e. grayscale. \n - Every image takes up 8 bit of data.","metadata":{}},{"cell_type":"markdown","source":"## Bounding Boxes","metadata":{}},{"cell_type":"code","source":"print('Total records for class information values:', train_class.shape[0])\nprint('------------------------------------------------------------------------')\nprint('Total number of bounding boxes:', train_labels.shape[0])\nprint('------------------------------------------------------------------------')\nbox_df = train_meta.groupby('patientId').size().reset_index(name='boxes')\ntrain_meta = pd.merge(train_meta, box_df, on='patientId')\n\nprint('Distribution of number of boxes in the lungs of each patient:')\nbox_df = box_df.groupby('boxes').size().reset_index(name='patients')\nbox_df = pd.DataFrame(box_df)\nprint(box_df)","metadata":{"execution":{"iopub.status.busy":"2022-11-20T23:24:58.965672Z","iopub.execute_input":"2022-11-20T23:24:58.966052Z","iopub.status.idle":"2022-11-20T23:24:59.014065Z","shell.execute_reply.started":"2022-11-20T23:24:58.966021Z","shell.execute_reply":"2022-11-20T23:24:59.012969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualizing Images with bounding boxes","metadata":{}},{"cell_type":"code","source":"train_meta.head()","metadata":{"execution":{"iopub.status.busy":"2022-11-20T23:43:17.158544Z","iopub.execute_input":"2022-11-20T23:43:17.159405Z","iopub.status.idle":"2022-11-20T23:43:17.191625Z","shell.execute_reply.started":"2022-11-20T23:43:17.159360Z","shell.execute_reply":"2022-11-20T23:43:17.190731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"vars = ['PatientAge','PatientSex','ImagePath']\n\ndef process_dicom_data(df, path):\n    # adding new columns to the imported DataFrame with Null values\n    for var in vars:\n        df[var] = None\n    images = os.listdir(path)\n    #looping through each dicom image, extract the information from it, and \n    # add it to the DataFrame\n    for i, img_name in enumerate(images):\n        imagePath = os.path.join(path,img_name)\n        img_data = pydicom.read_file(imagePath)\n        idx = (df['patientId']==img_data.PatientID)\n        df.loc[idx,'PatientAge'] = pd.to_numeric(img_data.PatientAge)\n        df.loc[idx,'PatientSex'] = img_data.PatientSex\n        df.loc[idx, 'ImagePath'] = str.format(imagePath)\n        \n\nprint('Saving the parsed data as DataFrame for a neat visualization')\nprocess_dicom_data(train_meta,'../input/rsna-pneumonia-detection-challenge/stage_2_train_images')\n\nprint('Now, we subset a sample from the training set DataFrame so we have one example of every type of image.')\nprint('We do this by successively grouping the training dataset by Target, class, and number of boxes')","metadata":{"execution":{"iopub.status.busy":"2022-11-20T23:25:15.690484Z","iopub.execute_input":"2022-11-20T23:25:15.690829Z","iopub.status.idle":"2022-11-20T23:33:14.812717Z","shell.execute_reply.started":"2022-11-20T23:25:15.690799Z","shell.execute_reply":"2022-11-20T23:33:14.811398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_df = train_meta.\\\n    groupby(['Target','class', 'boxes']).\\\n    apply(lambda x: x[x['patientId']==x.sample(1)['patientId'].values[0]]).\\\n    reset_index(drop=True)\n\n\nprint('After that we plot the subsetted sample of images along with the boundings boxes on them, if any')\nfig, m_axs = plt.subplots(3, 2, figsize = (20, 20))\nfor c_ax, (c_path, c_rows) in zip(m_axs.flatten(),\n                    sample_df.groupby(['ImagePath'])):\n    c_dicom = pydicom.read_file(c_path)\n    c_ax.imshow(c_dicom.pixel_array, cmap='bone')\n    c_ax.set_title('{class}'.format(**c_rows.iloc[0,:]))\n    for i, (_, c_row) in enumerate(c_rows.dropna().iterrows()):\n        c_ax.plot(c_row['x'], c_row['y'], 's', label='{class}'.format(**c_row))\n        c_ax.add_patch(Rectangle(xy=(c_row['x'], c_row['y']),\n                                width=c_row['width'],\n                                height=c_row['height'], \n                                 alpha = 0.5))\n        if i==0: c_ax.legend()\n            \nprint('Thus our EDA is complete')","metadata":{"execution":{"iopub.status.busy":"2022-11-20T23:42:35.036270Z","iopub.execute_input":"2022-11-20T23:42:35.037409Z","iopub.status.idle":"2022-11-20T23:42:35.068054Z","shell.execute_reply.started":"2022-11-20T23:42:35.037371Z","shell.execute_reply":"2022-11-20T23:42:35.066908Z"},"trusted":true},"execution_count":null,"outputs":[]}]}