{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"#%tensorflow_version 2.x\nimport tensorflow\ntensorflow.__version__\n## ignore warnings\nimport warnings\nwarnings.filterwarnings('ignore')\n\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom matplotlib.patches import Rectangle\nfrom tqdm import tqdm, tqdm_notebook\nimport seaborn as sns\nimport pydicom as dcm\nfrom glob import glob\nfrom skimage.transform import resize\nfrom skimage import io, measure\nimport cv2, random\n\nimport tensorflow as tf\nfrom tensorflow import keras","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-01-29T08:43:40.827017Z","iopub.execute_input":"2022-01-29T08:43:40.827500Z","iopub.status.idle":"2022-01-29T08:43:42.199913Z","shell.execute_reply.started":"2022-01-29T08:43:40.827455Z","shell.execute_reply":"2022-01-29T08:43:42.199142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!ls ../input","metadata":{"_uuid":"1b6c13edf09c7d80145173edad847c34db1ba5f2","execution":{"iopub.status.busy":"2022-01-29T08:43:42.200898Z","iopub.execute_input":"2022-01-29T08:43:42.201106Z","iopub.status.idle":"2022-01-29T08:43:42.987871Z","shell.execute_reply.started":"2022-01-29T08:43:42.201067Z","shell.execute_reply":"2022-01-29T08:43:42.986971Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The several key items in this folder:\n* `stage_2_train_labels.csv`: CSV file containing training set patientIds and  labels (including bounding boxes)\n* `stage_2_detailed_class_info.csv`: CSV file containing detailed labels (explored further below)\n* `stage_2_train_images/`:  directory containing training set raw image (DICOM) files\n\nLet's go ahead and take a look at the first labels CSV file first:","metadata":{"_uuid":"88fb25d6864f224a0995a28140de532f80cb3e6b"}},{"cell_type":"code","source":"train_labels= pd.read_csv('../input/stage_2_train_labels.csv')\nprint('First five rows of Training set:\\n', train_labels.head())\nprint(train_labels.iloc[0])","metadata":{"_uuid":"b4c54974f191b477e46370c79254f75657ed2a85","scrolled":true,"execution":{"iopub.status.busy":"2022-01-29T08:43:42.989365Z","iopub.execute_input":"2022-01-29T08:43:42.989650Z","iopub.status.idle":"2022-01-29T08:43:43.115771Z","shell.execute_reply.started":"2022-01-29T08:43:42.989575Z","shell.execute_reply":"2022-01-29T08:43:43.115015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you can see, each row in the CSV file contains a `patientId` (one unique value per patient), a target (either 0 or 1 for absence or presence of pneumonia, respectively) and the corresponding abnormality bounding box defined by the upper-left hand corner (x, y) coordinate and its corresponding width and height. In this particular case, the patient does *not* have pneumonia and so the corresponding bounding box information is set to `NaN`. See an example case with pnuemonia here:","metadata":{"_uuid":"5435a871306c834a2282ae867a86861f9b80a610"}},{"cell_type":"code","source":"print(train_labels.iloc[4])","metadata":{"scrolled":true,"_uuid":"98c60a2ddf5bba070c0908d9bb705c0c5976ac7a","execution":{"iopub.status.busy":"2022-01-29T08:43:50.033090Z","iopub.execute_input":"2022-01-29T08:43:50.033684Z","iopub.status.idle":"2022-01-29T08:43:50.039464Z","shell.execute_reply.started":"2022-01-29T08:43:50.033622Z","shell.execute_reply":"2022-01-29T08:43:50.038443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Some information about the data field present in the 'stage_2_train_labels.csv' are:\n*   **patientId** - A patientId. Each patientId corresponds to a unique image (which we will see a little bit later) \n*   **x** - The upper-left x coordinate of the bounding box\n*   **y** - The upper-left y coordinate of the bounding box\n*   **width** - The width of the bounding box\n*   **height** - The height of the bounding box\n*   **Target** - The binary Target indicating whether this sample has evidence of pneumonia or not.","metadata":{}},{"cell_type":"markdown","source":"One important thing to keep in mind is that a given `patientId` may have **multiple** boxes if more than one area of pneumonia is detected (see below for example images).","metadata":{"_uuid":"7ad152504ec74be4ac1992ca669da7ad3ce3aac2"}},{"cell_type":"code","source":"# Number of entries in Train label dataframe:\nprint('The train_label dataframe has {} rows and {} columns.'.format(train_labels.shape[0], train_labels.shape[1]))","metadata":{"execution":{"iopub.status.busy":"2022-01-29T08:43:54.586780Z","iopub.execute_input":"2022-01-29T08:43:54.587402Z","iopub.status.idle":"2022-01-29T08:43:54.592238Z","shell.execute_reply.started":"2022-01-29T08:43:54.587330Z","shell.execute_reply":"2022-01-29T08:43:54.591248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Number of duplicates in patientId:\nprint('Number of unique patientId are: {}'.format(train_labels['patientId'].nunique()))","metadata":{"execution":{"iopub.status.busy":"2022-01-29T08:43:59.276720Z","iopub.execute_input":"2022-01-29T08:43:59.277280Z","iopub.status.idle":"2022-01-29T08:43:59.292582Z","shell.execute_reply.started":"2022-01-29T08:43:59.277231Z","shell.execute_reply":"2022-01-29T08:43:59.291811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thus, the dataset contains information about 30227 patients. Out of these 26684 patients, some of them have multiple entries in the dataset.","metadata":{}},{"cell_type":"code","source":"train_labels.drop_duplicates('patientId').shape[0]\n","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:07:31.207136Z","iopub.execute_input":"2022-01-29T09:07:31.207458Z","iopub.status.idle":"2022-01-29T09:07:31.224700Z","shell.execute_reply.started":"2022-01-29T09:07:31.207396Z","shell.execute_reply":"2022-01-29T09:07:31.223836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"one = train_labels[train_labels.Target == 1].drop_duplicates('patientId').shape[0]\nzero = train_labels[train_labels.Target == 0].drop_duplicates('patientId').shape[0]\ntotal = train_labels.drop_duplicates('patientId').shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:14:18.699654Z","iopub.execute_input":"2022-01-29T09:14:18.699968Z","iopub.status.idle":"2022-01-29T09:14:18.728175Z","shell.execute_reply.started":"2022-01-29T09:14:18.699912Z","shell.execute_reply":"2022-01-29T09:14:18.727602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'No of entries which has Pneumonia: {one} i.e., {round(one/total*100, 0)}%')\nprint(f'No of entries which don\\'t have Pneumonia: {zero} i.e., {round(zero/total, 0)}%')\n_ = train_labels.drop_duplicates('patientId').drop_duplicates('patientId')['Target'].value_counts().plot(kind = 'pie', autopct = '%.0f%%', labels = ['Negative', 'Positive'], figsize = (10, 6))","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:17:03.271765Z","iopub.execute_input":"2022-01-29T09:17:03.272091Z","iopub.status.idle":"2022-01-29T09:17:03.423997Z","shell.execute_reply.started":"2022-01-29T09:17:03.272034Z","shell.execute_reply":"2022-01-29T09:17:03.423136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thus, from the above pie chart it is clear that out of unique 26684 entries in the dataset, there are 20672 (i.e., 77%) entries in the dataset which corresponds to the entries of the patient Not having Pnuemonia whereas 6012 (i.e., 23%) entries corresponds to Positive case of Pneumonia.","metadata":{}},{"cell_type":"code","source":"# Checking nulls in bounding box columns:\nprint('Number of nulls in bounding box columns: {}'.format(train_labels[['x', 'y', 'width', 'height']].isnull().sum().to_dict()))","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:20:14.011210Z","iopub.execute_input":"2022-01-29T09:20:14.011552Z","iopub.status.idle":"2022-01-29T09:20:14.021973Z","shell.execute_reply.started":"2022-01-29T09:20:14.011496Z","shell.execute_reply":"2022-01-29T09:20:14.021056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thus, we can see that number of nulls in bounding box columns are equal to the number of 0's we have in the Target column.","metadata":{}},{"cell_type":"code","source":"bounding_box = train_labels.groupby('patientId').size().to_frame('number_of_boxes').reset_index()\ntrain_labels = train_labels.merge(bounding_box, on = 'patientId', how = 'left')\nprint('Number of patientIds per bounding box in the dataset: ')\n(bounding_box.groupby('number_of_boxes').size().to_frame('number_of_patientId').reset_index().set_index('number_of_boxes').sort_values(by = 'number_of_boxes'))","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:20:17.148832Z","iopub.execute_input":"2022-01-29T09:20:17.149108Z","iopub.status.idle":"2022-01-29T09:20:17.222982Z","shell.execute_reply.started":"2022-01-29T09:20:17.149064Z","shell.execute_reply":"2022-01-29T09:20:17.222246Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thus, there are 23286 unique patients which have only one entry in the dataset. It also has the patientsbounding box, 3266 with 2 bounding box, 119 with 3 bounding box and 13 with 4 bounding box coordinates.","metadata":{}},{"cell_type":"markdown","source":"- **stage_2_detailed_class_info.csv**\n\nIt provides detailed information about the type of positive or negative class for each image.\n","metadata":{}},{"cell_type":"code","source":"class_labels = pd.read_csv('../input/stage_2_detailed_class_info.csv')\nprint('First five rows of Class label dataset are:\\n', class_labels.head())","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:20:21.150524Z","iopub.execute_input":"2022-01-29T09:20:21.151081Z","iopub.status.idle":"2022-01-29T09:20:21.227548Z","shell.execute_reply.started":"2022-01-29T09:20:21.151035Z","shell.execute_reply":"2022-01-29T09:20:21.226887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Some information about the data field present in the 'stage_2_detailed_class_info.csv' are:\n*   **patientId** - A patientId. Each patientId corresponds to a unique image\n*   **class** - Have three values depending what is the current state of the patient's lung: 'No Lung Opacity / Not Normal', 'Normal' and 'Lung Opacity'.\n\n","metadata":{}},{"cell_type":"code","source":"# Number of entries in class_label dataframe:\nprint('The class_label dataframe has {} rows and {} columns.'.format(class_labels.shape[0], class_labels.shape[1]))","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:20:25.720692Z","iopub.execute_input":"2022-01-29T09:20:25.721300Z","iopub.status.idle":"2022-01-29T09:20:25.725400Z","shell.execute_reply.started":"2022-01-29T09:20:25.721247Z","shell.execute_reply":"2022-01-29T09:20:25.724814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Number of duplicates in patients:\nprint('Number of unique patientId are: {}'.format(class_labels['patientId'].nunique()))","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:20:29.960468Z","iopub.execute_input":"2022-01-29T09:20:29.961060Z","iopub.status.idle":"2022-01-29T09:20:29.975373Z","shell.execute_reply.started":"2022-01-29T09:20:29.961008Z","shell.execute_reply":"2022-01-29T09:20:29.974478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thus, the dataset contains information about 26684 patients (which is same as that of the train_labels dataframe).","metadata":{}},{"cell_type":"code","source":"def get_feature_distribution(data, feature):\n  # Count for each label\n  label_counts = data[feature].value_counts()\n  # Count the number of items in each class\n  total_samples = len(data)\n  print(\"Feature: {}\".format(feature))\n  for i in range(len(label_counts)):\n    label = label_counts.index[i]\n    count = label_counts.values[i]\n    percent = int((count / total_samples) * 10000) / 100\n    print(\"{:<30s}: {} which is {}% of the total data in the dataset\".format(label, count, percent))","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:20:35.589449Z","iopub.execute_input":"2022-01-29T09:20:35.589754Z","iopub.status.idle":"2022-01-29T09:20:35.594745Z","shell.execute_reply.started":"2022-01-29T09:20:35.589692Z","shell.execute_reply":"2022-01-29T09:20:35.593916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"get_feature_distribution(class_labels, 'class')","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:20:48.911578Z","iopub.execute_input":"2022-01-29T09:20:48.912151Z","iopub.status.idle":"2022-01-29T09:20:48.920181Z","shell.execute_reply.started":"2022-01-29T09:20:48.912105Z","shell.execute_reply":"2022-01-29T09:20:48.919451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"figsize = (10, 6)\n_ = class_labels['class'].value_counts().sort_index(ascending = False).plot(kind = 'pie', autopct = '%.0f%%').set_ylabel('')","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:21:03.371848Z","iopub.execute_input":"2022-01-29T09:21:03.372325Z","iopub.status.idle":"2022-01-29T09:21:03.481019Z","shell.execute_reply.started":"2022-01-29T09:21:03.372263Z","shell.execute_reply":"2022-01-29T09:21:03.480147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Checking nulls in class_labels:\nprint('Number of nulls in class columns: {}'.format(class_labels['class'].isnull().sum()))","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:21:07.890867Z","iopub.execute_input":"2022-01-29T09:21:07.891292Z","iopub.status.idle":"2022-01-29T09:21:07.897565Z","shell.execute_reply.started":"2022-01-29T09:21:07.891197Z","shell.execute_reply":"2022-01-29T09:21:07.896814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thus, none of the columns in class_labels has an empty row. ","metadata":{}},{"cell_type":"code","source":"# Checking whether each patientId has only one type of class or not\nclass_labels.groupby(['patientId'])['class'].nunique().max()","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:21:11.170583Z","iopub.execute_input":"2022-01-29T09:21:11.171006Z","iopub.status.idle":"2022-01-29T09:21:11.223236Z","shell.execute_reply.started":"2022-01-29T09:21:11.170958Z","shell.execute_reply":"2022-01-29T09:21:11.222246Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thus, we can say that each patientId is associated with only 1 class.","metadata":{}},{"cell_type":"code","source":"# Merging the two dataset - 'train_labels' and 'class_labels':\ntraining_data = pd.concat([train_labels, class_labels['class']], axis = 1)\nprint('After merging, the dataset looks like: \\n')\ntraining_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:21:14.630462Z","iopub.execute_input":"2022-01-29T09:21:14.630933Z","iopub.status.idle":"2022-01-29T09:21:14.658748Z","shell.execute_reply.started":"2022-01-29T09:21:14.630729Z","shell.execute_reply":"2022-01-29T09:21:14.658010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('After merge, the dataset has {} rows and {} columns.'.format(training_data.shape[0], training_data.shape[1]))","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:21:18.370669Z","iopub.execute_input":"2022-01-29T09:21:18.370964Z","iopub.status.idle":"2022-01-29T09:21:18.375521Z","shell.execute_reply.started":"2022-01-29T09:21:18.370917Z","shell.execute_reply":"2022-01-29T09:21:18.374721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Target and Class","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(nrows = 1, figsize = (12, 6))\ntemp = training_data.groupby('Target')['class'].value_counts()\ndata_target_class = pd.DataFrame(data = {'Values': temp.values}, index = temp.index).reset_index()\nsns.barplot(ax = ax, x = 'Target', y = 'Values', hue = 'class', data = data_target_class, palette = 'Set3')\nplt.title('Class and Target for Chest Exams')","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:21:24.639134Z","iopub.execute_input":"2022-01-29T09:21:24.639678Z","iopub.status.idle":"2022-01-29T09:21:24.965748Z","shell.execute_reply.started":"2022-01-29T09:21:24.639630Z","shell.execute_reply":"2022-01-29T09:21:24.964457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thus, **Target = 1** is associated with only **class = Lung Opacity** whereas **Target = 0** is associated with only **class = No Lung Opacity / Not Normal** as well as **Normal**.","metadata":{}},{"cell_type":"markdown","source":"#### Bounding Box Distribution","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 1, figsize = (7, 7))\ntarget_1 = training_data[training_data['Target'] == 1]\ntarget_sample = target_1.sample(2000)\ntarget_sample['xc'] = target_sample['x'] + target_sample['width'] / 2\ntarget_sample['yc'] = target_sample['y'] + target_sample['height'] / 2\nplt.title('Centers of Lung Opacity Rectangles (brown) over rectangles (yellow)\\nSample Size: 2000')\ntarget_sample.plot.scatter(x = 'xc', y = 'yc', xlim = (0, 1024), ylim = (0, 1024), ax = ax, alpha = 0.8, marker = '.', color = 'brown')\n\nfor i, crt_sample in target_sample.iterrows():\n    ax.add_patch(Rectangle(xy=(crt_sample['x'], crt_sample['y']),\n                width=crt_sample['width'],height=crt_sample['height'],alpha=3.5e-3, color=\"yellow\"))","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:21:30.501745Z","iopub.execute_input":"2022-01-29T09:21:30.502019Z","iopub.status.idle":"2022-01-29T09:21:35.105206Z","shell.execute_reply.started":"2022-01-29T09:21:30.501975Z","shell.execute_reply":"2022-01-29T09:21:35.104262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thus, we can see that the centers for the bounding box are spread out evenly across the Lungs. Though a large portion of the bounding box have their centers at the centers of the Lung, but some centers of the box are also located at the edges of lung. ","metadata":{}},{"cell_type":"markdown","source":"# Overview of DICOM files and medical images\n\nMedical images are stored in a special format known as DICOM files (`*.dcm`). They contain a combination of header metadata as well as underlying raw image arrays for pixel data. In Python, one popular library to access and manipulate DICOM files is the `pydicom` module. To use the `pydicom` library, first find the DICOM file for a given `patientId` by simply looking for the matching file in the `stage_2_train_images/` folder, and the use the `pydicom.read_file()` method to load the data:","metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a"}},{"cell_type":"code","source":"sample_patientId = train_labels['patientId'][0]\ndcm_file = '../input/stage_2_train_images/'+'{}.dcm'.format(sample_patientId)\ndcm_data = dcm.read_file(dcm_file)\n\nprint('Metadata of the image consists of \\n', dcm_data)\n","metadata":{"_uuid":"5f2c15162a0d1390624b42ef94d4f9e260be56ac","execution":{"iopub.status.busy":"2022-01-29T09:21:35.106430Z","iopub.execute_input":"2022-01-29T09:21:35.106641Z","iopub.status.idle":"2022-01-29T09:21:35.130834Z","shell.execute_reply.started":"2022-01-29T09:21:35.106602Z","shell.execute_reply":"2022-01-29T09:21:35.130064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most of the standard headers containing patient identifable information have been anonymized (removed) so we are left with a relatively sparse set of metadata. The primary field we will be accessing is the underlying pixel data as follows:\n\nFrom the above sample we can see that dicom file contains some of the information that can be used for further analysis such as sex, age, body part examined (which should be mostly chest), view position and modality. Size of this image is 1024 x 1024 (rows x columns).","metadata":{"_uuid":"2cfbe9eb43f4e4922c42481739046943767f765c"}},{"cell_type":"code","source":"import os","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:21:37.510629Z","iopub.execute_input":"2022-01-29T09:21:37.510918Z","iopub.status.idle":"2022-01-29T09:21:37.514866Z","shell.execute_reply.started":"2022-01-29T09:21:37.510872Z","shell.execute_reply":"2022-01-29T09:21:37.514040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Number of images in training images folders are: {}.'.format(len(os.listdir('../input/stage_2_train_images'))))","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:21:39.559232Z","iopub.execute_input":"2022-01-29T09:21:39.559575Z","iopub.status.idle":"2022-01-29T09:21:40.924671Z","shell.execute_reply.started":"2022-01-29T09:21:39.559502Z","shell.execute_reply":"2022-01-29T09:21:40.923579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thus, we can see that in the training images folder we have just 26684 images which is same as that of unique patientId's present in either of the csv files. Thus, we can say that **each of the unique patientId's present in either of the csv files corresponds to an image present in the folder**.","metadata":{}},{"cell_type":"code","source":"training_image_path = '../input/stage_2_train_images/'\n\nimages = pd.DataFrame({'path': glob(os.path.join(training_image_path, '*.dcm'))})\nimages['patientId'] = images['path'].map(lambda x:os.path.splitext(os.path.basename(x))[0])\nprint('Columns in the training images dataframe: {}'.format(list(images.columns)))","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:21:42.750602Z","iopub.execute_input":"2022-01-29T09:21:42.751155Z","iopub.status.idle":"2022-01-29T09:21:42.902258Z","shell.execute_reply.started":"2022-01-29T09:21:42.751106Z","shell.execute_reply":"2022-01-29T09:21:42.901532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Merging the images dataframe with training_data dataframe\ntraining_data = training_data.merge(images, on = 'patientId', how = 'left')\nprint('After merging the two dataframe, the training_data has {} rows and {} columns.'.format(training_data.shape[0], training_data.shape[1]))","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:21:45.551517Z","iopub.execute_input":"2022-01-29T09:21:45.551801Z","iopub.status.idle":"2022-01-29T09:21:45.585742Z","shell.execute_reply.started":"2022-01-29T09:21:45.551753Z","shell.execute_reply":"2022-01-29T09:21:45.584708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('The training_data dataframe as of now stands like\\n')\ntraining_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:21:47.951491Z","iopub.execute_input":"2022-01-29T09:21:47.952041Z","iopub.status.idle":"2022-01-29T09:21:47.981055Z","shell.execute_reply.started":"2022-01-29T09:21:47.951995Z","shell.execute_reply":"2022-01-29T09:21:47.980344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"columns_to_add = ['Modality', 'PatientAge', 'PatientSex', 'BodyPartExamined', 'ViewPosition', 'ConversionType', 'Rows', 'Columns', 'PixelSpacing']\n\ndef parse_dicom_data(data_df, data_path):\n  for col in columns_to_add:\n    data_df[col] = None\n  image_names = os.listdir('../input/stage_2_train_images/')\n  \n  for i, img_name in tqdm_notebook(enumerate(image_names)):\n    imagepath = os.path.join('../input/stage_2_train_images/', img_name)\n    data_img = dcm.read_file(imagepath)\n    idx = (data_df['patientId'] == data_img.PatientID)\n    data_df.loc[idx, 'Modality'] = data_img.Modality\n    data_df.loc[idx, 'PatientAge'] = pd.to_numeric(data_img.PatientAge)\n    data_df.loc[idx, 'PatientSex'] = data_img.PatientSex\n    data_df.loc[idx, 'BodyPartExamined'] = data_img.BodyPartExamined\n    data_df.loc[idx, 'ViewPosition'] = data_img.ViewPosition\n    data_df.loc[idx, 'ConversionType'] = data_img.ConversionType\n    data_df.loc[idx, 'Rows'] = data_img.Rows\n    data_df.loc[idx, 'Columns'] = data_img.Columns\n    data_df.loc[idx, 'PixelSpacing'] = str.format(\"{:4.3f}\", data_img.PixelSpacing[0])","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:21:54.759851Z","iopub.execute_input":"2022-01-29T09:21:54.760440Z","iopub.status.idle":"2022-01-29T09:21:54.766602Z","shell.execute_reply.started":"2022-01-29T09:21:54.760385Z","shell.execute_reply":"2022-01-29T09:21:54.766029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"parse_dicom_data(training_data, '../input/stage_2_train_images/')","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:21:58.478455Z","iopub.execute_input":"2022-01-29T09:21:58.478998Z","iopub.status.idle":"2022-01-29T09:41:44.645143Z","shell.execute_reply.started":"2022-01-29T09:21:58.478950Z","shell.execute_reply":"2022-01-29T09:41:44.644175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('So after parsing the information from the dicom images, our training_data dataframe has {} rows and {} columns and it looks like:\\n'.format(training_data.shape[0], training_data.shape[1]))\ntraining_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:41:44.646732Z","iopub.execute_input":"2022-01-29T09:41:44.646983Z","iopub.status.idle":"2022-01-29T09:41:44.693877Z","shell.execute_reply.started":"2022-01-29T09:41:44.646937Z","shell.execute_reply":"2022-01-29T09:41:44.692924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Saving the training_data for further use:\ntraining_data.to_pickle('training_data.pkl')","metadata":{"execution":{"iopub.status.busy":"2022-01-15T06:34:08.988374Z","iopub.execute_input":"2022-01-15T06:34:08.988651Z","iopub.status.idle":"2022-01-15T06:34:09.295832Z","shell.execute_reply.started":"2022-01-15T06:34:08.988589Z","shell.execute_reply":"2022-01-15T06:34:09.29448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### EDA on this saved training data:","metadata":{}},{"cell_type":"markdown","source":"#### Modality","metadata":{}},{"cell_type":"code","source":"print('Modality for the images obtained is: {} \\n'.format(training_data['Modality'].unique()[0]))","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:41:44.694942Z","iopub.execute_input":"2022-01-29T09:41:44.695284Z","iopub.status.idle":"2022-01-29T09:41:44.704140Z","shell.execute_reply.started":"2022-01-29T09:41:44.695236Z","shell.execute_reply":"2022-01-29T09:41:44.703424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Body Part Examined","metadata":{}},{"cell_type":"code","source":"print('The images obtained are of {} areas.'.format(training_data['BodyPartExamined'].unique()[0]))","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:41:44.705436Z","iopub.execute_input":"2022-01-29T09:41:44.705653Z","iopub.status.idle":"2022-01-29T09:41:44.717512Z","shell.execute_reply.started":"2022-01-29T09:41:44.705612Z","shell.execute_reply":"2022-01-29T09:41:44.716447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Understanding Different Positions","metadata":{}},{"cell_type":"code","source":"get_feature_distribution(training_data.drop_duplicates('patientId'), 'ViewPosition')","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:43:16.669194Z","iopub.execute_input":"2022-01-29T09:43:16.669545Z","iopub.status.idle":"2022-01-29T09:43:16.701356Z","shell.execute_reply.started":"2022-01-29T09:43:16.669493Z","shell.execute_reply":"2022-01-29T09:43:16.700282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As seen above, two View Positions that are in the training dataset are AP (Anterior/Posterior) and PA (Posterior/Anterior). These type of X-rays are mostly used to obtain the front-view. Apart from front-view, a lateral image is usually taken to complement the front-view.\n- **Posterior/Anterior (PA)**: Here the chest radiograph is acquired by passing the X-Ray beam from the patient's posterior (back) part of the chest  to the anterior (front) part. While obtaining the image patient is asked to stand with their chest against the film. In this image, the hear is on the right side of the image as one looks at it. These are of higher quality and assess the heart size more accurately\n- **Anterior/Posterior (AP)**: At times it is not possible for radiographers to acquire a PA chest X-ray. This is usually because the patient is too unwell to stand. In these images the size of Heart is exaggerated.","metadata":{}},{"cell_type":"code","source":"print('The distribution of View Position when there is an evidence of Pneumonia:\\n')\n_ = training_data.drop_duplicates('patientId').loc[training_data['Target'] == 1, 'ViewPosition'].value_counts().sort_index(ascending = False).plot(kind = 'pie', autopct = '%.0f%%').set_ylabel('')","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:43:52.219403Z","iopub.execute_input":"2022-01-29T09:43:52.219979Z","iopub.status.idle":"2022-01-29T09:43:52.335395Z","shell.execute_reply.started":"2022-01-29T09:43:52.219922Z","shell.execute_reply":"2022-01-29T09:43:52.334473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Plot x and y centers of bounding box')\nbboxes = training_data[training_data['Target'] == 1]\nbboxes['xw'] = bboxes['x'] + bboxes['width']/2\nbboxes['yh'] = bboxes['y'] + bboxes['height']/2\n\ng = sns.jointplot(x = bboxes['xw'], y = bboxes['yh'], data = bboxes,\n                  kind = 'hex', alpha = 0.5, size = 8)\nplt.suptitle('Bounding Box location when there is an evidence of Pneumonia')\nplt.tight_layout()\nplt.subplots_adjust(top = 0.95)","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:44:08.419237Z","iopub.execute_input":"2022-01-29T09:44:08.419841Z","iopub.status.idle":"2022-01-29T09:44:09.582748Z","shell.execute_reply.started":"2022-01-29T09:44:08.419788Z","shell.execute_reply":"2022-01-29T09:44:09.581811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def bboxes_scatter(data, color_point, color_window, text):\n  fig, ax = plt.subplots(1, 1, figsize = (7, 7))\n  plt.title('Plotting centers of Lung Opacity\\n{}'.format(text))\n  data.plot.scatter(x = 'xw', y = 'yh', xlim = (0, 1024), ylim = (0, 1024), ax = ax, alpha = 0.8, marker = \".\", color = color_point)\n  for i, crt_sample in data.iterrows():\n    ax.add_patch(Rectangle(xy = (crt_sample['x'], crt_sample['y']), width = crt_sample['width'], height = crt_sample['height'], alpha = 3.5e-3, color = color_window))","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:44:16.174545Z","iopub.execute_input":"2022-01-29T09:44:16.175117Z","iopub.status.idle":"2022-01-29T09:44:16.179835Z","shell.execute_reply.started":"2022-01-29T09:44:16.175066Z","shell.execute_reply":"2022-01-29T09:44:16.179291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_PA = bboxes[bboxes['ViewPosition'] == 'PA'].sample(1000)\ndata_AP = bboxes[bboxes['ViewPosition'] == 'AP'].sample(1000)\n\nbboxes_scatter(data_PA, 'green', 'yellow', 'ViewPosition = PA')","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:44:23.199502Z","iopub.execute_input":"2022-01-29T09:44:23.200066Z","iopub.status.idle":"2022-01-29T09:44:25.428731Z","shell.execute_reply.started":"2022-01-29T09:44:23.200016Z","shell.execute_reply":"2022-01-29T09:44:25.428079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bboxes_scatter(data_AP, 'blue', 'red', 'ViewPosition = AP')","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:44:29.738283Z","iopub.execute_input":"2022-01-29T09:44:29.738894Z","iopub.status.idle":"2022-01-29T09:44:32.326263Z","shell.execute_reply.started":"2022-01-29T09:44:29.738840Z","shell.execute_reply":"2022-01-29T09:44:32.324799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that the centers of the box are spread across the entire region of the Lungs. Both of the cases (PA and AP) seem to have outliers in them.  ","metadata":{}},{"cell_type":"markdown","source":"#### Conversion Type","metadata":{}},{"cell_type":"code","source":"print('Conversion Type for the data in Training Data: ', training_data['ConversionType'].unique()[0])","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:44:39.939354Z","iopub.execute_input":"2022-01-29T09:44:39.939653Z","iopub.status.idle":"2022-01-29T09:44:39.948886Z","shell.execute_reply.started":"2022-01-29T09:44:39.939607Z","shell.execute_reply":"2022-01-29T09:44:39.947981Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Rows and Columns","metadata":{}},{"cell_type":"code","source":"print(f'The training images has {training_data.Rows.unique()[0]} rows and {training_data.Columns.unique()[0]} columns.')","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:44:48.019224Z","iopub.execute_input":"2022-01-29T09:44:48.019536Z","iopub.status.idle":"2022-01-29T09:44:48.025830Z","shell.execute_reply.started":"2022-01-29T09:44:48.019489Z","shell.execute_reply":"2022-01-29T09:44:48.024977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Patient Sex","metadata":{}},{"cell_type":"code","source":"def drawgraphs(data_file, columns, hue = False, width = 15, showdistribution = True):\n  if (hue):\n    print('Creating graph for: {} and {}'.format(columns, hue))\n  else:  \n    print('Creating graph for : {}'.format(columns))\n  length = len(columns) * 6\n  total = float(len(data_file))\n\n  fig, axes = plt.subplots(nrows = len(columns) if len(columns) > 1 else 1, ncols = 1, figsize = (width, length))\n  for index, content in enumerate(columns):\n    plt.title(content)\n\n    currentaxes = 0\n    if (len(columns) > 1):\n      currentaxes = axes[index]\n    else:\n      currentaxes = axes\n\n    if (hue):\n      sns.countplot(x = columns[index], data = data_file, ax = currentaxes, hue = hue)\n    else:\n      sns.countplot(x = columns[index], data = data_file, ax = currentaxes)\n\n    if(showdistribution):\n      for p in (currentaxes.patches):\n        height = p.get_height()\n        if (height > 0 and total > 0):\n          currentaxes.text(p.get_x() + p.get_width()/2., height + 3, '{:1.2f}%'.format(100*height/total), ha = \"center\")","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:44:53.168486Z","iopub.execute_input":"2022-01-29T09:44:53.168817Z","iopub.status.idle":"2022-01-29T09:44:53.176442Z","shell.execute_reply.started":"2022-01-29T09:44:53.168745Z","shell.execute_reply":"2022-01-29T09:44:53.175221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"get_feature_distribution(training_data.drop_duplicates('patientId'), 'PatientSex')","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:45:11.463517Z","iopub.execute_input":"2022-01-29T09:45:11.463972Z","iopub.status.idle":"2022-01-29T09:45:11.493048Z","shell.execute_reply.started":"2022-01-29T09:45:11.463928Z","shell.execute_reply":"2022-01-29T09:45:11.492226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"drawgraphs(data_file = training_data.drop_duplicates('patientId'), columns = ['PatientSex'], hue = 'class', width = 20, showdistribution = True)","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:45:24.044225Z","iopub.execute_input":"2022-01-29T09:45:24.044538Z","iopub.status.idle":"2022-01-29T09:45:24.448320Z","shell.execute_reply.started":"2022-01-29T09:45:24.044488Z","shell.execute_reply":"2022-01-29T09:45:24.447417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"drawgraphs(data_file = training_data.drop_duplicates('patientId'), columns = ['PatientSex'], hue = 'Target', width = 20, showdistribution = True)","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:45:35.359501Z","iopub.execute_input":"2022-01-29T09:45:35.360077Z","iopub.status.idle":"2022-01-29T09:45:35.711249Z","shell.execute_reply.started":"2022-01-29T09:45:35.360028Z","shell.execute_reply":"2022-01-29T09:45:35.710331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thus, we can see that the number of Male patients suffering from Pneumonia is greater when compared with that of Females.","metadata":{}},{"cell_type":"code","source":"data_male = bboxes[bboxes['PatientSex'] == 'M'].sample(1000)\ndata_female = bboxes[bboxes['PatientSex'] == 'F'].sample(1000)\n\nbboxes_scatter(data_male, \"darkblue\", \"blue\", \"Patients Sex: Male\")","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:45:48.559641Z","iopub.execute_input":"2022-01-29T09:45:48.560105Z","iopub.status.idle":"2022-01-29T09:45:51.133460Z","shell.execute_reply.started":"2022-01-29T09:45:48.560048Z","shell.execute_reply":"2022-01-29T09:45:51.132396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bboxes_scatter(data_female, \"red\", \"magenta\", \"Patients Sex: Female\")","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:45:51.918426Z","iopub.execute_input":"2022-01-29T09:45:51.918726Z","iopub.status.idle":"2022-01-29T09:45:54.441869Z","shell.execute_reply.started":"2022-01-29T09:45:51.918671Z","shell.execute_reply":"2022-01-29T09:45:54.440424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The centres for the bounding box in both the cases are spread out evenly across the entire lung with a slight number of outliers.","metadata":{}},{"cell_type":"markdown","source":"#### Patient Age","metadata":{}},{"cell_type":"code","source":"print('The minimum and maximum recorded age of the patients are {} and {} respectively.'.format(training_data['PatientAge'].min(), training_data['PatientAge'].max()))","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:46:02.559446Z","iopub.execute_input":"2022-01-29T09:46:02.559739Z","iopub.status.idle":"2022-01-29T09:46:02.565941Z","shell.execute_reply.started":"2022-01-29T09:46:02.559691Z","shell.execute_reply":"2022-01-29T09:46:02.565123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"age_25 = np.percentile(training_data['PatientAge'], 25)\nage_75 = np.percentile(training_data['PatientAge'], 75)\niqr_age = age_75 - age_25\ncutoff_age = 1.5 * iqr_age\n\nlow_lim_age = age_25 - cutoff_age\nupp_lim_age = age_75 + cutoff_age\n\noutlier_age = [x for x in training_data['PatientAge'] if x < low_lim_age or x > upp_lim_age]\nprint('The number of outliers in `PatientAge` out of 30277 records are: ', len(outlier_age))\nprint('\\nThe ages which are in the outlier categories are:', outlier_age)\n\nfig = plt.figure(figsize = (10, 6))\nsns.boxplot(training_data['PatientAge'], orient = 'h').set_title('Outliers in PatientAge')","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:46:07.388929Z","iopub.execute_input":"2022-01-29T09:46:07.390152Z","iopub.status.idle":"2022-01-29T09:46:07.620055Z","shell.execute_reply.started":"2022-01-29T09:46:07.390085Z","shell.execute_reply":"2022-01-29T09:46:07.619267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thus, we can say that the ages like 148, 150, 151, 153 and 155 are mistakes. We can trim these outlier values to a somewhat lower value say 100 so that the max age of the patient will be 100.","metadata":{}},{"cell_type":"code","source":"drawgraphs(data_file = training_data.drop_duplicates('patientId'), columns = ['PatientAge'], width = 20, showdistribution = True)","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:46:26.229055Z","iopub.execute_input":"2022-01-29T09:46:26.229539Z","iopub.status.idle":"2022-01-29T09:46:27.354106Z","shell.execute_reply.started":"2022-01-29T09:46:26.229489Z","shell.execute_reply":"2022-01-29T09:46:27.353346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here, we can see that the maximum number of patient are of 58 years old. In order to have a more clear idea, we will introduce a new column where the patients will be placed in an age group like (0, 10), (10, 20) etc.","metadata":{}},{"cell_type":"code","source":"print('Removing the outliers from `PatientAge`')\ntraining_data['PatientAge'] = training_data['PatientAge'].clip(training_data['PatientAge'].min(), 100)\ntraining_data['PatientAge'].describe().astype(int)","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:46:46.809496Z","iopub.execute_input":"2022-01-29T09:46:46.809933Z","iopub.status.idle":"2022-01-29T09:46:46.825052Z","shell.execute_reply.started":"2022-01-29T09:46:46.809893Z","shell.execute_reply":"2022-01-29T09:46:46.824404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Distribution of `PatientAge`: Overall and Target = 1')\nfig = plt.figure(figsize = (10, 6))\n\nax = fig.add_subplot(121)\ng = (sns.distplot(training_data['PatientAge']).set_title('Distribution of PatientAge'))\n\nax = fig.add_subplot(122)\ng = (sns.distplot(training_data.drop_duplicates('patientId').loc[training_data['Target'] == 1, 'PatientAge']).set_title('Distribution of PatientAge vs PnuemoniaEvidence'))","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:46:58.149191Z","iopub.execute_input":"2022-01-29T09:46:58.149495Z","iopub.status.idle":"2022-01-29T09:46:58.528261Z","shell.execute_reply.started":"2022-01-29T09:46:58.149448Z","shell.execute_reply":"2022-01-29T09:46:58.527716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"custom_array = np.linspace(0, 100, 11)\ntraining_data['PatientAgeBins'] = pd.cut(training_data['PatientAge'], custom_array)\ntraining_data.drop_duplicates('patientId')['PatientAgeBins'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:47:21.399408Z","iopub.execute_input":"2022-01-29T09:47:21.399742Z","iopub.status.idle":"2022-01-29T09:47:21.434269Z","shell.execute_reply.started":"2022-01-29T09:47:21.399683Z","shell.execute_reply":"2022-01-29T09:47:21.433450Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thus, we can see that the maximum number of patients belong to the age group of (50, 60] whereas the least belong to (90, 100]","metadata":{}},{"cell_type":"code","source":"print('After adding the bin column, the dataset turns out to be:\\n')\ntraining_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:47:34.219249Z","iopub.execute_input":"2022-01-29T09:47:34.219756Z","iopub.status.idle":"2022-01-29T09:47:34.260376Z","shell.execute_reply.started":"2022-01-29T09:47:34.219690Z","shell.execute_reply":"2022-01-29T09:47:34.259497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"drawgraphs(data_file = training_data.drop_duplicates('patientId'), columns = ['PatientAgeBins'], width = 20, showdistribution = True)","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:47:41.038543Z","iopub.execute_input":"2022-01-29T09:47:41.038978Z","iopub.status.idle":"2022-01-29T09:47:41.419819Z","shell.execute_reply.started":"2022-01-29T09:47:41.038937Z","shell.execute_reply":"2022-01-29T09:47:41.418498Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"drawgraphs(data_file = training_data.drop_duplicates('patientId'), columns = ['PatientAgeBins'], hue = 'PatientSex', width = 20, showdistribution = True)","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:47:47.999901Z","iopub.execute_input":"2022-01-29T09:47:48.000521Z","iopub.status.idle":"2022-01-29T09:47:48.495021Z","shell.execute_reply.started":"2022-01-29T09:47:48.000447Z","shell.execute_reply":"2022-01-29T09:47:48.494316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"drawgraphs(data_file = training_data.drop_duplicates('patientId'), columns = ['PatientAgeBins'], hue = 'class', width = 20, showdistribution = True)","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:47:55.359263Z","iopub.execute_input":"2022-01-29T09:47:55.359630Z","iopub.status.idle":"2022-01-29T09:47:55.935406Z","shell.execute_reply.started":"2022-01-29T09:47:55.359581Z","shell.execute_reply":"2022-01-29T09:47:55.934511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"drawgraphs(data_file = training_data.drop_duplicates('patientId'), columns = ['PatientAgeBins'], hue = 'Target', width = 20, showdistribution = True)","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:48:01.999283Z","iopub.execute_input":"2022-01-29T09:48:01.999592Z","iopub.status.idle":"2022-01-29T09:48:02.465371Z","shell.execute_reply.started":"2022-01-29T09:48:01.999544Z","shell.execute_reply":"2022-01-29T09:48:02.464507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above three plots we can infer that the maximum percentage of Male and Females, Patients with “No Lung Opacity / Not Normal” as well as “Lung Opacity” classes and Patients having Pneumonia all lies in the [50, 60] age group.","metadata":{}},{"cell_type":"code","source":"data_age_19 = bboxes[bboxes['PatientAge'] < 20]\ndata_age_20_34 = bboxes[(bboxes['PatientAge'] >= 20) & (bboxes['PatientAge'] < 35)]\ndata_age_35_49 = bboxes[(bboxes['PatientAge'] >= 35) & (bboxes['PatientAge'] < 50)]\ndata_age_50_64 = bboxes[(bboxes['PatientAge'] >= 50) & (bboxes['PatientAge'] < 65)]\ndata_age_65 = bboxes[bboxes['PatientAge'] >= 65]\n\nbboxes_scatter(data_age_19,'blue', 'red', 'Patient Age: 1-19 years')","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:48:07.298505Z","iopub.execute_input":"2022-01-29T09:48:07.298805Z","iopub.status.idle":"2022-01-29T09:48:09.397510Z","shell.execute_reply.started":"2022-01-29T09:48:07.298761Z","shell.execute_reply":"2022-01-29T09:48:09.396495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bboxes_scatter(data_age_20_34, 'blue', 'red', 'Patient Age: 20-34 years')","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:48:13.098510Z","iopub.execute_input":"2022-01-29T09:48:13.099104Z","iopub.status.idle":"2022-01-29T09:48:18.005606Z","shell.execute_reply.started":"2022-01-29T09:48:13.099053Z","shell.execute_reply":"2022-01-29T09:48:18.004681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bboxes_scatter(data_age_35_49, 'blue', 'red', 'Patient Age: 35-49 years')","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:48:18.006977Z","iopub.execute_input":"2022-01-29T09:48:18.007322Z","iopub.status.idle":"2022-01-29T09:48:23.551639Z","shell.execute_reply.started":"2022-01-29T09:48:18.007239Z","shell.execute_reply":"2022-01-29T09:48:23.550618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bboxes_scatter(data_age_50_64, 'blue', 'red', 'Patient Age: 50-64 years')","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:48:24.958292Z","iopub.execute_input":"2022-01-29T09:48:24.958775Z","iopub.status.idle":"2022-01-29T09:48:31.614437Z","shell.execute_reply.started":"2022-01-29T09:48:24.958725Z","shell.execute_reply":"2022-01-29T09:48:31.613364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bboxes_scatter(data_age_65, 'blue', 'red', 'Patient Age: 65+ years')","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:48:31.615694Z","iopub.execute_input":"2022-01-29T09:48:31.615962Z","iopub.status.idle":"2022-01-29T09:48:34.931902Z","shell.execute_reply.started":"2022-01-29T09:48:31.615909Z","shell.execute_reply":"2022-01-29T09:48:34.931259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Plotting DICOM Images","metadata":{}},{"cell_type":"code","source":"def show_dicom_images(data, df, img_path):\n  img_data = list(data.T.to_dict().values())\n  f, ax = plt.subplots(3, 3, figsize = (16, 18))\n  \n  for i, row in enumerate(img_data):\n    image = row['patientId'] + '.dcm'\n    path = os.path.join(img_path, image)\n    data = dcm.read_file(path)\n    rows = df[df['patientId'] == row['patientId']]\n    age = rows.PatientAge.unique().tolist()[0]\n    sex = data.PatientSex\n    part = data.BodyPartExamined\n    vp = data.ViewPosition\n    modality = data.Modality\n    data_img = dcm.dcmread(path)\n    ax[i//3, i%3].imshow(data_img.pixel_array, cmap = plt.cm.bone)\n    ax[i//3, i%3].axis('off')\n    ax[i//3, i%3].set_title('ID: {}\\nAge: {}, Sex: {}, Part: {}, VP: {}, Modality: {}\\nTarget: {}, Class: {}\\nWindow: {}:{}:{}:{}'\\\n                            .format(row['patientId'], age, sex, part,\n                                    vp, modality, row['Target'],\n                                    row['class'], row['x'],\n                                    row['y'], row['width'],\n                                    row['height']))\n    box_data = list(rows.T.to_dict().values())\n    \n    for j, row in enumerate(box_data):\n      ax[i//3, i%3].add_patch(Rectangle(xy = (row['x'], row['y']),\n                                        width = row['width'], height = row['height'],\n                                        color = 'blue', alpha = 0.15))\n  plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:48:35.728698Z","iopub.execute_input":"2022-01-29T09:48:35.729134Z","iopub.status.idle":"2022-01-29T09:48:35.739699Z","shell.execute_reply.started":"2022-01-29T09:48:35.729081Z","shell.execute_reply":"2022-01-29T09:48:35.738748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Target = 0","metadata":{}},{"cell_type":"code","source":"show_dicom_images(data = training_data.loc[(training_data['Target'] == 0)].sample(9),\n                  df = training_data, img_path = '../input/stage_2_train_images/')","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:48:41.378222Z","iopub.execute_input":"2022-01-29T09:48:41.378873Z","iopub.status.idle":"2022-01-29T09:48:43.200767Z","shell.execute_reply.started":"2022-01-29T09:48:41.378797Z","shell.execute_reply":"2022-01-29T09:48:43.199801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As the above subplots are of the images which belong to either \"Normal\" or \"No Lung Opacity / Not Normal\", hence no bounding box is observed.","metadata":{}},{"cell_type":"markdown","source":"- Target = 1","metadata":{}},{"cell_type":"code","source":"show_dicom_images(data = training_data.loc[(training_data['Target'] == 1)].sample(9),\n                  df = training_data, img_path = '../input/stage_2_train_images/')","metadata":{"execution":{"iopub.status.busy":"2022-01-29T09:48:47.818375Z","iopub.execute_input":"2022-01-29T09:48:47.818734Z","iopub.status.idle":"2022-01-29T09:48:49.594432Z","shell.execute_reply.started":"2022-01-29T09:48:47.818689Z","shell.execute_reply":"2022-01-29T09:48:49.593733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the above subplots, we can see that the area covered by the box (in blue colour) depicts the area of interest i.e., the area in which the opacity is observed in the Lungs.","metadata":{}},{"cell_type":"markdown","source":"### Conclusion\n- The training dataset (both of the csv files and the training image folder) contains information of 26684 patients (unique)\n- Out of these 26684 unique patients some of these have multiple entries in the both of the csv files\n- Most of the recorded patient belong to Target = 0 (i.e., they don't have Pneumonia)\n- Some of the patients have more than one bounding box. The maximum being 4\n- The classes \"No Lung Opacity / Not Normal\" and \"Normal\" is associated with Target = 0 whereas \"Lung Opacity\" belong to Target = 1\n- The images are present in dicom format, from which information like PatientAge, PatientSex, ViewPosition etc are obtained\n- There are two ways from which images were obtained: AP and PA. The age ranges from 1-155 (which were further clipped to 100)\n- The centers of the bounding box are spread out over the entire region of the lungs. But there are some centers which are outliers.","metadata":{}}]}