{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"<h1><center><font size=\"6\">RSNA Pneumonia Detection EDA</font></center></h1>\n\n<center><img src=\"https://www.rsna.org/images/rsna/home/line_r.svg\" width=\"500\"></img></center>\n\n# <a id='0'>Content</a>\n\n- <a href='#1'>Introduction</a>  \n- <a href='#2'>Prepare the data analysis</a>  \n    -<a href='#21'>Load packages</a>  \n     -<a href='#21'>Load the data</a>  \n- <a href='#3'>Data exploration</a>   \n    -<a href='#31'>Missing data</a>  \n    -<a href='#32'>Merge train and class info data</a>  \n    -<a href='#33'>Explore DICOM data</a>  \n    -<a href='#34'>Add meta information from DICOM data</a>  \n    -<a href='#35'>Modality</a>  \n    -<a href='#36'>Body Part Examined</a>  \n    -<a href='#37'>View Position</a>  \n    -<a href='#38'>Conversion Type</a>  \n    -<a href='#39'>Rows and Columns</a>  \n    -<a href='#310'>Patient Age</a>  \n    -<a href='#311'>Patient Sex</a>  \n- <a href='#4'>Conclusions</a>    \n- <a href='#5'>References</a>    \n"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"from datetime import datetime\ndt_string = datetime.now().strftime(\"%d/%m/%Y %H:%M:%S\")\nprint(f\"Updated {dt_string} (GMT)\")","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","collapsed":true,"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":false},"cell_type":"markdown","source":"# <a id=\"1\">Introduction</a>  \n\nThis Kernel objective is to explore the dataset for RSNA Pneumonia Detection Challenge.   \n\nWe start by exploring the DICOM data, we extract then meta information from the DICOM files and visualize the various features of the DICOM images, grouped by age, sex.\n\nThe Kernel was modified to work with **stage_2** data instead of **stage_1** data.\n\n"},{"metadata":{"_uuid":"9de3f37dddf934a3c93a7923f1ac614091c4f4e6"},"cell_type":"markdown","source":"# <a id=\"2\">Prepare the data analysis</a>  \n\n## <a id=\"21\">Load packages</a>\n"},{"metadata":{"trusted":true,"_uuid":"7fc8fbd9605836d0ec56983a7bc17c443665cd3e"},"cell_type":"code","source":"import pandas as pd \nimport numpy as np\nimport matplotlib\nimport matplotlib.pyplot as plt\nfrom tqdm import tqdm_notebook\nfrom matplotlib.patches import Rectangle\nimport seaborn as sns\nimport pydicom as dcm\n%matplotlib inline \nIS_LOCAL = False\nimport os\nif(IS_LOCAL):\n    PATH=\"../input/rsna-pneumonia-detection-challenge\"\nelse:\n    PATH=\"../input/\"\nprint(os.listdir(PATH))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e2adaf3e4b5d3f0e489cc26127fd7d9b14c47603"},"cell_type":"markdown","source":"<a href=\"#0\"><font size=\"1\">Go to top</font></a>\n\n\n## <a id=\"22\">Load the data</a>\n\nLet's load the tabular data. There are two files:\n* Detailed class info;  \n* Train labels."},{"metadata":{"trusted":true,"_uuid":"4926dab56f167471a7da87c6ee0175e88bddf058"},"cell_type":"code","source":"class_info_df = pd.read_csv(PATH+'/stage_2_detailed_class_info.csv')\ntrain_labels_df = pd.read_csv(PATH+'/stage_2_train_labels.csv')                         ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1481b79b8a76543a497a4c062df66bdbfb173304"},"cell_type":"code","source":"print(f\"Detailed class info -  rows: {class_info_df.shape[0]}, columns: {class_info_df.shape[1]}\")\nprint(f\"Train labels -  rows: {train_labels_df.shape[0]}, columns: {train_labels_df.shape[1]}\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"65b465959b97dc93b9f9b5018533fd8e86363455"},"cell_type":"markdown","source":"Let's explore the two loaded files. We will take out a 5 rows samples from each dataset."},{"metadata":{"trusted":true,"_uuid":"e39a26a766d706c9f74cd863eebc8726028f084a"},"cell_type":"code","source":"class_info_df.sample(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a02c1a87ad86a6c7c37a1717a15d6a4ffac03393"},"cell_type":"code","source":"train_labels_df.sample(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"61af15d21af073ae6031d6148dcd567650eca3df"},"cell_type":"markdown","source":"In **class detailed info** dataset are given the detailed information about the type of positive or negative class associated with a certain patient.  \n\nIn **train labels** dataset are given the patient ID and the window (x min, y min, width and height of the) containing evidence of pneumonia."},{"metadata":{"_uuid":"14177e6457d4d4c5df81a7e47f80dece57811a44"},"cell_type":"markdown","source":"<a href=\"#0\"><font size=\"1\">Go to top</font></a>\n\n# <a id=\"1\">Data exploration</a>  \n\nLet's explore the data further."},{"metadata":{"_uuid":"a7d2dcc8dbf6b2dbae102f1584bfa135bcbc6b20"},"cell_type":"markdown","source":"## <a id=\"31\">Missing data</a>\n\nLet's check missing information in the two datasets. "},{"metadata":{"trusted":true,"_uuid":"7a5a0aa8ddd21e868269fe28456e8c8900766be6"},"cell_type":"code","source":"def missing_data(data):\n    total = data.isnull().sum().sort_values(ascending = False)\n    percent = (data.isnull().sum()/data.isnull().count()*100).sort_values(ascending = False)\n    return np.transpose(pd.concat([total, percent], axis=1, keys=['Total', 'Percent']))\nmissing_data(train_labels_df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d7999e02e08b06e4e7729b370eb06f9803ccc722"},"cell_type":"code","source":"missing_data(class_info_df)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"84de2d7400d906a929bbd3b2bc1d4ed658b63a91"},"cell_type":"markdown","source":"The percent missing for x,y, height and width in train labels represents the percent of the target **0** (not **Lung opacity**).\n\nLet's check the class distribution from class detailed info."},{"metadata":{"trusted":true,"_uuid":"985ceeec9997ae2b636b54aa51f95ee75f75559a"},"cell_type":"code","source":"f, ax = plt.subplots(1,1, figsize=(6,4))\ntotal = float(len(class_info_df))\nsns.countplot(class_info_df['class'],order = class_info_df['class'].value_counts().index, palette='Set3')\nfor p in ax.patches:\n    height = p.get_height()\n    ax.text(p.get_x()+p.get_width()/2.,\n            height + 3,\n            '{:1.2f}%'.format(100*height/total),\n            ha=\"center\") \nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"935f048f6cb6ca0c45a3f502524fd6400d004ba5"},"cell_type":"markdown","source":"Let's look into more details to the classes."},{"metadata":{"trusted":true,"_uuid":"da131894089e7441984a29a6d133b0aaca3a39d5"},"cell_type":"code","source":"def get_feature_distribution(data, feature):\n    # Get the count for each label\n    label_counts = data[feature].value_counts()\n\n    # Get total number of samples\n    total_samples = len(data)\n\n    # Count the number of items in each class\n    print(\"Feature: {}\".format(feature))\n    for i in range(len(label_counts)):\n        label = label_counts.index[i]\n        count = label_counts.values[i]\n        percent = int((count / total_samples) * 10000) / 100\n        print(\"{:<30s}:   {} or {}%\".format(label, count, percent))\n\nget_feature_distribution(class_info_df, 'class')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7dc14703d2711917f26612ef8d8dbb2710b5d480"},"cell_type":"markdown","source":"**No Lung Opacity / Not Normal** and **Normal** have together the same percent (**69.077%**) as the percent of missing values for target window in class details information.   \n\nIn the train set, the percent of data with value for **Target = 1** is therefore **30.92%**.   \n"},{"metadata":{"_uuid":"052730fa7f52f5a578c798385d5943f085bb7c4f"},"cell_type":"markdown","source":"<a href=\"#0\"><font size=\"1\">Go to top</font></a>\n\n## <a id=\"32\">Merge train and class detail info data</a>   \n\nLet's merge now the two datasets, using Patient ID as the merge criteria."},{"metadata":{"trusted":true,"_uuid":"dd96bd47c390fae9dd39f057558d5ed2cf0c633e"},"cell_type":"code","source":"train_class_df = train_labels_df.merge(class_info_df, left_on='patientId', right_on='patientId', how='inner')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2d892d10b9864feb14798fdb535748b7319b58a6"},"cell_type":"code","source":"train_class_df.sample(5)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"04bc5318c1b0ea3d51e1874c4bfdb3dfe11f1acb"},"cell_type":"markdown","source":"### Target and class  \n\nLet's plot the number of examinations for each class detected, grouped by Target value."},{"metadata":{"trusted":true,"_uuid":"6c3434da48be54ef8a3efad278ad60d74785d41f"},"cell_type":"code","source":"fig, ax = plt.subplots(nrows=1,figsize=(12,6))\ntmp = train_class_df.groupby('Target')['class'].value_counts()\ndf = pd.DataFrame(data={'Exams': tmp.values}, index=tmp.index).reset_index()\nsns.barplot(ax=ax,x = 'Target', y='Exams',hue='class',data=df, palette='Set3')\nplt.title(\"Chest exams class and Target\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f7d2efc1d98d4cb9879be1a8273553c913867b72"},"cell_type":"markdown","source":"All chest examinations with`Target` = **1** (pathology detected) associated with `class`:  **Lung Opacity**.    \n\nThe chest examinations with `Target` = **0** (no pathology detected) are either of `class`: **Normal** or `class`: **No Lung Opacity / Not Normal**."},{"metadata":{"_uuid":"3cafc590c43b46d6e74c53cb2447b4c056395ddd"},"cell_type":"markdown","source":"### Detected Lung Opacity window   \n\nFor the class **Lung Opacity**, corresponding to values of **Target = 1**, we plot the density of **x**, **y**, **width** and **height**.\n\n"},{"metadata":{"trusted":true,"_uuid":"5cb4e12983877f248c62131a8d0496a27fce2260"},"cell_type":"code","source":"target1 = train_class_df[train_class_df['Target']==1]\nsns.set_style('whitegrid')\nplt.figure()\nfig, ax = plt.subplots(2,2,figsize=(12,12))\nsns.distplot(target1['x'],kde=True,bins=50, color=\"red\", ax=ax[0,0])\nsns.distplot(target1['y'],kde=True,bins=50, color=\"blue\", ax=ax[0,1])\nsns.distplot(target1['width'],kde=True,bins=50, color=\"green\", ax=ax[1,0])\nsns.distplot(target1['height'],kde=True,bins=50, color=\"magenta\", ax=ax[1,1])\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ca3f5587a1bd07ba64a967ddde9028fdfc12954a"},"cell_type":"markdown","source":"We can plot also the center of the rectangles points in the plane x0y.   The centers of the rectangles are the points $$x_c = x + \\frac{width}{2}$$ and $$y_c = y + \\frac{height}{2}$$.\n\nWe will show a sample of center points superposed with the corresponding sample of the rectangles.\nThe rectangles are created using the method described in Kevin's Kernel <a href=\"#4\">[1]</a>."},{"metadata":{"trusted":true,"_uuid":"4ce30c9295dfcb20e33d13ba634953a1ec4b20ae"},"cell_type":"code","source":"fig, ax = plt.subplots(1,1,figsize=(7,7))\ntarget_sample = target1.sample(2000)\ntarget_sample['xc'] = target_sample['x'] + target_sample['width'] / 2\ntarget_sample['yc'] = target_sample['y'] + target_sample['height'] / 2\nplt.title(\"Centers of Lung Opacity rectangles (brown) over rectangles (yellow)\\nSample size: 2000\")\ntarget_sample.plot.scatter(x='xc', y='yc', xlim=(0,1024), ylim=(0,1024), ax=ax, alpha=0.8, marker=\".\", color=\"brown\")\nfor i, crt_sample in target_sample.iterrows():\n    ax.add_patch(Rectangle(xy=(crt_sample['x'], crt_sample['y']),\n                width=crt_sample['width'],height=crt_sample['height'],alpha=3.5e-3, color=\"yellow\"))\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"aaf16ced175bb48c8de944a1bbd28974ac6b250f"},"cell_type":"markdown","source":"We follow with the exploration of the DICOM data."},{"metadata":{"_uuid":"644e5457ebf328d902f6998d50ef93e6fcca6ce2"},"cell_type":"markdown","source":"<a href=\"#0\"><font size=\"1\">Go to top</font></a>\n\n## <a id=\"33\">Explore DICOM data</a>  \n\nLet's read now the DICOM data in the train set. The image path is as following:"},{"metadata":{"trusted":true,"_uuid":"8d73bdf1b233cd88a14616db27f1daac3eda12db"},"cell_type":"code","source":"image_sample_path = os.listdir(PATH+'/stage_2_train_images')[:5]\nprint(image_sample_path)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"93a67a19c088b808b7f96cde64facb9f1c901673"},"cell_type":"markdown","source":"The files names are the patients IDs.    \nLet's check how many images are in the train and test folders."},{"metadata":{"trusted":true,"_uuid":"abd3bfa1be2b6cd64aba895de970e6a7d3ae8785"},"cell_type":"code","source":"image_train_path = os.listdir(PATH+'/stage_2_train_images')\nimage_test_path = os.listdir(PATH+'/stage_2_test_images')\nprint(\"Number of images in train set:\", len(image_train_path),\"\\nNumber of images in test set:\", len(image_test_path))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4b0c9749410611a6c752a192a0bb42bebb83d03a"},"cell_type":"markdown","source":"\n\nOnly a reduced number of images are present in the training set (**26684**), compared with the number of  images in the train_df data (**30227**).  \n\nIt might be that we do have duplicated entries in the train and class datasets. Let's check this.\n\n### Check duplicates in train dataset\n"},{"metadata":{"trusted":true,"_uuid":"50116316cedd9ff369fda744fdd921343af403e5"},"cell_type":"code","source":"print(\"Unique patientId in  train_class_df: \", train_class_df['patientId'].nunique())      ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0f26c6193b856ffd91c220bf538e382e86fd66c8"},"cell_type":"markdown","source":"We confirmed that the number of *unique* **patientsId** are equal with the number of DICOM images in the train set.  \n\nLet's see what entries are duplicated. We want to check how are these distributed accross classes and Target value."},{"metadata":{"trusted":true,"_uuid":"e48bcc88bc3c6e9eedd24a832e48812b9a8c44d6"},"cell_type":"code","source":"tmp = train_class_df.groupby(['patientId','Target', 'class'])['patientId'].count()\ndf = pd.DataFrame(data={'Exams': tmp.values}, index=tmp.index).reset_index()\ntmp = df.groupby(['Exams','Target','class']).count()\ndf2 = pd.DataFrame(data=tmp.values, index=tmp.index).reset_index()\ndf2.columns = ['Exams', 'Target','Class', 'Entries']\ndf2","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"31f140995ddade3f195c5028fbf700d4419cda28"},"cell_type":"code","source":"fig, ax = plt.subplots(nrows=1,figsize=(12,6))\nsns.barplot(ax=ax,x = 'Target', y='Entries', hue='Exams',data=df2, palette='Set2')\nplt.title(\"Chest exams class and Target\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3b2272d96831da2a5989f97740ebb1dc0e571656"},"cell_type":"markdown","source":"\nLet's now extract one image and process the DICOM information. "},{"metadata":{"_uuid":"a830372471a5e330b11cfe240b48d300b2785997"},"cell_type":"markdown","source":"### DICOM meta data"},{"metadata":{"trusted":true,"_uuid":"094370d9a2b5eea948fd56f64063f724cd01a3c2"},"cell_type":"code","source":"samplePatientID = list(train_class_df[:3].T.to_dict().values())[0]['patientId']\nsamplePatientID = samplePatientID+'.dcm'\ndicom_file_path = os.path.join(PATH,\"stage_2_train_images/\",samplePatientID)\ndicom_file_dataset = dcm.read_file(dicom_file_path)\ndicom_file_dataset","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c0d0ceb5696feface3e3810cc9374269331d041a"},"cell_type":"markdown","source":"We can observe that we do have available some useful information in the DICOM metadata with predictive value, for example:   \n* Patient sex;   \n* Patient age;  \n* Modality;  \n* Body part examined;  \n* View position;  \n* Rows & Columns;  \n* Pixel Spacing.  \n"},{"metadata":{"_uuid":"a3fb73323279669a1df5fca8ad204ddfdbb5678c"},"cell_type":"markdown","source":"Let's sample few images having the **Target = 1**.\n\n### Plot DICOM images with Target = 1"},{"metadata":{"trusted":true,"_uuid":"c27f207fd6d286704435dff873103c7cc5d14046"},"cell_type":"code","source":"def show_dicom_images(data):\n    img_data = list(data.T.to_dict().values())\n    f, ax = plt.subplots(3,3, figsize=(16,18))\n    for i,data_row in enumerate(img_data):\n        patientImage = data_row['patientId']+'.dcm'\n        imagePath = os.path.join(PATH,\"stage_2_train_images/\",patientImage)\n        data_row_img_data = dcm.read_file(imagePath)\n        modality = data_row_img_data.Modality\n        age = data_row_img_data.PatientAge\n        sex = data_row_img_data.PatientSex\n        data_row_img = dcm.dcmread(imagePath)\n        ax[i//3, i%3].imshow(data_row_img.pixel_array, cmap=plt.cm.bone) \n        ax[i//3, i%3].axis('off')\n        ax[i//3, i%3].set_title('ID: {}\\nModality: {} Age: {} Sex: {} Target: {}\\nClass: {}\\nWindow: {}:{}:{}:{}'.format(\n                data_row['patientId'],\n                modality, age, sex, data_row['Target'], data_row['class'], \n                data_row['x'],data_row['y'],data_row['width'],data_row['height']))\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fef222f43b41a6d4833435d8b214ce5fea1a0327"},"cell_type":"code","source":"show_dicom_images(train_class_df[train_class_df['Target']==1].sample(9))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"776b024b034b1d48fe147ec60205ad0346a05e0b"},"cell_type":"markdown","source":"We would like to represent the images with the overlay boxes superposed. For this, we will need first to parse the whole dataset with **Target = 1** and gather all coordinates of the windows showing a **Lung Opacity** on the same image.  The simples method is show in <a href='#5'>[1]</a> and we will adapt our rendering from this method."},{"metadata":{"trusted":true,"_uuid":"472ffc50121e70c31d14e21513a75b92247a15db"},"cell_type":"code","source":"def show_dicom_images_with_boxes(data):\n    img_data = list(data.T.to_dict().values())\n    f, ax = plt.subplots(3,3, figsize=(16,18))\n    for i,data_row in enumerate(img_data):\n        patientImage = data_row['patientId']+'.dcm'\n        imagePath = os.path.join(PATH,\"stage_2_train_images/\",patientImage)\n        data_row_img_data = dcm.read_file(imagePath)\n        modality = data_row_img_data.Modality\n        age = data_row_img_data.PatientAge\n        sex = data_row_img_data.PatientSex\n        data_row_img = dcm.dcmread(imagePath)\n        ax[i//3, i%3].imshow(data_row_img.pixel_array, cmap=plt.cm.bone) \n        ax[i//3, i%3].axis('off')\n        ax[i//3, i%3].set_title('ID: {}\\nModality: {} Age: {} Sex: {} Target: {}\\nClass: {}'.format(\n                data_row['patientId'],modality, age, sex, data_row['Target'], data_row['class']))\n        rows = train_class_df[train_class_df['patientId']==data_row['patientId']]\n        box_data = list(rows.T.to_dict().values())\n        for j, row in enumerate(box_data):\n            ax[i//3, i%3].add_patch(Rectangle(xy=(row['x'], row['y']),\n                        width=row['width'],height=row['height'], \n                        color=\"yellow\",alpha = 0.1))   \n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"40b5add52aa95cc20a6c3112dfb9dce581c69e4b"},"cell_type":"code","source":"show_dicom_images_with_boxes(train_class_df[train_class_df['Target']==1].sample(9))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fb28bff001aa9a2f4f7f6f1551181eb7ddfa79dd"},"cell_type":"markdown","source":"For some of the images with **Target=1**, we might see multiple areas (boxes/rectangles) with **Lung Opacity**.\n\nLet's sample few images having the **Target = 0**.   \n\n### Plot DICOM images with Target = 0\n"},{"metadata":{"trusted":true,"_uuid":"3d42e01bc5a9a8fc75be3930c155262bd4023081"},"cell_type":"code","source":"show_dicom_images(train_class_df[train_class_df['Target']==0].sample(9))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f4eca5a5ffee6e1a4980048e95f18841f60b9db5"},"cell_type":"markdown","source":"<a href=\"#0\"><font size=\"1\">Go to top</font></a>   \n\n\n## <a id=\"34\">Add meta information from DICOM data</a>\n\n\n### Train data\n\nWe will parse the DICOM meta information and add it to the train dataset. We will do the same with the test data."},{"metadata":{"trusted":true,"_uuid":"106f514e8e19a72969e9de0e18a7eedd0483537d"},"cell_type":"code","source":"vars = ['Modality', 'PatientAge', 'PatientSex', 'BodyPartExamined', 'ViewPosition', 'ConversionType', 'Rows', 'Columns', 'PixelSpacing']\n\ndef process_dicom_data(data_df, data_path):\n    for var in vars:\n        data_df[var] = None\n    image_names = os.listdir(PATH+data_path)\n    for i, img_name in tqdm_notebook(enumerate(image_names)):\n        imagePath = os.path.join(PATH,data_path,img_name)\n        data_row_img_data = dcm.read_file(imagePath)\n        idx = (data_df['patientId']==data_row_img_data.PatientID)\n        data_df.loc[idx,'Modality'] = data_row_img_data.Modality\n        data_df.loc[idx,'PatientAge'] = pd.to_numeric(data_row_img_data.PatientAge)\n        data_df.loc[idx,'PatientSex'] = data_row_img_data.PatientSex\n        data_df.loc[idx,'BodyPartExamined'] = data_row_img_data.BodyPartExamined\n        data_df.loc[idx,'ViewPosition'] = data_row_img_data.ViewPosition\n        data_df.loc[idx,'ConversionType'] = data_row_img_data.ConversionType\n        data_df.loc[idx,'Rows'] = data_row_img_data.Rows\n        data_df.loc[idx,'Columns'] = data_row_img_data.Columns  \n        data_df.loc[idx,'PixelSpacing'] = str.format(\"{:4.3f}\",data_row_img_data.PixelSpacing[0]) ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2d773e5079f9ec5099e616229846de20f36e45f9"},"cell_type":"code","source":"process_dicom_data(train_class_df,'stage_2_train_images/')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ff7946e29da111447e20f448dfda05dc0e50b8b2"},"cell_type":"markdown","source":"### Test data\n\nWe will create as well a test dataset with similar information."},{"metadata":{"trusted":true,"_uuid":"ca9b3af29b8fa022f6b2e8e258c1b530d3ccf19f"},"cell_type":"code","source":"test_class_df = pd.read_csv(PATH+'/stage_2_sample_submission.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"51bfd1b13d694c7ee68e06249f1f03f18ad6033f"},"cell_type":"code","source":"test_class_df = test_class_df.drop('PredictionString',1)\nprocess_dicom_data(test_class_df,'stage_2_test_images/')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8c491c2ef586b101ad66094b3747d68602923df9"},"cell_type":"markdown","source":"<a href=\"#0\"><font size=\"1\">Go to top</font></a>\n\n\n## <a id=\"35\">Modality</a>\n\nLet's check how many modalities are used. Both train and test set are checked."},{"metadata":{"trusted":true,"_uuid":"2ce3a4a7f1863bb803eb177b41f149f6c4c6001b"},"cell_type":"code","source":"print(\"Modalities: train:\",train_class_df['Modality'].unique(), \"test:\", test_class_df['Modality'].unique())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"064353b381b82804f0c38fa720c2e60243cca9f3"},"cell_type":"markdown","source":"The meaning of this modality is **CR** - **Computer Radiography**  <a href='#4'>[2]</a> <a href='#4'>[3]</a>.\n\n\n## <a id=\"36\">Body Part Examined</a>\n\nLet's check if other body parts than 'CHEST' appears in the data."},{"metadata":{"trusted":true,"_uuid":"d7a0a37a81190789b1588767f3efbca01e84c979"},"cell_type":"code","source":"print(\"Body Part Examined: train:\",train_class_df['BodyPartExamined'].unique(), \"test:\", test_class_df['BodyPartExamined'].unique())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0996d3e4ea751f94ffa43e019cd8d2165ee912ef"},"cell_type":"markdown","source":"## <a id=\"37\">View Position</a>\n\nView Position is a radiographic view associated with the Patient Position. Let's check the View Positions distribution for the both datasets.\n\n\n"},{"metadata":{"trusted":true,"_uuid":"21230651bb6a5f50918ddb4edcc6a36aeab4a497"},"cell_type":"code","source":"print(\"View Position: train:\",train_class_df['ViewPosition'].unique(), \"test:\", test_class_df['ViewPosition'].unique())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cad8148a0f731a17b690ef01a6be64b22ee17afd"},"cell_type":"markdown","source":"### Train dataset  \n\nLet's get into more details for the train dataset. First, let's check the distribution of PA and AP."},{"metadata":{"trusted":true,"_uuid":"2cb7386167ecef3384c80310b0745ce608b5b4e9"},"cell_type":"code","source":"get_feature_distribution(train_class_df,'ViewPosition')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"41541a5c84dc5327b8a3462a736a7696787293c0"},"cell_type":"markdown","source":"Both **AP** and **PA** body positions are present in the data.  The meaning of these view positions are <a href='#4'>[2]</a> <a href='#4'>[3]</a>:\n* **AP** - Anterior/Posterior;    \n* **PA** - Posterior/Anterior.    \n\n\nLet's check, for the training data presenting **Lung Opacity**, the distribution of the window for both View Positions. We create a function to represent the distribution of the window centers and windows."},{"metadata":{"trusted":true,"_uuid":"c23d252244fe9e2f85284cfea640cb6edda63fcb"},"cell_type":"code","source":"def plot_window(data,color_point, color_window,text):\n    fig, ax = plt.subplots(1,1,figsize=(7,7))\n    plt.title(\"Centers of Lung Opacity rectangles over rectangles\\n{}\".format(text))\n    data.plot.scatter(x='xc', y='yc', xlim=(0,1024), ylim=(0,1024), ax=ax, alpha=0.8, marker=\".\", color=color_point)\n    for i, crt_sample in data.iterrows():\n        ax.add_patch(Rectangle(xy=(crt_sample['x'], crt_sample['y']),\n            width=crt_sample['width'],height=crt_sample['height'],alpha=3.5e-3, color=color_window))\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"140dbe24e5050fcbd16c7aae2e9f99fed4c37211"},"cell_type":"markdown","source":"We sample a subset of the train data with **Target = 1**. We calculate as well the center of the windows with **Lung Opacity**.   We then select from this sample the data with the two view position, to plot the window distribution separatelly."},{"metadata":{"trusted":true,"_uuid":"8aa19a176ddf1c3da8de0119aa77a63bb050762a"},"cell_type":"code","source":"target1 = train_class_df[train_class_df['Target']==1]\n\ntarget_sample = target1.sample(2000)\ntarget_sample['xc'] = target_sample['x'] + target_sample['width'] / 2\ntarget_sample['yc'] = target_sample['y'] + target_sample['height'] / 2\n\ntarget_ap = target_sample[target_sample['ViewPosition']=='AP']\ntarget_pa = target_sample[target_sample['ViewPosition']=='PA']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6af9f4b8861a99ab7eb4e4160bba2eb00b92a216"},"cell_type":"code","source":"plot_window(target_ap,'green', 'yellow', 'Patient View Position: AP')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8ae43185dbec3b113607c982669f2342402c20cd"},"cell_type":"code","source":"plot_window(target_pa,'blue', 'red', 'Patient View Position: PA')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3f3e9dd690c2d304df9f99c90823caf56dc3a6b4"},"cell_type":"markdown","source":"### Test dataset  \n\nLet's check the distribution of AP and PA positions for the test set."},{"metadata":{"trusted":true,"_uuid":"68b3a1a9bfa778a1e9734265af1835f560edbba4"},"cell_type":"code","source":"get_feature_distribution(test_class_df,'ViewPosition')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8565e89770e5aa7fd6f078e640953d5960829cd3"},"cell_type":"markdown","source":"## <a id=\"38\">Conversion Type</a>\n\nLet's check the Conversion Type data."},{"metadata":{"trusted":true,"_uuid":"e6aefea2217de5101d746dcc0e90ed96088f7863"},"cell_type":"code","source":"print(\"Conversion Type: train:\",train_class_df['ConversionType'].unique(), \"test:\", test_class_df['ConversionType'].unique())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2c842986fffa3d113f50e77b9da0331425af9d2d"},"cell_type":"markdown","source":"Both train and test have only **WSD** Conversion Type Data. The meaning of this Conversion Type is **WSD**: **Workstation**.\n\n## <a id=\"39\">Rows and Columns</a>"},{"metadata":{"trusted":true,"_uuid":"805739ffca6581ec0abd6db5dfcda3702446c51c"},"cell_type":"code","source":"print(\"Rows: train:\",train_class_df['Rows'].unique(), \"test:\", test_class_df['Rows'].unique())\nprint(\"Columns: train:\",train_class_df['Columns'].unique(), \"test:\", test_class_df['Columns'].unique())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d46c7d7b265af0159fba9155a66a1207549724ee"},"cell_type":"markdown","source":"Only {Rows:Columns} {1024:1024} are present in both train and test.  \n\n\n<a href=\"#0\"><font size=\"1\">Go to top</font></a>"},{"metadata":{"_uuid":"8f3fd85400247f0e85248f7613566613e60bfab6"},"cell_type":"markdown","source":"## <a id=\"310\">Patient Age</a>\n\nLet's examine now the data for the Patient Age for the train set.\n\n### Train dataset"},{"metadata":{"trusted":true,"_uuid":"fdd0dbc01ecaed750f748e17d5fe10336a7f699e"},"cell_type":"code","source":"tmp = train_class_df.groupby(['Target', 'PatientAge'])['patientId'].count()\ndf = pd.DataFrame(data={'Exams': tmp.values}, index=tmp.index).reset_index()\ntmp = df.groupby(['Exams','Target', 'PatientAge']).count()\ndf2 = pd.DataFrame(data=tmp.values, index=tmp.index).reset_index()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1a6941cf162c8726eb8fcc7ae795b341a00e0465"},"cell_type":"code","source":"tmp = train_class_df.groupby(['class', 'PatientAge'])['patientId'].count()\ndf1 = pd.DataFrame(data={'Exams': tmp.values}, index=tmp.index).reset_index()\ntmp = df1.groupby(['Exams','class', 'PatientAge']).count()\ndf3 = pd.DataFrame(data=tmp.values, index=tmp.index).reset_index()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9f7d38a118f74d6017c355c24f5f45fd0530278b"},"cell_type":"code","source":"fig, (ax) = plt.subplots(nrows=1,figsize=(16,6))\nsns.barplot(ax=ax, x = 'PatientAge', y='Exams', hue='Target',data=df2)\nplt.title(\"Train set: Chest exams Age and Target\")\nplt.xticks(rotation=90)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ba2fa0a92aecd55a8054e0b3b5cfe191059afba5"},"cell_type":"code","source":"fig, (ax) = plt.subplots(nrows=1,figsize=(16,6))\nsns.barplot(ax=ax, x = 'PatientAge', y='Exams', hue='class',data=df3)\nplt.title(\"Train set: Chest exams Age and class\")\nplt.xticks(rotation=90)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ceff498702b9fb6f22a8e8375ff77bbf35aa0e5b"},"cell_type":"markdown","source":"\n**Note**: most probably, the values of age 148 to 155 are mistakes.   \n\nLet's group the ages in 5 groups (0-19, 20-34, 35-49, 50-64 and 65+). "},{"metadata":{"trusted":true,"_uuid":"626c21ac5233f6df32c3b957b33e44deaa5514c2"},"cell_type":"code","source":"target_age1 = target_sample[target_sample['PatientAge'] < 20]\ntarget_age2 = target_sample[(target_sample['PatientAge'] >=20) & (target_sample['PatientAge'] < 35)]\ntarget_age3 = target_sample[(target_sample['PatientAge'] >=35) & (target_sample['PatientAge'] < 50)]\ntarget_age4 = target_sample[(target_sample['PatientAge'] >=50) & (target_sample['PatientAge'] < 65)]\ntarget_age5 = target_sample[target_sample['PatientAge'] >= 65]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"eb3c9b8d8db4623a2a77ab215a3b9add9dc10edd"},"cell_type":"markdown","source":"Let's show the distribution of windows for the 5 age groups."},{"metadata":{"trusted":true,"_uuid":"b4576a36d04585475747d165889f9ad432582c69"},"cell_type":"code","source":"plot_window(target_age1,'blue', 'red', 'Patient Age: 1-19 years')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9fcd06249e1ecc6ba8b69197f4ec41452a82ecf2"},"cell_type":"code","source":"plot_window(target_age2,'blue', 'red', 'Patient Age: 20-34 years')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b4e75bbc844b3e5d1891c7cab054145bf3e9b4ec"},"cell_type":"code","source":"plot_window(target_age3,'blue', 'red', 'Patient Age: 35-49 years')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c1c26e31009356230e4784fbbc9fb9bbb16b0755"},"cell_type":"code","source":"plot_window(target_age4,'blue', 'red', 'Patient Age: 50-65 years')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f1b6afbade9f5ed7376f2b0b3c77ff273aab69f9"},"cell_type":"code","source":"plot_window(target_age5,'blue', 'red', 'Patient Age: 65+ years')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4b014b3fa6637d132ece71d70fd8e6acb7491aff"},"cell_type":"markdown","source":"Let's check also the distribution of patient age for the test data set.\n\n### Test dataset"},{"metadata":{"trusted":true,"_uuid":"788fe9f3f2f47284c01a14222c4b573905a7d839"},"cell_type":"code","source":"fig, (ax) = plt.subplots(nrows=1,figsize=(16,6))\nsns.countplot(test_class_df['PatientAge'], ax=ax)\nplt.title(\"Test set: Patient Age\")\nplt.xticks(rotation=90)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3a3b257544ffa125445c9ff9be7ebcc49144c700"},"cell_type":"markdown","source":"<a href=\"#0\"><font size=\"1\">Go to top</font></a>\n\n\n## <a id=\"311\">Patient Sex</a>\n\nLet's examine now the data for the Patient Sex.   \n\n### Train dataset\n\nWe represent the number of Exams for each Patient Sex, grouped by value of Target."},{"metadata":{"trusted":true,"_uuid":"faa8d113dee81cfc3f458c619a86deffeaa934ee"},"cell_type":"code","source":"tmp = train_class_df.groupby(['Target', 'PatientSex'])['patientId'].count()\ndf = pd.DataFrame(data={'Exams': tmp.values}, index=tmp.index).reset_index()\ntmp = df.groupby(['Exams','Target', 'PatientSex']).count()\ndf2 = pd.DataFrame(data=tmp.values, index=tmp.index).reset_index()\nfig, ax = plt.subplots(nrows=1,figsize=(6,6))\nsns.barplot(ax=ax, x = 'PatientSex', y='Exams', hue='Target',data=df2)\nplt.title(\"Train set: Patient Sex and Target\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7571981004180feaa2fba353b1f829e48c8d6df0"},"cell_type":"markdown","source":"We represent the number of Exams for each Patient Sex, grouped by value of  class."},{"metadata":{"trusted":true,"_uuid":"fff96d98a9b65431e408583fa6fe5f977b522365"},"cell_type":"code","source":"tmp = train_class_df.groupby(['class', 'PatientSex'])['patientId'].count()\ndf1 = pd.DataFrame(data={'Exams': tmp.values}, index=tmp.index).reset_index()\ntmp = df1.groupby(['Exams','class', 'PatientSex']).count()\ndf3 = pd.DataFrame(data=tmp.values, index=tmp.index).reset_index()\nfig, (ax) = plt.subplots(nrows=1,figsize=(6,6))\nsns.barplot(ax=ax, x = 'PatientSex', y='Exams', hue='class',data=df3)\nplt.title(\"Train set: Patient Sex and class\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1a195adec52b30461f2ae9a8d08302b3ec3c2e8e"},"cell_type":"markdown","source":"Let's plot as well the distribution of  window with Lung Opacity, separatelly for the female and male patients. We will reuse the sample with **Target = 1** for which we calculated also the center of the window."},{"metadata":{"trusted":true,"_uuid":"446bab321673830094bb2d51cbcfca352601ce7c"},"cell_type":"code","source":"target_female = target_sample[target_sample['PatientSex']=='F']\ntarget_male = target_sample[target_sample['PatientSex']=='M']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b4b6cbc70be1d00b587f4756a571feaaafa960bf"},"cell_type":"code","source":"plot_window(target_female,\"red\", \"magenta\",\"Patients Sex: Female\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1ee23aee589c3afb566e612c2420ac3f20d6b5f6"},"cell_type":"code","source":"plot_window(target_male,\"darkblue\", \"blue\", \"Patients Sex: Male\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6af774c2021f12736b2c274c1bf712d8e38c43f7"},"cell_type":"markdown","source":"Let's check as well the distribution of Patient Sex for the test data.   \n\n### Test dataset"},{"metadata":{"trusted":true,"_uuid":"b1804efc476ffebd618df4ce3f4edf35ca7221a9"},"cell_type":"code","source":"sns.countplot(test_class_df['PatientSex'])\nplt.title(\"Test set: Patient Sex\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f1e617210b9cbe91f0279a569cc727c16114d082"},"cell_type":"markdown","source":"<a href=\"#0\"><font size=\"1\">Go to top</font></a>\n\n\n# <a id='4'>Conclusions</a>   \n\nAfter exploring the data, both the tabular and DICOM data, we were able to:  \n- discover duplications in the tabular data;  \n- explore the DICOM images;  \n- extract meta information from the DICOM data;  \n- add features to the tabular data from the meta information in DICOM data;  \n- further analyze the distribution of the data with the newly added features from DICOM metadata;  \n\nAll these findings are useful as preliminary work for building a model.\n\n<a href=\"#0\"><font size=\"1\">Go to top</font></a>"},{"metadata":{"_uuid":"ceb579201491448bb489133975bb31d7e83d4707"},"cell_type":"markdown","source":"# <a id='5'>References</a>  \n\n\n[1] Kevin Mader, Lung Opacity Overview, https://www.kaggle.com/kmader/lung-opacity-overview  \n[2] Modality Specific Modules, DICOM Standard,  http://dicom.nema.org/medical/dicom/2014c/output/chtml/part03/sect_C.8.html  \n[3] DICOM Standard, https://www.dicomstandard.org/     \n[4] Getting Started with Pydicom, https://pydicom.github.io/pydicom/stable/getting_started.html   \n[5] ITKPYthon package, https://itkpythonpackage.readthedocs.io/en/latest/   \n[6] DICOM in Python: Importing medical image data into NumPy with PyDICOM and VTK, https://pyscience.wordpress.com/2014/09/08/dicom-in-python-importing-medical-image-data-into-numpy-with-pydicom-and-vtk/  \n[7] DICOM Processing and Segmentation in Python, https://www.raddq.com/dicom-processing-segmentation-visualization-in-python/  \n[8] DICOM Standard Browser, https://dicom.innolitics.com/ciods  \n[9] How can I read a DICOM image in Python, https://www.quora.com/How-can-I-read-a-DICOM-image-in-Python  \n[10] DICOM read example in Python, https://www.programcreek.com/python/example/97517/dicom.read_file    \n[11] DICOM in Python, https://github.com/pydicom   \n\n\n<a href=\"#0\"><font size=\"1\">Go to top</font></a>\n"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}