{"cells":[{"metadata":{},"cell_type":"markdown","source":"This is my first notebook for OSIC Fibrosis Progression challenge. Here I have done the EDA for both the tabular and image data.\n\nReferences : https://www.kaggle.com/jagdmir/insightful-eda-on-meta-data-dicom-files","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Import Libraries","execution_count":null},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nplt.rcParams['figure.figsize'] = (10,8)\nimport os\nimport glob\nimport pydicom","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Load Datasets","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/train.csv')\ntest = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/test.csv')\nsub = pd.read_csv('../input/osic-pulmonary-fibrosis-progression/sample_submission.csv')\n\ntrain_dir = '../input/osic-pulmonary-fibrosis-progression/train'\ntest_dir = '../input/osic-pulmonary-fibrosis-progression/test'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Checking dimensions of all datasets\n\ntrain.shape, test.shape, sub.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Exploring train data\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Exploring test data\ntest","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Explore sample submission\nsub.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As we can see from the format of submission file, we need to predict the FVC mesasurement for each patient for each week, along with the confidence value.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Checking the number of records in sample submission for a patient\nsub_patient1 = sub[sub['Patient_Week'].str.contains('ID00419637202311204720264')]\nsub_patient1.info()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Each patient has 146 entries across weeks for which we need to predict the FVC. Total number of patients in 5 and the total number of submissions will be 146*5 = 730.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Checking the range of values for week \nsub_patient1.head(1), sub_patient1.tail(1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It is clear that submission file consists of records for each patient from week -12 to 133.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Checking for null values\ntrain.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Check number of unique patients in train data\ntrain['Patient'].nunique()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Number of unqiue patients in train dataset is 176.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#Check types of features available in the dataset\ntrain.dtypes","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Excluding Patient column, we have total 5 independent features and 1 target feature, that is FVC.\nOut of 5 independent features, two are categorical (Sex, SmokingStatus) and rest are numeric features (Weeks, Percent,Age).\n\nLets explore all these features in the next section.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Data Exploration","execution_count":null},{"metadata":{},"cell_type":"markdown","source":" **Distribution of FVC**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#plt.hist(train['FVC'])\nsns.distplot(train['FVC'],hist=False,color='darkred')\nplt.title('FVC Distribution')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"FVC ranges from 1000 to 5000 mostly. Very few cases go beyond 5000. The peak values occur at around 2500.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Frequency of Patient Visits**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df = train.groupby('Patient').count()['Weeks'].value_counts()\ndf","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Most of the patient have got their FVC checked 9 times.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Categorical Variables : Sex, SmokingStatus**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Lets check number of males/females in the dataset\nsizes = [len(train[train['Sex']=='Male']), len(train[train['Sex']=='Female'])]\nexplode = (0.1,0) #explde first slice\ncolors = ['blue','pink']\n\nplt.pie(sizes, explode=explode, labels= train['Sex'].unique(), colors= colors, autopct= '%1.1f%%', \n        shadow=True, startangle=140)\nplt.axis('equal')\nplt.title('Pie Chart for Gender Distribution')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"79% of the patients are male.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Checking Smoking Status**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train['SmokingStatus'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"More than 70% of the records are of people who are either ex-smokers or currently smoke.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Distribution of Smoking Status**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"sizes = [len(train[train['SmokingStatus']=='Ex-smoker']), len(train[train['SmokingStatus']=='Never smoked']), \n        len(train[train['SmokingStatus']=='Currently smokes'])]\nexplode = (0,0,0.1)\ncolors = ['blue','red','green']\nplt.pie(sizes, explode=explode, colors = colors, labels= train['SmokingStatus'].unique(), autopct='%1.1f%%', \n        shadow=True, startangle=140)\n\nplt.axis('equal')\nplt.title('Pie Chart for Smoking Status distribution')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Age-Wise Distribution**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train['Age'].min(), train['Age'].max()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"f, (ax1, ax2) = plt.subplots(1, 2, figsize = (16,6))\n\n#Patient age group\nageGroupLabel = 'Below 60', '60-70', '70-80', 'Above 80'\n\nbelow60 = len(train[train.Age<60])\nsixty_to_seventy = len(train[(train['Age']>=60) & (train['Age']<= 70)])\nseventy_to_eighty = len(train[(train['Age']>70) & (train['Age']<= 80)])\nabove80 = len(train[train.Age>80])\n\npatientNumbers = [below60, sixty_to_seventy, seventy_to_eighty,above80]\nexplode = (0,0,0,0.1)\ncolors = ['green','indigo','blue','red']\n\n#Draw the pie chart\nax1.pie(patientNumbers, explode=explode, colors= colors, labels= ageGroupLabel, autopct = '%1.2f',startangle = 90)\n\n#Aspect ratio\nax1.axis('equal')\n\n# Distribution plot for age\nsns.distplot(train['Age'],hist= False, color='red')\nplt.suptitle('Age Distribution')\nplt.show()\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Majority of the patients belong to 60 to 80 age bracket. Very few patients are above 80.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Week-wise Distribution**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train['Weeks'].min(), train['Weeks'].max()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Dividing weeks into bins of 10 to visualize\n\nbelow10 = len(train[train['Weeks']<=10])\neleven_20 = len(train[(train['Weeks']>=11) & (train['Weeks']<= 20)])\ntwentyone_30 = len(train[(train['Weeks']>20) & (train['Weeks']<= 30)])\nthirtyone_40 = len(train[(train['Weeks']>30) & (train['Weeks']<= 40)])\nfortyone_50 = len(train[(train['Weeks']>40) & (train['Weeks']<= 50)])\nfiftyone_60 = len(train[(train['Weeks']>50) & (train['Weeks']<= 60)])\nsixtyone_70 = len(train[(train['Weeks']>60) & (train['Weeks']<= 70)])\nseventyone_80 = len(train[(train['Weeks']>70) & (train['Weeks']<= 80)])\neightyone_90 = len(train[(train['Weeks']>80) & (train['Weeks']<= 90)])\nninetyone_100 = len(train[(train['Weeks']>90) & (train['Weeks']<= 100)])\nhundredone_110 = len(train[(train['Weeks']>100) & (train['Weeks']<= 110)])\nhundredten_120 = len(train[(train['Weeks']>110) & (train['Weeks']<= 120)])\nabove120 = len(train[train.Weeks>120])\n\nsizes = [below10, eleven_20, twentyone_30, thirtyone_40, fortyone_50, fiftyone_60,sixtyone_70,\n         seventyone_80,eightyone_90,ninetyone_100,hundredone_110,hundredten_120,above120]\nlabels =  ['below10','eleven_20', 'twentyone_30', 'thirtyone_40', 'fortyone_50', 'fiftyone_60','sixtyone_70',\n'seventyone_80','eightyone_90',\n'ninetyone_100','hundredone_110','hundredten_120','above120']\n\nfig1, (ax1, ax2)= plt.subplots(1,2,figsize=(15, 10))\ntheme = plt.get_cmap('prism')\nax1.set_prop_cycle(\"color\", [theme(1. * i / len(sizes)) for i in range(len(sizes))])\n_, _ = ax1.pie(sizes, startangle=90)\n\nax1.axis('equal')\ntotal = sum(sizes)\nax1.legend(\n    loc='upper left',\n    labels=['%s, %1.1f%%' % (\n        l, (float(s) / total) * 100) for l, s in zip(labels, sizes)],\n    prop={'size': 11},\n    bbox_to_anchor=(0.0, 1),\n    bbox_transform=fig1.transFigure\n)\n\n# distribution plot for Weeks\nsns.distplot(train['Weeks'], hist = False, color = \"indigo\")\nplt.suptitle(\"Weeks Distribution\")\n\nplt.show()\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It is clear that almost 50% of the patients got their FVC checked between week -5 to 25.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Pair Plot**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Heatmap\ncorrMatrix = train.corr()\nmask = np.triu(corrMatrix)\nsns.heatmap(corrMatrix,annot=True,cmap='coolwarm',\n           linewidths=1, cbar=False, fmt='.1f',\n           mask=mask)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can say that FVC and Percent have a decent positive correlation. It seems FVC is more when the percent value is higher.\nFVC is also correlated with Age inversly.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# pair plot\nsns.pairplot(train)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Above plots backs the observation that FVC has strong correlation with Percent. None of the other features show\nsignificant correlation with FVC.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Drawing pairplots with different colors for Sex and Smoking Status**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.pairplot(train,hue='SmokingStatus')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.pairplot(train,hue='Sex')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Checking pairplot for a random sample**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_100 = train[train['Patient']==train['Patient'][100]]\nsns.pairplot(df_100)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can say now that our understanding is correct that FVC and Percent have decent correlation.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Mean FVC for Males vs Females**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"grp_sex  = train.groupby('Sex')\n\n#draw a plot to display mean of FVC for males and females\nsplot = sns.barplot(x=train['Sex'].unique(),y= grp_sex['FVC'].mean())\n\nplt.xlabel(\"Sex\",fontsize=30)\nplt.ylabel(\"Mean FVC\",fontsize=30)\nplt.title(\"Mean FVC Males vs Mean FVC Femailes\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Mean FVC is lower for males, but that can be because we do not have enough data for females","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Mean FVC for Diffetent SmokingStatus**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"grp_smoking = train.groupby('SmokingStatus')\n\nsplot = sns.barplot(x= train['SmokingStatus'].unique(),y=grp_smoking['FVC'].mean())\n\nplt.xlabel(\"Smoking Status\",fontsize=30)\nplt.ylabel(\"Mean FVC\",fontsize=30)\nplt.title(\"Mean FVC for different Smoking categiries\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Mean FVC is least for patient who are smokers currently, which makes sense. But mean FVC for ex-smokers is higher than non smokers, which seems misleading.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Mean FVC - Week Wise**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(12,8))\n\n#Creating bins for weeks\ntrain['Weeks_bins'] = pd.cut(train['Weeks'],13, duplicates='drop')\n\n#Group the data by bins\ngrp_weeks = train.groupby('Weeks_bins')\n\n#Drawing barplots\nsplot = sns.barplot(x=train['Weeks_bins'].unique(),y=grp_weeks['FVC'].mean())\n\nplt.xlabel(\"Weeks Bins\",fontsize=30)\nplt.ylabel(\"Mean FVC\",fontsize=30)\nplt.title(\"Mean FVC for different Weeks bins\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There is no much variation in mean FVC for different weeks","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Mean FVC - Age Wise**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#Creating bins\ntrain['Age_bins'] = pd.cut(train['Age'],4,duplicates='drop')\n\n#Group by Age bins\ngrp_Age = train.groupby('Age_bins')\n\n#Drawing plot\nsplot = sns.barplot(x= train['Age_bins'].unique(),y=grp_Age['FVC'].mean())\n\nplt.xlabel(\"Age Bins\",fontsize=30)\nplt.ylabel(\"Mean FVC\",fontsize=30)\nplt.title(\"Mean FVC for different Age bins\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Again FVC is evenly distributed for all age groups","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**FVC vs Weeks Plot for Randomly Selected Patients**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"p_id = list(train['Patient'].sample(3))\np_id","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Drawing the distribution of FVC over weeks for a randomly selected patient 1\n\nplt.plot(train[train['Patient']==p_id[0]].Weeks, train[train['Patient']==p_id[0]].FVC,\n            color='darkblue')\n    \nplt.xlabel('Weeks',fontsize=30)\nplt.ylabel('FVC',fontsize=30)\nplt.title('FVC for Patient: '+p_id[0])\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Drawing the distribution of FVC over weeks for a randomly selected patient 2\n\nplt.plot(train[train['Patient']==p_id[1]].Weeks, train[train['Patient']==p_id[1]].FVC,\n            color='darkgreen')\n    \nplt.xlabel('Weeks',fontsize=30)\nplt.ylabel('FVC',fontsize=30)\nplt.title('FVC for Patient: '+p_id[1])\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Drawing the distribution of FVC over weeks for a randomly selected patient 1\n\nplt.plot(train[train['Patient']==p_id[2]].Weeks, train[train['Patient']==p_id[2]].FVC,\n            color='purple')\n    \nplt.xlabel('Weeks',fontsize=30)\nplt.ylabel('FVC',fontsize=30)\nplt.title('FVC for Patient: '+p_id[2])\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Above plots give us better understanding of the relationship between weeks and FVC. Randomly selected patients are showing overall downward trend in FVC as the weeks progress.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# DICOM Files","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"DICOM stands for **Digital Imaging and Communications**. It is the standard for management of medical images. These files consist of images along with patient information. \n\nDCM files are basically 2-D slices of a single CT scan. Both the train and test directories have sub-directories for patients. Each sub-directory has multiple dicom images for a single patient, which make the CT scan for that patient.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Checking the contents of train directory\np_sizes = [] #list of number of dcm files present for each patient\n\nfor dir in os.listdir(train_dir):\n    print('Patient {} has {} scans'.format(dir, len(os.listdir(train_dir+ \"/\"+ dir))))\n    p_sizes.append(len(os.listdir(train_dir+\"/\"+dir)))\n    \nprint('----------')\nprint('Total number of patients {}. Total DCM files {}'.format(len(os.listdir(train_dir)),sum(p_sizes)))\n\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We have total 176 directories, one directory per patient. Total number is dicom images is 33026","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#Visualizing DICOM count per patient\np = sns.color_palette()\nplt.hist(p_sizes,color=p[2])\nplt.xlabel('Count of DCM files')\nplt.ylabel('Number of patients')\nplt.title('Histogram of DICOM count per patient-Training Data')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Most of the patients have between 20 to 620 DICOM files under their sub-directories. Very few patients have more than 800 files.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Checking the contents of test directory\np_sizes = [] #list of number of dcm files present for each patient\n\nfor dir in os.listdir(test_dir):\n    print('Patient {} has {} scans'.format(dir, len(os.listdir(test_dir+ \"/\"+ dir))))\n    p_sizes.append(len(os.listdir(test_dir+\"/\"+dir)))\n    \nprint('----------')\nprint('Total number of patients {}. Total DCM files {}'.format(len(os.listdir(test_dir)),sum(p_sizes)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Test directory has 5 sub-folders, one for each patient. These folders have 1261 files in total.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#Visualizing DICOM count per patient - Test Data\np = sns.color_palette()\nplt.hist(p_sizes,color=p[4])\nplt.xlabel('Count of DCM files')\nplt.ylabel('Number of patients')\nplt.title('Histogram of DICOM count per patient-Test Data')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Test directory has around 25 to 400 images per patient.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Checking the Size of DICOM Images**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"sizes = [os.path.getsize(dcm)/1000000 for dcm in glob.glob(train_dir+\"/*/*.dcm\")]\n\nprint('DCM File sizes: min {:.3} MB, max {:.3} MB, avg {:.3} MB, std {:.3} MB'.format(\n        np.min(sizes),np.max(sizes),np.mean(sizes),np.std(sizes)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can see images are in range of 0.2 MB to 3.4 MB. Average image size if 0.7 MB.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Loading a Random DICOM Image**","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"DICOM files can be read and processed using pydicom package. DICOM files allow storing of metadata along with pixel values. ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#read a dcm file for patient ID00368637202296470751086\ndcm = '/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00368637202296470751086/270.dcm'\nprint('Filename: {}'.format(dcm))\ndcm = pydicom.read_file(dcm)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We have read the image in the dcm variable. We can print this variable to see the information related to this image.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"print(dcm)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can retrieve image as a numpy array by calling dcm.pixel_array. We can replace -2000's, which are essentially NAs with Os.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#display the image read above\nimg = dcm.pixel_array\nimg[img==-2000] = 0\n\nplt.axis('off')\nplt.imshow(img)\nplt.show()\n\nplt.axis('off')\nplt.imshow(-img) #invert colors with - \nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Lets Plot Some More Random Images**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#Helper function\ndef dicom_to_image(filename):\n    dcm = pydicom.read_file(filename)\n    img = dcm.pixel_array\n    img[img==-2000] == 0\n    return img","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Lets display 20 images at random\nfiles = glob.glob(train_dir+\"/*/*.dcm\")\n\nf,plots = plt.subplots(4,5, sharex='col', sharey='row',figsize=(10,8))\n\nfor i in range(20):\n    plots[i//5, i%5].axis('off')\n    plots[i//5, i%5].imshow(dicom_to_image(np.random.choice(files)), cmap=plt.cm.bone)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can try to reconstruct the layers of body from which the images were taken, by taking a single patient and sorting the scans by Slice Location.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#function to sort patient dcm images\ndef get_slice_location(dcm):\n    return float(dcm[0x0020, 0x1041].value)\n\n#Return a list of images for given patient in ascending order of slice location\ndef load_patient(patient_id):\n    files = glob.glob(train_dir+'/'+patient_id+'/*.dcm')\n    imgs={}\n    for f in files:\n        dcm = pydicom.read_file(f)\n        img = dcm.pixel_array\n        img[img== -2000] =0\n        sl= get_slice_location(dcm)\n        imgs[sl] = img\n        \n    \n    sorted_imgs = [x[1] for x in sorted(imgs.items(), key=lambda x: x[0])]\n    return sorted_imgs","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#display all the images for patient ID00210637202257228694086\npat = load_patient('ID00210637202257228694086')\nf, plots = plt.subplots(31, 10, sharex='col', sharey='row', figsize=(10, 31))\nfor i in range(303):\n    plots[i // 10, i % 10].axis('off')\n    plots[i // 10, i % 10].imshow(pat[i], cmap=plt.cm.bone)\n    ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Stakcking all Images of a Patient to Create a 3D Volume**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import matplotlib.animation as animation\nfrom IPython.display import HTML","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Stack up all 2D slices to make up the 3D Volume\ndef load_scan(patient_name):\n    \n    patient_directory = sorted(os.listdir(f'../input/osic-pulmonary-fibrosis-progression/train/{patient_name}'), key=(lambda f: int(f.split('.')[0])))\n    volume = np.zeros((len(patient_directory), 512, 512))\n\n    for i, img in enumerate(patient_directory):\n        img_slice = pydicom.dcmread(f'../input/osic-pulmonary-fibrosis-progression/train/{patient_name}/{img}')\n        volume[i] = img_slice.pixel_array\n            \n    return volume\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_scan = load_scan('ID00368637202296470751086')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(8, 8))\n\nimgs = []\nfor ps in patient_scan:\n    img = plt.imshow(ps, animated=True, cmap=plt.cm.bone)\n    plt.axis('off')\n    imgs.append([img])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"vid = animation.ArtistAnimation(fig, imgs, interval=25, blit=False, repeat_delay=1000)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# lets play the video \nHTML(vid.to_html5_video())","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}