{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"## MileStone 1 of Capstone project (RSNA-pneumonia-detection-challenge)"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"import numpy as np \nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom matplotlib.patches import Rectangle\nimport seaborn as sns\nimport gc\nimport glob\nimport os\nimport cv2\nimport pydicom\n\nimport warnings\nwarnings.simplefilter(action = 'ignore')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Loading CSV files"},{"metadata":{"trusted":true},"cell_type":"code","source":"detailed_df = pd.read_csv('/kaggle/input/rsna-pneumonia-detection-challenge/stage_2_detailed_class_info.csv')\ntrain_df = pd.read_csv('/kaggle/input/rsna-pneumonia-detection-challenge/stage_2_train_labels.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"## shape of detailed_df\ndetailed_df.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"## shape of train_df\ntrain_df.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"detailed_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Merging the data tables detailed_df and train_df"},{"metadata":{"trusted":true},"cell_type":"code","source":"df = pd.merge(left = detailed_df, right = train_df, how = 'left', on = 'patientId')\ndf = df.drop_duplicates()\ndf.info()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### It is clealy evident that the above data table contains lots of null values.\n\n## Summary on the values, types and null values:"},{"metadata":{"trusted":true},"cell_type":"code","source":"df.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Distribution of classes"},{"metadata":{},"cell_type":"markdown","source":"The following output shows that the nearly 2/3 of the patients do not have pneumonia (with target value = 0) and 1/3 of the patients have pneumonia (with target value =1)"},{"metadata":{"trusted":true},"cell_type":"code","source":"pd.pivot_table(df,index=[\"Target\"], values=['patientId'], aggfunc='count')\n\n# alternative approach\n# train_df['Target'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Distribution of patients in each class\nThere are 9555 patients in the category '**Lung Opacity**' and 11821 in '**No Lung Opacity / Not Normal**' category and 8851 are in **Normal** category"},{"metadata":{"trusted":true},"cell_type":"code","source":"pd.pivot_table(df,index=[\"class\"], values=['patientId'], aggfunc='count')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### The classes \"No Lung Opacity / Not Normal\", \"Normal\", and \"Lung Opacity\" are in the proportion of 39%, 29% and 32% respectively."},{"metadata":{"trusted":true},"cell_type":"code","source":"df[\"class\"].value_counts().plot(kind='pie',autopct='%1.0f%%', shadow=True, subplots=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### It is also clear from the below output that the patients who do not have pnuemonia do not have the bounding box coordinates"},{"metadata":{"trusted":true},"cell_type":"code","source":"pd.pivot_table(df,index=[\"Target\"], aggfunc='count')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Count of patients having single row and more than single rows"},{"metadata":{"trusted":true},"cell_type":"code","source":"df['patientId'].value_counts().value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Patients who do not have pneumonia has only one record in the table"},{"metadata":{"trusted":true},"cell_type":"code","source":"df[df['Target'] == 0]['patientId'].value_counts().value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.countplot(x = 'class', hue = 'Target', data = df)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Preprocessing - Filling the null values"},{"metadata":{"trusted":true},"cell_type":"code","source":"df.fillna(0.0)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Correlation between the variables\nThere is a strong colleation between height and width variables"},{"metadata":{"trusted":true},"cell_type":"code","source":"df.corr()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.jointplot(x = 'width', y = 'height', data = df, kind=\"reg\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### EDA with the header values from the dataframe\n\nCreating a data frame with all of their appropriate header values from the dicom file takes long time as there are 30277 records. Hence, the EDA analysis is done on a subset of randomly chosen 1000 records by keeping the same proportion of the classes.\n(i.e) The classes Not Normal, Normal, Lunge Opacity are in a proportion 39%, 29%, and 32% respectively.\n\n* Number of rows of Not Normal class = 39% of 1000 = 390 rows\n* Number of rows of Normal class = 29% of 1000 = 290 rows\n* Number of rows of Lunge Opacity class = 32% of 1000 = 320 rows\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"df_Not_Normal = df[df['class']=='No Lung Opacity / Not Normal'].sample(n=390)\ndf_Normal = df[df['class']=='Normal'].sample(n=290)\ndf_Lunge_Opacity = df[df['class']=='Lung Opacity'].sample(n=320)\nframes = [df_Not_Normal, df_Normal, df_Lunge_Opacity]\n\ndicom_df = pd.concat(frames)\n\ndicom_df.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def process_dicom_data(data_df):\n    for n, pid in enumerate(data_df['patientId'].unique()):        \n        dcm_file = '/kaggle/input/rsna-pneumonia-detection-challenge/stage_2_train_images/%s.dcm' % pid\n        dcm_data = pydicom.read_file(dcm_file)        \n        idx = (data_df['patientId']==dcm_data.PatientID)\n        data_df.loc[idx,'Modality'] = dcm_data.Modality\n        data_df.loc[idx,'PatientAge'] = pd.to_numeric(dcm_data.PatientAge)\n        data_df.loc[idx,'PatientSex'] = dcm_data.PatientSex\n        data_df.loc[idx,'BodyPartExamined'] = dcm_data.BodyPartExamined\n        data_df.loc[idx,'ViewPosition'] = dcm_data.ViewPosition\n        \n    return data_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dicom_df = process_dicom_data(dicom_df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# converting PatientAge to int as it is in float\ndicom_df = dicom_df.astype({\"PatientAge\": int})\ndicom_df.fillna(0.0, inplace=True)\ndicom_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### There are 995 unique patient rows exist"},{"metadata":{"trusted":true},"cell_type":"code","source":"dicom_df.nunique()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Now Visualizing the data along with their dicom header values"},{"metadata":{},"cell_type":"markdown","source":"### Patient's age proportion in the detection"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize = (30, 10))\nsns.countplot(x = 'PatientAge', hue = 'Target', data = dicom_df)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Patient's gender proportion in the detection"},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.countplot(x = 'PatientSex', hue = 'Target', data = dicom_df)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### With respect to view proportion"},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.countplot(x = 'ViewPosition', hue = 'Target', data = dicom_df);","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dicom_df = dicom_df.drop('Target', axis=1)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dicom_df['PatientSex'].astype('category')\ndicom_df['ViewPosition'].astype('category')\ndicom_df['PatientSex'] = np.where(dicom_df[\"PatientSex\"].str.contains(\"M\"), 1, 0)\ndicom_df['ViewPosition'] = np.where(dicom_df[\"ViewPosition\"].str.contains(\"AP\"), 1, 0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dicom_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Apart from the correlation between the width and height,there is no strong correlation between the other variables in the dataframe"},{"metadata":{"trusted":true},"cell_type":"code","source":"dicom_df.corr()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Visualizing the dicom images"},{"metadata":{"trusted":true},"cell_type":"code","source":"def show_dicom_image(data_df):\n        img_data = list(data_df.T.to_dict().values())\n        f, ax = plt.subplots(2,2, figsize=(16,18))\n        for i,data_row in enumerate(img_data):\n            pid = data_row['patientId']\n            dcm_file = '/kaggle/input/rsna-pneumonia-detection-challenge/stage_2_train_images/%s.dcm' % pid\n            dcm_data = pydicom.read_file(dcm_file)                    \n            ax[i//2, i%2].imshow(dcm_data.pixel_array, cmap=plt.cm.bone)\n            ax[i//2, i%2].set_title('ID: {}\\n Age: {} Sex: {}'.format(\n                data_row['patientId'],dcm_data.PatientAge, dcm_data.PatientSex))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Showing some random dicom images of a patients who have Pnuemonia"},{"metadata":{"trusted":true},"cell_type":"code","source":"show_dicom_image(df[df['Target']==1].sample(n=4))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Showing some random dicom images of a patient who do not have Pnuemonia, however with class ***No Lung Opacity / Not Normal***"},{"metadata":{"trusted":true},"cell_type":"code","source":"show_dicom_image(df[ (df['Target']==0) & (df['class']=='No Lung Opacity / Not Normal')].sample(n=4))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Showing some random dicom images of a patients who do not have Pnuemonia, however with class ***Normal***"},{"metadata":{"trusted":true},"cell_type":"code","source":"show_dicom_image(df[ (df['Target']==0) & (df['class']=='Normal')].sample(n=4))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def show_dicome_with_boundingbox(data_df):\n    img_data = list(data_df.T.to_dict().values())\n    f, ax = plt.subplots(2,2, figsize=(16,18))\n    for i,data_row in enumerate(img_data):\n        pid = data_row['patientId']\n        dcm_file = '/kaggle/input/rsna-pneumonia-detection-challenge/stage_2_train_images/%s.dcm' % pid\n        dcm_data = pydicom.read_file(dcm_file)                    \n        ax[i//2, i%2].imshow(dcm_data.pixel_array, cmap=plt.cm.bone)\n        ax[i//2, i%2].set_title('ID: {}\\n Age: {} Sex: {}'.format(\n                data_row['patientId'],dcm_data.PatientAge, dcm_data.PatientSex))\n        rows = data_df[data_df['patientId']==data_row['patientId']]\n        box_data = list(rows.T.to_dict().values())        \n        for j, row in enumerate(box_data):            \n            x,y,width,height = row['x'], row['y'],row['width'],row['height']\n            rectangle = Rectangle(xy=(x,y),width=width, height=height, color=\"red\",alpha = 0.1)\n            ax[i//2, i%2].add_patch(rectangle)            ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"show_dicome_with_boundingbox(df[df['Target']==1].sample(n=4))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Building the pneumonia detection model using CNN"},{"metadata":{"trusted":true},"cell_type":"code","source":"from keras.models import Sequential\nfrom keras.layers import Conv2D\nfrom keras.layers import MaxPooling2D\nfrom keras.layers import Flatten\nfrom keras.layers import Dense\nfrom keras.layers.core import Flatten, Dense, Dropout\nfrom keras.layers.convolutional import Convolution2D, MaxPooling2D, ZeroPadding2D\nfrom keras.optimizers import SGD\nfrom keras.preprocessing.image import ImageDataGenerator","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"IMAGE_SIZE = [224, 224]\n\ntrain_path = '/kaggle/input/rsna-pneumonia-detection-challenge/stage_2_train_images/'\ntest_path = '/kaggle/input/rsna-pneumonia-detection-challenge/stage_2_test_images/'","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}