{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**Problem we want to solve:**\nThis is an object detection and classification problem. In this competition, we are identifying and localizing COVID-19 abnormalities on chest radiographs. \n\nDataset information\nThe train dataset comprises 6,334 chest scans in DICOM format, which were de-identified to protect patient privacy. All images were labeled by a panel of experienced radiologists for the presence of opacities as well as overall appearance.\n\nNote that all images are stored in paths with the form study/series/image. The study ID here relates directly to the study-level predictions, and the image ID is the ID used for image-level predictions.\n\nThe hidden test dataset is of roughly the same scale as the training dataset.\n\nFiles\ntrain_study_level.csv - the train study-level metadata, with one row for each study, including correct labels.\ntrain_image_level.csv - the train image-level metadata, with one row for each image, including both correct labels and any bounding boxes in a dictionary format. Some images in both test and train have multiple bounding boxes.\nsample_submission.csv - a sample submission file containing all image- and study-level IDs.\nColumns\ntrain_study_level.csv\n\nid - unique study identifier\nNegative for Pneumonia - 1 if the study is negative for pneumonia, 0 otherwise\nTypical Appearance - 1 if the study has this appearance, 0 otherwise\nIndeterminate Appearance  - 1 if the study has this appearance, 0 otherwise\nAtypical Appearance  - 1 if the study has this appearance, 0 otherwise\ntrain_image_level.csv\n\nid - unique image identifier\nboxes - bounding boxes in easily-readable dictionary format\nlabel - the correct prediction label for the provided bounding boxes\n\n**DICM (digital imaging and communications in medicine) files:**\nThe data are DICM (digital imaging and communications in medicine) files. A DICOM file consists of a header and image data sets packed into a single file. For more information about DICM, I advise you to read “Managing DICOM images: Tips and tricks for the radiologist” article (https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3354356/).\n\n**DICOM Data description:**\nWe have train data with (6054 directories) and test data with (1214 directories). But in total we have 6334 unique values which refer to no. of patients. \nThe DICOM image file constitute the header which stores demographic information about the patient, acquisition parameters for the imaging study, image dimensions, matrix size, color space, and a host of additional non-intensity information required by the computer to correctly display the image. After the header, a single attribute (7FE0) that contains all the pixel intensity data for the image which is stored as a long series of 0s and 1s.\nTo have look on the demographic information, we need to read the data and for this we going to use Pydicom package.\n\n","metadata":{"execution":{"iopub.status.busy":"2021-06-20T08:53:51.872116Z","iopub.execute_input":"2021-06-20T08:53:51.872912Z","iopub.status.idle":"2021-06-20T08:53:51.886597Z","shell.execute_reply.started":"2021-06-20T08:53:51.872741Z","shell.execute_reply":"2021-06-20T08:53:51.88468Z"}}},{"cell_type":"code","source":"#pip install pydicom\n#pip install pylibjpeg pylibjpeg-libjpeg pylibjpeg-openjpeg","metadata":{"execution":{"iopub.status.busy":"2021-07-25T08:02:37.950738Z","iopub.execute_input":"2021-07-25T08:02:37.95136Z","iopub.status.idle":"2021-07-25T08:02:37.955628Z","shell.execute_reply.started":"2021-07-25T08:02:37.951267Z","shell.execute_reply":"2021-07-25T08:02:37.954901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pydicom\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\n\n\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n#reading all the DICM files in train folder and test folder\nimport os\n\n","metadata":{"execution":{"iopub.status.busy":"2021-07-25T08:02:37.956868Z","iopub.execute_input":"2021-07-25T08:02:37.957268Z","iopub.status.idle":"2021-07-25T08:02:38.291311Z","shell.execute_reply.started":"2021-07-25T08:02:37.957237Z","shell.execute_reply":"2021-07-25T08:02:38.29017Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\ntrain=[]\ntest = []\npath = ('/kaggle/input/siim-covid19-detection')\n\nfor dirname, _, filenames in os.walk(path + '/train'):\n    for filename in filenames:\n        train.append(os.path.join(dirname, filename)) \n        \n        \n        \nfor dirname, _, filenames in os.walk('/kaggle/input/siim-covid19-detection/test'):\n    for filename in filenames:\n        test.append(os.path.join(dirname, filename))       \n        \n        \n        \n\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2021-07-25T08:02:38.293884Z","iopub.execute_input":"2021-07-25T08:02:38.294369Z","iopub.status.idle":"2021-07-25T08:03:09.311576Z","shell.execute_reply.started":"2021-07-25T08:02:38.294323Z","shell.execute_reply":"2021-07-25T08:03:09.310545Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#read the first file in the train folder\nds = pydicom.read_file(train[0])\nprint(ds)","metadata":{"execution":{"iopub.status.busy":"2021-07-25T08:03:09.313162Z","iopub.execute_input":"2021-07-25T08:03:09.31346Z","iopub.status.idle":"2021-07-25T08:03:09.721523Z","shell.execute_reply.started":"2021-07-25T08:03:09.313431Z","shell.execute_reply":"2021-07-25T08:03:09.720538Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If we want to reach to pixel array to plot the image, we can use ds.pixel_array as below","metadata":{}},{"cell_type":"code","source":"\nplt.imshow(ds.pixel_array, cmap='gray') ","metadata":{"execution":{"iopub.status.busy":"2021-07-25T08:03:09.722799Z","iopub.execute_input":"2021-07-25T08:03:09.723105Z","iopub.status.idle":"2021-07-25T08:03:10.592023Z","shell.execute_reply.started":"2021-07-25T08:03:09.723075Z","shell.execute_reply":"2021-07-25T08:03:10.591027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Converting x-ray to a meaning image with Hounsfield units:**\n\nA \"Modality LUT \" allows the transformation of manufacturer-dependent pixel values into manufacturer-independent pixel values (e.g., Hounsfield units for CT images). The \"Modality LUT\" may be contained within an image, a Presentation State (that references an image), or as a \"Standalone Modality LUT\" (that references an image). The \"Rescale Slope\" (0028,1053) and \"Rescale Intercept\" (0028,1052) elements are used to describe the transformation when it is linear, while the \"Modality LUT Sequence\" (0028,3000) element is used to describe non-linear transformations. In both cases it is implied that the transformation can only occur for grayscale data, that is, images with \"Photometric Interpretation\" values of \"MONOCHROME1\" or \"MONOCHROME2\".\nhttps://www.leadtools.com/help/sdk/v21/dh/to/working-with-dicom-lut.html\nhttps://www.kaggle.com/raddar/convert-dicom-to-np-array-the-correct-way","metadata":{"execution":{"iopub.status.busy":"2021-06-22T17:15:15.703405Z","iopub.execute_input":"2021-06-22T17:15:15.703722Z","iopub.status.idle":"2021-06-22T17:15:15.710654Z","shell.execute_reply.started":"2021-06-22T17:15:15.703694Z","shell.execute_reply":"2021-06-22T17:15:15.709669Z"}}},{"cell_type":"code","source":"\ndef read_xray(path, voi_lut = True, fix_monochrome = True):\n    dicom = pydicom.dcmread(path)\n    \n    # VOI LUT (if available by DICOM device) is used to transform raw DICOM data to \"human-friendly\" view\n    if voi_lut:\n        data = apply_voi_lut(dicom.pixel_array, dicom)\n    else:\n        data = dicom.pixel_array\n               \n    # depending on this value, X-ray may look inverted - fix that:\n    if fix_monochrome and dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = np.amax(data) - data\n        \n    data = data - np.min(data)\n    data = data / np.max(data)\n    data = (data * 255).astype(np.uint8)\n        \n    return data","metadata":{"execution":{"iopub.status.busy":"2021-07-25T08:03:10.593412Z","iopub.execute_input":"2021-07-25T08:03:10.593716Z","iopub.status.idle":"2021-07-25T08:03:10.601063Z","shell.execute_reply.started":"2021-07-25T08:03:10.593685Z","shell.execute_reply":"2021-07-25T08:03:10.599482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img = read_xray(train[0])\nplt.figure(figsize = (6,6))\nplt.imshow(img, 'gray')\n","metadata":{"execution":{"iopub.status.busy":"2021-07-25T08:03:10.603112Z","iopub.execute_input":"2021-07-25T08:03:10.603536Z","iopub.status.idle":"2021-07-25T08:03:11.54993Z","shell.execute_reply.started":"2021-07-25T08:03:10.603492Z","shell.execute_reply":"2021-07-25T08:03:11.548691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we are going to study the other csv file included in input data set","metadata":{}},{"cell_type":"code","source":"train_study = pd.read_csv(path+'/train_study_level.csv')\ntrain_image = pd.read_csv(path+'/train_image_level.csv')\n","metadata":{"execution":{"iopub.status.busy":"2021-07-25T08:03:11.551432Z","iopub.execute_input":"2021-07-25T08:03:11.551802Z","iopub.status.idle":"2021-07-25T08:03:11.610855Z","shell.execute_reply.started":"2021-07-25T08:03:11.551743Z","shell.execute_reply":"2021-07-25T08:03:11.609841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_study.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-25T08:03:11.613542Z","iopub.execute_input":"2021-07-25T08:03:11.613901Z","iopub.status.idle":"2021-07-25T08:03:11.638824Z","shell.execute_reply.started":"2021-07-25T08:03:11.613869Z","shell.execute_reply":"2021-07-25T08:03:11.637879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_study.shape","metadata":{"execution":{"iopub.status.busy":"2021-07-25T08:03:11.640486Z","iopub.execute_input":"2021-07-25T08:03:11.640816Z","iopub.status.idle":"2021-07-25T08:03:11.64704Z","shell.execute_reply.started":"2021-07-25T08:03:11.64076Z","shell.execute_reply":"2021-07-25T08:03:11.645791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cases_no = train_study[['Negative for Pneumonia', 'Typical Appearance', 'Indeterminate Appearance', 'Atypical Appearance']].values.sum(axis = 0)\n","metadata":{"execution":{"iopub.status.busy":"2021-07-25T08:03:11.648491Z","iopub.execute_input":"2021-07-25T08:03:11.648889Z","iopub.status.idle":"2021-07-25T08:03:11.662074Z","shell.execute_reply.started":"2021-07-25T08:03:11.648784Z","shell.execute_reply":"2021-07-25T08:03:11.660977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.bar( ['Negative for Pneumonia', 'Typical Appearance', 'Indeterminate Appearance', 'Atypical Appearance'],cases_no)\nplt.xticks(rotation=45)\nplt.ylabel('Frequency')","metadata":{"execution":{"iopub.status.busy":"2021-07-25T08:03:11.663533Z","iopub.execute_input":"2021-07-25T08:03:11.663876Z","iopub.status.idle":"2021-07-25T08:03:11.806946Z","shell.execute_reply.started":"2021-07-25T08:03:11.663844Z","shell.execute_reply":"2021-07-25T08:03:11.805996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_image.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-25T08:03:11.808666Z","iopub.execute_input":"2021-07-25T08:03:11.80904Z","iopub.status.idle":"2021-07-25T08:03:11.822485Z","shell.execute_reply.started":"2021-07-25T08:03:11.808998Z","shell.execute_reply":"2021-07-25T08:03:11.821441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_image.shape","metadata":{"execution":{"iopub.status.busy":"2021-07-25T08:03:11.823692Z","iopub.execute_input":"2021-07-25T08:03:11.824022Z","iopub.status.idle":"2021-07-25T08:03:11.830465Z","shell.execute_reply.started":"2021-07-25T08:03:11.823992Z","shell.execute_reply":"2021-07-25T08:03:11.829718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From above we realize that no. of study less than no. of images because each study can have more than one image. ID in train_study file and StudyInstanceUID in train_image file are same and for this we are going to merge both files depending on this variabel  ","metadata":{}},{"cell_type":"code","source":"train_study['id'] = train_study['id'].str.replace('_study',\"\")","metadata":{"execution":{"iopub.status.busy":"2021-07-25T08:03:11.831401Z","iopub.execute_input":"2021-07-25T08:03:11.831698Z","iopub.status.idle":"2021-07-25T08:03:11.848023Z","shell.execute_reply.started":"2021-07-25T08:03:11.831659Z","shell.execute_reply":"2021-07-25T08:03:11.847078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_study.rename({'id': 'StudyInstanceUID'})","metadata":{"execution":{"iopub.status.busy":"2021-07-25T08:04:19.947162Z","iopub.execute_input":"2021-07-25T08:04:19.947536Z","iopub.status.idle":"2021-07-25T08:04:19.967466Z","shell.execute_reply.started":"2021-07-25T08:04:19.947502Z","shell.execute_reply":"2021-07-25T08:04:19.966641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output1 = pd.merge(train_study, train_image, \n                   on='id', \n                   how='inner')","metadata":{"execution":{"iopub.status.busy":"2021-07-25T08:08:42.771546Z","iopub.execute_input":"2021-07-25T08:08:42.771937Z","iopub.status.idle":"2021-07-25T08:08:42.795557Z","shell.execute_reply.started":"2021-07-25T08:08:42.771906Z","shell.execute_reply":"2021-07-25T08:08:42.794431Z"},"trusted":true},"execution_count":null,"outputs":[]}]}