{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# SIIM-FISABIO-RSNA COVID-19 Detection: A simple EDA \n\nIn this competition, we are provided with DICOM images of chest X-ray radiographs, and we are asked to identify and localize COVID-19 abnormalities. This is important because typical diagnosis of COVID-19 requires molecular testing (polymerase chain reaction) requires several hours, while chest radiographs can be obtained in minutes, but it is hard to distinguish between COVID-19 pneumonia and other other viral and bacterial pneumonias. Therefore, in this competition, be hope to develop AI that that eventually help radiologists diagnose the millions of COVID-19 patients more confidently and quickly.\n\nI'll provide a quick and simple EDA to help you get started with this very interesting competition!","metadata":{}},{"cell_type":"markdown","source":"# Imports\nLet's start out by setting up our environment by importing the required modules:","metadata":{}},{"cell_type":"code","source":"# Thanks to https://www.kaggle.com/awsaf49/pydicom-conda-helper for pydicom files\n\n# !wget 'https://anaconda.org/conda-forge/libjpeg-turbo/2.1.0/download/linux-64/libjpeg-turbo-2.1.0-h7f98852_0.tar.bz2' -q\n# !wget 'https://anaconda.org/conda-forge/libgcc-ng/9.3.0/download/linux-64/libgcc-ng-9.3.0-h2828fa1_19.tar.bz2' -q\n# !wget 'https://anaconda.org/conda-forge/gdcm/2.8.9/download/linux-64/gdcm-2.8.9-py37h500ead1_1.tar.bz2' -q\n# !wget 'https://anaconda.org/conda-forge/conda/4.10.1/download/linux-64/conda-4.10.1-py37h89c1867_0.tar.bz2' -q\n# !wget 'https://anaconda.org/conda-forge/certifi/2020.12.5/download/linux-64/certifi-2020.12.5-py37h89c1867_1.tar.bz2' -q\n# !wget 'https://anaconda.org/conda-forge/openssl/1.1.1k/download/linux-64/openssl-1.1.1k-h7f98852_0.tar.bz2' -q\n\n!conda install 'libjpeg-turbo-2.1.0-h7f98852_0.tar.bz2' -c conda-forge -y\n!conda install 'libgcc-ng-9.3.0-h2828fa1_19.tar.bz2' -c conda-forge -y\n!conda install 'gdcm-2.8.9-py37h500ead1_1.tar.bz2' -c conda-forge -y\n!conda install 'conda-4.10.1-py37h89c1867_0.tar.bz2' -c conda-forge -y\n!conda install 'certifi-2020.12.5-py37h89c1867_1.tar.bz2' -c conda-forge -y\n!conda install 'openssl-1.1.1k-h7f98852_0.tar.bz2' -c conda-forge -y","metadata":{"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nimport glob\nfrom tqdm.notebook import tqdm\nimport matplotlib.pyplot as plt\nfrom skimage import exposure\nimport cv2\nimport warnings\nfrom fastai.vision.all import *\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:24:42.969734Z","iopub.execute_input":"2021-07-14T10:24:42.970774Z","iopub.status.idle":"2021-07-14T10:24:45.309909Z","shell.execute_reply.started":"2021-07-14T10:24:42.970662Z","shell.execute_reply":"2021-07-14T10:24:45.308639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.options.display.max_columns = 500\npd.options.display.max_rows=1000","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:24:45.319262Z","iopub.execute_input":"2021-07-14T10:24:45.319832Z","iopub.status.idle":"2021-07-14T10:24:45.325921Z","shell.execute_reply.started":"2021-07-14T10:24:45.319785Z","shell.execute_reply":"2021-07-14T10:24:45.324060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from fastai.basics import *\nfrom fastai.callback.all import *\nfrom fastai.vision.all import *\nfrom fastai.medical.imaging import *\n\nimport pydicom\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:24:45.328175Z","iopub.execute_input":"2021-07-14T10:24:45.328691Z","iopub.status.idle":"2021-07-14T10:24:45.824278Z","shell.execute_reply.started":"2021-07-14T10:24:45.328643Z","shell.execute_reply":"2021-07-14T10:24:45.822990Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# A look at the provided data","metadata":{}},{"cell_type":"markdown","source":"Let's check what data is available to us:","metadata":{}},{"cell_type":"code","source":"dataset_path = Path('../input/siim-covid19-detection')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2021-07-14T10:24:48.968433Z","iopub.execute_input":"2021-07-14T10:24:48.968819Z","iopub.status.idle":"2021-07-14T10:24:48.975462Z","shell.execute_reply.started":"2021-07-14T10:24:48.968783Z","shell.execute_reply":"2021-07-14T10:24:48.974229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset_path.ls()","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:24:53.468258Z","iopub.execute_input":"2021-07-14T10:24:53.468707Z","iopub.status.idle":"2021-07-14T10:24:53.483820Z","shell.execute_reply.started":"2021-07-14T10:24:53.468659Z","shell.execute_reply":"2021-07-14T10:24:53.482556Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that we have:\n\n* `train_study_level.csv` - the train study-level metadata, with one row for each study, including correct labels.\n* `train_image_level.csv` - the train image-level metadata, with one row for each image, including both correct labels and any bounding boxes in a dictionary format. Some images in both test and train have multiple bounding boxes.\n* `sample_submission.csv` - a sample submission file containing all image- and study-level IDs.\n* `train` folder - comprises 6,334 chest scans in DICOM format, stored in paths with the form `study`/`series`/`image`\n* `test` folder \n\nThe hidden test dataset is of roughly the same scale as the training dataset.\n","metadata":{}},{"cell_type":"markdown","source":"# A look at the CSVs\n\nLet's check the `train_study_level.csv` file:","metadata":{}},{"cell_type":"code","source":"train_study_df = pd.read_csv(dataset_path/'train_study_level.csv')\ntrain_image_df = pd.read_csv(dataset_path/'train_image_level.csv')","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:24:54.722131Z","iopub.execute_input":"2021-07-14T10:24:54.722551Z","iopub.status.idle":"2021-07-14T10:24:54.788448Z","shell.execute_reply.started":"2021-07-14T10:24:54.722516Z","shell.execute_reply":"2021-07-14T10:24:54.787061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_study_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:24:56.013957Z","iopub.execute_input":"2021-07-14T10:24:56.014364Z","iopub.status.idle":"2021-07-14T10:24:56.034655Z","shell.execute_reply.started":"2021-07-14T10:24:56.014332Z","shell.execute_reply":"2021-07-14T10:24:56.033425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_image_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:24:56.961870Z","iopub.execute_input":"2021-07-14T10:24:56.962319Z","iopub.status.idle":"2021-07-14T10:24:56.977058Z","shell.execute_reply.started":"2021-07-14T10:24:56.962288Z","shell.execute_reply":"2021-07-14T10:24:56.975664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's look at the unique labels:","metadata":{}},{"cell_type":"code","source":"train_study_df.shape","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:25:03.835421Z","iopub.execute_input":"2021-07-14T10:25:03.835821Z","iopub.status.idle":"2021-07-14T10:25:03.847015Z","shell.execute_reply.started":"2021-07-14T10:25:03.835789Z","shell.execute_reply":"2021-07-14T10:25:03.845327Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"study_classes = ['Negative for Pneumonia', 'Typical Appearance', 'Indeterminate Appearance', 'Atypical Appearance']\ntrain_study_df[study_classes].value_counts()","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:25:04.805521Z","iopub.execute_input":"2021-07-14T10:25:04.805976Z","iopub.status.idle":"2021-07-14T10:25:04.824016Z","shell.execute_reply.started":"2021-07-14T10:25:04.805903Z","shell.execute_reply":"2021-07-14T10:25:04.822456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you can see, at the study-level, we are predicting the following classes:\n* Negative for Pneumonia\n* Typical Appearance\n* Indeterminate Appearance\n* Atypical Appearance\n\nThis here is a standard multi-label classification problem. In the training set, interestingly they are not multi-label, but it is mentioned that:\n> Studies in the test set may contain more than one label.\n\nLet's look at the distribution:\n","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize = (10,5))\nplt.bar([1,2,3,4], train_study_df[study_classes].values.sum(axis=0))\nplt.xticks([1,2,3,4],study_classes)\nplt.ylabel('Frequency')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:25:06.845259Z","iopub.execute_input":"2021-07-14T10:25:06.845714Z","iopub.status.idle":"2021-07-14T10:25:07.024535Z","shell.execute_reply.started":"2021-07-14T10:25:06.845681Z","shell.execute_reply":"2021-07-14T10:25:07.023017Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's now look at `train_image_level.csv`:","metadata":{}},{"cell_type":"markdown","source":"We have our bounding box labels provided in the `label` column. The format is as follows:\n\n`[class ID] [confidence score] [bounding box]`\n\n* class ID - either `opacity` or `none`\n* confidence score - confidence from your neural network model. If none, the confidence is `1`.\n* bounding box - typical `xmin ymin xmax ymax` format. If class ID is none, the bounding box is `1 0 0 1 1`.\n\nThe bounding boxes are also provided in easily readable dictionary format in column `boxes`, and the study that each image is a part of is provided in`StudyInstanceUID`.\n\nLet's quick look at the distribution of opacity vs none:","metadata":{}},{"cell_type":"code","source":"train_image_df['split_label'] = train_image_df.label.apply(lambda x: [x.split()[offs:offs+6] for offs in range(0, len(x.split()), 6) ]) # start, stop, step","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:25:09.462197Z","iopub.execute_input":"2021-07-14T10:25:09.462618Z","iopub.status.idle":"2021-07-14T10:25:09.494463Z","shell.execute_reply.started":"2021-07-14T10:25:09.462585Z","shell.execute_reply":"2021-07-14T10:25:09.493341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_image_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:25:18.753698Z","iopub.execute_input":"2021-07-14T10:25:18.754199Z","iopub.status.idle":"2021-07-14T10:25:18.774491Z","shell.execute_reply.started":"2021-07-14T10:25:18.754150Z","shell.execute_reply":"2021-07-14T10:25:18.773187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# this is only the distribution of image level labels: opacity and none\nclasses_freq = []\nfor i in range(len(train_image_df)):\n    for j in train_image_df.iloc[i].split_label: classes_freq.append(j[0])\nplt.hist(classes_freq)\nplt.ylabel('Frequency')\n\nprint(classes_freq[:10])","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:25:28.096907Z","iopub.execute_input":"2021-07-14T10:25:28.097410Z","iopub.status.idle":"2021-07-14T10:25:29.396451Z","shell.execute_reply.started":"2021-07-14T10:25:28.097377Z","shell.execute_reply":"2021-07-14T10:25:29.395234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"we see more labels than the number of images because some images are tagged with more than 1 label because they have more than 1 nodes in them","metadata":{"execution":{"iopub.status.busy":"2021-06-22T16:06:34.428623Z","iopub.execute_input":"2021-06-22T16:06:34.429075Z","iopub.status.idle":"2021-06-22T16:06:34.432213Z","shell.execute_reply.started":"2021-06-22T16:06:34.429027Z","shell.execute_reply":"2021-06-22T16:06:34.431187Z"}}},{"cell_type":"markdown","source":"Let's also look at the distribution of the bounding box areas:","metadata":{}},{"cell_type":"code","source":"train_image_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:25:32.333546Z","iopub.execute_input":"2021-07-14T10:25:32.333975Z","iopub.status.idle":"2021-07-14T10:25:32.353782Z","shell.execute_reply.started":"2021-07-14T10:25:32.333905Z","shell.execute_reply":"2021-07-14T10:25:32.352582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bbox_areas = []\nfor i in range(len(train_image_df)):\n    for j in train_image_df.iloc[i].split_label:\n        bbox_areas.append((float(j[4])-float(j[2]))*(float(j[5])-float(j[3])))\nplt.hist(bbox_areas)\nplt.ylabel('Frequency')\n\nprint(bbox_areas[:10])","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:26:18.959531Z","iopub.execute_input":"2021-07-14T10:26:18.959902Z","iopub.status.idle":"2021-07-14T10:26:20.126586Z","shell.execute_reply.started":"2021-07-14T10:26:18.959870Z","shell.execute_reply":"2021-07-14T10:26:20.125019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# A look at the images\n\nOkay, let's now look at some example images:","metadata":{}},{"cell_type":"code","source":"# https://www.kaggle.com/raddar/convert-dicom-to-np-array-the-correct-way\n    \ndef dicom2array(path, voi_lut=False, fix_monochrome=True):\n    dicom = pydicom.read_file(path)\n    # VOI LUT (if available by DICOM device) is used to\n    # transform raw DICOM data to \"human-friendly\" view\n    if voi_lut:\n        data = apply_voi_lut(dicom.pixel_array, dicom)\n    else:\n        data = dicom.pixel_array\n    # depending on this value, X-ray may look inverted - fix that:\n    if fix_monochrome and dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = np.amax(data) - data\n    data = data - np.min(data)\n    data = data / np.max(data)\n    data = (data * 255).astype(np.uint8)\n    return data\n        \n    \ndef plot_img(img, size=(7, 7), is_rgb=True, title=\"\", cmap='gray'):\n    plt.figure(figsize=size)\n    plt.imshow(img, cmap=cmap)\n    plt.suptitle(title)\n    plt.show()\n\n\ndef plot_imgs(imgs, cols=4, size=7, is_rgb=True, title=\"\", cmap='gray', img_size=(500,500)):\n    rows = len(imgs)//cols + 1\n    fig = plt.figure(figsize=(cols*size, rows*size))\n    for i, img in enumerate(imgs):\n        if img_size is not None:\n            img = cv2.resize(img, img_size)\n        fig.add_subplot(rows, cols, i+1)\n        plt.imshow(img, cmap=cmap)\n    plt.suptitle(title)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:26:37.939123Z","iopub.execute_input":"2021-07-14T10:26:37.939666Z","iopub.status.idle":"2021-07-14T10:26:37.972350Z","shell.execute_reply.started":"2021-07-14T10:26:37.939605Z","shell.execute_reply":"2021-07-14T10:26:37.970838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dicom_paths = get_dicom_files(dataset_path/'train')[:10]\nimgs = [dicom2array(path) for path in dicom_paths]\nplot_imgs(imgs)","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:26:43.132413Z","iopub.execute_input":"2021-07-14T10:26:43.132811Z","iopub.status.idle":"2021-07-14T10:27:12.767664Z","shell.execute_reply.started":"2021-07-14T10:26:43.132778Z","shell.execute_reply":"2021-07-14T10:27:12.766555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's actually look at how many images are available per study:","metadata":{}},{"cell_type":"code","source":"num_images_per_study = []\nfor i in (dataset_path/'train').ls():\n    num_images_per_study.append(len(get_dicom_files(i)))\n    if len(get_dicom_files(i)) > 2:\n        print(f'Study {i} had {len(get_dicom_files(i))} images')","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:27:21.226587Z","iopub.execute_input":"2021-07-14T10:27:21.227022Z","iopub.status.idle":"2021-07-14T10:27:33.523221Z","shell.execute_reply.started":"2021-07-14T10:27:21.226983Z","shell.execute_reply":"2021-07-14T10:27:33.521995Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.hist(num_images_per_study)","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:27:33.525340Z","iopub.execute_input":"2021-07-14T10:27:33.525851Z","iopub.status.idle":"2021-07-14T10:27:33.787409Z","shell.execute_reply.started":"2021-07-14T10:27:33.525789Z","shell.execute_reply":"2021-07-14T10:27:33.786236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def image_path(row):\n    study_path = dataset_path/'train'/row.StudyInstanceUID\n    for i in get_dicom_files(study_path):\n        if row.id.split('_')[0] == i.stem: return i \n        \ntrain_image_df['image_path'] = train_image_df.apply(image_path, axis=1)","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:27:35.007763Z","iopub.execute_input":"2021-07-14T10:27:35.008246Z","iopub.status.idle":"2021-07-14T10:27:42.954959Z","shell.execute_reply.started":"2021-07-14T10:27:35.008197Z","shell.execute_reply":"2021-07-14T10:27:42.953705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_image_df['image_path'].iloc[0].stem, train_image_df['image_path'].iloc[0]","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:27:42.956720Z","iopub.execute_input":"2021-07-14T10:27:42.957236Z","iopub.status.idle":"2021-07-14T10:27:42.968466Z","shell.execute_reply.started":"2021-07-14T10:27:42.957186Z","shell.execute_reply":"2021-07-14T10:27:42.966994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_image_df.head(1)","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:34:56.076000Z","iopub.execute_input":"2021-07-14T10:34:56.076403Z","iopub.status.idle":"2021-07-14T10:34:56.093108Z","shell.execute_reply.started":"2021-07-14T10:34:56.076355Z","shell.execute_reply":"2021-07-14T10:34:56.091676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"imgs = []\nimage_paths = train_image_df['image_path'].values\n\n# map label_id to specify color\nthickness = 10\nscale = 1\n\n\nfor i in range(8):\n    image_path = random.choice(image_paths)\n    print(image_path)\n    img = dicom2array(path=image_path)\n    img = cv2.resize(img, None, fx=1/scale, fy=1/scale)\n    img = np.stack([img, img, img], axis=-1)\n    for i in train_image_df.loc[train_image_df['image_path'] == image_path].split_label.values[0]:\n        if i[0] == 'opacity':\n            img = cv2.rectangle(img,\n                                (int(float(i[2])/scale), int(float(i[3])/scale)),\n                                (int(float(i[4])/scale), int(float(i[5])/scale)),\n                                [255,0,0], thickness)\n    \n    img = cv2.resize(img, (500,500))\n    imgs.append(img)\n    \nplot_imgs(imgs, cmap='gray')","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:37:52.053555Z","iopub.execute_input":"2021-07-14T10:37:52.053987Z","iopub.status.idle":"2021-07-14T10:37:56.975457Z","shell.execute_reply.started":"2021-07-14T10:37:52.053917Z","shell.execute_reply":"2021-07-14T10:37:56.973990Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Combined analysis","metadata":{}},{"cell_type":"code","source":"train_image_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:38:18.723373Z","iopub.execute_input":"2021-07-14T10:38:18.723978Z","iopub.status.idle":"2021-07-14T10:38:18.763925Z","shell.execute_reply.started":"2021-07-14T10:38:18.723888Z","shell.execute_reply":"2021-07-14T10:38:18.762767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_label(row):\n    for c in train_study_df.columns:\n        if row[c] == 1:\n            return str.lower(c.split(\" \")[0])\n\ntrain_study_df['study_label'] = train_study_df.apply(get_label, axis=1)","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:38:28.343213Z","iopub.execute_input":"2021-07-14T10:38:28.343616Z","iopub.status.idle":"2021-07-14T10:38:28.519592Z","shell.execute_reply.started":"2021-07-14T10:38:28.343584Z","shell.execute_reply":"2021-07-14T10:38:28.518419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_study_df['StudyInstanceUID'] = train_study_df['id'].apply(lambda x: x.split('_')[0])","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:38:30.507658Z","iopub.execute_input":"2021-07-14T10:38:30.508104Z","iopub.status.idle":"2021-07-14T10:38:30.519687Z","shell.execute_reply.started":"2021-07-14T10:38:30.508055Z","shell.execute_reply":"2021-07-14T10:38:30.518346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_study_df = train_study_df.rename(columns={'id': 'study_id'})\ntrain_image_df = train_image_df.rename(columns={'id': 'image_id'})\ntrain_study_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:38:39.895303Z","iopub.execute_input":"2021-07-14T10:38:39.895685Z","iopub.status.idle":"2021-07-14T10:38:39.917676Z","shell.execute_reply.started":"2021-07-14T10:38:39.895652Z","shell.execute_reply":"2021-07-14T10:38:39.916500Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_image_df.shape, train_study_df.shape","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:38:52.004453Z","iopub.execute_input":"2021-07-14T10:38:52.004936Z","iopub.status.idle":"2021-07-14T10:38:52.017277Z","shell.execute_reply.started":"2021-07-14T10:38:52.004894Z","shell.execute_reply":"2021-07-14T10:38:52.015105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_image_df['ImageInstanceUID'] = train_image_df['image_id'].apply(lambda x: x.split('_')[0])\ntrain_image_df.head(2)","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:39:31.301154Z","iopub.execute_input":"2021-07-14T10:39:31.301576Z","iopub.status.idle":"2021-07-14T10:39:31.324365Z","shell.execute_reply.started":"2021-07-14T10:39:31.301542Z","shell.execute_reply":"2021-07-14T10:39:31.323067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_num_opacities(labels):\n    num_opacities = 0\n    for i in labels:\n        if i[0] == 'opacity':\n            num_opacities += 1\n    return num_opacities","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:40:57.578522Z","iopub.execute_input":"2021-07-14T10:40:57.578925Z","iopub.status.idle":"2021-07-14T10:40:57.585728Z","shell.execute_reply.started":"2021-07-14T10:40:57.578890Z","shell.execute_reply":"2021-07-14T10:40:57.584308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_df = train_image_df.merge(train_study_df, on='StudyInstanceUID')\ncombined_df['num_opacities'] = combined_df.split_label.apply(get_num_opacities)\ncombined_df.to_csv('study_image_combined_df.csv', index=False)\ncombined_df.head(2)","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:43:24.346564Z","iopub.execute_input":"2021-07-14T10:43:24.346965Z","iopub.status.idle":"2021-07-14T10:43:24.530274Z","shell.execute_reply.started":"2021-07-14T10:43:24.346907Z","shell.execute_reply":"2021-07-14T10:43:24.529030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_study_df.shape, train_image_df.shape, combined_df.shape","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:43:32.433089Z","iopub.execute_input":"2021-07-14T10:43:32.433532Z","iopub.status.idle":"2021-07-14T10:43:32.443273Z","shell.execute_reply.started":"2021-07-14T10:43:32.433498Z","shell.execute_reply":"2021-07-14T10:43:32.441846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# image_level_dicom_data = pd.DataFrame.from_dicoms(combined_df.image_path.values)\n# len(image_level_dicom_data)\n                          \n# train_image_dicom_df = pd.DataFrame(image_level_dicom_data)\n# train_image_dicom_df['ImageInstanceUID'] = train_image_dicom_df.fname.apply(lambda x: Path(x).stem)\n# train_image_dicom_df.to_csv('image_dicom_data.csv', index=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# train_image_dicom_df[train_image_dicom_df['StudyDate']=='284b97d038b7']\n\n# study_image_dicom_combined_df = combined_df.merge(train_image_dicom_df, on='ImageInstanceUID')\n# study_image_dicom_combined_df.drop('StudyInstanceUID_x', axis=1)\n# study_image_dicom_combined_df = study_image_dicom_combined_df.rename(columns={'StudyInstanceUID_x': 'StudyInstanceUID'})\n# study_image_dicom_combined_df['num_opacities'] = study_image_dicom_combined_df.split_label.apply(get_num_opacities)\n# study_image_dicom_combined_df.to_csv('study_image_dicom_combined_data.csv', index=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Image level analysis","metadata":{}},{"cell_type":"code","source":"# combined_df = pd.read_csv('../input/siimcovidcsvfiles/study_image_dicom_combined_data.csv')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_df.head(2)","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:44:03.522930Z","iopub.execute_input":"2021-07-14T10:44:03.523403Z","iopub.status.idle":"2021-07-14T10:44:03.544701Z","shell.execute_reply.started":"2021-07-14T10:44:03.523357Z","shell.execute_reply":"2021-07-14T10:44:03.543543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_df['study_label'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:44:10.458108Z","iopub.execute_input":"2021-07-14T10:44:10.458482Z","iopub.status.idle":"2021-07-14T10:44:10.474335Z","shell.execute_reply.started":"2021-07-14T10:44:10.458449Z","shell.execute_reply":"2021-07-14T10:44:10.473233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_df['study_label'].count()","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:44:13.371869Z","iopub.execute_input":"2021-07-14T10:44:13.372341Z","iopub.status.idle":"2021-07-14T10:44:13.384735Z","shell.execute_reply.started":"2021-07-14T10:44:13.372292Z","shell.execute_reply":"2021-07-14T10:44:13.383056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_df['ImageInstanceUID'].unique().shape, combined_df['StudyInstanceUID'].unique().shape ","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:44:18.187331Z","iopub.execute_input":"2021-07-14T10:44:18.187739Z","iopub.status.idle":"2021-07-14T10:44:18.201029Z","shell.execute_reply.started":"2021-07-14T10:44:18.187705Z","shell.execute_reply":"2021-07-14T10:44:18.199305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(combined_df[(combined_df['num_opacities']==0) &\n                              (combined_df['study_label']=='negative')]['study_label'].value_counts(), \n\n combined_df[(combined_df['num_opacities']==0) &\n                              (combined_df['study_label']=='negative')]['StudyInstanceUID'].unique().shape)","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:44:26.201707Z","iopub.execute_input":"2021-07-14T10:44:26.202146Z","iopub.status.idle":"2021-07-14T10:44:26.224763Z","shell.execute_reply.started":"2021-07-14T10:44:26.202113Z","shell.execute_reply":"2021-07-14T10:44:26.223354Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As expected, all the negative labelled images have no bounding boxes associated with them. On top of that, we have 60 duplicate images 1736 - 1676 = 60\n","metadata":{}},{"cell_type":"markdown","source":"# Bad and duplicate data\nLet's look at the images which have no bounding box but are labelled as not negative \n","metadata":{}},{"cell_type":"code","source":"combined_df['StudyInstanceUID'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:44:37.428378Z","iopub.execute_input":"2021-07-14T10:44:37.428821Z","iopub.status.idle":"2021-07-14T10:44:37.445801Z","shell.execute_reply.started":"2021-07-14T10:44:37.428788Z","shell.execute_reply":"2021-07-14T10:44:37.444286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"studies_with_more_than_1_image = \\\ncombined_df[combined_df['StudyInstanceUID'].isin([s for s, i in combined_df['StudyInstanceUID'].value_counts().items() if i>1])]","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:44:46.617802Z","iopub.execute_input":"2021-07-14T10:44:46.618690Z","iopub.status.idle":"2021-07-14T10:44:46.643849Z","shell.execute_reply.started":"2021-07-14T10:44:46.618642Z","shell.execute_reply":"2021-07-14T10:44:46.642597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(studies_with_more_than_1_image['StudyInstanceUID'].unique().shape, \nstudies_with_more_than_1_image.study_label.value_counts(), \nstudies_with_more_than_1_image.shape)","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:45:03.339471Z","iopub.execute_input":"2021-07-14T10:45:03.339840Z","iopub.status.idle":"2021-07-14T10:45:03.350725Z","shell.execute_reply.started":"2021-07-14T10:45:03.339806Z","shell.execute_reply":"2021-07-14T10:45:03.349504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 232 unique studies with more than 1 image, and we can also see the division of duplicate images accross these studies.\n\nWe are dealing with 2 types of bad data\n- Duplicate images in studies which don't add much value\n- Images which are labelled as non-negative but don't have proper bounding boxes\n- Images which are both, i.e are duplicated in a study and have no bounding boxes s","metadata":{"execution":{"iopub.status.busy":"2021-06-20T11:42:24.020758Z","iopub.execute_input":"2021-06-20T11:42:24.021116Z","iopub.status.idle":"2021-06-20T11:42:24.027073Z","shell.execute_reply.started":"2021-06-20T11:42:24.021079Z","shell.execute_reply":"2021-06-20T11:42:24.025651Z"}}},{"cell_type":"code","source":"combined_df[(combined_df['num_opacities']==0) &\n                              (combined_df['study_label']!='negative')]['study_label'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:46:20.310600Z","iopub.execute_input":"2021-07-14T10:46:20.311071Z","iopub.status.idle":"2021-07-14T10:46:20.328305Z","shell.execute_reply.started":"2021-07-14T10:46:20.311035Z","shell.execute_reply":"2021-07-14T10:46:20.326890Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"non_negative_no_bb_images = combined_df[(combined_df['num_opacities']==0) &\n                              (combined_df['study_label']!='negative')]\nnon_negative_no_bb_images.shape","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:46:36.940761Z","iopub.execute_input":"2021-07-14T10:46:36.941296Z","iopub.status.idle":"2021-07-14T10:46:36.966760Z","shell.execute_reply.started":"2021-07-14T10:46:36.941250Z","shell.execute_reply":"2021-07-14T10:46:36.965316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Okay, so now we have**\n- non_negative_no_bb_images: images which are part of studies which are non negative. But these images have no bounding boxes in them which is odd\n- combined_df: study and image combined data where we map all the study level labels to all the corresponding images","metadata":{}},{"cell_type":"code","source":"# separate the images which have proper labels and have co-ordinates if applicable vs the images under suspicion\ncombined_df_new = combined_df.drop(non_negative_no_bb_images.index)","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:47:16.696636Z","iopub.execute_input":"2021-07-14T10:47:16.697068Z","iopub.status.idle":"2021-07-14T10:47:16.705591Z","shell.execute_reply.started":"2021-07-14T10:47:16.697036Z","shell.execute_reply":"2021-07-14T10:47:16.704215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_df.shape, combined_df_new.shape, non_negative_no_bb_images.shape","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:47:23.177768Z","iopub.execute_input":"2021-07-14T10:47:23.178232Z","iopub.status.idle":"2021-07-14T10:47:23.185916Z","shell.execute_reply.started":"2021-07-14T10:47:23.178200Z","shell.execute_reply":"2021-07-14T10:47:23.184657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are studies with multiple images which have no bounding boxes and out of these studies some are captured in our new dataframe.","metadata":{}},{"cell_type":"code","source":"non_negative_no_bb_images['StudyInstanceUID'].value_counts()[:25], non_negative_no_bb_images['StudyInstanceUID'].unique().shape, non_negative_no_bb_images.shape","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:47:47.532329Z","iopub.execute_input":"2021-07-14T10:47:47.532819Z","iopub.status.idle":"2021-07-14T10:47:47.544432Z","shell.execute_reply.started":"2021-07-14T10:47:47.532785Z","shell.execute_reply":"2021-07-14T10:47:47.543098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 304 images which are a part of 261 studies which don't have bounding boxes","metadata":{}},{"cell_type":"code","source":"# len((set(non_negative_no_bb_images['StudyInstanceUID']).union(set(studies_with_more_than_1_image['StudyInstanceUID'].unique()))).difference(\n# (set(non_negative_no_bb_images['StudyInstanceUID']).intersection(set(studies_with_more_than_1_image['StudyInstanceUID'].unique())))))\n\n\nprint(\"Number of unique non-negative studies with no BB: {}\".format(\n    len(set(non_negative_no_bb_images['StudyInstanceUID']).intersection(\n        set(studies_with_more_than_1_image['StudyInstanceUID'].unique())))))\n\nprint(\"Total unique studies with more than 1 image: {}\".format(\n    studies_with_more_than_1_image['StudyInstanceUID'].unique().shape))\n\nlen(\"Total negative studies with duplicate images: {}\".format(\n    (set(studies_with_more_than_1_image['StudyInstanceUID'].unique())).difference(\n    (set(non_negative_no_bb_images['StudyInstanceUID'].unique())).intersection(\n        set(studies_with_more_than_1_image['StudyInstanceUID'].unique())))))\n","metadata":{"execution":{"iopub.status.busy":"2021-07-14T10:51:56.679915Z","iopub.execute_input":"2021-07-14T10:51:56.680368Z","iopub.status.idle":"2021-07-14T10:51:56.695513Z","shell.execute_reply.started":"2021-07-14T10:51:56.680335Z","shell.execute_reply":"2021-07-14T10:51:56.694062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Analyzing negative labelled studies with duplicate images\nWe will need to separate duplicate images with negative labels","metadata":{}},{"cell_type":"code","source":"multi_image_negative_study_df = combined_df[combined_df['StudyInstanceUID'].isin((set(studies_with_more_than_1_image['StudyInstanceUID'].unique())).difference(\n    (set(non_negative_no_bb_images['StudyInstanceUID']).intersection(set(studies_with_more_than_1_image['StudyInstanceUID'].unique())))))]\n\nmulti_image_negative_study_df.study_label.value_counts(), multi_image_negative_study_df.StudyInstanceUID.unique().shape","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:01:24.889788Z","iopub.execute_input":"2021-07-14T11:01:24.890255Z","iopub.status.idle":"2021-07-14T11:01:24.906347Z","shell.execute_reply.started":"2021-07-14T11:01:24.890220Z","shell.execute_reply":"2021-07-14T11:01:24.904802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The previous 60 duplicate images can be found here. We'll need to manually select the images we want to use for these studies","metadata":{}},{"cell_type":"code","source":"bad_images = []\nfor study_id in multi_image_negative_study_df.StudyInstanceUID.unique():\n    print(study_id)\n    rows = multi_image_negative_study_df[multi_image_negative_study_df['StudyInstanceUID']==study_id]\n    print(rows.ImageInstanceUID)\n    try:\n        rows = multi_image_negative_study_df[multi_image_negative_study_df['StudyInstanceUID']==study_id]\n        dcmimg = [pydicom.dcmread(i) for i in rows.image_path]\n        row_cols = [(item.Rows, item.Columns) for item in dcmimg]\n        imgs = [dicom2array(path) for path in rows.image_path]\n        print(row_cols)\n        img_avrages = [im.mean() for im in imgs]\n        print(img_avrages)\n        plot_imgs(imgs)\n        \n    except Exception as e:\n        bad_images.append((study_id, list(rows.ImageInstanceUID), list(rows.image_path)))\n        print(e)\n        print(f\"check {study_id} manually\")","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:08:15.982906Z","iopub.execute_input":"2021-07-14T11:08:15.983322Z","iopub.status.idle":"2021-07-14T11:09:04.802028Z","shell.execute_reply.started":"2021-07-14T11:08:15.983289Z","shell.execute_reply":"2021-07-14T11:09:04.800996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 113 negative labeled images which are repeated under 53 studies. We will need to manually select the image that we want to use in each of the studies","metadata":{}},{"cell_type":"code","source":"# need help figuring out which images to remove and how to go about this\nduplicate_images_negative_label_images = ['0d4d6acc9ed3','93a881fb3292','cdd9e3aaf45a','68ad4b624a6d','cbf0a27f993e','b61f3493c551','0b020a7aff0a','3e7b2ffc97db','ace7a9702770','e96133d06736','9108cdfd43dc','e897ef5c203c','d180fed57716','f208dc529d16','ea2688741043','21518ca15050','bdd3115879aa','a1fa5f79671d','59bc532be971','b0866caa201a','ea516e218fe6','93301812b0e7','a2ee4b862182','d9456aadecbe']\nlen(duplicate_images_negative_label_images)","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:18:33.066291Z","iopub.execute_input":"2021-07-14T11:18:33.066720Z","iopub.status.idle":"2021-07-14T11:18:33.076789Z","shell.execute_reply.started":"2021-07-14T11:18:33.066687Z","shell.execute_reply":"2021-07-14T11:18:33.075364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(combined_df_new.shape)\ncombined_df_new = combined_df_new.drop(\n    combined_df_new[combined_df_new.ImageInstanceUID.isin(duplicate_images_negative_label_images)].index)\nprint(combined_df_new.shape)","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:20:53.923213Z","iopub.execute_input":"2021-07-14T11:20:53.923691Z","iopub.status.idle":"2021-07-14T11:20:53.944872Z","shell.execute_reply.started":"2021-07-14T11:20:53.923625Z","shell.execute_reply":"2021-07-14T11:20:53.943291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Understandig images with no opacities but non negative label","metadata":{}},{"cell_type":"code","source":"non_negative_no_bb_images.StudyInstanceUID.unique().shape, non_negative_no_bb_images.shape","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:21:39.254381Z","iopub.execute_input":"2021-07-14T11:21:39.254784Z","iopub.status.idle":"2021-07-14T11:21:39.264290Z","shell.execute_reply.started":"2021-07-14T11:21:39.254751Z","shell.execute_reply":"2021-07-14T11:21:39.262564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"common_studies = non_negative_no_bb_images[non_negative_no_bb_images['StudyInstanceUID'].isin(combined_df_new['StudyInstanceUID'])]\n\nprint(\"Images with no BB which are a part of non-negative study which has more than 1 image: {}\".format(common_studies.shape))\nprint(\"NonNegative studie with more than 1 image: {}\".format(common_studies['StudyInstanceUID'].unique().shape))\nprint(\"TotalUnique images here: {}\".format(common_studies['ImageInstanceUID'].unique().shape))","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:26:54.089104Z","iopub.execute_input":"2021-07-14T11:26:54.089524Z","iopub.status.idle":"2021-07-14T11:26:54.100709Z","shell.execute_reply.started":"2021-07-14T11:26:54.089481Z","shell.execute_reply":"2021-07-14T11:26:54.099137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"common_studies.study_label.value_counts(), common_studies.StudyInstanceUID.unique().shape","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:22:40.385053Z","iopub.execute_input":"2021-07-14T11:22:40.385472Z","iopub.status.idle":"2021-07-14T11:22:40.396500Z","shell.execute_reply.started":"2021-07-14T11:22:40.385438Z","shell.execute_reply":"2021-07-14T11:22:40.394787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here, we get common studies. These are studies which are labelled with a `non-negative` tag and have images in which some images have opacities marked and some don't.\n\nWe can see that out of the 261 studies with images with no bounding boxes, 177 of those studies have atleast 1 image with a bounding box and we've captured that image in our dataframe `combined_df_new` along with the study details.","metadata":{}},{"cell_type":"code","source":"# in the original combined dataframe, we can see that there are 177 studies with more than 1 image. We capture the images for these 177 studies\n# which have atleast 1 bounding box. We can ignore the rest of the images for these studies since a study will have 1 label and all the images under\n# that study should have the same label\ncombined_df_new[combined_df_new['StudyInstanceUID'].isin(common_studies['StudyInstanceUID'])]['StudyInstanceUID'].shape","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:23:18.430716Z","iopub.execute_input":"2021-07-14T11:23:18.431187Z","iopub.status.idle":"2021-07-14T11:23:18.441361Z","shell.execute_reply.started":"2021-07-14T11:23:18.431153Z","shell.execute_reply":"2021-07-14T11:23:18.439827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that the 177 unique studies that we have in common (which have some images with bounding boxes and some with no bounding boxes) are covered in the dataframe where we separated out the bad data.","metadata":{}},{"cell_type":"markdown","source":"Now, we've separated the images with no bounding boxes and which are non negative with the images which have bounding boxes \n\nSo, we have 177 studies which have images with an opacity and images without an opacity. These 177 have multiple images attached to them and we've filtered out the images with no opacity out from the DF. Let's check how many images the remaining studies have in the DF","metadata":{}},{"cell_type":"code","source":"uncommon_studies = non_negative_no_bb_images[~non_negative_no_bb_images['StudyInstanceUID'].isin(common_studies['StudyInstanceUID'])]\nuncommon_studies.StudyInstanceUID.unique().shape, uncommon_studies.shape","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:29:36.020225Z","iopub.execute_input":"2021-07-14T11:29:36.020599Z","iopub.status.idle":"2021-07-14T11:29:36.030152Z","shell.execute_reply.started":"2021-07-14T11:29:36.020567Z","shell.execute_reply":"2021-07-14T11:29:36.029027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"uncommon_studies.study_label.value_counts()","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:29:40.640695Z","iopub.execute_input":"2021-07-14T11:29:40.641125Z","iopub.status.idle":"2021-07-14T11:29:40.651306Z","shell.execute_reply.started":"2021-07-14T11:29:40.641092Z","shell.execute_reply":"2021-07-14T11:29:40.649446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"uncommon_studies[uncommon_studies['study_label']=='typical']","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:37:04.730990Z","iopub.execute_input":"2021-07-14T11:37:04.731387Z","iopub.status.idle":"2021-07-14T11:37:04.747010Z","shell.execute_reply.started":"2021-07-14T11:37:04.731354Z","shell.execute_reply":"2021-07-14T11:37:04.745635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"uncommon_studies = uncommon_studies.drop(1793)","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:37:02.233228Z","iopub.execute_input":"2021-07-14T11:37:02.233642Z","iopub.status.idle":"2021-07-14T11:37:02.240904Z","shell.execute_reply.started":"2021-07-14T11:37:02.233610Z","shell.execute_reply":"2021-07-14T11:37:02.239547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 86 atypical images wth no opacity defined and 1 typical image with no opacity defined. the latter seems like an anomaly and we can probably remove that and use the other 86 atypical ones. Let's also make sure that these are not already accounted for","metadata":{}},{"cell_type":"code","source":"set(uncommon_studies.StudyInstanceUID).intersection(set(combined_df_new.StudyInstanceUID))","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:37:08.938316Z","iopub.execute_input":"2021-07-14T11:37:08.938740Z","iopub.status.idle":"2021-07-14T11:37:08.947593Z","shell.execute_reply.started":"2021-07-14T11:37:08.938707Z","shell.execute_reply":"2021-07-14T11:37:08.946160Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"uncommon_studies.StudyInstanceUID.value_counts()[:5]","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:37:11.128561Z","iopub.execute_input":"2021-07-14T11:37:11.129042Z","iopub.status.idle":"2021-07-14T11:37:11.139091Z","shell.execute_reply.started":"2021-07-14T11:37:11.129006Z","shell.execute_reply":"2021-07-14T11:37:11.137632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bad_images = []\nfor study_id in [\"0d9709b3af74\", \"784afbafee30\"]:\n    print(study_id)\n    rows = uncommon_studies[uncommon_studies['StudyInstanceUID']==study_id]\n    print(rows.ImageInstanceUID)\n    try:\n        rows = uncommon_studies[uncommon_studies['StudyInstanceUID']==study_id]\n        dcmimg = [pydicom.dcmread(i) for i in rows.image_path]\n        row_cols = [(item.Rows, item.Columns) for item in dcmimg]\n        imgs = [dicom2array(path) for path in rows.image_path]\n        print(row_cols)\n        print([im.mean() for im in imgs])\n        plot_imgs(imgs)\n    except Exception as e:\n        bad_images.append((study_id, list(rows.ImageInstanceUID), list(rows.image_path)))\n        print(e)\n        print(f\"check {study_id} manually\")","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:31:00.680761Z","iopub.execute_input":"2021-07-14T11:31:00.681209Z","iopub.status.idle":"2021-07-14T11:31:04.476959Z","shell.execute_reply.started":"2021-07-14T11:31:00.681175Z","shell.execute_reply":"2021-07-14T11:31:04.475822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"uncommon_studies = uncommon_studies.drop(\n    uncommon_studies[uncommon_studies['ImageInstanceUID'].isin([\"efc93a3917b6\", \"830063223a31\"])].index)","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:31:27.391699Z","iopub.execute_input":"2021-07-14T11:31:27.392142Z","iopub.status.idle":"2021-07-14T11:31:27.399132Z","shell.execute_reply.started":"2021-07-14T11:31:27.392109Z","shell.execute_reply":"2021-07-14T11:31:27.397851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"uncommon_studies.shape","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:37:25.801864Z","iopub.execute_input":"2021-07-14T11:37:25.802296Z","iopub.status.idle":"2021-07-14T11:37:25.811280Z","shell.execute_reply.started":"2021-07-14T11:37:25.802263Z","shell.execute_reply":"2021-07-14T11:37:25.809706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_df_new.shape","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:37:31.440276Z","iopub.execute_input":"2021-07-14T11:37:31.440678Z","iopub.status.idle":"2021-07-14T11:37:31.449256Z","shell.execute_reply.started":"2021-07-14T11:37:31.440645Z","shell.execute_reply":"2021-07-14T11:37:31.447969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_df = pd.concat([combined_df_new, uncommon_studies])","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:38:32.889971Z","iopub.execute_input":"2021-07-14T11:38:32.890394Z","iopub.status.idle":"2021-07-14T11:38:32.901826Z","shell.execute_reply.started":"2021-07-14T11:38:32.890359Z","shell.execute_reply":"2021-07-14T11:38:32.900687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_df.shape, train_image_df.shape, studies_with_more_than_1_image.shape","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:39:36.501317Z","iopub.execute_input":"2021-07-14T11:39:36.501718Z","iopub.status.idle":"2021-07-14T11:39:36.512265Z","shell.execute_reply.started":"2021-07-14T11:39:36.501684Z","shell.execute_reply":"2021-07-14T11:39:36.510643Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_df.to_csv('cleaned_data-v1.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:40:15.383920Z","iopub.execute_input":"2021-07-14T11:40:15.384305Z","iopub.status.idle":"2021-07-14T11:40:15.524668Z","shell.execute_reply.started":"2021-07-14T11:40:15.384274Z","shell.execute_reply":"2021-07-14T11:40:15.523568Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, to get all the images that are useful, we need to select from the above data frames.","metadata":{}},{"cell_type":"code","source":"# let's see the remaining duplicate images available to us\n\nbad_images = []\nfor study_id, cont in final_df.StudyInstanceUID.value_counts()[:30].items():\n    print(study_id)\n    rows = final_df[final_df['StudyInstanceUID']==study_id]\n    print(rows.ImageInstanceUID)\n    try:\n        rows = final_df[final_df['StudyInstanceUID']==study_id]\n        imgs = [dicom2array(path) for path in rows.image_path]\n        plot_imgs(imgs)\n        dcmimg = [pydicom.dcmread(i) for i in rows.image_path]\n        row_cols = [(item.Rows, item.Columns) for item in dcmimg]\n        print(row_cols)\n    except Exception as e:\n        bad_images.append((study_id, list(rows.ImageInstanceUID), list(rows.image_path)))\n        print(e)\n        print(f\"check {study_id} manually\")","metadata":{"execution":{"iopub.status.busy":"2021-07-14T11:41:53.583345Z","iopub.execute_input":"2021-07-14T11:41:53.583752Z","iopub.status.idle":"2021-07-14T11:42:22.589605Z","shell.execute_reply.started":"2021-07-14T11:41:53.583712Z","shell.execute_reply":"2021-07-14T11:42:22.588036Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most of thse images are different enough for us to consider them.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}