{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# SIIM COVID-19: Process to PNG/JPG 640/512 with bboxes\nData source: **SIIM-FISABIO-RSNA COVID-19 Detection** competition: https://www.kaggle.com/competitions/siim-covid19-detection\n\nImage processing includes:\n- reading DICOM, correcting BW and VOI_LUT\n- removing empty borders, \n- square crop at the center attempting to keep boxes uncropped, \n- resizing down to 512x512 pixels size, \n- histogram equalization, \n- transforming to 8-bits,\n- saving images to PNG or JPG using a good quality=95\n\nAnnotation processing: \n- recalculating bboxes to match new crop and size, \n- add human-readable labels from study level,\n- add some DICOM fields.\n\nAnnotation file 'test_proc.csv' is for the test images. Column 'hwtblr' contains a list of values for the: original_height, original_width, crop_top_line, crop_bottom_line, crop_left_pixel, crop_right_pixel. This data will help to re-calculate bboxes to the original image size.\n\nImage clean up process:\n- some Studies contain multiple image duplicates with only one image containing boxes: remove other duplicates with no boxes\n- some Studies labeled as positive, but contains no boxes - remove all images from these studies\n- manual review of some Studies for badly cropped images, unusual positions, obvious opacities, etc. \n- spotted some bad images\n- drop images: positive without boxes\n- drop images: negative with boxes\n- drop images: study duplicates\n\n### Upvote, thanks!","metadata":{}},{"cell_type":"markdown","source":"# Acknowlegdements\nI have started this notebook from this: https://www.kaggle.com/code/xhlulu/siim-covid-19-convert-to-jpg-256px, but later stepped away to another direction.\n\nInspiration for image processing came from: https://www.kaggle.com/code/davidbroberts/export-processed-jpg-512/notebook and https://www.kaggle.com/code/raddar/convert-dicom-to-np-array-the-correct-way","metadata":{}},{"cell_type":"markdown","source":"## TODO\n- border crop can be smarter than simply check for equal pixels, some borders contain garbage","metadata":{}},{"cell_type":"code","source":"!pip install python-gdcm -q","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:55:58.288547Z","iopub.execute_input":"2022-11-08T00:55:58.289279Z","iopub.status.idle":"2022-11-08T00:56:08.499904Z","shell.execute_reply.started":"2022-11-08T00:55:58.289175Z","shell.execute_reply":"2022-11-08T00:56:08.498613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from fastai.vision.all import *\nfrom fastai.medical.imaging import *\nimport pydicom\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\nimport pandas as pd\nimport skimage\nfrom skimage import transform, exposure, io","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:56:08.501711Z","iopub.execute_input":"2022-11-08T00:56:08.502003Z","iopub.status.idle":"2022-11-08T00:56:12.656360Z","shell.execute_reply.started":"2022-11-08T00:56:08.501970Z","shell.execute_reply":"2022-11-08T00:56:12.655328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_items = get_dicom_files('../input/siim-covid19-detection/train')\nprint(f'Got {len(train_items)} DICOM files from the TRAIN path')\ntest_items = get_dicom_files('../input/siim-covid19-detection/test')\nprint(f'Got {len(test_items)} DICOM files from the TEST path')","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:56:12.657989Z","iopub.execute_input":"2022-11-08T00:56:12.658424Z","iopub.status.idle":"2022-11-08T00:58:18.796432Z","shell.execute_reply.started":"2022-11-08T00:56:12.658378Z","shell.execute_reply":"2022-11-08T00:58:18.795227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# train_items = train_items[:12] # for test\n# test_items = train_items[:12] # for test","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:58:18.798155Z","iopub.execute_input":"2022-11-08T00:58:18.798469Z","iopub.status.idle":"2022-11-08T00:58:18.812365Z","shell.execute_reply.started":"2022-11-08T00:58:18.798437Z","shell.execute_reply":"2022-11-08T00:58:18.811202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Preprocess annotations","metadata":{}},{"cell_type":"code","source":"# read image and study annotations\nsdf = pd.read_csv('../input/siim-covid19-detection/train_study_level.csv')\nidf = pd.read_csv('../input/siim-covid19-detection/train_image_level.csv')","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:58:18.813984Z","iopub.execute_input":"2022-11-08T00:58:18.814426Z","iopub.status.idle":"2022-11-08T00:58:18.909001Z","shell.execute_reply.started":"2022-11-08T00:58:18.814361Z","shell.execute_reply":"2022-11-08T00:58:18.908105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# create human-readable classes\ndef lab_y( r ):\n    if r['Negative for Pneumonia']: return 'Neg'\n    if r['Typical Appearance']: return 'Typ'\n    if r['Atypical Appearance']: return 'Atyp'\n    if r['Indeterminate Appearance']: return 'Indet'\n\nsdf['label_y'] = sdf.apply(lab_y, axis=1)\nsdf.id = sdf.apply(lambda r: r.id.replace('_study',''), axis=1)\nsdf = sdf.set_index('id')","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:58:18.910323Z","iopub.execute_input":"2022-11-08T00:58:18.910896Z","iopub.status.idle":"2022-11-08T00:58:19.102377Z","shell.execute_reply.started":"2022-11-08T00:58:18.910860Z","shell.execute_reply":"2022-11-08T00:58:19.101551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"idf.drop(columns=['label'],inplace=True) # irrelevant column\n# add text labels from study csv\nidf['label_y'] = idf.apply(lambda r: sdf.loc[r.StudyInstanceUID].label_y, axis=1)\nidf.id = idf.apply(lambda r: r.id.replace('_image',''), axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:58:19.103439Z","iopub.execute_input":"2022-11-08T00:58:19.103726Z","iopub.status.idle":"2022-11-08T00:58:20.337921Z","shell.execute_reply.started":"2022-11-08T00:58:19.103699Z","shell.execute_reply":"2022-11-08T00:58:20.336814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# calculate a box, bounding all boxes\n# it will help to avoid/minimise crop of boxes\ndef box_all(r): # input: boxes from dataframe\n    boxes = r.boxes\n    if type(boxes) == list:\n        boxl = boxes\n        w = 'w'\n        h = 'h'\n    else: \n        if pd.isnull(boxes): return boxes\n        boxl = eval(boxes)\n        w = 'width'\n        h = 'height'\n    minx = min([b['x'] for b in boxl])\n    miny = min([b['y'] for b in boxl])\n    maxx = round(max([b['x']+b[w] for b in boxl]), 2)\n    maxy = round(max([b['y']+b[h] for b in boxl]), 2)\n    return [minx, miny, maxx, maxy]\n\nidf['bb_min_xy_max_xy'] = idf.apply(box_all, axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:58:20.341238Z","iopub.execute_input":"2022-11-08T00:58:20.341739Z","iopub.status.idle":"2022-11-08T00:58:20.596001Z","shell.execute_reply.started":"2022-11-08T00:58:20.341690Z","shell.execute_reply":"2022-11-08T00:58:20.595005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Process and convert DICOM images","metadata":{}},{"cell_type":"code","source":"# size of the output image\nDSIZE = 640 #512 \next = '.png' # '.jpg'","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:58:20.597565Z","iopub.execute_input":"2022-11-08T00:58:20.597877Z","iopub.status.idle":"2022-11-08T00:58:20.601303Z","shell.execute_reply.started":"2022-11-08T00:58:20.597846Z","shell.execute_reply":"2022-11-08T00:58:20.600534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Functions to remove borders and do maximum center crop, trying to keep boxes\ndef topbot_borders(pixels):\n    t = 0\n    b = 0\n    h = pixels.shape[0]\n    for i in range(h):\n        if not np.all(pixels[i] == pixels[i][0]):\n            t = i\n            break\n              \n    for i in range(h-1, 0, -1):\n        if not np.all(pixels[i] == pixels[i][0]):\n            b = i\n            break\n    return t, b\n\ndef remove_borders(pixels, bb_min_xy_max_xy):\n    h = pixels.shape[0]\n    w = pixels.shape[1]\n    t,b = topbot_borders(pixels)\n    pixels = pixels[t:b,:] \n    pixels = np.rot90(pixels)\n    r,l = topbot_borders(pixels)\n    pixels = pixels[r:l,:]\n    pixels = np.rot90(pixels, 3)\n    l = w-l\n    r = w-r\n    # crop square shape\n    nh = pixels.shape[0]\n    nw = pixels.shape[1]\n    d = abs(nh - nw)\n    st = d // 2\n    sp = d-st\n    if nh == nw: pass\n    elif nh > nw:\n        # check to include boxes\n        if bb_min_xy_max_xy is not None:\n            if bb_min_xy_max_xy[0] < st:\n                st = bb_min_xy_max_xy[0]\n                sp = d-st\n            elif bb_min_xy_max_xy[2] > nw-sp:\n                sp = nw - bb_min_xy_max_xy[2]\n                st = d - sp\n        pixels = pixels[st:-sp,:]\n        t += st\n        b -= sp\n    elif nw > nh:\n        # check to include boxes\n        if bb_min_xy_max_xy is not None:\n            if bb_min_xy_max_xy[1] < st:\n                st = bb_min_xy_max_xy[1]\n                sp = d-st\n            elif bb_min_xy_max_xy[3] > nh-sp:\n                sp = nh - bb_min_xy_max_xy[3]\n                st = d - sp\n        pixels = pixels[:,st:-sp]\n        l += st\n        r -= sp\n    return pixels, [h,w,t,b,l,r]","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:58:20.602636Z","iopub.execute_input":"2022-11-08T00:58:20.603229Z","iopub.status.idle":"2022-11-08T00:58:20.617284Z","shell.execute_reply.started":"2022-11-08T00:58:20.603194Z","shell.execute_reply":"2022-11-08T00:58:20.616066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# process DICOM files\ndef proc_dcm(items, folder, bdf, voi_lut = True):\n\n    proc_df = pd.DataFrame(columns =['id','hwtblr', 'Modality', 'ImagerPixelSpacing', 'StudyDate', 'PatientSex', 'BodyPartExamined'])\n\n    for ti in items:\n        id = ti.stem\n        img = ti.dcmread()\n        if voi_lut:\n            pixels = apply_voi_lut(img.pixel_array, img)\n        else:\n            pixels = img.pixel_array\n        #pixels = img.pixel_array\n        bb_mimax = None\n        if bdf is not None:\n            if id in bdf.index:\n                bb_mimax = bdf.loc[id].bb_min_xy_max_xy\n        pixels,hwtblr = remove_borders(pixels, bb_mimax)\n        if img.PhotometricInterpretation == \"MONOCHROME1\":\n            pixels = np.amax(pixels) - pixels\n        pixels = skimage.transform.resize(pixels, (DSIZE, DSIZE))\n        pixels = skimage.exposure.equalize_hist(pixels,nbins=256)\n        #pixels = skimage.filters.unsharp_mask(pixels, radius=5, amount=1)\n        pixels = (pixels * 255).astype(np.uint8)\n        # fig, ax = plt.subplots(figsize=(12,12))\n        # im = ax.imshow(pixels, cmap='gray')\n        # plt.show()\n\n        #skimage.io.imsave(folder + '/' + id + ext, pixels, quality=95)\n        skimage.io.imsave(folder + '/' + id + ext, pixels)\n        proc_df = proc_df.append({'id': id, 'hwtblr': hwtblr, 'Modality': img.Modality,\n            'ImagerPixelSpacing': img.ImagerPixelSpacing, 'StudyDate': img.StudyDate, \n            'PatientSex': img.PatientSex, 'BodyPartExamined': img.BodyPartExamined}, ignore_index=True)\n    return proc_df","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:58:20.618590Z","iopub.execute_input":"2022-11-08T00:58:20.618935Z","iopub.status.idle":"2022-11-08T00:58:20.635554Z","shell.execute_reply.started":"2022-11-08T00:58:20.618887Z","shell.execute_reply":"2022-11-08T00:58:20.634205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Processing train images')\n!mkdir train\ntrain_proc_df = proc_dcm(train_items, 'train', idf )","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:58:20.636849Z","iopub.execute_input":"2022-11-08T00:58:20.637217Z","iopub.status.idle":"2022-11-08T00:58:30.697627Z","shell.execute_reply.started":"2022-11-08T00:58:20.637183Z","shell.execute_reply":"2022-11-08T00:58:30.696065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Processing test images')\n!mkdir test\ntest_proc_df = proc_dcm(test_items, 'test', None )","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:58:30.700310Z","iopub.execute_input":"2022-11-08T00:58:30.701315Z","iopub.status.idle":"2022-11-08T00:58:37.400226Z","shell.execute_reply.started":"2022-11-08T00:58:30.701241Z","shell.execute_reply":"2022-11-08T00:58:37.398510Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Process annotation files\nAnnotation processing: recalculating bboxes to match new crop and size, add human-readable labels from study level.","metadata":{}},{"cell_type":"code","source":"# save and index train image crop data\ntrain_proc_df.to_csv('train_proc.csv',index=False) # don't need to save this file, for debug purpose\ntrain_proc_df = train_proc_df.set_index('id')\ntrain_proc_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:58:37.402717Z","iopub.execute_input":"2022-11-08T00:58:37.403267Z","iopub.status.idle":"2022-11-08T00:58:37.439219Z","shell.execute_reply.started":"2022-11-08T00:58:37.403208Z","shell.execute_reply":"2022-11-08T00:58:37.438223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# save dataframe with test image crop data\ntest_proc_df.to_csv('test_proc.csv',index=False) \ntest_proc_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:58:37.440444Z","iopub.execute_input":"2022-11-08T00:58:37.440727Z","iopub.status.idle":"2022-11-08T00:58:37.460902Z","shell.execute_reply.started":"2022-11-08T00:58:37.440699Z","shell.execute_reply":"2022-11-08T00:58:37.459998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Recalculate boxes for the new crop and new size","metadata":{}},{"cell_type":"code","source":"def calc_boxes(o, bdf):\n    # o is row of the train annotation file\n    # bdf is the dataframe of crop results\n    if pd.isnull(o.boxes): return\n    if o.id not in bdf.index: return o.boxes\n    boxl = eval(o.boxes)\n    hwtblr = bdf.loc[o.id].hwtblr\n    new_boxl = []\n    for bb in boxl:\n        cw = hwtblr[5] - hwtblr[4] # width after crop\n        ch = hwtblr[3] - hwtblr[2] # height after crop\n        \n        bbx = bb['x'] - hwtblr[4] # left crop\n        if bbx < 0: bbx = 0\n        bbx = round((bbx * DSIZE) / cw, 2)\n            \n        bby = bb['y'] - hwtblr[2] # top crop\n        if bby < 0: bby = 0\n        bby = round((bby * DSIZE) / ch, 2)\n            \n        bbw = round((bb['width'] * DSIZE) / cw, 2)\n        if bbx+bbw > DSIZE: bbw = DSIZE-bbx\n        \n        bbh = round((bb['height'] * DSIZE) / ch, 2)\n        if bby+bbh > DSIZE: bbh = DSIZE-bby\n        \n        bbd = {'x':bbx, 'y':bby, 'w':bbw, 'h':bbh}\n        new_boxl.append(bbd)\n    #print(o.id,new_boxl)\n    return new_boxl ","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:58:37.462036Z","iopub.execute_input":"2022-11-08T00:58:37.462304Z","iopub.status.idle":"2022-11-08T00:58:37.471420Z","shell.execute_reply.started":"2022-11-08T00:58:37.462277Z","shell.execute_reply":"2022-11-08T00:58:37.470569Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# replace boxes with updated values\nidf.boxes = idf.apply(lambda r: calc_boxes(r, train_proc_df), axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:58:37.472599Z","iopub.execute_input":"2022-11-08T00:58:37.472916Z","iopub.status.idle":"2022-11-08T00:58:37.652730Z","shell.execute_reply.started":"2022-11-08T00:58:37.472886Z","shell.execute_reply":"2022-11-08T00:58:37.651810Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Output train annotation file","metadata":{}},{"cell_type":"code","source":"idf = idf.set_index('id')\nidf = idf.join(train_proc_df)\n\n# update boxes outline with new box values\nidf['bb_min_xy_max_xy'] = idf.apply(box_all, axis=1)\n\n# save processed CSV for future use\nidf.to_csv('train_512.csv', index=True)\nidf.head()","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:58:37.653891Z","iopub.execute_input":"2022-11-08T00:58:37.654174Z","iopub.status.idle":"2022-11-08T00:58:38.133834Z","shell.execute_reply.started":"2022-11-08T00:58:37.654146Z","shell.execute_reply":"2022-11-08T00:58:38.132698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Visualization","metadata":{}},{"cell_type":"code","source":"# read and prepare original train annotation file\ntr_df = pd.read_csv('../input/siim-covid19-detection/train_image_level.csv')\ntr_df.id = tr_df.apply(lambda r: r.id.replace('_image',''), axis=1)\ntr_df = tr_df.set_index('id')","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:58:38.187429Z","iopub.execute_input":"2022-11-08T00:58:38.188014Z","iopub.status.idle":"2022-11-08T00:58:38.306726Z","shell.execute_reply.started":"2022-11-08T00:58:38.187969Z","shell.execute_reply":"2022-11-08T00:58:38.305753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# read processed train annotation file with bboxes for the images with size 512x512\njdf = pd.read_csv('train_512.csv')\njdf = jdf.set_index('id')","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:58:55.451552Z","iopub.execute_input":"2022-11-08T00:58:55.451920Z","iopub.status.idle":"2022-11-08T00:58:55.481926Z","shell.execute_reply.started":"2022-11-08T00:58:55.451890Z","shell.execute_reply":"2022-11-08T00:58:55.481186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Labels distribution before cleaning","metadata":{}},{"cell_type":"code","source":"def show_dist(df):\n    labels = list(df.label_y.unique())\n    dist = [sum(df.label_y == lab) for lab in labels]\n    distd = dict(zip(labels,dist))\n    explode = [0.05 for _ in range(len(labels))]\n    plt.pie(dist, labels = labels, startangle = 270, explode=explode, autopct='%1.1f%%')\n    plt.legend()\n    plt.show()\n    return distd\n    \nidist = show_dist(jdf)\nprint(idist)","metadata":{"execution":{"iopub.status.busy":"2022-11-08T01:32:47.719465Z","iopub.execute_input":"2022-11-08T01:32:47.719893Z","iopub.status.idle":"2022-11-08T01:32:47.896306Z","shell.execute_reply.started":"2022-11-08T01:32:47.719857Z","shell.execute_reply":"2022-11-08T01:32:47.895294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def find_dcm_path(id, dicom_files):\n    for it in dicom_files:\n        if id == it.stem: return it","metadata":{"execution":{"iopub.status.busy":"2022-11-08T00:58:58.072655Z","iopub.execute_input":"2022-11-08T00:58:58.073409Z","iopub.status.idle":"2022-11-08T00:58:58.078483Z","shell.execute_reply.started":"2022-11-08T00:58:58.073340Z","shell.execute_reply":"2022-11-08T00:58:58.077645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# this function will plot DICOM on the left, JPG/PNG on the right\nimport matplotlib.image as mpimg\nimport matplotlib.pyplot as plt\nimport matplotlib.image as mpimg\nfrom matplotlib.patches import Rectangle\nfrom matplotlib import patches\n\njpg_folder = './train'\n\ndef plot_test( vids, figsize ):\n    ni = len(vids)\n    if ni < 2:\n        print('Input list containing more than 1 ID')\n        return\n    fig, ax = plt.subplots(ni,2, figsize=( figsize, (figsize*ni)//2 ))\n #   fig, ax = plt.subplots(ni,2, figsize=(figsize,figsize))\n    fig.suptitle('DICOM on the left, JPG/PNG on the right', fontsize=16)\n    \n    fig.tight_layout()\n    #plt.subplots_adjust(hspace=0, wspace=0)\n    \n    for vin in range(ni): \n        iid = vids[vin]\n        # original DICOM on the left\n        img = find_dcm_path(iid, train_items).dcmread()\n        pixels = img.pixel_array\n        if img.PhotometricInterpretation == \"MONOCHROME1\":\n            pixels = np.amax(pixels) - pixels\n        ax[vin,0].imshow(pixels, cmap='gray')\n        ax[vin,0].set_title(iid + '  Study:' + tr_df.loc[iid].StudyInstanceUID)\n        if pd.notnull(tr_df.loc[iid].boxes):\n            bxs = eval(tr_df.loc[iid].boxes)\n            for bx in bxs:\n                rct = patches.Rectangle( (bx['x'], bx['y']), bx['width'], bx['height'], edgecolor=\"yellow\", facecolor=\"none\", lw=2 )\n                ax[vin,0].add_patch(rct)\n\n        # processed JPG on the right\n        jmg = mpimg.imread(jpg_folder + '/' + iid + ext)\n        ax[vin,1].imshow(jmg, cmap='gray')\n        ax[vin,1].set_title(iid)\n        if pd.notnull(jdf.loc[iid].boxes):\n            bxs = eval(jdf.loc[iid].boxes)\n            #print(iid, bxs)\n            for bx in bxs:\n                rct = patches.Rectangle( (bx['x'], bx['y']), bx['w'], bx['h'], edgecolor=\"yellow\", facecolor=\"none\", lw=2 )\n                ax[vin,1].add_patch(rct)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-11-08T01:16:26.836778Z","iopub.execute_input":"2022-11-08T01:16:26.837160Z","iopub.status.idle":"2022-11-08T01:16:26.849924Z","shell.execute_reply.started":"2022-11-08T01:16:26.837129Z","shell.execute_reply":"2022-11-08T01:16:26.849064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plotting random samples - re-run  to see more\nnsamples = 3\nsample_idl = [idx for idx in tr_df.sample(nsamples).index]\n#sample_idl = ['29b23a11d1e4', '89fd7f185d77']\nplot_test(sample_idl,12)","metadata":{"execution":{"iopub.status.busy":"2022-11-08T01:16:29.675424Z","iopub.execute_input":"2022-11-08T01:16:29.676036Z","iopub.status.idle":"2022-11-08T01:16:33.699879Z","shell.execute_reply.started":"2022-11-08T01:16:29.676000Z","shell.execute_reply":"2022-11-08T01:16:33.698877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Preparing a list of bad images\nSome hints found here: https://www.kaggle.com/code/dschettler8845/covid-detection-studies-with-multiple-images-viz/notebook","metadata":{}},{"cell_type":"code","source":"# helper function\ndef get_study(study_id):\n    # returns all images in this study\n    sidx = jdf[jdf.StudyInstanceUID == study_id].index\n    simgs = [idx for idx in sidx]\n    #print(simgs)\n    return simgs","metadata":{"execution":{"iopub.status.busy":"2022-11-08T01:51:49.292616Z","iopub.execute_input":"2022-11-08T01:51:49.292973Z","iopub.status.idle":"2022-11-08T01:51:49.298341Z","shell.execute_reply.started":"2022-11-08T01:51:49.292943Z","shell.execute_reply":"2022-11-08T01:51:49.297317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# for Study with multiple images - keep only image with boxes\n# if no boxes found - mark Study for the manual check\n\nmanual_check = []\nmanual_reason = []\nids_to_remove = []\npositive_without_boxes = 0\nfor st_id in sdf.index:\n    si = get_study(st_id)\n    if len(si) > 1: \n        with_box = False\n        for iid in si:\n            if pd.notnull(jdf.loc[iid].boxes): \n                if with_box: # second image with box in the same study\n                    manual_check.append(st_id) # review this study manually\n                    manual_reason.append('multipleBoxes') # multiple images containing boxes\n                    break\n                si.remove(iid) # except the one with box,\n                ids_to_remove += si # all other mark as 'bad'\n                with_box = True\n                \n        if not with_box: \n            if sdf.loc[st_id]['Negative for Pneumonia'] == 0:\n                # this study is positive, but contains no boxes - remove all\n                ids_to_remove += si # all images in study mark as 'bad'\n                positive_without_boxes += 1\n            else:\n                manual_check.append(st_id) # review this study manually\n                manual_reason.append('noBoxes')\n                \nprint(f'Removed {positive_without_boxes} studies without boxes, but labeled positive')\nprint( f\"Duplicates to remove:{len(ids_to_remove)}; Studies to check:{len(manual_check)}; without boxes:{manual_reason.count('noBoxes')}\" )\nmanual_reason.count('noBoxes')                ","metadata":{"execution":{"iopub.status.busy":"2022-11-08T02:19:06.864375Z","iopub.execute_input":"2022-11-08T02:19:06.864790Z","iopub.status.idle":"2022-11-08T02:19:11.815538Z","shell.execute_reply.started":"2022-11-08T02:19:06.864756Z","shell.execute_reply":"2022-11-08T02:19:11.814310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Badly cropped, unusual positions, obvious opacities, etc. \n# I prefer to skip sample in doubt rather than keep confusing data.\nbad_studies = [\n    '2edd69dd0934', '350cd603e64d', '3757a5a70aa8', '3e2d6faac5dd', '3f1e5e240522', \n    '4c1e5ca11b6e', '4c45ac349e3a', '4fe89b3c7225', '7416b5cbc531', '8c93b1d9363b', \n    '9855cbfdcb71', 'b185438c915c', 'ba27bcbd6881', 'bafb7e87a1b9', 'bddc88dfb40f', \n    'bf6f81be0705', 'c36cde17cd04', 'c6af19161721', 'cd082bbb2799', 'd172b5d65314', \n    'df07c0223107', 'e126d9d23457', 'e4b2c5f23e08', 'e4b50e7402c3', 'f07ec852f8c2', \n    'f109811743c5', 'f2b77c3c70c5', 'f6ffe212deeb',    \n]\nprint(f'Bad studies after manual review:{len(bad_studies)}')","metadata":{"execution":{"iopub.status.busy":"2022-11-08T02:19:11.817262Z","iopub.execute_input":"2022-11-08T02:19:11.817571Z","iopub.status.idle":"2022-11-08T02:19:11.824734Z","shell.execute_reply.started":"2022-11-08T02:19:11.817539Z","shell.execute_reply":"2022-11-08T02:19:11.823247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# remove all images from the \"bad\" Studies\nfor st_id in bad_studies:\n    si = get_study(st_id)\n    ids_to_remove += si # all images mark as 'bad'","metadata":{"execution":{"iopub.status.busy":"2022-11-08T02:19:12.740177Z","iopub.execute_input":"2022-11-08T02:19:12.740649Z","iopub.status.idle":"2022-11-08T02:19:12.771119Z","shell.execute_reply.started":"2022-11-08T02:19:12.740610Z","shell.execute_reply":"2022-11-08T02:19:12.769690Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# other bad images noticed\nbad_images = [ '6d496f79fe93', '06f6423be3f9', '8568c0c16033', '7e99b9eb897b', '7ba4420fbb89']\nids_to_remove += bad_images\nlen(ids_to_remove)","metadata":{"execution":{"iopub.status.busy":"2022-11-08T02:19:14.583146Z","iopub.execute_input":"2022-11-08T02:19:14.583601Z","iopub.status.idle":"2022-11-08T02:19:14.591441Z","shell.execute_reply.started":"2022-11-08T02:19:14.583562Z","shell.execute_reply":"2022-11-08T02:19:14.589993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Making cleaner annotation file","metadata":{}},{"cell_type":"code","source":"cdf = pd.DataFrame()\nnull_boxes = 0\nfalse_boxes = 0\nfor idx in jdf.index:\n    if idx in ids_to_remove: # skip bad image\n        continue\n    r = jdf.loc[idx]\n    if pd.isnull(r.boxes) and r.label_y != 'Neg':\n        null_boxes += 1\n        continue\n    if pd.notnull(r.boxes) and r.label_y == 'Neg':\n        false_boxes += 1\n        continue\n    cdf = cdf.append(r)\n\nprint(f'Amount of null_boxes detected:{null_boxes}; false_boxes:{false_boxes}')\nprint(f'Amount of records removed:{len(jdf) - len(cdf)}')","metadata":{"execution":{"iopub.status.busy":"2022-11-08T02:27:45.269414Z","iopub.execute_input":"2022-11-08T02:27:45.269874Z","iopub.status.idle":"2022-11-08T02:28:22.561484Z","shell.execute_reply.started":"2022-11-08T02:27:45.269832Z","shell.execute_reply":"2022-11-08T02:28:22.560301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# we still may have duplicates from Studies with multiple images - remove duplicates\nbefore = len(cdf)\ncdf = cdf.drop_duplicates(subset=['StudyInstanceUID'], keep='first')\nprint(f'Amount of duplicate records removed:{before - len(cdf)}')","metadata":{"execution":{"iopub.status.busy":"2022-11-08T02:29:26.742314Z","iopub.execute_input":"2022-11-08T02:29:26.742770Z","iopub.status.idle":"2022-11-08T02:29:26.757156Z","shell.execute_reply.started":"2022-11-08T02:29:26.742730Z","shell.execute_reply":"2022-11-08T02:29:26.755996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cdf.to_csv('train_512_clean.csv', index_label = 'id')","metadata":{"execution":{"iopub.status.busy":"2022-11-08T02:35:18.060632Z","iopub.execute_input":"2022-11-08T02:35:18.061009Z","iopub.status.idle":"2022-11-08T02:35:18.098275Z","shell.execute_reply.started":"2022-11-08T02:35:18.060978Z","shell.execute_reply":"2022-11-08T02:35:18.097058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"t = pd.read_csv('train_512_clean.csv')\nt.head()","metadata":{"execution":{"iopub.status.busy":"2022-11-08T02:35:20.992219Z","iopub.execute_input":"2022-11-08T02:35:20.992658Z","iopub.status.idle":"2022-11-08T02:35:21.033871Z","shell.execute_reply.started":"2022-11-08T02:35:20.992622Z","shell.execute_reply":"2022-11-08T02:35:21.032539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Labels distribution after clean","metadata":{}},{"cell_type":"code","source":"fdist = show_dist(cdf)\nprint(fdist)\nprint('From:')\nprint(idist)","metadata":{"execution":{"iopub.status.busy":"2022-11-08T02:31:45.742948Z","iopub.execute_input":"2022-11-08T02:31:45.743343Z","iopub.status.idle":"2022-11-08T02:31:45.916700Z","shell.execute_reply.started":"2022-11-08T02:31:45.743300Z","shell.execute_reply":"2022-11-08T02:31:45.915709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Have a nice day!\nPlease, upvote, thanks!","metadata":{}}]}