{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# SOLUTION PIPELINE\n\n1. *Image level Detection algorithm* - [FISABIO-Covid19-Yolov5m-img512](https://www.kaggle.com/meenakshiramaswamy/fisabio-covid19-yolov5m-img512) - Yolov5m model with imagesize 512, labels 'opacity' , 'none' used for Yolov5m model training\n1. *Study level classification algorithm* - [SIIM-COVID_img512_EF7_TPU](https://www.kaggle.com/meenakshiramaswamy/siim-covid-img512-ef7-tpu) - TPU model EfficientNet B7 used with 5 folds on competition training images\n1. [SIIM-COVID_img512_extdata_gpu](https://www.kaggle.com/meenakshiramaswamy/siim-covid-img512-extdata-gpu) - GPU model EfficeintNet B0 used with 5 folds on external data provided by RICORD [x-ray dataset](https://www.kaggle.com/raddar/ricord-covid19-xray-positive-tests)\n1. [SIIM-COVID_TF_metadata_NN](https://www.kaggle.com/meenakshiramaswamy/siim-covid-tf-metadata-nn) - simple NN model for metadata\n1. [SIIM-COVID-Submit-img512-ef](https://www.kaggle.com/meenakshiramaswamy/siim-covid-submit-img512-ef) - Inference kernel to predict both study level and image level labels on public and private test data","metadata":{}},{"cell_type":"markdown","source":"Credit to below great kernels","metadata":{}},{"cell_type":"markdown","source":"# REFERENCES\n\n* https://www.kaggle.com/h053473666/siim-cov19-efnb7-yolov5-infer\n* https://www.kaggle.com/salmaneunus/siim-covid-19-detection-submisssion-2/notebook\n* https://www.kaggle.com/awsaf49/siim-covid-19-effnetb6-study-level-infer/execution\n* https://www.kaggle.com/mistag/data-stratified-tfrecords-for-classification\n* https://www.kaggle.com/ayuraj/train-covid-19-detection-using-yolov5\n* https://www.kaggle.com/espsiyam/train-and-detect-covid-19-using-yolov5x","metadata":{}},{"cell_type":"markdown","source":"# Classification Model Reference kernel\n\nhttps://www.kaggle.com/ajax0564/tpu-eff-b5-cbam","metadata":{}},{"cell_type":"markdown","source":"# Discussion threads\n\n1. https://www.kaggle.com/c/siim-covid19-detection/discussion/240250#1315782\n1. https://www.kaggle.com/c/siim-covid19-detection/discussion/241238\n1. https://www.kaggle.com/c/siim-covid19-detection/discussion/240329\n1. https://www.kaggle.com/c/siim-covid19-detection/discussion/246597\n1. https://www.kaggle.com/c/siim-covid19-detection/discussion/246536\n","metadata":{}},{"cell_type":"markdown","source":"# Submission format\n\n**Id,PredictionString**\n* 2b95d54e4be65_study,negative 1 0 0 1 1\n* 2b95d54e4be66_study,typical 1 0 0 1 1\n* 2b95d54e4be67_study,indeterminate 1 0 0 1 1 atypical 1 0 0 1 1\n* 2b95d54e4be68_image,none 1 0 0 1 1\n* 2b95d54e4be69_image,opacity 0.5 100 100 200 200 opacity 0.7 10 10 20 20\netc.\n\n1. Since the image level prediction includes bbox and the prediction format is ***opacity/none conf_score xmin ymin xmax ymax*.**  we use detection alogorithm like Yolo, MMDetection, Detectron, RetinaNet etc \n       *     \"negative\", \"typical\", \"indeterminate\", \"atypical\"\n1. The study level prediction includes one pixel bounding box and the prediction format is ***class ID confidence_score 0 0 1 1***","metadata":{}},{"cell_type":"code","source":"!cp ../input/gdcm-conda-install/gdcm.tar .\n!tar -xvzf gdcm.tar\n!conda install --offline ./gdcm/gdcm-2.8.9-py37h71b2a6d_0.tar.bz2\nprint(\"done\")","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:36:26.249176Z","iopub.execute_input":"2021-08-08T16:36:26.249628Z","iopub.status.idle":"2021-08-08T16:36:50.626419Z","shell.execute_reply.started":"2021-08-08T16:36:26.249542Z","shell.execute_reply":"2021-08-08T16:36:50.625436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#import cuml\n#print('RAPIDS version',cuml.__version__)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-08-08T16:36:50.628239Z","iopub.execute_input":"2021-08-08T16:36:50.628613Z","iopub.status.idle":"2021-08-08T16:36:50.63433Z","shell.execute_reply.started":"2021-08-08T16:36:50.628572Z","shell.execute_reply":"2021-08-08T16:36:50.632631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import sys\nsys.path.append('../input/efficientnet-keras-dataset/efficientnet_kaggle')\nfrom efficientnet.tfkeras import *","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:36:50.637861Z","iopub.execute_input":"2021-08-08T16:36:50.638178Z","iopub.status.idle":"2021-08-08T16:36:55.125644Z","shell.execute_reply.started":"2021-08-08T16:36:50.638141Z","shell.execute_reply":"2021-08-08T16:36:55.124769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pydicom\nfrom pydicom.pixel_data_handlers.util import convert_color_space, apply_voi_lut\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:36:55.127195Z","iopub.execute_input":"2021-08-08T16:36:55.127554Z","iopub.status.idle":"2021-08-08T16:36:55.349074Z","shell.execute_reply.started":"2021-08-08T16:36:55.127517Z","shell.execute_reply":"2021-08-08T16:36:55.348273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd, numpy as np, glob\nimport math, re, tqdm, os, gc, cv2\nfrom time import time\nfrom sklearn.model_selection import KFold\nfrom sklearn.metrics import classification_report, roc_auc_score\n\n\n#from operator import itemgetter\nfrom matplotlib import pyplot as plt\nimport tensorflow as tf\nimport tensorflow.keras.backend as K\nimport tensorflow_addons as tfa\nfrom PIL import Image\nimport efficientnet.tfkeras as efn\nimport shutil, torch","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:36:55.350443Z","iopub.execute_input":"2021-08-08T16:36:55.350802Z","iopub.status.idle":"2021-08-08T16:36:57.096874Z","shell.execute_reply.started":"2021-08-08T16:36:55.350763Z","shell.execute_reply":"2021-08-08T16:36:57.096056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.set_option('display.max_columns', 100)\npd.set_option('display.max_colwidth' , 150)","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:36:57.098076Z","iopub.execute_input":"2021-08-08T16:36:57.098429Z","iopub.status.idle":"2021-08-08T16:36:57.10517Z","shell.execute_reply.started":"2021-08-08T16:36:57.098395Z","shell.execute_reply":"2021-08-08T16:36:57.104465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#combined_train_df = pd.read_csv('../input/preprocessessiimcovid19ds/combined_train_df.csv')\n#combined_train_df.head(5)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-08-08T16:36:57.10644Z","iopub.execute_input":"2021-08-08T16:36:57.107019Z","iopub.status.idle":"2021-08-08T16:36:57.114486Z","shell.execute_reply.started":"2021-08-08T16:36:57.10698Z","shell.execute_reply":"2021-08-08T16:36:57.113609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#train_df = combined_train_df[combined_train_df.image_label != 'none']\n#train_df.shape","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-08-08T16:36:57.117281Z","iopub.execute_input":"2021-08-08T16:36:57.117652Z","iopub.status.idle":"2021-08-08T16:36:57.122851Z","shell.execute_reply.started":"2021-08-08T16:36:57.117615Z","shell.execute_reply":"2021-08-08T16:36:57.122071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"subm_df = pd.read_csv('../input/siim-covid19-detection/sample_submission.csv')\nsubm_df.shape\n","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:36:57.126168Z","iopub.execute_input":"2021-08-08T16:36:57.126498Z","iopub.status.idle":"2021-08-08T16:36:57.146492Z","shell.execute_reply.started":"2021-08-08T16:36:57.126454Z","shell.execute_reply":"2021-08-08T16:36:57.145779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#subm_df[subm_df.id=='2fb11712bc93_study']","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-08-08T16:36:57.148576Z","iopub.execute_input":"2021-08-08T16:36:57.148815Z","iopub.status.idle":"2021-08-08T16:36:57.153245Z","shell.execute_reply.started":"2021-08-08T16:36:57.148791Z","shell.execute_reply":"2021-08-08T16:36:57.152452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_dir = '../input/siim-covid19-detection/test/'\n#test_images_lst = glob.glob(test_dir+'*/*/*.dcm')\n#len(test_images_lst)\n\nlabel_cols = ['0','1','2','3']","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:36:57.154635Z","iopub.execute_input":"2021-08-08T16:36:57.155165Z","iopub.status.idle":"2021-08-08T16:36:57.16116Z","shell.execute_reply.started":"2021-08-08T16:36:57.155128Z","shell.execute_reply":"2021-08-08T16:36:57.160411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"filepaths = glob.glob('/kaggle/input/siim-covid19-detection/test/**/*dcm',recursive=True)\ntest_df = pd.DataFrame({'filepath':filepaths,})\ntest_df['image_id'] = test_df.filepath.map(lambda x: x.split('/')[-1].replace('.dcm', '')+'_image')\ntest_df['study_id'] = test_df.filepath.map(lambda x: x.split('/')[-3].replace('.dcm', '')+'_study')\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:36:57.162399Z","iopub.execute_input":"2021-08-08T16:36:57.16289Z","iopub.status.idle":"2021-08-08T16:37:02.494355Z","shell.execute_reply.started":"2021-08-08T16:36:57.162851Z","shell.execute_reply":"2021-08-08T16:37:02.493597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.image_id.nunique(), test_df.study_id.nunique()","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:02.495751Z","iopub.execute_input":"2021-08-08T16:37:02.496105Z","iopub.status.idle":"2021-08-08T16:37:02.505713Z","shell.execute_reply.started":"2021-08-08T16:37:02.49607Z","shell.execute_reply":"2021-08-08T16:37:02.504792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.study_id.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:02.506938Z","iopub.execute_input":"2021-08-08T16:37:02.507484Z","iopub.status.idle":"2021-08-08T16:37:02.517812Z","shell.execute_reply.started":"2021-08-08T16:37:02.507445Z","shell.execute_reply":"2021-08-08T16:37:02.517026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.shape","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:02.519069Z","iopub.execute_input":"2021-08-08T16:37:02.519636Z","iopub.status.idle":"2021-08-08T16:37:02.52678Z","shell.execute_reply.started":"2021-08-08T16:37:02.519601Z","shell.execute_reply":"2021-08-08T16:37:02.525726Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"study_len = subm_df.shape[0] - len(test_df)","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:02.528206Z","iopub.execute_input":"2021-08-08T16:37:02.528684Z","iopub.status.idle":"2021-08-08T16:37:02.537169Z","shell.execute_reply.started":"2021-08-08T16:37:02.52864Z","shell.execute_reply":"2021-08-08T16:37:02.536396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print (f'There are total {subm_df.shape[0]} rows in public submission, out of which {len(test_df)} images, {study_len} studies are to be predicted'  )","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:02.539969Z","iopub.execute_input":"2021-08-08T16:37:02.540327Z","iopub.status.idle":"2021-08-08T16:37:02.547061Z","shell.execute_reply.started":"2021-08-08T16:37:02.540291Z","shell.execute_reply":"2021-08-08T16:37:02.546206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**NOTE: Some studies may have more than one image**","metadata":{}},{"cell_type":"code","source":"#save_dir = f'./resized_test/study/'\n#os.makedirs(save_dir, exist_ok=True)\n\ntest_img_dir = f'./resized_test/image/'\nos.makedirs(test_img_dir, exist_ok=True)","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:02.548506Z","iopub.execute_input":"2021-08-08T16:37:02.548864Z","iopub.status.idle":"2021-08-08T16:37:02.555024Z","shell.execute_reply.started":"2021-08-08T16:37:02.548811Z","shell.execute_reply":"2021-08-08T16:37:02.554249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# function to convert dicom to array\n#https://www.geeksforgeeks.org/clahe-histogram-eqalization-opencv/\nCLIP_LIMIT = 2.5\nGRID_SIZE = (8,8)\n\ndef dicom2array(path, voi_lut=True, fix_monochrome=True, use_clahe=True):\n    dicom = pydicom.read_file(path)\n    # VOI LUT (if available by DICOM device) is used to\n    # transform raw DICOM data to \"human-friendly\" view\n    print (path)\n    if voi_lut:\n        data = apply_voi_lut(dicom.pixel_array, dicom)\n    else:\n        data = dicom.pixel_array\n        \n    data = data - np.min(data)\n    data = data / np.max(data)\n    data = (data * 255).astype(np.uint8)\n    \n    # depending on this value, X-ray may look inverted - fix that:\n    if fix_monochrome and dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = np.amax(data) - data\n    #print (data)\n    if use_clahe:\n        clahe = cv2.createCLAHE(clipLimit=CLIP_LIMIT, tileGridSize=GRID_SIZE)\n        climg = clahe.apply(data.astype('uint8'))\n        data = Image.fromarray(climg.astype('uint8'), 'L')\n    else:\n        data = Image.fromarray(data.astype('uint8'), 'L')\n    \n    return data","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:02.55659Z","iopub.execute_input":"2021-08-08T16:37:02.557202Z","iopub.status.idle":"2021-08-08T16:37:02.56704Z","shell.execute_reply.started":"2021-08-08T16:37:02.557157Z","shell.execute_reply":"2021-08-08T16:37:02.566313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def resize(array, size, keep_ratio=False, resample=Image.LANCZOS):\n    # Original from: https://www.kaggle.com/xhlulu/vinbigdata-process-and-resize-to-image\n    #im = Image.fromarray(array)\n    im = array\n    if keep_ratio:\n        im.thumbnail((size, size), resample)\n    else:\n        im = im.resize((size, size), resample)\n    \n    return im","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:02.56845Z","iopub.execute_input":"2021-08-08T16:37:02.568839Z","iopub.status.idle":"2021-08-08T16:37:02.576746Z","shell.execute_reply.started":"2021-08-08T16:37:02.568804Z","shell.execute_reply":"2021-08-08T16:37:02.575852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if subm_df.shape[0] == 2477:\n    PRIVATE_SUB = False\nelse:\n    PRIVATE_SUB = True\n\nif  not PRIVATE_SUB: # public sub\n    test_lst = test_df.filepath.iloc[:10].values\nelse:\n    test_lst = test_df.filepath.values","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:02.577791Z","iopub.execute_input":"2021-08-08T16:37:02.578028Z","iopub.status.idle":"2021-08-08T16:37:02.586226Z","shell.execute_reply.started":"2021-08-08T16:37:02.578005Z","shell.execute_reply":"2021-08-08T16:37:02.585456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"file = glob.glob('/kaggle/input/siim-covid19-detection/test/10e407d29b98/*/*.dcm')[0]\nfig = plt.figure(figsize=(15,15))\naxes = fig.add_subplot(1, 2, 1)\n\nxray = dicom2array(file, use_clahe = False)\nimg = resize(xray, size=768) \naxes.set_title('Original')\nplt.imshow(img, cmap='gray')\naxes = fig.add_subplot(1, 2, 2)\nxray = dicom2array(file, use_clahe=True)\naxes.set_title('CLAHE')\nplt.imshow(img, cmap='gray');","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:02.587479Z","iopub.execute_input":"2021-08-08T16:37:02.588026Z","iopub.status.idle":"2021-08-08T16:37:03.817325Z","shell.execute_reply.started":"2021-08-08T16:37:02.587988Z","shell.execute_reply":"2021-08-08T16:37:03.813174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''\nstudy_id = []\nimage_path = []\n\n\nfor i, file in tqdm.tqdm_notebook(enumerate(test_lst)):\n\n    #print (i, file)\n        # set keep_ratio=True to have original aspect ratio\n    xray = dicom2array( file)\n    im = resize(xray, size=512)  \n    study = file.split('/')[-3] + '_study.jpg'\n    im.save(os.path.join(save_dir, study))\n    study_id.append(study)\n\n    image = file.split('/')[-1] \n    #print (image)\n    image = image.replace('.dcm','_image.jpg') \n    image_path.append( test_img_dir + image )\n'''","metadata":{"_kg_hide-output":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-08-08T16:37:03.822033Z","iopub.execute_input":"2021-08-08T16:37:03.822321Z","iopub.status.idle":"2021-08-08T16:37:03.828235Z","shell.execute_reply.started":"2021-08-08T16:37:03.822294Z","shell.execute_reply":"2021-08-08T16:37:03.827349Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''\nsubm_study_df = pd.DataFrame([study_id, image_path]).T\nsubm_study_df.columns=['id','image_path']\nsubm_study_df[label_cols] = 0\nsubm_study_df.head()\nsubm_study_df.shape\n'''","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-08-08T16:37:03.831923Z","iopub.execute_input":"2021-08-08T16:37:03.832203Z","iopub.status.idle":"2021-08-08T16:37:03.838213Z","shell.execute_reply.started":"2021-08-08T16:37:03.832179Z","shell.execute_reply":"2021-08-08T16:37:03.837064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_id = []\ndim0 = []\ndim1 = []","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:03.839853Z","iopub.execute_input":"2021-08-08T16:37:03.840282Z","iopub.status.idle":"2021-08-08T16:37:03.845421Z","shell.execute_reply.started":"2021-08-08T16:37:03.84023Z","shell.execute_reply":"2021-08-08T16:37:03.844471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> **Resize and save the image**","metadata":{}},{"cell_type":"code","source":"#if  not PRIVATE_SUB: # public sub\nfor i, file in tqdm.tqdm_notebook(enumerate(test_lst)):\n    #print (file)\n        # set keep_ratio=True to have original aspect ratio\n    \n    xray = dicom2array(file, use_clahe = False)\n    \n    dim0.append(xray.size[0])\n    dim1.append(xray.size[1])\n    im = resize(xray, size=768) \n    image = file.split('/')[-1] \n    #print (image)\n    image = image.replace('.dcm','_image.jpg') \n\n    #image = image.replace('.dcm','.jpg') \n    #print (image)\n    im.save(os.path.join(test_img_dir, image))\n    image_id.append(image.replace('.jpg',''))\n    #dim0.append(im.size[0])\n    #dim1.append(im.size[1])","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-08-08T16:37:03.846939Z","iopub.execute_input":"2021-08-08T16:37:03.847485Z","iopub.status.idle":"2021-08-08T16:37:10.216539Z","shell.execute_reply.started":"2021-08-08T16:37:03.84745Z","shell.execute_reply":"2021-08-08T16:37:10.215573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"subm_image_df = pd.DataFrame.from_dict({'image_id': image_id, 'dim0': dim0, 'dim1': dim1})\n#test_meta[label_cols] = 0","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.217864Z","iopub.execute_input":"2021-08-08T16:37:10.218212Z","iopub.status.idle":"2021-08-08T16:37:10.225592Z","shell.execute_reply.started":"2021-08-08T16:37:10.218175Z","shell.execute_reply":"2021-08-08T16:37:10.224776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"subm_image_df['image_path'] = subm_image_df.apply(lambda row : (test_img_dir + row['image_id']+'.jpg'), axis = 1)\nsubm_image_df[label_cols] = 0","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.226743Z","iopub.execute_input":"2021-08-08T16:37:10.227273Z","iopub.status.idle":"2021-08-08T16:37:10.242875Z","shell.execute_reply.started":"2021-08-08T16:37:10.227224Z","shell.execute_reply":"2021-08-08T16:37:10.242026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#subm_image_df","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.24389Z","iopub.execute_input":"2021-08-08T16:37:10.244217Z","iopub.status.idle":"2021-08-08T16:37:10.249786Z","shell.execute_reply.started":"2021-08-08T16:37:10.244183Z","shell.execute_reply":"2021-08-08T16:37:10.24897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df = pd.merge(test_df, subm_image_df, on = 'image_id', how = 'left')\ntest_df = test_df if PRIVATE_SUB else test_df[:10]","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.250903Z","iopub.execute_input":"2021-08-08T16:37:10.251245Z","iopub.status.idle":"2021-08-08T16:37:10.268495Z","shell.execute_reply.started":"2021-08-08T16:37:10.251209Z","shell.execute_reply":"2021-08-08T16:37:10.267767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head(10)","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.269578Z","iopub.execute_input":"2021-08-08T16:37:10.26991Z","iopub.status.idle":"2021-08-08T16:37:10.292066Z","shell.execute_reply.started":"2021-08-08T16:37:10.269876Z","shell.execute_reply":"2021-08-08T16:37:10.291011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.shape","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.293798Z","iopub.execute_input":"2021-08-08T16:37:10.294204Z","iopub.status.idle":"2021-08-08T16:37:10.300356Z","shell.execute_reply.started":"2021-08-08T16:37:10.29416Z","shell.execute_reply":"2021-08-08T16:37:10.29935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#shutil.remove './resized_test/image/*'","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-08-08T16:37:10.301977Z","iopub.execute_input":"2021-08-08T16:37:10.302354Z","iopub.status.idle":"2021-08-08T16:37:10.30735Z","shell.execute_reply.started":"2021-08-08T16:37:10.302318Z","shell.execute_reply":"2021-08-08T16:37:10.306288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# EXTRACT META DATA","metadata":{}},{"cell_type":"code","source":"def get_observation_data(path):\n    '''Get information from the .dcm files.\n    path: complete path to the .dcm file'''\n\n    image_data = pydicom.read_file(path, stop_before_pixels=True)\n    \n\n\n    # Dictionary to store the information from the image\n    observation_data = {\n        \"FileNumber\" : path.split('/')[5],\n        \"Rows\" : image_data.Rows,\n        \"Columns\" : image_data.Columns,\n        \"PatientID\" : image_data.PatientID,\n        \"BodyPartExamined\" : image_data.get('BodyPartExamined',\"(missing)\"),\n        \"SliceThickness\" : int(image_data.get('SliceThickness',0)),\n        \"KVP\" : int(image_data.get('KVP',0)),\n        'Manufacturer Model Name' : str(image_data.get('ManufacturersModelName',\"(missing)\")),\n        \"Manufacturer\" : str(image_data.get(\"Manufacturer\",\"(missing)\")),\n        \"DistanceSourceToDetector\" : int(image_data.get('DistanceSourceToDetector',0)),\n        \"DistanceSourceToPatient\" : int(image_data.get('DistanceSourceToPatient',0)),\n        \"GantryDetectorTilt\" : int(image_data.get('GantryDetectorTilt',0)),\n        \"TableHeight\" : int(image_data.get('TableHeight',0)),\n        \"De-identification Method\" : str(image_data.get(\"De-identification Method \",\"(missing)\")),\n        \"RotationDirection\" : image_data.get('RotationDirection',\"(missing)\"),\n        \"XRayTubeCurrent\" : int(image_data.get('XRayTubeCurrent',0)),\n        \"GeneratorPower\" : image_data.get('GeneratorPower',\"(missing)\"), #int(image_data.GeneratorPower),\n        \"ConvolutionKernel\" : image_data.get('ConvolutionKernel',\"(missing)\"),\n        \"PatientPosition\" : image_data.get('PatientPosition',\"(missing)\"),\n        \"ImagePositionPatient\" :image_data.get('ImagePositionPatient',\"(missing)\"),\n        \"ImageOrientationPatient\" : str(image_data.get('ImageOrientationPatient',\"(missing)\")),\n        \"PhotometricInterpretation\" : image_data.get('PhotometricInterpretation',\"(missing)\"),\n        \"ImageType\" : str(image_data.get('ImageType',\"(missing)\")),\n        \"PixelSpacing\" : str(image_data.get('PixelSpacing',\"(missing)\")),\n        \"WindowCenter\" : image_data.get('WindowCenter',\"(Missing)\"),\n        \"WindowWidth\" : image_data.get('WindowWidth',\"(missing)\"),\n        \"Modality\" : image_data.get('Modality',\"(missing)\"),\n        \n        \"PixelPaddingValue\" : int(image_data.get('PixelPaddingValue',0)),\n        \"SamplesPerPixel\" : image_data.get('SamplesPerPixel',\"(missing)\"),\n        \"SliceLocation\" : image_data.get('SliceLocation',\"(missing)\"),\n        \"Instance Number\" : image_data.InstanceNumber,  # .dcm file name\n        \"BitsAllocated\" : image_data.BitsAllocated,\n        \"BitsStored\" : image_data.BitsStored,\n        \"HighBit\" : image_data.HighBit,\n        \"PixelRepresentation\" : image_data.get('PixelRepresentation',\"(missing)\"),\n        \"RescaleIntercept\" : int(image_data.get('RescaleIntercept',0)),\n        \"RescaleSlope\" : int(image_data.get('RescaleSlope',0)),\n        \"RescaleType\" : image_data.get('RescaleType',\"(missing)\"),\n        \"StudyInstanceUID\" :   image_data.get('StudyInstanceUID',\"(missing)\")   ,     \n        \"SeriesInstanceUID\" :  image_data.get('SeriesInstanceUID',\"(missing)\"),\n        \"SOPInstanceUID\" :  image_data.get('SOPInstanceUID',\"(missing)\")\n    }\n    \n    return observation_data","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.308827Z","iopub.execute_input":"2021-08-08T16:37:10.30931Z","iopub.status.idle":"2021-08-08T16:37:10.325803Z","shell.execute_reply.started":"2021-08-08T16:37:10.309253Z","shell.execute_reply":"2021-08-08T16:37:10.324961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"arr_meta_data =[]\nfor i, filename in tqdm.notebook.tqdm(enumerate(test_df.filepath)):#len(train_df):\n    arr_meta_data.append(get_observation_data(filename))","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.328617Z","iopub.execute_input":"2021-08-08T16:37:10.328855Z","iopub.status.idle":"2021-08-08T16:37:10.398963Z","shell.execute_reply.started":"2021-08-08T16:37:10.328833Z","shell.execute_reply":"2021-08-08T16:37:10.398193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dcm_meta_data = pd.DataFrame(arr_meta_data)\n","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.400335Z","iopub.execute_input":"2021-08-08T16:37:10.400888Z","iopub.status.idle":"2021-08-08T16:37:10.410183Z","shell.execute_reply.started":"2021-08-08T16:37:10.400849Z","shell.execute_reply":"2021-08-08T16:37:10.409177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\ndcm_meta_data['image_id'] = dcm_meta_data.SOPInstanceUID.map(lambda x: x+'_image')\ndcm_meta_data['study_id'] = dcm_meta_data.StudyInstanceUID.map(lambda x: x+'_study')","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.411715Z","iopub.execute_input":"2021-08-08T16:37:10.412076Z","iopub.status.idle":"2021-08-08T16:37:10.423594Z","shell.execute_reply.started":"2021-08-08T16:37:10.41204Z","shell.execute_reply":"2021-08-08T16:37:10.422612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dcm_meta_data.head()","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.42645Z","iopub.execute_input":"2021-08-08T16:37:10.426733Z","iopub.status.idle":"2021-08-08T16:37:10.467422Z","shell.execute_reply.started":"2021-08-08T16:37:10.426699Z","shell.execute_reply":"2021-08-08T16:37:10.466515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_test_df = pd.merge(test_df,\n                             dcm_meta_data [['PatientID', 'StudyInstanceUID', 'SeriesInstanceUID',\n                                             'SOPInstanceUID','BodyPartExamined',\n                                             'PhotometricInterpretation','ImageType','Modality',\n                                             'FileNumber','Rows','Columns',\n                                            'SamplesPerPixel','Instance Number',\n                                            'BitsAllocated','BitsStored','HighBit','image_id','study_id'\n                                            ]],                 \n                             on = ['image_id','study_id'],\n                             how = 'inner')\n\ntest_df.shape[0],dcm_meta_data.shape[0], combined_test_df.shape[0]","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.468981Z","iopub.execute_input":"2021-08-08T16:37:10.469535Z","iopub.status.idle":"2021-08-08T16:37:10.488078Z","shell.execute_reply.started":"2021-08-08T16:37:10.469355Z","shell.execute_reply":"2021-08-08T16:37:10.487344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_test_df = combined_test_df.reset_index(drop = True)\ncombined_test_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.490095Z","iopub.execute_input":"2021-08-08T16:37:10.490659Z","iopub.status.idle":"2021-08-08T16:37:10.520143Z","shell.execute_reply.started":"2021-08-08T16:37:10.490618Z","shell.execute_reply":"2021-08-08T16:37:10.519048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"combined_test_df.info()","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.521754Z","iopub.execute_input":"2021-08-08T16:37:10.522122Z","iopub.status.idle":"2021-08-08T16:37:10.54221Z","shell.execute_reply.started":"2021-08-08T16:37:10.522083Z","shell.execute_reply":"2021-08-08T16:37:10.541273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nfrom sklearn.preprocessing import LabelEncoder\n# creating instance of labelencoder\nle = LabelEncoder()\ncombined_test_df['filenumber_encoded'] = le.fit_transform(combined_test_df['FileNumber'])\n\ncombined_test_df['pat_id_transformed']= le.fit_transform(combined_test_df['PatientID'])\n\ncombined_test_df['imgtype_transformed']= le.fit_transform(combined_test_df['ImageType'])\n\ncombined_test_df['moda_transformed']= le.fit_transform(combined_test_df['Modality'])\n\ncombined_test_df['bp_transformed']= le.fit_transform(combined_test_df['BodyPartExamined'])\n\ncombined_test_df['ph_intp_transformed']= le.fit_transform(combined_test_df['PhotometricInterpretation'])\n\ncombined_test_df['SopID_transformed']= le.fit_transform(combined_test_df['SOPInstanceUID'])\n\ncombined_test_df['SerID_transformed']= le.fit_transform(combined_test_df['SeriesInstanceUID'])\n\ncombined_test_df['StdID_transformed']= le.fit_transform(combined_test_df['StudyInstanceUID'])","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.543633Z","iopub.execute_input":"2021-08-08T16:37:10.54401Z","iopub.status.idle":"2021-08-08T16:37:10.55664Z","shell.execute_reply.started":"2021-08-08T16:37:10.543971Z","shell.execute_reply":"2021-08-08T16:37:10.555698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_meta = combined_test_df[['pat_id_transformed',\n               'bp_transformed',\n               'ph_intp_transformed',\n                'imgtype_transformed',\n               'moda_transformed',\n               'SopID_transformed',\n               'SerID_transformed',\n               'StdID_transformed',\n               'Rows',\n               'Columns',\n               'BitsAllocated',\n               'BitsStored',\n               'HighBit',\n                'filenumber_encoded',\n               'dim0',\n               'dim1',\n              'SamplesPerPixel',\n               'Instance Number']]","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.558242Z","iopub.execute_input":"2021-08-08T16:37:10.559127Z","iopub.status.idle":"2021-08-08T16:37:10.56845Z","shell.execute_reply.started":"2021-08-08T16:37:10.559066Z","shell.execute_reply":"2021-08-08T16:37:10.567664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_meta","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.56994Z","iopub.execute_input":"2021-08-08T16:37:10.570375Z","iopub.status.idle":"2021-08-08T16:37:10.592895Z","shell.execute_reply.started":"2021-08-08T16:37:10.570336Z","shell.execute_reply":"2021-08-08T16:37:10.591806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# SETUP","metadata":{}},{"cell_type":"code","source":"DEVICE = 'GPU'\n\nif DEVICE == \"TPU\":\n    print(\"connecting to TPU...\")\n    try:\n        tpu = tf.distribute.cluster_resolver.TPUClusterResolver()\n        print('Running on TPU ', tpu.master())\n    except ValueError:\n        print(\"Could not connect to TPU\")\n        tpu = None\n\n    if tpu:\n        try:\n            print(\"initializing  TPU ...\")\n            tf.config.experimental_connect_to_cluster(tpu)\n            tf.tpu.experimental.initialize_tpu_system(tpu)\n            strategy = tf.distribute.experimental.TPUStrategy(tpu)\n            print(\"TPU initialized\")\n        except _:\n            print(\"failed to initialize TPU\")\n    else:\n        DEVICE = \"GPU\"\n\nif DEVICE != \"TPU\":\n    print(\"Using default strategy for CPU and single GPU\")\n    strategy = tf.distribute.get_strategy()\n\nif DEVICE == \"GPU\":\n    print(\"Num GPUs Available: \", len(tf.config.experimental.list_physical_devices('GPU')))\n    \n\nAUTO = tf.data.experimental.AUTOTUNE\nREPLICAS = strategy.num_replicas_in_sync\nprint(f'REPLICAS: {REPLICAS}')","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.594575Z","iopub.execute_input":"2021-08-08T16:37:10.595183Z","iopub.status.idle":"2021-08-08T16:37:10.661874Z","shell.execute_reply.started":"2021-08-08T16:37:10.59513Z","shell.execute_reply":"2021-08-08T16:37:10.660386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# CONFIG","metadata":{}},{"cell_type":"code","source":"#basic\nEFFNET = 7  # Efnet B7 used for training data\nEFFNET_META = 0 #Efnet B0 used for external data\nSEED = 42       \nIMAGE_SIZE = [768, 768]               \nBATCH_SIZE = 8\n\nTRANSFORM = False\n\nsat  = (0.7, 1.3)\ncont = (0.8, 1.2)\nbri  =  0.1\nROT_    = 0.0\nSHR_    = 2.0\nHZOOM_  = 8.0\nWZOOM_  = 8.0\nHSHIFT_ = 8.0\nWSHIFT_ = 8.0\n","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.663152Z","iopub.execute_input":"2021-08-08T16:37:10.663529Z","iopub.status.idle":"2021-08-08T16:37:10.670256Z","shell.execute_reply.started":"2021-08-08T16:37:10.663491Z","shell.execute_reply":"2021-08-08T16:37:10.669229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"name2label = {\n'negative': 0,\n 'indeterminate': 1,\n'atypical': 2,\n 'typical': 3}","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.671656Z","iopub.execute_input":"2021-08-08T16:37:10.672004Z","iopub.status.idle":"2021-08-08T16:37:10.677761Z","shell.execute_reply.started":"2021-08-08T16:37:10.671967Z","shell.execute_reply":"2021-08-08T16:37:10.676689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"name2label.items()\nlabel2name = {v:k for k, v in name2label.items()}","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.679314Z","iopub.execute_input":"2021-08-08T16:37:10.679786Z","iopub.status.idle":"2021-08-08T16:37:10.684985Z","shell.execute_reply.started":"2021-08-08T16:37:10.679749Z","shell.execute_reply":"2021-08-08T16:37:10.684029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''\nfrom numba import cuda\n\ncuda.select_device(0)\ncuda.close()\ncuda.select_device(0)\n'''\n","metadata":{"_kg_hide-output":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-08-08T16:37:10.686612Z","iopub.execute_input":"2021-08-08T16:37:10.687003Z","iopub.status.idle":"2021-08-08T16:37:10.695364Z","shell.execute_reply.started":"2021-08-08T16:37:10.686968Z","shell.execute_reply":"2021-08-08T16:37:10.694553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1. Predict image class with BBoxes using Yolov5","metadata":{}},{"cell_type":"code","source":"# using custom head with yolov5x\nos.listdir('../input/fisabio-covid19-yolov5m-img512/runs/train/exp/weights')","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.697344Z","iopub.execute_input":"2021-08-08T16:37:10.697582Z","iopub.status.idle":"2021-08-08T16:37:10.729019Z","shell.execute_reply.started":"2021-08-08T16:37:10.69756Z","shell.execute_reply":"2021-08-08T16:37:10.728203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"weights_dir = '/kaggle/input/fisabio-covid19-yolov5m-img512/runs/train/exp/weights/best.pt'\n\n#shutil.copytree('/kaggle/input/yolov5-official-v31-dataset/yolov5', './yolov5')\nshutil.copytree('../input/yolov5-repo/','./yolov5')\n#os.chdir('./yolov5') # install dependencies\nimage_dir = './resized_test/image/' \n\nimport torch\n#from IPython.display import Image, clear_output  # to display images\n\nfrom IPython.display import Image, clear_output  # to display images\n\nclear_output()\nprint('Setup complete. Using torch %s %s' % (torch.__version__, torch.cuda.get_device_properties(0) if torch.cuda.is_available() else 'CPU'))","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:10.730339Z","iopub.execute_input":"2021-08-08T16:37:10.730687Z","iopub.status.idle":"2021-08-08T16:37:11.043734Z","shell.execute_reply.started":"2021-08-08T16:37:10.73065Z","shell.execute_reply":"2021-08-08T16:37:11.042738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#os.chdir('./')../input/yolov5-repo/","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-08-08T16:37:11.045158Z","iopub.execute_input":"2021-08-08T16:37:11.045542Z","iopub.status.idle":"2021-08-08T16:37:11.049724Z","shell.execute_reply.started":"2021-08-08T16:37:11.045504Z","shell.execute_reply":"2021-08-08T16:37:11.048701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#./yolov5/detect.py\n!python ./yolov5/yolov5-master/detect.py --weights {weights_dir} --source './resized_test/image' --img {IMAGE_SIZE[0]} --conf 0.1 --iou-thres 0.5 --save-txt --save-conf","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:11.051218Z","iopub.execute_input":"2021-08-08T16:37:11.051605Z","iopub.status.idle":"2021-08-08T16:37:21.228036Z","shell.execute_reply.started":"2021-08-08T16:37:11.051549Z","shell.execute_reply":"2021-08-08T16:37:21.226946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Image('runs/detect/exp/f785f9c6bbf7_image.jpg')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-08-08T16:37:21.229813Z","iopub.execute_input":"2021-08-08T16:37:21.230196Z","iopub.status.idle":"2021-08-08T16:37:21.237193Z","shell.execute_reply.started":"2021-08-08T16:37:21.230153Z","shell.execute_reply":"2021-08-08T16:37:21.236429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# The submisison requires xmin, ymin, xmax, ymax format. \n# YOLOv5 returns x_center, y_center, width, height\ndef correct_bbox_format(bboxes):\n    correct_bboxes = []\n    for b in bboxes:\n        xc, yc = int(np.round(b[0]*IMAGE_SIZE[0])), int(np.round(b[1]*IMAGE_SIZE[0]))\n        w, h = int(np.round(b[2]*IMAGE_SIZE[0])), int(np.round(b[3]*IMAGE_SIZE[0]))\n\n        xmin = xc - int(np.round(w/2))\n        xmax = xc + int(np.round(w/2))\n        ymin = yc - int(np.round(h/2))\n        ymax = yc + int(np.round(h/2))\n        \n        correct_bboxes.append([xmin, xmax, ymin, ymax])\n        \n    return correct_bboxes\n\n# Read the txt file generated by YOLOv5 during inference and extract \n# confidence and bounding box coordinates.\ndef get_conf_bboxes(file_path):\n    confidence = []\n    bboxes = []\n    with open(file_path, 'r') as file:\n        for line in file:\n            preds = line.strip('\\n').split(' ')\n            preds = list(map(float, preds))\n            confidence.append(preds[-1])\n            bboxes.append(preds[1:-1])\n    return confidence, bboxes","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:21.238499Z","iopub.execute_input":"2021-08-08T16:37:21.238884Z","iopub.status.idle":"2021-08-08T16:37:21.2632Z","shell.execute_reply.started":"2021-08-08T16:37:21.238846Z","shell.execute_reply":"2021-08-08T16:37:21.262344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualize one of the predicted test image with bbox","metadata":{}},{"cell_type":"code","source":"PRED_PATH = 'runs/detect/exp/labels'\nprediction_files = os.listdir(PRED_PATH)\nprint('Number of test images predicted as opacity: ', len(prediction_files))","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:21.266833Z","iopub.execute_input":"2021-08-08T16:37:21.267284Z","iopub.status.idle":"2021-08-08T16:37:21.27531Z","shell.execute_reply.started":"2021-08-08T16:37:21.26724Z","shell.execute_reply":"2021-08-08T16:37:21.273942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#os.listdir('runs/detect/exp/labels')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-08-08T16:37:21.277247Z","iopub.execute_input":"2021-08-08T16:37:21.277677Z","iopub.status.idle":"2021-08-08T16:37:21.28495Z","shell.execute_reply.started":"2021-08-08T16:37:21.277636Z","shell.execute_reply":"2021-08-08T16:37:21.28402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"subm_image_df.rename(columns={'image_id':'id'}, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:21.288391Z","iopub.execute_input":"2021-08-08T16:37:21.288723Z","iopub.status.idle":"2021-08-08T16:37:21.296818Z","shell.execute_reply.started":"2021-08-08T16:37:21.288689Z","shell.execute_reply":"2021-08-08T16:37:21.295813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Prediction loop for submission\npredictions = []\n\nfor i, row in tqdm.tqdm_notebook(enumerate(subm_image_df.id)):\n\n    #print(i)\n    #id_name  = row.split('.')[0]\n    id_name = row.split('_')[0]\n    #print (id_name)\n    #id_level = row.id.split('_')[-1]\n\n    if f'{id_name}_image.txt' in prediction_files:\n        # opacity label\n        confidence, bboxes = get_conf_bboxes(f'{PRED_PATH}/{id_name}_image.txt')\n        bboxes = correct_bbox_format(bboxes)\n        pred_string = ''\n        for j, conf in enumerate(confidence):\n            pred_string += f'opacity {conf} ' + ' '.join(map(str, bboxes[j])) + ' '\n        predictions.append(pred_string[:-1]) \n    else:\n        predictions.append(\"none 1 0 0 1 1\")","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:21.298435Z","iopub.execute_input":"2021-08-08T16:37:21.298853Z","iopub.status.idle":"2021-08-08T16:37:21.345276Z","shell.execute_reply.started":"2021-08-08T16:37:21.298812Z","shell.execute_reply":"2021-08-08T16:37:21.344309Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"subm_image_df['PredictionString'] = predictions\nsubm_image_df = subm_image_df[['id','PredictionString']]\n#subm_image_df.rename(columns={'image_id':'id'}, inplace = True)\nsubm_image_df","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:21.346686Z","iopub.execute_input":"2021-08-08T16:37:21.347037Z","iopub.status.idle":"2021-08-08T16:37:21.363613Z","shell.execute_reply.started":"2021-08-08T16:37:21.347Z","shell.execute_reply":"2021-08-08T16:37:21.362771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Predict Study labels with TF models....5 FOLDS","metadata":{}},{"cell_type":"code","source":"# Free space\nshutil.rmtree('./yolov5')\nshutil.rmtree('./runs')\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:21.365078Z","iopub.execute_input":"2021-08-08T16:37:21.365584Z","iopub.status.idle":"2021-08-08T16:37:21.591608Z","shell.execute_reply.started":"2021-08-08T16:37:21.365547Z","shell.execute_reply":"2021-08-08T16:37:21.590497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# CLASSIFICATION MODEL","metadata":{}},{"cell_type":"code","source":"class SpatialAttentionModule(tf.keras.layers.Layer):\n    def __init__(self, kernel_size=3):\n        '''\n        paper: https://arxiv.org/abs/1807.06521\n        code: https://gist.github.com/innat/99888fa8065ecbf3ae2b297e5c10db70\n        '''\n        super(SpatialAttentionModule, self).__init__()\n        self.conv1 = tf.keras.layers.Conv2D(64, kernel_size=kernel_size, \n                                            use_bias=False, \n                                            kernel_initializer='he_normal',\n                                            strides=1, padding='same', \n                                            activation=tf.nn.relu6)\n        self.conv2 = tf.keras.layers.Conv2D(32, kernel_size=kernel_size, \n                                            use_bias=False, \n                                            kernel_initializer='he_normal',\n                                            strides=1, padding='same', \n                                            activation=tf.nn.relu6)\n        self.conv3 = tf.keras.layers.Conv2D(16, kernel_size=kernel_size, \n                                            use_bias=False, \n                                            kernel_initializer='he_normal',\n                                            strides=1, padding='same', \n                                            activation=tf.nn.relu6)\n        self.conv4 = tf.keras.layers.Conv2D(1, kernel_size=kernel_size,  \n                                            use_bias=False,\n                                            kernel_initializer='he_normal',\n                                            strides=1, padding='same', \n                                            activation=tf.math.sigmoid)\n\n    def call(self, inputs):\n        avg_out = tf.reduce_mean(inputs, axis=3)\n        max_out = tf.reduce_max(inputs,  axis=3)\n        x = tf.stack([avg_out, max_out], axis=3) \n        x = self.conv1(x)\n        x = self.conv2(x)\n        x = self.conv3(x)\n        return self.conv4(x)\n    \n# A custom layer\nclass ChannelAttentionModule(tf.keras.layers.Layer):\n    def __init__(self, ratio=8):\n        '''\n        paper: https://arxiv.org/abs/1807.06521\n        code: https://gist.github.com/innat/99888fa8065ecbf3ae2b297e5c10db70\n        '''\n        super(ChannelAttentionModule, self).__init__()\n        self.ratio = ratio\n        self.gapavg = tf.keras.layers.GlobalAveragePooling2D()\n        self.gmpmax = tf.keras.layers.GlobalMaxPooling2D()\n        \n    def build(self, input_shape):\n        self.conv1 = tf.keras.layers.Conv2D(input_shape[-1]//self.ratio, \n                                            kernel_size=1, \n                                            strides=1, padding='same',\n                                            use_bias=True, activation=tf.nn.relu)\n    \n        self.conv2 = tf.keras.layers.Conv2D(input_shape[-1], \n                                            kernel_size=1, \n                                            strides=1, padding='same',\n                                            use_bias=True, activation=tf.nn.relu)\n        super(ChannelAttentionModule, self).build(input_shape)\n\n    def call(self, inputs):\n        # compute gap and gmp pooling \n        gapavg = self.gapavg(inputs)\n        gmpmax = self.gmpmax(inputs)\n        gapavg = tf.keras.layers.Reshape((1, 1, gapavg.shape[1]))(gapavg)   \n        gmpmax = tf.keras.layers.Reshape((1, 1, gmpmax.shape[1]))(gmpmax)   \n        # forward passing to the respected layers\n        gapavg_out = self.conv2(self.conv1(gapavg))\n        gmpmax_out = self.conv2(self.conv1(gmpmax))\n        return tf.math.sigmoid(gapavg_out + gmpmax_out)\n    \n    def get_output_shape_for(self, input_shape):\n        return self.compute_output_shape(input_shape)\n\n    def compute_output_shape(self, input_shape):\n        output_len = input_shape[3]\n        return (input_shape[0], output_len)\n# Original Src: https://github.com/bfelbo/DeepMoji/blob/master/deepmoji/attlayer.py\n# Adoped and Modified: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/77269#454482\nclass AttentionWeightedAverage2D(tf.keras.layers.Layer):\n    def __init__(self, **kwargs):\n        self.init = tf.keras.initializers.get('uniform')\n        super(AttentionWeightedAverage2D, self).__init__(** kwargs)\n\n    def build(self, input_shape):\n        self.input_spec = [tf.keras.layers.InputSpec(ndim=4)]\n        assert len(input_shape) == 4\n        self.W = self.add_weight(shape=(input_shape[3], 1),\n                                 name='{}_W'.format(self.name),\n                                 initializer=self.init)\n        self._trainable_weights = [self.W]\n        super(AttentionWeightedAverage2D, self).build(input_shape)\n\n    def call(self, x):\n        # computes a probability distribution over the timesteps\n        # uses 'max trick' for numerical stability\n        # reshape is done to avoid issue with Tensorflow\n        # and 2-dimensional weights\n        logits  = K.dot(x, self.W)\n        x_shape = K.shape(x)\n        logits  = K.reshape(logits, (x_shape[0], x_shape[1], x_shape[2]))\n        ai      = K.exp(logits - K.max(logits, axis=[1,2], keepdims=True))\n        \n        att_weights    = ai / (K.sum(ai, axis=[1,2], keepdims=True) + K.epsilon())\n        weighted_input = x * K.expand_dims(att_weights)\n        result         = K.sum(weighted_input, axis=[1,2])\n        return result\n\n    def get_output_shape_for(self, input_shape):\n        return self.compute_output_shape(input_shape)\n\n    def compute_output_shape(self, input_shape):\n        output_len = input_shape[3]\n        return (input_shape[0], output_len)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-08-08T16:37:21.593191Z","iopub.execute_input":"2021-08-08T16:37:21.593611Z","iopub.status.idle":"2021-08-08T16:37:21.622074Z","shell.execute_reply.started":"2021-08-08T16:37:21.593568Z","shell.execute_reply":"2021-08-08T16:37:21.621033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"EFNS = [efn.EfficientNetB0, efn.EfficientNetB1, efn.EfficientNetB2, efn.EfficientNetB3, \n        efn.EfficientNetB4, efn.EfficientNetB5, efn.EfficientNetB6, efn.EfficientNetB7]\n\ndef build_model(dim, ef):\n    #dim = [*dim, 3]\n    inp = tf.keras.layers.Input(shape=(*dim,3))\n    #inp = tf.keras.layers.Input(shape=dim)\n    base = EFNS[ef](input_shape=(*dim,3),weights='imagenet',include_top=False)\n    x = base(inp)\n    #\n    CAN  = ChannelAttentionModule()\n    SPN = SpatialAttentionModule()\n    AWG  = AttentionWeightedAverage2D()\n    canx   = CAN(x)*x\n    spnx   = SPN(canx)*canx\n    gapx = tf.keras.layers.GlobalAveragePooling2D()(spnx)\n    wvgx   = tf.keras.layers.GlobalAveragePooling2D()(SPN(canx))\n    avg = tf.keras.layers.Average()([gapx, wvgx])\n    awg = AWG(x)\n    x = tf.keras.layers.Add()([avg, awg])\n    x = tf.keras.layers.BatchNormalization()(x)\n    x = tf.keras.layers.Dense(32, activation = 'relu')(x)\n    \n    \n    #x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    x = tf.keras.layers.Dense(4,activation='softmax')(x)\n    auc = tf.keras.metrics.AUC(curve='ROC',\n                               multi_label=True)   \n    f1  = tfa.metrics.F1Score(num_classes=4,average='macro',threshold=None)\n    \n    model = tf.keras.Model(inputs=inp,outputs=x)\n    opt = tf.keras.optimizers.Adam(learning_rate=0.001)\n    loss = tf.keras.losses.CategoricalCrossentropy(label_smoothing=0.05) \n    model.compile(optimizer=opt,loss=loss,metrics=[auc, f1])\n    return model\n\n\nmodel = build_model(dim=IMAGE_SIZE,ef=EFFNET)\n#model.summary()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-08-08T16:37:21.623859Z","iopub.execute_input":"2021-08-08T16:37:21.624576Z","iopub.status.idle":"2021-08-08T16:37:35.941126Z","shell.execute_reply.started":"2021-08-08T16:37:21.624509Z","shell.execute_reply":"2021-08-08T16:37:35.940287Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"EFNS = [efn.EfficientNetB0, efn.EfficientNetB1, efn.EfficientNetB2, efn.EfficientNetB3, \n        efn.EfficientNetB4, efn.EfficientNetB5, efn.EfficientNetB6, efn.EfficientNetB7]\n\ndef build_model(dim, ef):\n    \n    model = tf.keras.Sequential([\n        EFNS[ef](\n            input_shape=(*dim, 3),\n            weights='imagenet',\n            include_top=False),\n        tf.keras.layers.GlobalAveragePooling2D(),\n        tf.keras.layers.Dense(4, activation='softmax')\n    ])\n    model.compile(\n        optimizer=tf.keras.optimizers.Adam(),\n        loss='categorical_crossentropy',\n        metrics=[tf.keras.metrics.AUC(multi_label=True)])\n    return model\n\nmodel = build_model(dim=IMAGE_SIZE,ef=EFFNET)","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:35.943042Z","iopub.execute_input":"2021-08-08T16:37:35.943689Z","iopub.status.idle":"2021-08-08T16:37:45.498632Z","shell.execute_reply.started":"2021-08-08T16:37:35.94365Z","shell.execute_reply":"2021-08-08T16:37:45.497797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[Model trained with external data](https://www.kaggle.com/meenakshiramaswamy/siim-covid-img512-extdata-gpu/)","metadata":{}},{"cell_type":"code","source":"# model with extdata\nEFNS = [efn.EfficientNetB0, efn.EfficientNetB1, efn.EfficientNetB2, efn.EfficientNetB3, \n        efn.EfficientNetB4, efn.EfficientNetB5, efn.EfficientNetB6, efn.EfficientNetB7]\n\ndef build_model_extdata(dim, ef):\n    #dim = [*dim, 3]\n    inp = tf.keras.layers.Input(shape=(*dim,3))\n    base = EFNS[ef](input_shape=(*dim,3),weights='imagenet',include_top=False)\n    x = base(inp)\n    \n    CAN  = ChannelAttentionModule()\n    SPN = SpatialAttentionModule()\n    AWG  = AttentionWeightedAverage2D()\n    canx   = CAN(x)*x\n    spnx   = SPN(canx)*canx\n    gapx = tf.keras.layers.GlobalAveragePooling2D()(spnx)\n    wvgx   = tf.keras.layers.GlobalAveragePooling2D()(SPN(canx))\n    avg = tf.keras.layers.Average()([gapx, wvgx])\n    awg = AWG(x)\n    x = tf.keras.layers.Add()([avg, awg])\n    x = tf.keras.layers.BatchNormalization()(x)\n    x = tf.keras.layers.Dense(32, activation = 'relu')(x)\n    \n    \n    #x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    x = tf.keras.layers.Dense(4,activation='sigmoid')(x)\n    auc = tf.keras.metrics.AUC(curve='ROC',\n                               multi_label=True)   \n    f1  = tfa.metrics.F1Score(num_classes=4,average='macro',threshold=None)\n    \n    model = tf.keras.Model(inputs=inp,outputs=x)\n    opt = tf.keras.optimizers.Adam(learning_rate=0.00001)\n    loss = tf.keras.losses.CategoricalCrossentropy(label_smoothing=0.05) \n    #loss = 'sparse_categorical_crossentropy'\n    #crs_ent = tf.keras.metrics.binary_accuracy\n    metrics=[tf.keras.metrics.AUC(multi_label=True)]\n    model.compile(optimizer=opt,loss=loss,metrics=[metrics])\n    return model\n\nmodel_extdata = build_model_extdata(dim=IMAGE_SIZE,ef=EFFNET_META)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-08-08T16:37:45.499851Z","iopub.execute_input":"2021-08-08T16:37:45.500171Z","iopub.status.idle":"2021-08-08T16:37:47.873221Z","shell.execute_reply.started":"2021-08-08T16:37:45.500138Z","shell.execute_reply":"2021-08-08T16:37:47.872182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# TF DATASET FUNCTIONS","metadata":{}},{"cell_type":"code","source":"def build_decoder(path, augment = True, target_size=IMAGE_SIZE, ext='jpg'):\n    \n    file_bytes = tf.io.read_file(path)\n    if ext == 'png':\n        img = tf.image.decode_png(file_bytes, channels=3)\n    elif ext in ['jpg', 'jpeg']:\n        img = tf.image.decode_jpeg(file_bytes, channels=3)\n    else:\n        raise ValueError(\"Image extension not supported\")\n\n    img = tf.cast(img, tf.float32)\n    img = tf.image.resize(img, target_size, method='area')\n\n    img = img/255.0\n\n    \n    if augment:\n        img = transform(img,DIM=dim) if TRANSFORM else img\n        img = tf.image.random_flip_left_right(img)\n        img = tf.image.random_hue(img, 0.01)\n        img = tf.image.random_saturation(img, sat[0], sat[1])\n        img = tf.image.random_contrast(img, cont[0], cont[1])\n        img = tf.image.random_brightness(img, bri)      \n\n    img = tf.reshape(img, [*target_size, 3])\n    \n    return img\n    \n\ndef get_test_dataset(paths, repeat = True,  labels=None):\n    #print (paths)\n    ds = tf.data.Dataset.from_tensor_slices(paths)\n    ds = ds.map(build_decoder, num_parallel_calls=AUTO)\n    #ds = build_decoder(slices)\n    #ds = ds.repeat() if repeat else dset\n    ds = ds.batch(BATCH_SIZE * REPLICAS)\n    ds = ds.prefetch(AUTO)\n    \n    return ds","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:47.874622Z","iopub.execute_input":"2021-08-08T16:37:47.874995Z","iopub.status.idle":"2021-08-08T16:37:47.885799Z","shell.execute_reply.started":"2021-08-08T16:37:47.874956Z","shell.execute_reply":"2021-08-08T16:37:47.884844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#test_img_paths = test_df.image_path.iloc[:10 if not PRIVATE_SUB else test_df.shape[0]]\ntest_ds = get_test_dataset(test_df.image_path.values)\nmodels_lst = glob.glob('../input/siim-covid-img512-ef7-model2-tpu/*.h5')","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:47.887139Z","iopub.execute_input":"2021-08-08T16:37:47.887519Z","iopub.status.idle":"2021-08-08T16:37:48.148954Z","shell.execute_reply.started":"2021-08-08T16:37:47.887481Z","shell.execute_reply":"2021-08-08T16:37:48.148117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Efnet B4 with img size 512 , 5 folds...\n\npred = []    \nfor i, weight in enumerate(models_lst):\n    #print (i, weight)\n    print (f'Predicting test using model {weight}....')\n    model = build_model(IMAGE_SIZE,EFFNET)\n\n    model.load_weights(weight)\n\n    pred = model.predict(test_ds, verbose=1) / len(models_lst)","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:37:48.155063Z","iopub.execute_input":"2021-08-08T16:37:48.155347Z","iopub.status.idle":"2021-08-08T16:39:27.77421Z","shell.execute_reply.started":"2021-08-08T16:37:48.15532Z","shell.execute_reply":"2021-08-08T16:39:27.773108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:27.778232Z","iopub.execute_input":"2021-08-08T16:39:27.778634Z","iopub.status.idle":"2021-08-08T16:39:27.794112Z","shell.execute_reply.started":"2021-08-08T16:39:27.778597Z","shell.execute_reply":"2021-08-08T16:39:27.793089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#test_df[label_cols] = pred\n#test_df.head(10)","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:27.795792Z","iopub.execute_input":"2021-08-08T16:39:27.796462Z","iopub.status.idle":"2021-08-08T16:39:27.804561Z","shell.execute_reply.started":"2021-08-08T16:39:27.796424Z","shell.execute_reply":"2021-08-08T16:39:27.803546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model trained with Extdata by RICORD\nhttps://www.kaggle.com/raddar/ricord-covid19-xray-positive-tests","metadata":{}},{"cell_type":"code","source":"model_lst_extdata = glob.glob('../input/siim-covid-img512-extdata-gpu/*.h5')","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:27.806302Z","iopub.execute_input":"2021-08-08T16:39:27.806895Z","iopub.status.idle":"2021-08-08T16:39:27.83428Z","shell.execute_reply.started":"2021-08-08T16:39:27.80686Z","shell.execute_reply":"2021-08-08T16:39:27.833352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Efnet B0 with img size 512 , 5 folds...\n\npred_extdata = []    \nfor i, weight in enumerate(model_lst_extdata):\n    #print (i, weight)\n    print (f'Predicting test using model trained from extdata {weight}....')\n    model_extdata = build_model_extdata(dim=IMAGE_SIZE,ef=EFFNET_META)  \n\n    model_extdata.load_weights(weight)\n\n    pred_extdata = model_extdata.predict(test_ds, verbose=1) / len(model_lst_extdata)","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:27.836957Z","iopub.execute_input":"2021-08-08T16:39:27.837296Z","iopub.status.idle":"2021-08-08T16:39:51.883678Z","shell.execute_reply.started":"2021-08-08T16:39:27.837248Z","shell.execute_reply":"2021-08-08T16:39:51.882781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_extdata","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:51.888202Z","iopub.execute_input":"2021-08-08T16:39:51.890289Z","iopub.status.idle":"2021-08-08T16:39:51.900843Z","shell.execute_reply.started":"2021-08-08T16:39:51.890213Z","shell.execute_reply":"2021-08-08T16:39:51.899813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_both = 0.65 * pred + 0.35 * pred_extdata\npred_both","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:51.906312Z","iopub.execute_input":"2021-08-08T16:39:51.908639Z","iopub.status.idle":"2021-08-08T16:39:51.920972Z","shell.execute_reply.started":"2021-08-08T16:39:51.908595Z","shell.execute_reply":"2021-08-08T16:39:51.91982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model trained with metadata","metadata":{}},{"cell_type":"code","source":"from keras.models import Sequential\nfrom tensorflow.keras.layers import Dense, Flatten, Dropout, Activation, concatenate\n\nfrom tensorflow.keras.models import Model\nfrom tensorflow.keras.callbacks import ModelCheckpoint, ReduceLROnPlateau, EarlyStopping, Callback","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:51.924921Z","iopub.execute_input":"2021-08-08T16:39:51.927111Z","iopub.status.idle":"2021-08-08T16:39:51.978227Z","shell.execute_reply.started":"2021-08-08T16:39:51.927069Z","shell.execute_reply":"2021-08-08T16:39:51.977139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def meta_model():\n    model = Sequential()\n    model.add(Dense(512, input_shape=(18,), activation='relu'))\n    model.add(Dense(512, activation='relu'))\n    model.add(Dense(512, activation='relu'))\n    model.add(Dense(4, activation='softmax'))\n    model.compile(optimizer='adam',\n                  loss=tf.keras.losses.CategoricalCrossentropy(from_logits=True),\n                  metrics=[\"accuracy\"])\n    return model\n\nmodel = meta_model()\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:51.98264Z","iopub.execute_input":"2021-08-08T16:39:51.984808Z","iopub.status.idle":"2021-08-08T16:39:52.052214Z","shell.execute_reply.started":"2021-08-08T16:39:51.984761Z","shell.execute_reply":"2021-08-08T16:39:52.051348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_lst_meta = glob.glob('../input/siim-covid-tf-metadata-nn/*.h5')\ntest_meta = np.asarray(test_meta)","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:52.053455Z","iopub.execute_input":"2021-08-08T16:39:52.05395Z","iopub.status.idle":"2021-08-08T16:39:52.066935Z","shell.execute_reply.started":"2021-08-08T16:39:52.053912Z","shell.execute_reply":"2021-08-08T16:39:52.066181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"models = []\n\nfor i, model in enumerate(model_lst_meta):\n    models.append(tf.keras.models.load_model(model))\n\ncombined_test_df[label_cols] = sum([model.predict(test_meta, verbose=1) for model in models]) / len(models)","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:52.068221Z","iopub.execute_input":"2021-08-08T16:39:52.068586Z","iopub.status.idle":"2021-08-08T16:39:53.182766Z","shell.execute_reply.started":"2021-08-08T16:39:52.06855Z","shell.execute_reply":"2021-08-08T16:39:53.18199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_meta = combined_test_df[label_cols]\npred_meta","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:53.18395Z","iopub.execute_input":"2021-08-08T16:39:53.184228Z","iopub.status.idle":"2021-08-08T16:39:53.199573Z","shell.execute_reply.started":"2021-08-08T16:39:53.184202Z","shell.execute_reply":"2021-08-08T16:39:53.198401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Weighted the results of extdata trained model prediction with competition data trained model prediction","metadata":{}},{"cell_type":"code","source":"WGT = 0.6\npred_avg = 0.65 * pred + 0.2 * pred_extdata + 0.15 * pred_meta\npred_avg","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:53.200888Z","iopub.execute_input":"2021-08-08T16:39:53.201307Z","iopub.status.idle":"2021-08-08T16:39:53.231969Z","shell.execute_reply.started":"2021-08-08T16:39:53.201247Z","shell.execute_reply":"2021-08-08T16:39:53.231141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check the extdata and our competition data\ntest_ext = test_df.copy() # ext\ntest_comp = test_df.copy() # comp\ntest_both = test_df.copy() # ext + comp\ntest_all_3 = test_df.copy() # + ext + xomp +meta","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:53.233367Z","iopub.execute_input":"2021-08-08T16:39:53.233704Z","iopub.status.idle":"2021-08-08T16:39:53.24109Z","shell.execute_reply.started":"2021-08-08T16:39:53.233669Z","shell.execute_reply":"2021-08-08T16:39:53.240248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_ext[label_cols] = pred_extdata\ntest_comp[label_cols] = pred\n#test_df[label_cols] = pred_avg\ntest_both[label_cols] = pred_both\ntest_all_3[label_cols] = pred_avg # all 3 (60, 20, 20 %)","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:53.242551Z","iopub.execute_input":"2021-08-08T16:39:53.243039Z","iopub.status.idle":"2021-08-08T16:39:53.255275Z","shell.execute_reply.started":"2021-08-08T16:39:53.243002Z","shell.execute_reply":"2021-08-08T16:39:53.254335Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display (test_ext.head(2), test_comp.head(2), test_both.head(2) , test_all_3.head(2))","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:53.256585Z","iopub.execute_input":"2021-08-08T16:39:53.256948Z","iopub.status.idle":"2021-08-08T16:39:53.306237Z","shell.execute_reply.started":"2021-08-08T16:39:53.256913Z","shell.execute_reply":"2021-08-08T16:39:53.305281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#study_df = test_ext.groupby(['study_id'])[label_cols].mean().reset_index()","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:53.307672Z","iopub.execute_input":"2021-08-08T16:39:53.308037Z","iopub.status.idle":"2021-08-08T16:39:53.312025Z","shell.execute_reply.started":"2021-08-08T16:39:53.308Z","shell.execute_reply":"2021-08-08T16:39:53.310848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"study_df = test_both.groupby(['study_id'])[label_cols].mean().reset_index()\nstudy_df.rename(columns={'study_id':'id'}, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:53.313542Z","iopub.execute_input":"2021-08-08T16:39:53.313975Z","iopub.status.idle":"2021-08-08T16:39:53.32791Z","shell.execute_reply.started":"2021-08-08T16:39:53.313932Z","shell.execute_reply":"2021-08-08T16:39:53.326847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_PredictionString(row, thr=0):\n    string = ''\n    for idx in range(4):\n        conf =  row[str(idx)]\n        if conf>thr:\n            string+=f'{label2name[idx]} {conf} 0 0 1 1 '\n    string = string.strip()\n    return string","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:53.329562Z","iopub.execute_input":"2021-08-08T16:39:53.330065Z","iopub.status.idle":"2021-08-08T16:39:53.336746Z","shell.execute_reply.started":"2021-08-08T16:39:53.330027Z","shell.execute_reply":"2021-08-08T16:39:53.335368Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"study_df['PredictionString'] = study_df.apply(get_PredictionString, axis=1)\nstudy_df = study_df.drop(label_cols, axis=1)\n#subm_study_df = subm_study_df.drop('image_path', axis=1)\nstudy_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:53.338414Z","iopub.execute_input":"2021-08-08T16:39:53.339029Z","iopub.status.idle":"2021-08-08T16:39:53.355866Z","shell.execute_reply.started":"2021-08-08T16:39:53.338989Z","shell.execute_reply":"2021-08-08T16:39:53.354687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"study_df.shape, subm_image_df.shape\nsub_df = pd.concat([study_df, subm_image_df])\nsub_df = sub_df.reset_index(drop = True)","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:53.357579Z","iopub.execute_input":"2021-08-08T16:39:53.358009Z","iopub.status.idle":"2021-08-08T16:39:53.365113Z","shell.execute_reply.started":"2021-08-08T16:39:53.357968Z","shell.execute_reply":"2021-08-08T16:39:53.364136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#import copy\n#sub_df1 = sub_df.copy()","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-08-08T16:39:53.366748Z","iopub.execute_input":"2021-08-08T16:39:53.367139Z","iopub.status.idle":"2021-08-08T16:39:53.373907Z","shell.execute_reply.started":"2021-08-08T16:39:53.367099Z","shell.execute_reply":"2021-08-08T16:39:53.372994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#sub_df['id'] = sub_df.apply(lambda row : ( row['id'].split('.')[0]), axis = 1)\n#sub_df.tail()","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-08-08T16:39:53.375672Z","iopub.execute_input":"2021-08-08T16:39:53.376521Z","iopub.status.idle":"2021-08-08T16:39:53.38313Z","shell.execute_reply.started":"2021-08-08T16:39:53.376443Z","shell.execute_reply":"2021-08-08T16:39:53.381627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#shutil.rmtree('./')\n\nshutil.rmtree('./gdcm')\nshutil.rmtree('./resized_test')\nos.remove('./gdcm.tar')","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:53.384805Z","iopub.execute_input":"2021-08-08T16:39:53.385231Z","iopub.status.idle":"2021-08-08T16:39:53.395973Z","shell.execute_reply.started":"2021-08-08T16:39:53.385189Z","shell.execute_reply":"2021-08-08T16:39:53.394932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_df.to_csv('submission.csv', index=False)\nprint (sub_df.shape)\nsub_df","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:53.397691Z","iopub.execute_input":"2021-08-08T16:39:53.39818Z","iopub.status.idle":"2021-08-08T16:39:53.41621Z","shell.execute_reply.started":"2021-08-08T16:39:53.398128Z","shell.execute_reply":"2021-08-08T16:39:53.415203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2021-08-08T16:39:53.417765Z","iopub.execute_input":"2021-08-08T16:39:53.418132Z","iopub.status.idle":"2021-08-08T16:39:54.137237Z","shell.execute_reply.started":"2021-08-08T16:39:53.418095Z","shell.execute_reply":"2021-08-08T16:39:54.136247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}