{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":75176,"databundleVersionId":8252256,"sourceType":"competition"},{"sourceId":182471918,"sourceType":"kernelVersion"}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Unit Circle in Chest X Ray Image","metadata":{}},{"cell_type":"markdown","source":"Thanks to the organizer and the community, you made this awesome learning possible, appreciated very much. This is to share few findings so far during the journey, and propose a method to do domain specific tasks, analogous to developing sawing machine in textile industry.","metadata":{}},{"cell_type":"markdown","source":"# Motivations\n\nWhen reviewing x-ray images, systematic approach has long been established. One of them is ABCDEF framework, which involves reviewing each of the areas in turn: A – Apices and angles; B – Bases and bones; C – Cardiac (behind and within the heart); D – Domes of diaphragm; E – Extra-thoracic; F – Foreign bodies (lines, tubes position)\n\nTraditional methods like edge detection and contour-based approaches can struggle to accurately locate VOI such as heart and bones in X ray images for several reasons, such as noise, low contrast, partial occlusion, pose variation.\n\nThis experiment, let's call it UCCX method, first find an opptimized circle around upper chest, then follow the systematic approach to check target areas one by one, traditional computer vision methods could unleash greater potential, industry knowedge could guild the kernel design of contemporary convolve algorithms for patten recogonition. \n\nWhen a computer performs tasks following industrial best practices, it's more likely to get FDA clearance and clinical validation, then its 1000-fold efficiency might make a difference in reality.","metadata":{}},{"cell_type":"markdown","source":"# Experiments","metadata":{}},{"cell_type":"markdown","source":"## setup environment\nTo run snippets, add one more input into Kaggle Notebook.  <br>\nIn Kaggle Notedbook Edit mode, in the upper right area, click \"Add input\", search for https://www.kaggle.com/code/jeff271/chest-x-ray-2024, click \"add\". <br>\nthen run this command, make <b>util</b> folder available in your working directory.","metadata":{}},{"cell_type":"code","source":"!ln -s /kaggle/input/chest-x-ray-2024/util /kaggle/working/util 2> /dev/null","metadata":{"execution":{"iopub.status.busy":"2024-06-10T02:19:20.085523Z","iopub.execute_input":"2024-06-10T02:19:20.086455Z","iopub.status.idle":"2024-06-10T02:19:21.266370Z","shell.execute_reply.started":"2024-06-10T02:19:20.086419Z","shell.execute_reply":"2024-06-10T02:19:21.264831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport sqlite3\nimport cv2\nfrom util.unit_circle import norm_xray,keypoints,get_hist\nimport matplotlib.pyplot as plt\nfrom matplotlib.patches import Rectangle,Circle\n%matplotlib inline\ndata_path = '/kaggle/input/amia-public-challenge-2024'\nDB = '/kaggle/working/util/amia.db'","metadata":{"execution":{"iopub.status.busy":"2024-06-10T02:19:28.033360Z","iopub.execute_input":"2024-06-10T02:19:28.034003Z","iopub.status.idle":"2024-06-10T02:19:28.044054Z","shell.execute_reply.started":"2024-06-10T02:19:28.033961Z","shell.execute_reply":"2024-06-10T02:19:28.042686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# visual demostration","metadata":{}},{"cell_type":"code","source":"# voi box [x1,y1,x2,y2], relative to center of unit circle \nbox_targets={} \nbox_targets['trachea']      = -0.6,-1.1,0.6,0.1\nbox_targets['left heart']   = -1.2,-0.2,0.1,1.2\nbox_targets['right heart']  = -0.1,-0.5,1.1,1.2\nbox_targets['heart']        = -0.9,-0.7,1.1,1.3","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def show(dk,v):\n    # random select 5 images. check Appendix-1 for bit more details.\n    con = sqlite3.connect(DB)\n    sql = 'SELECT t.image_id,s.dim0 as height,s.dim1 as width from \\\n    (select image_id from train_csv order by random() limit 3) as t join img_size_csv s on t.image_id=s.image_id'\n    ds = pd.read_sql(sql,con)\n    con.close()    \n    \n    for ix,row in ds.iterrows():\n        fn = os.path.join(data_path,'train','train',row.image_id+'.png')\n        gray = cv2.imread(fn,cv2.IMREAD_GRAYSCALE)\n\n        print(\"voi:\",dk,f\"sample {ix}\",fn)\n        \n        # restore height width ratio\n        h,w = gray.shape;     hw = row.height/row.width;     h_,w_ = h,int(w*hw)    \n        gray = cv2.resize(gray,(h_,w_),interpolation=cv2.INTER_AREA)\n        # resize to 128\n        k = 128./max(h_,w_); h_,w_ = int(h_*k),int(w_*k)\n        img = cv2.resize(gray,(h_,w_),interpolation=cv2.INTER_AREA)\n        \n        # deal with low contrast and noise challenges\n        normed_img = norm_xray(img) \n        \n        # get the anchor for further explorations in target areas\n        xc,yc,rc = keypoints(normed_img)\n    \n        # target area, ref to the unit circle (xc,yc,rc)\n        box_target = xc+rc*v[0],yc+rc*v[1],xc+rc*v[2],yc+rc*v[3]\n        # scale back\n        box = x1,y1,x2,y2 = [int(x/k) for x in box_target]\n        img_target_fullsize = gray[y1:y2,x1:x2]\n        \n        # resize to 128 for norm\n        h_,w_ = img_target_fullsize.shape\n        k = 128./max(h_,w_); h_,w_ = int(h_*k),int(w_*k)\n        img_target = cv2.resize(img_target_fullsize,(w_,h_),interpolation=cv2.INTER_AREA)\n        normed_tgt = norm_xray(img_target) \n        \n        # center in img_target and normed_tgt\n        tgt_x,tgt_y = -int(np.round(img_target.shape[1]/(v[2]-v[0])*v[0])),-int(np.round(img_target.shape[0]/(v[3]-v[1])*v[1]))\n        \n        _,otsu_target = cv2.threshold(img_target, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)\n    \n        # visualize\n        fig,ax = plt.subplots(2,5,figsize=(11,3.6),gridspec_kw={'height_ratios':[2,1]})\n        \n        for i,mx in enumerate([img,normed_img,img_target,normed_tgt,otsu_target]):    \n            ax[0,i].imshow(mx,cmap='gray')\n            idx,pdf,cdf=get_hist(mx)\n            ax[1,i].plot(idx,pdf); ax[1,i].get_yaxis().set_visible(False)\n            aw=ax[1,i].twinx(); aw.plot(idx,cdf,'--'); aw.get_yaxis().set_visible(False)\n    \n            if i<2:\n                box_target\n                ax[0,i].scatter(xc,yc,c='r')\n                ax[0,i].add_patch( Circle((xc,yc),rc,color='r',fill=False))\n                x1,y1,x2,y2 = list(map(int,box_target))\n                ax[0,i].add_patch( Rectangle((x1,y1),x2-x1,y2-y1,fc='none',color='red',linewidth = 1, linestyle='dotted') )\n            elif i<5:\n                ax[0,i].scatter(tgt_x,tgt_y,c='r')\n                \n        plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-06-10T02:35:43.701742Z","iopub.execute_input":"2024-06-10T02:35:43.702212Z","iopub.status.idle":"2024-06-10T02:35:43.725022Z","shell.execute_reply.started":"2024-06-10T02:35:43.702178Z","shell.execute_reply":"2024-06-10T02:35:43.723833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# target: trachea","metadata":{}},{"cell_type":"code","source":"show(k:='trachea',box_targets[k])","metadata":{"execution":{"iopub.status.busy":"2024-06-10T02:35:49.145028Z","iopub.execute_input":"2024-06-10T02:35:49.145441Z","iopub.status.idle":"2024-06-10T02:36:06.681831Z","shell.execute_reply.started":"2024-06-10T02:35:49.145412Z","shell.execute_reply":"2024-06-10T02:36:06.680700Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# target: left heart","metadata":{}},{"cell_type":"code","source":"show(k:='left heart',box_targets[k])","metadata":{"execution":{"iopub.status.busy":"2024-06-10T02:36:17.768735Z","iopub.execute_input":"2024-06-10T02:36:17.769138Z","iopub.status.idle":"2024-06-10T02:36:35.322636Z","shell.execute_reply.started":"2024-06-10T02:36:17.769106Z","shell.execute_reply":"2024-06-10T02:36:35.321544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# target: right heart","metadata":{}},{"cell_type":"code","source":"show(k:='right heart',box_targets[k])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# target: heart","metadata":{}},{"cell_type":"code","source":"show(k:='heart',box_targets[k])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# clustering","metadata":{}},{"cell_type":"markdown","source":"## heart size","metadata":{}},{"cell_type":"markdown","source":"### ...","metadata":{}},{"cell_type":"markdown","source":"# play for fun","metadata":{}},{"cell_type":"markdown","source":"### polynomial equation for the shape of heart","metadata":{}},{"cell_type":"markdown","source":"# Findings - Todo\nThis is draft, far from completion...","metadata":{}},{"cell_type":"markdown","source":"# Appendix","metadata":{}},{"cell_type":"markdown","source":"### snippets to create amia.db","metadata":{}},{"cell_type":"code","source":"# inital attempts to explore the datasets. a lot unnecessary tries.\ndef init_db():\n    # what image files are there\n    lst=[]\n    for dirpath, dnames, fnames in os.walk(os.path.join(data_path,'train','train')):\n        for f in fnames:\n            if f.endswith('.png'):\n                im = Image.open(os.path.join(data_path,'train','train',f))\n                width, height = im.size\n                im = np.array(im,np.float32)\n                m,n,s,k = np.mean(im),np.median(im),np.std(im),skew(im.reshape(-1))\n                lst += [(f[:-4],width,height,m,n,s,k)]\n    ds = pd.DataFrame(lst)\n    ds.columns=['img_name','width','height','mean','median','std','skew']\n    ds.to_sql('train_img_names',con,if_exists='replace')\n\n    lst=[]\n    for dirpath, dnames, fnames in os.walk(os.path.join(data_path,'test','test')):\n        for f in fnames:\n            if f.endswith('.png'):\n                im = Image.open(os.path.join(data_path,'test','test',f))\n                width, height = im.size\n                im = np.array(im,np.float32)\n                m,n,s,k = np.mean(im),np.median(im),np.std(im),skew(im.reshape(-1))\n                lst += [(f[:-4],width,height,m,n,s,k)]\n    ds = pd.DataFrame(lst)\n    ds.columns=['img_name','width','height','mean','median','std','skew']\n    ds.to_sql('test_img_names',con,if_exists='replace')\n\n    # csv to db.\n    for f in ['train','test','img_size']:\n        csv = pd.read_csv(os.path.join(data_path,f+'.csv'))\n        csv.to_sql(f+'_csv',con,if_exists='replace')","metadata":{},"execution_count":null,"outputs":[]}]}