{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style=\"text-align: center; font-family: Verdana; font-size: 32px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; font-variant: small-caps; letter-spacing: 3px; color: #468282; background-color: #ffffff;\">HuBMAP + HPA - Hacking the Human Body</h1>\n<h2 style=\"text-align: center; font-family: Verdana; font-size: 24px; font-style: normal; font-weight: bold; text-decoration: underline; text-transform: none; letter-spacing: 2px; color: navy; background-color: #ffffff;\">Segment multi-organ functional tissue units</h2>\n\n<img src=\"https://storage.googleapis.com/kaggle-competitions/kaggle/34547/logos/header.png\"> \n\n> # 📌Introduction: \n>> A lot of segmentation competition is getting hosted in kaggle recently. This competition is very similar to the previous segmentation competitions, eg: [Sartorius CIS](https://www.kaggle.com/competitions/sartorius-cell-instance-segmentation/overview), and ongoing [UW-Madison GI Tract Image Segmentation](https://www.kaggle.com/competitions/uw-madison-gi-tract-image-segmentation) also felt similar [but no so much], from the POV of medical data, segmentation problem statement. You might start with different models used in those above competitions, and learn a few frameworks used like MMdet, detectron2 etc. The evaluation metric is mean Dice coefficient, don't really know why mean?[maybe because mean over all the segments found in a single image]. But other than that Dice coefficient is a very popular metric for image segmentation. if you don't know, you might want to check out this [NB](https://www.kaggle.com/code/yerramvarun/understanding-dice-coefficient). Use models like Unet, SegNet, Enet etc. type encoder decoder model for baseline, then move on to more complex models and pre-processing and post-processing techniques. Follow the augmentations used in the previous competitions, and do trial and error for fitting those augmentations to the model, or come up with some new one. \n\n>> Data for this competition comes from two different consortiums, the Human Protein Atlas (HPA) and Human BioMolecular Atlas Program (HUBMAP). As mentioned in the Data tab, one of the main challenges of this competition will be adapting models to function properly when presented with data collected using a different protocol. Because among the three datasets, the training set contains data from public HPAs, the public test set is a combination of private HPAs and HuBMAP data, and the private test set contains only HuBMAP data. Image resolution is high, though the number of training images are very small(351).\n\n>> I plan to publish three different parts which will include data prep, model training, model inference. Its a very basic version of the NB, will try to improve over time.","metadata":{}},{"cell_type":"markdown","source":"\n\n<h3 style=\"text-align:center; background-color:#C8FF33;padding:40px;border-radius: 30px;\">\n<p style=\"text-align: left\">Also see these notebooks:</p>\n    <p style=\"text-align: left\"><b>* HuBMAP+HPA multiOrgan Segmentation 1/3 [data prep]</b></p>\n    <p style=\"text-align: left\"><b><a href=\"https://www.kaggle.com/code/soumya9977/hubmap-vanilla-unet-w-b-pytorch-2-3-train\"> &nbsp; HuBMAP: Vanilla Unet + W&B + Pytorch 2/3 [train]</a></b></p>\n    <p style=\"text-align: left\"><b><a href=\"https://www.kaggle.com/code/soumya9977/hubmap-vanilla-unet-pytorch-3-3-inference\"> &nbsp; HuBMAP: Vanilla Unet Pytorch 3/3 [inference]</a></b></p>\n</h3>\n\n## Please _DO_ upvote!","metadata":{}},{"cell_type":"markdown","source":"# Imports","metadata":{}},{"cell_type":"code","source":"!pip install -q tifffile","metadata":{"execution":{"iopub.status.busy":"2022-07-23T13:43:00.803508Z","iopub.execute_input":"2022-07-23T13:43:00.804340Z","iopub.status.idle":"2022-07-23T13:43:14.270706Z","shell.execute_reply.started":"2022-07-23T13:43:00.804216Z","shell.execute_reply":"2022-07-23T13:43:14.269408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nimport cv2\nimport matplotlib.pyplot as plt\nimport json\nfrom tqdm import tqdm\nimport random\nfrom tifffile import imread","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-23T13:43:14.272586Z","iopub.execute_input":"2022-07-23T13:43:14.272997Z","iopub.status.idle":"2022-07-23T13:43:14.667288Z","shell.execute_reply.started":"2022-07-23T13:43:14.272958Z","shell.execute_reply":"2022-07-23T13:43:14.666208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DIR = \"../input/hubmap-organ-segmentation\"\ntrain_df = pd.read_csv(os.path.join(DIR,\"train.csv\"))\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T13:43:14.668816Z","iopub.execute_input":"2022-07-23T13:43:14.669941Z","iopub.status.idle":"2022-07-23T13:43:15.018841Z","shell.execute_reply.started":"2022-07-23T13:43:14.669901Z","shell.execute_reply":"2022-07-23T13:43:15.017715Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T13:43:15.021482Z","iopub.execute_input":"2022-07-23T13:43:15.021857Z","iopub.status.idle":"2022-07-23T13:43:15.053205Z","shell.execute_reply.started":"2022-07-23T13:43:15.021824Z","shell.execute_reply":"2022-07-23T13:43:15.051746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15,5))\ntrain_df[\"organ\"].value_counts().plot(kind='bar', color='green')","metadata":{"execution":{"iopub.status.busy":"2022-07-23T13:43:15.055527Z","iopub.execute_input":"2022-07-23T13:43:15.056093Z","iopub.status.idle":"2022-07-23T13:43:15.303922Z","shell.execute_reply.started":"2022-07-23T13:43:15.056040Z","shell.execute_reply":"2022-07-23T13:43:15.302670Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# source: https://www.kaggle.com/code/julian3833/sartorius-starter-torch-mask-r-cnn-lb-0-273 w/ a bit change\ndef rle_decode(mask_rle, shape, color=1):\n    '''\n    mask_rle: run-length as string formated (start length)\n    shape: (height,width) of array to return \n    Returns numpy array, 1 - mask, 0 - background\n    '''\n    s = mask_rle.split()\n    starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\n    starts -= 1\n    ends = starts + lengths\n    img = np.zeros(shape[0] * shape[1], dtype=np.float32)\n    for lo, hi in zip(starts, ends):\n        img[lo : hi] = color\n    return img.reshape(shape).T","metadata":{"execution":{"iopub.status.busy":"2022-07-23T13:43:15.305920Z","iopub.execute_input":"2022-07-23T13:43:15.306667Z","iopub.status.idle":"2022-07-23T13:43:15.316825Z","shell.execute_reply.started":"2022-07-23T13:43:15.306617Z","shell.execute_reply":"2022-07-23T13:43:15.315750Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Checking the training data:","metadata":{}},{"cell_type":"code","source":"# train_df[\"rle\"].iloc[0]\n\nfor i in np.random.choice(200,5):\n    rle_img = rle_decode(train_df[\"rle\"].iloc[i],(train_df[\"img_height\"].iloc[i],train_df[\"img_width\"].iloc[i]))\n    img_dir = os.path.join(\"../input/hubmap-organ-segmentation/train_images\" , str(train_df[\"id\"].iloc[i]) + '.tiff')\n#     print(img_dir)\n    img = plt.imread(img_dir)\n\n    plt.figure(figsize=(16,18))\n    plt.subplot(1,2,1)\n    plt.imshow(img)\n    plt.title(f\"id: {i} image\")\n\n    plt.subplot(1,2,2)\n    plt.imshow(rle_img)\n    plt.title(f\"id: {i} mask\");","metadata":{"execution":{"iopub.status.busy":"2022-07-23T13:43:15.318636Z","iopub.execute_input":"2022-07-23T13:43:15.319405Z","iopub.status.idle":"2022-07-23T13:43:29.769475Z","shell.execute_reply.started":"2022-07-23T13:43:15.319372Z","shell.execute_reply":"2022-07-23T13:43:29.768438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Comparing `train_annotations` folder data with RLE: ","metadata":{}},{"cell_type":"markdown","source":"- Reading the RLE from `.csv` and creating the mask.","metadata":{}},{"cell_type":"code","source":"i = 0\nrle_img = rle_decode(train_df[\"rle\"].iloc[i],(train_df[\"img_height\"].iloc[i],train_df[\"img_width\"].iloc[i]))\nimg_dir = os.path.join(\"../input/hubmap-organ-segmentation/train_images\" , str(train_df[\"id\"].iloc[i]) + '.tiff')\n#     print(img_dir)\nimg = plt.imread(img_dir)\n\nplt.figure(figsize=(16,18))\nplt.subplot(1,2,1)\nplt.imshow(img)\n\nplt.subplot(1,2,2)\nplt.imshow(rle_img);","metadata":{"execution":{"iopub.status.busy":"2022-07-23T13:43:29.770844Z","iopub.execute_input":"2022-07-23T13:43:29.771174Z","iopub.status.idle":"2022-07-23T13:43:32.652050Z","shell.execute_reply.started":"2022-07-23T13:43:29.771144Z","shell.execute_reply":"2022-07-23T13:43:32.650309Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Reading the polygon ploints from json.","metadata":{}},{"cell_type":"code","source":"with open(\"../input/hubmap-organ-segmentation/train_annotations/10044.json\") as rle_json:\n    data = json.load(rle_json)\n    \nprint(data.__len__())","metadata":{"execution":{"iopub.status.busy":"2022-07-23T13:43:32.653983Z","iopub.execute_input":"2022-07-23T13:43:32.654924Z","iopub.status.idle":"2022-07-23T13:43:32.672005Z","shell.execute_reply.started":"2022-07-23T13:43:32.654861Z","shell.execute_reply":"2022-07-23T13:43:32.670723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image = np.zeros((3000,3000))\nfor i in range(len(data)):\n    image = cv2.fillPoly(image, pts = [np.array(data[i])], color =(255,255,255))\n\nplt.figure(figsize=(16,18))\nplt.subplot(1,2,1)\nplt.imshow(img)\n\nplt.subplot(1,2,2)\nplt.imshow(image);","metadata":{"execution":{"iopub.status.busy":"2022-07-23T13:43:32.677123Z","iopub.execute_input":"2022-07-23T13:43:32.677599Z","iopub.status.idle":"2022-07-23T13:43:35.092831Z","shell.execute_reply.started":"2022-07-23T13:43:32.677562Z","shell.execute_reply":"2022-07-23T13:43:35.091407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Both are same","metadata":{}},{"cell_type":"markdown","source":"## Overlay of Mask for Different Class:","metadata":{}},{"cell_type":"code","source":"folder1 = \"/kaggle/working/train_masks\"\nfolder2 = \"/kaggle/working/train_masks_np\"\n\nif not os.path.isdir(folder1):\n    os.mkdir(folder1)\n    \nif not os.path.isdir(folder2):\n    os.mkdir(folder2)\n    \n    \nfor i in tqdm(range(len(train_df))):\n    rle_img = rle_decode(train_df[\"rle\"].iloc[i],(train_df[\"img_height\"].iloc[i],train_df[\"img_width\"].iloc[i]))\n    f_name1 = os.path.join(folder1, str(train_df[\"id\"].iloc[i])+'.png')\n    f_name2 = os.path.join(folder2, str(train_df[\"id\"].iloc[i])+'.npy')\n\n    cv2.imwrite(f_name1,rle_img)\n    np.save(f_name2, rle_img)\n    ","metadata":{"execution":{"iopub.status.busy":"2022-07-23T13:43:35.094726Z","iopub.execute_input":"2022-07-23T13:43:35.095214Z","iopub.status.idle":"2022-07-23T13:44:41.577697Z","shell.execute_reply.started":"2022-07-23T13:43:35.095165Z","shell.execute_reply":"2022-07-23T13:44:41.576636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Saving the masks in folder:","metadata":{}},{"cell_type":"code","source":"TRAIN_DIR = DIR + \"/train_images/\"\nfunc = lambda x: TRAIN_DIR + str(x) + \".tiff\"\ntrain_df[\"img_path\"] = train_df[\"id\"].apply(func)\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T13:44:41.579102Z","iopub.execute_input":"2022-07-23T13:44:41.579530Z","iopub.status.idle":"2022-07-23T13:44:41.602916Z","shell.execute_reply.started":"2022-07-23T13:44:41.579496Z","shell.execute_reply":"2022-07-23T13:44:41.602051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"kidney_files = random.choices(train_df[train_df[\"organ\"]=='kidney']['img_path'].to_list(),k=1)\nprostate_files = random.choices(train_df[train_df[\"organ\"]=='prostate']['img_path'].to_list(),k=1)\nlargeint_files = random.choices(train_df[train_df[\"organ\"]=='largeintestine']['img_path'].to_list(),k=1)\nsplieen_files = random.choices(train_df[train_df[\"organ\"]=='spleen']['img_path'].to_list(),k=1)\nlung_files = random.choices(train_df[train_df[\"organ\"]=='lung']['img_path'].to_list(),k=1)\n\norgan_list = np.unique(train_df[\"organ\"]).tolist()\nmain_list = [kidney_files,prostate_files, largeint_files, splieen_files, lung_files,]","metadata":{"execution":{"iopub.status.busy":"2022-07-23T13:44:41.604493Z","iopub.execute_input":"2022-07-23T13:44:41.604891Z","iopub.status.idle":"2022-07-23T13:44:42.698308Z","shell.execute_reply.started":"2022-07-23T13:44:41.604859Z","shell.execute_reply":"2022-07-23T13:44:42.696634Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# lung_files[0].split('/')[-1].split('.')[0]\nMASK_DIR = \"./train_masks_np/\"\n\nfor i in range(5):\n#     print(main_list[i][0])\n    img = cv2.resize(imread(main_list[i][0]),(3000,3000))\n    mask_file = MASK_DIR + main_list[i][0].split('/')[-1].split('.')[0]+'.npy'\n#     print(mask_file)\n    mask = np.resize(np.load(mask_file),(3000,3000))\n    mask = mask.reshape(mask.shape[0],mask.shape[1],1)\n    zero_mask = np.zeros((3000,3000,3))\n\n    zero_mask[:,:,0] = mask[:,:,0]\n    zero_mask[:,:,1] = mask[:,:,0]\n    zero_mask[:,:,2] = mask[:,:,0]\n    \n    zero_mask[:,:,0] = zero_mask[:,:,0] * 0\n    zero_mask[:,:,1] = zero_mask[:,:,1] * 255 \n    zero_mask[:,:,2] = zero_mask[:,:,2] * 0 \n\n    zero_mask = zero_mask.astype(np.uint8)\n    \n    combo = cv2.addWeighted(img, 0.7, zero_mask, 0.3, 0.0)\n    \n    plt.figure(figsize=(14,16))\n    \n    ax1 = plt.subplot(1,3,1)\n    plt.imshow(img)\n    ax1.set_title(f\"{organ_list[i]}: image\")\n    \n    \n    ax2 = plt.subplot(1,3,2)\n    plt.imshow(zero_mask)\n    ax2.set_title(f\"{organ_list[i]}: mask\")\n    \n    ax3 = plt.subplot(1,3,3)\n    plt.imshow(combo)\n    ax3.set_title(f\"{organ_list[i]}: overlay\")\n    \n    plt.show()\n#     break","metadata":{"execution":{"iopub.status.busy":"2022-07-23T13:44:42.700144Z","iopub.execute_input":"2022-07-23T13:44:42.700907Z","iopub.status.idle":"2022-07-23T13:45:05.868484Z","shell.execute_reply.started":"2022-07-23T13:44:42.700865Z","shell.execute_reply":"2022-07-23T13:45:05.867242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.imshow(np.load(os.path.join(folder2,\"10610.npy\")))","metadata":{"execution":{"iopub.status.busy":"2022-07-23T13:45:05.870004Z","iopub.execute_input":"2022-07-23T13:45:05.870459Z","iopub.status.idle":"2022-07-23T13:45:07.096265Z","shell.execute_reply.started":"2022-07-23T13:45:05.870424Z","shell.execute_reply":"2022-07-23T13:45:07.095391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.imshow(plt.imread(os.path.join(folder1,\"10610.png\")))","metadata":{"execution":{"iopub.status.busy":"2022-07-23T13:45:07.097362Z","iopub.execute_input":"2022-07-23T13:45:07.098143Z","iopub.status.idle":"2022-07-23T13:45:08.166594Z","shell.execute_reply.started":"2022-07-23T13:45:07.098102Z","shell.execute_reply":"2022-07-23T13:45:08.165451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"os.listdir(\"./train_masks_np\").__len__(), os.listdir(\"./train_masks\").__len__()","metadata":{"execution":{"iopub.status.busy":"2022-07-23T13:45:08.168489Z","iopub.execute_input":"2022-07-23T13:45:08.168857Z","iopub.status.idle":"2022-07-23T13:45:08.177210Z","shell.execute_reply.started":"2022-07-23T13:45:08.168824Z","shell.execute_reply":"2022-07-23T13:45:08.175971Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> # ⭕ WORK IN PROGRESS ! ! !\n<p align=\"center\">\n<img src=\"https://media.giphy.com/media/xThuWu82QD3pj4wvEQ/giphy.gif\" width=\"300\">\n</p>","metadata":{}}]}