{"metadata":{"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":52279,"databundleVersionId":5822112,"sourceType":"competition"},{"sourceId":8380012,"sourceType":"datasetVersion","datasetId":4983219},{"sourceId":8558800,"sourceType":"datasetVersion","datasetId":5115459}],"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.10.13"},"papermill":{"default_parameters":{},"duration":27079.611541,"end_time":"2024-05-19T23:16:05.498967","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2024-05-19T15:44:45.887426","version":"2.5.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<a id=toc></a>\n<h1 style=\"padding: 35px;color:white;margin:10;font-size:200%;text-align:center;display:fill;border-radius:10px;overflow:hidden;background-image: url(https://i.postimg.cc/j2bBmHWx/Py-Torch-Gradient.jpg); background-size: 100% auto;background-position: 0px 0px; \n\"><span style='color:white'><b>UNET| Blood Vessels in Kidney Semantic Segmentation</b></span></h1>\n\n<br>\n\n<center>\n    <figure>\n        <img src=\"https://i.pinimg.com/originals/a4/7b/9a/a47b9a5a523cb4d49f79e2d73fee636d.gif\" alt =\"Human\" style='width:70%;'>\n    </figure>\n</center>\n\n<br>\n\n## 🎯 Introduction\nThis is a notebook for a past Kaggle competition [**HuBMAP - Hacking the Human Vasculature**.](https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature) The goal of this competition is to detect Blood Vessels from images of kidney tissue taken by a microscope, and the detection mask shall have [IoU (Intersection over Union)](https://learnopencv.com/intersection-over-union-iou-in-object-detection-and-segmentation/) greater than 0.6. The Kaggle score is calculated by Average Precision Over Confidence which is the same as [Open Images 2019 - Instance Segmentation](https://www.kaggle.com/c/open-images-2019-instance-segmentation/overview/evaluation). \n\nThis HuBMAP competition seems to be difficult for a few reasons. First, it looks almost impossible to correctly identify blood vessels by non-experts(refer to EDA). It would be hard for deep learning too. Second, only 1633 label data are given for images which is insufficient for such a complex segmentation. Another reason is an imbalanced label: only 5% of image pixels are positive. Therefore, I defined the following targets for this project. \n\n**Target**\n\n* To develop an effective training loop which includes **data augmentation** and **custom loss function**\n\n<br>\n\n<hr>\n","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.execute_input":"2024-05-16T18:47:11.764973Z","iopub.status.busy":"2024-05-16T18:47:11.764175Z","iopub.status.idle":"2024-05-16T18:47:27.231138Z","shell.execute_reply":"2024-05-16T18:47:27.230241Z","shell.execute_reply.started":"2024-05-16T18:47:11.764943Z"},"papermill":{"duration":0.029291,"end_time":"2024-05-19T15:44:48.594406","exception":false,"start_time":"2024-05-19T15:44:48.565115","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<center><div style='color:#ffffff;\n           display:inline-block;\n           padding: 5px 5px 5px 5px;\n           border-radius:5px;\n           background-color:#78D1E1;\n           font-size:100%;'><a href=#toc style='text-decoration: none; color:#03001C;'>⬆️ Back To Top</a></div></center>\n\n<a id='1'></a>\n# 1 | Import important libraries\n<div style=\"padding: 4px;color:white;margin:10;font-size:200%;text-align:center;display:fill;border-radius:10px;overflow:hidden;background-image: url(https://i.postimg.cc/j2bBmHWx/Py-Torch-Gradient.jpg); background-size: 100% auto;\"></div>\n\n<br>","metadata":{"papermill":{"duration":0.018542,"end_time":"2024-05-19T15:44:48.635600","exception":false,"start_time":"2024-05-19T15:44:48.617058","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\nimport cv2\nimport seaborn as sns\nimport matplotlib .pyplot as plt\nimport gc\nimport time\nimport math\nimport json\nfrom tensorflow.keras.utils import plot_model\nimport tensorflow as tf\nfrom keras import backend as K\nfrom tensorflow.keras import layers\nfrom tensorflow.keras import Model\nfrom tensorflow import keras\nimport warnings\nimport random\nwarnings.filterwarnings(\"ignore\")","metadata":{"papermill":{"duration":12.908352,"end_time":"2024-05-19T15:45:01.563140","exception":false,"start_time":"2024-05-19T15:44:48.654788","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:12:49.759251Z","iopub.execute_input":"2024-05-31T09:12:49.759631Z","iopub.status.idle":"2024-05-31T09:13:07.136179Z","shell.execute_reply.started":"2024-05-31T09:12:49.759588Z","shell.execute_reply":"2024-05-31T09:13:07.135044Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<center><div style='color:#ffffff;\n           display:inline-block;\n           padding: 5px 5px 5px 5px;\n           border-radius:5px;\n           background-color:#78D1E1;\n           font-size:100%;'><a href=#toc style='text-decoration: none; color:#03001C;'>⬆️ Back To Top</a></div></center>\n\n<a id='1'></a>\n# 2 | Data Exploration\n<div style=\"padding: 4px;color:white;margin:10;font-size:200%;text-align:center;display:fill;border-radius:10px;overflow:hidden;background-image: url(https://i.postimg.cc/j2bBmHWx/Py-Torch-Gradient.jpg); background-size: 100% auto;\"></div>\n\n<br>\n<h3>Data Source</h3>\n\nhttps://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data","metadata":{"papermill":{"duration":0.018564,"end_time":"2024-05-19T15:45:01.600565","exception":false,"start_time":"2024-05-19T15:45:01.582001","status":"completed"},"tags":[]}},{"cell_type":"code","source":"masks_dir = \"/kaggle/input/hubmap-human-vasculature-dataset-512512/HuPMap/masks/\"\nimages_dir = \"/kaggle/input/hubmap-human-vasculature-dataset-512512/HuPMap/images/\"\ntiles_fpath = \"/kaggle/input/hubmap-human-vasculature-dataset-512512/HuPMap/kidney_tiles.csv\"\ntest_folder = \"/kaggle/input/hubmap-hacking-the-human-vasculature/test/\"","metadata":{"papermill":{"duration":0.026142,"end_time":"2024-05-19T15:45:01.645357","exception":false,"start_time":"2024-05-19T15:45:01.619215","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:13:07.137640Z","iopub.execute_input":"2024-05-31T09:13:07.138296Z","iopub.status.idle":"2024-05-31T09:13:07.143685Z","shell.execute_reply.started":"2024-05-31T09:13:07.138265Z","shell.execute_reply":"2024-05-31T09:13:07.142456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"kidney_df = pd.read_csv(tiles_fpath)\nprint(kidney_df.shape)\nkidney_df.head()","metadata":{"papermill":{"duration":0.080442,"end_time":"2024-05-19T15:45:01.744328","exception":false,"start_time":"2024-05-19T15:45:01.663886","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:13:41.048961Z","iopub.execute_input":"2024-05-31T09:13:41.049416Z","iopub.status.idle":"2024-05-31T09:13:41.112578Z","shell.execute_reply.started":"2024-05-31T09:13:41.049384Z","shell.execute_reply":"2024-05-31T09:13:41.111377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ignore blank masks, and select some features\nkidney_df = kidney_df[kidney_df['annotated'] == 1]\nkidney_df = kidney_df[(kidney_df['blood_vessel'] > 0) | (kidney_df['unsure'] > 0)]\nkidney_df = kidney_df[['id', 'source_wsi', 'dataset', 'dataset_wsi', 'blood_vessel', 'glomerulus', 'unsure']]\n\nprint(kidney_df.shape)\nkidney_df.head()","metadata":{"papermill":{"duration":0.043153,"end_time":"2024-05-19T15:45:01.806527","exception":false,"start_time":"2024-05-19T15:45:01.763374","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:13:45.924626Z","iopub.execute_input":"2024-05-31T09:13:45.925684Z","iopub.status.idle":"2024-05-31T09:13:45.952577Z","shell.execute_reply.started":"2024-05-31T09:13:45.925640Z","shell.execute_reply":"2024-05-31T09:13:45.951282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"time1 = time.time()\nL = 512\nNP = len(kidney_df)  # Assuming NP is the number of images in your dataframe\nX_train = np.zeros((NP, L, L, 3), dtype=np.uint8)\nfor i in range(NP):\n    # Load images from NumPy files\n    row = kidney_df.iloc[i]\n    img_path = images_dir + row['id'] + \".npy\"  # Assuming 'id' is the column containing image filenames\n    img = np.load(img_path)  # Load the image using np.load()\n    X_train[i, :, :, :] = img\n    \ntime2 = time.time()\ntime3 = np.round(time2 - time1)\nprint(time3, \"sec\")\n","metadata":{"papermill":{"duration":18.54549,"end_time":"2024-05-19T15:45:20.371241","exception":false,"start_time":"2024-05-19T15:45:01.825751","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:13:52.814233Z","iopub.execute_input":"2024-05-31T09:13:52.814652Z","iopub.status.idle":"2024-05-31T09:14:18.507624Z","shell.execute_reply.started":"2024-05-31T09:13:52.814620Z","shell.execute_reply":"2024-05-31T09:14:18.506405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def read_mask(path, channel=2):\n    # loading mask\n    mask = np.load(path)\n    \n    # Select the specified channels\n    if channel == 1:\n        mask = mask[:, :, 0]\n    elif channel == 2:\n        # Sum the pixel values of selected channels along the last axis (channel axis)\n        selected_channels = mask[:, :, [0, 2]]\n        mask = np.sum(selected_channels, axis=2)\n    else:\n        pass\n    \n    # expanding dimension if needed\n    if len(mask.shape) != 3:\n        mask = np.expand_dims(mask, axis=-1)\n        mask = np.where(mask > 0, 1, 0).astype(np.uint8)\n        \n    return mask","metadata":{"papermill":{"duration":0.028314,"end_time":"2024-05-19T15:45:20.419242","exception":false,"start_time":"2024-05-19T15:45:20.390928","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:14:18.509486Z","iopub.execute_input":"2024-05-31T09:14:18.509852Z","iopub.status.idle":"2024-05-31T09:14:18.516284Z","shell.execute_reply.started":"2024-05-31T09:14:18.509824Z","shell.execute_reply":"2024-05-31T09:14:18.515408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"time1 = time.time()\nL = 512\nNP = len(kidney_df)  # Assuming NP is the number of images in your dataframe\nlabel_mat = np.zeros((NP, L, L, 1), dtype=np.uint8)  # Initialize label_mat with zeros\n\nfor i in range(NP):\n    # Load images from NumPy files\n    row = kidney_df.iloc[i]\n    mask_path = masks_dir + row['id'] + \".npy\"  # Assuming 'id' is the column containing image filenames\n    \n    # Read the mask using the read_mask function\n    mask = read_mask(mask_path, channel=2)  # Assuming you want to use the second channel\n    \n    # Assign the modified mask to label_mat\n    label_mat[i, :, :, :] = mask\n\ntime2 = time.time()\ntime3 = np.round(time2 - time1)\nprint(time3, \"sec\")","metadata":{"papermill":{"duration":16.807064,"end_time":"2024-05-19T15:45:37.245933","exception":false,"start_time":"2024-05-19T15:45:20.438869","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:14:18.517405Z","iopub.execute_input":"2024-05-31T09:14:18.517886Z","iopub.status.idle":"2024-05-31T09:14:41.956186Z","shell.execute_reply.started":"2024-05-31T09:14:18.517861Z","shell.execute_reply":"2024-05-31T09:14:41.955012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(label_mat.shape)","metadata":{"papermill":{"duration":0.026563,"end_time":"2024-05-19T15:45:37.292174","exception":false,"start_time":"2024-05-19T15:45:37.265611","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:14:41.960428Z","iopub.execute_input":"2024-05-31T09:14:41.960810Z","iopub.status.idle":"2024-05-31T09:14:41.966821Z","shell.execute_reply.started":"2024-05-31T09:14:41.960779Z","shell.execute_reply":"2024-05-31T09:14:41.965418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def overlay_mask(image, mask, opacity=0.80):\n    if np.max(mask) == 0:\n        return image.astype(np.uint8)  # Return the original image if the mask is blank (all zeros)\n    \n    alpha = mask[:, :, 0] * opacity  # Extract the single channel from the mask & Adjust the opacity by multiplying with a factor\n    alpha = alpha[:, :, np.newaxis]   # Add a third dimension to make it compatible with the image\n    result = alpha * mask + (1 - alpha) * image\n    \n    return result.astype(np.uint8)","metadata":{"papermill":{"duration":0.026924,"end_time":"2024-05-19T15:45:37.338247","exception":false,"start_time":"2024-05-19T15:45:37.311323","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:14:41.968138Z","iopub.execute_input":"2024-05-31T09:14:41.968560Z","iopub.status.idle":"2024-05-31T09:14:41.978062Z","shell.execute_reply.started":"2024-05-31T09:14:41.968532Z","shell.execute_reply":"2024-05-31T09:14:41.977169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1,3, figsize = (10, 10))\n      \n_overlay = overlay_mask(X_train[0], label_mat[0])\n\nax[0].imshow(X_train[0])\nax[1].imshow(_overlay)\nax[2].imshow(label_mat[0], cmap='gray')\n\nax[0].set_title(\"Image \")\nax[1].set_title(\"Overlayed Image\")\nax[2].set_title(\"pixel wise label\")","metadata":{"papermill":{"duration":0.777212,"end_time":"2024-05-19T15:45:38.134639","exception":false,"start_time":"2024-05-19T15:45:37.357427","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:14:41.979194Z","iopub.execute_input":"2024-05-31T09:14:41.979976Z","iopub.status.idle":"2024-05-31T09:14:42.811319Z","shell.execute_reply.started":"2024-05-31T09:14:41.979936Z","shell.execute_reply":"2024-05-31T09:14:42.809835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<center><div style='color:#ffffff;\n           display:inline-block;\n           padding: 5px 5px 5px 5px;\n           border-radius:5px;\n           background-color:#78D1E1;\n           font-size:100%;'><a href=#toc style='text-decoration: none; color:#03001C;'>⬆️ Back To Top</a></div></center>\n\n<a id='1'></a>\n# 3 | EDA\n<div style=\"padding: 4px;color:white;margin:10;font-size:200%;text-align:center;display:fill;border-radius:10px;overflow:hidden;background-image: url(https://i.postimg.cc/j2bBmHWx/Py-Torch-Gradient.jpg); background-size: 100% auto;\"></div>\n\n<br>","metadata":{"papermill":{"duration":0.022605,"end_time":"2024-05-19T15:45:38.179477","exception":false,"start_time":"2024-05-19T15:45:38.156872","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"\n# 3.1 | Basic Information","metadata":{"papermill":{"duration":0.022481,"end_time":"2024-05-19T15:45:38.224562","exception":false,"start_time":"2024-05-19T15:45:38.202081","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# To dispaly data distribution\nselect = [\"dataset\", \"source_wsi\", \"dataset_wsi\"]\n\nsum_data_df = kidney_df[select].groupby(select[0:-1]).count()\nsum_data_df.plot(kind = \"bar\", title = \"count of (dataset, source wsi)\", color=[\"#4682B4\", \"#4682B4\"])\nplt.show()","metadata":{"papermill":{"duration":0.32821,"end_time":"2024-05-19T15:45:38.575452","exception":false,"start_time":"2024-05-19T15:45:38.247242","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:14:42.813026Z","iopub.execute_input":"2024-05-31T09:14:42.813388Z","iopub.status.idle":"2024-05-31T09:14:43.109081Z","shell.execute_reply.started":"2024-05-31T09:14:42.813361Z","shell.execute_reply":"2024-05-31T09:14:43.107719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# To dispaly data distribution\nselect = [\"id\", \"dataset\"]\n\nsum_data_df = kidney_df[select].groupby(select[-1]).count()\nsum_data_df.plot(kind = \"bar\", title = \"count of dataset\")\nplt.show()","metadata":{"papermill":{"duration":0.272712,"end_time":"2024-05-19T15:45:38.870836","exception":false,"start_time":"2024-05-19T15:45:38.598124","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:14:43.110668Z","iopub.execute_input":"2024-05-31T09:14:43.111000Z","iopub.status.idle":"2024-05-31T09:14:43.344993Z","shell.execute_reply.started":"2024-05-31T09:14:43.110973Z","shell.execute_reply":"2024-05-31T09:14:43.343857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_photo(img, title = \"\"):\n    \n    N = img.shape[0]\n    \n    NC = 5\n    NR =  math.ceil(N/NC)\n    fig, ax = plt.subplots(NR, NC, figsize = (12, NR*2.3))\n    \n    for k in range(N):\n        i =  int(k/NC)\n        j = k % NC\n        \n        if N <= 5:\n            ax[j].imshow(img[k],cmap='gray')\n            ax[j].tick_params(left = False, right = False , labelleft = False ,\n                        labelbottom = False, bottom = False)\n        else:\n            ax[i,j].imshow(img[k],cmap='gray')\n            ax[i,j].tick_params(left = False, right = False , labelleft = False ,\n                        labelbottom = False, bottom = False)\n    \n    if title != \"\":\n        if N <= 5:\n            ax[2].set_title(title)\n        else:\n            ax[0,2].set_title(title)","metadata":{"papermill":{"duration":0.035852,"end_time":"2024-05-19T15:45:38.929958","exception":false,"start_time":"2024-05-19T15:45:38.894106","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:14:43.346632Z","iopub.execute_input":"2024-05-31T09:14:43.347097Z","iopub.status.idle":"2024-05-31T09:14:43.357965Z","shell.execute_reply.started":"2024-05-31T09:14:43.347057Z","shell.execute_reply":"2024-05-31T09:14:43.356632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Function to overlay masks and plot\ndef plot_with_masks(images, masks, title=\"\"):\n    overlaid_images = [overlay_mask(image, mask) for image, mask in zip(images, masks)]\n    plot_photo(np.array(overlaid_images), title)","metadata":{"papermill":{"duration":0.031499,"end_time":"2024-05-19T15:45:38.984136","exception":false,"start_time":"2024-05-19T15:45:38.952637","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:14:43.362224Z","iopub.execute_input":"2024-05-31T09:14:43.362624Z","iopub.status.idle":"2024-05-31T09:14:43.372508Z","shell.execute_reply.started":"2024-05-31T09:14:43.362568Z","shell.execute_reply":"2024-05-31T09:14:43.371328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3.2 | Images and Labels\n\nNext, images and labels are visualized by each dataset and source (person). The size and quantity of blood vessels are different between sources. ","metadata":{"papermill":{"duration":0.022946,"end_time":"2024-05-19T15:45:39.029865","exception":false,"start_time":"2024-05-19T15:45:39.006919","status":"completed"},"tags":[]}},{"cell_type":"code","source":"#This code select some samples from each dataset and source\nnp.random.seed(1)\n\nsamples_ds = []\nfor ds in [1,2]:\n    filter1 = kidney_df[\"dataset\"]==ds\n    samples = []\n    for wsi in [1,2,3,4]:\n        \n        if ds ==1:\n            if wsi in [1,2]:\n                filter2 = filter1 & (kidney_df[\"source_wsi\"] == wsi)\n                s_idx = np.random.choice(np.where(filter2)[0], 50)\n                \n        else:\n            filter2 = filter1 & (kidney_df[\"source_wsi\"] == wsi)\n            s_idx = np.random.choice(np.where(filter2)[0], 50)\n            \n        samples.append(s_idx)\n        \n    samples_ds.append(samples)","metadata":{"papermill":{"duration":0.037476,"end_time":"2024-05-19T15:45:39.132244","exception":false,"start_time":"2024-05-19T15:45:39.094768","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:14:43.374081Z","iopub.execute_input":"2024-05-31T09:14:43.374498Z","iopub.status.idle":"2024-05-31T09:14:43.389470Z","shell.execute_reply.started":"2024-05-31T09:14:43.374459Z","shell.execute_reply":"2024-05-31T09:14:43.388166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Dataset1**\n\nIn dataset 1, source 1 contains a larger size but smaller number of blood vessels, whereas source 2 has many smaller size blood vessels. ","metadata":{"papermill":{"duration":0.022793,"end_time":"2024-05-19T15:45:39.177750","exception":false,"start_time":"2024-05-19T15:45:39.154957","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Plotting the images with overlaid masks\nplot_photo(X_train[samples_ds[0][0][0:5]], \"dataset 1, source 1\")\nplot_with_masks(X_train[samples_ds[0][0][0:5]], label_mat[samples_ds[0][0][0:5]], \"dataset 1, source 1 with masks\")\nplot_photo(label_mat[samples_ds[0][0][0:5]], \"dataset 1, source 1, Label\")\nplot_photo(X_train[samples_ds[0][1][0:5]], \"dataset 1, source 2\")\nplot_with_masks(X_train[samples_ds[0][1][0:5]], label_mat[samples_ds[0][1][0:5]], \"dataset 1, source 2 with masks\")\nplot_photo(label_mat[samples_ds[0][1][0:5]], \"dataset 1, source 2, Label\")","metadata":{"papermill":{"duration":6.101514,"end_time":"2024-05-19T15:45:45.302075","exception":false,"start_time":"2024-05-19T15:45:39.200561","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:14:43.391764Z","iopub.execute_input":"2024-05-31T09:14:43.392320Z","iopub.status.idle":"2024-05-31T09:14:49.899689Z","shell.execute_reply.started":"2024-05-31T09:14:43.392235Z","shell.execute_reply":"2024-05-31T09:14:49.898508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Dataset 2**\n\nThe images and labels of dataset 2 look similar to dataset 1, but their color is slightly different. ","metadata":{"papermill":{"duration":0.049988,"end_time":"2024-05-19T15:45:45.404473","exception":false,"start_time":"2024-05-19T15:45:45.354485","status":"completed"},"tags":[]}},{"cell_type":"code","source":"plot_photo(X_train[samples_ds[1][0][0:5]], \"dataset 2, source 1\")\nplot_with_masks(X_train[samples_ds[1][0][0:5]], label_mat[samples_ds[1][0][0:5]], \"dataset 2, source 1 with masks\")\nplot_photo(label_mat[samples_ds[1][0][0:5]], \"dataset 2, source 1, Label\")\n\nplot_photo(X_train[samples_ds[1][1][0:5]], \"dataset 2, source 2\")\nplot_with_masks(X_train[samples_ds[1][1][0:5]], label_mat[samples_ds[1][1][0:5]], \"dataset 2, source 2 with masks\")\nplot_photo(label_mat[samples_ds[1][1][0:5]], \"dataset 2, source 2, Label\")\n\nplot_photo(X_train[samples_ds[1][2][0:5]], \"dataset 2, source 3\")\nplot_with_masks(X_train[samples_ds[1][2][0:5]], label_mat[samples_ds[1][2][0:5]], \"dataset 2, source 3 with masks\")\nplot_photo(label_mat[samples_ds[1][2][0:5]], \"dataset 2, source 3, Label\")\n\nplot_photo(X_train[samples_ds[1][3][0:5]], \"dataset 2, source 4\")\nplot_with_masks(X_train[samples_ds[1][3][0:5]], label_mat[samples_ds[1][3][0:5]], \"dataset 2, source 4 with masks\")\nplot_photo(label_mat[samples_ds[1][3][0:5]], \"dataset 2, source 4, Label\")","metadata":{"papermill":{"duration":12.719297,"end_time":"2024-05-19T15:45:58.173887","exception":false,"start_time":"2024-05-19T15:45:45.454590","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:14:49.901300Z","iopub.execute_input":"2024-05-31T09:14:49.901760Z","iopub.status.idle":"2024-05-31T09:15:02.932699Z","shell.execute_reply.started":"2024-05-31T09:14:49.901721Z","shell.execute_reply":"2024-05-31T09:15:02.931590Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def label_stat(Y_mat):\n    \n    #n_labels = number of labels\n    # labels_mat= 2D matrix in which label of connected areas are assigned. \n    \n    n_labels, labels_mat = cv2.connectedComponents(Y_mat)\n    size_list = []\n    for i in range(1, n_labels):\n        #Calculate number of pixels for each label\n        filter1 = labels_mat == i\n        size = filter1.sum()\n        size_list.append(size)\n\n    return n_labels - 1, np.mean(size_list)\n\n\ndef density_plot(dataset, n_source):\n    \n    fig, ax = plt.subplots(1,n_source, figsize = (3.5*n_source,2.7))\n    \n    for i in range(n_source):\n        filter1 = (kidney_df[\"dataset\"] == dataset) & (kidney_df[\"source_wsi\"] == i + 1)\n        sns.kdeplot(kidney_df[\"positive_ratio\"].loc[filter1], ax = ax[i])\n        ax[i].set_xlim(0, 0.5)\n        ax[i].set_title(\"(dataset, source) = (\" +str(dataset) + \", \" + str(i+1)  + \")\" )\n        if i > 0:\n            ax[i].set_ylabel(\"\")\n            \n    plt.show()\n    \n\ndef scatter_plot(dataset, n_source):\n    fig, ax = plt.subplots(1,n_source, figsize = (3.5*n_source,2.7))\n\n    for i in range(n_source):\n        filter1 = (kidney_df[\"dataset\"] == dataset) & (kidney_df[\"source_wsi\"] == i + 1)\n        sns.scatterplot(data = kidney_df.loc[filter1], x =  \"n_BloodVessel\", y = \"meanSize_BloodVessel\", ax = ax[i], alpha = 0.4)\n        ax[i].set_xlim(0, 35)\n        ax[i].set_ylim(0, 15)\n        if i > 0:\n            ax[i].set_ylabel(\"\")\n        ax[i].set_title(\"(dataset, source) = (\" +str(dataset) + \", \" + str(i+1)  + \")\" )\n    \n    ax[0].set_ylabel(\"mean size (1000 pixel)\")\n    plt.show()\n    \n","metadata":{"papermill":{"duration":0.121443,"end_time":"2024-05-19T15:45:58.401896","exception":false,"start_time":"2024-05-19T15:45:58.280453","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:15:02.934196Z","iopub.execute_input":"2024-05-31T09:15:02.934562Z","iopub.status.idle":"2024-05-31T09:15:02.947612Z","shell.execute_reply.started":"2024-05-31T09:15:02.934532Z","shell.execute_reply":"2024-05-31T09:15:02.946418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_label_list = []\nmean_size_list = []\nratio_list = []\nfor i in range(NP):\n    n_label, mean_size = label_stat(label_mat[i])\n    n_label_list.append(n_label)\n    mean_size_list.append(mean_size)\n    ratio_list.append(np.mean(label_mat[i]))","metadata":{"papermill":{"duration":4.892741,"end_time":"2024-05-19T15:46:03.403296","exception":false,"start_time":"2024-05-19T15:45:58.510555","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:15:02.949081Z","iopub.execute_input":"2024-05-31T09:15:02.949432Z","iopub.status.idle":"2024-05-31T09:15:08.547165Z","shell.execute_reply.started":"2024-05-31T09:15:02.949403Z","shell.execute_reply":"2024-05-31T09:15:08.546116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"kidney_df[\"n_BloodVessel\"] = n_label_list\nkidney_df[\"meanSize_BloodVessel\"] = np.array(mean_size_list)/1000\nkidney_df[\"positive_ratio\"] = ratio_list","metadata":{"papermill":{"duration":0.120938,"end_time":"2024-05-19T15:46:03.635830","exception":false,"start_time":"2024-05-19T15:46:03.514892","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:15:08.548669Z","iopub.execute_input":"2024-05-31T09:15:08.549543Z","iopub.status.idle":"2024-05-31T09:15:08.559572Z","shell.execute_reply.started":"2024-05-31T09:15:08.549500Z","shell.execute_reply":"2024-05-31T09:15:08.558231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Ratio of positive pixel\n\nOverall, only 5.0% of pixels are positive (blood vessels). When it is calculated by each dataset and source, dataset 1 has more positive pixels than dataset 2.","metadata":{"papermill":{"duration":0.10868,"end_time":"2024-05-19T15:46:03.852217","exception":false,"start_time":"2024-05-19T15:46:03.743537","status":"completed"},"tags":[]}},{"cell_type":"code","source":"pct = np.round(kidney_df[\"positive_ratio\"].mean()*100, 1)\nprint(pct, \"% of pixels are positive in all data\")","metadata":{"papermill":{"duration":0.116711,"end_time":"2024-05-19T15:46:04.076406","exception":false,"start_time":"2024-05-19T15:46:03.959695","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:15:08.560915Z","iopub.execute_input":"2024-05-31T09:15:08.561765Z","iopub.status.idle":"2024-05-31T09:15:08.574906Z","shell.execute_reply.started":"2024-05-31T09:15:08.561731Z","shell.execute_reply":"2024-05-31T09:15:08.573703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"kidney_df[[\"dataset\", \"source_wsi\", \"positive_ratio\"]].groupby([\"dataset\", \"source_wsi\"]).mean()","metadata":{"papermill":{"duration":0.134875,"end_time":"2024-05-19T15:46:04.318274","exception":false,"start_time":"2024-05-19T15:46:04.183399","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:15:08.578781Z","iopub.execute_input":"2024-05-31T09:15:08.579535Z","iopub.status.idle":"2024-05-31T09:15:08.605483Z","shell.execute_reply.started":"2024-05-31T09:15:08.579499Z","shell.execute_reply":"2024-05-31T09:15:08.604135Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The ratio of positive pixels is calculated by each image, and the distribution is plotted in the following density plots.","metadata":{"papermill":{"duration":0.127779,"end_time":"2024-05-19T15:46:04.578028","exception":false,"start_time":"2024-05-19T15:46:04.450249","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"**Dataset 1**","metadata":{"papermill":{"duration":0.108045,"end_time":"2024-05-19T15:46:04.796169","exception":false,"start_time":"2024-05-19T15:46:04.688124","status":"completed"},"tags":[]}},{"cell_type":"code","source":"density_plot(1, 2)","metadata":{"papermill":{"duration":0.547861,"end_time":"2024-05-19T15:46:05.451330","exception":false,"start_time":"2024-05-19T15:46:04.903469","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:15:09.330061Z","iopub.execute_input":"2024-05-31T09:15:09.330446Z","iopub.status.idle":"2024-05-31T09:15:09.756253Z","shell.execute_reply.started":"2024-05-31T09:15:09.330420Z","shell.execute_reply":"2024-05-31T09:15:09.755047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Dataset 2**","metadata":{"papermill":{"duration":0.111605,"end_time":"2024-05-19T15:46:05.672038","exception":false,"start_time":"2024-05-19T15:46:05.560433","status":"completed"},"tags":[]}},{"cell_type":"code","source":"density_plot(2, 4)","metadata":{"papermill":{"duration":0.89092,"end_time":"2024-05-19T15:46:06.672154","exception":false,"start_time":"2024-05-19T15:46:05.781234","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:15:09.758312Z","iopub.execute_input":"2024-05-31T09:15:09.758787Z","iopub.status.idle":"2024-05-31T09:15:10.593710Z","shell.execute_reply.started":"2024-05-31T09:15:09.758747Z","shell.execute_reply":"2024-05-31T09:15:10.592693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Size and quantity of Blood Vessels\n\nDataset 1 source 1 has a different distribution from others. It has less number of blood vessels, but their size is larger. Please note that **one dot corresponds to one image** in scatterplots.","metadata":{"papermill":{"duration":0.107683,"end_time":"2024-05-19T15:46:06.888562","exception":false,"start_time":"2024-05-19T15:46:06.780879","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"**dataset1**","metadata":{"papermill":{"duration":0.109427,"end_time":"2024-05-19T15:46:07.105813","exception":false,"start_time":"2024-05-19T15:46:06.996386","status":"completed"},"tags":[]}},{"cell_type":"code","source":"scatter_plot(1, 2)","metadata":{"papermill":{"duration":0.529361,"end_time":"2024-05-19T15:46:07.743774","exception":false,"start_time":"2024-05-19T15:46:07.214413","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:15:11.550918Z","iopub.execute_input":"2024-05-31T09:15:11.551336Z","iopub.status.idle":"2024-05-31T09:15:11.981823Z","shell.execute_reply.started":"2024-05-31T09:15:11.551306Z","shell.execute_reply":"2024-05-31T09:15:11.980438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**dataset2**","metadata":{"papermill":{"duration":0.107888,"end_time":"2024-05-19T15:46:07.961794","exception":false,"start_time":"2024-05-19T15:46:07.853906","status":"completed"},"tags":[]}},{"cell_type":"code","source":"scatter_plot(2, 4)","metadata":{"papermill":{"duration":0.861901,"end_time":"2024-05-19T15:46:08.932541","exception":false,"start_time":"2024-05-19T15:46:08.070640","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:15:11.984047Z","iopub.execute_input":"2024-05-31T09:15:11.984621Z","iopub.status.idle":"2024-05-31T09:15:12.778247Z","shell.execute_reply.started":"2024-05-31T09:15:11.984561Z","shell.execute_reply":"2024-05-31T09:15:12.777199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<center><div style='color:#ffffff;\n           display:inline-block;\n           padding: 5px 5px 5px 5px;\n           border-radius:5px;\n           background-color:#78D1E1;\n           font-size:100%;'><a href=#toc style='text-decoration: none; color:#03001C;'>⬆️ Back To Top</a></div></center>\n\n<a id='1'></a>\n# 4 | Implement UNET Architecture\n<div style=\"padding: 4px;color:white;margin:10;font-size:200%;text-align:center;display:fill;border-radius:10px;overflow:hidden;background-image: url(https://i.postimg.cc/j2bBmHWx/Py-Torch-Gradient.jpg); background-size: 100% auto;\"></div>\n\n<br> ","metadata":{"papermill":{"duration":0.110178,"end_time":"2024-05-19T15:46:09.154820","exception":false,"start_time":"2024-05-19T15:46:09.044642","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# 4.1 | U-NET\n\nU-net is similar to convolutional autoencoders that have a downsampling decoder and an upsampling encoder. What makes U-net unique is **skip connection from the decoder to the encoder at each level**. This structure looks similar to [V-model](https://en.wikipedia.org/wiki/V-model). \n\n`LeftBlock` written in the below code has `Conv2D` followed by `BatchNormalization` and `Relu`. It downsamples input by `MaxPooling2D` upon request. `RightBlock` upsamples input by `Conv2DTranspose` and it is concatenated with a skip connection from `LeftBlock`. It is simple when plotted in the figure. ","metadata":{"papermill":{"duration":0.111998,"end_time":"2024-05-19T15:46:09.379063","exception":false,"start_time":"2024-05-19T15:46:09.267065","status":"completed"},"tags":[]}},{"cell_type":"code","source":"\ndef LeftBlock(channel, X, ksize=3, downsample=True):\n    if downsample:\n        X = layers.MaxPooling2D(pool_size=(2, 2), strides=(2, 2))(X)\n\n    X = layers.Conv2D(channel, kernel_size=ksize, strides=1, padding=\"same\")(X)\n    X = layers.BatchNormalization()(X)\n    X = layers.ReLU()(X)\n\n    # Add more convolutional layers to increase model capacity\n    X = layers.Conv2D(channel, kernel_size=ksize, strides=1, padding=\"same\")(X)\n    X = layers.BatchNormalization()(X)\n    X = layers.ReLU()(X)\n    \n    X = layers.Conv2D(channel, kernel_size=ksize, strides=1, padding=\"same\")(X)\n    X = layers.BatchNormalization()(X)\n    X = layers.ReLU()(X)\n    \n    X = layers.Conv2D(channel, kernel_size=ksize, strides=1, padding=\"same\")(X)\n    X = layers.BatchNormalization()(X)\n    X = layers.ReLU()(X)\n\n    return X\n\ndef RightBlock(channel, X, ksize=3, X_skip=None, upsample=True):\n    if upsample:\n        X = layers.Conv2DTranspose(channel, kernel_size=3, strides=2, padding=\"same\")(X)\n\n    if X_skip is not None:\n        X = layers.Concatenate()([X, X_skip])\n\n    X = layers.Conv2D(channel, kernel_size=ksize, strides=1, padding=\"same\")(X)\n    X = layers.BatchNormalization()(X)\n    X = layers.ReLU()(X)\n\n    # Add more convolutional layers to increase model capacity\n    X = layers.Conv2D(channel, kernel_size=ksize, strides=1, padding=\"same\")(X)\n    X = layers.BatchNormalization()(X)\n    X = layers.ReLU()(X)\n    \n    X = layers.Conv2D(channel, kernel_size=ksize, strides=1, padding=\"same\")(X)\n    X = layers.BatchNormalization()(X)\n    X = layers.ReLU()(X)\n    \n    X = layers.Conv2D(channel, kernel_size=ksize, strides=1, padding=\"same\")(X)\n    X = layers.BatchNormalization()(X)\n    X = layers.ReLU()(X)\n\n    return X","metadata":{"papermill":{"duration":0.128043,"end_time":"2024-05-19T15:46:09.618113","exception":false,"start_time":"2024-05-19T15:46:09.490070","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:15:12.780008Z","iopub.execute_input":"2024-05-31T09:15:12.780313Z","iopub.status.idle":"2024-05-31T09:15:12.793471Z","shell.execute_reply.started":"2024-05-31T09:15:12.780288Z","shell.execute_reply":"2024-05-31T09:15:12.791592Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras import layers, Model\n\ndef create_UNET():\n    L = 512\n    Input = layers.Input(shape=(L, L, 3))\n\n    X0 = layers.Rescaling(scale=1./127.5, offset=-1)(Input)\n\n    KS = 3  # Reduced kernel size for optimization\n    channel = 16  # Reduced number of initial channels for optimization\n\n    X1 = LeftBlock(channel, X0, ksize=KS, downsample=False)  # 512\n\n    channel *= 2\n    X2 = LeftBlock(channel, X1, ksize=KS, downsample=True)  # 256\n\n    channel *= 2\n    X3 = LeftBlock(channel, X2, ksize=KS, downsample=True)  # 128\n\n    channel *= 2\n    X4 = LeftBlock(channel, X3, ksize=KS, downsample=True)  # 64\n\n    XR = RightBlock(channel, X4, ksize=KS, X_skip=X3, upsample=True)  # 128\n\n    channel = int(channel / 2)\n    XR = RightBlock(channel, XR, ksize=KS, X_skip=X2, upsample=True)  # 256\n\n    channel = int(channel / 2)\n    XR = RightBlock(channel, XR, ksize=KS, X_skip=X1, upsample=True)  # 512\n\n    channel = 1\n    XR = layers.Conv2D(channel, kernel_size=1, strides=1, padding=\"same\")(XR)\n\n    model = Model(inputs=Input, outputs=XR)\n\n    return model","metadata":{"papermill":{"duration":0.124631,"end_time":"2024-05-19T15:46:09.853667","exception":false,"start_time":"2024-05-19T15:46:09.729036","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:15:13.763064Z","iopub.execute_input":"2024-05-31T09:15:13.763438Z","iopub.status.idle":"2024-05-31T09:15:13.772361Z","shell.execute_reply.started":"2024-05-31T09:15:13.763411Z","shell.execute_reply":"2024-05-31T09:15:13.771472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Example of changing weight initialization\nmodel_Unet = create_UNET()\nmodel_Unet.summary()","metadata":{"papermill":{"duration":1.404415,"end_time":"2024-05-19T15:46:11.376617","exception":false,"start_time":"2024-05-19T15:46:09.972202","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:15:13.786191Z","iopub.execute_input":"2024-05-31T09:15:13.786857Z","iopub.status.idle":"2024-05-31T09:15:14.609509Z","shell.execute_reply.started":"2024-05-31T09:15:13.786822Z","shell.execute_reply":"2024-05-31T09:15:14.608295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tf.keras.utils.plot_model(model_Unet)\n","metadata":{"papermill":{"duration":2.291921,"end_time":"2024-05-19T15:46:13.782437","exception":false,"start_time":"2024-05-19T15:46:11.490516","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:15:14.611509Z","iopub.execute_input":"2024-05-31T09:15:14.611845Z","iopub.status.idle":"2024-05-31T09:15:17.245161Z","shell.execute_reply.started":"2024-05-31T09:15:14.611818Z","shell.execute_reply":"2024-05-31T09:15:17.244198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<center><div style='color:#ffffff;\n           display:inline-block;\n           padding: 5px 5px 5px 5px;\n           border-radius:5px;\n           background-color:#78D1E1;\n           font-size:100%;'><a href=#toc style='text-decoration: none; color:#03001C;'>⬆️ Back To Top</a></div></center>\n\n<a id='1'></a>\n# 5 | Training and Hyperparameter Tuning\n<div style=\"padding: 4px;color:white;margin:10;font-size:200%;text-align:center;display:fill;border-radius:10px;overflow:hidden;background-image: url(https://i.postimg.cc/j2bBmHWx/Py-Torch-Gradient.jpg); background-size: 100% auto;\"></div>\n\n<br>","metadata":{"papermill":{"duration":0.127684,"end_time":"2024-05-19T15:46:14.039118","exception":false,"start_time":"2024-05-19T15:46:13.911434","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# 5.1 | Split Dataset\n\n80% of images are randomly selected for training. ","metadata":{"papermill":{"duration":0.129243,"end_time":"2024-05-19T15:46:14.299241","exception":false,"start_time":"2024-05-19T15:46:14.169998","status":"completed"},"tags":[]}},{"cell_type":"code","source":"batch_size = 4","metadata":{"papermill":{"duration":0.154726,"end_time":"2024-05-19T15:46:14.586120","exception":false,"start_time":"2024-05-19T15:46:14.431394","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:15:17.246802Z","iopub.execute_input":"2024-05-31T09:15:17.247583Z","iopub.status.idle":"2024-05-31T09:15:17.252689Z","shell.execute_reply.started":"2024-05-31T09:15:17.247543Z","shell.execute_reply":"2024-05-31T09:15:17.251188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"np.random.seed(1)\nn_data = NP\nn_train = int(n_data*0.8)\n\nall_idx = np.arange(n_data)\nnp.random.shuffle(all_idx)\n\ntrain_idx = all_idx[:n_train]\nval_idx = all_idx[n_train:]\n\nprint(\"N of data for training = \", train_idx.shape[0]) \nprint(\"N of data for validation = \", val_idx.shape[0])","metadata":{"papermill":{"duration":0.137846,"end_time":"2024-05-19T15:46:14.858001","exception":false,"start_time":"2024-05-19T15:46:14.720155","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:15:19.881956Z","iopub.execute_input":"2024-05-31T09:15:19.882356Z","iopub.status.idle":"2024-05-31T09:15:19.889752Z","shell.execute_reply.started":"2024-05-31T09:15:19.882324Z","shell.execute_reply":"2024-05-31T09:15:19.888478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# train_ds = tf.data.Dataset.from_tensor_slices((X_train[train_idx], label_mat[train_idx])).shuffle(1000).batch(batch_size)\n\nX_val = X_train[val_idx]\nY_val = label_mat[val_idx]\nn_val = X_val.shape[0]\n","metadata":{"papermill":{"duration":0.23844,"end_time":"2024-05-19T15:46:15.225151","exception":false,"start_time":"2024-05-19T15:46:14.986711","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:15:21.404040Z","iopub.execute_input":"2024-05-31T09:15:21.404796Z","iopub.status.idle":"2024-05-31T09:15:21.545051Z","shell.execute_reply.started":"2024-05-31T09:15:21.404745Z","shell.execute_reply":"2024-05-31T09:15:21.543225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5.2 | Oversample technique","metadata":{"papermill":{"duration":0.127781,"end_time":"2024-05-19T15:46:15.481921","exception":false,"start_time":"2024-05-19T15:46:15.354140","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import numpy as np\nfrom sklearn.utils import shuffle\n\ndef oversample_data_generator(X_train2, label_mat2, batch_size=4500, random_state=123):\n    L = 512\n\n    # Identify the minority class samples (assuming label 1 is the minority class)\n    minority_indices = np.where(label_mat2 == 1)[0]\n\n    while True:\n        # Set the random seed\n        np.random.seed(random_state)\n\n        # Randomly oversample the minority class to match the number of majority class samples\n        oversampled_minority_indices = np.random.choice(minority_indices, size=batch_size // 2, replace=True)\n # Combine oversampled minority and majority indices\n        selected_indices = np.concatenate([oversampled_minority_indices, np.where(label_mat2 == 0)[0]])\n\n        # Shuffle the combined indices\n        selected_indices = shuffle(selected_indices, random_state=random_state)\n\n        for i in range(0, len(selected_indices), batch_size):\n            batch_indices = selected_indices[i:i + batch_size]\n            X_batch = X_train2[batch_indices]\n            label_mat_batch = label_mat2[batch_indices]\n\n            # Assuming label_mat_batch is binary (0 or 1)\n            y_batch = label_mat_batch\n\n            yield X_batch, y_batch","metadata":{"papermill":{"duration":0.230848,"end_time":"2024-05-19T15:46:15.840298","exception":false,"start_time":"2024-05-19T15:46:15.609450","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:15:53.218327Z","iopub.execute_input":"2024-05-31T09:15:53.218711Z","iopub.status.idle":"2024-05-31T09:15:53.262447Z","shell.execute_reply.started":"2024-05-31T09:15:53.218683Z","shell.execute_reply":"2024-05-31T09:15:53.261051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_resampled, y_resampled  = next(oversample_data_generator(X_train[train_idx], label_mat[train_idx]))","metadata":{"papermill":{"duration":38.410782,"end_time":"2024-05-19T15:46:54.379690","exception":false,"start_time":"2024-05-19T15:46:15.968908","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:15:55.224566Z","iopub.execute_input":"2024-05-31T09:15:55.224985Z","iopub.status.idle":"2024-05-31T09:16:35.565037Z","shell.execute_reply.started":"2024-05-31T09:15:55.224952Z","shell.execute_reply":"2024-05-31T09:16:35.563258Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(y_resampled),len(X_resampled)","metadata":{"papermill":{"duration":0.137835,"end_time":"2024-05-19T15:46:54.659794","exception":false,"start_time":"2024-05-19T15:46:54.521959","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:16:35.568658Z","iopub.execute_input":"2024-05-31T09:16:35.569047Z","iopub.status.idle":"2024-05-31T09:16:35.578658Z","shell.execute_reply.started":"2024-05-31T09:16:35.569015Z","shell.execute_reply":"2024-05-31T09:16:35.576833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Concatenate the original and resampled data along with their labels\nX_combined = np.concatenate([X_train[train_idx], X_resampled])\ny_combined = np.concatenate([label_mat[train_idx], y_resampled])\ntrain_ds = tf.data.Dataset.from_tensor_slices((X_combined, y_combined)).shuffle(2000).batch(batch_size)","metadata":{"papermill":{"duration":16.248779,"end_time":"2024-05-19T15:47:11.037670","exception":false,"start_time":"2024-05-19T15:46:54.788891","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:16:35.580224Z","iopub.execute_input":"2024-05-31T09:16:35.580710Z","iopub.status.idle":"2024-05-31T09:16:48.431551Z","shell.execute_reply.started":"2024-05-31T09:16:35.580673Z","shell.execute_reply":"2024-05-31T09:16:48.429990Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(X_combined),len(y_combined),len(train_ds)","metadata":{"papermill":{"duration":0.139329,"end_time":"2024-05-19T15:47:11.308664","exception":false,"start_time":"2024-05-19T15:47:11.169335","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:16:48.434222Z","iopub.execute_input":"2024-05-31T09:16:48.434775Z","iopub.status.idle":"2024-05-31T09:16:48.444438Z","shell.execute_reply.started":"2024-05-31T09:16:48.434731Z","shell.execute_reply":"2024-05-31T09:16:48.443041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1,3, figsize = (10, 10))\n      \n_overlay = overlay_mask(X_resampled[0], y_resampled[0])\n\nax[0].imshow(X_combined[0])\nax[1].imshow(_overlay)\nax[2].imshow(y_combined[0], cmap='gray')\n\nax[0].set_title(\"Image (oversampled)\")\nax[1].set_title(\"Overlayed (oversampled) Image\")\nax[2].set_title(\"pixel wise label (oversampled)\")","metadata":{"execution":{"iopub.status.busy":"2024-05-31T09:17:26.562981Z","iopub.execute_input":"2024-05-31T09:17:26.563396Z","iopub.status.idle":"2024-05-31T09:17:27.362494Z","shell.execute_reply.started":"2024-05-31T09:17:26.563363Z","shell.execute_reply":"2024-05-31T09:17:27.359713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5.3 | Vessel Segmentation Loss Function. ","metadata":{"papermill":{"duration":0.127814,"end_time":"2024-05-19T15:47:11.565151","exception":false,"start_time":"2024-05-19T15:47:11.437337","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import keras.backend as K\nimport tensorflow as tf\n\ndef vessel_segmentation_loss(y_true, y_pred):\n    \"\"\"\n    Computes the loss for semantic segmentation of blood vessels in the kidney.\n    \"\"\"\n    # Compute pixel-wise cross entropy loss\n    ce_loss = tf.nn.sigmoid_cross_entropy_with_logits(labels=y_true, logits=y_pred)\n    ce_loss = tf.reduce_mean(ce_loss)\n    \n    # Compute dice loss to encourage better segmentation of vessels\n    smooth = 1e-5\n    y_pred_sigmoid = tf.sigmoid(y_pred)\n    y_true_binary = tf.cast(tf.equal(y_true, 1), tf.float32)\n    intersection = tf.reduce_sum(y_pred_sigmoid * y_true_binary)\n    union = tf.reduce_sum(y_pred_sigmoid) + tf.reduce_sum(y_true_binary)\n    dice_loss = 1 - (2 * intersection + smooth) / (union + smooth)\n    \n    # Combine cross entropy and dice losses\n    total_loss = ce_loss + dice_loss\n    \n    return total_loss","metadata":{"papermill":{"duration":0.139093,"end_time":"2024-05-19T15:47:11.832465","exception":false,"start_time":"2024-05-19T15:47:11.693372","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:18:05.727974Z","iopub.execute_input":"2024-05-31T09:18:05.728501Z","iopub.status.idle":"2024-05-31T09:18:05.742640Z","shell.execute_reply.started":"2024-05-31T09:18:05.728450Z","shell.execute_reply":"2024-05-31T09:18:05.741594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5.4 | Image Augmentation\n\nThe function `augment` is applied to both the image and the label, because when the image is flipped, the label shall follow it. The first part of augmentation is **horizontal and vertical flip** at probability 0.5 respectively. Next is image rotation. **The rotation angle is selected from 0, 90, 180, and 270**. Finally, **the image and the label are zoomed** at 0.3 probability. It selects the upper left pixel randomly. Then the length is selected. It is clipped and resized to 512x512. ## 6.3 Image Augmentation\n\nThe funciton `augment` is applied to both image and label, because when image is flipped, label shall follow it. First part of augmentation is **horizontal and vertical flip** at probability 0.5 respectively. Next is image rotation. **The rotation angle is selected from 0, 90, 180 and 270**. Finally, **the image and the label is zoomed** at 0.3 probability. It select uppler left pixel randomly. Then the length is selected. It is clipped and resized to 512x512. ","metadata":{"papermill":{"duration":0.127705,"end_time":"2024-05-19T15:47:12.089269","exception":false,"start_time":"2024-05-19T15:47:11.961564","status":"completed"},"tags":[]}},{"cell_type":"code","source":"\n\n#Image Augment Function\ndef augment(X, Y):\n    \n    #1. random flip--------------\n    #1.1 horizontal\n    if tf.random.uniform(shape=[1]) > 0.5:\n        X = X[:,:,::-1]\n        Y = Y[:,:,::-1]\n        \n    #1.2 vertical\n    if tf.random.uniform(shape=[1]) > 0.5:\n        X = X[:,::-1]\n        Y = Y[:,::-1]\n    \n    #2. rotation------------------\n    #2.1 set angle \n    angle = np.random.randint(4) \n    X = tf.image.rot90(X, k=angle)\n    Y = tf.image.rot90(Y, k=angle)\n    \n    #3. zoom -----------------------------\n    L = 512\n    if tf.random.uniform(shape=[1]) < 0.3:\n        \n        #left upper\n        x1 = np.random.choice(np.arange(0, int(L*0.2)), 1)[0]\n        y1 = np.random.choice(np.arange(0, int(L*0.2)), 1)[0]\n        \n        #Length\n        L2 = min(L-x1, L-y1)\n        L3 = np.random.choice(np.arange(int(L2*0.8), L2), 1)[0]\n        \n        X = X[:,y1:(y1+L3), x1:(x1+L3),:] \n        Y = Y[:,y1:(y1+L3), x1:(x1+L3),:] \n        \n        X = tf.cast(tf.image.resize(X,[L, L]), dtype = tf.uint8)\n        Y = tf.cast(tf.image.resize(Y,[L, L]), dtype = tf.uint8)\n    \n    return X, Y","metadata":{"papermill":{"duration":0.144979,"end_time":"2024-05-19T15:47:12.363573","exception":false,"start_time":"2024-05-19T15:47:12.218594","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:18:05.746307Z","iopub.execute_input":"2024-05-31T09:18:05.747001Z","iopub.status.idle":"2024-05-31T09:18:05.764490Z","shell.execute_reply.started":"2024-05-31T09:18:05.746960Z","shell.execute_reply":"2024-05-31T09:18:05.762952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def apply_augmentation(image, label):\n    # Apply augmentation to the image and label\n    augmented_image, augmented_label = augment(image, label)\n    return augmented_image, augmented_label","metadata":{"papermill":{"duration":0.137443,"end_time":"2024-05-19T15:47:12.638486","exception":false,"start_time":"2024-05-19T15:47:12.501043","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:18:05.765897Z","iopub.execute_input":"2024-05-31T09:18:05.766276Z","iopub.status.idle":"2024-05-31T09:18:05.779748Z","shell.execute_reply.started":"2024-05-31T09:18:05.766241Z","shell.execute_reply":"2024-05-31T09:18:05.778468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_ds_augmented = train_ds.map(apply_augmentation)","metadata":{"papermill":{"duration":0.611814,"end_time":"2024-05-19T15:47:13.380708","exception":false,"start_time":"2024-05-19T15:47:12.768894","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:18:09.786332Z","iopub.execute_input":"2024-05-31T09:18:09.786856Z","iopub.status.idle":"2024-05-31T09:18:10.317593Z","shell.execute_reply.started":"2024-05-31T09:18:09.786813Z","shell.execute_reply":"2024-05-31T09:18:10.316248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5.5 | Hyperparameter Tuning\n**General**\n\n* adam optimizer with learning rate = 0.0002. (much better than default lr 0.001) with batch size = 4\n* weight of negative loss (false positive) in `custom_loss` = 3\n* best epoch: 20 ~ 30 \n","metadata":{"papermill":{"duration":0.12865,"end_time":"2024-05-19T15:47:13.639564","exception":false,"start_time":"2024-05-19T15:47:13.510914","status":"completed"},"tags":[]}},{"cell_type":"code","source":"optimizer =  tf.keras.optimizers.Adam(learning_rate=0.0002)","metadata":{"papermill":{"duration":0.144011,"end_time":"2024-05-19T15:47:13.912322","exception":false,"start_time":"2024-05-19T15:47:13.768311","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:18:15.482981Z","iopub.execute_input":"2024-05-31T09:18:15.483375Z","iopub.status.idle":"2024-05-31T09:18:15.495844Z","shell.execute_reply.started":"2024-05-31T09:18:15.483348Z","shell.execute_reply":"2024-05-31T09:18:15.494554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras.callbacks import ModelCheckpoint, EarlyStopping ,ReduceLROnPlateau\n# Define early stopping callback\nearly_stopping = EarlyStopping(monitor='val_loss', patience=5, verbose=1, restore_best_weights=True)\n\n# Define model checkpoint callback to save the best model\nmodel_checkpoint = ModelCheckpoint('best_model.keras', monitor='val_loss', save_best_only=True)\n# Define learning rate reduction callback\nreduce_lr = ReduceLROnPlateau(monitor='val_loss', factor=0.2, patience=3, min_lr=1e-7)\nmodel_Unet.compile(optimizer=optimizer, loss=vessel_segmentation_loss, metrics=['accuracy'])","metadata":{"papermill":{"duration":0.143566,"end_time":"2024-05-19T15:47:14.185242","exception":false,"start_time":"2024-05-19T15:47:14.041676","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:18:20.253159Z","iopub.execute_input":"2024-05-31T09:18:20.253535Z","iopub.status.idle":"2024-05-31T09:18:20.268078Z","shell.execute_reply.started":"2024-05-31T09:18:20.253509Z","shell.execute_reply":"2024-05-31T09:18:20.267016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Train the model using the callbacks\nhistory = model_Unet.fit(train_ds_augmented, epochs=100, validation_data=(X_val, Y_val), callbacks=[model_checkpoint,reduce_lr]) ","metadata":{"execution":{"iopub.execute_input":"2024-05-19T15:47:14.453644Z","iopub.status.busy":"2024-05-19T15:47:14.453223Z","iopub.status.idle":"2024-05-19T22:59:31.067544Z","shell.execute_reply":"2024-05-19T22:59:31.066574Z"},"papermill":{"duration":25940.921835,"end_time":"2024-05-19T22:59:35.235311","exception":false,"start_time":"2024-05-19T15:47:14.313476","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\ndef plot_history(history):\n    fig, ax = plt.subplots(figsize=(5, 4))\n    ax.plot(history.history[\"loss\"], label=\"train loss\")\n    ax.plot(history.history[\"val_loss\"], label=\"val loss\")\n    ax.set_title(\"Loss\")\n    ax.legend()\n    ax.grid()\n\n# Assuming you have already trained your model and obtained the `history` object\nplot_history(history)\n","metadata":{"execution":{"iopub.execute_input":"2024-05-19T23:00:00.824708Z","iopub.status.busy":"2024-05-19T23:00:00.824365Z","iopub.status.idle":"2024-05-19T23:00:01.067964Z","shell.execute_reply":"2024-05-19T23:00:01.067095Z"},"papermill":{"duration":13.166934,"end_time":"2024-05-19T23:00:01.069957","exception":false,"start_time":"2024-05-19T22:59:47.903023","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<center><div style='color:#ffffff;\n           display:inline-block;\n           padding: 5px 5px 5px 5px;\n           border-radius:5px;\n           background-color:#78D1E1;\n           font-size:100%;'><a href=#toc style='text-decoration: none; color:#03001C;'>⬆️ Back To Top</a></div></center>\n\n<a id='1'></a>\n# 6 | Performance Analysis\n<div style=\"padding: 4px;color:white;margin:10;font-size:200%;text-align:center;display:fill;border-radius:10px;overflow:hidden;background-image: url(https://i.postimg.cc/j2bBmHWx/Py-Torch-Gradient.jpg); background-size: 100% auto;\"></div>\n\n<br>","metadata":{"papermill":{"duration":12.579078,"end_time":"2024-05-19T23:00:26.045342","exception":false,"start_time":"2024-05-19T23:00:13.466264","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"this is the best weight from training this model and we use different metrics to perform the model so we downloaded the link for this model is https://www.kaggle.com/code/norhanawad/hupmap-unet-semantic-segmentation-with-oversample","metadata":{}},{"cell_type":"code","source":"model_Unet.load_weights('/kaggle/input/hupmap-unet-oversample/best_model.keras')","metadata":{"execution":{"iopub.status.busy":"2024-05-31T09:18:42.724336Z","iopub.execute_input":"2024-05-31T09:18:42.725224Z","iopub.status.idle":"2024-05-31T09:18:43.911811Z","shell.execute_reply.started":"2024-05-31T09:18:42.725185Z","shell.execute_reply":"2024-05-31T09:18:43.910334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# model.predict method\ndef predict_masks(model, images):\n    # Ensure the input image has the right shape for prediction\n    if len(images.shape) == 3:\n        images = np.expand_dims(images, axis=0)  # Add batch dimension if needed\n    # Predict the probabilities for the input image\n    prob = model.predict(images)\n\n    # Return the predicted mask\n    return prob\ndef calculate_metrics(y_true, y_pred, threshold):\n    y_pred_binary = (y_pred > threshold).astype(np.uint8)\n\n    # True Positives, False Positives, False Negatives, True Negatives\n    TP = np.sum((y_true == 1) & (y_pred_binary == 1))\n    FP = np.sum((y_true == 0) & (y_pred_binary == 1))\n    TN = np.sum((y_true == 0) & (y_pred_binary == 0))\n    FN = np.sum((y_true == 1) & (y_pred_binary == 0))\n\n    # Dice coefficient\n    dice_denominator = 2 * TP + FP + FN\n    dice = (2 * TP) / dice_denominator if dice_denominator != 0 else 1\n\n    # Intersection over Union (IoU)\n    iou_denominator = TP + FP + FN\n    iou = TP / iou_denominator if iou_denominator != 0 else 1\n\n    # Precision\n    precision = TP / (TP + FP) if (TP + FP) != 0 else 1\n\n    # Recall\n    recall = TP / (TP + FN) if (TP + FN) != 0 else 1\n    \n  # F1 Score\n    f1_score = 2 * ((precision * recall) / (precision + recall)) if (precision + recall) != 0 else 1\n\n    # Confidence\n    binary_mask = y_pred > threshold\n    confidence_scores = y_pred.flatten()\n    binary_mask_flat = binary_mask.flatten()\n    blood_vessel_confidences = confidence_scores[binary_mask_flat]\n    confidence = np.mean(blood_vessel_confidences)\n\n    return dice, iou, precision, recall, f1_score, confidence\n\n\ndef metrics_dataframe(Y, Y_hat, threshold=0.5):\n    n_val = len(Y)\n    df_object = {}\n    df_object['dice'] = []\n    df_object['iou'] = []\n    df_object['precision'] = []\n    df_object['recall'] = []    \n    df_object['f1_score'] = []\n    df_object['confidence'] = []\n    df_object['threshold'] = threshold\n\n    for i in range(Y.shape[0]):\n        y_true = Y[i, :, :, 0]\n        y_pred = Y_hat[i, :, :, 0] > threshold\n        metrics = calculate_metrics(y_true,y_pred, threshold)\n        df_object['dice'].append(metrics[0])\n        df_object['iou'].append(metrics[1])\n        df_object['precision'].append(metrics[2])\n        df_object['recall'].append(metrics[3])        \n        df_object['f1_score'].append(metrics[4])\n        df_object['confidence'].append(metrics[5])\n        \n    return pd.DataFrame(df_object)","metadata":{"papermill":{"duration":12.633858,"end_time":"2024-05-19T23:00:51.311502","exception":false,"start_time":"2024-05-19T23:00:38.677644","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:20:42.051927Z","iopub.execute_input":"2024-05-31T09:20:42.052354Z","iopub.status.idle":"2024-05-31T09:20:42.067999Z","shell.execute_reply.started":"2024-05-31T09:20:42.052325Z","shell.execute_reply":"2024-05-31T09:20:42.066980Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# making prediction\nY_hat = predict_masks(model_Unet, X_val)","metadata":{"papermill":{"duration":20.356347,"end_time":"2024-05-19T23:01:24.537056","exception":false,"start_time":"2024-05-19T23:01:04.180709","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:20:52.355036Z","iopub.execute_input":"2024-05-31T09:20:52.355803Z","iopub.status.idle":"2024-05-31T09:26:35.280259Z","shell.execute_reply.started":"2024-05-31T09:20:52.355767Z","shell.execute_reply":"2024-05-31T09:26:35.279106Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def print_metrics(threshold=0.5):\n    n = len(Y_val)\n    iou, dice = [], []\n    \n    for i in range(Y_val.shape[0]):\n        metrics = calculate_metrics(Y_val[i], Y_hat[i], threshold)\n        iou.append(metrics[1])\n        dice.append(metrics[0])\n    \n    iou = np.round(np.mean(iou ) * 100, 4)\n    dice = np.round(np.mean(dice) * 100, 4)\n    \n    print(\"threshold {}% - IoU score {}% - Dice coefficient {}%\"\n          .format(threshold*100, np.mean(iou),np.mean(dice)))   \n    \n    return threshold, dice\nmax_dice = -1\nbest_threshold = 0\nfor i in range(50, 100, 1):\n    threshold, dice = print_metrics(i / 100)\n    if dice > max_dice:\n        max_dice = dice\n        best_threshold = threshold","metadata":{"papermill":{"duration":697.008693,"end_time":"2024-05-19T23:13:14.274850","exception":false,"start_time":"2024-05-19T23:01:37.266157","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:26:35.282848Z","iopub.execute_input":"2024-05-31T09:26:35.283589Z","iopub.status.idle":"2024-05-31T09:27:01.454373Z","shell.execute_reply.started":"2024-05-31T09:26:35.283548Z","shell.execute_reply":"2024-05-31T09:27:01.453076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metrics_df = metrics_dataframe(Y_val, Y_hat, best_threshold)\nmetrics_df.to_csv('metrics_dataframe.csv', index=False)\nmetrics_df.mean()","metadata":{"execution":{"iopub.status.busy":"2024-05-31T09:27:01.456017Z","iopub.execute_input":"2024-05-31T09:27:01.456679Z","iopub.status.idle":"2024-05-31T09:27:02.024838Z","shell.execute_reply.started":"2024-05-31T09:27:01.456630Z","shell.execute_reply":"2024-05-31T09:27:02.023878Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<center><div style='color:#ffffff;\n           display:inline-block;\n           padding: 5px 5px 5px 5px;\n           border-radius:5px;\n           background-color:#78D1E1;\n           font-size:100%;'><a href=#toc style='text-decoration: none; color:#03001C;'>⬆️ Back To Top</a></div></center>\n\n<a id='1'></a>\n# 7 | Predicted labels\n<div style=\"padding: 4px;color:white;margin:10;font-size:200%;text-align:center;display:fill;border-radius:10px;overflow:hidden;background-image: url(https://i.postimg.cc/j2bBmHWx/Py-Torch-Gradient.jpg); background-size: 100% auto;\"></div>\n\n<br>\nUNET is trained well, it must be able to predict labels with good accuracy. The function `plot_result` visualizes an image, its label, and prediction results ","metadata":{"papermill":{"duration":12.700882,"end_time":"2024-05-19T23:14:30.407961","exception":false,"start_time":"2024-05-19T23:14:17.707079","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def plot_result(X, Y_true, Y_pred, cutoff = 0.5):\n    \n    N = X.shape[0]\n    \n    fig, ax = plt.subplots(N,4, figsize = (10, 3*N))\n    \n    for k in range(N):\n        \n        cutoff_img1 = (Y_pred[k,:,:,0] > cutoff).astype(int)\n\n        true_img = np.zeros((512, 512, 3), dtype = np.uint8)\n        true_img[:,:,1] = Y_true[k,:,:,0]*200\n        \n        cutoff1 = np.zeros((512, 512, 3), dtype = np.uint8)\n        \n        cutoff1[:,:,0] = cutoff_img1*230\n        \n        cutoff1[:,:,1] = cutoff_img1*50\n        cutoff1[:,:,2] = cutoff_img1*50\n        \n        diff_photo1 = cutoff1.copy()\n        diff_photo1[:,:,1] += (Y_true[k,:,:,0]*200).astype(np.uint8)\n        \n        ax[k, 0].imshow(X[k])\n        ax[k, 1].imshow(true_img, cmap = \"gray\")\n        ax[k, 2].imshow(cutoff1, cmap = \"gray\")\n        ax[k, 3].imshow(diff_photo1)\n        \n        for j in range(4):\n            ax[k,j].set_xticks([])\n            ax[k,j].set_yticks([])\n            \n        if k == 0:\n                ax[k, 0].set_title(\"val img\")\n                ax[k, 1].set_title(\"true label\")\n                ax[k, 2].set_title(\"model (cutoff at {})\".format(cutoff))\n                ax[k, 3].set_title(\"Compare (Y:tp)\")","metadata":{"papermill":{"duration":12.631739,"end_time":"2024-05-19T23:14:55.834321","exception":false,"start_time":"2024-05-19T23:14:43.202582","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:27:02.027179Z","iopub.execute_input":"2024-05-31T09:27:02.027897Z","iopub.status.idle":"2024-05-31T09:27:02.038918Z","shell.execute_reply.started":"2024-05-31T09:27:02.027863Z","shell.execute_reply":"2024-05-31T09:27:02.037703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Results**\n\nTen of the prediction results are shown below. Both models predict somewhat well, although it is not a simple task. The true label is **green** and the predicted label is **red**. The predicted label is created by cutting off the predicted probability at 0.8. When both are plotted in the same image (refer to **Compare (Y: tp)**), the true positive becomes **yellow**. Precisions of prediction are varied, some are very well predicted but some are not. The next question would be \"Does the result depend on the dataset?\" ","metadata":{"papermill":{"duration":12.554651,"end_time":"2024-05-19T23:15:21.097746","exception":false,"start_time":"2024-05-19T23:15:08.543095","status":"completed"},"tags":[]}},{"cell_type":"code","source":"n_val = len(X_val)\nval_sample = np.random.choice(n_val, 10)\nplot_result(X_val[val_sample], Y_val[val_sample], Y_hat[val_sample], cutoff=best_threshold)","metadata":{"papermill":{"duration":16.322883,"end_time":"2024-05-19T23:15:50.003288","exception":false,"start_time":"2024-05-19T23:15:33.680405","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-31T09:30:09.992026Z","iopub.execute_input":"2024-05-31T09:30:09.992414Z","iopub.status.idle":"2024-05-31T09:30:14.176976Z","shell.execute_reply.started":"2024-05-31T09:30:09.992387Z","shell.execute_reply":"2024-05-31T09:30:14.175623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}