{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style=\"font-family: Verdana; font-size: 28px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; letter-spacing: 3px; background-color: #CCCCFF; color: black;\"><center><br>Understanding Evaluation Metrics in Medical Image Segmentation </center></h1>\n                                                      \n<center><img src = \"https://drive.google.com/uc?id=1pbIvjTlhGywfhiMTqcsdOB5LSHlklM90\"/></center>   \n\n<h5 style=\"text-align: center; font-family: Verdana; font-size: 12px; font-style: normal; font-weight: bold; text-decoration: None; text-transform: none; letter-spacing: 1px; color: black; background-color: #ffffff;\">CREATED BY: NGHI HUYNH</h5>","metadata":{}},{"cell_type":"markdown","source":"<p id=\"toc\"></p>\n<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style=\"font-family: Verdana; font-size: 24px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; letter-spacing: 3px; background-color: #CCCCFF; color: black;\" role=\"tab\" aria-controls=\"home\"><center><br>CONTENTS</center></h2>\n\n<h3 style=\"text-indent: 10vw; font-family: Verdana; font-size: 16px; font-style: normal; font-weight: normal; text-decoration: none; text-transform: none; letter-spacing: 2px; color: black; background-color: #ffffff;\"><a href=\"#motive\">0&nbsp;&nbsp;&nbsp;&nbsp;MOTIVATION</a></h3>\n\n---\n\n<h3 style=\"text-indent: 10vw; font-family: Verdana; font-size: 16px; font-style: normal; font-weight: normal; text-decoration: none; text-transform: none; letter-spacing: 2px; color: black; background-color: #ffffff;\"><a href=\"#ss\">1&nbsp;&nbsp;&nbsp;&nbsp;PRECISION & RECALL</a></h3>\n\n---\n\n<h3 style=\"text-indent: 10vw; font-family: Verdana; font-size: 16px; font-style: normal; font-weight: normal; text-decoration: none; text-transform: none; letter-spacing: 2px; color: black; background-color: #ffffff;\"><a href=\"#acc\">2&nbsp;&nbsp;&nbsp;&nbsp;ACCURACY/RAND INDEX</a></h3>\n\n---\n\n<h3 style=\"text-indent: 10vw; font-family: Verdana; font-size: 16px; font-style: normal; font-weight: normal; text-decoration: none; text-transform: none; letter-spacing: 2px; color: black; background-color: #ffffff;\"><a href=\"#dice\">3&nbsp;&nbsp;&nbsp;&nbsp;DICE COEFFICIENT (F1-SCORE)</a></h3>\n\n---\n\n<h3 style=\"text-indent: 10vw; font-family: Verdana; font-size: 16px; font-style: normal; font-weight: normal; text-decoration: none; text-transform: none; letter-spacing: 2px; color: black; background-color: #ffffff;\"><a href=\"#iou\">4&nbsp;&nbsp;&nbsp;&nbsp;JACCARD INDEX (IoU)</a></h3>\n\n---\n\n<h3 style=\"text-indent: 10vw; font-family: Verdana; font-size: 16px; font-style: normal; font-weight: normal; text-decoration: none; text-transform: none; letter-spacing: 2px; color: black; background-color: #ffffff;\"><a href=\"#conclusion\">5&nbsp;&nbsp;&nbsp;&nbsp;TESTING + CONCLUSION</a></h3>\n\n---","metadata":{}},{"cell_type":"markdown","source":"<h3 style='background:#F08080; border:0; color:black'><center><br>If you find this notebook useful, do give me an upvote, it motivates me a lot.<br><br> This notebook is still a work in progress. Keep checking for further developments!😊</center></h3>","metadata":{}},{"cell_type":"markdown","source":"<a id=\"motive\"></a>\n\n<h2 style=\"font-family: Verdana; font-size: 24px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; letter-spacing: 3px; background-color: #CCCCFF; color: black;\" id=\"motive\"><left><br>&nbsp;0. MOTIVATION <a href=\"#toc\">&#10514;</a><br></left> </h2>\n\n#### *Now that you have a trained model for your segmentation task? But how do you know if your segmentation model performs well? In other words, how do we evaluate our model performance?*\n\n\n**Evaluation metrics** are the answer!\n\nIn this notebook, I will provide an overview of appropriate, most common evaluation metrics, demonstrate their interpretation and implementation, and propose a guideline to properly evaluate medical image segmentation performance to increase research reliability and reproducibility in the field.\n\n**Goal:** to score the similarity between the predicted (prediction) and annotated segmentation (ground truth)\n\n5 evaluation metrics:\n* **Precision and Recall (Sensitivity)**\n* **Accuracy/Rand index**\n* **Dice coefficient**\n* **Jaccard index (IoU)**\n\nAll presented metrics are based on the computation of a **confusion matrix** for a binary segmentation mask, which contains the number of true positive (TP), false positive (FP), true negative (TN), and false negative (FN) predictions. The value ranges of all presented metrics span from zero (worst) to one (best).\n\n#### *What is a confusion matrix? And why is it important?*\n\nA picture is worth a thousand words!\n\n![](https://drive.google.com/uc?id=1IDnNqQoXlhRBbkFQZuVKrGzQFkcCeDHf)\n\nLet's interpret these terms in our context! Our goal in this challenge is to segment a mask of FTUs for each histopathological image. Each mask is represented either by 0-background or 1-FTU. So, \n\n* **TP (True Positive)**: represents the number of FTU pixels that have been properly classified as FTU\n* **FP (False Positive)**: represents the number background pixels being misclassified as FTUs (due to misalignment)\n* **FN (False Negative)**: represents the number of FTU pixels being misclassified as background\n* **TN(True Negative)**: represents the number of background pixels that have been properly classified as background\n\n> ### *Let's dive deeper into each evaluation metric, and build our understanding from simple illustrations!*","metadata":{}},{"cell_type":"markdown","source":"# Imports\n---","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport pandas as pd\nimport numpy as np\nimport cv2\nimport os\nimport gc\nimport tifffile as tiff\nfrom sklearn.metrics import recall_score, accuracy_score, precision_score, f1_score, jaccard_score","metadata":{"execution":{"iopub.status.busy":"2022-08-15T18:34:08.582122Z","iopub.execute_input":"2022-08-15T18:34:08.582622Z","iopub.status.idle":"2022-08-15T18:34:10.055186Z","shell.execute_reply.started":"2022-08-15T18:34:08.582522Z","shell.execute_reply":"2022-08-15T18:34:10.054222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Helper functions\n---","metadata":{}},{"cell_type":"code","source":"## We need to decode the mask from encoding column of train.csv\n## https://www.kaggle.com/paulorzp/rle-functions-run-lenght-encode-decode\ndef mask2rle(img):\n    '''\n    img: numpy array, 1 - mask, 0 - background\n    Returns run length as string formated\n    '''\n    pixels= img.T.flatten()\n    pixels = np.concatenate([[0], pixels, [0]])\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 1\n    runs[1::2] -= runs[::2]\n    return ' '.join(str(x) for x in runs)\n \ndef rle2mask(mask_rle, shape):\n    '''\n    mask_rle: run-length as string formated (start length)\n    shape: (width,height) of array to return \n    Returns numpy array, 1 - mask, 0 - background\n\n    '''\n    s = mask_rle.split()\n    starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\n    starts -= 1\n    ends = starts + lengths\n    #print(starts, ends)\n    img = np.zeros(shape[0]*shape[1], dtype=np.uint8)\n    for lo, hi in zip(starts, ends):\n        img[lo:hi] = 1\n    return img.reshape(shape).T\n\n# https://www.kaggle.com/code/yerramvarun/understanding-dice-coefficient\n# We are just shifting the images towards the bottom to keep it simple\ndef return_shifted(mask, shift=5):\n    nmask = np.zeros((mask.shape[0]+shift, mask.shape[1]))\n    nmask[shift:, :] = mask\n    nmask = nmask[:-shift, :]\n    return nmask","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-15T18:34:44.717652Z","iopub.execute_input":"2022-08-15T18:34:44.718082Z","iopub.status.idle":"2022-08-15T18:34:44.731187Z","shell.execute_reply.started":"2022-08-15T18:34:44.718040Z","shell.execute_reply":"2022-08-15T18:34:44.730198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# inspired by: https://www.kaggle.com/code/dschettler8845/visual-guide-to-understanding-segmentation-metrics/notebook\ndef grid_imshow(mask):\n    ax = plt.gca()\n    plt.imshow(mask, cmap=\"gray\" if len(mask.shape)==2 else None, interpolation=\"none\", vmin=0, vmax=1, aspect=\"equal\")\n    \n    if mask.shape[1] < 20:\n        # Major ticks\n        ax.set_xticks(np.arange(0, mask.shape[1], 1))\n        ax.set_yticks(np.arange(0, mask.shape[0], 1))\n\n        # Labels for major ticks\n        ax.set_xticklabels(np.arange(1, mask.shape[1]+1, 1))\n        ax.set_yticklabels(np.arange(1, mask.shape[0]+1, 1))\n\n        # Minor ticks\n        ax.set_xticks(np.arange(-.5, mask.shape[1], 1), minor=True)\n        ax.set_yticks(np.arange(-.5, mask.shape[0], 1), minor=True)\n        \n        # Gridlines based on minor ticks\n        ax.grid(which='minor', color='blue', linestyle='-', linewidth=2)\n\n    else:\n        # Major ticks\n        ax.set_xticks(np.arange(0, mask.shape[1], 500))\n        ax.set_yticks(np.arange(0, mask.shape[0], 500))\n\n        # Labels for major ticks\n        ax.set_xticklabels(np.arange(0, mask.shape[1], 500))\n        ax.set_yticklabels(np.arange(0, mask.shape[0], 500))\n\n        # Gridlines based on major ticks\n        ax.grid(which='major', color='blue', linestyle='-', linewidth=2)\n\ndef compare_masks(gt_mask, pred_mask, gt_title=\"Ground Truth Mask\", pred_title=\"Prediction Mask\", _figshape=(10,10)):\n    plt.figure(figsize=_figshape)\n\n    plt.subplot(1,2,1)    \n    grid_imshow(gt_mask)\n    plt.title(gt_title, fontweight=\"bold\")\n\n    plt.subplot(1,2,2)\n    grid_imshow(pred_mask)\n    plt.title(pred_title, fontweight=\"bold\")\n\n    plt.tight_layout()\n    plt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-15T18:34:51.302914Z","iopub.execute_input":"2022-08-15T18:34:51.303306Z","iopub.status.idle":"2022-08-15T18:34:51.317062Z","shell.execute_reply.started":"2022-08-15T18:34:51.303272Z","shell.execute_reply":"2022-08-15T18:34:51.316144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def highlight(row):\n    df = lambda x: ['background: #CCCCFF' if x.name in row\n                        else '' for i in x]\n    return df\n\ndef metrics_table(gt_masks, pred_masks):\n    metrics = {'Precision':[],'Recall':[],'Accuracy':[],'Dice':[],'IoU':[]}\n    for i, (mask, pred) in enumerate(zip(gt_masks, pred_masks)):\n        metrics['Precision'].append(precision_score_(mask, pred))\n        metrics['Recall'].append(recall_score_(mask, pred))\n        metrics['Accuracy'].append(accuracy(mask, pred))\n        metrics['Dice'].append(dice_coef(mask, pred))\n        metrics['IoU'].append(iou(mask, pred))\n    df = pd.DataFrame.from_dict(metrics)\n    df.columns = ['Precision', 'Recall', 'Accuracy', 'Dice', 'IoU']\n    return df","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-08-15T18:34:53.896031Z","iopub.execute_input":"2022-08-15T18:34:53.896422Z","iopub.status.idle":"2022-08-15T18:34:53.905661Z","shell.execute_reply.started":"2022-08-15T18:34:53.896388Z","shell.execute_reply":"2022-08-15T18:34:53.904708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Ground truth masks and prediction masks\n---","metadata":{}},{"cell_type":"markdown","source":"## Generate a simple ground truth + predicted masks:","metadata":{}},{"cell_type":"code","source":"DEMO_SHAPE = (5,5)\ngt_mask_0 = np.zeros(DEMO_SHAPE, dtype=np.float32)\nfor i in range(DEMO_SHAPE[0]):\n    gt_mask_0[i:i+2,i:i+2] = 1.0\npred_mask_0 = gt_mask_0.copy()\npred_mask_0[1:4,2:4] = 0.","metadata":{"execution":{"iopub.status.busy":"2022-08-15T18:35:08.359389Z","iopub.execute_input":"2022-08-15T18:35:08.359804Z","iopub.status.idle":"2022-08-15T18:35:08.367055Z","shell.execute_reply.started":"2022-08-15T18:35:08.359768Z","shell.execute_reply":"2022-08-15T18:35:08.365698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Retrieve ground truth + predicted masks from the challenge:","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv('../input/hubmap-organ-segmentation/train.csv')\ndf[10:15].drop('rle', axis=1).style.apply(highlight([10,12]), axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-15T18:35:11.426433Z","iopub.execute_input":"2022-08-15T18:35:11.426826Z","iopub.status.idle":"2022-08-15T18:35:11.861248Z","shell.execute_reply.started":"2022-08-15T18:35:11.426794Z","shell.execute_reply":"2022-08-15T18:35:11.859974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#img_1 = tiff.imread('../input/hubmap-organ-segmentation/train_images/10912.tiff')\ngt_mask_1 = rle2mask(df.loc[10].rle, (3000,3000))\n#img_2 = tiff.imread('../input/hubmap-organ-segmentation/train_images/10992.tiff')\ngt_mask_2 = rle2mask(df.loc[12].rle, (3000,3000))","metadata":{"execution":{"iopub.status.busy":"2022-08-15T18:35:14.498192Z","iopub.execute_input":"2022-08-15T18:35:14.498950Z","iopub.status.idle":"2022-08-15T18:35:14.519445Z","shell.execute_reply.started":"2022-08-15T18:35:14.498905Z","shell.execute_reply":"2022-08-15T18:35:14.518546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_mask_1 = return_shifted(gt_mask_1, shift=500)\npred_mask_2 = return_shifted(gt_mask_2, shift=100)","metadata":{"execution":{"iopub.status.busy":"2022-08-15T18:35:18.181729Z","iopub.execute_input":"2022-08-15T18:35:18.182153Z","iopub.status.idle":"2022-08-15T18:35:18.415349Z","shell.execute_reply.started":"2022-08-15T18:35:18.182116Z","shell.execute_reply":"2022-08-15T18:35:18.414258Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"ss\"></a>\n\n<h2 style=\"font-family: Verdana; font-size: 24px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; letter-spacing: 3px; background-color: #CCCCFF; color: black;\" id=\"ss\"><left><br>&nbsp;1. PRECISION & RECALL <a href=\"#toc\">&#10514;</a><br></left> </h2>\n\nFor pixel classification: \n\n**Precision** score is the number of true positive results divided by the number of all positive results\n\n#### $$Precision = \\frac{TP}{TP + FP}$$\n\n**Recall** score, also known as *Sensitivity* or *true positive rate*, is the number of true positive results divided by the number of all samples that should have been identified as positive\n#### $$Recall = \\frac{TP}{TP + FN}$$\n---\n","metadata":{}},{"cell_type":"code","source":"def precision_score_(groundtruth_mask, pred_mask):\n    intersect = np.sum(pred_mask*groundtruth_mask)\n    total_pixel_pred = np.sum(pred_mask)\n    precision = np.mean(intersect/total_pixel_pred)\n    return round(precision, 3)","metadata":{"execution":{"iopub.status.busy":"2022-08-15T18:35:21.078960Z","iopub.execute_input":"2022-08-15T18:35:21.080186Z","iopub.status.idle":"2022-08-15T18:35:21.085304Z","shell.execute_reply.started":"2022-08-15T18:35:21.080141Z","shell.execute_reply":"2022-08-15T18:35:21.084325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def recall_score_(groundtruth_mask, pred_mask):\n    intersect = np.sum(pred_mask*groundtruth_mask)\n    total_pixel_truth = np.sum(groundtruth_mask)\n    recall = np.mean(intersect/total_pixel_truth)\n    return round(recall, 3)","metadata":{"execution":{"iopub.status.busy":"2022-08-15T18:35:23.568317Z","iopub.execute_input":"2022-08-15T18:35:23.568730Z","iopub.status.idle":"2022-08-15T18:35:23.574487Z","shell.execute_reply.started":"2022-08-15T18:35:23.568696Z","shell.execute_reply":"2022-08-15T18:35:23.573320Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"acc\"></a>\n\n<h2 style=\"font-family: Verdana; font-size: 24px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; letter-spacing: 3px; background-color: #CCCCFF; color: black;\" id=\"acc\"><left><br>&nbsp;2. ACCURACY/RAND INDEX <a href=\"#toc\">&#10514;</a><br></left> </h2>\n\n**Accuracy** score, also known as *Rand index* is the number of correct predictions, consisting of correct positive and negative predictions divided by the total number of predictions.\n\n#### $$Accuracy = \\frac{TP + TN}{TP + TN + FN + FP}$$\n\n---","metadata":{}},{"cell_type":"code","source":"def accuracy(groundtruth_mask, pred_mask):\n    intersect = np.sum(pred_mask*groundtruth_mask)\n    union = np.sum(pred_mask) + np.sum(groundtruth_mask) - intersect\n    xor = np.sum(groundtruth_mask==pred_mask)\n    acc = np.mean(xor/(union + xor - intersect))\n    return round(acc, 3)","metadata":{"execution":{"iopub.status.busy":"2022-08-15T18:35:26.024204Z","iopub.execute_input":"2022-08-15T18:35:26.024609Z","iopub.status.idle":"2022-08-15T18:35:26.031477Z","shell.execute_reply.started":"2022-08-15T18:35:26.024575Z","shell.execute_reply":"2022-08-15T18:35:26.030556Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"dice\"></a>\n\n<h2 style=\"font-family: Verdana; font-size: 24px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; letter-spacing: 3px; background-color: #CCCCFF; color: black;\" id=\"dice\"><left><br>&nbsp;3. DICE COEFFICIENT (F1-SCORE) <a href=\"#toc\">&#10514;</a><br></left> </h2>\n\n\nF-measure, also called F-score: one of the most widespread scores for performance measuring in computer vision and in MIS (Medical Image Segmentation). \n\nDice coefficient is calculated from the **precision** and **recall** of a prediction. Then, it scores the overlap between predicted segmentation and ground truth. It also penalize false positives, which is a common factor in highly class imbalanced datasets like MIS. \n\nBased on the F-measure, there are two popular utilized metrics in MIS: \n\n* The **Intersection-over-Union (IoU)**, also known as **Jaccard index** or Jaccard similarity coefficient\n\n* The **Dice similarity coefficient (DSC)**, also known as **F1-score** or Sørensen-Dice index: most used metric in the large majority of scientific publictions for MIS evaluation\n\n=> The difference between the two metrics is that the IoU penalizes under- and over-segmentation more than DSC. \n\n## Dice coefficient (for HuBMAP+HPA competition)\n\n**Dice coefficient** = **F1 score**: a harmonic mean of precision and recall. In other words, it is calculated by 2\\*intersection divided by the total number of pixel in both images. Note that it's possible to adjust the F-score to give more importance to precision over recall, or vice-versa. Refer to [F-beta score](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.fbeta_score.html#sklearn.metrics.fbeta_score) for more info.\n\n#### $$Dice = \\frac{2TP}{2TP + FP + FN}$$\n\n![](https://drive.google.com/uc?id=1xRya3Z9IA6t3K48OZBx7XQ2-S7QjVJsJ)\n\n---","metadata":{}},{"cell_type":"code","source":"def dice_coef(groundtruth_mask, pred_mask):\n    intersect = np.sum(pred_mask*groundtruth_mask)\n    total_sum = np.sum(pred_mask) + np.sum(groundtruth_mask)\n    dice = np.mean(2*intersect/total_sum)\n    return round(dice, 3) #round up to 3 decimal places","metadata":{"execution":{"iopub.status.busy":"2022-08-15T18:35:29.935714Z","iopub.execute_input":"2022-08-15T18:35:29.936881Z","iopub.status.idle":"2022-08-15T18:35:29.943661Z","shell.execute_reply.started":"2022-08-15T18:35:29.936822Z","shell.execute_reply":"2022-08-15T18:35:29.942551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"iou\"></a>\n\n<h2 style=\"font-family: Verdana; font-size: 24px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; letter-spacing: 3px; background-color: #CCCCFF; color: black;\" id=\"iou\"><left><br>&nbsp;4. Jaccard Index (IoU) <a href=\"#toc\">&#10514;</a><br></left> </h2>\n\n**Jaccard index**, also known as **Intersection over Union (IoU)**, is the area of the intersection over union of the predicted segmentation and the ground truth\n\n#### $$IoU = \\frac{TP}{TP + FP + FN}$$\n\n![](https://drive.google.com/uc?id=18aEXA1l5gAzee1NuHbZ5SsJNt9YKif5x)\n\n---","metadata":{}},{"cell_type":"code","source":"def iou(groundtruth_mask, pred_mask):\n    intersect = np.sum(pred_mask*groundtruth_mask)\n    union = np.sum(pred_mask) + np.sum(groundtruth_mask) - intersect\n    iou = np.mean(intersect/union)\n    return round(iou, 3)","metadata":{"execution":{"iopub.status.busy":"2022-08-15T18:35:33.228875Z","iopub.execute_input":"2022-08-15T18:35:33.229310Z","iopub.status.idle":"2022-08-15T18:35:33.234821Z","shell.execute_reply.started":"2022-08-15T18:35:33.229267Z","shell.execute_reply":"2022-08-15T18:35:33.233403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"conclusion\"></a>\n\n<h2 style=\"font-family: Verdana; font-size: 24px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; letter-spacing: 3px; background-color: #CCCCFF; color: black;\" id=\"conclusion\"><left><br>&nbsp;5. TESTING + CONCLUSION <a href=\"#toc\">&#10514;</a><br></left> </h2>\n","metadata":{}},{"cell_type":"markdown","source":"# Testing\n---\nLet's put it all together and see how one metric is more suitable for a segmentation task than another!","metadata":{}},{"cell_type":"markdown","source":"## Simple ground truth + predicted masks","metadata":{}},{"cell_type":"code","source":"compare_masks(gt_mask_0, pred_mask_0)","metadata":{"execution":{"iopub.status.busy":"2022-08-15T18:35:37.285455Z","iopub.execute_input":"2022-08-15T18:35:37.286233Z","iopub.status.idle":"2022-08-15T18:35:37.707455Z","shell.execute_reply.started":"2022-08-15T18:35:37.286191Z","shell.execute_reply":"2022-08-15T18:35:37.706306Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metrics_table([gt_mask_0], [pred_mask_0])","metadata":{"execution":{"iopub.status.busy":"2022-08-15T18:35:40.443946Z","iopub.execute_input":"2022-08-15T18:35:40.445105Z","iopub.status.idle":"2022-08-15T18:35:40.464549Z","shell.execute_reply.started":"2022-08-15T18:35:40.445049Z","shell.execute_reply":"2022-08-15T18:35:40.463656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Observations:\n\n* Precision score = 1! What's an awesome prediction! But, wait! Isn't the predicted mask differed from the ground truth by 5 pixels? Is there anything wrong in this metric? Let's recall that precision score is the total number of true positives divided by the total number of all positive results in the prediction. Coincidently, those numbers are equal (=8), Thus, the precision score = 1. \n* Recall = IoU\n* Dice > IoU","metadata":{}},{"cell_type":"markdown","source":"## Ground truth + predicted masks from the challenge","metadata":{}},{"cell_type":"code","source":"compare_masks(gt_mask_1, pred_mask_1)\ncompare_masks(gt_mask_2, pred_mask_2)","metadata":{"execution":{"iopub.status.busy":"2022-08-15T18:35:44.387905Z","iopub.execute_input":"2022-08-15T18:35:44.388747Z","iopub.status.idle":"2022-08-15T18:35:45.758390Z","shell.execute_reply.started":"2022-08-15T18:35:44.388700Z","shell.execute_reply":"2022-08-15T18:35:45.757349Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metrics_table([gt_mask_1, gt_mask_2], [pred_mask_1, pred_mask_2])","metadata":{"execution":{"iopub.status.busy":"2022-08-15T18:35:48.755802Z","iopub.execute_input":"2022-08-15T18:35:48.756216Z","iopub.status.idle":"2022-08-15T18:35:49.963284Z","shell.execute_reply.started":"2022-08-15T18:35:48.756182Z","shell.execute_reply":"2022-08-15T18:35:49.961962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Observations: \n* High accuracy, but low Dice and IoU score. Why is it? Let's recall how the accuracy score is calculated. The accuracy score is the total number of correctly predicted pixels (1-FTU, 0-background) divided by the total number of predictions. We can notice that there is only a small percentage of pixels of the FTUs compared to the background. Thus, if we use the accuracy score, those correctly classified backgrounds will dominate the whole numerator term and result in an illegitimate high scoring.\n* Precision = Recall = Dice\n* Dice coefficient > IoU","metadata":{}},{"cell_type":"markdown","source":"# Conclusion\n---\n\nIn this notebook, I've demonstrated 5 evaluation metrics in Medical Image Segmentation (MIS).\n\n* **Precision and Recall (Sensitivity)**\n* **Accuracy/Rand index**\n* **Dice coefficient**\n* **Jaccard index (IoU)**\n\nHowever, it's strongly **discouraged** to use **accuracy** in MIS since MIS has highly imbalanced classes between ROIs vs. background. In other words, medical segmentation image usually contains a **small percentage** of pixels in the **ROIs**, whereas the remaining image is all annotated as background. Accuracy score includes true negative results, thus will always result in an illegitimate high scoring.\n\nIn contrast, **dice coefficient** and **IoU** are the most commonly used metrics for semantic segmentation because both metrics **penalize** false positives, which is a common factor in highly class imbalanced datasets like MIS. However, choosing dice coefficient over IoU or vice versa is based on specific use cases of the task.\n\n# References: \n* [Metrics to evaluate your semantic segmentation](https://towardsdatascience.com/metrics-to-evaluate-your-semantic-segmentation-model-6bcb99639aa2)\n* [What is the F-score?](https://deepai.org/machine-learning-glossary-and-terms/f-score)\n* [Classification: Precision and Recall](https://developers.google.com/machine-learning/crash-course/classification/precision-and-recall)\n* [Understanding DICE COEFFICIENT](https://www.kaggle.com/code/yerramvarun/understanding-dice-coefficient)\n* [Visual Guide to Understanding Segmentation Metrics](https://www.kaggle.com/code/dschettler8845/visual-guide-to-understanding-segmentation-metrics)","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}