{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\nThis notebook is inspired by the discussion (https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/419309). \n\nAccording to the [competition info](https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/overview/evaluation), the evaluation code is public available.\n\nThe idea is to make the validation simple:\n\n- Run your inference pipeline on your validation set and write the results in the submission format.\n- Upload the validation output in submission format as a dataset and update the path in this notebook.\n- Execute this notebook to get the metrics.\n\nSince the upload-evaluate process may be annoying, I have put notes for how to run the same evaluation locally. ","metadata":{}},{"cell_type":"markdown","source":"# Install Dependencies\nFor local test, you can install the dependencies and comment out this block.\n\n`protobuf-compiler` is required in my ubuntu env to run `protoc`","metadata":{}},{"cell_type":"code","source":"# Eval Dependencies\n!pip install -q /kaggle/input/oid-eval-dependencies/tensorflow_io-0.31.0-cp310-cp310-manylinux_2_12_x86_64.manylinux2010_x86_64.whl\n!pip install -q /kaggle/input/pycocotools-206/wheels/pycocotools-2.0.6-cp310-cp310-linux_x86_64.whl\n\n!protoc --version","metadata":{"execution":{"iopub.status.busy":"2023-07-10T19:50:21.177376Z","iopub.execute_input":"2023-07-10T19:50:21.178446Z","iopub.status.idle":"2023-07-10T19:51:28.692141Z","shell.execute_reply.started":"2023-07-10T19:50:21.178409Z","shell.execute_reply":"2023-07-10T19:51:28.690681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Install tensorflow models for the OpenImages Chanllenge Evaluator\nIn order to use the evaluator, we need to run the `protoc` command for the proto files, which is not allowed in the read only kaggle input dataset. Therefore, we copy the input to the working directory.\n\nFor local test, you can just clone the `tensorflow models` repo from https://github.com/tensorflow/models and run the `protoc object_detection/protos/*.proto --python_out=.` in the `research` sub dir in your local repo. ","metadata":{}},{"cell_type":"code","source":"!cp -r  /kaggle/input/tf-models-v213 /kaggle/working/tf-models-v213\n!cd /kaggle/working/tf-models-v213/research && protoc object_detection/protos/*.proto --python_out=.","metadata":{"execution":{"iopub.status.busy":"2023-07-10T19:51:28.695821Z","iopub.execute_input":"2023-07-10T19:51:28.696364Z","iopub.status.idle":"2023-07-10T19:51:36.175361Z","shell.execute_reply.started":"2023-07-10T19:51:28.696297Z","shell.execute_reply":"2023-07-10T19:51:36.173700Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Import OpenImages Challenge Evaluator\nFor local test, update the `TF_MODEL_REPO_PATH` to your cloned repo path. This block should not throw any error if you have all the dependencies and finish the `protoc` command. ","metadata":{}},{"cell_type":"code","source":"import sys\n\n# Change this to your cloned tensorflow models repo\nTF_MODEL_REPO_PATH = \"/kaggle/working/tf-models-v213/research\"\n\nsys.path.append(TF_MODEL_REPO_PATH)\nfrom object_detection.utils.object_detection_evaluation import (\n    OpenImagesInstanceSegmentationChallengeEvaluator,\n)\nfrom object_detection.metrics.oid_challenge_evaluation_utils import (\n    build_groundtruth_dictionary,\n    build_predictions_dictionary,\n)","metadata":{"execution":{"iopub.status.busy":"2023-07-10T19:51:36.177330Z","iopub.execute_input":"2023-07-10T19:51:36.177723Z","iopub.status.idle":"2023-07-10T19:51:36.188241Z","shell.execute_reply.started":"2023-07-10T19:51:36.177689Z","shell.execute_reply":"2023-07-10T19:51:36.187041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tqdm import tqdm\nimport pandas as pd\nimport cv2\nimport json","metadata":{"execution":{"iopub.status.busy":"2023-07-10T19:51:36.189861Z","iopub.execute_input":"2023-07-10T19:51:36.190855Z","iopub.status.idle":"2023-07-10T19:51:36.198783Z","shell.execute_reply.started":"2023-07-10T19:51:36.190821Z","shell.execute_reply":"2023-07-10T19:51:36.197521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Global Configs","metadata":{}},{"cell_type":"code","source":"# Class, label conversion directories. Probably don't need to change these unless you need to add unsure class.\nOID_CATEGORIES = [{\"id\": 1, \"name\": \"blood_vessel\"}, {\"id\": 2, \"name\": \"glomerulus\"}]\nOID_CLASS_LABEL_MAP = {\"blood_vessel\": 1, \"glomerulus\": 2}\n\nSUBM_TO_CLASS_LABEL = {\"0\": \"blood_vessel\", \"1\": \"glomerulus\"}\nGT_CATEGORIES = [\"blood_vessel\", \"glomerulus\"]\n\n\n# The validation output generated by you inference pipeline in the submission format. I have provided a small sample from dataset 1 for testing.\n# Update this to your output file.\nVAL_SUBM_PATH = \"/kaggle/input/hubmap-eval-ds1-sample-subm/sample_submission_ds1_val_fold1.csv\"\n\n# Input image directory and annoations from the competition dataset. Only need to be updated for local evaluation.\nTRAIN_IMAGE_DIR = \"/kaggle/input/hubmap-hacking-the-human-vasculature/train\" \nPOLYGON_JSON_PATH = \"/kaggle/input/hubmap-hacking-the-human-vasculature/polygons.jsonl\"\n\n# Evaluation IOU thresholds. 0.5 and 0.75 are common in other evaluators and they are only used for checking the correctness. Use 0.6 for the competition evaluation. \nOID_IOUS = [0.5,0.6,0.75]\n# OID_IOUS = [0.6]\n","metadata":{"execution":{"iopub.status.busy":"2023-07-10T19:51:36.201871Z","iopub.execute_input":"2023-07-10T19:51:36.202224Z","iopub.status.idle":"2023-07-10T19:51:36.213222Z","shell.execute_reply.started":"2023-07-10T19:51:36.202196Z","shell.execute_reply":"2023-07-10T19:51:36.212056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Utils","metadata":{}},{"cell_type":"code","source":"import base64\nimport numpy as np\nfrom pycocotools import _mask as coco_mask\nimport typing as t\nimport zlib\n\ndef encode_binary_mask(mask: np.ndarray) -> t.Text:\n  \"\"\"Converts a binary mask into OID challenge encoding ascii text.\"\"\"\n\n  # check input mask --\n  if mask.dtype != bool:\n    raise ValueError(\n        \"encode_binary_mask expects a binary mask, received dtype == %s\" %\n        mask.dtype)\n\n  mask = np.squeeze(mask)\n  if len(mask.shape) != 2:\n    raise ValueError(\n        \"encode_binary_mask expects a 2d mask, received shape == %s\" %\n        mask.shape)\n\n  # convert input mask to expected COCO API input --\n  mask_to_encode = mask.reshape(mask.shape[0], mask.shape[1], 1)\n  mask_to_encode = mask_to_encode.astype(np.uint8)\n  mask_to_encode = np.asfortranarray(mask_to_encode)\n\n  # RLE encode mask --\n  encoded_mask = coco_mask.encode(mask_to_encode)[0][\"counts\"]\n\n  # compress and base64 encoding --\n  binary_str = zlib.compress(encoded_mask, zlib.Z_BEST_COMPRESSION)\n  base64_str = base64.b64encode(binary_str)\n  return base64_str","metadata":{"_kg_hide-output":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-10T19:51:36.214737Z","iopub.execute_input":"2023-07-10T19:51:36.215052Z","iopub.status.idle":"2023-07-10T19:51:36.229303Z","shell.execute_reply.started":"2023-07-10T19:51:36.215025Z","shell.execute_reply":"2023-07-10T19:51:36.228494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Utils for generating the ground truth in OpenImages format\nModified from: \n- https://www.kaggle.com/code/benihime91/hubmap-2023-create-coco-annotations\n- https://www.kaggle.com/code/ammarnassanalhajali/hubmap-2023-k-fold-cv-coco-dataset-generator","metadata":{}},{"cell_type":"code","source":"def coordinates_to_masks(coordinates, shape):\n    masks = []\n    for coord in coordinates:\n        mask = np.zeros(shape, dtype=np.uint8)\n        cv2.fillPoly(mask, [np.array(coord)], 1)\n        masks.append(mask)\n    return masks\n\ndef binary_mask_to_rle(binary_mask):\n    rle = {'counts': [], 'size': list(binary_mask.shape)}\n    counts = rle.get('counts')\n    for i, (value, elements) in enumerate(itertools.groupby(binary_mask.ravel(order='F'))):\n        if i == 0 and value == 1:\n            counts.append(0)\n        counts.append(len(list(elements)))\n    return rle\n\n\ndef generate_oid_gt_df(images_ids, polygon_json_path):\n    polygons_data = []        \n    \n    num_val_images = len(images_ids)\n    num_val_images_in_polygons = 0\n    \n    # Load polygons for the images in the validation list.\n    with open(polygon_json_path, \"r\") as file:\n        for line in file:\n            item = json.loads(line)\n            if item[\"id\"] in images_ids:\n                polygons_data.append(json.loads(line))\n                num_val_images_in_polygons +=1\n    assert num_val_images == num_val_images_in_polygons\n    \n    image_id_list = []\n    label_name = []\n    xmin = []\n    xmax = []\n    ymin = []\n    ymax = []\n    mask = []\n    mask_size = []\n    \n    IMG_WIDTH = 512\n    IMG_HEIGHT = 512\n\n    for item in tqdm(polygons_data):\n        image_id = item[\"id\"]\n        anns = item[\"annotations\"]\n        empty = True\n        for an in anns:\n            category_type = an[\"type\"]\n            if category_type in GT_CATEGORIES:\n                empty = False\n                segmentation = an[\"coordinates\"]\n                mask_img = coordinates_to_masks(segmentation, (IMG_HEIGHT, IMG_WIDTH))[0]\n                ys, xs = np.where(mask_img)\n                x1, x2 = min(xs), max(xs)\n                y1, y2 = min(ys), max(ys)\n\n                mask_img = mask_img.astype(bool)\n                mask_size.append(mask_img.sum())\n                encoded_mask = encode_binary_mask(mask_img).decode()\n\n                image_id_list.append(image_id)\n                label_name.append(category_type)\n                xmin.append(int(x1) / IMG_WIDTH)\n                xmax.append(int(x2) / IMG_WIDTH)\n                ymin.append(int(y1) / IMG_HEIGHT)\n                ymax.append(int(y2) / IMG_HEIGHT)\n                mask.append(encoded_mask)\n\n        # Make sure empty images are also added to the ground truth dataset for false positive evaluation.\n        if empty:\n            image_id_list.append(image_id)\n            label_name.append(categories_list[0])\n            xmin.append(None)\n            xmax.append(None)\n            ymin.append(None)\n            ymax.append(None)\n            mask.append(None)\n\n    num = len(image_id_list)\n    is_group_of = [0] * num\n    confidence = [1] * num\n    image_width = [IMG_WIDTH] * num\n    image_height = [IMG_HEIGHT] * num\n\n    df_data = {\n        \"ImageID\": image_id_list,\n        \"LabelName\": label_name,\n        \"XMin\": xmin,\n        \"XMax\": xmax,\n        \"YMin\": ymin,\n        \"YMax\": ymax,\n        \"IsGroupOf\": is_group_of,\n        \"Confidence\": confidence,\n        \"ImageWidth\": image_width,\n        \"ImageHeight\": image_height,\n        \"Mask\": mask,\n    }\n    return pd.DataFrame(df_data)","metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"execution":{"iopub.status.busy":"2023-07-10T19:51:36.230812Z","iopub.execute_input":"2023-07-10T19:51:36.231111Z","iopub.status.idle":"2023-07-10T19:51:36.254844Z","shell.execute_reply.started":"2023-07-10T19:51:36.231085Z","shell.execute_reply":"2023-07-10T19:51:36.253951Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Convert the submission file to the format for OpenImages Challenge Evaluation","metadata":{}},{"cell_type":"code","source":"def subm_to_pred_df(subm_df):\n    ids = []\n    label = []\n    image_width = []\n    image_height = []\n    score = []\n    mask = []\n    for index, row in subm_df.iterrows():\n        pred_strs = row[\"prediction_string\"].split(\" \")\n        labels_str = pred_strs[0::3]\n        scores = pred_strs[1::3]\n        img_masks = pred_strs[2::3]\n        assert len(scores) == len(img_masks)\n        for i in range(len(scores)):\n            ids.append(row[\"id\"])\n            label.append(SUBM_TO_CLASS_LABEL[labels_str[i]])\n            image_width.append(row[\"width\"])\n            image_height.append(row[\"height\"])\n            score.append(scores[i])\n            mask.append(img_masks[i])\n\n    pred_df = pd.DataFrame(\n        {\n            \"ImageID\": ids,\n            \"LabelName\": label,\n            \"ImageWidth\": image_width,\n            \"ImageHeight\": image_height,\n            \"Score\": score,\n            \"Mask\": mask,\n        }\n    )\n\n    return pred_df","metadata":{"execution":{"iopub.status.busy":"2023-07-10T19:51:36.256163Z","iopub.execute_input":"2023-07-10T19:51:36.256827Z","iopub.status.idle":"2023-07-10T19:51:36.272080Z","shell.execute_reply.started":"2023-07-10T19:51:36.256797Z","shell.execute_reply":"2023-07-10T19:51:36.270780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Evaluation function \nModified from: https://github.com/tensorflow/models/blob/master/research/object_detection/metrics/oid_challenge_evaluation.py","metadata":{}},{"cell_type":"code","source":"def get_oid_metrics(gt_df, pred_df, iou_thresholds=[0.5, 0.6, 0.75]):\n    # print the number of imageID in gt_df and pred_df\n    print(\"Number gt_df imageID: \", len(gt_df.groupby(\"ImageID\")))\n    print(\"Number of pred_df imageID: \", len(pred_df.groupby(\"ImageID\")))\n\n    gt_df = gt_df.loc[gt_df[\"ImageID\"].isin(pred_df[\"ImageID\"].unique())]\n    gt_df.rename(columns={\"Confidence\": \"ConfidenceImageLabel\"}, inplace=True)\n\n    gt_dicts = {}\n    pred_dicts = {}\n    print(\"Building groundtruth and prediction dictionary for oid evaluator\")\n    for _, groundtruth in tqdm(\n        enumerate(gt_df.groupby(\"ImageID\")), total=len(gt_df.groupby(\"ImageID\"))\n    ):\n        image_id, image_groundtruth = groundtruth\n        groundtruth_dictionary = build_groundtruth_dictionary(\n            image_groundtruth, OID_CLASS_LABEL_MAP\n        )\n        gt_dicts[image_id] = groundtruth_dictionary\n        prediction_dictionary = build_predictions_dictionary(\n            pred_df.loc[pred_df[\"ImageID\"] == image_id], OID_CLASS_LABEL_MAP\n        )\n        pred_dicts[image_id] = prediction_dictionary\n\n    print(\"Calculating metrics for different iou_thresholds\")\n    metrics_list = []\n    for iou_threshold in iou_thresholds:\n        challenge_evaluator = OpenImagesInstanceSegmentationChallengeEvaluator(\n            OID_CATEGORIES, matching_iou_threshold=iou_threshold\n        )\n        print(\"Iou_threshold: \", iou_threshold)\n        for _, groundtruth in tqdm(\n            enumerate(gt_df.groupby(\"ImageID\")), total=len(gt_df.groupby(\"ImageID\"))\n        ):\n            image_id, image_groundtruth = groundtruth\n\n            challenge_evaluator.add_single_ground_truth_image_info(\n                image_id, gt_dicts[image_id]\n            )\n            image_id, image_groundtruth = groundtruth\n            challenge_evaluator.add_single_detected_image_info(\n                image_id, pred_dicts[image_id]\n            )\n\n        metrics = challenge_evaluator.evaluate()\n        metrics_list.append(metrics)\n    return metrics_list","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-10T19:51:36.274023Z","iopub.execute_input":"2023-07-10T19:51:36.274499Z","iopub.status.idle":"2023-07-10T19:51:36.293324Z","shell.execute_reply.started":"2023-07-10T19:51:36.274456Z","shell.execute_reply":"2023-07-10T19:51:36.292405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Main","metadata":{}},{"cell_type":"code","source":"# Read the inference result for your validation set in the competition submission format\nval_subm_df = pd.read_csv(VAL_SUBM_PATH) \nval_subm_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-10T19:51:36.295057Z","iopub.execute_input":"2023-07-10T19:51:36.296108Z","iopub.status.idle":"2023-07-10T19:51:36.334638Z","shell.execute_reply.started":"2023-07-10T19:51:36.296075Z","shell.execute_reply":"2023-07-10T19:51:36.333674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the image list from the validation output\nval_image_ids = val_subm_df[\"id\"].to_list()\nprint(len(val_image_ids))","metadata":{"execution":{"iopub.status.busy":"2023-07-10T19:51:36.335892Z","iopub.execute_input":"2023-07-10T19:51:36.336920Z","iopub.status.idle":"2023-07-10T19:51:36.342188Z","shell.execute_reply.started":"2023-07-10T19:51:36.336887Z","shell.execute_reply":"2023-07-10T19:51:36.341406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Generate the ground truth dataframe in OpenImages Challenge format for the validation images.\n# You can also dump this to a csv file and reload it to avoid running the generation function again.\ngt_df = generate_oid_gt_df(val_image_ids, POLYGON_JSON_PATH)\ngt_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-10T19:51:36.343623Z","iopub.execute_input":"2023-07-10T19:51:36.344180Z","iopub.status.idle":"2023-07-10T19:51:42.806969Z","shell.execute_reply.started":"2023-07-10T19:51:36.344151Z","shell.execute_reply":"2023-07-10T19:51:42.805765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Convert the validation output to OpenImages Challenge submission format\npred_df = subm_to_pred_df(val_subm_df)\npred_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-10T19:51:42.808474Z","iopub.execute_input":"2023-07-10T19:51:42.808892Z","iopub.status.idle":"2023-07-10T19:51:42.996687Z","shell.execute_reply.started":"2023-07-10T19:51:42.808854Z","shell.execute_reply":"2023-07-10T19:51:42.995576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate OpenImages metrics.\nmetrics_list = get_oid_metrics(gt_df, pred_df, iou_thresholds=OID_IOUS)","metadata":{"execution":{"iopub.status.busy":"2023-07-10T19:51:43.000734Z","iopub.execute_input":"2023-07-10T19:51:43.001087Z","iopub.status.idle":"2023-07-10T19:53:11.459893Z","shell.execute_reply.started":"2023-07-10T19:51:43.001057Z","shell.execute_reply":"2023-07-10T19:53:11.458702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Print the metrics\nfor metrics in metrics_list:\n    for item in metrics:\n        print(\"--\", item, metrics[item])\n    print(\" \")","metadata":{"execution":{"iopub.status.busy":"2023-07-10T19:53:11.461601Z","iopub.execute_input":"2023-07-10T19:53:11.462022Z","iopub.status.idle":"2023-07-10T19:53:11.469393Z","shell.execute_reply.started":"2023-07-10T19:53:11.461983Z","shell.execute_reply":"2023-07-10T19:53:11.468243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Notes","metadata":{}},{"cell_type":"markdown","source":"## Validation of results\n- I compared the results using the COCO `mAP_50` and `mAP_75` calculated from mmdetection and there are small differences.\n    - 0.681 -> 0.6823\n    - 0.245 -> 0.2429\n- See this discussion for more info about the difference: https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/419309#2316979\n\n- The mmdetection COCO metrics for the model:\n\n```\n+--------------+-------+--------+--------+-------+-------+-------+\n| category     | mAP   | mAP_50 | mAP_75 | mAP_s | mAP_m | mAP_l |\n+--------------+-------+--------+--------+-------+-------+-------+\n| blood_vessel | 0.319 | 0.681  | 0.245  | 0.261 | 0.397 | 0.525 |\n+--------------+-------+--------+--------+-------+-------+-------+\n```\n\n- OpenImages Metrics calculated from this notebook:\n\n```\n-- OpenImagesInstanceSegmentationChallenge_Precision/mAP@0.5IOU 0.3411267599635575\n-- OpenImagesInstanceSegmentationChallenge_PerformanceByCategory/AP@0.5IOU/blood_vessel 0.682253519927115\n-- OpenImagesInstanceSegmentationChallenge_PerformanceByCategory/AP@0.5IOU/glomerulus 0.0\n \n-- OpenImagesInstanceSegmentationChallenge_Precision/mAP@0.6IOU 0.2909974285533541\n-- OpenImagesInstanceSegmentationChallenge_PerformanceByCategory/AP@0.6IOU/blood_vessel 0.5819948571067082\n-- OpenImagesInstanceSegmentationChallenge_PerformanceByCategory/AP@0.6IOU/glomerulus 0.0\n \n-- OpenImagesInstanceSegmentationChallenge_Precision/mAP@0.75IOU 0.12145058973617051\n-- OpenImagesInstanceSegmentationChallenge_PerformanceByCategory/AP@0.75IOU/blood_vessel 0.24290117947234102\n-- OpenImagesInstanceSegmentationChallenge_PerformanceByCategory/AP@0.75IOU/glomerulus 0.0\n```\n- **Be careful about overfitting to WSI1 and WSI2.**\n- You may want to control the number of validation images(<500) since the metrics calculation requires some amount of memory. It should be totally fine with the 422 images from dataset 1.\n","metadata":{}},{"cell_type":"markdown","source":"\n### Let me know if there is a mistake in this notebook.","metadata":{}}]}