{
  "id": 416452,
  "title": "How far can we push SAM? [LB 0.372]",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/discussion/416452",
  "author_name": "",
  "post_date": "2023-06-11T14:11:14.689777300Z",
  "votes": 11,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I'm a big fan of <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's <a href=\"https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332941\" target=\"_blank\">public</a> <a href=\"https://www.kaggle.com/competitions/champs-scalar-coupling/discussion/93972\" target=\"_blank\">solution</a> <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333\" target=\"_blank\">threads</a>, and always find them to be full of useful information, so in that spirit I will try a solution and keep track of it publicly, at least for a while.     </p>\n<p>Since Meta released their <a href=\"https://github.com/facebookresearch/segment-anything/tree/main\" target=\"_blank\">SegmentAnything Model</a> (SAM), I've been looking for an excuse to play around with it, so that's what I'll do here.   </p>\n<p>So the interesting thing about SAM, is that it's been designed to be prompted, i.e. given an image and a point/box/mask, it will return a mask. In the <a href=\"https://ai.facebook.com/research/publications/segment-anything/\" target=\"_blank\">SAM paper </a>, they point out this can be done by hand or programmatically. </p>\n<p>Naturally the latter is more relevant here, so how about we train a lightweight object detection model (e.g. YOLO) to find objects, and then use SAM to refine the masks? The weakness here might be that we rely quite heavily on the object detector to do a good job first. Is it even necessary? As SAM can do <a href=\"https://github.com/facebookresearch/segment-anything/blob/main/segment_anything/automatic_mask_generator.py\" target=\"_blank\">automatic mask generation</a>. </p>\n<p>My first attempt <a href=\"https://www.kaggle.com/code/fnands/yolov7-sam-inference-only\" target=\"_blank\">is here</a>, and naively takes a mask and box from YOLOv7 as prompts, and then tries do predict a segmentation masks. <br>\nSo far, it does worse than just using using the YOLOv7 masks directly (LB 0.25). </p>\n<p>Next, some fine-tuning of SAM? </p>",
  "messages": [
    {
      "id": "2296060",
      "postDate": "06/11/2023 14:11:14",
      "content": "<p>I'm a big fan of <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's <a href=\"https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332941\" target=\"_blank\">public</a> <a href=\"https://www.kaggle.com/competitions/champs-scalar-coupling/discussion/93972\" target=\"_blank\">solution</a> <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333\" target=\"_blank\">threads</a>, and always find them to be full of useful information, so in that spirit I will try a solution and keep track of it publicly, at least for a while.     </p>\n<p>Since Meta released their <a href=\"https://github.com/facebookresearch/segment-anything/tree/main\" target=\"_blank\">SegmentAnything Model</a> (SAM), I've been looking for an excuse to play around with it, so that's what I'll do here.   </p>\n<p>So the interesting thing about SAM, is that it's been designed to be prompted, i.e. given an image and a point/box/mask, it will return a mask. In the <a href=\"https://ai.facebook.com/research/publications/segment-anything/\" target=\"_blank\">SAM paper </a>, they point out this can be done by hand or programmatically. </p>\n<p>Naturally the latter is more relevant here, so how about we train a lightweight object detection model (e.g. YOLO) to find objects, and then use SAM to refine the masks? The weakness here might be that we rely quite heavily on the object detector to do a good job first. Is it even necessary? As SAM can do <a href=\"https://github.com/facebookresearch/segment-anything/blob/main/segment_anything/automatic_mask_generator.py\" target=\"_blank\">automatic mask generation</a>. </p>\n<p>My first attempt <a href=\"https://www.kaggle.com/code/fnands/yolov7-sam-inference-only\" target=\"_blank\">is here</a>, and naively takes a mask and box from YOLOv7 as prompts, and then tries do predict a segmentation masks. <br>\nSo far, it does worse than just using using the YOLOv7 masks directly (LB 0.25). </p>\n<p>Next, some fine-tuning of SAM? </p>",
      "rawMarkdown": "I'm a big fan of @hengck23 's [public](https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332941) [solution](https://www.kaggle.com/competitions/champs-scalar-coupling/discussion/93972) [threads](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333), and always find them to be full of useful information, so in that spirit I will try a solution and keep track of it publicly, at least for a while.     \n\nSince Meta released their [SegmentAnything Model](https://github.com/facebookresearch/segment-anything/tree/main) (SAM), I've been looking for an excuse to play around with it, so that's what I'll do here.   \n\nSo the interesting thing about SAM, is that it's been designed to be prompted, i.e. given an image and a point/box/mask, it will return a mask. In the [SAM paper ](https://ai.facebook.com/research/publications/segment-anything/), they point out this can be done by hand or programmatically. \n\nNaturally the latter is more relevant here, so how about we train a lightweight object detection model (e.g. YOLO) to find objects, and then use SAM to refine the masks? The weakness here might be that we rely quite heavily on the object detector to do a good job first. Is it even necessary? As SAM can do [automatic mask generation](https://github.com/facebookresearch/segment-anything/blob/main/segment_anything/automatic_mask_generator.py). \n\nMy first attempt [is here](https://www.kaggle.com/code/fnands/yolov7-sam-inference-only), and naively takes a mask and box from YOLOv7 as prompts, and then tries do predict a segmentation masks. \nSo far, it does worse than just using using the YOLOv7 masks directly (LB 0.25). \n\nNext, some fine-tuning of SAM?",
      "votes": null
    },
    {
      "id": "2305674",
      "postDate": "06/16/2023 20:46:58",
      "content": "<p>Adding dilation 0.148 -&gt; 0.195, without any other changes. <br>\nGot the tip from: <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/416901\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/416901</a></p>",
      "rawMarkdown": "Adding dilation 0.148 -> 0.195, without any other changes. \nGot the tip from: https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/416901",
      "votes": null
    },
    {
      "id": "2307105",
      "postDate": "06/17/2023 20:14:41",
      "content": "<p>Fine-tuning the mask decoder on dataset 1 takes me from 0.195 -&gt; 0.364. <br>\nWas a bit of a pain in the ass to get training, but clipping the gradients helped a lot. <br>\nHaven't implemented CV, which is bad practice. Will do so now to get an idea of the difference between CV and LB. </p>",
      "rawMarkdown": "Fine-tuning the mask decoder on dataset 1 takes me from 0.195 -> 0.364. \nWas a bit of a pain in the ass to get training, but clipping the gradients helped a lot. \nHaven't implemented CV, which is bad practice. Will do so now to get an idea of the difference between CV and LB.",
      "votes": null
    },
    {
      "id": "2307435",
      "postDate": "06/18/2023 07:33:20",
      "content": "<p>Funnily, training on dataset 1 and 2 lowers the LB score compared to training on dataset 1 only. 0.364 -&gt; 0.352. </p>\n<p>This does make sense, especially when considering we are trying to get the annotation quality to match dataset 1 as closely as possible, so fine tuning the segmentation part on only dataset 1 might be the way to go. </p>",
      "rawMarkdown": "Funnily, training on dataset 1 and 2 lowers the LB score compared to training on dataset 1 only. 0.364 -> 0.352. \n\nThis does make sense, especially when considering we are trying to get the annotation quality to match dataset 1 as closely as possible, so fine tuning the segmentation part on only dataset 1 might be the way to go.",
      "votes": null
    },
    {
      "id": "2307532",
      "postDate": "06/18/2023 08:56:03",
      "content": "<p>Training the prompt encoder as well seems to make basically no difference…</p>",
      "rawMarkdown": "Training the prompt encoder as well seems to make basically no difference...",
      "votes": null
    },
    {
      "id": "2307741",
      "postDate": "06/18/2023 11:44:14",
      "content": "<p>Going from the ViT-b to the ViT-l size of SAM only nets a 0.001 increase (0.364 -&gt; 0.365), which might just be noise. </p>",
      "rawMarkdown": "Going from the ViT-b to the ViT-l size of SAM only nets a 0.001 increase (0.364 -> 0.365), which might just be noise.",
      "votes": null
    },
    {
      "id": "2308670",
      "postDate": "06/19/2023 06:11:05",
      "content": "<p>Fine tuning the encoder as well for the ViT-b size gets me to 0.365 -&gt; 0.372 on the LB. Takes a bit of work to get it in the notebook GPU memory, so batch size=1 and 16-bit training. </p>\n<p>In related news, has someone found a reliable CV metric? <br>\nI'm using the <a href=\"https://torchmetrics.readthedocs.io/en/stable/detection/mean_average_precision.html\" target=\"_blank\">torchmetrics implementation</a> and get a mAP at IOU threshold 0.6 of 0.512, which is pretty far off from my LB score. </p>",
      "rawMarkdown": "Fine tuning the encoder as well for the ViT-b size gets me to 0.365 -> 0.372 on the LB. Takes a bit of work to get it in the notebook GPU memory, so batch size=1 and 16-bit training. \n\nIn related news, has someone found a reliable CV metric? \nI'm using the [torchmetrics implementation](https://torchmetrics.readthedocs.io/en/stable/detection/mean_average_precision.html) and get a mAP at IOU threshold 0.6 of 0.512, which is pretty far off from my LB score.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2305674,
      "author_name": "fnands",
      "author_url": "",
      "post_date": "06/16/2023 20:46:58",
      "content": "<p>Adding dilation 0.148 -&gt; 0.195, without any other changes. <br>\nGot the tip from: <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/416901\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/416901</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2307105,
      "author_name": "fnands",
      "author_url": "",
      "post_date": "06/17/2023 20:14:41",
      "content": "<p>Fine-tuning the mask decoder on dataset 1 takes me from 0.195 -&gt; 0.364. <br>\nWas a bit of a pain in the ass to get training, but clipping the gradients helped a lot. <br>\nHaven't implemented CV, which is bad practice. Will do so now to get an idea of the difference between CV and LB. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2307435,
      "author_name": "fnands",
      "author_url": "",
      "post_date": "06/18/2023 07:33:20",
      "content": "<p>Funnily, training on dataset 1 and 2 lowers the LB score compared to training on dataset 1 only. 0.364 -&gt; 0.352. </p>\n<p>This does make sense, especially when considering we are trying to get the annotation quality to match dataset 1 as closely as possible, so fine tuning the segmentation part on only dataset 1 might be the way to go. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2307532,
      "author_name": "fnands",
      "author_url": "",
      "post_date": "06/18/2023 08:56:03",
      "content": "<p>Training the prompt encoder as well seems to make basically no difference…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2307741,
      "author_name": "fnands",
      "author_url": "",
      "post_date": "06/18/2023 11:44:14",
      "content": "<p>Going from the ViT-b to the ViT-l size of SAM only nets a 0.001 increase (0.364 -&gt; 0.365), which might just be noise. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2308670,
      "author_name": "fnands",
      "author_url": "",
      "post_date": "06/19/2023 06:11:05",
      "content": "<p>Fine tuning the encoder as well for the ViT-b size gets me to 0.365 -&gt; 0.372 on the LB. Takes a bit of work to get it in the notebook GPU memory, so batch size=1 and 16-bit training. </p>\n<p>In related news, has someone found a reliable CV metric? <br>\nI'm using the <a href=\"https://torchmetrics.readthedocs.io/en/stable/detection/mean_average_precision.html\" target=\"_blank\">torchmetrics implementation</a> and get a mAP at IOU threshold 0.6 of 0.512, which is pretty far off from my LB score. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2296060": "I'm a big fan of @hengck23 's [public](https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332941) [solution](https://www.kaggle.com/competitions/champs-scalar-coupling/discussion/93972) [threads](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333), and always find them to be full of useful information, so in that spirit I will try a solution and keep track of it publicly, at least for a while.     \n\nSince Meta released their [SegmentAnything Model](https://github.com/facebookresearch/segment-anything/tree/main) (SAM), I've been looking for an excuse to play around with it, so that's what I'll do here.   \n\nSo the interesting thing about SAM, is that it's been designed to be prompted, i.e. given an image and a point/box/mask, it will return a mask. In the [SAM paper ](https://ai.facebook.com/research/publications/segment-anything/), they point out this can be done by hand or programmatically. \n\nNaturally the latter is more relevant here, so how about we train a lightweight object detection model (e.g. YOLO) to find objects, and then use SAM to refine the masks? The weakness here might be that we rely quite heavily on the object detector to do a good job first. Is it even necessary? As SAM can do [automatic mask generation](https://github.com/facebookresearch/segment-anything/blob/main/segment_anything/automatic_mask_generator.py). \n\nMy first attempt [is here](https://www.kaggle.com/code/fnands/yolov7-sam-inference-only), and naively takes a mask and box from YOLOv7 as prompts, and then tries do predict a segmentation masks. \nSo far, it does worse than just using using the YOLOv7 masks directly (LB 0.25). \n\nNext, some fine-tuning of SAM?",
    "2305674": "Adding dilation 0.148 -> 0.195, without any other changes. \nGot the tip from: https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/416901",
    "2307105": "Fine-tuning the mask decoder on dataset 1 takes me from 0.195 -> 0.364. \nWas a bit of a pain in the ass to get training, but clipping the gradients helped a lot. \nHaven't implemented CV, which is bad practice. Will do so now to get an idea of the difference between CV and LB.",
    "2307435": "Funnily, training on dataset 1 and 2 lowers the LB score compared to training on dataset 1 only. 0.364 -> 0.352. \n\nThis does make sense, especially when considering we are trying to get the annotation quality to match dataset 1 as closely as possible, so fine tuning the segmentation part on only dataset 1 might be the way to go.",
    "2307532": "Training the prompt encoder as well seems to make basically no difference...",
    "2307741": "Going from the ViT-b to the ViT-l size of SAM only nets a 0.001 increase (0.364 -> 0.365), which might just be noise.",
    "2308670": "Fine tuning the encoder as well for the ViT-b size gets me to 0.365 -> 0.372 on the LB. Takes a bit of work to get it in the notebook GPU memory, so batch size=1 and 16-bit training. \n\nIn related news, has someone found a reliable CV metric? \nI'm using the [torchmetrics implementation](https://torchmetrics.readthedocs.io/en/stable/detection/mean_average_precision.html) and get a mAP at IOU threshold 0.6 of 0.512, which is pretty far off from my LB score."
  },
  "source": "meta"
}