{
  "id": 420764,
  "title": "onmipose demo: unet-based instance segmentation",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/discussion/420764",
  "author_name": "hengck23",
  "post_date": "2023-07-02T12:23:30.994000",
  "votes": 10,
  "comment_count": 14,
  "views": 0,
  "content": "<p>this method is quite promising. hence i am starting a new thread on it<br>\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F82351c00fc0cd7979ced374ae5190108%2FSelection_999(2454).png?generation=1688300870548277&amp;alt=media\" alt=\"\"></p>\n<p>you can check a notebbook code on it<br>\npart1:<br>\n<a href=\"https://www.kaggle.com/code/hengck23/unet-instance-segmentation-onmipose-part1\" target=\"_blank\">https://www.kaggle.com/code/hengck23/unet-instance-segmentation-onmipose-part1</a><br>\npart2:<br>\n<a href=\"https://www.kaggle.com/code/hengck23/unet-instance-segmentation-onmipose-part2\" target=\"_blank\">https://www.kaggle.com/code/hengck23/unet-instance-segmentation-onmipose-part2</a><br>\npart3:<br>\nto be updated (learn a neutral network to do the post processing)</p>",
  "messages": [
    {
      "id": 2326806,
      "postDate": "2023-07-02T12:23:30.993Z",
      "content": "<p>this method is quite promising. hence i am starting a new thread on it<br>\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F82351c00fc0cd7979ced374ae5190108%2FSelection_999(2454).png?generation=1688300870548277&amp;alt=media\" alt=\"\"></p>\n<p>you can check a notebbook code on it<br>\npart1:<br>\n<a href=\"https://www.kaggle.com/code/hengck23/unet-instance-segmentation-onmipose-part1\" target=\"_blank\">https://www.kaggle.com/code/hengck23/unet-instance-segmentation-onmipose-part1</a><br>\npart2:<br>\n<a href=\"https://www.kaggle.com/code/hengck23/unet-instance-segmentation-onmipose-part2\" target=\"_blank\">https://www.kaggle.com/code/hengck23/unet-instance-segmentation-onmipose-part2</a><br>\npart3:<br>\nto be updated (learn a neutral network to do the post processing)</p>",
      "rawMarkdown": "this method is quite promising. hence i am starting a new thread on it\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F82351c00fc0cd7979ced374ae5190108%2FSelection_999(2454).png?generation=1688300870548277&alt=media)\n\nyou can check a notebbook code on it\npart1:\nhttps://www.kaggle.com/code/hengck23/unet-instance-segmentation-onmipose-part1\npart2:\nhttps://www.kaggle.com/code/hengck23/unet-instance-segmentation-onmipose-part2\npart3:\nto be updated (learn a neutral network to do the post processing)",
      "votes": 9
    },
    {
      "id": 2327377,
      "postDate": "2023-07-02T23:08:05.740Z",
      "content": "<p>not sure about yolo (but i think so be similar), but here in unet, the predicted mask is mostly smaller than the truth mask.<br>\nthis explains why dilation helps?</p>\n<p>this is results of inconsistent labelling, i think</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe4dbc166bcbd78eb84604137addf0a12%2FSelection_999(2486).png?generation=1688339250820661&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "not sure about yolo (but i think so be similar), but here in unet, the predicted mask is mostly smaller than the truth mask.\nthis explains why dilation helps?\n\nthis is results of inconsistent labelling, i think\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe4dbc166bcbd78eb84604137addf0a12%2FSelection_999(2486).png?generation=1688339250820661&alt=media)",
      "votes": 1,
      "replies": [
        {
          "id": 2327741,
          "postDate": "2023-07-03T06:24:04.630Z",
          "content": "<p>Interesting, then this means we are supposed to optimize performance of bbox prediction? i.e. locate where the instance is then use a better segmentation algorithm.</p>",
          "rawMarkdown": "Interesting, then this means we are supposed to optimize performance of bbox prediction? i.e. locate where the instance is then use a better segmentation algorithm."
        }
      ]
    },
    {
      "id": 2327325,
      "postDate": "2023-07-02T21:03:04.107Z",
      "content": "<p>detect intstances is not the problem on postprocessing, but because of the metric used in this competition define proper confidence is a challenge, I was able reach only 0.325 with unet on LB</p>",
      "rawMarkdown": "detect intstances is not the problem on postprocessing, but because of the metric used in this competition define proper confidence is a challenge, I was able reach only 0.325 with unet on LB",
      "votes": 1,
      "replies": [
        {
          "id": 2327366,
          "postDate": "2023-07-02T22:45:51.627Z",
          "content": "<p>quite a lot of object masks are touching each other, so you definitely need to seprate them as instances.</p>\n<p>you also need a head in unet to predict iou or some way to rank or output confidence score.<br>\nfor example.</p>\n<pre><code>let ground truthbe\n\ntruth =  a WxH image with =0 means pixel is background, =1 means pixel is  object 1, =2 \npredict = . same format .\n\nU,C = np.unique (truth, =)\n\nduring training, you need   some dynamic matching\niou_truth = zeros()\n i  range(1, predict.max()+1):\n    u,c = np.unique (  truth[==i], =)\n   .choose max count k.\n   iou = count k/ count K\n   iou_truth[==i] = iou\n\n\n\niou_predict = conv_layer (feature)\nloss = F.binary_cross_entropy_with_logits(iou_predict,iou_truth)\n\n\n\n</code></pre>",
          "rawMarkdown": "quite a lot of object masks are touching each other, so you definitely need to seprate them as instances.\n\nyou also need a head in unet to predict iou or some way to rank or output confidence score.\nfor example.\n```\nlet ground truth instance be\n\ntruth = .... a WxH image with value=0 means pixel is background, value=1 means pixel is from object 1, value=2 ....\npredict = ... same format ...\n\nU,C = np.unique (truth, return_counts=True)\n\nduring training, you need to do some dynamic matching\niou_truth = zeros()\nfor i in range(1, predict.max()+1):\n    u,c = np.unique (  truth[redict==i], return_counts=True)\n   ...choose max count k...\n   iou = count k/ count K\n   iou_truth[predict==i] = iou\n\n    \n#one example is to predict iou per pixel\niou_predict = conv_layer (feature)\nloss = F.binary_cross_entropy_with_logits(iou_predict,iou_truth)\n\n\n#a better solution is to predict iou per instance object, e.g. pool over some region\n\n```\n",
          "replies": [
            {
              "id": 2327371,
              "postDate": "2023-07-02T22:50:25.220Z",
              "content": "<p>a another solution is input the results of unet to some single instance segmentation network like mask-rcnn or train a SAM-like mask decoder head with e.g. instance distance transform as prompt</p>",
              "rawMarkdown": "a another solution is input the results of unet to some single instance segmentation network like mask-rcnn or train a SAM-like mask decoder head with e.g. instance distance transform as prompt"
            },
            {
              "id": 2327998,
              "postDate": "2023-07-03T09:39:04.913Z",
              "content": "<p>as an example for iou predictor</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fdc8a55261891d8056d72f30368a8e4ac%2FSelection_999(2504).png?generation=1688377027269280&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://arxiv.org/pdf/2107.06278.pdf\" target=\"_blank\">https://arxiv.org/pdf/2107.06278.pdf</a><br>\n\"Per-Pixel Classification is Not All You Need for Semantic Segmentation\" - Bowen Cheng</p>",
              "rawMarkdown": "as an example for iou predictor\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fdc8a55261891d8056d72f30368a8e4ac%2FSelection_999(2504).png?generation=1688377027269280&alt=media)\n\nhttps://arxiv.org/pdf/2107.06278.pdf\n\"Per-Pixel Classification is Not All You Need for Semantic Segmentation\" - Bowen Cheng"
            }
          ]
        }
      ]
    },
    {
      "id": 2326916,
      "postDate": "2023-07-02T14:05:20.703Z",
      "content": "<p><img src=\"https://i.ibb.co/QpGsRN2/Selection-999-2452.png\" alt=\"https://i.ibb.co/QpGsRN2/Selection-999-2452.png\"></p>\n<p>Background Reading<br>\nFor information on onmipose, check these</p>\n<p>code: <a href=\"https://github.com/kevinjohncutler/omnipose\" target=\"_blank\">https://github.com/kevinjohncutler/omnipose</a><br>\ndoc: <a href=\"https://omnipose.readthedocs.io/\" target=\"_blank\">https://omnipose.readthedocs.io/</a><br>\npaper: <a href=\"https://www.nature.com/articles/s41592-022-01639-4\" target=\"_blank\">https://www.nature.com/articles/s41592-022-01639-4</a><br>\n\"Omnipose: a high-precision morphology-independent solution for bacterial cell segmentation\" - Kevin J. Cutler</p>",
      "rawMarkdown": "![https://i.ibb.co/QpGsRN2/Selection-999-2452.png](https://i.ibb.co/QpGsRN2/Selection-999-2452.png)\n\nBackground Reading\nFor information on onmipose, check these\n\ncode: https://github.com/kevinjohncutler/omnipose\ndoc: https://omnipose.readthedocs.io/\npaper: https://www.nature.com/articles/s41592-022-01639-4\n\"Omnipose: a high-precision morphology-independent solution for bacterial cell segmentation\" - Kevin J. Cutler",
      "votes": 1,
      "replies": [
        {
          "id": 2327235,
          "postDate": "2023-07-02T18:43:49.680Z",
          "content": "<p>early results</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5579a7bead0ea9a3855eaf483ad8c292%2FSelection_999(2485).png?generation=1688323533434447&amp;alt=media\" alt=\"\"></p>\n<p>it can detect large, small, cluttered, elongated, complex/non-straight objects.<br>\nBut the seresnext101 encoder seems to be not strong enough (not enough context).</p>\n<p>i will try a strong VIT encoder tmr.</p>",
          "rawMarkdown": "early results\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5579a7bead0ea9a3855eaf483ad8c292%2FSelection_999(2485).png?generation=1688323533434447&alt=media)\n\nit can detect large, small, cluttered, elongated, complex/non-straight objects.\nBut the seresnext101 encoder seems to be not strong enough (not enough context).\n\ni will try a strong VIT encoder tmr."
        }
      ]
    },
    {
      "id": 2335893,
      "postDate": "2023-07-08T23:34:17.677Z",
      "content": "<p>i used previous kaggle datasciencebowl 2018 winner solution to seprarte touching instance in segmentation + public note unet-plus-plus</p>\n<p><a href=\"https://www.kaggle.com/code/hidngnguyna/baseline-unet-semantic-as-instance-segmentation/notebook\" target=\"_blank\">https://www.kaggle.com/code/hidngnguyna/baseline-unet-semantic-as-instance-segmentation/notebook</a><br>\n<a href=\"https://www.kaggle.com/competitions/data-science-bowl-2018/discussion/54741\" target=\"_blank\">https://www.kaggle.com/competitions/data-science-bowl-2018/discussion/54741</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7d3867ab74e4357323f11f0b21aca0cb%2FSelection_999(2594).png?generation=1688859241725503&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2b0ed74ffb465fa720ca4d69d32c2c37%2FSelection_999(2593).png?generation=1688859255439643&amp;alt=media\" alt=\"\"></p>\n<p>according to datasciencebowl 2018 winner solution, you can use watershed to postprocess (to convert <br>\nsemantic to instance mask)</p>\n<hr>\n<p>to compute score, you can use</p>\n<pre><code>conf = (semantic_prob*instance_mask).() / instance_mask.() \n\n  \n\nconf = - (unsure_prob*instance_mask).() / instance_mask.() \n\n\nself.semantic = sigmoid of  : background, vessel\nself.part = ... softmax of  : background, seed, border\nself.unsure = ... softmax of  : background, unsure, sure\n</code></pre>",
      "rawMarkdown": "i used previous kaggle datasciencebowl 2018 winner solution to seprarte touching instance in segmentation + public note unet-plus-plus\n\nhttps://www.kaggle.com/code/hidngnguyna/baseline-unet-semantic-as-instance-segmentation/notebook\nhttps://www.kaggle.com/competitions/data-science-bowl-2018/discussion/54741\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7d3867ab74e4357323f11f0b21aca0cb%2FSelection_999(2594).png?generation=1688859241725503&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2b0ed74ffb465fa720ca4d69d32c2c37%2FSelection_999(2593).png?generation=1688859255439643&alt=media)\n\naccording to datasciencebowl 2018 winner solution, you can use watershed to postprocess (to convert \nsemantic to instance mask)\n\n\n---\nto compute score, you can use\n\n```\nconf = (semantic_prob*instance_mask).sum() / instance_mask.sum() \n\nor  \n\nconf = 1- (unsure_prob*instance_mask).sum() / instance_mask.sum() \n\n#assume you have the following segmentation head\nself.semantic = sigmoid of 1 class: background, vessel\nself.part = ... softmax of 3 class: background, seed, border\nself.unsure = ... softmax of 3 class: background, unsure, sure\n\n```"
    },
    {
      "id": 2328817,
      "postDate": "2023-07-03T23:16:07.933Z",
      "content": "<p>i suddenly recalled an old related kaggle competition:</p>\n<p><a href=\"https://github.com/ternaus/TernausNetV2\" target=\"_blank\">https://github.com/ternaus/TernausNetV2</a></p>\n<p>after they separate the instances from semantic mask, they crop the features and predict iou:</p>\n<pre><code>crop = roi(feature_map) \niou = iou_head(crop_instance, crop_feature_map)\n\nthere are sets of (roi, iou_head) for </code></pre>\n<p>`</p>",
      "rawMarkdown": "i suddenly recalled an old related kaggle competition:\n\nhttps://github.com/ternaus/TernausNetV2\n\nafter they separate the instances from semantic mask, they crop the features and predict iou:\n\n```\ncrop = roi(instance, feature_map) #modify from torchvision rpn roi align\niou = iou_head(crop_instance, crop_feature_map)\n\nthere are sets of (roi, iou_head) for differently-sized instances\n\n````"
    },
    {
      "id": 2328804,
      "postDate": "2023-07-03T22:57:10.533Z",
      "content": "<p>results for convnext base encoder<br>\n(i choose convext becuase it can do self-supervised conv-MAE learning, see convnext-v2 paper)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0bf589cc649dc32c67da0865c4f0040b%2FSelection_999(2535).png?generation=1688424663534110&amp;alt=media\" alt=\"\"></p>\n<p>i am surprise that the model is not able to floodfill waht seems to be \"simple region\" (constant  intensity)</p>\n<p>it seems that the issue is not context?<br>\nyou can see from above that i almost detect all objects (even very small ones) at the validation, but detected objects are much smaller in size.</p>\n<p>some interesting fact about using convnext encoder:</p>\n<ul>\n<li>you can actually freeze it entirely and just train the decoder (blazing fast). then finetune end-to-end at last the end with small learning rate</li>\n<li>typically i need 400 to 500 epoches (even for seresnext101)</li>\n<li>results are pretty good and would even be better  if i can solve \"unknown missing floodfill issue\"</li>\n</ul>",
      "rawMarkdown": "results for convnext base encoder\n(i choose convext becuase it can do self-supervised conv-MAE learning, see convnext-v2 paper)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0bf589cc649dc32c67da0865c4f0040b%2FSelection_999(2535).png?generation=1688424663534110&alt=media)\n\ni am surprise that the model is not able to floodfill waht seems to be \"simple region\" (constant  intensity)\n\nit seems that the issue is not context?\nyou can see from above that i almost detect all objects (even very small ones) at the validation, but detected objects are much smaller in size.\n\nsome interesting fact about using convnext encoder:\n- you can actually freeze it entirely and just train the decoder (blazing fast). then finetune end-to-end at last the end with small learning rate\n- typically i need 400 to 500 epoches (even for seresnext101)\n- results are pretty good and would even be better  if i can solve \"unknown missing floodfill issue\""
    },
    {
      "id": 2328094,
      "postDate": "2023-07-03T10:58:30.893Z",
      "content": "<p>instead of heuristics for post processing, you can learn a mask decoding head to convert flow to instance mask.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7940430ef9d10e6060ee048c4b32269b%2FSelection_999(2510).png?generation=1688381905786325&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "instead of heuristics for post processing, you can learn a mask decoding head to convert flow to instance mask.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7940430ef9d10e6060ee048c4b32269b%2FSelection_999(2510).png?generation=1688381905786325&alt=media)",
      "replies": [
        {
          "id": 2328113,
          "postDate": "2023-07-03T11:05:53.520Z",
          "content": "<p>i just realise \"dynamic conv kernel\" of box free instance segmentation head (e.g. solov2, YOLACT) is related to transformer and attnetion.</p>\n<p>such segmentation head predict the \"weight of convolution kernel\" from test image (not memorised from train image) and use it extract feature for instance mask prediction.</p>\n<p>isn't it:</p>\n<p>query = dynamic weight<br>\nkey = image feature<br>\ndot-product attnetion = convolution + gate</p>",
          "rawMarkdown": "i just realise \"dynamic conv kernel\" of box free instance segmentation head (e.g. solov2, YOLACT) is related to transformer and attnetion.\n\nsuch segmentation head predict the \"weight of convolution kernel\" from test image (not memorised from train image) and use it extract feature for instance mask prediction.\n\nisn't it:\n\nquery = dynamic weight\nkey = image feature\ndot-product attnetion = convolution + gate\n\n \n"
        }
      ]
    },
    {
      "id": 2357686,
      "postDate": "2023-07-25T04:06:31.760Z",
      "content": "<p>Great work！</p>",
      "rawMarkdown": "Great work！",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2327377,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-07-02T23:08:05.740000",
      "content": "<p>not sure about yolo (but i think so be similar), but here in unet, the predicted mask is mostly smaller than the truth mask.<br>\nthis explains why dilation helps?</p>\n<p>this is results of inconsistent labelling, i think</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe4dbc166bcbd78eb84604137addf0a12%2FSelection_999(2486).png?generation=1688339250820661&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2327741,
          "author_name": "Feng Qilong",
          "author_url": "",
          "post_date": "2023-07-03T06:24:04.630000",
          "content": "<p>Interesting, then this means we are supposed to optimize performance of bbox prediction? i.e. locate where the instance is then use a better segmentation algorithm.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2327325,
      "author_name": "Kostiantyn Maksymov",
      "author_url": "",
      "post_date": "2023-07-02T21:03:04.107000",
      "content": "<p>detect intstances is not the problem on postprocessing, but because of the metric used in this competition define proper confidence is a challenge, I was able reach only 0.325 with unet on LB</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2327366,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-07-02T22:45:51.627000",
          "content": "<p>quite a lot of object masks are touching each other, so you definitely need to seprate them as instances.</p>\n<p>you also need a head in unet to predict iou or some way to rank or output confidence score.<br>\nfor example.</p>\n<pre><code>let ground truthbe\n\ntruth =  a WxH image with =0 means pixel is background, =1 means pixel is  object 1, =2 \npredict = . same format .\n\nU,C = np.unique (truth, =)\n\nduring training, you need   some dynamic matching\niou_truth = zeros()\n i  range(1, predict.max()+1):\n    u,c = np.unique (  truth[==i], =)\n   .choose max count k.\n   iou = count k/ count K\n   iou_truth[==i] = iou\n\n\n\niou_predict = conv_layer (feature)\nloss = F.binary_cross_entropy_with_logits(iou_predict,iou_truth)\n\n\n\n</code></pre>",
          "votes": 0,
          "replies": [
            {
              "id": 2327371,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-07-02T22:50:25.220000",
              "content": "<p>a another solution is input the results of unet to some single instance segmentation network like mask-rcnn or train a SAM-like mask decoder head with e.g. instance distance transform as prompt</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2327998,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-07-03T09:39:04.913000",
              "content": "<p>as an example for iou predictor</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fdc8a55261891d8056d72f30368a8e4ac%2FSelection_999(2504).png?generation=1688377027269280&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://arxiv.org/pdf/2107.06278.pdf\" target=\"_blank\">https://arxiv.org/pdf/2107.06278.pdf</a><br>\n\"Per-Pixel Classification is Not All You Need for Semantic Segmentation\" - Bowen Cheng</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2326916,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-07-02T14:05:20.703000",
      "content": "<p><img src=\"https://i.ibb.co/QpGsRN2/Selection-999-2452.png\" alt=\"https://i.ibb.co/QpGsRN2/Selection-999-2452.png\"></p>\n<p>Background Reading<br>\nFor information on onmipose, check these</p>\n<p>code: <a href=\"https://github.com/kevinjohncutler/omnipose\" target=\"_blank\">https://github.com/kevinjohncutler/omnipose</a><br>\ndoc: <a href=\"https://omnipose.readthedocs.io/\" target=\"_blank\">https://omnipose.readthedocs.io/</a><br>\npaper: <a href=\"https://www.nature.com/articles/s41592-022-01639-4\" target=\"_blank\">https://www.nature.com/articles/s41592-022-01639-4</a><br>\n\"Omnipose: a high-precision morphology-independent solution for bacterial cell segmentation\" - Kevin J. Cutler</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2327235,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-07-02T18:43:49.680000",
          "content": "<p>early results</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5579a7bead0ea9a3855eaf483ad8c292%2FSelection_999(2485).png?generation=1688323533434447&amp;alt=media\" alt=\"\"></p>\n<p>it can detect large, small, cluttered, elongated, complex/non-straight objects.<br>\nBut the seresnext101 encoder seems to be not strong enough (not enough context).</p>\n<p>i will try a strong VIT encoder tmr.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2335893,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-07-08T23:34:17.677000",
      "content": "<p>i used previous kaggle datasciencebowl 2018 winner solution to seprarte touching instance in segmentation + public note unet-plus-plus</p>\n<p><a href=\"https://www.kaggle.com/code/hidngnguyna/baseline-unet-semantic-as-instance-segmentation/notebook\" target=\"_blank\">https://www.kaggle.com/code/hidngnguyna/baseline-unet-semantic-as-instance-segmentation/notebook</a><br>\n<a href=\"https://www.kaggle.com/competitions/data-science-bowl-2018/discussion/54741\" target=\"_blank\">https://www.kaggle.com/competitions/data-science-bowl-2018/discussion/54741</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7d3867ab74e4357323f11f0b21aca0cb%2FSelection_999(2594).png?generation=1688859241725503&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2b0ed74ffb465fa720ca4d69d32c2c37%2FSelection_999(2593).png?generation=1688859255439643&amp;alt=media\" alt=\"\"></p>\n<p>according to datasciencebowl 2018 winner solution, you can use watershed to postprocess (to convert <br>\nsemantic to instance mask)</p>\n<hr>\n<p>to compute score, you can use</p>\n<pre><code>conf = (semantic_prob*instance_mask).() / instance_mask.() \n\n  \n\nconf = - (unsure_prob*instance_mask).() / instance_mask.() \n\n\nself.semantic = sigmoid of  : background, vessel\nself.part = ... softmax of  : background, seed, border\nself.unsure = ... softmax of  : background, unsure, sure\n</code></pre>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2328817,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-07-03T23:16:07.933000",
      "content": "<p>i suddenly recalled an old related kaggle competition:</p>\n<p><a href=\"https://github.com/ternaus/TernausNetV2\" target=\"_blank\">https://github.com/ternaus/TernausNetV2</a></p>\n<p>after they separate the instances from semantic mask, they crop the features and predict iou:</p>\n<pre><code>crop = roi(feature_map) \niou = iou_head(crop_instance, crop_feature_map)\n\nthere are sets of (roi, iou_head) for </code></pre>\n<p>`</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2328804,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-07-03T22:57:10.533000",
      "content": "<p>results for convnext base encoder<br>\n(i choose convext becuase it can do self-supervised conv-MAE learning, see convnext-v2 paper)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0bf589cc649dc32c67da0865c4f0040b%2FSelection_999(2535).png?generation=1688424663534110&amp;alt=media\" alt=\"\"></p>\n<p>i am surprise that the model is not able to floodfill waht seems to be \"simple region\" (constant  intensity)</p>\n<p>it seems that the issue is not context?<br>\nyou can see from above that i almost detect all objects (even very small ones) at the validation, but detected objects are much smaller in size.</p>\n<p>some interesting fact about using convnext encoder:</p>\n<ul>\n<li>you can actually freeze it entirely and just train the decoder (blazing fast). then finetune end-to-end at last the end with small learning rate</li>\n<li>typically i need 400 to 500 epoches (even for seresnext101)</li>\n<li>results are pretty good and would even be better  if i can solve \"unknown missing floodfill issue\"</li>\n</ul>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2328094,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-07-03T10:58:30.893000",
      "content": "<p>instead of heuristics for post processing, you can learn a mask decoding head to convert flow to instance mask.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7940430ef9d10e6060ee048c4b32269b%2FSelection_999(2510).png?generation=1688381905786325&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2328113,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-07-03T11:05:53.520000",
          "content": "<p>i just realise \"dynamic conv kernel\" of box free instance segmentation head (e.g. solov2, YOLACT) is related to transformer and attnetion.</p>\n<p>such segmentation head predict the \"weight of convolution kernel\" from test image (not memorised from train image) and use it extract feature for instance mask prediction.</p>\n<p>isn't it:</p>\n<p>query = dynamic weight<br>\nkey = image feature<br>\ndot-product attnetion = convolution + gate</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2357686,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-07-25T04:06:31.760000",
      "content": "<p>Great work！</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2326806": "this method is quite promising. hence i am starting a new thread on it\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F82351c00fc0cd7979ced374ae5190108%2FSelection_999(2454).png?generation=1688300870548277&alt=media)\n\nyou can check a notebbook code on it\npart1:\nhttps://www.kaggle.com/code/hengck23/unet-instance-segmentation-onmipose-part1\npart2:\nhttps://www.kaggle.com/code/hengck23/unet-instance-segmentation-onmipose-part2\npart3:\nto be updated (learn a neutral network to do the post processing)",
    "2327377": "not sure about yolo (but i think so be similar), but here in unet, the predicted mask is mostly smaller than the truth mask.\nthis explains why dilation helps?\n\nthis is results of inconsistent labelling, i think\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe4dbc166bcbd78eb84604137addf0a12%2FSelection_999(2486).png?generation=1688339250820661&alt=media)",
    "2327325": "detect intstances is not the problem on postprocessing, but because of the metric used in this competition define proper confidence is a challenge, I was able reach only 0.325 with unet on LB",
    "2326916": "![https://i.ibb.co/QpGsRN2/Selection-999-2452.png](https://i.ibb.co/QpGsRN2/Selection-999-2452.png)\n\nBackground Reading\nFor information on onmipose, check these\n\ncode: https://github.com/kevinjohncutler/omnipose\ndoc: https://omnipose.readthedocs.io/\npaper: https://www.nature.com/articles/s41592-022-01639-4\n\"Omnipose: a high-precision morphology-independent solution for bacterial cell segmentation\" - Kevin J. Cutler",
    "2335893": "i used previous kaggle datasciencebowl 2018 winner solution to seprarte touching instance in segmentation + public note unet-plus-plus\n\nhttps://www.kaggle.com/code/hidngnguyna/baseline-unet-semantic-as-instance-segmentation/notebook\nhttps://www.kaggle.com/competitions/data-science-bowl-2018/discussion/54741\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7d3867ab74e4357323f11f0b21aca0cb%2FSelection_999(2594).png?generation=1688859241725503&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2b0ed74ffb465fa720ca4d69d32c2c37%2FSelection_999(2593).png?generation=1688859255439643&alt=media)\n\naccording to datasciencebowl 2018 winner solution, you can use watershed to postprocess (to convert \nsemantic to instance mask)\n\n\n---\nto compute score, you can use\n\n```\nconf = (semantic_prob*instance_mask).sum() / instance_mask.sum() \n\nor  \n\nconf = 1- (unsure_prob*instance_mask).sum() / instance_mask.sum() \n\n#assume you have the following segmentation head\nself.semantic = sigmoid of 1 class: background, vessel\nself.part = ... softmax of 3 class: background, seed, border\nself.unsure = ... softmax of 3 class: background, unsure, sure\n\n```",
    "2328817": "i suddenly recalled an old related kaggle competition:\n\nhttps://github.com/ternaus/TernausNetV2\n\nafter they separate the instances from semantic mask, they crop the features and predict iou:\n\n```\ncrop = roi(instance, feature_map) #modify from torchvision rpn roi align\niou = iou_head(crop_instance, crop_feature_map)\n\nthere are sets of (roi, iou_head) for differently-sized instances\n\n````",
    "2328804": "results for convnext base encoder\n(i choose convext becuase it can do self-supervised conv-MAE learning, see convnext-v2 paper)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0bf589cc649dc32c67da0865c4f0040b%2FSelection_999(2535).png?generation=1688424663534110&alt=media)\n\ni am surprise that the model is not able to floodfill waht seems to be \"simple region\" (constant  intensity)\n\nit seems that the issue is not context?\nyou can see from above that i almost detect all objects (even very small ones) at the validation, but detected objects are much smaller in size.\n\nsome interesting fact about using convnext encoder:\n- you can actually freeze it entirely and just train the decoder (blazing fast). then finetune end-to-end at last the end with small learning rate\n- typically i need 400 to 500 epoches (even for seresnext101)\n- results are pretty good and would even be better  if i can solve \"unknown missing floodfill issue\"",
    "2328094": "instead of heuristics for post processing, you can learn a mask decoding head to convert flow to instance mask.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7940430ef9d10e6060ee048c4b32269b%2FSelection_999(2510).png?generation=1688381905786325&alt=media)",
    "2357686": "Great work！"
  }
}