{
  "id": 293874,
  "title": "What I tried and did not work?",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/293874",
  "author_name": "",
  "post_date": "2021-12-07T11:49:06.079021700Z",
  "votes": 31,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I joined this competition purely to learn about various approaches in instance segmentation (because I never trained a IS model before). While it seems like the top of the LB are based on Mask-RCNN, I tried some other approaches as well… just for fun 😄😄 (and to have an excuse why I failed to train high-performing models lol)</p>\n<p>Over the past weekends, I tried several approaches that does not require RPN with bounding box regression for the following reasons:</p>\n<ul>\n<li>in COCO, there's not that much samples with two masks of the same class having the same bounding-box. So, my intuition is Mask-RCNN will struggle in situations where a bunch of cells of the same semantic class are cramped into the same bounding-box.</li>\n<li>Sometimes, the masks are too irregular / too thin for Mask-RCNN to pick up.</li>\n</ul>\n<p>Here are what I tried during the past weekends:</p>\n<h2>MaskFormer</h2>\n<p><a href=\"https://arxiv.org/abs/2107.06278\" target=\"_blank\">[paper]</a> <a href=\"https://github.com/facebookresearch/MaskFormer\" target=\"_blank\">[github]</a></p>\n<p>I tried MaskFormer (not Mask2Former) for this competition, by re-formulating it as a panoptic segmentation problem. The model failed to converge. I did not have time to debug, but my hunch is that the instance-specific masks in MaskFormer are formed using a query (basically 1 query = 1 cell). If you increase the number of queries, the whole thing will consume an astronomical amount of memory. The original paper uses only &lt;100 queries, with the assumption that in COCO dataset there's only that much semantically different objects in the image. So, with 400+ instances of the same class, the dataset in this competition kind of breaks all of the assumptions that COCO-based methods have.</p>\n<h2>SOLO v2</h2>\n<p><a href=\"https://arxiv.org/abs/2003.10152\" target=\"_blank\">[paper]</a> <a href=\"https://github.com/aim-uofa/AdelaiDet\" target=\"_blank\">[github]</a></p>\n<p>I had pretty high hope for this approach. Instead of having an RPN and then segmenting proposed regions, this approach predicts the mask for specific locations. However, the resulting model struggles to reach 0.2 (while my Mask-RCNN models can easily reach 0.3+). After some EDA, I noticed that the masks are not bounded nicely, but spreads around (e.g. the predicted mask for position <code>(100, 100)</code> will contain some pixels at position <code>(400, 500)</code>, which is faaaar away. I guess more regularization is needed. The code of this approach contains a few bugs that will break if you have 500+ instances. I'll make a PR to that repo soon, but if you want to try this right now - feel free to pm me.</p>\n<h2>PointRend</h2>\n<p><a href=\"https://arxiv.org/abs/1912.08193\" target=\"_blank\">[paper]</a> <a href=\"https://github.com/facebookresearch/detectron2/tree/main/projects/PointRend\" target=\"_blank\">[github]</a></p>\n<p>This one is just an improvement on top of Mask-RCNN. However, I trained it from scratch instead of fine-tuning with a trained MRCNN. I guess that's why I failed to get higher scores than 0.25 with this one.</p>\n<hr>\n<p>I guess the conclusion here is that vanilla implementation of all those methods based on COCO will not work nicely in this comp. You'll need to adapt them for this specific dataset.</p>",
  "messages": [
    {
      "id": "1610655",
      "postDate": "12/07/2021 11:49:06",
      "content": "<p>I joined this competition purely to learn about various approaches in instance segmentation (because I never trained a IS model before). While it seems like the top of the LB are based on Mask-RCNN, I tried some other approaches as well… just for fun 😄😄 (and to have an excuse why I failed to train high-performing models lol)</p>\n<p>Over the past weekends, I tried several approaches that does not require RPN with bounding box regression for the following reasons:</p>\n<ul>\n<li>in COCO, there's not that much samples with two masks of the same class having the same bounding-box. So, my intuition is Mask-RCNN will struggle in situations where a bunch of cells of the same semantic class are cramped into the same bounding-box.</li>\n<li>Sometimes, the masks are too irregular / too thin for Mask-RCNN to pick up.</li>\n</ul>\n<p>Here are what I tried during the past weekends:</p>\n<h2>MaskFormer</h2>\n<p><a href=\"https://arxiv.org/abs/2107.06278\" target=\"_blank\">[paper]</a> <a href=\"https://github.com/facebookresearch/MaskFormer\" target=\"_blank\">[github]</a></p>\n<p>I tried MaskFormer (not Mask2Former) for this competition, by re-formulating it as a panoptic segmentation problem. The model failed to converge. I did not have time to debug, but my hunch is that the instance-specific masks in MaskFormer are formed using a query (basically 1 query = 1 cell). If you increase the number of queries, the whole thing will consume an astronomical amount of memory. The original paper uses only &lt;100 queries, with the assumption that in COCO dataset there's only that much semantically different objects in the image. So, with 400+ instances of the same class, the dataset in this competition kind of breaks all of the assumptions that COCO-based methods have.</p>\n<h2>SOLO v2</h2>\n<p><a href=\"https://arxiv.org/abs/2003.10152\" target=\"_blank\">[paper]</a> <a href=\"https://github.com/aim-uofa/AdelaiDet\" target=\"_blank\">[github]</a></p>\n<p>I had pretty high hope for this approach. Instead of having an RPN and then segmenting proposed regions, this approach predicts the mask for specific locations. However, the resulting model struggles to reach 0.2 (while my Mask-RCNN models can easily reach 0.3+). After some EDA, I noticed that the masks are not bounded nicely, but spreads around (e.g. the predicted mask for position <code>(100, 100)</code> will contain some pixels at position <code>(400, 500)</code>, which is faaaar away. I guess more regularization is needed. The code of this approach contains a few bugs that will break if you have 500+ instances. I'll make a PR to that repo soon, but if you want to try this right now - feel free to pm me.</p>\n<h2>PointRend</h2>\n<p><a href=\"https://arxiv.org/abs/1912.08193\" target=\"_blank\">[paper]</a> <a href=\"https://github.com/facebookresearch/detectron2/tree/main/projects/PointRend\" target=\"_blank\">[github]</a></p>\n<p>This one is just an improvement on top of Mask-RCNN. However, I trained it from scratch instead of fine-tuning with a trained MRCNN. I guess that's why I failed to get higher scores than 0.25 with this one.</p>\n<hr>\n<p>I guess the conclusion here is that vanilla implementation of all those methods based on COCO will not work nicely in this comp. You'll need to adapt them for this specific dataset.</p>",
      "rawMarkdown": "I joined this competition purely to learn about various approaches in instance segmentation (because I never trained a IS model before). While it seems like the top of the LB are based on Mask-RCNN, I tried some other approaches as well... just for fun 😄😄 (and to have an excuse why I failed to train high-performing models lol)\n\nOver the past weekends, I tried several approaches that does not require RPN with bounding box regression for the following reasons:\n- in COCO, there's not that much samples with two masks of the same class having the same bounding-box. So, my intuition is Mask-RCNN will struggle in situations where a bunch of cells of the same semantic class are cramped into the same bounding-box.\n- Sometimes, the masks are too irregular / too thin for Mask-RCNN to pick up.\n\nHere are what I tried during the past weekends:\n\n## MaskFormer\n[[paper]](https://arxiv.org/abs/2107.06278) [[github]](https://github.com/facebookresearch/MaskFormer)\n\nI tried MaskFormer (not Mask2Former) for this competition, by re-formulating it as a panoptic segmentation problem. The model failed to converge. I did not have time to debug, but my hunch is that the instance-specific masks in MaskFormer are formed using a query (basically 1 query = 1 cell). If you increase the number of queries, the whole thing will consume an astronomical amount of memory. The original paper uses only <100 queries, with the assumption that in COCO dataset there's only that much semantically different objects in the image. So, with 400+ instances of the same class, the dataset in this competition kind of breaks all of the assumptions that COCO-based methods have.\n\n## SOLO v2\n[[paper]](https://arxiv.org/abs/2003.10152) [[github]](https://github.com/aim-uofa/AdelaiDet)\n\nI had pretty high hope for this approach. Instead of having an RPN and then segmenting proposed regions, this approach predicts the mask for specific locations. However, the resulting model struggles to reach 0.2 (while my Mask-RCNN models can easily reach 0.3+). After some EDA, I noticed that the masks are not bounded nicely, but spreads around (e.g. the predicted mask for position `(100, 100)` will contain some pixels at position `(400, 500)`, which is faaaar away. I guess more regularization is needed. The code of this approach contains a few bugs that will break if you have 500+ instances. I'll make a PR to that repo soon, but if you want to try this right now - feel free to pm me.\n\n## PointRend\n[[paper]](https://arxiv.org/abs/1912.08193) [[github]](https://github.com/facebookresearch/detectron2/tree/main/projects/PointRend)\n\nThis one is just an improvement on top of Mask-RCNN. However, I trained it from scratch instead of fine-tuning with a trained MRCNN. I guess that's why I failed to get higher scores than 0.25 with this one.\n\n------------------------------------------------------------\n\nI guess the conclusion here is that vanilla implementation of all those methods based on COCO will not work nicely in this comp. You'll need to adapt them for this specific dataset.",
      "votes": null
    },
    {
      "id": "1610707",
      "postDate": "12/07/2021 13:00:07",
      "content": "<p>I use IoU loss as bboxes loss, the performance dropped significantly. (maskrcnn)</p>",
      "rawMarkdown": "I use IoU loss as bboxes loss, the performance dropped significantly. (maskrcnn)",
      "votes": null
    },
    {
      "id": "1610790",
      "postDate": "12/07/2021 14:24:41",
      "content": "<p>I've had similar experience with Solov2. Reached around .2 after I made it work with 500+ masks</p>",
      "rawMarkdown": "I've had similar experience with Solov2. Reached around .2 after I made it work with 500+ masks",
      "votes": null
    },
    {
      "id": "1610871",
      "postDate": "12/07/2021 15:15:23",
      "content": "<p>anyone tried  resnest200 from this repo <a href=\"https://github.com/sartorius-research/LIVECell/tree/main/model\" target=\"_blank\">https://github.com/sartorius-research/LIVECell/tree/main/model</a><br>\nI can't fit to my machine with batch size of 1 due to SynBN</p>",
      "rawMarkdown": "anyone tried  resnest200 from this repo https://github.com/sartorius-research/LIVECell/tree/main/model\nI can't fit to my machine with batch size of 1 due to SynBN",
      "votes": null
    },
    {
      "id": "1612133",
      "postDate": "12/08/2021 15:25:05",
      "content": "<p>I tried to train resnest200 but i am not able to make a successful submission with it yet. Try to reduce MIN_SIZE_TRAIN to smaller values. In my case i tried 520 and 560 and it works fine within the GPU</p>",
      "rawMarkdown": "I tried to train resnest200 but i am not able to make a successful submission with it yet. Try to reduce MIN_SIZE_TRAIN to smaller values. In my case i tried 520 and 560 and it works fine within the GPU",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1610707,
      "author_name": "blueboy97",
      "author_url": "",
      "post_date": "12/07/2021 13:00:07",
      "content": "<p>I use IoU loss as bboxes loss, the performance dropped significantly. (maskrcnn)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1610790,
      "author_name": "slawekbiel",
      "author_url": "",
      "post_date": "12/07/2021 14:24:41",
      "content": "<p>I've had similar experience with Solov2. Reached around .2 after I made it work with 500+ masks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1610871,
      "author_name": "ptran1203",
      "author_url": "",
      "post_date": "12/07/2021 15:15:23",
      "content": "<p>anyone tried  resnest200 from this repo <a href=\"https://github.com/sartorius-research/LIVECell/tree/main/model\" target=\"_blank\">https://github.com/sartorius-research/LIVECell/tree/main/model</a><br>\nI can't fit to my machine with batch size of 1 due to SynBN</p>",
      "votes": null,
      "replies": [
        {
          "id": 1612133,
          "author_name": "osamir",
          "author_url": "",
          "post_date": "12/08/2021 15:25:05",
          "content": "<p>I tried to train resnest200 but i am not able to make a successful submission with it yet. Try to reduce MIN_SIZE_TRAIN to smaller values. In my case i tried 520 and 560 and it works fine within the GPU</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1610655": "I joined this competition purely to learn about various approaches in instance segmentation (because I never trained a IS model before). While it seems like the top of the LB are based on Mask-RCNN, I tried some other approaches as well... just for fun 😄😄 (and to have an excuse why I failed to train high-performing models lol)\n\nOver the past weekends, I tried several approaches that does not require RPN with bounding box regression for the following reasons:\n- in COCO, there's not that much samples with two masks of the same class having the same bounding-box. So, my intuition is Mask-RCNN will struggle in situations where a bunch of cells of the same semantic class are cramped into the same bounding-box.\n- Sometimes, the masks are too irregular / too thin for Mask-RCNN to pick up.\n\nHere are what I tried during the past weekends:\n\n## MaskFormer\n[[paper]](https://arxiv.org/abs/2107.06278) [[github]](https://github.com/facebookresearch/MaskFormer)\n\nI tried MaskFormer (not Mask2Former) for this competition, by re-formulating it as a panoptic segmentation problem. The model failed to converge. I did not have time to debug, but my hunch is that the instance-specific masks in MaskFormer are formed using a query (basically 1 query = 1 cell). If you increase the number of queries, the whole thing will consume an astronomical amount of memory. The original paper uses only <100 queries, with the assumption that in COCO dataset there's only that much semantically different objects in the image. So, with 400+ instances of the same class, the dataset in this competition kind of breaks all of the assumptions that COCO-based methods have.\n\n## SOLO v2\n[[paper]](https://arxiv.org/abs/2003.10152) [[github]](https://github.com/aim-uofa/AdelaiDet)\n\nI had pretty high hope for this approach. Instead of having an RPN and then segmenting proposed regions, this approach predicts the mask for specific locations. However, the resulting model struggles to reach 0.2 (while my Mask-RCNN models can easily reach 0.3+). After some EDA, I noticed that the masks are not bounded nicely, but spreads around (e.g. the predicted mask for position `(100, 100)` will contain some pixels at position `(400, 500)`, which is faaaar away. I guess more regularization is needed. The code of this approach contains a few bugs that will break if you have 500+ instances. I'll make a PR to that repo soon, but if you want to try this right now - feel free to pm me.\n\n## PointRend\n[[paper]](https://arxiv.org/abs/1912.08193) [[github]](https://github.com/facebookresearch/detectron2/tree/main/projects/PointRend)\n\nThis one is just an improvement on top of Mask-RCNN. However, I trained it from scratch instead of fine-tuning with a trained MRCNN. I guess that's why I failed to get higher scores than 0.25 with this one.\n\n------------------------------------------------------------\n\nI guess the conclusion here is that vanilla implementation of all those methods based on COCO will not work nicely in this comp. You'll need to adapt them for this specific dataset.",
    "1610707": "I use IoU loss as bboxes loss, the performance dropped significantly. (maskrcnn)",
    "1610790": "I've had similar experience with Solov2. Reached around .2 after I made it work with 500+ masks",
    "1610871": "anyone tried  resnest200 from this repo https://github.com/sartorius-research/LIVECell/tree/main/model\nI can't fit to my machine with batch size of 1 due to SynBN",
    "1612133": "I tried to train resnest200 but i am not able to make a successful submission with it yet. Try to reduce MIN_SIZE_TRAIN to smaller values. In my case i tried 520 and 560 and it works fine within the GPU"
  },
  "source": "meta"
}