{
  "id": 296955,
  "title": "Some question on how to train mask-rcnn if any expert can help",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/296955",
  "author_name": "",
  "post_date": "2021-12-24T14:37:00.641106700Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi all. I am new to instance segmentation and the method of Mask RCNN.</p>\n<p>I wonder if anyone can briefly outline the steps needed to train Mask-RCNN, that will be very helpful.</p>\n<p>Lets take the region proposal network as an example. For every output feature map, the network has to learn whether an object is present by using different sized anchors. So the anchor with the highest IoU/IoU above some threshold with the groundtruth box will be labeled as positive, and negative if IoU is low. This is how we can train the region proposal network and is only one part of the Mask RCNN network.</p>\n<p>So my question is, if we want to fine tune mask-RCNN, do we code these things ourselves? Whats the inputs and whats the targets? Do we have to build the groundtruth anchor boxes, label them as positive or negative e.t.c, or pytorch/tensorflow will have some built in method to do it? </p>\n<p>Thanks so much for the help!!!</p>",
  "messages": [
    {
      "id": "1628127",
      "postDate": "12/24/2021 14:37:00",
      "content": "<p>Hi all. I am new to instance segmentation and the method of Mask RCNN.</p>\n<p>I wonder if anyone can briefly outline the steps needed to train Mask-RCNN, that will be very helpful.</p>\n<p>Lets take the region proposal network as an example. For every output feature map, the network has to learn whether an object is present by using different sized anchors. So the anchor with the highest IoU/IoU above some threshold with the groundtruth box will be labeled as positive, and negative if IoU is low. This is how we can train the region proposal network and is only one part of the Mask RCNN network.</p>\n<p>So my question is, if we want to fine tune mask-RCNN, do we code these things ourselves? Whats the inputs and whats the targets? Do we have to build the groundtruth anchor boxes, label them as positive or negative e.t.c, or pytorch/tensorflow will have some built in method to do it? </p>\n<p>Thanks so much for the help!!!</p>",
      "rawMarkdown": "Hi all. I am new to instance segmentation and the method of Mask RCNN.\n\nI wonder if anyone can briefly outline the steps needed to train Mask-RCNN, that will be very helpful.\n\nLets take the region proposal network as an example. For every output feature map, the network has to learn whether an object is present by using different sized anchors. So the anchor with the highest IoU/IoU above some threshold with the groundtruth box will be labeled as positive, and negative if IoU is low. This is how we can train the region proposal network and is only one part of the Mask RCNN network.\n\nSo my question is, if we want to fine tune mask-RCNN, do we code these things ourselves? Whats the inputs and whats the targets? Do we have to build the groundtruth anchor boxes, label them as positive or negative e.t.c, or pytorch/tensorflow will have some built in method to do it? \n\nThanks so much for the help!!!",
      "votes": null
    },
    {
      "id": "1628388",
      "postDate": "12/24/2021 19:10:58",
      "content": "<p>Well, according to the name of my team, I am not an expert but I can give a try to answer your question(s).</p>\n<p>Basically  if you are doing instance segmentation, you should use a framework that is already organizing the entire pipeline (backbone, FPN, RPN, heads). For pytorch, currently the two most frequently used frameworks are mmdetection (<a href=\"https://github.com/open-mmlab/mmdetection\" target=\"_blank\">https://github.com/open-mmlab/mmdetection</a>) and Detectron2 (<a href=\"https://github.com/facebookresearch/detectron2)\" target=\"_blank\">https://github.com/facebookresearch/detectron2)</a>. For tensorflow, it's mostly done with the TF Object Detection API (<a href=\"https://github.com/tensorflow/models/tree/master/research/object_detection\" target=\"_blank\">https://github.com/tensorflow/models/tree/master/research/object_detection</a>) but I don't have any experience with it.</p>\n<p>Since these frameworks are doing all the internal job, the inputs are simply the images and the targets are  1) object classes,  2) bounding boxes (instance object detection) 3) +- masks (instance segmentation). The easiest way to define these is by using Coco format .json files. All frameworks are directly supporting this annotation format. It can be useful to do transfer learning by  using pretrained weights (from Imagenet for backbone, and/or Coco for backbones, RPN and heads, or other pretraining methods)</p>\n<p>About the anchors boxes, it's usually a good idea to match the box sizes to the range of sizes of the object that you aim to detect/segment. I hope this answers your question.</p>",
      "rawMarkdown": "Well, according to the name of my team, I am not an expert but I can give a try to answer your question(s).\n\nBasically  if you are doing instance segmentation, you should use a framework that is already organizing the entire pipeline (backbone, FPN, RPN, heads). For pytorch, currently the two most frequently used frameworks are mmdetection (https://github.com/open-mmlab/mmdetection) and Detectron2 (https://github.com/facebookresearch/detectron2). For tensorflow, it's mostly done with the TF Object Detection API (https://github.com/tensorflow/models/tree/master/research/object_detection) but I don't have any experience with it.\n\nSince these frameworks are doing all the internal job, the inputs are simply the images and the targets are  1) object classes,  2) bounding boxes (instance object detection) 3) +- masks (instance segmentation). The easiest way to define these is by using Coco format .json files. All frameworks are directly supporting this annotation format. It can be useful to do transfer learning by  using pretrained weights (from Imagenet for backbone, and/or Coco for backbones, RPN and heads, or other pretraining methods)\n\nAbout the anchors boxes, it's usually a good idea to match the box sizes to the range of sizes of the object that you aim to detect/segment. I hope this answers your question.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1628388,
      "author_name": "alexandrecc",
      "author_url": "",
      "post_date": "12/24/2021 19:10:58",
      "content": "<p>Well, according to the name of my team, I am not an expert but I can give a try to answer your question(s).</p>\n<p>Basically  if you are doing instance segmentation, you should use a framework that is already organizing the entire pipeline (backbone, FPN, RPN, heads). For pytorch, currently the two most frequently used frameworks are mmdetection (<a href=\"https://github.com/open-mmlab/mmdetection\" target=\"_blank\">https://github.com/open-mmlab/mmdetection</a>) and Detectron2 (<a href=\"https://github.com/facebookresearch/detectron2)\" target=\"_blank\">https://github.com/facebookresearch/detectron2)</a>. For tensorflow, it's mostly done with the TF Object Detection API (<a href=\"https://github.com/tensorflow/models/tree/master/research/object_detection\" target=\"_blank\">https://github.com/tensorflow/models/tree/master/research/object_detection</a>) but I don't have any experience with it.</p>\n<p>Since these frameworks are doing all the internal job, the inputs are simply the images and the targets are  1) object classes,  2) bounding boxes (instance object detection) 3) +- masks (instance segmentation). The easiest way to define these is by using Coco format .json files. All frameworks are directly supporting this annotation format. It can be useful to do transfer learning by  using pretrained weights (from Imagenet for backbone, and/or Coco for backbones, RPN and heads, or other pretraining methods)</p>\n<p>About the anchors boxes, it's usually a good idea to match the box sizes to the range of sizes of the object that you aim to detect/segment. I hope this answers your question.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1628127": "Hi all. I am new to instance segmentation and the method of Mask RCNN.\n\nI wonder if anyone can briefly outline the steps needed to train Mask-RCNN, that will be very helpful.\n\nLets take the region proposal network as an example. For every output feature map, the network has to learn whether an object is present by using different sized anchors. So the anchor with the highest IoU/IoU above some threshold with the groundtruth box will be labeled as positive, and negative if IoU is low. This is how we can train the region proposal network and is only one part of the Mask RCNN network.\n\nSo my question is, if we want to fine tune mask-RCNN, do we code these things ourselves? Whats the inputs and whats the targets? Do we have to build the groundtruth anchor boxes, label them as positive or negative e.t.c, or pytorch/tensorflow will have some built in method to do it? \n\nThanks so much for the help!!!",
    "1628388": "Well, according to the name of my team, I am not an expert but I can give a try to answer your question(s).\n\nBasically  if you are doing instance segmentation, you should use a framework that is already organizing the entire pipeline (backbone, FPN, RPN, heads). For pytorch, currently the two most frequently used frameworks are mmdetection (https://github.com/open-mmlab/mmdetection) and Detectron2 (https://github.com/facebookresearch/detectron2). For tensorflow, it's mostly done with the TF Object Detection API (https://github.com/tensorflow/models/tree/master/research/object_detection) but I don't have any experience with it.\n\nSince these frameworks are doing all the internal job, the inputs are simply the images and the targets are  1) object classes,  2) bounding boxes (instance object detection) 3) +- masks (instance segmentation). The easiest way to define these is by using Coco format .json files. All frameworks are directly supporting this annotation format. It can be useful to do transfer learning by  using pretrained weights (from Imagenet for backbone, and/or Coco for backbones, RPN and heads, or other pretraining methods)\n\nAbout the anchors boxes, it's usually a good idea to match the box sizes to the range of sizes of the object that you aim to detect/segment. I hope this answers your question."
  },
  "source": "meta"
}