{
  "id": 289895,
  "title": "Several backbone for instance segmentation",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/289895",
  "author_name": "",
  "post_date": "2021-11-22T15:00:10.744030500Z",
  "votes": 10,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I wonder what is the appropriate backbone for instance (or semantic) segmentation.<br>\nThus I surveyed the paper briefly.<br>\nHere is the shortlist of 6 papers with excerpt of abstract.<br>\nThose papers are listed in historical order(If published in the same year, lexicographical order is used). <br>\nWe cannot try all but they may give us some insight to choose a backbone.</p>\n<p><strong>2020</strong> </p>\n<ul>\n<li><a href=\"http://vigir.missouri.edu/~gdesouza/Research/Conference_CDs/IEEE_WCCI_2020/IJCNN/Papers/N-21295.pdf\" target=\"_blank\">A comparative analysis of multi-backbone Mask R-CNN for surgical tools detection</a></li>\n</ul>\n<blockquote>\n  <p>Real-time surgical tool segmentation and tracking based on convolutional neural networks (CNN) has gained increasing interest in the field of mini-invasive surgery. In fact, the application of this novel artificial vision technologies allows both to reduce surgical risks and to increase patient safety. Moreover, these types of models can be used both to track the tools and detect markers or external artefacts in a real-time video stream. Multiple object detection and instance segmentation can be addressed efficiently by leveraging region-based CNN models. Thus, this work provides a comparison among state-of-the-art multi-backbone Mask R-CNNs to solve these tasks. Moreover, we show that such models can serve as a basis for tracking algorithms. The models were trained and tested with a data-set of 4955 manually annotated images, validated by 3 experts in the field. We tested 12 different combinations of CNN backbones and training hyperparameters. The results show that it is possible to employ a modern CNN to tackle the surgical tool detection problem, with the best-performing Mask R-CNN configuration achieving 87% Average Precision (AP) at Intersection over Union (IOU) 0.5.</p>\n</blockquote>\n<p><strong>2020</strong> </p>\n<ul>\n<li><a href=\"https://ieeexplore.ieee.org/document/9264148\" target=\"_blank\">A New Backbone Network for Instance Segmentation: Application on a Semiconductor Process Inspection</a></li>\n</ul>\n<blockquote>\n  <p>In this paper, we propose Instance Segmentation Detector (ISD) to extract the enhanced feature-maps under the situations where training dataset is limited in the specific industry domain such as semiconductor photo lithography inspection. ISD is used as a new backbone network of state-of-the-art Mask R-CNN framework for instance segmentation. ISD consists of four dense blocks and four transition layers. Each dense block in ISD has the shortcut connection and the concatenation of the feature-maps produced in layer with dynamic growth rate. ISD is trained from scratch without using recently approached transfer learning method. Additionally, ISD is trained with image dataset pre-processed by means of the specific designed image filter to extract the better enhanced feature map of Convolutional Neural Network (CNN). In ISD, one of the key principles is the compactness, plays a critical role for addressing real time problem and for application on resource bounded devices. To validate the model, this paper uses the real image collected from the computer vision system embedded in the currently operating semiconductor manufacturing equipment. ISD achieves consistently better results than state-of-the-art methods at the standard mean average precision. Specifically, our ISD outperforms baseline method DenseNet, while requiring only 1/4 parameters. We also observe that ISD can achieve comparable better results than ResNet, with only much smaller 1/268 parameters, using no extra data or pre-trained models.</p>\n</blockquote>\n<p><strong>2020</strong></p>\n<ul>\n<li><a href=\"https://arxiv.org/abs/1909.03625\" target=\"_blank\">CBNet: A Novel Composite Backbone Network Architecture for Object Detection</a></li>\n</ul>\n<blockquote>\n  <p>In existing CNN based detectors, the backbone network is a very important component for basic feature1 extraction, and the performance of the detectors highly depends on it. In this paper, we aim to achieve better detection performance by building a more powerful backbone from existing ones like ResNet and ResNeXt. Specifically, we propose a novel strategy for assembling multiple identical backbones by composite connections between the adjacent backbones, to form a more powerful backbone named Composite Backbone Network (CBNet). In this way, CBNet iteratively feeds the output features of the previous backbone, namely high-level features, as part of input features to the succeeding backbone, in a stage-by-stage fashion, and finally the feature maps of the last backbone (named Lead Backbone) are used for object detection. We show that CBNet can be very easily integrated into most state-of-the-art detectors and significantly improve their performances. For example, it boosts the mAP of FPN, Mask R-CNN and Cascade R-CNN on the COCO dataset by about 1.5 to 3.0 points. Moreover, experimental results show that the instance segmentation results can be improved as well. Specifically, by simply integrating the proposed CBNet into the baseline detector Cascade Mask R-CNN, we achieve a new state-of-the-art result on COCO dataset (mAP of 53.3) with a single model, which demonstrates great effectivenessof the proposed CBNet architecture. Code will be made available at <a href=\"https://github.com/PKUbahuangliuhe/CBNet\" target=\"_blank\">https://github.com/PKUbahuangliuhe/CBNet</a>.</p>\n</blockquote>\n<p><strong>2020</strong></p>\n<ul>\n<li><a href=\"https://iopscience.iop.org/article/10.1088/1742-6596/1544/1/012196/pdf\" target=\"_blank\">Comparison of Backbones for Semantic Segmentation Network </a></li>\n</ul>\n<blockquote>\n  <p>As for the classification network that is constantly emerging with each passing  day, different classification network as the backbone of the semantic segmentation  network may show different performance. This paper selected the road extraction data  set of CVPR DeepGlobe, and compared the performance differences of VGG-16 as the  backbone of Unet, ResNet34, ResNet101 and Xception as the backbone of AD-LinkNet.  When VGG-16 is used as the backbone of the semantic segmentation network, it  performs better in the face of long and wide road extraction. As the backbone of the semantic segmentation network, ResNet has a higher ability to extract small roads. <br>\n  When Xception is used as the backbone of the semantic segmentation network, it not only retains the characteristics of ResNet34, but also can effectively deal with the complex situation of extracting target covered by occlusions.</p>\n</blockquote>\n<p><strong>2020</strong></p>\n<ul>\n<li><a href=\"https://arxiv.org/pdf/1912.05027.pdf\" target=\"_blank\">SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization</a></li>\n</ul>\n<blockquote>\n  <p>Convolutional neural networks typically encode an input image into a series of intermediate features with decreasing resolutions. While this structure is suited to classification tasks, it does not perform well for tasks requiring simultaneous recognition and localization (e.g., object detection). The encoder-decoder architectures are proposed to resolve this by applying a decoder network onto a backbone model designed for classification tasks. In this paper, we argue encoder-decoder architecture is ineffective in generating strong multi-scale features because of the scale-decreased backbone. We propose SpineNet, a backbone with scale-permuted intermediate features and cross-scale connections that is learned on an object detection task by Neural Architecture Search. Using similar building blocks, SpineNet models outperform ResNet-FPN models by ∼3% AP at various scales while using 10-20% fewer FLOPs. In particular, SpineNet-190 achieves 52.5% AP with a Mask R-CNN detector and achieves 52.1% AP with a RetinaNet detector on COCO for a single model without test-time augmentation, significantly outperforms prior art of detectors. SpineNet can transfer to classification tasks, achieving 5% top-1 accuracy improvement on a challenging iNaturalist fine-grained dataset. Code is at: <a href=\"https://github.com/tensorflow/tpu/tree/master/models/official/detection\" target=\"_blank\">https://github.com/tensorflow/tpu/tree/master/models/official/detection</a>.</p>\n</blockquote>\n<p><strong>2018</strong></p>\n<ul>\n<li><a href=\"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&amp;arnumber=8531594\" target=\"_blank\">Exploring New Backbone and Attention Module for Semantic Segmentation in Street Scenes</a></li>\n</ul>\n<blockquote>\n  <p>Semantic segmentation, as dense pixel-wise classification task, played an important tache in scene understanding. There are two main challenges in many state-of-the-art works: 1) most backbone of segmentation models that often were extracted from pretrained classification models generated poor performance in small categories because they were lacking in spatial information and 2) the gap of combination between high-level and low-level features in segmentation models has led to inaccurate predictions.<br>\n  To handle these challenges, in this paper, we proposed a new tailored backbone and attention select module for segmentation tasks. Specifically, our new backbone was modified from the original Resnet, which can yield better segmentation performance. Attention select module employed spatial and channel self-attention mechanism to reinforce the propagation of contextual features, which can aggregate semantic and spatial information simultaneously. In addition, based on our new backbone and attention select module, we further proposed our segmentation model for street scenes understanding. We conducted a series of ablation studies on two public benchmarks, including Cityscapes and CamVid dataset to demonstrate the effectiveness of our proposals. Our model achieved a mIoU score of 71.5% on the Cityscapes test set with only fine annotation data and 60.1% on the CamVid test set.</p>\n</blockquote>",
  "messages": [
    {
      "id": "1591693",
      "postDate": "11/22/2021 15:00:10",
      "content": "<p>I wonder what is the appropriate backbone for instance (or semantic) segmentation.<br>\nThus I surveyed the paper briefly.<br>\nHere is the shortlist of 6 papers with excerpt of abstract.<br>\nThose papers are listed in historical order(If published in the same year, lexicographical order is used). <br>\nWe cannot try all but they may give us some insight to choose a backbone.</p>\n<p><strong>2020</strong> </p>\n<ul>\n<li><a href=\"http://vigir.missouri.edu/~gdesouza/Research/Conference_CDs/IEEE_WCCI_2020/IJCNN/Papers/N-21295.pdf\" target=\"_blank\">A comparative analysis of multi-backbone Mask R-CNN for surgical tools detection</a></li>\n</ul>\n<blockquote>\n  <p>Real-time surgical tool segmentation and tracking based on convolutional neural networks (CNN) has gained increasing interest in the field of mini-invasive surgery. In fact, the application of this novel artificial vision technologies allows both to reduce surgical risks and to increase patient safety. Moreover, these types of models can be used both to track the tools and detect markers or external artefacts in a real-time video stream. Multiple object detection and instance segmentation can be addressed efficiently by leveraging region-based CNN models. Thus, this work provides a comparison among state-of-the-art multi-backbone Mask R-CNNs to solve these tasks. Moreover, we show that such models can serve as a basis for tracking algorithms. The models were trained and tested with a data-set of 4955 manually annotated images, validated by 3 experts in the field. We tested 12 different combinations of CNN backbones and training hyperparameters. The results show that it is possible to employ a modern CNN to tackle the surgical tool detection problem, with the best-performing Mask R-CNN configuration achieving 87% Average Precision (AP) at Intersection over Union (IOU) 0.5.</p>\n</blockquote>\n<p><strong>2020</strong> </p>\n<ul>\n<li><a href=\"https://ieeexplore.ieee.org/document/9264148\" target=\"_blank\">A New Backbone Network for Instance Segmentation: Application on a Semiconductor Process Inspection</a></li>\n</ul>\n<blockquote>\n  <p>In this paper, we propose Instance Segmentation Detector (ISD) to extract the enhanced feature-maps under the situations where training dataset is limited in the specific industry domain such as semiconductor photo lithography inspection. ISD is used as a new backbone network of state-of-the-art Mask R-CNN framework for instance segmentation. ISD consists of four dense blocks and four transition layers. Each dense block in ISD has the shortcut connection and the concatenation of the feature-maps produced in layer with dynamic growth rate. ISD is trained from scratch without using recently approached transfer learning method. Additionally, ISD is trained with image dataset pre-processed by means of the specific designed image filter to extract the better enhanced feature map of Convolutional Neural Network (CNN). In ISD, one of the key principles is the compactness, plays a critical role for addressing real time problem and for application on resource bounded devices. To validate the model, this paper uses the real image collected from the computer vision system embedded in the currently operating semiconductor manufacturing equipment. ISD achieves consistently better results than state-of-the-art methods at the standard mean average precision. Specifically, our ISD outperforms baseline method DenseNet, while requiring only 1/4 parameters. We also observe that ISD can achieve comparable better results than ResNet, with only much smaller 1/268 parameters, using no extra data or pre-trained models.</p>\n</blockquote>\n<p><strong>2020</strong></p>\n<ul>\n<li><a href=\"https://arxiv.org/abs/1909.03625\" target=\"_blank\">CBNet: A Novel Composite Backbone Network Architecture for Object Detection</a></li>\n</ul>\n<blockquote>\n  <p>In existing CNN based detectors, the backbone network is a very important component for basic feature1 extraction, and the performance of the detectors highly depends on it. In this paper, we aim to achieve better detection performance by building a more powerful backbone from existing ones like ResNet and ResNeXt. Specifically, we propose a novel strategy for assembling multiple identical backbones by composite connections between the adjacent backbones, to form a more powerful backbone named Composite Backbone Network (CBNet). In this way, CBNet iteratively feeds the output features of the previous backbone, namely high-level features, as part of input features to the succeeding backbone, in a stage-by-stage fashion, and finally the feature maps of the last backbone (named Lead Backbone) are used for object detection. We show that CBNet can be very easily integrated into most state-of-the-art detectors and significantly improve their performances. For example, it boosts the mAP of FPN, Mask R-CNN and Cascade R-CNN on the COCO dataset by about 1.5 to 3.0 points. Moreover, experimental results show that the instance segmentation results can be improved as well. Specifically, by simply integrating the proposed CBNet into the baseline detector Cascade Mask R-CNN, we achieve a new state-of-the-art result on COCO dataset (mAP of 53.3) with a single model, which demonstrates great effectivenessof the proposed CBNet architecture. Code will be made available at <a href=\"https://github.com/PKUbahuangliuhe/CBNet\" target=\"_blank\">https://github.com/PKUbahuangliuhe/CBNet</a>.</p>\n</blockquote>\n<p><strong>2020</strong></p>\n<ul>\n<li><a href=\"https://iopscience.iop.org/article/10.1088/1742-6596/1544/1/012196/pdf\" target=\"_blank\">Comparison of Backbones for Semantic Segmentation Network </a></li>\n</ul>\n<blockquote>\n  <p>As for the classification network that is constantly emerging with each passing  day, different classification network as the backbone of the semantic segmentation  network may show different performance. This paper selected the road extraction data  set of CVPR DeepGlobe, and compared the performance differences of VGG-16 as the  backbone of Unet, ResNet34, ResNet101 and Xception as the backbone of AD-LinkNet.  When VGG-16 is used as the backbone of the semantic segmentation network, it  performs better in the face of long and wide road extraction. As the backbone of the semantic segmentation network, ResNet has a higher ability to extract small roads. <br>\n  When Xception is used as the backbone of the semantic segmentation network, it not only retains the characteristics of ResNet34, but also can effectively deal with the complex situation of extracting target covered by occlusions.</p>\n</blockquote>\n<p><strong>2020</strong></p>\n<ul>\n<li><a href=\"https://arxiv.org/pdf/1912.05027.pdf\" target=\"_blank\">SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization</a></li>\n</ul>\n<blockquote>\n  <p>Convolutional neural networks typically encode an input image into a series of intermediate features with decreasing resolutions. While this structure is suited to classification tasks, it does not perform well for tasks requiring simultaneous recognition and localization (e.g., object detection). The encoder-decoder architectures are proposed to resolve this by applying a decoder network onto a backbone model designed for classification tasks. In this paper, we argue encoder-decoder architecture is ineffective in generating strong multi-scale features because of the scale-decreased backbone. We propose SpineNet, a backbone with scale-permuted intermediate features and cross-scale connections that is learned on an object detection task by Neural Architecture Search. Using similar building blocks, SpineNet models outperform ResNet-FPN models by ∼3% AP at various scales while using 10-20% fewer FLOPs. In particular, SpineNet-190 achieves 52.5% AP with a Mask R-CNN detector and achieves 52.1% AP with a RetinaNet detector on COCO for a single model without test-time augmentation, significantly outperforms prior art of detectors. SpineNet can transfer to classification tasks, achieving 5% top-1 accuracy improvement on a challenging iNaturalist fine-grained dataset. Code is at: <a href=\"https://github.com/tensorflow/tpu/tree/master/models/official/detection\" target=\"_blank\">https://github.com/tensorflow/tpu/tree/master/models/official/detection</a>.</p>\n</blockquote>\n<p><strong>2018</strong></p>\n<ul>\n<li><a href=\"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&amp;arnumber=8531594\" target=\"_blank\">Exploring New Backbone and Attention Module for Semantic Segmentation in Street Scenes</a></li>\n</ul>\n<blockquote>\n  <p>Semantic segmentation, as dense pixel-wise classification task, played an important tache in scene understanding. There are two main challenges in many state-of-the-art works: 1) most backbone of segmentation models that often were extracted from pretrained classification models generated poor performance in small categories because they were lacking in spatial information and 2) the gap of combination between high-level and low-level features in segmentation models has led to inaccurate predictions.<br>\n  To handle these challenges, in this paper, we proposed a new tailored backbone and attention select module for segmentation tasks. Specifically, our new backbone was modified from the original Resnet, which can yield better segmentation performance. Attention select module employed spatial and channel self-attention mechanism to reinforce the propagation of contextual features, which can aggregate semantic and spatial information simultaneously. In addition, based on our new backbone and attention select module, we further proposed our segmentation model for street scenes understanding. We conducted a series of ablation studies on two public benchmarks, including Cityscapes and CamVid dataset to demonstrate the effectiveness of our proposals. Our model achieved a mIoU score of 71.5% on the Cityscapes test set with only fine annotation data and 60.1% on the CamVid test set.</p>\n</blockquote>",
      "rawMarkdown": "I wonder what is the appropriate backbone for instance (or semantic) segmentation.\nThus I surveyed the paper briefly.\nHere is the shortlist of 6 papers with excerpt of abstract.\nThose papers are listed in historical order(If published in the same year, lexicographical order is used). \nWe cannot try all but they may give us some insight to choose a backbone.\n\n\n**2020** \n- [A comparative analysis of multi-backbone Mask R-CNN for surgical tools detection](http://vigir.missouri.edu/~gdesouza/Research/Conference_CDs/IEEE_WCCI_2020/IJCNN/Papers/N-21295.pdf)\n> Real-time surgical tool segmentation and tracking based on convolutional neural networks (CNN) has gained increasing interest in the field of mini-invasive surgery. In fact, the application of this novel artificial vision technologies allows both to reduce surgical risks and to increase patient safety. Moreover, these types of models can be used both to track the tools and detect markers or external artefacts in a real-time video stream. Multiple object detection and instance segmentation can be addressed efficiently by leveraging region-based CNN models. Thus, this work provides a comparison among state-of-the-art multi-backbone Mask R-CNNs to solve these tasks. Moreover, we show that such models can serve as a basis for tracking algorithms. The models were trained and tested with a data-set of 4955 manually annotated images, validated by 3 experts in the field. We tested 12 different combinations of CNN backbones and training hyperparameters. The results show that it is possible to employ a modern CNN to tackle the surgical tool detection problem, with the best-performing Mask R-CNN configuration achieving 87% Average Precision (AP) at Intersection over Union (IOU) 0.5.\n\n\n**2020** \n- [A New Backbone Network for Instance Segmentation: Application on a Semiconductor Process Inspection](https://ieeexplore.ieee.org/document/9264148)\n> In this paper, we propose Instance Segmentation Detector (ISD) to extract the enhanced feature-maps under the situations where training dataset is limited in the specific industry domain such as semiconductor photo lithography inspection. ISD is used as a new backbone network of state-of-the-art Mask R-CNN framework for instance segmentation. ISD consists of four dense blocks and four transition layers. Each dense block in ISD has the shortcut connection and the concatenation of the feature-maps produced in layer with dynamic growth rate. ISD is trained from scratch without using recently approached transfer learning method. Additionally, ISD is trained with image dataset pre-processed by means of the specific designed image filter to extract the better enhanced feature map of Convolutional Neural Network (CNN). In ISD, one of the key principles is the compactness, plays a critical role for addressing real time problem and for application on resource bounded devices. To validate the model, this paper uses the real image collected from the computer vision system embedded in the currently operating semiconductor manufacturing equipment. ISD achieves consistently better results than state-of-the-art methods at the standard mean average precision. Specifically, our ISD outperforms baseline method DenseNet, while requiring only 1/4 parameters. We also observe that ISD can achieve comparable better results than ResNet, with only much smaller 1/268 parameters, using no extra data or pre-trained models.\n\n\n**2020**\n- [CBNet: A Novel Composite Backbone Network Architecture for Object Detection](https://arxiv.org/abs/1909.03625)\n> In existing CNN based detectors, the backbone network is a very important component for basic feature1 extraction, and the performance of the detectors highly depends on it. In this paper, we aim to achieve better detection performance by building a more powerful backbone from existing ones like ResNet and ResNeXt. Specifically, we propose a novel strategy for assembling multiple identical backbones by composite connections between the adjacent backbones, to form a more powerful backbone named Composite Backbone Network (CBNet). In this way, CBNet iteratively feeds the output features of the previous backbone, namely high-level features, as part of input features to the succeeding backbone, in a stage-by-stage fashion, and finally the feature maps of the last backbone (named Lead Backbone) are used for object detection. We show that CBNet can be very easily integrated into most state-of-the-art detectors and significantly improve their performances. For example, it boosts the mAP of FPN, Mask R-CNN and Cascade R-CNN on the COCO dataset by about 1.5 to 3.0 points. Moreover, experimental results show that the instance segmentation results can be improved as well. Specifically, by simply integrating the proposed CBNet into the baseline detector Cascade Mask R-CNN, we achieve a new state-of-the-art result on COCO dataset (mAP of 53.3) with a single model, which demonstrates great effectivenessof the proposed CBNet architecture. Code will be made available at https://github.com/PKUbahuangliuhe/CBNet.\n\n\n**2020**\n- [Comparison of Backbones for Semantic Segmentation Network ](https://iopscience.iop.org/article/10.1088/1742-6596/1544/1/012196/pdf)\n> As for the classification network that is constantly emerging with each passing  day, different classification network as the backbone of the semantic segmentation  network may show different performance. This paper selected the road extraction data  set of CVPR DeepGlobe, and compared the performance differences of VGG-16 as the  backbone of Unet, ResNet34, ResNet101 and Xception as the backbone of AD-LinkNet.  When VGG-16 is used as the backbone of the semantic segmentation network, it  performs better in the face of long and wide road extraction. As the backbone of the semantic segmentation network, ResNet has a higher ability to extract small roads. \nWhen Xception is used as the backbone of the semantic segmentation network, it not only retains the characteristics of ResNet34, but also can effectively deal with the complex situation of extracting target covered by occlusions.\n\n\n**2020**\n- [SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization](https://arxiv.org/pdf/1912.05027.pdf)\n> Convolutional neural networks typically encode an input image into a series of intermediate features with decreasing resolutions. While this structure is suited to classification tasks, it does not perform well for tasks requiring simultaneous recognition and localization (e.g., object detection). The encoder-decoder architectures are proposed to resolve this by applying a decoder network onto a backbone model designed for classification tasks. In this paper, we argue encoder-decoder architecture is ineffective in generating strong multi-scale features because of the scale-decreased backbone. We propose SpineNet, a backbone with scale-permuted intermediate features and cross-scale connections that is learned on an object detection task by Neural Architecture Search. Using similar building blocks, SpineNet models outperform ResNet-FPN models by ∼3% AP at various scales while using 10-20% fewer FLOPs. In particular, SpineNet-190 achieves 52.5% AP with a Mask R-CNN detector and achieves 52.1% AP with a RetinaNet detector on COCO for a single model without test-time augmentation, significantly outperforms prior art of detectors. SpineNet can transfer to classification tasks, achieving 5% top-1 accuracy improvement on a challenging iNaturalist fine-grained dataset. Code is at: https://github.com/tensorflow/tpu/tree/master/models/official/detection.\n\n\n**2018**\n- [Exploring New Backbone and Attention Module for Semantic Segmentation in Street Scenes](https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=8531594)\n> Semantic segmentation, as dense pixel-wise classification task, played an important tache in scene understanding. There are two main challenges in many state-of-the-art works: 1) most backbone of segmentation models that often were extracted from pretrained classification models generated poor performance in small categories because they were lacking in spatial information and 2) the gap of combination between high-level and low-level features in segmentation models has led to inaccurate predictions.\nTo handle these challenges, in this paper, we proposed a new tailored backbone and attention select module for segmentation tasks. Specifically, our new backbone was modified from the original Resnet, which can yield better segmentation performance. Attention select module employed spatial and channel self-attention mechanism to reinforce the propagation of contextual features, which can aggregate semantic and spatial information simultaneously. In addition, based on our new backbone and attention select module, we further proposed our segmentation model for street scenes understanding. We conducted a series of ablation studies on two public benchmarks, including Cityscapes and CamVid dataset to demonstrate the effectiveness of our proposals. Our model achieved a mIoU score of 71.5% on the Cityscapes test set with only fine annotation data and 60.1% on the CamVid test set.",
      "votes": null
    },
    {
      "id": "1592864",
      "postDate": "11/23/2021 12:28:27",
      "content": "<p>Nice work！！</p>",
      "rawMarkdown": "Nice work！！",
      "votes": null
    },
    {
      "id": "1592872",
      "postDate": "11/23/2021 12:35:42",
      "content": "<p>Thank you. Arigatou:)</p>",
      "rawMarkdown": "Thank you. Arigatou:)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1592864,
      "author_name": "manyuli",
      "author_url": "",
      "post_date": "11/23/2021 12:28:27",
      "content": "<p>Nice work！！</p>",
      "votes": null,
      "replies": [
        {
          "id": 1592872,
          "author_name": "osamurai",
          "author_url": "",
          "post_date": "11/23/2021 12:35:42",
          "content": "<p>Thank you. Arigatou:)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1591693": "I wonder what is the appropriate backbone for instance (or semantic) segmentation.\nThus I surveyed the paper briefly.\nHere is the shortlist of 6 papers with excerpt of abstract.\nThose papers are listed in historical order(If published in the same year, lexicographical order is used). \nWe cannot try all but they may give us some insight to choose a backbone.\n\n\n**2020** \n- [A comparative analysis of multi-backbone Mask R-CNN for surgical tools detection](http://vigir.missouri.edu/~gdesouza/Research/Conference_CDs/IEEE_WCCI_2020/IJCNN/Papers/N-21295.pdf)\n> Real-time surgical tool segmentation and tracking based on convolutional neural networks (CNN) has gained increasing interest in the field of mini-invasive surgery. In fact, the application of this novel artificial vision technologies allows both to reduce surgical risks and to increase patient safety. Moreover, these types of models can be used both to track the tools and detect markers or external artefacts in a real-time video stream. Multiple object detection and instance segmentation can be addressed efficiently by leveraging region-based CNN models. Thus, this work provides a comparison among state-of-the-art multi-backbone Mask R-CNNs to solve these tasks. Moreover, we show that such models can serve as a basis for tracking algorithms. The models were trained and tested with a data-set of 4955 manually annotated images, validated by 3 experts in the field. We tested 12 different combinations of CNN backbones and training hyperparameters. The results show that it is possible to employ a modern CNN to tackle the surgical tool detection problem, with the best-performing Mask R-CNN configuration achieving 87% Average Precision (AP) at Intersection over Union (IOU) 0.5.\n\n\n**2020** \n- [A New Backbone Network for Instance Segmentation: Application on a Semiconductor Process Inspection](https://ieeexplore.ieee.org/document/9264148)\n> In this paper, we propose Instance Segmentation Detector (ISD) to extract the enhanced feature-maps under the situations where training dataset is limited in the specific industry domain such as semiconductor photo lithography inspection. ISD is used as a new backbone network of state-of-the-art Mask R-CNN framework for instance segmentation. ISD consists of four dense blocks and four transition layers. Each dense block in ISD has the shortcut connection and the concatenation of the feature-maps produced in layer with dynamic growth rate. ISD is trained from scratch without using recently approached transfer learning method. Additionally, ISD is trained with image dataset pre-processed by means of the specific designed image filter to extract the better enhanced feature map of Convolutional Neural Network (CNN). In ISD, one of the key principles is the compactness, plays a critical role for addressing real time problem and for application on resource bounded devices. To validate the model, this paper uses the real image collected from the computer vision system embedded in the currently operating semiconductor manufacturing equipment. ISD achieves consistently better results than state-of-the-art methods at the standard mean average precision. Specifically, our ISD outperforms baseline method DenseNet, while requiring only 1/4 parameters. We also observe that ISD can achieve comparable better results than ResNet, with only much smaller 1/268 parameters, using no extra data or pre-trained models.\n\n\n**2020**\n- [CBNet: A Novel Composite Backbone Network Architecture for Object Detection](https://arxiv.org/abs/1909.03625)\n> In existing CNN based detectors, the backbone network is a very important component for basic feature1 extraction, and the performance of the detectors highly depends on it. In this paper, we aim to achieve better detection performance by building a more powerful backbone from existing ones like ResNet and ResNeXt. Specifically, we propose a novel strategy for assembling multiple identical backbones by composite connections between the adjacent backbones, to form a more powerful backbone named Composite Backbone Network (CBNet). In this way, CBNet iteratively feeds the output features of the previous backbone, namely high-level features, as part of input features to the succeeding backbone, in a stage-by-stage fashion, and finally the feature maps of the last backbone (named Lead Backbone) are used for object detection. We show that CBNet can be very easily integrated into most state-of-the-art detectors and significantly improve their performances. For example, it boosts the mAP of FPN, Mask R-CNN and Cascade R-CNN on the COCO dataset by about 1.5 to 3.0 points. Moreover, experimental results show that the instance segmentation results can be improved as well. Specifically, by simply integrating the proposed CBNet into the baseline detector Cascade Mask R-CNN, we achieve a new state-of-the-art result on COCO dataset (mAP of 53.3) with a single model, which demonstrates great effectivenessof the proposed CBNet architecture. Code will be made available at https://github.com/PKUbahuangliuhe/CBNet.\n\n\n**2020**\n- [Comparison of Backbones for Semantic Segmentation Network ](https://iopscience.iop.org/article/10.1088/1742-6596/1544/1/012196/pdf)\n> As for the classification network that is constantly emerging with each passing  day, different classification network as the backbone of the semantic segmentation  network may show different performance. This paper selected the road extraction data  set of CVPR DeepGlobe, and compared the performance differences of VGG-16 as the  backbone of Unet, ResNet34, ResNet101 and Xception as the backbone of AD-LinkNet.  When VGG-16 is used as the backbone of the semantic segmentation network, it  performs better in the face of long and wide road extraction. As the backbone of the semantic segmentation network, ResNet has a higher ability to extract small roads. \nWhen Xception is used as the backbone of the semantic segmentation network, it not only retains the characteristics of ResNet34, but also can effectively deal with the complex situation of extracting target covered by occlusions.\n\n\n**2020**\n- [SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization](https://arxiv.org/pdf/1912.05027.pdf)\n> Convolutional neural networks typically encode an input image into a series of intermediate features with decreasing resolutions. While this structure is suited to classification tasks, it does not perform well for tasks requiring simultaneous recognition and localization (e.g., object detection). The encoder-decoder architectures are proposed to resolve this by applying a decoder network onto a backbone model designed for classification tasks. In this paper, we argue encoder-decoder architecture is ineffective in generating strong multi-scale features because of the scale-decreased backbone. We propose SpineNet, a backbone with scale-permuted intermediate features and cross-scale connections that is learned on an object detection task by Neural Architecture Search. Using similar building blocks, SpineNet models outperform ResNet-FPN models by ∼3% AP at various scales while using 10-20% fewer FLOPs. In particular, SpineNet-190 achieves 52.5% AP with a Mask R-CNN detector and achieves 52.1% AP with a RetinaNet detector on COCO for a single model without test-time augmentation, significantly outperforms prior art of detectors. SpineNet can transfer to classification tasks, achieving 5% top-1 accuracy improvement on a challenging iNaturalist fine-grained dataset. Code is at: https://github.com/tensorflow/tpu/tree/master/models/official/detection.\n\n\n**2018**\n- [Exploring New Backbone and Attention Module for Semantic Segmentation in Street Scenes](https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=8531594)\n> Semantic segmentation, as dense pixel-wise classification task, played an important tache in scene understanding. There are two main challenges in many state-of-the-art works: 1) most backbone of segmentation models that often were extracted from pretrained classification models generated poor performance in small categories because they were lacking in spatial information and 2) the gap of combination between high-level and low-level features in segmentation models has led to inaccurate predictions.\nTo handle these challenges, in this paper, we proposed a new tailored backbone and attention select module for segmentation tasks. Specifically, our new backbone was modified from the original Resnet, which can yield better segmentation performance. Attention select module employed spatial and channel self-attention mechanism to reinforce the propagation of contextual features, which can aggregate semantic and spatial information simultaneously. In addition, based on our new backbone and attention select module, we further proposed our segmentation model for street scenes understanding. We conducted a series of ablation studies on two public benchmarks, including Cityscapes and CamVid dataset to demonstrate the effectiveness of our proposals. Our model achieved a mIoU score of 71.5% on the Cityscapes test set with only fine annotation data and 60.1% on the CamVid test set.",
    "1592864": "Nice work！！",
    "1592872": "Thank you. Arigatou:)"
  },
  "source": "meta"
}