{
  "id": 284135,
  "title": "Summary of every SOTA instance segmentation paper of the past 2 years ",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/284135",
  "author_name": "",
  "post_date": "2021-10-29T19:54:29.165964700Z",
  "votes": 75,
  "comment_count": 4,
  "views": 0,
  "content": "<p>'<br>\nHey Guys, </p>\n<p>I started going through some recent relevant SOTA papers for ideas, here is my summary so far: (</p>\n<hr>\n<p><strong>End-to-End Semi-Supervised Object Detection with Soft Teacher | 2021</strong></p>\n<ul>\n<li>Models: Soft Teacher + Swin-L (HTC++, multi-scale)</li>\n<li>propose a soft teacher mechanism where the classification loss of each unlabeled bounding box is weighed by the classification score produced by the teacher network</li>\n<li>As well as a box jittering approach to select reliable pseudo boxes for the learning of box regression</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2106.09018v3.pdf\" target=\"_blank\">https://arxiv.org/pdf/2106.09018v3.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>CBNetV2: A Composite Backbone Network Architecture for Object Detection | 2021</strong></p>\n<ul>\n<li>Models: Dual-Swin-L (HTC, multi-scale)</li>\n<li>Try out many existing open-sourced pre-trained backbones.</li>\n<li>Can be used as a nice review.<br>\nread more</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2107.00420v6.pdf\" target=\"_blank\">https://arxiv.org/pdf/2107.00420v6.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>Focal Self-attention for Local-Global Interactions in Vision Transformers | 2021</strong></p>\n<ul>\n<li>Models: Focal-L (HTC++, multi-scale)</li>\n<li>Present focal self-attention for vision transformers, each token attends the closest surrounding tokens at fine granularity but the tokens far away at coarse granularity</li>\n<li>Propose a variant of Vision Transformer models, substantial improvements over the current state-of-the-art Swin Transformers for 6 different object detection methods.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2107.00641v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/2107.00641v1.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>Swin Transformer: Hierarchical Vision Transformer using Shifted Windows | 2021</strong></p>\n<ul>\n<li>Models: Swin-L (HTC++, multi scale)</li>\n<li>Abstract: Hierarchical Transformer whose representation is computed with shifted windows. The shifted windowing scheme brings greater efficiency by limiting self-attention computation to non-overlapping local windows while also allowing for cross-window connection.</li>\n<li>Spinoffs:</li>\n<li>Used as a backbone for Self-Supervised Learning: Transformer-SSL</li>\n<li>Add Swin MLP, which is an adaption of Swin Transformer by replacing all multi-head self-attention (MHSA) blocks by MLP layers (more precisely it is a group linear layer). The shifted window configuration can also significantly improve the performance of vanilla MLP architectures.</li>\n<li>Soft Teacher: An end-to-end semi-supervisd object detection method, achieving a new record on the COCO test-dev.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2103.14030v2.pdf\" target=\"_blank\">https://arxiv.org/pdf/2103.14030v2.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>Instances as Queries | 2021</strong></p>\n<ul>\n<li>Models: QueryInst (single scale)</li>\n<li>This paper proposed a simple and effective query based instance segmentation method driven by parallel supervision on dynamic mask heads, which outperforms previous arts in terms of both accuracy and speed.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2105.01928v3.pdf\" target=\"_blank\">https://arxiv.org/pdf/2105.01928v3.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation | 2020</strong></p>\n<ul>\n<li>Models: Cascade Eff-B7 NAS-FPN (1280, self-training Copy Paste, single-scale)</li>\n<li>Show that the simple mechanism of pasting objects randomly is good enough and can provide solid gains on top of strong baselines.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2012.07177v2.pdf\" target=\"_blank\">https://arxiv.org/pdf/2012.07177v2.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>DetectoRS: Detecting Objects with Recursive Feature Pyramid and Switchable Atrous Convolution | 2020</strong></p>\n<ul>\n<li>Models: DetectoRS (ResNeXt-101-64x4d, multi-scale)</li>\n<li>Abstract: Propose Recursive Feature Pyramid, to incorporate extra feedback connections from Feature Pyramid Networks into the bottom-up backbone layers and Switchable Atrous Convolutions on the micro-level.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2006.02334v2.pdf\" target=\"_blank\">https://arxiv.org/pdf/2006.02334v2.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>SOLQ: Segmenting Objects by Learning Queries | 2021</strong></p>\n<ul>\n<li>Models: SOLQ (Swin-L, single scale)</li>\n<li>Propose an end-to-end framework for instance segmentation. Based on DETR, the sifference is that they segment objects by learning unified queries.</li>\n<li>During training phase, the mask vectors encoded are supervised by the compression coding of raw spatial masks</li>\n<li>Shot that the joint learning of unified query representation can greatly improve the detection performance of DETR.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2106.02351v3.pdf\" target=\"_blank\">https://arxiv.org/pdf/2106.02351v3.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization | 2019</strong></p>\n<ul>\n<li>Models: Mask R-CNN (SpineNet-190, 1536x1536)</li>\n<li>Propose SpineNet, a backbone with scale-permuted intermediate features and cross-scale connections that is learned on an object detection task by Neural Architecture Search.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/1912.05027v3.pdf\" target=\"_blank\">https://arxiv.org/pdf/1912.05027v3.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>Global Context Networks | 2020</strong></p>\n<ul>\n<li>Models: GCNet (ResNeXt-101 + DCN + cascade + GC r4)</li>\n<li>Replace the one-layer transformation function of the non-local block of NLNet by a two-layer bottleneck, which reduces the parameter number considerably.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2012.13375v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/2012.13375v1.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>CBNet: A Novel Composite Backbone Network Architecture for Object Detection | 2019</strong></p>\n<ul>\n<li>Models: Cascade Mask R-CNN (ResNeXt152, CBNet)</li>\n<li>Propose a novel strategy for assembling multiple identical backbones by composite connections between the adjacent backbones, to form a more powerful backbone named Composite Backbone Network (CBNet).</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/1909.03625v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/1909.03625v1.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>ResNeSt: Split-Attention Networks | 2020</strong></p>\n<ul>\n<li>Models: ResNeSt101</li>\n<li>Present a modularized architecture, which applies the channel-wise attention on different network branches to leverage their success in capturing cross-feature interactions and learning diverse representations. Our design results in a simple and unified computation block, which can be parameterized using only a few variables.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2004.08955v2.pdf\" target=\"_blank\">https://arxiv.org/pdf/2004.08955v2.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>Path Aggregation Network for Instance Segmentation | 2018</strong></p>\n<ul>\n<li>Models: PANet</li>\n<li>Presents adaptive feature pooling, which links feature grid and all feature levels to make useful information in each feature level propagate directly to following proposal subnetworks.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/1803.01534v4.pdf\" target=\"_blank\">https://arxiv.org/pdf/1803.01534v4.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>CenterMask : Real-Time Anchor-Free Instance Segmentation | 2019</strong></p>\n<ul>\n<li>Models: CenterMask + VoVNet99</li>\n<li>Adds a spatial attention-guided mask branch to anchor-free one stage object detector in the same vein with Mask R-CNN.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/1911.06667v6.pdf\" target=\"_blank\">https://arxiv.org/pdf/1911.06667v6.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>SOLOv2: Dynamic and Fast Instance Segmentation | 2020</strong></p>\n<ul>\n<li>Models: SOLOv2(Res-DCN-101-FPN)</li>\n<li>Dynamically learning the mask head of the object segmenter such that the mask head is conditioned on the location. The mask branch is decoupled into a mask kernel branch and mask feature branch, which are responsible for learning the convolution kernel and the convolved features respectively.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2003.10152v3.pdf\" target=\"_blank\">https://arxiv.org/pdf/2003.10152v3.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>Deep Occlusion-Aware Instance Segmentation with Overlapping BiLayers | 2021</strong></p>\n<ul>\n<li>Models: BCNet(ResNeXt-101 + FPN+ FCOS)</li>\n<li>Model image formation as composition of two overlapping layers, and propose Bilayer Convolutional Network (BCNet), where the top GCN layer detects the occluding objects (occluder) and the bottom GCN layer infers partially occluded instance. In general: Bilayer = Conv -&gt; GCN -&gt; FNC</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2103.12340v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/2103.12340v1.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>BlendMask: Top-Down Meets Bottom-Up for Instance Segmentation | 2020</strong></p>\n<ul>\n<li>Models: BlendMask (ResNet-101 + DCN interval=3)</li>\n<li>Improved mask prediction by effectively combining instance-level information with semantic information with lower-level fine-granularity The main contribution is a blender module which draws inspiration from both top-down and bottom-up instance segmentation approaches.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2001.00309v3.pdf\" target=\"_blank\">https://arxiv.org/pdf/2001.00309v3.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>K-Net: Towards Unified Image Segmentation | 2021</strong></p>\n<ul>\n<li>Models: K-Net-N256 (ResNet-101)</li>\n<li>Propose a kernel update strategy that enables each kernel dynamic and conditional on its meaningful group in the input image.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2106.14855v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/2106.14855v1.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>D2Det: Towards High Quality Object Detection and Instance Segmentation | 2020</strong></p>\n<ul>\n<li>Models: D2Det (ResNet-101, single-scale test)</li>\n<li>For precise localization, they introduced a dense local regression that predicts multiple dense box offsets for an object proposal.</li>\n<li>The dense local regression is not limited to a quantized set of keypoints within a fixed region and has the ability to regress position-sensitive real number dense</li>\n<li>Thedense local regression is further improved by a binary overlap prediction strategy that reduces the influence of background region on the final box regression.</li>\n<li>For accurate classification, they introduced a discriminative RoI pooling scheme that samples from various sub-regions of a proposal and performs adaptive weighting to obtain discriminative features.</li>\n<li>Link: <a href=\"http://openaccess.thecvf.com/content_CVPR_2020/papers/Cao_D2Det_Towards_High_Quality_Object_Detection_and_Instance_Segmentation_CVPR_2020_paper.pdf\" target=\"_blank\">http://openaccess.thecvf.com/content_CVPR_2020/papers/Cao_D2Det_Towards_High_Quality_Object_Detection_and_Instance_Segmentation_CVPR_2020_paper.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>ISTR: End-to-End Instance Segmentation with Transformers | 2021</strong></p>\n<ul>\n<li>Models: ISTR (ResNet101-FPN-3x, single-scale)</li>\n<li>Proposed an instance segmentation Transformer, it predicts low-dimensional mask embeddings, and matches them with ground truth mask embeddings for the set loss.</li>\n<li>Also ir concurrently conducts detection and segmentation with a recurrent refinement strategy, which provides a new way to achieve instance segmentation compared to the existing top-down and bottom-up frameworks.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2105.00637v2.pdf\" target=\"_blank\">https://arxiv.org/pdf/2105.00637v2.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>SipMask: Spatial Information Preservation for Fast Image and Video Instance Segmentation | 2020</strong></p>\n<ul>\n<li>Models: SipMask (ResNet-101, single-scale test)</li>\n<li>Proposed a fast single-stage instance segmentation method that preserves instance-specific spatial information by separating mask prediction of an instance to different sub-regions of a detected bounding-box.</li>\n<li>A novel modeule: light-weight spatial preservation (SP) module that generates a separate set of spatial coefficients for each sub-region within a bounding-box, leading to improved mask predictions. </li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2007.14772v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/2007.14772v1.pdf</a></li>\n</ul>\n<hr>\n<p>'</p>\n<p>Happy Kaggling! </p>",
  "messages": [
    {
      "id": "1564940",
      "postDate": "10/29/2021 19:54:29",
      "content": "<p>'<br>\nHey Guys, </p>\n<p>I started going through some recent relevant SOTA papers for ideas, here is my summary so far: (</p>\n<hr>\n<p><strong>End-to-End Semi-Supervised Object Detection with Soft Teacher | 2021</strong></p>\n<ul>\n<li>Models: Soft Teacher + Swin-L (HTC++, multi-scale)</li>\n<li>propose a soft teacher mechanism where the classification loss of each unlabeled bounding box is weighed by the classification score produced by the teacher network</li>\n<li>As well as a box jittering approach to select reliable pseudo boxes for the learning of box regression</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2106.09018v3.pdf\" target=\"_blank\">https://arxiv.org/pdf/2106.09018v3.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>CBNetV2: A Composite Backbone Network Architecture for Object Detection | 2021</strong></p>\n<ul>\n<li>Models: Dual-Swin-L (HTC, multi-scale)</li>\n<li>Try out many existing open-sourced pre-trained backbones.</li>\n<li>Can be used as a nice review.<br>\nread more</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2107.00420v6.pdf\" target=\"_blank\">https://arxiv.org/pdf/2107.00420v6.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>Focal Self-attention for Local-Global Interactions in Vision Transformers | 2021</strong></p>\n<ul>\n<li>Models: Focal-L (HTC++, multi-scale)</li>\n<li>Present focal self-attention for vision transformers, each token attends the closest surrounding tokens at fine granularity but the tokens far away at coarse granularity</li>\n<li>Propose a variant of Vision Transformer models, substantial improvements over the current state-of-the-art Swin Transformers for 6 different object detection methods.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2107.00641v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/2107.00641v1.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>Swin Transformer: Hierarchical Vision Transformer using Shifted Windows | 2021</strong></p>\n<ul>\n<li>Models: Swin-L (HTC++, multi scale)</li>\n<li>Abstract: Hierarchical Transformer whose representation is computed with shifted windows. The shifted windowing scheme brings greater efficiency by limiting self-attention computation to non-overlapping local windows while also allowing for cross-window connection.</li>\n<li>Spinoffs:</li>\n<li>Used as a backbone for Self-Supervised Learning: Transformer-SSL</li>\n<li>Add Swin MLP, which is an adaption of Swin Transformer by replacing all multi-head self-attention (MHSA) blocks by MLP layers (more precisely it is a group linear layer). The shifted window configuration can also significantly improve the performance of vanilla MLP architectures.</li>\n<li>Soft Teacher: An end-to-end semi-supervisd object detection method, achieving a new record on the COCO test-dev.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2103.14030v2.pdf\" target=\"_blank\">https://arxiv.org/pdf/2103.14030v2.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>Instances as Queries | 2021</strong></p>\n<ul>\n<li>Models: QueryInst (single scale)</li>\n<li>This paper proposed a simple and effective query based instance segmentation method driven by parallel supervision on dynamic mask heads, which outperforms previous arts in terms of both accuracy and speed.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2105.01928v3.pdf\" target=\"_blank\">https://arxiv.org/pdf/2105.01928v3.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation | 2020</strong></p>\n<ul>\n<li>Models: Cascade Eff-B7 NAS-FPN (1280, self-training Copy Paste, single-scale)</li>\n<li>Show that the simple mechanism of pasting objects randomly is good enough and can provide solid gains on top of strong baselines.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2012.07177v2.pdf\" target=\"_blank\">https://arxiv.org/pdf/2012.07177v2.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>DetectoRS: Detecting Objects with Recursive Feature Pyramid and Switchable Atrous Convolution | 2020</strong></p>\n<ul>\n<li>Models: DetectoRS (ResNeXt-101-64x4d, multi-scale)</li>\n<li>Abstract: Propose Recursive Feature Pyramid, to incorporate extra feedback connections from Feature Pyramid Networks into the bottom-up backbone layers and Switchable Atrous Convolutions on the micro-level.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2006.02334v2.pdf\" target=\"_blank\">https://arxiv.org/pdf/2006.02334v2.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>SOLQ: Segmenting Objects by Learning Queries | 2021</strong></p>\n<ul>\n<li>Models: SOLQ (Swin-L, single scale)</li>\n<li>Propose an end-to-end framework for instance segmentation. Based on DETR, the sifference is that they segment objects by learning unified queries.</li>\n<li>During training phase, the mask vectors encoded are supervised by the compression coding of raw spatial masks</li>\n<li>Shot that the joint learning of unified query representation can greatly improve the detection performance of DETR.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2106.02351v3.pdf\" target=\"_blank\">https://arxiv.org/pdf/2106.02351v3.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization | 2019</strong></p>\n<ul>\n<li>Models: Mask R-CNN (SpineNet-190, 1536x1536)</li>\n<li>Propose SpineNet, a backbone with scale-permuted intermediate features and cross-scale connections that is learned on an object detection task by Neural Architecture Search.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/1912.05027v3.pdf\" target=\"_blank\">https://arxiv.org/pdf/1912.05027v3.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>Global Context Networks | 2020</strong></p>\n<ul>\n<li>Models: GCNet (ResNeXt-101 + DCN + cascade + GC r4)</li>\n<li>Replace the one-layer transformation function of the non-local block of NLNet by a two-layer bottleneck, which reduces the parameter number considerably.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2012.13375v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/2012.13375v1.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>CBNet: A Novel Composite Backbone Network Architecture for Object Detection | 2019</strong></p>\n<ul>\n<li>Models: Cascade Mask R-CNN (ResNeXt152, CBNet)</li>\n<li>Propose a novel strategy for assembling multiple identical backbones by composite connections between the adjacent backbones, to form a more powerful backbone named Composite Backbone Network (CBNet).</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/1909.03625v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/1909.03625v1.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>ResNeSt: Split-Attention Networks | 2020</strong></p>\n<ul>\n<li>Models: ResNeSt101</li>\n<li>Present a modularized architecture, which applies the channel-wise attention on different network branches to leverage their success in capturing cross-feature interactions and learning diverse representations. Our design results in a simple and unified computation block, which can be parameterized using only a few variables.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2004.08955v2.pdf\" target=\"_blank\">https://arxiv.org/pdf/2004.08955v2.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>Path Aggregation Network for Instance Segmentation | 2018</strong></p>\n<ul>\n<li>Models: PANet</li>\n<li>Presents adaptive feature pooling, which links feature grid and all feature levels to make useful information in each feature level propagate directly to following proposal subnetworks.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/1803.01534v4.pdf\" target=\"_blank\">https://arxiv.org/pdf/1803.01534v4.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>CenterMask : Real-Time Anchor-Free Instance Segmentation | 2019</strong></p>\n<ul>\n<li>Models: CenterMask + VoVNet99</li>\n<li>Adds a spatial attention-guided mask branch to anchor-free one stage object detector in the same vein with Mask R-CNN.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/1911.06667v6.pdf\" target=\"_blank\">https://arxiv.org/pdf/1911.06667v6.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>SOLOv2: Dynamic and Fast Instance Segmentation | 2020</strong></p>\n<ul>\n<li>Models: SOLOv2(Res-DCN-101-FPN)</li>\n<li>Dynamically learning the mask head of the object segmenter such that the mask head is conditioned on the location. The mask branch is decoupled into a mask kernel branch and mask feature branch, which are responsible for learning the convolution kernel and the convolved features respectively.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2003.10152v3.pdf\" target=\"_blank\">https://arxiv.org/pdf/2003.10152v3.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>Deep Occlusion-Aware Instance Segmentation with Overlapping BiLayers | 2021</strong></p>\n<ul>\n<li>Models: BCNet(ResNeXt-101 + FPN+ FCOS)</li>\n<li>Model image formation as composition of two overlapping layers, and propose Bilayer Convolutional Network (BCNet), where the top GCN layer detects the occluding objects (occluder) and the bottom GCN layer infers partially occluded instance. In general: Bilayer = Conv -&gt; GCN -&gt; FNC</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2103.12340v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/2103.12340v1.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>BlendMask: Top-Down Meets Bottom-Up for Instance Segmentation | 2020</strong></p>\n<ul>\n<li>Models: BlendMask (ResNet-101 + DCN interval=3)</li>\n<li>Improved mask prediction by effectively combining instance-level information with semantic information with lower-level fine-granularity The main contribution is a blender module which draws inspiration from both top-down and bottom-up instance segmentation approaches.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2001.00309v3.pdf\" target=\"_blank\">https://arxiv.org/pdf/2001.00309v3.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>K-Net: Towards Unified Image Segmentation | 2021</strong></p>\n<ul>\n<li>Models: K-Net-N256 (ResNet-101)</li>\n<li>Propose a kernel update strategy that enables each kernel dynamic and conditional on its meaningful group in the input image.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2106.14855v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/2106.14855v1.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>D2Det: Towards High Quality Object Detection and Instance Segmentation | 2020</strong></p>\n<ul>\n<li>Models: D2Det (ResNet-101, single-scale test)</li>\n<li>For precise localization, they introduced a dense local regression that predicts multiple dense box offsets for an object proposal.</li>\n<li>The dense local regression is not limited to a quantized set of keypoints within a fixed region and has the ability to regress position-sensitive real number dense</li>\n<li>Thedense local regression is further improved by a binary overlap prediction strategy that reduces the influence of background region on the final box regression.</li>\n<li>For accurate classification, they introduced a discriminative RoI pooling scheme that samples from various sub-regions of a proposal and performs adaptive weighting to obtain discriminative features.</li>\n<li>Link: <a href=\"http://openaccess.thecvf.com/content_CVPR_2020/papers/Cao_D2Det_Towards_High_Quality_Object_Detection_and_Instance_Segmentation_CVPR_2020_paper.pdf\" target=\"_blank\">http://openaccess.thecvf.com/content_CVPR_2020/papers/Cao_D2Det_Towards_High_Quality_Object_Detection_and_Instance_Segmentation_CVPR_2020_paper.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>ISTR: End-to-End Instance Segmentation with Transformers | 2021</strong></p>\n<ul>\n<li>Models: ISTR (ResNet101-FPN-3x, single-scale)</li>\n<li>Proposed an instance segmentation Transformer, it predicts low-dimensional mask embeddings, and matches them with ground truth mask embeddings for the set loss.</li>\n<li>Also ir concurrently conducts detection and segmentation with a recurrent refinement strategy, which provides a new way to achieve instance segmentation compared to the existing top-down and bottom-up frameworks.</li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2105.00637v2.pdf\" target=\"_blank\">https://arxiv.org/pdf/2105.00637v2.pdf</a></li>\n</ul>\n<hr>\n<hr>\n<p><strong>SipMask: Spatial Information Preservation for Fast Image and Video Instance Segmentation | 2020</strong></p>\n<ul>\n<li>Models: SipMask (ResNet-101, single-scale test)</li>\n<li>Proposed a fast single-stage instance segmentation method that preserves instance-specific spatial information by separating mask prediction of an instance to different sub-regions of a detected bounding-box.</li>\n<li>A novel modeule: light-weight spatial preservation (SP) module that generates a separate set of spatial coefficients for each sub-region within a bounding-box, leading to improved mask predictions. </li>\n<li>Link: <a href=\"https://arxiv.org/pdf/2007.14772v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/2007.14772v1.pdf</a></li>\n</ul>\n<hr>\n<p>'</p>\n<p>Happy Kaggling! </p>",
      "rawMarkdown": "'\nHey Guys, \n\nI started going through some recent relevant SOTA papers for ideas, here is my summary so far: (\n\n____\n**End-to-End Semi-Supervised Object Detection with Soft Teacher | 2021**\n- Models: Soft Teacher + Swin-L (HTC++, multi-scale)\n- propose a soft teacher mechanism where the classification loss of each unlabeled bounding box is weighed by the classification score produced by the teacher network\n- As well as a box jittering approach to select reliable pseudo boxes for the learning of box regression\n- Link: https://arxiv.org/pdf/2106.09018v3.pdf\n____\n____\n**CBNetV2: A Composite Backbone Network Architecture for Object Detection | 2021**\n- Models: Dual-Swin-L (HTC, multi-scale)\n- Try out many existing open-sourced pre-trained backbones.\n- Can be used as a nice review.\nread more\n- Link: https://arxiv.org/pdf/2107.00420v6.pdf\n____\n____\n**Focal Self-attention for Local-Global Interactions in Vision Transformers | 2021**\n- Models: Focal-L (HTC++, multi-scale)\n- Present focal self-attention for vision transformers, each token attends the closest surrounding tokens at fine granularity but the tokens far away at coarse granularity\n- Propose a variant of Vision Transformer models, substantial improvements over the current state-of-the-art Swin Transformers for 6 different object detection methods.\n- Link: https://arxiv.org/pdf/2107.00641v1.pdf\n____\n____\n**Swin Transformer: Hierarchical Vision Transformer using Shifted Windows | 2021**\n- Models: Swin-L (HTC++, multi scale)\n- Abstract: Hierarchical Transformer whose representation is computed with shifted windows. The shifted windowing scheme brings greater efficiency by limiting self-attention computation to non-overlapping local windows while also allowing for cross-window connection.\n- Spinoffs:\n- Used as a backbone for Self-Supervised Learning: Transformer-SSL\n- Add Swin MLP, which is an adaption of Swin Transformer by replacing all multi-head self-attention (MHSA) blocks by MLP layers (more precisely it is a group linear layer). The shifted window configuration can also significantly improve the performance of vanilla MLP architectures.\n- Soft Teacher: An end-to-end semi-supervisd object detection method, achieving a new record on the COCO test-dev.\n- Link: https://arxiv.org/pdf/2103.14030v2.pdf\n____\n____\n**Instances as Queries | 2021**\n- Models: QueryInst (single scale)\n- This paper proposed a simple and effective query based instance segmentation method driven by parallel supervision on dynamic mask heads, which outperforms previous arts in terms of both accuracy and speed.\n- Link: https://arxiv.org/pdf/2105.01928v3.pdf\n____\n____\n**Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation | 2020**\n- Models: Cascade Eff-B7 NAS-FPN (1280, self-training Copy Paste, single-scale)\n- Show that the simple mechanism of pasting objects randomly is good enough and can provide solid gains on top of strong baselines.\n- Link: https://arxiv.org/pdf/2012.07177v2.pdf\n____\n____\n**DetectoRS: Detecting Objects with Recursive Feature Pyramid and Switchable Atrous Convolution | 2020**\n- Models: DetectoRS (ResNeXt-101-64x4d, multi-scale)\n- Abstract: Propose Recursive Feature Pyramid, to incorporate extra feedback connections from Feature Pyramid Networks into the bottom-up backbone layers and Switchable Atrous Convolutions on the micro-level.\n- Link: https://arxiv.org/pdf/2006.02334v2.pdf\n____\n\n____\n**SOLQ: Segmenting Objects by Learning Queries | 2021**\n- Models: SOLQ (Swin-L, single scale)\n- Propose an end-to-end framework for instance segmentation. Based on DETR, the sifference is that they segment objects by learning unified queries.\n- During training phase, the mask vectors encoded are supervised by the compression coding of raw spatial masks\n- Shot that the joint learning of unified query representation can greatly improve the detection performance of DETR.\n- Link: https://arxiv.org/pdf/2106.02351v3.pdf\n____\n____\n**SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization | 2019**\n- Models: Mask R-CNN (SpineNet-190, 1536x1536)\n- Propose SpineNet, a backbone with scale-permuted intermediate features and cross-scale connections that is learned on an object detection task by Neural Architecture Search.\n- Link: https://arxiv.org/pdf/1912.05027v3.pdf\n____\n____\n**Global Context Networks | 2020**\n- Models: GCNet (ResNeXt-101 + DCN + cascade + GC r4)\n- Replace the one-layer transformation function of the non-local block of NLNet by a two-layer bottleneck, which reduces the parameter number considerably.\n- Link: https://arxiv.org/pdf/2012.13375v1.pdf\n____\n____\n**CBNet: A Novel Composite Backbone Network Architecture for Object Detection | 2019**\n- Models: Cascade Mask R-CNN (ResNeXt152, CBNet)\n- Propose a novel strategy for assembling multiple identical backbones by composite connections between the adjacent backbones, to form a more powerful backbone named Composite Backbone Network (CBNet).\n- Link: https://arxiv.org/pdf/1909.03625v1.pdf\n____\n____\n**ResNeSt: Split-Attention Networks | 2020**\n- Models: ResNeSt101\n- Present a modularized architecture, which applies the channel-wise attention on different network branches to leverage their success in capturing cross-feature interactions and learning diverse representations. Our design results in a simple and unified computation block, which can be parameterized using only a few variables.\n- Link: https://arxiv.org/pdf/2004.08955v2.pdf\n____\n____\n**Path Aggregation Network for Instance Segmentation | 2018**\n- Models: PANet\n- Presents adaptive feature pooling, which links feature grid and all feature levels to make useful information in each feature level propagate directly to following proposal subnetworks.\n- Link: https://arxiv.org/pdf/1803.01534v4.pdf\n____\n____\n**CenterMask : Real-Time Anchor-Free Instance Segmentation | 2019**\n- Models: CenterMask + VoVNet99\n- Adds a spatial attention-guided mask branch to anchor-free one stage object detector in the same vein with Mask R-CNN.\n- Link: https://arxiv.org/pdf/1911.06667v6.pdf\n____\n____\n**SOLOv2: Dynamic and Fast Instance Segmentation | 2020**\n- Models: SOLOv2(Res-DCN-101-FPN)\n- Dynamically learning the mask head of the object segmenter such that the mask head is conditioned on the location. The mask branch is decoupled into a mask kernel branch and mask feature branch, which are responsible for learning the convolution kernel and the convolved features respectively.\n- Link: https://arxiv.org/pdf/2003.10152v3.pdf\n____\n____\n**Deep Occlusion-Aware Instance Segmentation with Overlapping BiLayers | 2021**\n- Models: BCNet(ResNeXt-101 + FPN+ FCOS)\n- Model image formation as composition of two overlapping layers, and propose Bilayer Convolutional Network (BCNet), where the top GCN layer detects the occluding objects (occluder) and the bottom GCN layer infers partially occluded instance. In general: Bilayer = Conv -> GCN -> FNC\n- Link: https://arxiv.org/pdf/2103.12340v1.pdf\n____\n____\n**BlendMask: Top-Down Meets Bottom-Up for Instance Segmentation | 2020**\n- Models: BlendMask (ResNet-101 + DCN interval=3)\n- Improved mask prediction by effectively combining instance-level information with semantic information with lower-level fine-granularity The main contribution is a blender module which draws inspiration from both top-down and bottom-up instance segmentation approaches.\n- Link: https://arxiv.org/pdf/2001.00309v3.pdf\n____\n____\n**K-Net: Towards Unified Image Segmentation | 2021**\n- Models: K-Net-N256 (ResNet-101)\n- Propose a kernel update strategy that enables each kernel dynamic and conditional on its meaningful group in the input image.\n- Link: https://arxiv.org/pdf/2106.14855v1.pdf\n____\n____\n**D2Det: Towards High Quality Object Detection and Instance Segmentation | 2020**\n- Models: D2Det (ResNet-101, single-scale test)\n- For precise localization, they introduced a dense local regression that predicts multiple dense box offsets for an object proposal.\n- The dense local regression is not limited to a quantized set of keypoints within a fixed region and has the ability to regress position-sensitive real number dense\n- Thedense local regression is further improved by a binary overlap prediction strategy that reduces the influence of background region on the final box regression.\n- For accurate classification, they introduced a discriminative RoI pooling scheme that samples from various sub-regions of a proposal and performs adaptive weighting to obtain discriminative features.\n- Link: http://openaccess.thecvf.com/content_CVPR_2020/papers/Cao_D2Det_Towards_High_Quality_Object_Detection_and_Instance_Segmentation_CVPR_2020_paper.pdf\n____\n____\n**ISTR: End-to-End Instance Segmentation with Transformers | 2021**\n- Models: ISTR (ResNet101-FPN-3x, single-scale)\n- Proposed an instance segmentation Transformer, it predicts low-dimensional mask embeddings, and matches them with ground truth mask embeddings for the set loss.\n- Also ir concurrently conducts detection and segmentation with a recurrent refinement strategy, which provides a new way to achieve instance segmentation compared to the existing top-down and bottom-up frameworks.\n- Link: https://arxiv.org/pdf/2105.00637v2.pdf\n____\n____\n**SipMask: Spatial Information Preservation for Fast Image and Video Instance Segmentation | 2020**\n- Models: SipMask (ResNet-101, single-scale test)\n- Proposed a fast single-stage instance segmentation method that preserves instance-specific spatial information by separating mask prediction of an instance to different sub-regions of a detected bounding-box.\n- A novel modeule: light-weight spatial preservation (SP) module that generates a separate set of spatial coefficients for each sub-region within a bounding-box, leading to improved mask predictions. \n- Link: https://arxiv.org/pdf/2007.14772v1.pdf\n____\n'\n\nHappy Kaggling!",
      "votes": null
    },
    {
      "id": "1565674",
      "postDate": "10/30/2021 20:34:09",
      "content": "<p>Nicely done! Thank you!</p>",
      "rawMarkdown": "Nicely done! Thank you!",
      "votes": null
    },
    {
      "id": "1568037",
      "postDate": "11/02/2021 12:50:52",
      "content": "<p>Thanks for the summary!</p>",
      "rawMarkdown": "Thanks for the summary!",
      "votes": null
    },
    {
      "id": "1569137",
      "postDate": "11/03/2021 09:17:42",
      "content": "<p>Thanks, that helps a lot !</p>",
      "rawMarkdown": "Thanks, that helps a lot !",
      "votes": null
    },
    {
      "id": "1584957",
      "postDate": "11/16/2021 22:59:46",
      "content": "<p>Thank you very much for your survey.<br>\nI have one minor comment.<br>\nThere is a paper published in 2018 in this list.<br>\nSo the title should be 'the past three years'.</p>",
      "rawMarkdown": "Thank you very much for your survey.\nI have one minor comment.\nThere is a paper published in 2018 in this list.\nSo the title should be 'the past three years'.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1565674,
      "author_name": "evancofsky",
      "author_url": "",
      "post_date": "10/30/2021 20:34:09",
      "content": "<p>Nicely done! Thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1568037,
      "author_name": "lawrenceroy",
      "author_url": "",
      "post_date": "11/02/2021 12:50:52",
      "content": "<p>Thanks for the summary!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1569137,
      "author_name": "barappear",
      "author_url": "",
      "post_date": "11/03/2021 09:17:42",
      "content": "<p>Thanks, that helps a lot !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1584957,
      "author_name": "osamurai",
      "author_url": "",
      "post_date": "11/16/2021 22:59:46",
      "content": "<p>Thank you very much for your survey.<br>\nI have one minor comment.<br>\nThere is a paper published in 2018 in this list.<br>\nSo the title should be 'the past three years'.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1564940": "'\nHey Guys, \n\nI started going through some recent relevant SOTA papers for ideas, here is my summary so far: (\n\n____\n**End-to-End Semi-Supervised Object Detection with Soft Teacher | 2021**\n- Models: Soft Teacher + Swin-L (HTC++, multi-scale)\n- propose a soft teacher mechanism where the classification loss of each unlabeled bounding box is weighed by the classification score produced by the teacher network\n- As well as a box jittering approach to select reliable pseudo boxes for the learning of box regression\n- Link: https://arxiv.org/pdf/2106.09018v3.pdf\n____\n____\n**CBNetV2: A Composite Backbone Network Architecture for Object Detection | 2021**\n- Models: Dual-Swin-L (HTC, multi-scale)\n- Try out many existing open-sourced pre-trained backbones.\n- Can be used as a nice review.\nread more\n- Link: https://arxiv.org/pdf/2107.00420v6.pdf\n____\n____\n**Focal Self-attention for Local-Global Interactions in Vision Transformers | 2021**\n- Models: Focal-L (HTC++, multi-scale)\n- Present focal self-attention for vision transformers, each token attends the closest surrounding tokens at fine granularity but the tokens far away at coarse granularity\n- Propose a variant of Vision Transformer models, substantial improvements over the current state-of-the-art Swin Transformers for 6 different object detection methods.\n- Link: https://arxiv.org/pdf/2107.00641v1.pdf\n____\n____\n**Swin Transformer: Hierarchical Vision Transformer using Shifted Windows | 2021**\n- Models: Swin-L (HTC++, multi scale)\n- Abstract: Hierarchical Transformer whose representation is computed with shifted windows. The shifted windowing scheme brings greater efficiency by limiting self-attention computation to non-overlapping local windows while also allowing for cross-window connection.\n- Spinoffs:\n- Used as a backbone for Self-Supervised Learning: Transformer-SSL\n- Add Swin MLP, which is an adaption of Swin Transformer by replacing all multi-head self-attention (MHSA) blocks by MLP layers (more precisely it is a group linear layer). The shifted window configuration can also significantly improve the performance of vanilla MLP architectures.\n- Soft Teacher: An end-to-end semi-supervisd object detection method, achieving a new record on the COCO test-dev.\n- Link: https://arxiv.org/pdf/2103.14030v2.pdf\n____\n____\n**Instances as Queries | 2021**\n- Models: QueryInst (single scale)\n- This paper proposed a simple and effective query based instance segmentation method driven by parallel supervision on dynamic mask heads, which outperforms previous arts in terms of both accuracy and speed.\n- Link: https://arxiv.org/pdf/2105.01928v3.pdf\n____\n____\n**Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation | 2020**\n- Models: Cascade Eff-B7 NAS-FPN (1280, self-training Copy Paste, single-scale)\n- Show that the simple mechanism of pasting objects randomly is good enough and can provide solid gains on top of strong baselines.\n- Link: https://arxiv.org/pdf/2012.07177v2.pdf\n____\n____\n**DetectoRS: Detecting Objects with Recursive Feature Pyramid and Switchable Atrous Convolution | 2020**\n- Models: DetectoRS (ResNeXt-101-64x4d, multi-scale)\n- Abstract: Propose Recursive Feature Pyramid, to incorporate extra feedback connections from Feature Pyramid Networks into the bottom-up backbone layers and Switchable Atrous Convolutions on the micro-level.\n- Link: https://arxiv.org/pdf/2006.02334v2.pdf\n____\n\n____\n**SOLQ: Segmenting Objects by Learning Queries | 2021**\n- Models: SOLQ (Swin-L, single scale)\n- Propose an end-to-end framework for instance segmentation. Based on DETR, the sifference is that they segment objects by learning unified queries.\n- During training phase, the mask vectors encoded are supervised by the compression coding of raw spatial masks\n- Shot that the joint learning of unified query representation can greatly improve the detection performance of DETR.\n- Link: https://arxiv.org/pdf/2106.02351v3.pdf\n____\n____\n**SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization | 2019**\n- Models: Mask R-CNN (SpineNet-190, 1536x1536)\n- Propose SpineNet, a backbone with scale-permuted intermediate features and cross-scale connections that is learned on an object detection task by Neural Architecture Search.\n- Link: https://arxiv.org/pdf/1912.05027v3.pdf\n____\n____\n**Global Context Networks | 2020**\n- Models: GCNet (ResNeXt-101 + DCN + cascade + GC r4)\n- Replace the one-layer transformation function of the non-local block of NLNet by a two-layer bottleneck, which reduces the parameter number considerably.\n- Link: https://arxiv.org/pdf/2012.13375v1.pdf\n____\n____\n**CBNet: A Novel Composite Backbone Network Architecture for Object Detection | 2019**\n- Models: Cascade Mask R-CNN (ResNeXt152, CBNet)\n- Propose a novel strategy for assembling multiple identical backbones by composite connections between the adjacent backbones, to form a more powerful backbone named Composite Backbone Network (CBNet).\n- Link: https://arxiv.org/pdf/1909.03625v1.pdf\n____\n____\n**ResNeSt: Split-Attention Networks | 2020**\n- Models: ResNeSt101\n- Present a modularized architecture, which applies the channel-wise attention on different network branches to leverage their success in capturing cross-feature interactions and learning diverse representations. Our design results in a simple and unified computation block, which can be parameterized using only a few variables.\n- Link: https://arxiv.org/pdf/2004.08955v2.pdf\n____\n____\n**Path Aggregation Network for Instance Segmentation | 2018**\n- Models: PANet\n- Presents adaptive feature pooling, which links feature grid and all feature levels to make useful information in each feature level propagate directly to following proposal subnetworks.\n- Link: https://arxiv.org/pdf/1803.01534v4.pdf\n____\n____\n**CenterMask : Real-Time Anchor-Free Instance Segmentation | 2019**\n- Models: CenterMask + VoVNet99\n- Adds a spatial attention-guided mask branch to anchor-free one stage object detector in the same vein with Mask R-CNN.\n- Link: https://arxiv.org/pdf/1911.06667v6.pdf\n____\n____\n**SOLOv2: Dynamic and Fast Instance Segmentation | 2020**\n- Models: SOLOv2(Res-DCN-101-FPN)\n- Dynamically learning the mask head of the object segmenter such that the mask head is conditioned on the location. The mask branch is decoupled into a mask kernel branch and mask feature branch, which are responsible for learning the convolution kernel and the convolved features respectively.\n- Link: https://arxiv.org/pdf/2003.10152v3.pdf\n____\n____\n**Deep Occlusion-Aware Instance Segmentation with Overlapping BiLayers | 2021**\n- Models: BCNet(ResNeXt-101 + FPN+ FCOS)\n- Model image formation as composition of two overlapping layers, and propose Bilayer Convolutional Network (BCNet), where the top GCN layer detects the occluding objects (occluder) and the bottom GCN layer infers partially occluded instance. In general: Bilayer = Conv -> GCN -> FNC\n- Link: https://arxiv.org/pdf/2103.12340v1.pdf\n____\n____\n**BlendMask: Top-Down Meets Bottom-Up for Instance Segmentation | 2020**\n- Models: BlendMask (ResNet-101 + DCN interval=3)\n- Improved mask prediction by effectively combining instance-level information with semantic information with lower-level fine-granularity The main contribution is a blender module which draws inspiration from both top-down and bottom-up instance segmentation approaches.\n- Link: https://arxiv.org/pdf/2001.00309v3.pdf\n____\n____\n**K-Net: Towards Unified Image Segmentation | 2021**\n- Models: K-Net-N256 (ResNet-101)\n- Propose a kernel update strategy that enables each kernel dynamic and conditional on its meaningful group in the input image.\n- Link: https://arxiv.org/pdf/2106.14855v1.pdf\n____\n____\n**D2Det: Towards High Quality Object Detection and Instance Segmentation | 2020**\n- Models: D2Det (ResNet-101, single-scale test)\n- For precise localization, they introduced a dense local regression that predicts multiple dense box offsets for an object proposal.\n- The dense local regression is not limited to a quantized set of keypoints within a fixed region and has the ability to regress position-sensitive real number dense\n- Thedense local regression is further improved by a binary overlap prediction strategy that reduces the influence of background region on the final box regression.\n- For accurate classification, they introduced a discriminative RoI pooling scheme that samples from various sub-regions of a proposal and performs adaptive weighting to obtain discriminative features.\n- Link: http://openaccess.thecvf.com/content_CVPR_2020/papers/Cao_D2Det_Towards_High_Quality_Object_Detection_and_Instance_Segmentation_CVPR_2020_paper.pdf\n____\n____\n**ISTR: End-to-End Instance Segmentation with Transformers | 2021**\n- Models: ISTR (ResNet101-FPN-3x, single-scale)\n- Proposed an instance segmentation Transformer, it predicts low-dimensional mask embeddings, and matches them with ground truth mask embeddings for the set loss.\n- Also ir concurrently conducts detection and segmentation with a recurrent refinement strategy, which provides a new way to achieve instance segmentation compared to the existing top-down and bottom-up frameworks.\n- Link: https://arxiv.org/pdf/2105.00637v2.pdf\n____\n____\n**SipMask: Spatial Information Preservation for Fast Image and Video Instance Segmentation | 2020**\n- Models: SipMask (ResNet-101, single-scale test)\n- Proposed a fast single-stage instance segmentation method that preserves instance-specific spatial information by separating mask prediction of an instance to different sub-regions of a detected bounding-box.\n- A novel modeule: light-weight spatial preservation (SP) module that generates a separate set of spatial coefficients for each sub-region within a bounding-box, leading to improved mask predictions. \n- Link: https://arxiv.org/pdf/2007.14772v1.pdf\n____\n'\n\nHappy Kaggling!",
    "1565674": "Nicely done! Thank you!",
    "1568037": "Thanks for the summary!",
    "1569137": "Thanks, that helps a lot !",
    "1584957": "Thank you very much for your survey.\nI have one minor comment.\nThere is a paper published in 2018 in this list.\nSo the title should be 'the past three years'."
  },
  "source": "meta"
}