{
  "id": 293799,
  "title": "Mask2Former: Masked-attention Mask Transformer for Universal Image Segmentation",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/293799",
  "author_name": "",
  "post_date": "2021-12-07T04:21:15.994310200Z",
  "votes": 11,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Image segmentation is about grouping pixels with different semantics, e.g., category or instance membership, where each choice of semantics defines a task. While only the semantics of each task differ, current research focuses on designing specialized architectures for each task. We present Masked-attention Mask Transformer (Mask2Former), a new architecture capable of addressing any image segmentation task (panoptic, instance, or semantic). Its key components include masked attention, which extracts localized features by constraining cross-attention within predicted mask regions. In addition to reducing the research effort by at least three times, it outperforms the best-specialized architectures by a significant margin on four popular datasets. Most notably, Mask2Former sets a new state-of-the-art for panoptic segmentation (57.8 PQ on COCO), instance segmentation (50.1 AP on COCO), and semantic segmentation (57.7 mIoU on ADE20K).</p>\n<p>Paper: <a href=\"https://arxiv.org/pdf/2112.01527v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/2112.01527v1.pdf</a></p>\n<p>Github: <a href=\"https://github.com/facebookresearch/Mask2Former\" target=\"_blank\">https://github.com/facebookresearch/Mask2Former</a></p>",
  "messages": [
    {
      "id": "1610240",
      "postDate": "12/07/2021 04:21:15",
      "content": "<p>Image segmentation is about grouping pixels with different semantics, e.g., category or instance membership, where each choice of semantics defines a task. While only the semantics of each task differ, current research focuses on designing specialized architectures for each task. We present Masked-attention Mask Transformer (Mask2Former), a new architecture capable of addressing any image segmentation task (panoptic, instance, or semantic). Its key components include masked attention, which extracts localized features by constraining cross-attention within predicted mask regions. In addition to reducing the research effort by at least three times, it outperforms the best-specialized architectures by a significant margin on four popular datasets. Most notably, Mask2Former sets a new state-of-the-art for panoptic segmentation (57.8 PQ on COCO), instance segmentation (50.1 AP on COCO), and semantic segmentation (57.7 mIoU on ADE20K).</p>\n<p>Paper: <a href=\"https://arxiv.org/pdf/2112.01527v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/2112.01527v1.pdf</a></p>\n<p>Github: <a href=\"https://github.com/facebookresearch/Mask2Former\" target=\"_blank\">https://github.com/facebookresearch/Mask2Former</a></p>",
      "rawMarkdown": "Image segmentation is about grouping pixels with different semantics, e.g., category or instance membership, where each choice of semantics defines a task. While only the semantics of each task differ, current research focuses on designing specialized architectures for each task. We present Masked-attention Mask Transformer (Mask2Former), a new architecture capable of addressing any image segmentation task (panoptic, instance, or semantic). Its key components include masked attention, which extracts localized features by constraining cross-attention within predicted mask regions. In addition to reducing the research effort by at least three times, it outperforms the best-specialized architectures by a significant margin on four popular datasets. Most notably, Mask2Former sets a new state-of-the-art for panoptic segmentation (57.8 PQ on COCO), instance segmentation (50.1 AP on COCO), and semantic segmentation (57.7 mIoU on ADE20K).\n\nPaper: https://arxiv.org/pdf/2112.01527v1.pdf\n\nGithub: https://github.com/facebookresearch/Mask2Former",
      "votes": null
    },
    {
      "id": "1612781",
      "postDate": "12/09/2021 10:09:30",
      "content": "<p>Thank you for sharing. BTW how do you know so early? Do you register some newsletters?<br>\nI register <a href=\"https://www.deeplearningweekly.com/\" target=\"_blank\">Deep Learning Weekly</a>. I know mask2transformer today by this newsletter.</p>",
      "rawMarkdown": "Thank you for sharing. BTW how do you know so early? Do you register some newsletters?\nI register [Deep Learning Weekly](https://www.deeplearningweekly.com/). I know mask2transformer today by this newsletter.",
      "votes": null
    },
    {
      "id": "1612833",
      "postDate": "12/09/2021 11:13:00",
      "content": "<p>Sorry I found it.<br>\n<a href=\"https://www.kaggle.com/general/71954\" target=\"_blank\">https://www.kaggle.com/general/71954</a></p>",
      "rawMarkdown": "Sorry I found it.\nhttps://www.kaggle.com/general/71954",
      "votes": null
    },
    {
      "id": "1614049",
      "postDate": "12/10/2021 14:26:58",
      "content": "<p>I visit paperswithcode.com regularly</p>",
      "rawMarkdown": "I visit paperswithcode.com regularly",
      "votes": null
    },
    {
      "id": "1614063",
      "postDate": "12/10/2021 14:51:01",
      "content": "<p>Thank you very much!</p>",
      "rawMarkdown": "Thank you very much!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1612781,
      "author_name": "osamurai",
      "author_url": "",
      "post_date": "12/09/2021 10:09:30",
      "content": "<p>Thank you for sharing. BTW how do you know so early? Do you register some newsletters?<br>\nI register <a href=\"https://www.deeplearningweekly.com/\" target=\"_blank\">Deep Learning Weekly</a>. I know mask2transformer today by this newsletter.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1612833,
          "author_name": "osamurai",
          "author_url": "",
          "post_date": "12/09/2021 11:13:00",
          "content": "<p>Sorry I found it.<br>\n<a href=\"https://www.kaggle.com/general/71954\" target=\"_blank\">https://www.kaggle.com/general/71954</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1614049,
          "author_name": "projdev",
          "author_url": "",
          "post_date": "12/10/2021 14:26:58",
          "content": "<p>I visit paperswithcode.com regularly</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1614063,
          "author_name": "osamurai",
          "author_url": "",
          "post_date": "12/10/2021 14:51:01",
          "content": "<p>Thank you very much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1610240": "Image segmentation is about grouping pixels with different semantics, e.g., category or instance membership, where each choice of semantics defines a task. While only the semantics of each task differ, current research focuses on designing specialized architectures for each task. We present Masked-attention Mask Transformer (Mask2Former), a new architecture capable of addressing any image segmentation task (panoptic, instance, or semantic). Its key components include masked attention, which extracts localized features by constraining cross-attention within predicted mask regions. In addition to reducing the research effort by at least three times, it outperforms the best-specialized architectures by a significant margin on four popular datasets. Most notably, Mask2Former sets a new state-of-the-art for panoptic segmentation (57.8 PQ on COCO), instance segmentation (50.1 AP on COCO), and semantic segmentation (57.7 mIoU on ADE20K).\n\nPaper: https://arxiv.org/pdf/2112.01527v1.pdf\n\nGithub: https://github.com/facebookresearch/Mask2Former",
    "1612781": "Thank you for sharing. BTW how do you know so early? Do you register some newsletters?\nI register [Deep Learning Weekly](https://www.deeplearningweekly.com/). I know mask2transformer today by this newsletter.",
    "1612833": "Sorry I found it.\nhttps://www.kaggle.com/general/71954",
    "1614049": "I visit paperswithcode.com regularly",
    "1614063": "Thank you very much!"
  },
  "source": "meta"
}