{
  "id": 470498,
  "title": "segmentation_models_pytorch/mmsegmentation/pytorch-3dunet, what framework are you using?",
  "url": "/competitions/blood-vessel-segmentation/discussion/470498",
  "author_name": "",
  "post_date": "2024-01-24T13:17:50.261145100Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>segmentation_models_pytorch is used by top public nodebooks.<br>\nnextvit officially uses mmsegmentation.<br>\npytorch-3dunet is a 3d segmentation framework.</p>\n<p>I tried pytorch-3dunet with pretrained weights from the official page (with training pipeline from the public nodebook), but the model didn't converge well.<br>\nI haven't tried mmsegmentation because I found a way to use nextvit with smp (the result is slightly better than se_resnext50_32x4d).</p>",
  "messages": [
    {
      "id": "2617808",
      "postDate": "01/24/2024 13:17:50",
      "content": "<p>segmentation_models_pytorch is used by top public nodebooks.<br>\nnextvit officially uses mmsegmentation.<br>\npytorch-3dunet is a 3d segmentation framework.</p>\n<p>I tried pytorch-3dunet with pretrained weights from the official page (with training pipeline from the public nodebook), but the model didn't converge well.<br>\nI haven't tried mmsegmentation because I found a way to use nextvit with smp (the result is slightly better than se_resnext50_32x4d).</p>",
      "rawMarkdown": "segmentation_models_pytorch is used by top public nodebooks.\nnextvit officially uses mmsegmentation.\npytorch-3dunet is a 3d segmentation framework.\n\nI tried pytorch-3dunet with pretrained weights from the official page (with training pipeline from the public nodebook), but the model didn't converge well.\nI haven't tried mmsegmentation because I found a way to use nextvit with smp (the result is slightly better than se_resnext50_32x4d).",
      "votes": null
    },
    {
      "id": "2618284",
      "postDate": "01/24/2024 16:48:42",
      "content": "<p>You found a way to use this <a href=\"https://github.com/bytedance/Next-ViT/blob/main/segmentation/nextvit.py\" target=\"_blank\">https://github.com/bytedance/Next-ViT/blob/main/segmentation/nextvit.py</a> ?<br>\nWhat do you mean?</p>",
      "rawMarkdown": "You found a way to use this https://github.com/bytedance/Next-ViT/blob/main/segmentation/nextvit.py ?\nWhat do you mean?",
      "votes": null
    },
    {
      "id": "2618573",
      "postDate": "01/24/2024 20:05:44",
      "content": "<p>You can use any encoder you want using smp.</p>\n<p>smp has a class called EncoderMixin that it uses to connect your custom encoder with their decoder. You have to return the skip connection tensors. For example here is a custom 2.5D encoder with 2D output with corrected skip connections using their class.<br>\nIf you want to use NextVit, you can probably figure it out using this class. </p>\n<p>Here are more information about it: <a href=\"https://smp.readthedocs.io/en/latest/insights.html#creating-your-own-encoder\" target=\"_blank\">https://smp.readthedocs.io/en/latest/insights.html#creating-your-own-encoder</a></p>\n<pre><code> (torch.nn.Module, EncoderMixin):\n\n     ():\n        ().__init__()\n        self._out_channels = [, , , , , ]\n        self._depth:  = \n\n        self._in_channels:  = \n\n        self.cnn3d = nn.Conv3d(, , kernel_size=, padding=)\n\n        self.backbone = timm.create_model(, pretrained=, features_only=, in_chans=)\n\n     ():\n        features_x = self.cnn3d(x.unsqueeze()).squeeze()\n\n        feat2, feat3, feat4, feat5, feat6 = self.backbone(features_x)\n\n         [x[:,,:,:], feat2, feat3, feat4, feat5, feat6]\n\nsmp.encoders.encoders[] = {\n    : MiddleSkipConnectionEncoder,\n     : {},\n    : {},\n}\n\nmodel = smp.Unet(encoder_name=, encoder_weights= ,classes=)\n</code></pre>",
      "rawMarkdown": "You can use any encoder you want using smp.\n\nsmp has a class called EncoderMixin that it uses to connect your custom encoder with their decoder. You have to return the skip connection tensors. For example here is a custom 2.5D encoder with 2D output with corrected skip connections using their class.\nIf you want to use NextVit, you can probably figure it out using this class. \n\nHere are more information about it: https://smp.readthedocs.io/en/latest/insights.html#creating-your-own-encoder\n\n```\nclass MiddleSkipConnectionEncoder(torch.nn.Module, EncoderMixin):\n\n    def __init__(self, **kwargs):\n        super().__init__()\n        self._out_channels = [1, 64, 96, 192, 384, 768]\n        self._depth: int = 5\n\n        self._in_channels: int = 1\n\n        self.cnn3d = nn.Conv3d(3, 32, kernel_size=3, padding=1)\n\n        self.backbone = timm.create_model('maxvit_small_tf_384', pretrained=False, features_only=True, in_chans=32)\n\n    def forward(self, x: torch.Tensor):\n        features_x = self.cnn3d(x.unsqueeze(2)).squeeze(2)\n\n        feat2, feat3, feat4, feat5, feat6 = self.backbone(features_x)\n\n        return [x[:,1,:,:], feat2, feat3, feat4, feat5, feat6]\n    \nsmp.encoders.encoders[\"225d_encoder\"] = {\n    \"encoder\": MiddleSkipConnectionEncoder,\n    'params' : {},\n    'pretrained_settings': {},\n}\n\nmodel = smp.Unet(encoder_name='225d_encoder', encoder_weights=None ,classes=1)\n```",
      "votes": null
    },
    {
      "id": "2618627",
      "postDate": "01/24/2024 21:25:56",
      "content": "<p>Thanks. But I mean, there is already a segmentation NextViT code. Would'nt be just copy and paste?<br>\nAnd there is several ckp with pretrained weights at the repo.</p>",
      "rawMarkdown": "Thanks. But I mean, there is already a segmentation NextViT code. Would'nt be just copy and paste?\nAnd there is several ckp with pretrained weights at the repo.",
      "votes": null
    },
    {
      "id": "2618977",
      "postDate": "01/25/2024 06:20:18",
      "content": "<p>Hi LIU, how much video memory and inference time do you need using nextvit based unet?</p>",
      "rawMarkdown": "Hi LIU, how much video memory and inference time do you need using nextvit based unet?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2618284,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "01/24/2024 16:48:42",
      "content": "<p>You found a way to use this <a href=\"https://github.com/bytedance/Next-ViT/blob/main/segmentation/nextvit.py\" target=\"_blank\">https://github.com/bytedance/Next-ViT/blob/main/segmentation/nextvit.py</a> ?<br>\nWhat do you mean?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2618573,
          "author_name": "janmpia",
          "author_url": "",
          "post_date": "01/24/2024 20:05:44",
          "content": "<p>You can use any encoder you want using smp.</p>\n<p>smp has a class called EncoderMixin that it uses to connect your custom encoder with their decoder. You have to return the skip connection tensors. For example here is a custom 2.5D encoder with 2D output with corrected skip connections using their class.<br>\nIf you want to use NextVit, you can probably figure it out using this class. </p>\n<p>Here are more information about it: <a href=\"https://smp.readthedocs.io/en/latest/insights.html#creating-your-own-encoder\" target=\"_blank\">https://smp.readthedocs.io/en/latest/insights.html#creating-your-own-encoder</a></p>\n<pre><code> (torch.nn.Module, EncoderMixin):\n\n     ():\n        ().__init__()\n        self._out_channels = [, , , , , ]\n        self._depth:  = \n\n        self._in_channels:  = \n\n        self.cnn3d = nn.Conv3d(, , kernel_size=, padding=)\n\n        self.backbone = timm.create_model(, pretrained=, features_only=, in_chans=)\n\n     ():\n        features_x = self.cnn3d(x.unsqueeze()).squeeze()\n\n        feat2, feat3, feat4, feat5, feat6 = self.backbone(features_x)\n\n         [x[:,,:,:], feat2, feat3, feat4, feat5, feat6]\n\nsmp.encoders.encoders[] = {\n    : MiddleSkipConnectionEncoder,\n     : {},\n    : {},\n}\n\nmodel = smp.Unet(encoder_name=, encoder_weights= ,classes=)\n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 2618627,
              "author_name": "sacuscreed",
              "author_url": "",
              "post_date": "01/24/2024 21:25:56",
              "content": "<p>Thanks. But I mean, there is already a segmentation NextViT code. Would'nt be just copy and paste?<br>\nAnd there is several ckp with pretrained weights at the repo.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2618977,
      "author_name": "tanxxx",
      "author_url": "",
      "post_date": "01/25/2024 06:20:18",
      "content": "<p>Hi LIU, how much video memory and inference time do you need using nextvit based unet?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2617808": "segmentation_models_pytorch is used by top public nodebooks.\nnextvit officially uses mmsegmentation.\npytorch-3dunet is a 3d segmentation framework.\n\nI tried pytorch-3dunet with pretrained weights from the official page (with training pipeline from the public nodebook), but the model didn't converge well.\nI haven't tried mmsegmentation because I found a way to use nextvit with smp (the result is slightly better than se_resnext50_32x4d).",
    "2618284": "You found a way to use this https://github.com/bytedance/Next-ViT/blob/main/segmentation/nextvit.py ?\nWhat do you mean?",
    "2618573": "You can use any encoder you want using smp.\n\nsmp has a class called EncoderMixin that it uses to connect your custom encoder with their decoder. You have to return the skip connection tensors. For example here is a custom 2.5D encoder with 2D output with corrected skip connections using their class.\nIf you want to use NextVit, you can probably figure it out using this class. \n\nHere are more information about it: https://smp.readthedocs.io/en/latest/insights.html#creating-your-own-encoder\n\n```\nclass MiddleSkipConnectionEncoder(torch.nn.Module, EncoderMixin):\n\n    def __init__(self, **kwargs):\n        super().__init__()\n        self._out_channels = [1, 64, 96, 192, 384, 768]\n        self._depth: int = 5\n\n        self._in_channels: int = 1\n\n        self.cnn3d = nn.Conv3d(3, 32, kernel_size=3, padding=1)\n\n        self.backbone = timm.create_model('maxvit_small_tf_384', pretrained=False, features_only=True, in_chans=32)\n\n    def forward(self, x: torch.Tensor):\n        features_x = self.cnn3d(x.unsqueeze(2)).squeeze(2)\n\n        feat2, feat3, feat4, feat5, feat6 = self.backbone(features_x)\n\n        return [x[:,1,:,:], feat2, feat3, feat4, feat5, feat6]\n    \nsmp.encoders.encoders[\"225d_encoder\"] = {\n    \"encoder\": MiddleSkipConnectionEncoder,\n    'params' : {},\n    'pretrained_settings': {},\n}\n\nmodel = smp.Unet(encoder_name='225d_encoder', encoder_weights=None ,classes=1)\n```",
    "2618627": "Thanks. But I mean, there is already a segmentation NextViT code. Would'nt be just copy and paste?\nAnd there is several ckp with pretrained weights at the repo.",
    "2618977": "Hi LIU, how much video memory and inference time do you need using nextvit based unet?"
  },
  "source": "meta"
}