{
  "id": 475080,
  "title": "9th Place Solution",
  "url": "/competitions/blood-vessel-segmentation/writeups/tereka-9th-place-solution",
  "author_name": "",
  "post_date": "2024-02-09T13:24:46.257Z",
  "votes": 30,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Thank you for host, I learned a lot from this competiton!<br>\nalso special thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> . I refer <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> discussion topics many times.</p>\n<h1>TL;DR</h1>\n<ul>\n<li>MaxVit tiny</li>\n<li>xy, xz, yz inference</li>\n<li>heavy regularization</li>\n<li>probability threshold</li>\n</ul>\n<h1>Solution</h1>\n<h2>Model</h2>\n<ul>\n<li>MaxVit tiny</li>\n</ul>\n<pre><code> (nn.Module):\n     ():\n        (SenUNetStem, self).__init__()\n        kwargs = (\n            in_chans=in_chans,\n            features_only=,\n            \n            pretrained=,\n            out_indices=((encoder_depth)),\n        )\n        self.conv_stem = Conv2dReLU(in_chans, , , use_layernorm=, padding=)\n        self.encoder = timm.create_model(encoder_name, **kwargs)\n        self._out_channels = [\n            ,\n        ] + self.encoder.feature_info.channels()\n\n        self.decoder = UnetDecoder(\n            encoder_channels=self._out_channels,\n            decoder_channels=decoder_channels,\n            n_blocks=encoder_depth,\n            use_batchnorm=decoder_use_batchnorm,\n            center=  encoder_name.startswith()  ,\n            attention_type=decoder_attention_type,\n        )\n\n        self.segmentation_head = SegmentationHead(\n            in_channels=decoder_channels[-] + ,\n            out_channels=classes,\n            activation=activation,\n            kernel_size=,\n        )\n\n        self.n_time = n_time\n        self.pickup_index = pickup_index\n\n     ():\n        B, C, H, W = x.shape\n        h = (H//)*\n        w = (W//)*\n        x = x[:,:,:h,:w]\n        stem = self.conv_stem(x)\n        features = self.encoder(x)        \n        features = [\n            stem,\n        ] + features\n\n        decoder_output = self.decoder(*features)\n        masks = self.segmentation_head(decoder_output)\n        masks = F.pad(masks,[,W-w,,H-h,,,,], mode=, value=)\n\n         masks[:,]\n\nSenUNetStem(\n            encoder_name=,\n            classes=,\n            activation=,\n        )\n</code></pre>\n<h2>Dataset</h2>\n<ul>\n<li>kidney1 and 3 dataset.</li>\n</ul>\n<h2>Training Tricks</h2>\n<p>I focus on regularization trick.<br>\nbecause this competiton have unstable cv, public is not related.<br>\nMoreover host shared public/Private LB information images, I hink it indicate unstable.</p>\n<ul>\n<li>EMA</li>\n<li>50epochs</li>\n<li>AdamW(Weight Decay 1e-2)</li>\n<li>CutMix(until 25ep)</li>\n<li>MixUp(until 25ep)</li>\n<li>DiceLoss(smooth_factor=0.1)</li>\n<li>Heavy Augmentation </li>\n</ul>\n<pre><code>    train_aug = A.Compose([\n        A.HorizontalFlip(=0.5),\n        A.VerticalFlip(=0.5),\n        A.RandomBrightness(=0.1, =0.7),\n        A.OneOf([\n                A.GaussNoise(var_limit=[10, 50]),\n                A.GaussianBlur(),\n                A.MotionBlur(),\n                A.MedianBlur(=3),\n                ], =0.4),\n        A.OneOf([\n            A.GridDistortion(=5, =0.3, =1.0),\n            A.OpticalDistortion(=1., =1.0)\n        ],=0.2),\n        A.ShiftScaleRotate(=0.7, =0.5, =0.2, =30),\n        A.CoarseDropout(=1, =0.25, =0.25),\n        ToTensorV2(=)\n    ])\n</code></pre>\n<ul>\n<li>Crop(512)</li>\n</ul>\n<h2>Inference</h2>\n<p>Inference is xy, xz, yz axis, and crop 512, stride 256.</p>\n<h2>Post-Process</h2>\n<p>Probability threshold(=sigmoid output). I didn't use percentile which method used many public notebook and past segmentation competiton(e.g. Volcano).<br>\nBecause I checked percentile threshold in local cv, it's not stable. I didn't use it.</p>\n<h2>Not worked</h2>\n<ul>\n<li>Bigger models(maxvit base, small)</li>\n<li>large size inference(1024), 512 is enough for this competiton.</li>\n<li>Rotate90</li>\n<li>Pretrained Other volumes(kidney_2/kidney_1_volumes)</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/code/tereka/simpleunet-xy-xz-yz-v2-nbp-b749ff/notebook\" target=\"_blank\">https://www.kaggle.com/code/tereka/simpleunet-xy-xz-yz-v2-nbp-b749ff/notebook</a></p>",
  "messages": [
    {
      "id": "2640615",
      "postDate": "02/07/2024 02:41:30",
      "content": "<p>Thank you for host, I learned a lot from this competiton!<br>\nalso special thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> . I refer <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> discussion topics many times.</p>\n<h1>TL;DR</h1>\n<ul>\n<li>MaxVit tiny</li>\n<li>xy, xz, yz inference</li>\n<li>heavy regularization</li>\n<li>probability threshold</li>\n</ul>\n<h1>Solution</h1>\n<h2>Model</h2>\n<ul>\n<li>MaxVit tiny</li>\n</ul>\n<pre><code> (nn.Module):\n     ():\n        (SenUNetStem, self).__init__()\n        kwargs = (\n            in_chans=in_chans,\n            features_only=,\n            \n            pretrained=,\n            out_indices=((encoder_depth)),\n        )\n        self.conv_stem = Conv2dReLU(in_chans, , , use_layernorm=, padding=)\n        self.encoder = timm.create_model(encoder_name, **kwargs)\n        self._out_channels = [\n            ,\n        ] + self.encoder.feature_info.channels()\n\n        self.decoder = UnetDecoder(\n            encoder_channels=self._out_channels,\n            decoder_channels=decoder_channels,\n            n_blocks=encoder_depth,\n            use_batchnorm=decoder_use_batchnorm,\n            center=  encoder_name.startswith()  ,\n            attention_type=decoder_attention_type,\n        )\n\n        self.segmentation_head = SegmentationHead(\n            in_channels=decoder_channels[-] + ,\n            out_channels=classes,\n            activation=activation,\n            kernel_size=,\n        )\n\n        self.n_time = n_time\n        self.pickup_index = pickup_index\n\n     ():\n        B, C, H, W = x.shape\n        h = (H//)*\n        w = (W//)*\n        x = x[:,:,:h,:w]\n        stem = self.conv_stem(x)\n        features = self.encoder(x)        \n        features = [\n            stem,\n        ] + features\n\n        decoder_output = self.decoder(*features)\n        masks = self.segmentation_head(decoder_output)\n        masks = F.pad(masks,[,W-w,,H-h,,,,], mode=, value=)\n\n         masks[:,]\n\nSenUNetStem(\n            encoder_name=,\n            classes=,\n            activation=,\n        )\n</code></pre>\n<h2>Dataset</h2>\n<ul>\n<li>kidney1 and 3 dataset.</li>\n</ul>\n<h2>Training Tricks</h2>\n<p>I focus on regularization trick.<br>\nbecause this competiton have unstable cv, public is not related.<br>\nMoreover host shared public/Private LB information images, I hink it indicate unstable.</p>\n<ul>\n<li>EMA</li>\n<li>50epochs</li>\n<li>AdamW(Weight Decay 1e-2)</li>\n<li>CutMix(until 25ep)</li>\n<li>MixUp(until 25ep)</li>\n<li>DiceLoss(smooth_factor=0.1)</li>\n<li>Heavy Augmentation </li>\n</ul>\n<pre><code>    train_aug = A.Compose([\n        A.HorizontalFlip(=0.5),\n        A.VerticalFlip(=0.5),\n        A.RandomBrightness(=0.1, =0.7),\n        A.OneOf([\n                A.GaussNoise(var_limit=[10, 50]),\n                A.GaussianBlur(),\n                A.MotionBlur(),\n                A.MedianBlur(=3),\n                ], =0.4),\n        A.OneOf([\n            A.GridDistortion(=5, =0.3, =1.0),\n            A.OpticalDistortion(=1., =1.0)\n        ],=0.2),\n        A.ShiftScaleRotate(=0.7, =0.5, =0.2, =30),\n        A.CoarseDropout(=1, =0.25, =0.25),\n        ToTensorV2(=)\n    ])\n</code></pre>\n<ul>\n<li>Crop(512)</li>\n</ul>\n<h2>Inference</h2>\n<p>Inference is xy, xz, yz axis, and crop 512, stride 256.</p>\n<h2>Post-Process</h2>\n<p>Probability threshold(=sigmoid output). I didn't use percentile which method used many public notebook and past segmentation competiton(e.g. Volcano).<br>\nBecause I checked percentile threshold in local cv, it's not stable. I didn't use it.</p>\n<h2>Not worked</h2>\n<ul>\n<li>Bigger models(maxvit base, small)</li>\n<li>large size inference(1024), 512 is enough for this competiton.</li>\n<li>Rotate90</li>\n<li>Pretrained Other volumes(kidney_2/kidney_1_volumes)</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/code/tereka/simpleunet-xy-xz-yz-v2-nbp-b749ff/notebook\" target=\"_blank\">https://www.kaggle.com/code/tereka/simpleunet-xy-xz-yz-v2-nbp-b749ff/notebook</a></p>",
      "rawMarkdown": "Thank you for host, I learned a lot from this competiton!\nalso special thanks to @hengck23 . I refer @hengck23 discussion topics many times.\n\n# TL;DR\n- MaxVit tiny\n- xy, xz, yz inference\n- heavy regularization\n- probability threshold\n\n# Solution\n## Model\n- MaxVit tiny\n\n```\nclass SenUNetStem(nn.Module):\n    def __init__(self, encoder_name=\"resnest26d\",output_stride=32,\n                 encoder_depth=5 , in_chans=1,\n        decoder_use_batchnorm: bool = True,\n        decoder_channels: List[int] = (256, 128, 64, 32, 16),\n        decoder_attention_type: Optional[str] = None, classes=1, activation=None):\n        super(SenUNetStem, self).__init__()\n        kwargs = dict(\n            in_chans=in_chans,\n            features_only=True,\n            # output_stride=output_stride,\n            pretrained=True,\n            out_indices=tuple(range(encoder_depth)),\n        )\n        self.conv_stem = Conv2dReLU(in_chans, 16, 3, use_layernorm=False, padding=1)\n        self.encoder = timm.create_model(encoder_name, **kwargs)\n        self._out_channels = [\n            32,\n        ] + self.encoder.feature_info.channels()\n\n        self.decoder = UnetDecoder(\n            encoder_channels=self._out_channels,\n            decoder_channels=decoder_channels,\n            n_blocks=encoder_depth,\n            use_batchnorm=decoder_use_batchnorm,\n            center=True if encoder_name.startswith(\"vgg\") else False,\n            attention_type=decoder_attention_type,\n        )\n\n        self.segmentation_head = SegmentationHead(\n            in_channels=decoder_channels[-1] + 16,\n            out_channels=classes,\n            activation=activation,\n            kernel_size=3,\n        )\n\n        self.n_time = n_time\n        self.pickup_index = pickup_index\n\n    def forward(self, x):\n        B, C, H, W = x.shape\n        h = (H//32)*32\n        w = (W//32)*32\n        x = x[:,:,:h,:w]\n        stem = self.conv_stem(x)\n        features = self.encoder(x)        \n        features = [\n            stem,\n        ] + features\n\n        decoder_output = self.decoder(*features)\n        masks = self.segmentation_head(decoder_output)\n        masks = F.pad(masks,[0,W-w,0,H-h,0,0,0,0], mode='constant', value=0)\n        \n        return masks[:,0]\n\nSenUNetStem(\n            encoder_name=\"maxvit_tiny_tf_512.in1k\",\n            classes=1,\n            activation=None,\n        )\n```\n\n## Dataset\n- kidney1 and 3 dataset.\n\n## Training Tricks\nI focus on regularization trick.\nbecause this competiton have unstable cv, public is not related.\nMoreover host shared public/Private LB information images, I hink it indicate unstable.\n\n- EMA\n- 50epochs\n- AdamW(Weight Decay 1e-2)\n- CutMix(until 25ep)\n- MixUp(until 25ep)\n- DiceLoss(smooth_factor=0.1)\n- Heavy Augmentation \n```\n    train_aug = A.Compose([\n        A.HorizontalFlip(p=0.5),\n        A.VerticalFlip(p=0.5),\n        A.RandomBrightness(limit=0.1, p=0.7),\n        A.OneOf([\n                A.GaussNoise(var_limit=[10, 50]),\n                A.GaussianBlur(),\n                A.MotionBlur(),\n                A.MedianBlur(blur_limit=3),\n                ], p=0.4),\n        A.OneOf([\n            A.GridDistortion(num_steps=5, distort_limit=0.3, p=1.0),\n            A.OpticalDistortion(distort_limit=1., p=1.0)\n        ],p=0.2),\n        A.ShiftScaleRotate(p=0.7, scale_limit=0.5, shift_limit=0.2, rotate_limit=30),\n        A.CoarseDropout(max_holes=1, max_height=0.25, max_width=0.25),\n        ToTensorV2(transpose_mask=True)\n    ])\n```\n- Crop(512)\n\n## Inference\nInference is xy, xz, yz axis, and crop 512, stride 256.\n\n## Post-Process\nProbability threshold(=sigmoid output). I didn't use percentile which method used many public notebook and past segmentation competiton(e.g. Volcano).\nBecause I checked percentile threshold in local cv, it's not stable. I didn't use it.\n\n## Not worked\n- Bigger models(maxvit base, small)\n- large size inference(1024), 512 is enough for this competiton.\n- Rotate90\n- Pretrained Other volumes(kidney_2/kidney_1_volumes)\n\nhttps://www.kaggle.com/code/tereka/simpleunet-xy-xz-yz-v2-nbp-b749ff/notebook",
      "votes": null
    },
    {
      "id": "2640670",
      "postDate": "02/07/2024 03:32:21",
      "content": "<p>Congratulations! Would you consider sharing your notebook?</p>",
      "rawMarkdown": "Congratulations! Would you consider sharing your notebook?",
      "votes": null
    },
    {
      "id": "2641491",
      "postDate": "02/07/2024 14:11:32",
      "content": "<p><a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a> congratz with gold zone and thank you for the code snippets. Could you share your ema implementation if it is not a secret, please? I've been running standard pytorch EMA implementation and was boosting the performance for CV, I wonder I did it correctly. This was my code:</p>\n<pre><code> torch_ema  ExponentialMovingAverage\n\nema = ExponentialMovingAverage(model.parameters(), decay=)\n\n\n            output = model(images, qmin, qmax)\n            loss = criterion(output, masks)\n            optimizer.zero_grad()\n            loss .backward()\n            optimizer.step()\n            ema.update()\n</code></pre>",
      "rawMarkdown": "tereka congratz with gold zone and thank you for the code snippets. Could you share your ema implementation if it is not a secret, please? I've been running standard pytorch EMA implementation and was boosting the performance for CV, I wonder I did it correctly. This was my code:\n\n```python\nfrom torch_ema import ExponentialMovingAverage\n\nema = ExponentialMovingAverage(model.parameters(), decay=0.995)\n\n# .... train loop\n            output = model(images, qmin, qmax)\n            loss = criterion(output, masks)\n            optimizer.zero_grad()\n            loss .backward()\n            optimizer.step()\n            ema.update()\n```",
      "votes": null
    },
    {
      "id": "2641631",
      "postDate": "02/07/2024 15:25:11",
      "content": "<p>thanks, I will share inference notebook after finalize.</p>",
      "rawMarkdown": "thanks, I will share inference notebook after finalize.",
      "votes": null
    },
    {
      "id": "2641635",
      "postDate": "02/07/2024 15:28:07",
      "content": "<p>Thanks, here is my EMA code.<br>\nI think your code(ema update position) is correct.</p>\n<pre><code># PyTorch Lightning Module\n    def training:\n        self.ema_model.update(self.model)\n\n :\n    def :\n        super(ModelEMA, self).\n        # make a copy  the model  accumulating moving average  weights\n        self. = deepcopy(model)\n        self..eval\n        self.decay = decay\n        self.device = device  # perform ema on different device from model  set\n         self.device is not None:\n            self..(device=device)\n\n    def :\n         torch.no:\n             ema_v, model_v  zip(self..state.values, model.state.values):\n                 self.device is not None:\n                    model_v = model_v.(device=self.device)\n                ema_v.copy)\n\n    def update(self, model):\n        self.m)\n\n    def set(self, model):\n        self.\n</code></pre>",
      "rawMarkdown": "Thanks, here is my EMA code.\nI think your code(ema update position) is correct.\n\n```\n# PyTorch Lightning Module\n    def training_step(self, batch, batch_idx):\n        self.ema_model.update(self.model)\n\nclass ModelEMA(nn.Module):\n    def __init__(self, model, decay=0.9999, device=None):\n        super(ModelEMA, self).__init__()\n        # make a copy of the model for accumulating moving average of weights\n        self.module = deepcopy(model)\n        self.module.eval()\n        self.decay = decay\n        self.device = device  # perform ema on different device from model if set\n        if self.device is not None:\n            self.module.to(device=device)\n\n    def _update(self, model, update_fn):\n        with torch.no_grad():\n            for ema_v, model_v in zip(self.module.state_dict().values(), model.state_dict().values()):\n                if self.device is not None:\n                    model_v = model_v.to(device=self.device)\n                ema_v.copy_(update_fn(ema_v, model_v))\n\n    def update(self, model):\n        self._update(model, update_fn=lambda e, m: self.decay * e + (1. - self.decay) * m)\n\n    def set(self, model):\n        self._update(model, update_fn=lambda e, m: m)\n```",
      "votes": null
    },
    {
      "id": "2641647",
      "postDate": "02/07/2024 15:39:08",
      "content": "<p><a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a> wow, I did not expect pytorch-lightning that's exactly what I needed. Thank you, you are the man!</p>",
      "rawMarkdown": "tereka wow, I did not expect pytorch-lightning that's exactly what I needed. Thank you, you are the man!",
      "votes": null
    },
    {
      "id": "2642444",
      "postDate": "02/08/2024 06:40:47",
      "content": "<p><a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a> , thank you for sharing the idea + code and congratulation with your 9th position.</p>\n<p>I would like to ask one very beginner's question confused me a while:</p>\n<p>How do you decide change the decoder to UnetDecoder?  Compared with MaxVit's original decoder, what is the benefit using this one?</p>",
      "rawMarkdown": "tereka , thank you for sharing the idea + code and congratulation with your 9th position.\n\nI would like to ask one very beginner's question confused me a while:\n\nHow do you decide change the decoder to UnetDecoder?  Compared with MaxVit's original decoder, what is the benefit using this one?",
      "votes": null
    },
    {
      "id": "2643161",
      "postDate": "02/08/2024 16:54:45",
      "content": "<p>thanks.<br>\nDoes Original decoder means smp decoder ?<br>\nI would like to input 1/2 x 1/2 image, original max vit 1/4.<br>\nI need to customize smp unet.</p>",
      "rawMarkdown": "thanks.\nDoes Original decoder means smp decoder ?\nI would like to input 1/2 x 1/2 image, original max vit 1/4.\nI need to customize smp unet.",
      "votes": null
    },
    {
      "id": "2644734",
      "postDate": "02/09/2024 16:32:56",
      "content": "<p>Thanks for sharing this class <a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a>!</p>\n<p>A quick general EMA question.. Does a decay rate of 0.9999 typically work well for you, or do you often tune this value?</p>",
      "rawMarkdown": "Thanks for sharing this class @tereka!\n\nA quick general EMA question.. Does a decay rate of 0.9999 typically work well for you, or do you often tune this value?",
      "votes": null
    },
    {
      "id": "2645065",
      "postDate": "02/09/2024 23:01:25",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> <br>\nthanks! EMA decay rate is 0.99 in this competition.<br>\nif EMA decay value is low, model convergence is very slow..</p>",
      "rawMarkdown": "brendanartley \nthanks! EMA decay rate is 0.99 in this competition.\nif EMA decay value is low, model convergence is very slow..",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2640670,
      "author_name": "ardaorcun",
      "author_url": "",
      "post_date": "02/07/2024 03:32:21",
      "content": "<p>Congratulations! Would you consider sharing your notebook?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2641631,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "02/07/2024 15:25:11",
          "content": "<p>thanks, I will share inference notebook after finalize.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2641491,
      "author_name": "sergiosaharovskiy",
      "author_url": "",
      "post_date": "02/07/2024 14:11:32",
      "content": "<p><a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a> congratz with gold zone and thank you for the code snippets. Could you share your ema implementation if it is not a secret, please? I've been running standard pytorch EMA implementation and was boosting the performance for CV, I wonder I did it correctly. This was my code:</p>\n<pre><code> torch_ema  ExponentialMovingAverage\n\nema = ExponentialMovingAverage(model.parameters(), decay=)\n\n\n            output = model(images, qmin, qmax)\n            loss = criterion(output, masks)\n            optimizer.zero_grad()\n            loss .backward()\n            optimizer.step()\n            ema.update()\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2641635,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "02/07/2024 15:28:07",
          "content": "<p>Thanks, here is my EMA code.<br>\nI think your code(ema update position) is correct.</p>\n<pre><code># PyTorch Lightning Module\n    def training:\n        self.ema_model.update(self.model)\n\n :\n    def :\n        super(ModelEMA, self).\n        # make a copy  the model  accumulating moving average  weights\n        self. = deepcopy(model)\n        self..eval\n        self.decay = decay\n        self.device = device  # perform ema on different device from model  set\n         self.device is not None:\n            self..(device=device)\n\n    def :\n         torch.no:\n             ema_v, model_v  zip(self..state.values, model.state.values):\n                 self.device is not None:\n                    model_v = model_v.(device=self.device)\n                ema_v.copy)\n\n    def update(self, model):\n        self.m)\n\n    def set(self, model):\n        self.\n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 2641647,
              "author_name": "sergiosaharovskiy",
              "author_url": "",
              "post_date": "02/07/2024 15:39:08",
              "content": "<p><a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a> wow, I did not expect pytorch-lightning that's exactly what I needed. Thank you, you are the man!</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2644734,
              "author_name": "brendanartley",
              "author_url": "",
              "post_date": "02/09/2024 16:32:56",
              "content": "<p>Thanks for sharing this class <a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a>!</p>\n<p>A quick general EMA question.. Does a decay rate of 0.9999 typically work well for you, or do you often tune this value?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2645065,
                  "author_name": "tereka",
                  "author_url": "",
                  "post_date": "02/09/2024 23:01:25",
                  "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> <br>\nthanks! EMA decay rate is 0.99 in this competition.<br>\nif EMA decay value is low, model convergence is very slow..</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2642444,
      "author_name": "yue300c",
      "author_url": "",
      "post_date": "02/08/2024 06:40:47",
      "content": "<p><a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a> , thank you for sharing the idea + code and congratulation with your 9th position.</p>\n<p>I would like to ask one very beginner's question confused me a while:</p>\n<p>How do you decide change the decoder to UnetDecoder?  Compared with MaxVit's original decoder, what is the benefit using this one?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2643161,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "02/08/2024 16:54:45",
          "content": "<p>thanks.<br>\nDoes Original decoder means smp decoder ?<br>\nI would like to input 1/2 x 1/2 image, original max vit 1/4.<br>\nI need to customize smp unet.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2640615": "Thank you for host, I learned a lot from this competiton!\nalso special thanks to @hengck23 . I refer @hengck23 discussion topics many times.\n\n# TL;DR\n- MaxVit tiny\n- xy, xz, yz inference\n- heavy regularization\n- probability threshold\n\n# Solution\n## Model\n- MaxVit tiny\n\n```\nclass SenUNetStem(nn.Module):\n    def __init__(self, encoder_name=\"resnest26d\",output_stride=32,\n                 encoder_depth=5 , in_chans=1,\n        decoder_use_batchnorm: bool = True,\n        decoder_channels: List[int] = (256, 128, 64, 32, 16),\n        decoder_attention_type: Optional[str] = None, classes=1, activation=None):\n        super(SenUNetStem, self).__init__()\n        kwargs = dict(\n            in_chans=in_chans,\n            features_only=True,\n            # output_stride=output_stride,\n            pretrained=True,\n            out_indices=tuple(range(encoder_depth)),\n        )\n        self.conv_stem = Conv2dReLU(in_chans, 16, 3, use_layernorm=False, padding=1)\n        self.encoder = timm.create_model(encoder_name, **kwargs)\n        self._out_channels = [\n            32,\n        ] + self.encoder.feature_info.channels()\n\n        self.decoder = UnetDecoder(\n            encoder_channels=self._out_channels,\n            decoder_channels=decoder_channels,\n            n_blocks=encoder_depth,\n            use_batchnorm=decoder_use_batchnorm,\n            center=True if encoder_name.startswith(\"vgg\") else False,\n            attention_type=decoder_attention_type,\n        )\n\n        self.segmentation_head = SegmentationHead(\n            in_channels=decoder_channels[-1] + 16,\n            out_channels=classes,\n            activation=activation,\n            kernel_size=3,\n        )\n\n        self.n_time = n_time\n        self.pickup_index = pickup_index\n\n    def forward(self, x):\n        B, C, H, W = x.shape\n        h = (H//32)*32\n        w = (W//32)*32\n        x = x[:,:,:h,:w]\n        stem = self.conv_stem(x)\n        features = self.encoder(x)        \n        features = [\n            stem,\n        ] + features\n\n        decoder_output = self.decoder(*features)\n        masks = self.segmentation_head(decoder_output)\n        masks = F.pad(masks,[0,W-w,0,H-h,0,0,0,0], mode='constant', value=0)\n        \n        return masks[:,0]\n\nSenUNetStem(\n            encoder_name=\"maxvit_tiny_tf_512.in1k\",\n            classes=1,\n            activation=None,\n        )\n```\n\n## Dataset\n- kidney1 and 3 dataset.\n\n## Training Tricks\nI focus on regularization trick.\nbecause this competiton have unstable cv, public is not related.\nMoreover host shared public/Private LB information images, I hink it indicate unstable.\n\n- EMA\n- 50epochs\n- AdamW(Weight Decay 1e-2)\n- CutMix(until 25ep)\n- MixUp(until 25ep)\n- DiceLoss(smooth_factor=0.1)\n- Heavy Augmentation \n```\n    train_aug = A.Compose([\n        A.HorizontalFlip(p=0.5),\n        A.VerticalFlip(p=0.5),\n        A.RandomBrightness(limit=0.1, p=0.7),\n        A.OneOf([\n                A.GaussNoise(var_limit=[10, 50]),\n                A.GaussianBlur(),\n                A.MotionBlur(),\n                A.MedianBlur(blur_limit=3),\n                ], p=0.4),\n        A.OneOf([\n            A.GridDistortion(num_steps=5, distort_limit=0.3, p=1.0),\n            A.OpticalDistortion(distort_limit=1., p=1.0)\n        ],p=0.2),\n        A.ShiftScaleRotate(p=0.7, scale_limit=0.5, shift_limit=0.2, rotate_limit=30),\n        A.CoarseDropout(max_holes=1, max_height=0.25, max_width=0.25),\n        ToTensorV2(transpose_mask=True)\n    ])\n```\n- Crop(512)\n\n## Inference\nInference is xy, xz, yz axis, and crop 512, stride 256.\n\n## Post-Process\nProbability threshold(=sigmoid output). I didn't use percentile which method used many public notebook and past segmentation competiton(e.g. Volcano).\nBecause I checked percentile threshold in local cv, it's not stable. I didn't use it.\n\n## Not worked\n- Bigger models(maxvit base, small)\n- large size inference(1024), 512 is enough for this competiton.\n- Rotate90\n- Pretrained Other volumes(kidney_2/kidney_1_volumes)\n\nhttps://www.kaggle.com/code/tereka/simpleunet-xy-xz-yz-v2-nbp-b749ff/notebook",
    "2640670": "Congratulations! Would you consider sharing your notebook?",
    "2641491": "tereka congratz with gold zone and thank you for the code snippets. Could you share your ema implementation if it is not a secret, please? I've been running standard pytorch EMA implementation and was boosting the performance for CV, I wonder I did it correctly. This was my code:\n\n```python\nfrom torch_ema import ExponentialMovingAverage\n\nema = ExponentialMovingAverage(model.parameters(), decay=0.995)\n\n# .... train loop\n            output = model(images, qmin, qmax)\n            loss = criterion(output, masks)\n            optimizer.zero_grad()\n            loss .backward()\n            optimizer.step()\n            ema.update()\n```",
    "2641631": "thanks, I will share inference notebook after finalize.",
    "2641635": "Thanks, here is my EMA code.\nI think your code(ema update position) is correct.\n\n```\n# PyTorch Lightning Module\n    def training_step(self, batch, batch_idx):\n        self.ema_model.update(self.model)\n\nclass ModelEMA(nn.Module):\n    def __init__(self, model, decay=0.9999, device=None):\n        super(ModelEMA, self).__init__()\n        # make a copy of the model for accumulating moving average of weights\n        self.module = deepcopy(model)\n        self.module.eval()\n        self.decay = decay\n        self.device = device  # perform ema on different device from model if set\n        if self.device is not None:\n            self.module.to(device=device)\n\n    def _update(self, model, update_fn):\n        with torch.no_grad():\n            for ema_v, model_v in zip(self.module.state_dict().values(), model.state_dict().values()):\n                if self.device is not None:\n                    model_v = model_v.to(device=self.device)\n                ema_v.copy_(update_fn(ema_v, model_v))\n\n    def update(self, model):\n        self._update(model, update_fn=lambda e, m: self.decay * e + (1. - self.decay) * m)\n\n    def set(self, model):\n        self._update(model, update_fn=lambda e, m: m)\n```",
    "2641647": "tereka wow, I did not expect pytorch-lightning that's exactly what I needed. Thank you, you are the man!",
    "2642444": "tereka , thank you for sharing the idea + code and congratulation with your 9th position.\n\nI would like to ask one very beginner's question confused me a while:\n\nHow do you decide change the decoder to UnetDecoder?  Compared with MaxVit's original decoder, what is the benefit using this one?",
    "2643161": "thanks.\nDoes Original decoder means smp decoder ?\nI would like to input 1/2 x 1/2 image, original max vit 1/4.\nI need to customize smp unet.",
    "2644734": "Thanks for sharing this class @tereka!\n\nA quick general EMA question.. Does a decay rate of 0.9999 typically work well for you, or do you often tune this value?",
    "2645065": "brendanartley \nthanks! EMA decay rate is 0.99 in this competition.\nif EMA decay value is low, model convergence is very slow.."
  },
  "source": "meta"
}