{
  "id": 475288,
  "title": "5th place solution - 3D interpolation is all you need (updated with code)",
  "url": "/competitions/blood-vessel-segmentation/writeups/ivan-panshin-5th-place-solution-3d-interpolation-i",
  "author_name": "",
  "post_date": "2024-02-16T15:22:51.410Z",
  "votes": 26,
  "comment_count": 10,
  "views": 0,
  "content": "<p>First of all, I want to express my deepest gratitude to the organizers of this competition. The reason I mostly compete in CV competitions is because I love it. Especially in medical competitions such as this one. So I'm really happy we managed to work with such a great technology (resolution is insane). </p>\n<h3>Validation</h3>\n<p>Initially I thought it's gonna be a challenge validation-wise, since we simply don't have enough data for reliable validation. To make the validation without leakage and similar to test, I decided to train my pipelines in 2 ways: </p>\n<ul>\n<li>Take kidney_1 as base for trainings, kidney_3 for validation</li>\n<li>Take kidney_3 as base for trainings, kidney_1 for validation</li>\n</ul>\n<p>Since test is densely annotated, I wanted to compute metrics only on densely annotated kidneys, which eliminated kidney_2 from the discussion. </p>\n<h3>Data</h3>\n<p>I always believe that data is the key. So I tried my hardest to utilize the additional datasets provided by organizers at <a href=\"https://human-organ-atlas.esrf.eu\" target=\"_blank\">Human Organ Atlas</a>. </p>\n<p>In the end, I decided to use to following with the help of pseudo-labeling:</p>\n<ul>\n<li>LADAF-2020-31 kidney</li>\n<li>LADAF-2020-27 spleen</li>\n</ul>\n<p>In other words, after examining the pseudo annotations on spleen, I realized that they are quite good and should serve as a good regularization method. </p>\n<p>Additionally, I tried to use heart + brain + lung. However, my models make semi-accurate predictions for lung, but horrible for heart + brain. So in the end I decided to stick with kidney + spleen. </p>\n<h3>Pseudo annotations</h3>\n<p>I would say there are 2 important points</p>\n<p>The first one is don't pseudo annotate everything right away. In order to create full pseudo annotations, I run a 4-step process:</p>\n<ul>\n<li>Train on kidney_1. Pseudo-annotate kidney_2</li>\n<li>Train on kidney_1 + kidney_2. Pseudo-annotate the 2020-31 kidney.</li>\n<li>Train on kidney_1 + kidney_2 + 2020-31-kidney. Pseudo-annotate 2020-27 spleen.</li>\n<li>Train on kidney_1 + kidney_2 + 2020-31-kidney + 2020-27 spleen. </li>\n</ul>\n<p>The second point is that don't use hard labels. In other words, don't apply thresholding to the predictions. Simply use soft labels (predictions are sigmoided to be in the range of [0,1]) for training. </p>\n<h3>Loss</h3>\n<p>My baseline go-to loss in semantic segmentation is <code>CE + Dice + Focal</code>. This worked quite well in this competition. However, since we have a surface metric, I wanted to weight the boundaries of masks more heavily. </p>\n<ul>\n<li>What didn't work: losses I found in open-source repositories (like Hausdorff Distance loss).</li>\n<li>What worked really well in terms of Surface Dice, FP and FN on validation: CE with x2 weights for boundaries. </li>\n</ul>\n<p>So in the end I decided to use <code>CE_boundaries + Dice + Focal</code> for most of my models, and <code>CE_boundaries + Twersky + Focal</code> for a single model.</p>\n<p>Twersky was focusing more on FN rather than FP, but more on that in the next section. </p>\n<pre><code> (torch.nn.modules.loss._Loss):\n     ():\n        ().__init__()\n        self.bound = EdgeEmphasisLoss(alpha=bound_alpha)\n        self.dice = smp.losses.DiceLoss(mode=)\n        self.focal = smp.losses.FocalLoss(mode=)\n        self.bound_weight = bound_weight\n        self.dice_weight = dice_weight\n        self.focal_weight = focal_weight\n\n     ():\n         (\n            self.bound_weight * self.bound(preds, gt, boundaries)\n            + self.dice_weight * self.dice(preds, gt)\n            + self.focal_weight * self.focal(preds, gt)\n        )\n\n (nn.Module):\n     ():\n        (EdgeEmphasisLoss, self).__init__()\n        self.alpha = alpha\n\n     ():\n        bce_loss = F.binary_cross_entropy_with_logits(inputs, targets, reduction=)\n\n        \n        weighted_loss = bce_loss * ( + self.alpha * boundaries)\n\n        \n         weighted_loss.mean()\n</code></pre>\n<h3>Preprocessing</h3>\n<p>After analyzing initial models and its errors, I realized that my biggest issue is FN, not FP. In other words, my models simply don't see some masks, mostly the small ones. </p>\n<p>So I decided to increase the resolution of my trainings with crops from 512x512 to 1024x1024. However, after a couple of hours of training it hit me: that doesn't make much sense. By going from 512x512 to 1024x1024 I don't really increase resolution (each pixels holds the same real-world size), just the context, and 512x512 seemed like a big-enough context already. </p>\n<p>Instead, I decided to do the following: </p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        self.upscale_factor = upscale_factor\n\n        self.model = Unet(\n            encoder_weights=encoder_weights,\n            encoder_name=encoder_name,\n            decoder_use_batchnorm=decoder_use_batchnorm,\n            in_channels=in_channels,\n            classes=classes,\n        )\n\n     ():\n        x = torch.nn.functional.interpolate(\n            x, (x.shape[-] * self.upscale_factor, x.shape[-] * self.upscale_factor), mode=\n        )\n        x = self.model(x)\n        x = torch.nn.functional.interpolate(\n            x, (x.shape[-] // self.upscale_factor, x.shape[-] // self.upscale_factor), mode=\n        )\n         x\n</code></pre>\n<p>This approached worked really well and I could clearly see improvements both on CV, and LB.</p>\n<h3>Models</h3>\n<p>I used only U-Net models from SMP with different backbones. Tried a lot of things, but for final ensembles decided to settle on the following:</p>\n<ul>\n<li>effnet_v2_s</li>\n<li>effnet_v2_m</li>\n<li>maxvit_base</li>\n<li>dpn68 </li>\n</ul>\n<p>Maxvit was trained on 512x512 crops, effent and dpn - on 512x512 with x2 interpolation. Crops were used from xy, xz, and yz axes. During inference, I use the same crops resolution with overlaps of crops_size / 2 (so that's 256). In other words, sliding window approach.</p>\n<p>Augmentation were medium-level in terms of intensity. </p>\n<pre><code> A.Compose(\n    [\n        A.ShiftScaleRotate(\n            p=,\n            shift_limit_x=(-, ),\n            shift_limit_y=(-, ),\n            scale_limit=(-, ),\n            rotate_limit=(-, ),\n            border_mode=cv2.BORDER_CONSTANT,\n            \n        ),\n        A.RandomBrightnessContrast(\n            brightness_limit=(-, ),\n            contrast_limit=(-, ),\n            p=,\n        ),\n        A.HorizontalFlip(),\n        A.VerticalFlip(),\n        A.OneOf(\n            [\n                A.GridDistortion(border_mode=cv2.BORDER_CONSTANT, distort_limit=),\n                A.ElasticTransform(border_mode=cv2.BORDER_CONSTANT),\n            ],\n            p=,\n        ),\n        AT.ToTensorV2(),\n    ],\n    )\n</code></pre>\n<h3>Post processing</h3>\n<p>I tried to use cc3d to remove small objects, it made weak models better, but no difference for ensemble.</p>\n<h3>Private resolution</h3>\n<p>Now, this part is really tricky. My huge thanks to the organizers for announcing the test resolutions. It sincerely warms my heart to see organizers interact with participants that much here on the forum. Really, thank you. </p>\n<p>One approach is not to do anything. You train your model on 50um/voxel, inference on 63um/voxel. Considering I use conv-based backbones (except for maxvit) that have some level of scale-invariance + have scale augs in validation, this might work.</p>\n<p>The second approach is to do rescaling. I believe the correct approach for rescaling is the following: </p>\n<pre><code> test_kidney == :\n    private_res = \n    public_res = \n\n    scale = private_res / public_res\n\n    d_original, h_original, w_original = test_kidney_image.shape\n    test_kidney_image = torch.tensor(test_kidney_image).view(, , d_original, h_original, w_original)\n    test_kidney_image = test_kidney_image.to(dtype=torch.float32)\n    test_kidney_image = torch.nn.functional.interpolate(test_kidney_image, (\n        (d_original*scale),\n        (h_original*scale),\n        (w_original*scale),\n    ), mode=).squeeze().numpy()\n</code></pre>\n<p>…</p>\n<pre><code>d_preds, h_preds, w_preds = preds_ensemble.shape \npreds_ensemble = preds_ensemble.view(, , d_preds, h_preds, w_preds)\npreds_ensemble = preds_ensemble.to(dtype=torch.float32)\n\npreds_ensemble = torch.nn.functional.interpolate(preds_ensemble, (\n    d_original,\n    h_original,\n    w_original,\n), mode=).squeeze()\n</code></pre>\n<p>So we do 3D resize instead of 2D one: re-scale image from 63um (private) to 50um (public + CV), compute predictions, and re-scale them back to 63um. Simply going for 2D would work as well, but theoretically you end up with different spatial and temporal resolutions in that case. </p>\n<p>This trick helped. To give a single point (I don't have much else): the same ensemble scores 0.634 on private without interpolation, and 0.670 - with interpolation. </p>\n<p>To be honest, I didn't think it would make that much difference. I tried the following experiment locally: </p>\n<ul>\n<li>Download kidney in 25um resolution. Compute predictions in 25um, interpolate them to 50um, compute metrics. This approach brought my 0.92 surface dice to 0.895. Which is quite good, considering we're talking about x2 interpolation in all 3 directions (that's 8 times less volume) and the fact that it's harder to detect small objects in smaller resolution.</li>\n<li>Download kidney in 25um resolution. Interpolate image to 50um, compute predictions, compute metrics. This approach essentially provided the same metrics as in the case of simply using 50um from organizers. </li>\n</ul>\n<p>So even though I didn't really think interpolation is that important, it also didn't hurt (I was afraid of interpolation artifacts), so I used it for both final subs. </p>\n<h3>Final subs</h3>\n<p>Both subs have an ensemble of 3 models, each inferenced on all 3 axes without TTA (TTA took too much time, and didn't really help on CV). </p>\n<ul>\n<li>First sub. CV: 0.84 (kidney_1), Public: 0.768. Private: 0.566<br>\n<code>Maxvit_ce_dice_focal</code> + <code>effnet_v2_s_ce_dice_focal</code> + <code>effnet_v2_m_ce_dice_focal</code> trained on kidney_3, validated on kidney_1. This approach didn't work that well on CV, and also on Public and Private. </li>\n<li>Second sub. CV: 0.923 (kidney_3), Public: 0.855. Private: 0.691<br>\n<code>Maxvit_ce_dice_focal</code> + <code>effnet_v2_s_ce_bounds_dice_focal</code> + <code>dpn_68_ce_bounds_twersky_focal</code>.</li>\n</ul>\n<p>Code:</p>\n<ul>\n<li>Inference notebook <a href=\"https://www.kaggle.com/code/ivanpan/final-submission/notebook\" target=\"_blank\">link</a></li>\n<li>Training code <a href=\"https://github.com/ivanpanshin/segment-vasculature-5th-place\" target=\"_blank\">link</a></li>\n</ul>",
  "messages": [
    {
      "id": "2641911",
      "postDate": "02/07/2024 18:35:55",
      "content": "<p>First of all, I want to express my deepest gratitude to the organizers of this competition. The reason I mostly compete in CV competitions is because I love it. Especially in medical competitions such as this one. So I'm really happy we managed to work with such a great technology (resolution is insane). </p>\n<h3>Validation</h3>\n<p>Initially I thought it's gonna be a challenge validation-wise, since we simply don't have enough data for reliable validation. To make the validation without leakage and similar to test, I decided to train my pipelines in 2 ways: </p>\n<ul>\n<li>Take kidney_1 as base for trainings, kidney_3 for validation</li>\n<li>Take kidney_3 as base for trainings, kidney_1 for validation</li>\n</ul>\n<p>Since test is densely annotated, I wanted to compute metrics only on densely annotated kidneys, which eliminated kidney_2 from the discussion. </p>\n<h3>Data</h3>\n<p>I always believe that data is the key. So I tried my hardest to utilize the additional datasets provided by organizers at <a href=\"https://human-organ-atlas.esrf.eu\" target=\"_blank\">Human Organ Atlas</a>. </p>\n<p>In the end, I decided to use to following with the help of pseudo-labeling:</p>\n<ul>\n<li>LADAF-2020-31 kidney</li>\n<li>LADAF-2020-27 spleen</li>\n</ul>\n<p>In other words, after examining the pseudo annotations on spleen, I realized that they are quite good and should serve as a good regularization method. </p>\n<p>Additionally, I tried to use heart + brain + lung. However, my models make semi-accurate predictions for lung, but horrible for heart + brain. So in the end I decided to stick with kidney + spleen. </p>\n<h3>Pseudo annotations</h3>\n<p>I would say there are 2 important points</p>\n<p>The first one is don't pseudo annotate everything right away. In order to create full pseudo annotations, I run a 4-step process:</p>\n<ul>\n<li>Train on kidney_1. Pseudo-annotate kidney_2</li>\n<li>Train on kidney_1 + kidney_2. Pseudo-annotate the 2020-31 kidney.</li>\n<li>Train on kidney_1 + kidney_2 + 2020-31-kidney. Pseudo-annotate 2020-27 spleen.</li>\n<li>Train on kidney_1 + kidney_2 + 2020-31-kidney + 2020-27 spleen. </li>\n</ul>\n<p>The second point is that don't use hard labels. In other words, don't apply thresholding to the predictions. Simply use soft labels (predictions are sigmoided to be in the range of [0,1]) for training. </p>\n<h3>Loss</h3>\n<p>My baseline go-to loss in semantic segmentation is <code>CE + Dice + Focal</code>. This worked quite well in this competition. However, since we have a surface metric, I wanted to weight the boundaries of masks more heavily. </p>\n<ul>\n<li>What didn't work: losses I found in open-source repositories (like Hausdorff Distance loss).</li>\n<li>What worked really well in terms of Surface Dice, FP and FN on validation: CE with x2 weights for boundaries. </li>\n</ul>\n<p>So in the end I decided to use <code>CE_boundaries + Dice + Focal</code> for most of my models, and <code>CE_boundaries + Twersky + Focal</code> for a single model.</p>\n<p>Twersky was focusing more on FN rather than FP, but more on that in the next section. </p>\n<pre><code> (torch.nn.modules.loss._Loss):\n     ():\n        ().__init__()\n        self.bound = EdgeEmphasisLoss(alpha=bound_alpha)\n        self.dice = smp.losses.DiceLoss(mode=)\n        self.focal = smp.losses.FocalLoss(mode=)\n        self.bound_weight = bound_weight\n        self.dice_weight = dice_weight\n        self.focal_weight = focal_weight\n\n     ():\n         (\n            self.bound_weight * self.bound(preds, gt, boundaries)\n            + self.dice_weight * self.dice(preds, gt)\n            + self.focal_weight * self.focal(preds, gt)\n        )\n\n (nn.Module):\n     ():\n        (EdgeEmphasisLoss, self).__init__()\n        self.alpha = alpha\n\n     ():\n        bce_loss = F.binary_cross_entropy_with_logits(inputs, targets, reduction=)\n\n        \n        weighted_loss = bce_loss * ( + self.alpha * boundaries)\n\n        \n         weighted_loss.mean()\n</code></pre>\n<h3>Preprocessing</h3>\n<p>After analyzing initial models and its errors, I realized that my biggest issue is FN, not FP. In other words, my models simply don't see some masks, mostly the small ones. </p>\n<p>So I decided to increase the resolution of my trainings with crops from 512x512 to 1024x1024. However, after a couple of hours of training it hit me: that doesn't make much sense. By going from 512x512 to 1024x1024 I don't really increase resolution (each pixels holds the same real-world size), just the context, and 512x512 seemed like a big-enough context already. </p>\n<p>Instead, I decided to do the following: </p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        self.upscale_factor = upscale_factor\n\n        self.model = Unet(\n            encoder_weights=encoder_weights,\n            encoder_name=encoder_name,\n            decoder_use_batchnorm=decoder_use_batchnorm,\n            in_channels=in_channels,\n            classes=classes,\n        )\n\n     ():\n        x = torch.nn.functional.interpolate(\n            x, (x.shape[-] * self.upscale_factor, x.shape[-] * self.upscale_factor), mode=\n        )\n        x = self.model(x)\n        x = torch.nn.functional.interpolate(\n            x, (x.shape[-] // self.upscale_factor, x.shape[-] // self.upscale_factor), mode=\n        )\n         x\n</code></pre>\n<p>This approached worked really well and I could clearly see improvements both on CV, and LB.</p>\n<h3>Models</h3>\n<p>I used only U-Net models from SMP with different backbones. Tried a lot of things, but for final ensembles decided to settle on the following:</p>\n<ul>\n<li>effnet_v2_s</li>\n<li>effnet_v2_m</li>\n<li>maxvit_base</li>\n<li>dpn68 </li>\n</ul>\n<p>Maxvit was trained on 512x512 crops, effent and dpn - on 512x512 with x2 interpolation. Crops were used from xy, xz, and yz axes. During inference, I use the same crops resolution with overlaps of crops_size / 2 (so that's 256). In other words, sliding window approach.</p>\n<p>Augmentation were medium-level in terms of intensity. </p>\n<pre><code> A.Compose(\n    [\n        A.ShiftScaleRotate(\n            p=,\n            shift_limit_x=(-, ),\n            shift_limit_y=(-, ),\n            scale_limit=(-, ),\n            rotate_limit=(-, ),\n            border_mode=cv2.BORDER_CONSTANT,\n            \n        ),\n        A.RandomBrightnessContrast(\n            brightness_limit=(-, ),\n            contrast_limit=(-, ),\n            p=,\n        ),\n        A.HorizontalFlip(),\n        A.VerticalFlip(),\n        A.OneOf(\n            [\n                A.GridDistortion(border_mode=cv2.BORDER_CONSTANT, distort_limit=),\n                A.ElasticTransform(border_mode=cv2.BORDER_CONSTANT),\n            ],\n            p=,\n        ),\n        AT.ToTensorV2(),\n    ],\n    )\n</code></pre>\n<h3>Post processing</h3>\n<p>I tried to use cc3d to remove small objects, it made weak models better, but no difference for ensemble.</p>\n<h3>Private resolution</h3>\n<p>Now, this part is really tricky. My huge thanks to the organizers for announcing the test resolutions. It sincerely warms my heart to see organizers interact with participants that much here on the forum. Really, thank you. </p>\n<p>One approach is not to do anything. You train your model on 50um/voxel, inference on 63um/voxel. Considering I use conv-based backbones (except for maxvit) that have some level of scale-invariance + have scale augs in validation, this might work.</p>\n<p>The second approach is to do rescaling. I believe the correct approach for rescaling is the following: </p>\n<pre><code> test_kidney == :\n    private_res = \n    public_res = \n\n    scale = private_res / public_res\n\n    d_original, h_original, w_original = test_kidney_image.shape\n    test_kidney_image = torch.tensor(test_kidney_image).view(, , d_original, h_original, w_original)\n    test_kidney_image = test_kidney_image.to(dtype=torch.float32)\n    test_kidney_image = torch.nn.functional.interpolate(test_kidney_image, (\n        (d_original*scale),\n        (h_original*scale),\n        (w_original*scale),\n    ), mode=).squeeze().numpy()\n</code></pre>\n<p>…</p>\n<pre><code>d_preds, h_preds, w_preds = preds_ensemble.shape \npreds_ensemble = preds_ensemble.view(, , d_preds, h_preds, w_preds)\npreds_ensemble = preds_ensemble.to(dtype=torch.float32)\n\npreds_ensemble = torch.nn.functional.interpolate(preds_ensemble, (\n    d_original,\n    h_original,\n    w_original,\n), mode=).squeeze()\n</code></pre>\n<p>So we do 3D resize instead of 2D one: re-scale image from 63um (private) to 50um (public + CV), compute predictions, and re-scale them back to 63um. Simply going for 2D would work as well, but theoretically you end up with different spatial and temporal resolutions in that case. </p>\n<p>This trick helped. To give a single point (I don't have much else): the same ensemble scores 0.634 on private without interpolation, and 0.670 - with interpolation. </p>\n<p>To be honest, I didn't think it would make that much difference. I tried the following experiment locally: </p>\n<ul>\n<li>Download kidney in 25um resolution. Compute predictions in 25um, interpolate them to 50um, compute metrics. This approach brought my 0.92 surface dice to 0.895. Which is quite good, considering we're talking about x2 interpolation in all 3 directions (that's 8 times less volume) and the fact that it's harder to detect small objects in smaller resolution.</li>\n<li>Download kidney in 25um resolution. Interpolate image to 50um, compute predictions, compute metrics. This approach essentially provided the same metrics as in the case of simply using 50um from organizers. </li>\n</ul>\n<p>So even though I didn't really think interpolation is that important, it also didn't hurt (I was afraid of interpolation artifacts), so I used it for both final subs. </p>\n<h3>Final subs</h3>\n<p>Both subs have an ensemble of 3 models, each inferenced on all 3 axes without TTA (TTA took too much time, and didn't really help on CV). </p>\n<ul>\n<li>First sub. CV: 0.84 (kidney_1), Public: 0.768. Private: 0.566<br>\n<code>Maxvit_ce_dice_focal</code> + <code>effnet_v2_s_ce_dice_focal</code> + <code>effnet_v2_m_ce_dice_focal</code> trained on kidney_3, validated on kidney_1. This approach didn't work that well on CV, and also on Public and Private. </li>\n<li>Second sub. CV: 0.923 (kidney_3), Public: 0.855. Private: 0.691<br>\n<code>Maxvit_ce_dice_focal</code> + <code>effnet_v2_s_ce_bounds_dice_focal</code> + <code>dpn_68_ce_bounds_twersky_focal</code>.</li>\n</ul>\n<p>Code:</p>\n<ul>\n<li>Inference notebook <a href=\"https://www.kaggle.com/code/ivanpan/final-submission/notebook\" target=\"_blank\">link</a></li>\n<li>Training code <a href=\"https://github.com/ivanpanshin/segment-vasculature-5th-place\" target=\"_blank\">link</a></li>\n</ul>",
      "rawMarkdown": "First of all, I want to express my deepest gratitude to the organizers of this competition. The reason I mostly compete in CV competitions is because I love it. Especially in medical competitions such as this one. So I'm really happy we managed to work with such a great technology (resolution is insane). \n\n### Validation \nInitially I thought it's gonna be a challenge validation-wise, since we simply don't have enough data for reliable validation. To make the validation without leakage and similar to test, I decided to train my pipelines in 2 ways: \n- Take kidney_1 as base for trainings, kidney_3 for validation\n- Take kidney_3 as base for trainings, kidney_1 for validation\n\nSince test is densely annotated, I wanted to compute metrics only on densely annotated kidneys, which eliminated kidney_2 from the discussion. \n\n### Data\nI always believe that data is the key. So I tried my hardest to utilize the additional datasets provided by organizers at [Human Organ Atlas](https://human-organ-atlas.esrf.eu). \n\nIn the end, I decided to use to following with the help of pseudo-labeling:\n- LADAF-2020-31 kidney\n- LADAF-2020-27 spleen\n\nIn other words, after examining the pseudo annotations on spleen, I realized that they are quite good and should serve as a good regularization method. \n\nAdditionally, I tried to use heart + brain + lung. However, my models make semi-accurate predictions for lung, but horrible for heart + brain. So in the end I decided to stick with kidney + spleen. \n\n### Pseudo annotations\nI would say there are 2 important points\n\nThe first one is don't pseudo annotate everything right away. In order to create full pseudo annotations, I run a 4-step process:\n- Train on kidney_1. Pseudo-annotate kidney_2\n- Train on kidney_1 + kidney_2. Pseudo-annotate the 2020-31 kidney.\n- Train on kidney_1 + kidney_2 + 2020-31-kidney. Pseudo-annotate 2020-27 spleen.\n- Train on kidney_1 + kidney_2 + 2020-31-kidney + 2020-27 spleen. \n\nThe second point is that don't use hard labels. In other words, don't apply thresholding to the predictions. Simply use soft labels (predictions are sigmoided to be in the range of [0,1]) for training. \n\n### Loss \n\nMy baseline go-to loss in semantic segmentation is `CE + Dice + Focal`. This worked quite well in this competition. However, since we have a surface metric, I wanted to weight the boundaries of masks more heavily. \n\n- What didn't work: losses I found in open-source repositories (like Hausdorff Distance loss).\n- What worked really well in terms of Surface Dice, FP and FN on validation: CE with x2 weights for boundaries. \n\nSo in the end I decided to use `CE_boundaries + Dice + Focal` for most of my models, and `CE_boundaries + Twersky + Focal` for a single model.\n\nTwersky was focusing more on FN rather than FP, but more on that in the next section. \n\n```python\nclass BoundDiceFocalLoss(torch.nn.modules.loss._Loss):\n    def __init__(self, bound_alpha=1.0, bound_weight, dice_weight, focal_weight):\n        super().__init__()\n        self.bound = EdgeEmphasisLoss(alpha=bound_alpha)\n        self.dice = smp.losses.DiceLoss(mode=\"binary\")\n        self.focal = smp.losses.FocalLoss(mode=\"binary\")\n        self.bound_weight = bound_weight\n        self.dice_weight = dice_weight\n        self.focal_weight = focal_weight\n\n    def forward(self, preds, gt, boundaries):\n        return (\n            self.bound_weight * self.bound(preds, gt, boundaries)\n            + self.dice_weight * self.dice(preds, gt)\n            + self.focal_weight * self.focal(preds, gt)\n        )\n\nclass EdgeEmphasisLoss(nn.Module):\n    def __init__(self, alpha=1.0):\n        super(EdgeEmphasisLoss, self).__init__()\n        self.alpha = alpha\n\n    def forward(self, inputs, targets, boundaries):\n        bce_loss = F.binary_cross_entropy_with_logits(inputs, targets, reduction=\"none\")\n\n        # Apply the edge weighting\n        weighted_loss = bce_loss * (1 + self.alpha * boundaries)\n\n        # Average over the batch\n        return weighted_loss.mean()\n```\n\n### Preprocessing\nAfter analyzing initial models and its errors, I realized that my biggest issue is FN, not FP. In other words, my models simply don't see some masks, mostly the small ones. \n\nSo I decided to increase the resolution of my trainings with crops from 512x512 to 1024x1024. However, after a couple of hours of training it hit me: that doesn't make much sense. By going from 512x512 to 1024x1024 I don't really increase resolution (each pixels holds the same real-world size), just the context, and 512x512 seemed like a big-enough context already. \n\nInstead, I decided to do the following: \n```python\n\nclass UnetUpscale(nn.Module):\n    def __init__(\n        self, encoder_name, decoder_use_batchnorm, in_channels, classes, upscale_factor, encoder_weights=\"imagenet\"\n    ):\n        super().__init__()\n        self.upscale_factor = upscale_factor\n\n        self.model = Unet(\n            encoder_weights=encoder_weights,\n            encoder_name=encoder_name,\n            decoder_use_batchnorm=decoder_use_batchnorm,\n            in_channels=in_channels,\n            classes=classes,\n        )\n\n    def forward(self, x):\n        x = torch.nn.functional.interpolate(\n            x, (x.shape[-2] * self.upscale_factor, x.shape[-1] * self.upscale_factor), mode=\"bilinear\"\n        )\n        x = self.model(x)\n        x = torch.nn.functional.interpolate(\n            x, (x.shape[-2] // self.upscale_factor, x.shape[-1] // self.upscale_factor), mode=\"bilinear\"\n        )\n        return x\n\n```\n\nThis approached worked really well and I could clearly see improvements both on CV, and LB.\n\n### Models\nI used only U-Net models from SMP with different backbones. Tried a lot of things, but for final ensembles decided to settle on the following:\n- effnet_v2_s\n- effnet_v2_m\n- maxvit_base\n- dpn68 \n\nMaxvit was trained on 512x512 crops, effent and dpn - on 512x512 with x2 interpolation. Crops were used from xy, xz, and yz axes. During inference, I use the same crops resolution with overlaps of crops_size / 2 (so that's 256). In other words, sliding window approach.\n\nAugmentation were medium-level in terms of intensity. \n\n```python\nreturn A.Compose(\n    [\n        A.ShiftScaleRotate(\n            p=0.7,\n            shift_limit_x=(-0.1, 0.1),\n            shift_limit_y=(-0.1, 0.1),\n            scale_limit=(-0.25, 0.25),\n            rotate_limit=(-25, 25),\n            border_mode=cv2.BORDER_CONSTANT,\n            # rotate_method=\"largest_box\",\n        ),\n        A.RandomBrightnessContrast(\n            brightness_limit=(-0.25, 0.25),\n            contrast_limit=(-0.25, 0.25),\n            p=0.5,\n        ),\n        A.HorizontalFlip(),\n        A.VerticalFlip(),\n        A.OneOf(\n            [\n                A.GridDistortion(border_mode=cv2.BORDER_CONSTANT, distort_limit=0.1),\n                A.ElasticTransform(border_mode=cv2.BORDER_CONSTANT),\n            ],\n            p=0.2,\n        ),\n        AT.ToTensorV2(),\n    ],\n    )\n\n```\n\n### Post processing\nI tried to use cc3d to remove small objects, it made weak models better, but no difference for ensemble.\n\n### Private resolution\nNow, this part is really tricky. My huge thanks to the organizers for announcing the test resolutions. It sincerely warms my heart to see organizers interact with participants that much here on the forum. Really, thank you. \n\nOne approach is not to do anything. You train your model on 50um/voxel, inference on 63um/voxel. Considering I use conv-based backbones (except for maxvit) that have some level of scale-invariance + have scale augs in validation, this might work.\n\nThe second approach is to do rescaling. I believe the correct approach for rescaling is the following: \n\n```python\nif test_kidney == 6:\n    private_res = 63.08\n    public_res = 50.0\n            \n    scale = private_res / public_res\n            \n    d_original, h_original, w_original = test_kidney_image.shape\n    test_kidney_image = torch.tensor(test_kidney_image).view(1, 1, d_original, h_original, w_original)\n    test_kidney_image = test_kidney_image.to(dtype=torch.float32)\n    test_kidney_image = torch.nn.functional.interpolate(test_kidney_image, (\n        int(d_original*scale),\n        int(h_original*scale),\n        int(w_original*scale),\n    ), mode='trilinear').squeeze().numpy()\n\n```\n...\n```python\n\nd_preds, h_preds, w_preds = preds_ensemble.shape \npreds_ensemble = preds_ensemble.view(1, 1, d_preds, h_preds, w_preds)\npreds_ensemble = preds_ensemble.to(dtype=torch.float32)\n            \npreds_ensemble = torch.nn.functional.interpolate(preds_ensemble, (\n    d_original,\n    h_original,\n    w_original,\n), mode='trilinear').squeeze()\n```\n\nSo we do 3D resize instead of 2D one: re-scale image from 63um (private) to 50um (public + CV), compute predictions, and re-scale them back to 63um. Simply going for 2D would work as well, but theoretically you end up with different spatial and temporal resolutions in that case. \n\nThis trick helped. To give a single point (I don't have much else): the same ensemble scores 0.634 on private without interpolation, and 0.670 - with interpolation. \n\nTo be honest, I didn't think it would make that much difference. I tried the following experiment locally: \n- Download kidney in 25um resolution. Compute predictions in 25um, interpolate them to 50um, compute metrics. This approach brought my 0.92 surface dice to 0.895. Which is quite good, considering we're talking about x2 interpolation in all 3 directions (that's 8 times less volume) and the fact that it's harder to detect small objects in smaller resolution.\n- Download kidney in 25um resolution. Interpolate image to 50um, compute predictions, compute metrics. This approach essentially provided the same metrics as in the case of simply using 50um from organizers. \n\nSo even though I didn't really think interpolation is that important, it also didn't hurt (I was afraid of interpolation artifacts), so I used it for both final subs. \n\n### Final subs\nBoth subs have an ensemble of 3 models, each inferenced on all 3 axes without TTA (TTA took too much time, and didn't really help on CV). \n- First sub. CV: 0.84 (kidney_1), Public: 0.768. Private: 0.566\n`Maxvit_ce_dice_focal` + `effnet_v2_s_ce_dice_focal` + `effnet_v2_m_ce_dice_focal` trained on kidney_3, validated on kidney_1. This approach didn't work that well on CV, and also on Public and Private. \n- Second sub. CV: 0.923 (kidney_3), Public: 0.855. Private: 0.691\n`Maxvit_ce_dice_focal` + `effnet_v2_s_ce_bounds_dice_focal` + `dpn_68_ce_bounds_twersky_focal`.\n\nCode:\n- Inference notebook [link](https://www.kaggle.com/code/ivanpan/final-submission/notebook)\n- Training code [link](https://github.com/ivanpanshin/segment-vasculature-5th-place)",
      "votes": null
    },
    {
      "id": "2641916",
      "postDate": "02/07/2024 18:44:26",
      "content": "<p>Awesome! Have you tried training other models like FPN, Unet++ ?</p>",
      "rawMarkdown": "Awesome! Have you tried training other models like FPN, Unet++ ?",
      "votes": null
    },
    {
      "id": "2641920",
      "postDate": "02/07/2024 18:48:14",
      "content": "<p>Thanks! </p>\n<p>Nope. I think it could have helped in the ensemble, but in my experience U-Net with skip-connections is pretty much always enough. Which is quite crazy considering it's pretty much 10 years old. </p>",
      "rawMarkdown": "Thanks! \n\nNope. I think it could have helped in the ensemble, but in my experience U-Net with skip-connections is pretty much always enough. Which is quite crazy considering it's pretty much 10 years old.",
      "votes": null
    },
    {
      "id": "2642139",
      "postDate": "02/08/2024 00:15:11",
      "content": "<p>Excellent solution, very insightful, congratulations! Thanks for the write up. What values did you use for bound_alpha, bound_weight, dice_weight, focal_weight in your loss function?</p>",
      "rawMarkdown": "Excellent solution, very insightful, congratulations! Thanks for the write up. What values did you use for bound_alpha, bound_weight, dice_weight, focal_weight in your loss function?",
      "votes": null
    },
    {
      "id": "2642596",
      "postDate": "02/08/2024 09:16:43",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ivanpan\" target=\"_blank\">@ivanpan</a> , Congratulations on this solo gold(prize) and grandmaster! </p>\n<p>I think the most important part in your solution is </p>\n<ol>\n<li><p>boundary CE loss and </p></li>\n<li><p>pseudo label on external data. </p></li>\n</ol>\n<p>May you show more ablation experiments results on these two key points? I am really interested to it. </p>",
      "rawMarkdown": "Hi @ivanpan , Congratulations on this solo gold(prize) and grandmaster! \n\nI think the most important part in your solution is \n\n1. boundary CE loss and \n\n2. pseudo label on external data. \n\nMay you show more ablation experiments results on these two key points? I am really interested to it.",
      "votes": null
    },
    {
      "id": "2642739",
      "postDate": "02/08/2024 11:28:14",
      "content": "<p>Congratulations on securing 5th position in this competition. Thanks for sharing the details of your approach. </p>",
      "rawMarkdown": "Congratulations on securing 5th position in this competition. Thanks for sharing the details of your approach.",
      "votes": null
    },
    {
      "id": "2642753",
      "postDate": "02/08/2024 11:41:19",
      "content": "<p>Hm. I can. I don't have a lot of data points, but I will share the best ones I can. </p>\n<ol>\n<li><p>I have a single data point here that I still have access to. <code>CE + Dice + Focal</code>: SD: 91.5, FP: 120K, FN: 160K. <code>CE_boundaries + Dice + Focal</code>: 92.3, FP: 120K, FN: 140K. These metrics are for kidney_3 used as validation. </p></li>\n<li><p>This is tricky. If I utilize pseudo label on external data based on trainings from kidney_1, then I don't see a lot of improvement either on CV, or LB. The killer feature is std. Even though you can get the same metrics on CV, and LB, by training with pseudo-annotations LB scores become muuuch more consistent. For example, my original trainings (without pseudo) score somewhere between 0.83-0.88 on Public. However, after I add pseudo to training, I don't think my LB was even below 0.87 for these models. </p></li>\n</ol>\n<p>Additionally, if I start training from kidney_3 (not kidney_1), the difference on CV is very evident (still doesn't work very well on Public or Private though).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2755695%2F678efcccf7ae440d9f77bbace7323466%2FScreenshot%202024-02-08%20at%2013.39.04.png?generation=1707392353733609&amp;alt=media\"></p>\n<p>Green - kidney_3, yellow - kidney_3 + kidney_2, blue - kidney_3 + kidney_2 + kidney_external </p>",
      "rawMarkdown": "Hm. I can. I don't have a lot of data points, but I will share the best ones I can. \n\n1. I have a single data point here that I still have access to. `CE + Dice + Focal`: SD: 91.5, FP: 120K, FN: 160K. `CE_boundaries + Dice + Focal`: 92.3, FP: 120K, FN: 140K. These metrics are for kidney_3 used as validation. \n\n2. This is tricky. If I utilize pseudo label on external data based on trainings from kidney_1, then I don't see a lot of improvement either on CV, or LB. The killer feature is std. Even though you can get the same metrics on CV, and LB, by training with pseudo-annotations LB scores become muuuch more consistent. For example, my original trainings (without pseudo) score somewhere between 0.83-0.88 on Public. However, after I add pseudo to training, I don't think my LB was even below 0.87 for these models. \n\nAdditionally, if I start training from kidney_3 (not kidney_1), the difference on CV is very evident (still doesn't work very well on Public or Private though).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2755695%2F678efcccf7ae440d9f77bbace7323466%2FScreenshot%202024-02-08%20at%2013.39.04.png?generation=1707392353733609&alt=media)\n\nGreen - kidney_3, yellow - kidney_3 + kidney_2, blue - kidney_3 + kidney_2 + kidney_external",
      "votes": null
    },
    {
      "id": "2642760",
      "postDate": "02/08/2024 11:46:31",
      "content": "<p>Sure thing. </p>\n<pre><code>bound_alpha=, \nbound_weight=, \ndice_weight=, \nfocal_weight=\n</code></pre>\n<p>Let me elaborate. The reason for bound_alpha=1 is to give boundaries twice as much importance as to other areas of the mask. I tried to set it higher (so even more importance to boundaries), but it was worse on CV and LB.</p>\n<p>The reason for other weights is the absolute value of losses during training. If you train a model with CE, and then Dice (separately) you will notice that on average Dice loss values are 2 orders of magnitude higher (that is, 100 times bigger). So I wanted to combine losses in such a way that their values are on the same scale. Additionally, CE and Focal have the same scale (expectedly, since Focal is just CE with bells and whistles), but I simply didn't want to give too much importance to Focal.</p>",
      "rawMarkdown": "Sure thing. \n\n```python\nbound_alpha=1.0, \nbound_weight=0.9, \ndice_weight=0.01, \nfocal_weight=0.09\n```\n\nLet me elaborate. The reason for bound_alpha=1 is to give boundaries twice as much importance as to other areas of the mask. I tried to set it higher (so even more importance to boundaries), but it was worse on CV and LB.\n\nThe reason for other weights is the absolute value of losses during training. If you train a model with CE, and then Dice (separately) you will notice that on average Dice loss values are 2 orders of magnitude higher (that is, 100 times bigger). So I wanted to combine losses in such a way that their values are on the same scale. Additionally, CE and Focal have the same scale (expectedly, since Focal is just CE with bells and whistles), but I simply didn't want to give too much importance to Focal.",
      "votes": null
    },
    {
      "id": "2643551",
      "postDate": "02/08/2024 22:29:24",
      "content": "<p>Thank you for sharing the weights for your custom loss code. I am attempting to incorporate some of the winning ideas into my solution and see if the score improves.</p>",
      "rawMarkdown": "Thank you for sharing the weights for your custom loss code. I am attempting to incorporate some of the winning ideas into my solution and see if the score improves.",
      "votes": null
    },
    {
      "id": "2643594",
      "postDate": "02/08/2024 23:57:56",
      "content": "<p>i did not try this, but instead of pesudo label, one can use denoising as pretraining</p>\n<p>Decoder Denoising Pretraining for Semantic Segmentation<br>\n<a href=\"https://arxiv.org/abs/2205.11423\" target=\"_blank\">https://arxiv.org/abs/2205.11423</a><br>\n|<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7a19a5a23c5a10197c4207a0e8e60f04%2FSelection_999(4970).png?generation=1707436670468343&amp;alt=media\"> </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2a9dafb8f997aa4ec06e67eb5471e49a%2FSelection_999(4972).png?generation=1707436823778926&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc05c94dfb42e51afc50d513480c9e9a3%2FSelection_999(4971).png?generation=1707436834634540&amp;alt=media\"></p>",
      "rawMarkdown": "i did not try this, but instead of pesudo label, one can use denoising as pretraining\n\nDecoder Denoising Pretraining for Semantic Segmentation\nhttps://arxiv.org/abs/2205.11423\n|![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7a19a5a23c5a10197c4207a0e8e60f04%2FSelection_999(4970).png?generation=1707436670468343&alt=media) \n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2a9dafb8f997aa4ec06e67eb5471e49a%2FSelection_999(4972).png?generation=1707436823778926&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc05c94dfb42e51afc50d513480c9e9a3%2FSelection_999(4971).png?generation=1707436834634540&alt=media)",
      "votes": null
    },
    {
      "id": "2644697",
      "postDate": "02/09/2024 16:18:21",
      "content": "<p>Thank you for this great reply. The comparison of pseudo label is similar to mine and make sense. </p>",
      "rawMarkdown": "Thank you for this great reply. The comparison of pseudo label is similar to mine and make sense.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2641916,
      "author_name": "romach",
      "author_url": "",
      "post_date": "02/07/2024 18:44:26",
      "content": "<p>Awesome! Have you tried training other models like FPN, Unet++ ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2641920,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "02/07/2024 18:48:14",
          "content": "<p>Thanks! </p>\n<p>Nope. I think it could have helped in the ensemble, but in my experience U-Net with skip-connections is pretty much always enough. Which is quite crazy considering it's pretty much 10 years old. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2642139,
      "author_name": "velangovan",
      "author_url": "",
      "post_date": "02/08/2024 00:15:11",
      "content": "<p>Excellent solution, very insightful, congratulations! Thanks for the write up. What values did you use for bound_alpha, bound_weight, dice_weight, focal_weight in your loss function?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2642760,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "02/08/2024 11:46:31",
          "content": "<p>Sure thing. </p>\n<pre><code>bound_alpha=, \nbound_weight=, \ndice_weight=, \nfocal_weight=\n</code></pre>\n<p>Let me elaborate. The reason for bound_alpha=1 is to give boundaries twice as much importance as to other areas of the mask. I tried to set it higher (so even more importance to boundaries), but it was worse on CV and LB.</p>\n<p>The reason for other weights is the absolute value of losses during training. If you train a model with CE, and then Dice (separately) you will notice that on average Dice loss values are 2 orders of magnitude higher (that is, 100 times bigger). So I wanted to combine losses in such a way that their values are on the same scale. Additionally, CE and Focal have the same scale (expectedly, since Focal is just CE with bells and whistles), but I simply didn't want to give too much importance to Focal.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2643551,
              "author_name": "velangovan",
              "author_url": "",
              "post_date": "02/08/2024 22:29:24",
              "content": "<p>Thank you for sharing the weights for your custom loss code. I am attempting to incorporate some of the winning ideas into my solution and see if the score improves.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2642596,
      "author_name": "forcewithme",
      "author_url": "",
      "post_date": "02/08/2024 09:16:43",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ivanpan\" target=\"_blank\">@ivanpan</a> , Congratulations on this solo gold(prize) and grandmaster! </p>\n<p>I think the most important part in your solution is </p>\n<ol>\n<li><p>boundary CE loss and </p></li>\n<li><p>pseudo label on external data. </p></li>\n</ol>\n<p>May you show more ablation experiments results on these two key points? I am really interested to it. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2642753,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "02/08/2024 11:41:19",
          "content": "<p>Hm. I can. I don't have a lot of data points, but I will share the best ones I can. </p>\n<ol>\n<li><p>I have a single data point here that I still have access to. <code>CE + Dice + Focal</code>: SD: 91.5, FP: 120K, FN: 160K. <code>CE_boundaries + Dice + Focal</code>: 92.3, FP: 120K, FN: 140K. These metrics are for kidney_3 used as validation. </p></li>\n<li><p>This is tricky. If I utilize pseudo label on external data based on trainings from kidney_1, then I don't see a lot of improvement either on CV, or LB. The killer feature is std. Even though you can get the same metrics on CV, and LB, by training with pseudo-annotations LB scores become muuuch more consistent. For example, my original trainings (without pseudo) score somewhere between 0.83-0.88 on Public. However, after I add pseudo to training, I don't think my LB was even below 0.87 for these models. </p></li>\n</ol>\n<p>Additionally, if I start training from kidney_3 (not kidney_1), the difference on CV is very evident (still doesn't work very well on Public or Private though).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2755695%2F678efcccf7ae440d9f77bbace7323466%2FScreenshot%202024-02-08%20at%2013.39.04.png?generation=1707392353733609&amp;alt=media\"></p>\n<p>Green - kidney_3, yellow - kidney_3 + kidney_2, blue - kidney_3 + kidney_2 + kidney_external </p>",
          "votes": null,
          "replies": [
            {
              "id": 2644697,
              "author_name": "forcewithme",
              "author_url": "",
              "post_date": "02/09/2024 16:18:21",
              "content": "<p>Thank you for this great reply. The comparison of pseudo label is similar to mine and make sense. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2642739,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "02/08/2024 11:28:14",
      "content": "<p>Congratulations on securing 5th position in this competition. Thanks for sharing the details of your approach. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2643594,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/08/2024 23:57:56",
      "content": "<p>i did not try this, but instead of pesudo label, one can use denoising as pretraining</p>\n<p>Decoder Denoising Pretraining for Semantic Segmentation<br>\n<a href=\"https://arxiv.org/abs/2205.11423\" target=\"_blank\">https://arxiv.org/abs/2205.11423</a><br>\n|<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7a19a5a23c5a10197c4207a0e8e60f04%2FSelection_999(4970).png?generation=1707436670468343&amp;alt=media\"> </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2a9dafb8f997aa4ec06e67eb5471e49a%2FSelection_999(4972).png?generation=1707436823778926&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc05c94dfb42e51afc50d513480c9e9a3%2FSelection_999(4971).png?generation=1707436834634540&amp;alt=media\"></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2641911": "First of all, I want to express my deepest gratitude to the organizers of this competition. The reason I mostly compete in CV competitions is because I love it. Especially in medical competitions such as this one. So I'm really happy we managed to work with such a great technology (resolution is insane). \n\n### Validation \nInitially I thought it's gonna be a challenge validation-wise, since we simply don't have enough data for reliable validation. To make the validation without leakage and similar to test, I decided to train my pipelines in 2 ways: \n- Take kidney_1 as base for trainings, kidney_3 for validation\n- Take kidney_3 as base for trainings, kidney_1 for validation\n\nSince test is densely annotated, I wanted to compute metrics only on densely annotated kidneys, which eliminated kidney_2 from the discussion. \n\n### Data\nI always believe that data is the key. So I tried my hardest to utilize the additional datasets provided by organizers at [Human Organ Atlas](https://human-organ-atlas.esrf.eu). \n\nIn the end, I decided to use to following with the help of pseudo-labeling:\n- LADAF-2020-31 kidney\n- LADAF-2020-27 spleen\n\nIn other words, after examining the pseudo annotations on spleen, I realized that they are quite good and should serve as a good regularization method. \n\nAdditionally, I tried to use heart + brain + lung. However, my models make semi-accurate predictions for lung, but horrible for heart + brain. So in the end I decided to stick with kidney + spleen. \n\n### Pseudo annotations\nI would say there are 2 important points\n\nThe first one is don't pseudo annotate everything right away. In order to create full pseudo annotations, I run a 4-step process:\n- Train on kidney_1. Pseudo-annotate kidney_2\n- Train on kidney_1 + kidney_2. Pseudo-annotate the 2020-31 kidney.\n- Train on kidney_1 + kidney_2 + 2020-31-kidney. Pseudo-annotate 2020-27 spleen.\n- Train on kidney_1 + kidney_2 + 2020-31-kidney + 2020-27 spleen. \n\nThe second point is that don't use hard labels. In other words, don't apply thresholding to the predictions. Simply use soft labels (predictions are sigmoided to be in the range of [0,1]) for training. \n\n### Loss \n\nMy baseline go-to loss in semantic segmentation is `CE + Dice + Focal`. This worked quite well in this competition. However, since we have a surface metric, I wanted to weight the boundaries of masks more heavily. \n\n- What didn't work: losses I found in open-source repositories (like Hausdorff Distance loss).\n- What worked really well in terms of Surface Dice, FP and FN on validation: CE with x2 weights for boundaries. \n\nSo in the end I decided to use `CE_boundaries + Dice + Focal` for most of my models, and `CE_boundaries + Twersky + Focal` for a single model.\n\nTwersky was focusing more on FN rather than FP, but more on that in the next section. \n\n```python\nclass BoundDiceFocalLoss(torch.nn.modules.loss._Loss):\n    def __init__(self, bound_alpha=1.0, bound_weight, dice_weight, focal_weight):\n        super().__init__()\n        self.bound = EdgeEmphasisLoss(alpha=bound_alpha)\n        self.dice = smp.losses.DiceLoss(mode=\"binary\")\n        self.focal = smp.losses.FocalLoss(mode=\"binary\")\n        self.bound_weight = bound_weight\n        self.dice_weight = dice_weight\n        self.focal_weight = focal_weight\n\n    def forward(self, preds, gt, boundaries):\n        return (\n            self.bound_weight * self.bound(preds, gt, boundaries)\n            + self.dice_weight * self.dice(preds, gt)\n            + self.focal_weight * self.focal(preds, gt)\n        )\n\nclass EdgeEmphasisLoss(nn.Module):\n    def __init__(self, alpha=1.0):\n        super(EdgeEmphasisLoss, self).__init__()\n        self.alpha = alpha\n\n    def forward(self, inputs, targets, boundaries):\n        bce_loss = F.binary_cross_entropy_with_logits(inputs, targets, reduction=\"none\")\n\n        # Apply the edge weighting\n        weighted_loss = bce_loss * (1 + self.alpha * boundaries)\n\n        # Average over the batch\n        return weighted_loss.mean()\n```\n\n### Preprocessing\nAfter analyzing initial models and its errors, I realized that my biggest issue is FN, not FP. In other words, my models simply don't see some masks, mostly the small ones. \n\nSo I decided to increase the resolution of my trainings with crops from 512x512 to 1024x1024. However, after a couple of hours of training it hit me: that doesn't make much sense. By going from 512x512 to 1024x1024 I don't really increase resolution (each pixels holds the same real-world size), just the context, and 512x512 seemed like a big-enough context already. \n\nInstead, I decided to do the following: \n```python\n\nclass UnetUpscale(nn.Module):\n    def __init__(\n        self, encoder_name, decoder_use_batchnorm, in_channels, classes, upscale_factor, encoder_weights=\"imagenet\"\n    ):\n        super().__init__()\n        self.upscale_factor = upscale_factor\n\n        self.model = Unet(\n            encoder_weights=encoder_weights,\n            encoder_name=encoder_name,\n            decoder_use_batchnorm=decoder_use_batchnorm,\n            in_channels=in_channels,\n            classes=classes,\n        )\n\n    def forward(self, x):\n        x = torch.nn.functional.interpolate(\n            x, (x.shape[-2] * self.upscale_factor, x.shape[-1] * self.upscale_factor), mode=\"bilinear\"\n        )\n        x = self.model(x)\n        x = torch.nn.functional.interpolate(\n            x, (x.shape[-2] // self.upscale_factor, x.shape[-1] // self.upscale_factor), mode=\"bilinear\"\n        )\n        return x\n\n```\n\nThis approached worked really well and I could clearly see improvements both on CV, and LB.\n\n### Models\nI used only U-Net models from SMP with different backbones. Tried a lot of things, but for final ensembles decided to settle on the following:\n- effnet_v2_s\n- effnet_v2_m\n- maxvit_base\n- dpn68 \n\nMaxvit was trained on 512x512 crops, effent and dpn - on 512x512 with x2 interpolation. Crops were used from xy, xz, and yz axes. During inference, I use the same crops resolution with overlaps of crops_size / 2 (so that's 256). In other words, sliding window approach.\n\nAugmentation were medium-level in terms of intensity. \n\n```python\nreturn A.Compose(\n    [\n        A.ShiftScaleRotate(\n            p=0.7,\n            shift_limit_x=(-0.1, 0.1),\n            shift_limit_y=(-0.1, 0.1),\n            scale_limit=(-0.25, 0.25),\n            rotate_limit=(-25, 25),\n            border_mode=cv2.BORDER_CONSTANT,\n            # rotate_method=\"largest_box\",\n        ),\n        A.RandomBrightnessContrast(\n            brightness_limit=(-0.25, 0.25),\n            contrast_limit=(-0.25, 0.25),\n            p=0.5,\n        ),\n        A.HorizontalFlip(),\n        A.VerticalFlip(),\n        A.OneOf(\n            [\n                A.GridDistortion(border_mode=cv2.BORDER_CONSTANT, distort_limit=0.1),\n                A.ElasticTransform(border_mode=cv2.BORDER_CONSTANT),\n            ],\n            p=0.2,\n        ),\n        AT.ToTensorV2(),\n    ],\n    )\n\n```\n\n### Post processing\nI tried to use cc3d to remove small objects, it made weak models better, but no difference for ensemble.\n\n### Private resolution\nNow, this part is really tricky. My huge thanks to the organizers for announcing the test resolutions. It sincerely warms my heart to see organizers interact with participants that much here on the forum. Really, thank you. \n\nOne approach is not to do anything. You train your model on 50um/voxel, inference on 63um/voxel. Considering I use conv-based backbones (except for maxvit) that have some level of scale-invariance + have scale augs in validation, this might work.\n\nThe second approach is to do rescaling. I believe the correct approach for rescaling is the following: \n\n```python\nif test_kidney == 6:\n    private_res = 63.08\n    public_res = 50.0\n            \n    scale = private_res / public_res\n            \n    d_original, h_original, w_original = test_kidney_image.shape\n    test_kidney_image = torch.tensor(test_kidney_image).view(1, 1, d_original, h_original, w_original)\n    test_kidney_image = test_kidney_image.to(dtype=torch.float32)\n    test_kidney_image = torch.nn.functional.interpolate(test_kidney_image, (\n        int(d_original*scale),\n        int(h_original*scale),\n        int(w_original*scale),\n    ), mode='trilinear').squeeze().numpy()\n\n```\n...\n```python\n\nd_preds, h_preds, w_preds = preds_ensemble.shape \npreds_ensemble = preds_ensemble.view(1, 1, d_preds, h_preds, w_preds)\npreds_ensemble = preds_ensemble.to(dtype=torch.float32)\n            \npreds_ensemble = torch.nn.functional.interpolate(preds_ensemble, (\n    d_original,\n    h_original,\n    w_original,\n), mode='trilinear').squeeze()\n```\n\nSo we do 3D resize instead of 2D one: re-scale image from 63um (private) to 50um (public + CV), compute predictions, and re-scale them back to 63um. Simply going for 2D would work as well, but theoretically you end up with different spatial and temporal resolutions in that case. \n\nThis trick helped. To give a single point (I don't have much else): the same ensemble scores 0.634 on private without interpolation, and 0.670 - with interpolation. \n\nTo be honest, I didn't think it would make that much difference. I tried the following experiment locally: \n- Download kidney in 25um resolution. Compute predictions in 25um, interpolate them to 50um, compute metrics. This approach brought my 0.92 surface dice to 0.895. Which is quite good, considering we're talking about x2 interpolation in all 3 directions (that's 8 times less volume) and the fact that it's harder to detect small objects in smaller resolution.\n- Download kidney in 25um resolution. Interpolate image to 50um, compute predictions, compute metrics. This approach essentially provided the same metrics as in the case of simply using 50um from organizers. \n\nSo even though I didn't really think interpolation is that important, it also didn't hurt (I was afraid of interpolation artifacts), so I used it for both final subs. \n\n### Final subs\nBoth subs have an ensemble of 3 models, each inferenced on all 3 axes without TTA (TTA took too much time, and didn't really help on CV). \n- First sub. CV: 0.84 (kidney_1), Public: 0.768. Private: 0.566\n`Maxvit_ce_dice_focal` + `effnet_v2_s_ce_dice_focal` + `effnet_v2_m_ce_dice_focal` trained on kidney_3, validated on kidney_1. This approach didn't work that well on CV, and also on Public and Private. \n- Second sub. CV: 0.923 (kidney_3), Public: 0.855. Private: 0.691\n`Maxvit_ce_dice_focal` + `effnet_v2_s_ce_bounds_dice_focal` + `dpn_68_ce_bounds_twersky_focal`.\n\nCode:\n- Inference notebook [link](https://www.kaggle.com/code/ivanpan/final-submission/notebook)\n- Training code [link](https://github.com/ivanpanshin/segment-vasculature-5th-place)",
    "2641916": "Awesome! Have you tried training other models like FPN, Unet++ ?",
    "2641920": "Thanks! \n\nNope. I think it could have helped in the ensemble, but in my experience U-Net with skip-connections is pretty much always enough. Which is quite crazy considering it's pretty much 10 years old.",
    "2642139": "Excellent solution, very insightful, congratulations! Thanks for the write up. What values did you use for bound_alpha, bound_weight, dice_weight, focal_weight in your loss function?",
    "2642596": "Hi @ivanpan , Congratulations on this solo gold(prize) and grandmaster! \n\nI think the most important part in your solution is \n\n1. boundary CE loss and \n\n2. pseudo label on external data. \n\nMay you show more ablation experiments results on these two key points? I am really interested to it.",
    "2642739": "Congratulations on securing 5th position in this competition. Thanks for sharing the details of your approach.",
    "2642753": "Hm. I can. I don't have a lot of data points, but I will share the best ones I can. \n\n1. I have a single data point here that I still have access to. `CE + Dice + Focal`: SD: 91.5, FP: 120K, FN: 160K. `CE_boundaries + Dice + Focal`: 92.3, FP: 120K, FN: 140K. These metrics are for kidney_3 used as validation. \n\n2. This is tricky. If I utilize pseudo label on external data based on trainings from kidney_1, then I don't see a lot of improvement either on CV, or LB. The killer feature is std. Even though you can get the same metrics on CV, and LB, by training with pseudo-annotations LB scores become muuuch more consistent. For example, my original trainings (without pseudo) score somewhere between 0.83-0.88 on Public. However, after I add pseudo to training, I don't think my LB was even below 0.87 for these models. \n\nAdditionally, if I start training from kidney_3 (not kidney_1), the difference on CV is very evident (still doesn't work very well on Public or Private though).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2755695%2F678efcccf7ae440d9f77bbace7323466%2FScreenshot%202024-02-08%20at%2013.39.04.png?generation=1707392353733609&alt=media)\n\nGreen - kidney_3, yellow - kidney_3 + kidney_2, blue - kidney_3 + kidney_2 + kidney_external",
    "2642760": "Sure thing. \n\n```python\nbound_alpha=1.0, \nbound_weight=0.9, \ndice_weight=0.01, \nfocal_weight=0.09\n```\n\nLet me elaborate. The reason for bound_alpha=1 is to give boundaries twice as much importance as to other areas of the mask. I tried to set it higher (so even more importance to boundaries), but it was worse on CV and LB.\n\nThe reason for other weights is the absolute value of losses during training. If you train a model with CE, and then Dice (separately) you will notice that on average Dice loss values are 2 orders of magnitude higher (that is, 100 times bigger). So I wanted to combine losses in such a way that their values are on the same scale. Additionally, CE and Focal have the same scale (expectedly, since Focal is just CE with bells and whistles), but I simply didn't want to give too much importance to Focal.",
    "2643551": "Thank you for sharing the weights for your custom loss code. I am attempting to incorporate some of the winning ideas into my solution and see if the score improves.",
    "2643594": "i did not try this, but instead of pesudo label, one can use denoising as pretraining\n\nDecoder Denoising Pretraining for Semantic Segmentation\nhttps://arxiv.org/abs/2205.11423\n|![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7a19a5a23c5a10197c4207a0e8e60f04%2FSelection_999(4970).png?generation=1707436670468343&alt=media) \n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2a9dafb8f997aa4ec06e67eb5471e49a%2FSelection_999(4972).png?generation=1707436823778926&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc05c94dfb42e51afc50d513480c9e9a3%2FSelection_999(4971).png?generation=1707436834634540&alt=media)",
    "2644697": "Thank you for this great reply. The comparison of pseudo label is similar to mine and make sense."
  },
  "source": "meta"
}