{
  "id": 475090,
  "title": "🥈24th Place Solution(A potential solution for achieving private 0.70)",
  "url": "/competitions/blood-vessel-segmentation/discussion/475090",
  "author_name": "siwooyong",
  "post_date": "2024-02-07T03:24:30.267000",
  "votes": 19,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>TLDR</h1>\n<p>The key factors in enhancing the model's robust performance included <strong>Tversky loss</strong>, <strong>increased inference size</strong>, and <strong>resolution augmentation</strong>. These elements ultimately played a significant role in the model's survival during shakeups. The final model is an ensemble composed of 2D U-Net models based on the RegNetY-016 architecture</p>\n<h1>Interesting Point</h1>\n<p>It was a truly challenging competition to secure a reliable validation set (which, unfortunately, I couldn't achieve). Upon reviewing the results, I discovered a significant difference between the public and private sets. Surprisingly, there were submissions from the past that would have made it into the gold zone. It's astonishing, considering I didn't think much of those submissions and didn't end up submitting them anyway.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8251891%2Fe5a2fef330678d164bc8ff54d06c3e34%2F2024-02-07%2021-58-54.png?generation=1707310804960891&amp;alt=media\"></p>\n<p>In my experiments, I found that giving a higher weight to the positive class in the Tversky loss (with a larger beta) resulted in better performance on the private leaderboard. However, for my final submission, where I trained with a smaller beta value in the Tversky loss, the model performed well on the public leaderboard but surprisingly poorly on the private leaderboard. <strong>This discrepancy may be attributed to the higher resolution of the private dataset, which likely contained more detailed input information. Consequently, it suggests that a lower threshold for the model logits might have been more appropriate in this context. Furthermore, scaling the image (1.2, 1.5, …) during inference seemed to dilute input information, reducing the impact of resolution.</strong></p>\n<h1>Data Processing</h1>\n<p>Three consecutive slices were used as input for the model, and training was performed with the original image size. Despite experimenting with an increased number of slices (5, 7, 9…), the results revealed a decrease in performance.</p>\n<p>In the early stages of the competition, I trained the model by resizing the data to a specific size. However, as I mentioned in the <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/463121\" target=\"_blank\">discussion</a>, the Resize function in the Albumentations library applies nearest interpolation to masks, causing significant noise in fine labels and resulting in a notable decrease in performance. Therefore, I opted to use the original image size.</p>\n<h1>Augmentation</h1>\n<p>Awaring of the different resolutions between public and private data, I aimed to create a model robust to resolution variations. Additionally, understanding the existence of resizing during the binning process, I employed blur augmentation. </p>\n<pre><code>def blur_augmentation(x):\n    h, w,  = x.shape\n     = ..uniform(, )\n\n    x = A.Resize(int(h*), int(w*))(=x)['']\n    x = A.Resize(h, w)(=x)['']\n     x\n</code></pre>\n<p>Moreover, since the channels were constructed by stacking depth, I applied the following augmentations.</p>\n<pre><code>def channel_augmentation(x, =0.5, =3):\n    assert x.shape[2]==n_channel\n     np.random.rand()&lt;prob:\n        x = np.flip(x, =2)\n    return x\n</code></pre>\n<p>Finally, given the inherent noise in the data annotations, I implemented strong cutout augmentation to prevent overfitting.</p>\n<pre><code>A.Cutout(=8, =128, =128, =0.8),\n</code></pre>\n<p>These approaches effectively contributed to performance improvement in both CV and LB.</p>\n<h1>Model</h1>\n<p>The model, like most others, employed a 2D U-Net architecture. For the CNN backbone, I utilized the lightweight model RegNetY-016. Other than that, the settings remained consistent with the default values in SMP (Segmentation Models PyTorch).</p>\n<p>Despite investing significant time in developing a 3D-based model, it failed to demonstrate notable score improvements in both CV and LB. Due to resource constraints, I shifted my focus to a 2D model.</p>\n<pre><code>class CustomModel(nn.Module):\n    def __init__(self):\n        super(CustomModel, self).__init__()\n\n        self.n_classes = 1\n        self.in_chans = 3\n\n        self.encoder = timm.create_model(\n            ,\n            =,\n            =,\n            =self.in_chans,\n        )\n        encoder_channels = tuple(\n            [self.in_chans]\n            + [\n                self.encoder.feature_info[i][]\n                 i  range(len(self.encoder.feature_info))\n            ]\n        )\n        self.decoder = UnetDecoder(\n            =encoder_channels,\n            decoder_channels=(256, 128, 64, 32, 16),\n            =5,\n            =,\n            =,\n            =None,\n        )\n\n        self.segmentation_head = SegmentationHead(\n            =16,\n            =self.n_classes,\n            =None,\n            =3,\n        )\n\n        self.train_loss = smp.losses.TverskyLoss(=, =0.1, =0.9)\n        self.test_loss = smp.losses.DiceLoss(=)\n\n\n    def forward(self, batch, =):\n\n        x_in = batch[]\n\n        enc_out = self.encoder(x_in)\n\n        decoder_out = self.decoder(*[x_in] + enc_out)\n        x_seg = self.segmentation_head(decoder_out)\n\n        output = {}\n        one_hot_mask = batch[][:, None]\n         training:\n            loss = self.train_loss(x_seg, one_hot_mask.float())\n        :\n            loss = self.test_loss(x_seg, one_hot_mask.float())\n\n        output[] = loss\n        output[] = nn.Sigmoid()(x_seg)[:, 0]\n\n        return output\n</code></pre>\n<h1>Train</h1>\n<ul>\n<li>Scheduler : lr_warmup_cosine_decay </li>\n<li>Warmup Ratio : 0.1</li>\n<li>Optimizer : AdamW </li>\n<li>Weight Decay : 0.01</li>\n<li>Epoch : 20</li>\n<li>Learning Rate : 2e-4</li>\n<li>Loss Function : TverskyLoss(mode='binary', alpha=0.1, beta=0.9)</li>\n<li>Batchsize : 4</li>\n<li>Gradient Accumulation : 4</li>\n</ul>\n<h1>Inference</h1>\n<p>Scaling the image size by 1.5x during inference consistently resulted in score improvements in CV, public LB, and private LB. This acted as a form of dilation, significantly reducing false negatives and enhancing the model's performance. <strong>Simply increasing the image size by 1.2 for inference resulted in a 0.1 improvement on the private leaderboard.</strong></p>\n<h1>Didn't Work</h1>\n<ul>\n<li>I attempted to enhance the utility of kidney2 and kidney3 through pseudo-labeling, but it did not result in significant score improvement.</li>\n<li>Efforts to create a more robust model using heavier augmentations did not yield substantial effects.</li>\n<li>Experimenting with larger CNN models led to issues of overfitting.</li>\n</ul>\n<h1>Code</h1>\n<p><a href=\"https://github.com/siwooyong/SenNet-HOA-Hacking-the-Human-Vasculature-in-3D\" target=\"_blank\">https://github.com/siwooyong/SenNet-HOA-Hacking-the-Human-Vasculature-in-3D</a></p>",
  "messages": [
    {
      "id": 2640657,
      "postDate": "2024-02-07T03:24:30.267Z",
      "content": "<h1>TLDR</h1>\n<p>The key factors in enhancing the model's robust performance included <strong>Tversky loss</strong>, <strong>increased inference size</strong>, and <strong>resolution augmentation</strong>. These elements ultimately played a significant role in the model's survival during shakeups. The final model is an ensemble composed of 2D U-Net models based on the RegNetY-016 architecture</p>\n<h1>Interesting Point</h1>\n<p>It was a truly challenging competition to secure a reliable validation set (which, unfortunately, I couldn't achieve). Upon reviewing the results, I discovered a significant difference between the public and private sets. Surprisingly, there were submissions from the past that would have made it into the gold zone. It's astonishing, considering I didn't think much of those submissions and didn't end up submitting them anyway.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8251891%2Fe5a2fef330678d164bc8ff54d06c3e34%2F2024-02-07%2021-58-54.png?generation=1707310804960891&amp;alt=media\"></p>\n<p>In my experiments, I found that giving a higher weight to the positive class in the Tversky loss (with a larger beta) resulted in better performance on the private leaderboard. However, for my final submission, where I trained with a smaller beta value in the Tversky loss, the model performed well on the public leaderboard but surprisingly poorly on the private leaderboard. <strong>This discrepancy may be attributed to the higher resolution of the private dataset, which likely contained more detailed input information. Consequently, it suggests that a lower threshold for the model logits might have been more appropriate in this context. Furthermore, scaling the image (1.2, 1.5, …) during inference seemed to dilute input information, reducing the impact of resolution.</strong></p>\n<h1>Data Processing</h1>\n<p>Three consecutive slices were used as input for the model, and training was performed with the original image size. Despite experimenting with an increased number of slices (5, 7, 9…), the results revealed a decrease in performance.</p>\n<p>In the early stages of the competition, I trained the model by resizing the data to a specific size. However, as I mentioned in the <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/463121\" target=\"_blank\">discussion</a>, the Resize function in the Albumentations library applies nearest interpolation to masks, causing significant noise in fine labels and resulting in a notable decrease in performance. Therefore, I opted to use the original image size.</p>\n<h1>Augmentation</h1>\n<p>Awaring of the different resolutions between public and private data, I aimed to create a model robust to resolution variations. Additionally, understanding the existence of resizing during the binning process, I employed blur augmentation. </p>\n<pre><code>def blur_augmentation(x):\n    h, w,  = x.shape\n     = ..uniform(, )\n\n    x = A.Resize(int(h*), int(w*))(=x)['']\n    x = A.Resize(h, w)(=x)['']\n     x\n</code></pre>\n<p>Moreover, since the channels were constructed by stacking depth, I applied the following augmentations.</p>\n<pre><code>def channel_augmentation(x, =0.5, =3):\n    assert x.shape[2]==n_channel\n     np.random.rand()&lt;prob:\n        x = np.flip(x, =2)\n    return x\n</code></pre>\n<p>Finally, given the inherent noise in the data annotations, I implemented strong cutout augmentation to prevent overfitting.</p>\n<pre><code>A.Cutout(=8, =128, =128, =0.8),\n</code></pre>\n<p>These approaches effectively contributed to performance improvement in both CV and LB.</p>\n<h1>Model</h1>\n<p>The model, like most others, employed a 2D U-Net architecture. For the CNN backbone, I utilized the lightweight model RegNetY-016. Other than that, the settings remained consistent with the default values in SMP (Segmentation Models PyTorch).</p>\n<p>Despite investing significant time in developing a 3D-based model, it failed to demonstrate notable score improvements in both CV and LB. Due to resource constraints, I shifted my focus to a 2D model.</p>\n<pre><code>class CustomModel(nn.Module):\n    def __init__(self):\n        super(CustomModel, self).__init__()\n\n        self.n_classes = 1\n        self.in_chans = 3\n\n        self.encoder = timm.create_model(\n            ,\n            =,\n            =,\n            =self.in_chans,\n        )\n        encoder_channels = tuple(\n            [self.in_chans]\n            + [\n                self.encoder.feature_info[i][]\n                 i  range(len(self.encoder.feature_info))\n            ]\n        )\n        self.decoder = UnetDecoder(\n            =encoder_channels,\n            decoder_channels=(256, 128, 64, 32, 16),\n            =5,\n            =,\n            =,\n            =None,\n        )\n\n        self.segmentation_head = SegmentationHead(\n            =16,\n            =self.n_classes,\n            =None,\n            =3,\n        )\n\n        self.train_loss = smp.losses.TverskyLoss(=, =0.1, =0.9)\n        self.test_loss = smp.losses.DiceLoss(=)\n\n\n    def forward(self, batch, =):\n\n        x_in = batch[]\n\n        enc_out = self.encoder(x_in)\n\n        decoder_out = self.decoder(*[x_in] + enc_out)\n        x_seg = self.segmentation_head(decoder_out)\n\n        output = {}\n        one_hot_mask = batch[][:, None]\n         training:\n            loss = self.train_loss(x_seg, one_hot_mask.float())\n        :\n            loss = self.test_loss(x_seg, one_hot_mask.float())\n\n        output[] = loss\n        output[] = nn.Sigmoid()(x_seg)[:, 0]\n\n        return output\n</code></pre>\n<h1>Train</h1>\n<ul>\n<li>Scheduler : lr_warmup_cosine_decay </li>\n<li>Warmup Ratio : 0.1</li>\n<li>Optimizer : AdamW </li>\n<li>Weight Decay : 0.01</li>\n<li>Epoch : 20</li>\n<li>Learning Rate : 2e-4</li>\n<li>Loss Function : TverskyLoss(mode='binary', alpha=0.1, beta=0.9)</li>\n<li>Batchsize : 4</li>\n<li>Gradient Accumulation : 4</li>\n</ul>\n<h1>Inference</h1>\n<p>Scaling the image size by 1.5x during inference consistently resulted in score improvements in CV, public LB, and private LB. This acted as a form of dilation, significantly reducing false negatives and enhancing the model's performance. <strong>Simply increasing the image size by 1.2 for inference resulted in a 0.1 improvement on the private leaderboard.</strong></p>\n<h1>Didn't Work</h1>\n<ul>\n<li>I attempted to enhance the utility of kidney2 and kidney3 through pseudo-labeling, but it did not result in significant score improvement.</li>\n<li>Efforts to create a more robust model using heavier augmentations did not yield substantial effects.</li>\n<li>Experimenting with larger CNN models led to issues of overfitting.</li>\n</ul>\n<h1>Code</h1>\n<p><a href=\"https://github.com/siwooyong/SenNet-HOA-Hacking-the-Human-Vasculature-in-3D\" target=\"_blank\">https://github.com/siwooyong/SenNet-HOA-Hacking-the-Human-Vasculature-in-3D</a></p>",
      "rawMarkdown": "# TLDR\nThe key factors in enhancing the model's robust performance included **Tversky loss**, **increased inference size**, and **resolution augmentation**. These elements ultimately played a significant role in the model's survival during shakeups. The final model is an ensemble composed of 2D U-Net models based on the RegNetY-016 architecture\n\n# Interesting Point\nIt was a truly challenging competition to secure a reliable validation set (which, unfortunately, I couldn't achieve). Upon reviewing the results, I discovered a significant difference between the public and private sets. Surprisingly, there were submissions from the past that would have made it into the gold zone. It's astonishing, considering I didn't think much of those submissions and didn't end up submitting them anyway.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8251891%2Fe5a2fef330678d164bc8ff54d06c3e34%2F2024-02-07%2021-58-54.png?generation=1707310804960891&alt=media)\n\nIn my experiments, I found that giving a higher weight to the positive class in the Tversky loss (with a larger beta) resulted in better performance on the private leaderboard. However, for my final submission, where I trained with a smaller beta value in the Tversky loss, the model performed well on the public leaderboard but surprisingly poorly on the private leaderboard. **This discrepancy may be attributed to the higher resolution of the private dataset, which likely contained more detailed input information. Consequently, it suggests that a lower threshold for the model logits might have been more appropriate in this context. Furthermore, scaling the image (1.2, 1.5, ...) during inference seemed to dilute input information, reducing the impact of resolution.**\n\n# Data Processing\nThree consecutive slices were used as input for the model, and training was performed with the original image size. Despite experimenting with an increased number of slices (5, 7, 9...), the results revealed a decrease in performance.\n\nIn the early stages of the competition, I trained the model by resizing the data to a specific size. However, as I mentioned in the [discussion](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/463121), the Resize function in the Albumentations library applies nearest interpolation to masks, causing significant noise in fine labels and resulting in a notable decrease in performance. Therefore, I opted to use the original image size.\n\n# Augmentation\nAwaring of the different resolutions between public and private data, I aimed to create a model robust to resolution variations. Additionally, understanding the existence of resizing during the binning process, I employed blur augmentation. \n\n    def blur_augmentation(x):\n        h, w, _ = x.shape\n        scale = np.random.uniform(0.5, 1.5)\n\n        x = A.Resize(int(h*scale), int(w*scale))(image=x)['image']\n        x = A.Resize(h, w)(image=x)['image']\n        return x\n\nMoreover, since the channels were constructed by stacking depth, I applied the following augmentations.\n\n    def channel_augmentation(x, prob=0.5, n_channel=3):\n        assert x.shape[2]==n_channel\n        if np.random.rand()<prob:\n            x = np.flip(x, axis=2)\n        return x\n\nFinally, given the inherent noise in the data annotations, I implemented strong cutout augmentation to prevent overfitting.\n\n    A.Cutout(num_holes=8, max_h_size=128, max_w_size=128, p=0.8),\n\nThese approaches effectively contributed to performance improvement in both CV and LB.\n\n\n# Model\nThe model, like most others, employed a 2D U-Net architecture. For the CNN backbone, I utilized the lightweight model RegNetY-016. Other than that, the settings remained consistent with the default values in SMP (Segmentation Models PyTorch).\n\nDespite investing significant time in developing a 3D-based model, it failed to demonstrate notable score improvements in both CV and LB. Due to resource constraints, I shifted my focus to a 2D model.\n\n    class CustomModel(nn.Module):\n        def __init__(self):\n            super(CustomModel, self).__init__()\n\n            self.n_classes = 1\n            self.in_chans = 3\n\n            self.encoder = timm.create_model(\n                'regnety_016',\n                pretrained=True,\n                features_only=True,\n                in_chans=self.in_chans,\n            )\n            encoder_channels = tuple(\n                [self.in_chans]\n                + [\n                    self.encoder.feature_info[i][\"num_chs\"]\n                    for i in range(len(self.encoder.feature_info))\n                ]\n            )\n            self.decoder = UnetDecoder(\n                encoder_channels=encoder_channels,\n                decoder_channels=(256, 128, 64, 32, 16),\n                n_blocks=5,\n                use_batchnorm=True,\n                center=False,\n                attention_type=None,\n            )\n\n            self.segmentation_head = SegmentationHead(\n                in_channels=16,\n                out_channels=self.n_classes,\n                activation=None,\n                kernel_size=3,\n            )\n\n            self.train_loss = smp.losses.TverskyLoss(mode='binary', alpha=0.1, beta=0.9)\n            self.test_loss = smp.losses.DiceLoss(mode='binary')\n\n\n        def forward(self, batch, training=False):\n\n            x_in = batch[\"input\"]\n\n            enc_out = self.encoder(x_in)\n\n            decoder_out = self.decoder(*[x_in] + enc_out)\n            x_seg = self.segmentation_head(decoder_out)\n\n            output = {}\n            one_hot_mask = batch[\"mask\"][:, None]\n            if training:\n                loss = self.train_loss(x_seg, one_hot_mask.float())\n            else:\n                loss = self.test_loss(x_seg, one_hot_mask.float())\n\n            output[\"loss\"] = loss\n            output['logit'] = nn.Sigmoid()(x_seg)[:, 0]\n\n            return output\n\n# Train\n\n* Scheduler : lr_warmup_cosine_decay \n* Warmup Ratio : 0.1\n* Optimizer : AdamW \n* Weight Decay : 0.01\n* Epoch : 20\n* Learning Rate : 2e-4\n* Loss Function : TverskyLoss(mode='binary', alpha=0.1, beta=0.9)\n* Batchsize : 4\n* Gradient Accumulation : 4\n\n# Inference\nScaling the image size by 1.5x during inference consistently resulted in score improvements in CV, public LB, and private LB. This acted as a form of dilation, significantly reducing false negatives and enhancing the model's performance. **Simply increasing the image size by 1.2 for inference resulted in a 0.1 improvement on the private leaderboard.**\n\n# Didn't Work\n* I attempted to enhance the utility of kidney2 and kidney3 through pseudo-labeling, but it did not result in significant score improvement.\n* Efforts to create a more robust model using heavier augmentations did not yield substantial effects.\n* Experimenting with larger CNN models led to issues of overfitting.\n\n# Code\nhttps://github.com/siwooyong/SenNet-HOA-Hacking-the-Human-Vasculature-in-3D",
      "votes": 19
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2640657": "# TLDR\nThe key factors in enhancing the model's robust performance included **Tversky loss**, **increased inference size**, and **resolution augmentation**. These elements ultimately played a significant role in the model's survival during shakeups. The final model is an ensemble composed of 2D U-Net models based on the RegNetY-016 architecture\n\n# Interesting Point\nIt was a truly challenging competition to secure a reliable validation set (which, unfortunately, I couldn't achieve). Upon reviewing the results, I discovered a significant difference between the public and private sets. Surprisingly, there were submissions from the past that would have made it into the gold zone. It's astonishing, considering I didn't think much of those submissions and didn't end up submitting them anyway.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8251891%2Fe5a2fef330678d164bc8ff54d06c3e34%2F2024-02-07%2021-58-54.png?generation=1707310804960891&alt=media)\n\nIn my experiments, I found that giving a higher weight to the positive class in the Tversky loss (with a larger beta) resulted in better performance on the private leaderboard. However, for my final submission, where I trained with a smaller beta value in the Tversky loss, the model performed well on the public leaderboard but surprisingly poorly on the private leaderboard. **This discrepancy may be attributed to the higher resolution of the private dataset, which likely contained more detailed input information. Consequently, it suggests that a lower threshold for the model logits might have been more appropriate in this context. Furthermore, scaling the image (1.2, 1.5, ...) during inference seemed to dilute input information, reducing the impact of resolution.**\n\n# Data Processing\nThree consecutive slices were used as input for the model, and training was performed with the original image size. Despite experimenting with an increased number of slices (5, 7, 9...), the results revealed a decrease in performance.\n\nIn the early stages of the competition, I trained the model by resizing the data to a specific size. However, as I mentioned in the [discussion](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/463121), the Resize function in the Albumentations library applies nearest interpolation to masks, causing significant noise in fine labels and resulting in a notable decrease in performance. Therefore, I opted to use the original image size.\n\n# Augmentation\nAwaring of the different resolutions between public and private data, I aimed to create a model robust to resolution variations. Additionally, understanding the existence of resizing during the binning process, I employed blur augmentation. \n\n    def blur_augmentation(x):\n        h, w, _ = x.shape\n        scale = np.random.uniform(0.5, 1.5)\n\n        x = A.Resize(int(h*scale), int(w*scale))(image=x)['image']\n        x = A.Resize(h, w)(image=x)['image']\n        return x\n\nMoreover, since the channels were constructed by stacking depth, I applied the following augmentations.\n\n    def channel_augmentation(x, prob=0.5, n_channel=3):\n        assert x.shape[2]==n_channel\n        if np.random.rand()<prob:\n            x = np.flip(x, axis=2)\n        return x\n\nFinally, given the inherent noise in the data annotations, I implemented strong cutout augmentation to prevent overfitting.\n\n    A.Cutout(num_holes=8, max_h_size=128, max_w_size=128, p=0.8),\n\nThese approaches effectively contributed to performance improvement in both CV and LB.\n\n\n# Model\nThe model, like most others, employed a 2D U-Net architecture. For the CNN backbone, I utilized the lightweight model RegNetY-016. Other than that, the settings remained consistent with the default values in SMP (Segmentation Models PyTorch).\n\nDespite investing significant time in developing a 3D-based model, it failed to demonstrate notable score improvements in both CV and LB. Due to resource constraints, I shifted my focus to a 2D model.\n\n    class CustomModel(nn.Module):\n        def __init__(self):\n            super(CustomModel, self).__init__()\n\n            self.n_classes = 1\n            self.in_chans = 3\n\n            self.encoder = timm.create_model(\n                'regnety_016',\n                pretrained=True,\n                features_only=True,\n                in_chans=self.in_chans,\n            )\n            encoder_channels = tuple(\n                [self.in_chans]\n                + [\n                    self.encoder.feature_info[i][\"num_chs\"]\n                    for i in range(len(self.encoder.feature_info))\n                ]\n            )\n            self.decoder = UnetDecoder(\n                encoder_channels=encoder_channels,\n                decoder_channels=(256, 128, 64, 32, 16),\n                n_blocks=5,\n                use_batchnorm=True,\n                center=False,\n                attention_type=None,\n            )\n\n            self.segmentation_head = SegmentationHead(\n                in_channels=16,\n                out_channels=self.n_classes,\n                activation=None,\n                kernel_size=3,\n            )\n\n            self.train_loss = smp.losses.TverskyLoss(mode='binary', alpha=0.1, beta=0.9)\n            self.test_loss = smp.losses.DiceLoss(mode='binary')\n\n\n        def forward(self, batch, training=False):\n\n            x_in = batch[\"input\"]\n\n            enc_out = self.encoder(x_in)\n\n            decoder_out = self.decoder(*[x_in] + enc_out)\n            x_seg = self.segmentation_head(decoder_out)\n\n            output = {}\n            one_hot_mask = batch[\"mask\"][:, None]\n            if training:\n                loss = self.train_loss(x_seg, one_hot_mask.float())\n            else:\n                loss = self.test_loss(x_seg, one_hot_mask.float())\n\n            output[\"loss\"] = loss\n            output['logit'] = nn.Sigmoid()(x_seg)[:, 0]\n\n            return output\n\n# Train\n\n* Scheduler : lr_warmup_cosine_decay \n* Warmup Ratio : 0.1\n* Optimizer : AdamW \n* Weight Decay : 0.01\n* Epoch : 20\n* Learning Rate : 2e-4\n* Loss Function : TverskyLoss(mode='binary', alpha=0.1, beta=0.9)\n* Batchsize : 4\n* Gradient Accumulation : 4\n\n# Inference\nScaling the image size by 1.5x during inference consistently resulted in score improvements in CV, public LB, and private LB. This acted as a form of dilation, significantly reducing false negatives and enhancing the model's performance. **Simply increasing the image size by 1.2 for inference resulted in a 0.1 improvement on the private leaderboard.**\n\n# Didn't Work\n* I attempted to enhance the utility of kidney2 and kidney3 through pseudo-labeling, but it did not result in significant score improvement.\n* Efforts to create a more robust model using heavier augmentations did not yield substantial effects.\n* Experimenting with larger CNN models led to issues of overfitting.\n\n# Code\nhttps://github.com/siwooyong/SenNet-HOA-Hacking-the-Human-Vasculature-in-3D"
  }
}