{
  "id": 475052,
  "title": "4th place solution. Boundary DoU Loss is all you need!",
  "url": "/competitions/blood-vessel-segmentation/writeups/igor-krashenyi-4th-place-solution-boundary-dou-los",
  "author_name": "",
  "post_date": "2024-02-09T16:36:35.977Z",
  "votes": 68,
  "comment_count": 35,
  "views": 0,
  "content": "<p>First of all, I would like to start my solution description with a few important words:</p>\n<p><em>I would like to thank the Armed Forces of Ukraine, the Security Service of Ukraine, Defence Intelligence of Ukraine, and the State Emergency Service of Ukraine for providing safety and security to participate in this great competition, complete this work, and help science, technology, and business not to stop but to move forward.</em></p>\n<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation</a> </li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/data\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/data</a></li>\n</ul>\n<h1>Overview of the approach:</h1>\n<p>My final model is a mixture of 2d and 3d models with d4 tta. For the 2d model, the multiview tta was applied. All models were trained in a 2-fold setup with kidney_2 and kidney_3_dense selected as validation sets. The ensembling was performed with equal weights for both 2d and 3d models.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F207760%2F0e15deeafd39981c337888fdaf27e2c9%2FScreenshot%202024-02-07%20at%2002.40.31.png?generation=1707266453616337&amp;alt=media\"></p>\n<h1>Details of the submission</h1>\n<h2>Data preparation and training data and validation scheme</h2>\n<p>All final (3d and 2d) models were trained on kidney_1_dense, kidney_2, kidney_3_dense, kidney_3_sparse and pseudo labels <a href=\"http://human-organ-atlas.esrf.eu\" target=\"_blank\">50um_LADAF-2020-31_kidney_pag-0.01_0.02_jp2_</a>. Initially, I used slice-wise normalization to normalize images but later switched to stack-wise normalization based on percentiles.</p>\n<p>The 2D model was trained in a multiview setup: all images were stacked in a tensor and sliced in different axes afterward. During the training, the set of augmentations and sampling strategy was crucial. The weighted sampling was based on sparsity percentage: dense samples had a weight of 1, while sparse samples had a weight equal to their sparsity. For pseudo labels,  I chose the same weight as for kidney_2, e.g.: </p>\n<pre><code>kidney_1_dense: , \nkidney_2: , \nkidney_3_dense: , \nkidney_3_sparse: , \n50um_LADAF--31_kidney_pag-_jp2_: . \n</code></pre>\n<p>The augmentation scheme was the next one, with a chance of 0.5 CutMix augmentation being applied. The cropping was performed from the same organ and the same projection axis. Afterward, on top of CutMix, the next augmentation pipeline was applied:</p>\n<pre><code>A.Compose(\n    [\n          A.PadIfNeeded(*crop_size),\n          A.CropNonEmptyMaskIfExists(*crop_size, p=),\n          A.ShiftScaleRotate(scale_limit=),\n          A.HorizontalFlip(p=),\n          A.VerticalFlip(p=),\n          A.RandomRotate90(p=),\n          A.OneOf([\n                A.RandomBrightnessContrast(), \n                A.RandomBrightness(), \n                A.RandomGamma(),\n          ],p=,),\n    ],p=,)\n</code></pre>\n<p>The crop size was set to 512. I’ve also tried higher resolution, but it performs +- the same result. </p>\n<p>I did some experiments with 2.5d approaches (3 and 5 channels), but it produced the same result or worse. </p>\n<p>The 3d model augmentation scheme contained only d4 augmentations and random crops. The cropping was performed with a 0.5 probability of an empty mask. This was motivated by false positives that appeared outside the kidney volume. This could be improved by incorporating the two-class 3D segmentation, but I didn’t have much time and resources to perform such an experiment. Thus, I decided to create a post-processing that would handle this. <br>\nThe crop size for the 3d model was 192x192x192.</p>\n<p>Both models were trained in a 2-fold setup where as validation, I used kidney_2 (fold_1) and kidney_3_dense (fold_0). Removal of kidney_1 from the training set caused performance degradation in performance in both CV and LB, so I dropped the fold_2 and didn't perform training in that setup.</p>\n<h2>Model setup</h2>\n<p>The best results I was able to get using the efficientnet family models with UnetPlusPlus decoder and SCSE attention from the segmentation_models_pytorch library. I’ve tried the resnet50 model, like it was mentioned in the discussion section, different transformers and seresnext models, but could overcome the performance of efficientnet-b5 (which performed the best on both CV and LB). On my local validation, the score I was able to get with efficientnet_b7 encoder and mit_b5 encoder, but on the LB the score was significantly lower.<br>\nThe training was performed for 30 epochs with a Cosine LR scheduler starting from 3e-4 to 1e-6. I saved the top 3 checkpoints and used the best-last checkpoint for the submission.<br>\nThe model efficientnet_b5_UnetPlusPlus trained in such a setup was able to score 0.878 on the public LB and 0.714 on the private LB at a 0.05 threshold. </p>\n<p>Here yellow is TP, green is FP, and red is FN.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F207760%2F0484711261976f944aca4341305752f2%2FindividualImage.png?generation=1707265123800657&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F207760%2F157fa4895ef81f6e1fc0dbb8e685c17d%2FindividualImage-2.png?generation=1707265135912886&amp;alt=media\"></p>\n<p>The 3d model was heavily inspired by the nnUnet model architecture and was pretty much the same. Instead of the native nnUnet model, I used DynUnet from the monai library with almost default configuration and trained in almost the same setup as for nnUnet. As the optimizer, I used SGD with initial LR 0.01 and the Cosine Annealing LR scheme instead of LinearLR and trained for 500 epochs with 2000 samples per epoch. <br>\nThis model scored 0.869 (0.868 and 0.866 -- 0 and 1 folds respectively) on the public LB and 0.694 on the private LB (0.758 and 0.663 -- 0 and 1 folds respectively). </p>\n<p>Both models were trained using the BoundaryDOULoss (<a href=\"https://arxiv.org/pdf/2308.00220.pdf)\" target=\"_blank\">https://arxiv.org/pdf/2308.00220.pdf)</a>, which performed the best. I’ve tried to modify it to perform better on sparse data but failed. </p>\n<h2>Pseudo labeling</h2>\n<p>Based on the preprint, I downloaded the additional data from <a href=\"http://human-organ-atlas.esrf.eu\" target=\"_blank\">http://human-organ-atlas.esrf.eu</a> site (2 datasets). It appeared, that one of the datasets overlaps with the kidney_3, so I dropped it to prevent leakage. I used the other one to generate pseudo labels. For pseudo labeling, I used an ensemble of 2d models (efficientnet-b5 and efficientnet-b6 with UnetPlusPlus) trained with the same setup but without CutMix. The correct setup of CutMix as well as the 3d model I was able to discover close to the competition deadline, so I didn’t retrain the original ensemble and stick to the first version of pseudos. </p>\n<h2>Inference setup and Post-processing</h2>\n<p>The inference for both models was performed using sliding_window_inference from monai library. Additionally, for 2d model I performed multi-view tta, which helped to detect small vessels and improve overall performance. </p>\n<p>For the 2d model, the crop size was 800 pix, while for the 3d – 256 pix with 0.25 overlap and Gaussian merging. All models used d4_transform from ttach library. I’ve forked the ttach repository and implemented the logic for 3d images, but the inference time increased significantly, and there was no major boost in performance, so I’ve sticked with 2d d4_transform for both 2d and 3d models :)</p>\n<p>As I mentioned before, the 3d model had decent performance on the non-empty cubes, while empty ones were confusing the model. To handle this issue, decided to experiment with post-processing. The idea was the next one: let's try to find ROI where the vessels were presented. Since the 2d model didn’t have such a problem I’ve decided to find a bounding polygon for vessels for each 2d slice. Having a mask of ROI, I multiplied it with 3d model predictions and got a boost from 0.869 to 0.881 public LB and 0.701 private LB for a single 3d model.</p>\n<p>Ensembling the 2d model and 3d model predictions with weights 1 and 1, I was able to improve the score from 0.881 to 0.884 on the public LB and 0.712 on the private LB.</p>\n<p>Another post-processing approach that I’ve tried is to use Canny filters from cv2 to segment the kidney. This segmentation algorithm was not perfect, but applying such post-processing boosted my score from 0.884 to 0.892 on the public LB while failing on the private LB, scoring just 0.313.</p>\n<h2>What didn’t work</h2>\n<ul>\n<li>nnUnet out of the box. At the beginning of the challenge and after the pre-print reading, I tried to reproduce the result with nnUnet. The local score was promising, but the LB was 0. My intuition behind this issue related to the data normalization and spacing (scale), but I didn’t try to fix it and decided to build my own solution.</li>\n<li>BCE and Focal Loss.</li>\n<li>Transformers in both 2d and 3d model</li>\n<li>Zoom and brightness augmentation for 3d images</li>\n<li>Pseudo on top of sparse datasets. I’ve tried to fulfill the sparsity of the dataset by pseudo labeling and aggregation, but it didn’t improve the score.</li>\n<li>Additional projections. I’ve performed experiments with 2d models and additional slices generated from the 3d stack, but LB performance dropped by 20% while CV was about the same. </li>\n<li>Auxiliary outputs such as distance transform or center of mass. </li>\n<li><strong>and the most important: validation</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F207760%2Fcd379db59f3575778ebc4eca63ef0189%2FScreenshot%202024-02-07%20at%2002.29.01.png?generation=1707265761683297&amp;alt=media\"></li>\n</ul>\n<p>P.S. If you were able to read all of this, the top score on the private LB was a simple mix of 2d and 2.5d models with 1 and 3 channels :) </p>\n<p>P.P.S. Thank you for reading!</p>\n<h2>Links</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/datasets/igorkrashenyi/50um-ladaf-2020-31-kidney-pag-0-01-0-02-jp2\" target=\"_blank\">Pseudo labels </a></li>\n<li><a href=\"https://github.com/burnmyletters/blood-vessel-segmentation-public\" target=\"_blank\">Source code</a></li>\n<li>Inference code <a href=\"https://www.kaggle.com/code/igorkrashenyi/4th-place-solution/notebook\" target=\"_blank\">https://www.kaggle.com/code/igorkrashenyi/4th-place-solution/notebook</a> + <a href=\"https://www.kaggle.com/code/igorkrashenyi/fork-of-multiview-2-5-sennet-hoa-inference-v3\" target=\"_blank\">https://www.kaggle.com/code/igorkrashenyi/fork-of-multiview-2-5-sennet-hoa-inference-v3</a></li>\n</ul>",
  "messages": [
    {
      "id": "2640468",
      "postDate": "02/07/2024 00:31:42",
      "content": "<p>First of all, I would like to start my solution description with a few important words:</p>\n<p><em>I would like to thank the Armed Forces of Ukraine, the Security Service of Ukraine, Defence Intelligence of Ukraine, and the State Emergency Service of Ukraine for providing safety and security to participate in this great competition, complete this work, and help science, technology, and business not to stop but to move forward.</em></p>\n<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation</a> </li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/data\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/data</a></li>\n</ul>\n<h1>Overview of the approach:</h1>\n<p>My final model is a mixture of 2d and 3d models with d4 tta. For the 2d model, the multiview tta was applied. All models were trained in a 2-fold setup with kidney_2 and kidney_3_dense selected as validation sets. The ensembling was performed with equal weights for both 2d and 3d models.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F207760%2F0e15deeafd39981c337888fdaf27e2c9%2FScreenshot%202024-02-07%20at%2002.40.31.png?generation=1707266453616337&amp;alt=media\"></p>\n<h1>Details of the submission</h1>\n<h2>Data preparation and training data and validation scheme</h2>\n<p>All final (3d and 2d) models were trained on kidney_1_dense, kidney_2, kidney_3_dense, kidney_3_sparse and pseudo labels <a href=\"http://human-organ-atlas.esrf.eu\" target=\"_blank\">50um_LADAF-2020-31_kidney_pag-0.01_0.02_jp2_</a>. Initially, I used slice-wise normalization to normalize images but later switched to stack-wise normalization based on percentiles.</p>\n<p>The 2D model was trained in a multiview setup: all images were stacked in a tensor and sliced in different axes afterward. During the training, the set of augmentations and sampling strategy was crucial. The weighted sampling was based on sparsity percentage: dense samples had a weight of 1, while sparse samples had a weight equal to their sparsity. For pseudo labels,  I chose the same weight as for kidney_2, e.g.: </p>\n<pre><code>kidney_1_dense: , \nkidney_2: , \nkidney_3_dense: , \nkidney_3_sparse: , \n50um_LADAF--31_kidney_pag-_jp2_: . \n</code></pre>\n<p>The augmentation scheme was the next one, with a chance of 0.5 CutMix augmentation being applied. The cropping was performed from the same organ and the same projection axis. Afterward, on top of CutMix, the next augmentation pipeline was applied:</p>\n<pre><code>A.Compose(\n    [\n          A.PadIfNeeded(*crop_size),\n          A.CropNonEmptyMaskIfExists(*crop_size, p=),\n          A.ShiftScaleRotate(scale_limit=),\n          A.HorizontalFlip(p=),\n          A.VerticalFlip(p=),\n          A.RandomRotate90(p=),\n          A.OneOf([\n                A.RandomBrightnessContrast(), \n                A.RandomBrightness(), \n                A.RandomGamma(),\n          ],p=,),\n    ],p=,)\n</code></pre>\n<p>The crop size was set to 512. I’ve also tried higher resolution, but it performs +- the same result. </p>\n<p>I did some experiments with 2.5d approaches (3 and 5 channels), but it produced the same result or worse. </p>\n<p>The 3d model augmentation scheme contained only d4 augmentations and random crops. The cropping was performed with a 0.5 probability of an empty mask. This was motivated by false positives that appeared outside the kidney volume. This could be improved by incorporating the two-class 3D segmentation, but I didn’t have much time and resources to perform such an experiment. Thus, I decided to create a post-processing that would handle this. <br>\nThe crop size for the 3d model was 192x192x192.</p>\n<p>Both models were trained in a 2-fold setup where as validation, I used kidney_2 (fold_1) and kidney_3_dense (fold_0). Removal of kidney_1 from the training set caused performance degradation in performance in both CV and LB, so I dropped the fold_2 and didn't perform training in that setup.</p>\n<h2>Model setup</h2>\n<p>The best results I was able to get using the efficientnet family models with UnetPlusPlus decoder and SCSE attention from the segmentation_models_pytorch library. I’ve tried the resnet50 model, like it was mentioned in the discussion section, different transformers and seresnext models, but could overcome the performance of efficientnet-b5 (which performed the best on both CV and LB). On my local validation, the score I was able to get with efficientnet_b7 encoder and mit_b5 encoder, but on the LB the score was significantly lower.<br>\nThe training was performed for 30 epochs with a Cosine LR scheduler starting from 3e-4 to 1e-6. I saved the top 3 checkpoints and used the best-last checkpoint for the submission.<br>\nThe model efficientnet_b5_UnetPlusPlus trained in such a setup was able to score 0.878 on the public LB and 0.714 on the private LB at a 0.05 threshold. </p>\n<p>Here yellow is TP, green is FP, and red is FN.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F207760%2F0484711261976f944aca4341305752f2%2FindividualImage.png?generation=1707265123800657&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F207760%2F157fa4895ef81f6e1fc0dbb8e685c17d%2FindividualImage-2.png?generation=1707265135912886&amp;alt=media\"></p>\n<p>The 3d model was heavily inspired by the nnUnet model architecture and was pretty much the same. Instead of the native nnUnet model, I used DynUnet from the monai library with almost default configuration and trained in almost the same setup as for nnUnet. As the optimizer, I used SGD with initial LR 0.01 and the Cosine Annealing LR scheme instead of LinearLR and trained for 500 epochs with 2000 samples per epoch. <br>\nThis model scored 0.869 (0.868 and 0.866 -- 0 and 1 folds respectively) on the public LB and 0.694 on the private LB (0.758 and 0.663 -- 0 and 1 folds respectively). </p>\n<p>Both models were trained using the BoundaryDOULoss (<a href=\"https://arxiv.org/pdf/2308.00220.pdf)\" target=\"_blank\">https://arxiv.org/pdf/2308.00220.pdf)</a>, which performed the best. I’ve tried to modify it to perform better on sparse data but failed. </p>\n<h2>Pseudo labeling</h2>\n<p>Based on the preprint, I downloaded the additional data from <a href=\"http://human-organ-atlas.esrf.eu\" target=\"_blank\">http://human-organ-atlas.esrf.eu</a> site (2 datasets). It appeared, that one of the datasets overlaps with the kidney_3, so I dropped it to prevent leakage. I used the other one to generate pseudo labels. For pseudo labeling, I used an ensemble of 2d models (efficientnet-b5 and efficientnet-b6 with UnetPlusPlus) trained with the same setup but without CutMix. The correct setup of CutMix as well as the 3d model I was able to discover close to the competition deadline, so I didn’t retrain the original ensemble and stick to the first version of pseudos. </p>\n<h2>Inference setup and Post-processing</h2>\n<p>The inference for both models was performed using sliding_window_inference from monai library. Additionally, for 2d model I performed multi-view tta, which helped to detect small vessels and improve overall performance. </p>\n<p>For the 2d model, the crop size was 800 pix, while for the 3d – 256 pix with 0.25 overlap and Gaussian merging. All models used d4_transform from ttach library. I’ve forked the ttach repository and implemented the logic for 3d images, but the inference time increased significantly, and there was no major boost in performance, so I’ve sticked with 2d d4_transform for both 2d and 3d models :)</p>\n<p>As I mentioned before, the 3d model had decent performance on the non-empty cubes, while empty ones were confusing the model. To handle this issue, decided to experiment with post-processing. The idea was the next one: let's try to find ROI where the vessels were presented. Since the 2d model didn’t have such a problem I’ve decided to find a bounding polygon for vessels for each 2d slice. Having a mask of ROI, I multiplied it with 3d model predictions and got a boost from 0.869 to 0.881 public LB and 0.701 private LB for a single 3d model.</p>\n<p>Ensembling the 2d model and 3d model predictions with weights 1 and 1, I was able to improve the score from 0.881 to 0.884 on the public LB and 0.712 on the private LB.</p>\n<p>Another post-processing approach that I’ve tried is to use Canny filters from cv2 to segment the kidney. This segmentation algorithm was not perfect, but applying such post-processing boosted my score from 0.884 to 0.892 on the public LB while failing on the private LB, scoring just 0.313.</p>\n<h2>What didn’t work</h2>\n<ul>\n<li>nnUnet out of the box. At the beginning of the challenge and after the pre-print reading, I tried to reproduce the result with nnUnet. The local score was promising, but the LB was 0. My intuition behind this issue related to the data normalization and spacing (scale), but I didn’t try to fix it and decided to build my own solution.</li>\n<li>BCE and Focal Loss.</li>\n<li>Transformers in both 2d and 3d model</li>\n<li>Zoom and brightness augmentation for 3d images</li>\n<li>Pseudo on top of sparse datasets. I’ve tried to fulfill the sparsity of the dataset by pseudo labeling and aggregation, but it didn’t improve the score.</li>\n<li>Additional projections. I’ve performed experiments with 2d models and additional slices generated from the 3d stack, but LB performance dropped by 20% while CV was about the same. </li>\n<li>Auxiliary outputs such as distance transform or center of mass. </li>\n<li><strong>and the most important: validation</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F207760%2Fcd379db59f3575778ebc4eca63ef0189%2FScreenshot%202024-02-07%20at%2002.29.01.png?generation=1707265761683297&amp;alt=media\"></li>\n</ul>\n<p>P.S. If you were able to read all of this, the top score on the private LB was a simple mix of 2d and 2.5d models with 1 and 3 channels :) </p>\n<p>P.P.S. Thank you for reading!</p>\n<h2>Links</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/datasets/igorkrashenyi/50um-ladaf-2020-31-kidney-pag-0-01-0-02-jp2\" target=\"_blank\">Pseudo labels </a></li>\n<li><a href=\"https://github.com/burnmyletters/blood-vessel-segmentation-public\" target=\"_blank\">Source code</a></li>\n<li>Inference code <a href=\"https://www.kaggle.com/code/igorkrashenyi/4th-place-solution/notebook\" target=\"_blank\">https://www.kaggle.com/code/igorkrashenyi/4th-place-solution/notebook</a> + <a href=\"https://www.kaggle.com/code/igorkrashenyi/fork-of-multiview-2-5-sennet-hoa-inference-v3\" target=\"_blank\">https://www.kaggle.com/code/igorkrashenyi/fork-of-multiview-2-5-sennet-hoa-inference-v3</a></li>\n</ul>",
      "rawMarkdown": "First of all, I would like to start my solution description with a few important words:\n\n*I would like to thank the Armed Forces of Ukraine, the Security Service of Ukraine, Defence Intelligence of Ukraine, and the State Emergency Service of Ukraine for providing safety and security to participate in this great competition, complete this work, and help science, technology, and business not to stop but to move forward.*\n\n# Context\n- Business context: https://www.kaggle.com/competitions/blood-vessel-segmentation \n- Data context: https://www.kaggle.com/competitions/blood-vessel-segmentation/data\n\n# Overview of the approach:\nMy final model is a mixture of 2d and 3d models with d4 tta. For the 2d model, the multiview tta was applied. All models were trained in a 2-fold setup with kidney_2 and kidney_3_dense selected as validation sets. The ensembling was performed with equal weights for both 2d and 3d models.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F207760%2F0e15deeafd39981c337888fdaf27e2c9%2FScreenshot%202024-02-07%20at%2002.40.31.png?generation=1707266453616337&alt=media)\n\n# Details of the submission\n\n## Data preparation and training data and validation scheme \n\nAll final (3d and 2d) models were trained on kidney_1_dense, kidney_2, kidney_3_dense, kidney_3_sparse and pseudo labels [50um_LADAF-2020-31_kidney_pag-0.01_0.02_jp2_](http://human-organ-atlas.esrf.eu). Initially, I used slice-wise normalization to normalize images but later switched to stack-wise normalization based on percentiles.\n\nThe 2D model was trained in a multiview setup: all images were stacked in a tensor and sliced in different axes afterward. During the training, the set of augmentations and sampling strategy was crucial. The weighted sampling was based on sparsity percentage: dense samples had a weight of 1, while sparse samples had a weight equal to their sparsity. For pseudo labels,  I chose the same weight as for kidney_2, e.g.: \n\n```python\nkidney_1_dense: 1, \nkidney_2: 0.65, \nkidney_3_dense: 1, \nkidney_3_sparse: 0.85, \n50um_LADAF-2020-31_kidney_pag-0.01_0.02_jp2_: 0.65. \n```\n\nThe augmentation scheme was the next one, with a chance of 0.5 CutMix augmentation being applied. The cropping was performed from the same organ and the same projection axis. Afterward, on top of CutMix, the next augmentation pipeline was applied:\n\n```python\nA.Compose(\n    [\n          A.PadIfNeeded(*crop_size),\n          A.CropNonEmptyMaskIfExists(*crop_size, p=1.0),\n          A.ShiftScaleRotate(scale_limit=0.2),\n          A.HorizontalFlip(p=0.5),\n          A.VerticalFlip(p=0.5),\n          A.RandomRotate90(p=0.5),\n          A.OneOf([\n                A.RandomBrightnessContrast(), \n                A.RandomBrightness(), \n                A.RandomGamma(),\n          ],p=1.0,),\n    ],p=1.0,)\n```\n\nThe crop size was set to 512. I’ve also tried higher resolution, but it performs +- the same result. \n\nI did some experiments with 2.5d approaches (3 and 5 channels), but it produced the same result or worse. \n\nThe 3d model augmentation scheme contained only d4 augmentations and random crops. The cropping was performed with a 0.5 probability of an empty mask. This was motivated by false positives that appeared outside the kidney volume. This could be improved by incorporating the two-class 3D segmentation, but I didn’t have much time and resources to perform such an experiment. Thus, I decided to create a post-processing that would handle this. \nThe crop size for the 3d model was 192x192x192.\n\nBoth models were trained in a 2-fold setup where as validation, I used kidney_2 (fold_1) and kidney_3_dense (fold_0). Removal of kidney_1 from the training set caused performance degradation in performance in both CV and LB, so I dropped the fold_2 and didn't perform training in that setup.\n\n## Model setup\nThe best results I was able to get using the efficientnet family models with UnetPlusPlus decoder and SCSE attention from the segmentation_models_pytorch library. I’ve tried the resnet50 model, like it was mentioned in the discussion section, different transformers and seresnext models, but could overcome the performance of efficientnet-b5 (which performed the best on both CV and LB). On my local validation, the score I was able to get with efficientnet_b7 encoder and mit_b5 encoder, but on the LB the score was significantly lower.\nThe training was performed for 30 epochs with a Cosine LR scheduler starting from 3e-4 to 1e-6. I saved the top 3 checkpoints and used the best-last checkpoint for the submission.\nThe model efficientnet_b5_UnetPlusPlus trained in such a setup was able to score 0.878 on the public LB and 0.714 on the private LB at a 0.05 threshold. \n\nHere yellow is TP, green is FP, and red is FN.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F207760%2F0484711261976f944aca4341305752f2%2FindividualImage.png?generation=1707265123800657&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F207760%2F157fa4895ef81f6e1fc0dbb8e685c17d%2FindividualImage-2.png?generation=1707265135912886&alt=media)\n\nThe 3d model was heavily inspired by the nnUnet model architecture and was pretty much the same. Instead of the native nnUnet model, I used DynUnet from the monai library with almost default configuration and trained in almost the same setup as for nnUnet. As the optimizer, I used SGD with initial LR 0.01 and the Cosine Annealing LR scheme instead of LinearLR and trained for 500 epochs with 2000 samples per epoch. \nThis model scored 0.869 (0.868 and 0.866 -- 0 and 1 folds respectively) on the public LB and 0.694 on the private LB (0.758 and 0.663 -- 0 and 1 folds respectively). \n\nBoth models were trained using the BoundaryDOULoss (https://arxiv.org/pdf/2308.00220.pdf), which performed the best. I’ve tried to modify it to perform better on sparse data but failed. \n\n## Pseudo labeling \n\nBased on the preprint, I downloaded the additional data from http://human-organ-atlas.esrf.eu site (2 datasets). It appeared, that one of the datasets overlaps with the kidney_3, so I dropped it to prevent leakage. I used the other one to generate pseudo labels. For pseudo labeling, I used an ensemble of 2d models (efficientnet-b5 and efficientnet-b6 with UnetPlusPlus) trained with the same setup but without CutMix. The correct setup of CutMix as well as the 3d model I was able to discover close to the competition deadline, so I didn’t retrain the original ensemble and stick to the first version of pseudos. \n\n## Inference setup and Post-processing \n\nThe inference for both models was performed using sliding_window_inference from monai library. Additionally, for 2d model I performed multi-view tta, which helped to detect small vessels and improve overall performance. \n\nFor the 2d model, the crop size was 800 pix, while for the 3d – 256 pix with 0.25 overlap and Gaussian merging. All models used d4_transform from ttach library. I’ve forked the ttach repository and implemented the logic for 3d images, but the inference time increased significantly, and there was no major boost in performance, so I’ve sticked with 2d d4_transform for both 2d and 3d models :)\n\nAs I mentioned before, the 3d model had decent performance on the non-empty cubes, while empty ones were confusing the model. To handle this issue, decided to experiment with post-processing. The idea was the next one: let's try to find ROI where the vessels were presented. Since the 2d model didn’t have such a problem I’ve decided to find a bounding polygon for vessels for each 2d slice. Having a mask of ROI, I multiplied it with 3d model predictions and got a boost from 0.869 to 0.881 public LB and 0.701 private LB for a single 3d model.\n\nEnsembling the 2d model and 3d model predictions with weights 1 and 1, I was able to improve the score from 0.881 to 0.884 on the public LB and 0.712 on the private LB.\n\nAnother post-processing approach that I’ve tried is to use Canny filters from cv2 to segment the kidney. This segmentation algorithm was not perfect, but applying such post-processing boosted my score from 0.884 to 0.892 on the public LB while failing on the private LB, scoring just 0.313.\n\n## What didn’t work\n- nnUnet out of the box. At the beginning of the challenge and after the pre-print reading, I tried to reproduce the result with nnUnet. The local score was promising, but the LB was 0. My intuition behind this issue related to the data normalization and spacing (scale), but I didn’t try to fix it and decided to build my own solution.\n- BCE and Focal Loss.\n- Transformers in both 2d and 3d model\n- Zoom and brightness augmentation for 3d images\n- Pseudo on top of sparse datasets. I’ve tried to fulfill the sparsity of the dataset by pseudo labeling and aggregation, but it didn’t improve the score.\n- Additional projections. I’ve performed experiments with 2d models and additional slices generated from the 3d stack, but LB performance dropped by 20% while CV was about the same. \n- Auxiliary outputs such as distance transform or center of mass. \n- **and the most important: validation**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F207760%2Fcd379db59f3575778ebc4eca63ef0189%2FScreenshot%202024-02-07%20at%2002.29.01.png?generation=1707265761683297&alt=media)\n\nP.S. If you were able to read all of this, the top score on the private LB was a simple mix of 2d and 2.5d models with 1 and 3 channels :) \n\nP.P.S. Thank you for reading!\n\n## Links\n- [Pseudo labels ](https://www.kaggle.com/datasets/igorkrashenyi/50um-ladaf-2020-31-kidney-pag-0-01-0-02-jp2 )\n- [Source code](https://github.com/burnmyletters/blood-vessel-segmentation-public)\n- Inference code https://www.kaggle.com/code/igorkrashenyi/4th-place-solution/notebook + https://www.kaggle.com/code/igorkrashenyi/fork-of-multiview-2-5-sennet-hoa-inference-v3",
      "votes": null
    },
    {
      "id": "2640481",
      "postDate": "02/07/2024 00:37:46",
      "content": "<p>Congratulations! 🎉</p>",
      "rawMarkdown": "Congratulations! 🎉",
      "votes": null
    },
    {
      "id": "2640483",
      "postDate": "02/07/2024 00:38:35",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/igorkrashenyi\" target=\"_blank\">@igorkrashenyi</a> this is great! Thanks for writing this :). Congrats on your solution!</p>\n<p>If I may ask, what do you mean stackwise normalization based on percentiles?</p>",
      "rawMarkdown": "Hey @igorkrashenyi this is great! Thanks for writing this :). Congrats on your solution!\n\nIf I may ask, what do you mean stackwise normalization based on percentiles?",
      "votes": null
    },
    {
      "id": "2640484",
      "postDate": "02/07/2024 00:38:38",
      "content": "<p>Wow incredible solution! Congrats on your placement! I thought about using 2.5d and 3d ensemble but never made the try. I am feeling like I had a lot of room for improvement on my augmentations as I was only using flips and contrast because some of my images were naturally different sizes! I will learn a lot from this one! Thank you for sharing!</p>",
      "rawMarkdown": "Wow incredible solution! Congrats on your placement! I thought about using 2.5d and 3d ensemble but never made the try. I am feeling like I had a lot of room for improvement on my augmentations as I was only using flips and contrast because some of my images were naturally different sizes! I will learn a lot from this one! Thank you for sharing!",
      "votes": null
    },
    {
      "id": "2640485",
      "postDate": "02/07/2024 00:38:52",
      "content": "<p>thanks for the write up and congrats for your top ranking and results!!!</p>\n<p>is there any comparison results with and without additional data?<br>\n(you mentioned \"I downloaded the additional data from <a href=\"http://human-organ-atlas.esrf.eu\" target=\"_blank\">http://human-organ-atlas.esrf.eu</a> \")</p>",
      "rawMarkdown": "thanks for the write up and congrats for your top ranking and results!!!\n\nis there any comparison results with and without additional data?\n(you mentioned \"I downloaded the additional data from http://human-organ-atlas.esrf.eu \")",
      "votes": null
    },
    {
      "id": "2640487",
      "postDate": "02/07/2024 00:42:27",
      "content": "<p>Congratulations!  Thx for sharing.😀</p>",
      "rawMarkdown": "Congratulations!  Thx for sharing.😀",
      "votes": null
    },
    {
      "id": "2640492",
      "postDate": "02/07/2024 00:48:08",
      "content": "<p>That's impressive, I've always wanted to try BD Loss but haven't succeeded yet.</p>",
      "rawMarkdown": "That's impressive, I've always wanted to try BD Loss but haven't succeeded yet.",
      "votes": null
    },
    {
      "id": "2640496",
      "postDate": "02/07/2024 00:51:25",
      "content": "<p>Yeah, so with the usage of pseudos  I was able to get around 2% on CV, 1% on the public LB, and 4% on private LB. </p>",
      "rawMarkdown": "Yeah, so with the usage of pseudos  I was able to get around 2% on CV, 1% on the public LB, and 4% on private LB.",
      "votes": null
    },
    {
      "id": "2640504",
      "postDate": "02/07/2024 00:54:37",
      "content": "<p>This normalization is based not on a single image but on a stack of images. <br>\nFirst, you calculate the stats for the stack (kidney_1, kidney_2, etc), each stack will have its own stats.</p>\n<p><code>xmin = np.percentile(volume, low)</code><br>\n<code>xmax = np.max([np.percentile(volume, high), 1])</code></p>\n<p>and afterward apply:<br>\n<code>\nimage = (image - xmin) / (xmax - xmin)</code><br>\n<code>image = np.clip(image, 0, 1)\n</code></p>",
      "rawMarkdown": "This normalization is based not on a single image but on a stack of images. \nFirst, you calculate the stats for the stack (kidney_1, kidney_2, etc), each stack will have its own stats.\n\n`xmin = np.percentile(volume, low)`\n`xmax = np.max([np.percentile(volume, high), 1])`\n\nand afterward apply:\n`\nimage = (image - xmin) / (xmax - xmin)`\n`image = np.clip(image, 0, 1)\n`",
      "votes": null
    },
    {
      "id": "2640528",
      "postDate": "02/07/2024 01:13:22",
      "content": "<p>thanks for the reply.</p>\n<p>i use boundary loss as well. This is for your reference.<br>\ni  implemented by BCE and interior weights and boundary weights .<br>\nfor boundary weights, it uses  both foreground and background pixels near the object boundary.<br>\n boundary loss produces my top performing models and stablised training as well.</p>\n<p>i note that towards the end of training, interior and boundary BCE completes against each other.<br>\nboundary BCE  prevents the collapse of the boundary predict. hence my validation boundary dice is very stable even after very very long training.   </p>",
      "rawMarkdown": "thanks for the reply.\n\ni use boundary loss as well. This is for your reference.\ni  implemented by BCE and interior weights and boundary weights .\nfor boundary weights, it uses  both foreground and background pixels near the object boundary.\n boundary loss produces my top performing models and stablised training as well.\n\ni note that towards the end of training, interior and boundary BCE completes against each other.\nboundary BCE  prevents the collapse of the boundary predict. hence my validation boundary dice is very stable even after very very long training.",
      "votes": null
    },
    {
      "id": "2640649",
      "postDate": "02/07/2024 03:22:21",
      "content": "<p>Congratulations on your good results. Thank you for sharing.</p>",
      "rawMarkdown": "Congratulations on your good results. Thank you for sharing.",
      "votes": null
    },
    {
      "id": "2640693",
      "postDate": "02/07/2024 03:49:42",
      "content": "<p>Congratulations! Thank you for sharing the results!❤️</p>",
      "rawMarkdown": "Congratulations! Thank you for sharing the results!❤️",
      "votes": null
    },
    {
      "id": "2640772",
      "postDate": "02/07/2024 04:46:45",
      "content": "<p>Ah, I see. Makes sense! Thank you </p>",
      "rawMarkdown": "Ah, I see. Makes sense! Thank you",
      "votes": null
    },
    {
      "id": "2641130",
      "postDate": "02/07/2024 10:00:43",
      "content": "<p>Congratulations, Thank you for sharing and explaining your approach. </p>",
      "rawMarkdown": "Congratulations, Thank you for sharing and explaining your approach.",
      "votes": null
    },
    {
      "id": "2641199",
      "postDate": "02/07/2024 10:45:13",
      "content": "<p><a href=\"https://www.kaggle.com/igorkrashenyi\" target=\"_blank\">@igorkrashenyi</a> congratz with solo gold and finally gm. By cropping technique you mean you did tiling 512x512 with overlap during inference, is that correct? What was the training time?</p>",
      "rawMarkdown": "igorkrashenyi congratz with solo gold and finally gm. By cropping technique you mean you did tiling 512x512 with overlap during inference, is that correct? What was the training time?",
      "votes": null
    },
    {
      "id": "2641260",
      "postDate": "02/07/2024 11:27:40",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/igorkrashenyi\" target=\"_blank\">@igorkrashenyi</a> , solid and creative work. Congratulations for the solo gold and grandmaster! </p>",
      "rawMarkdown": "Hi @igorkrashenyi , solid and creative work. Congratulations for the solo gold and grandmaster!",
      "votes": null
    },
    {
      "id": "2641298",
      "postDate": "02/07/2024 12:08:33",
      "content": "<p>Thanks!<br>\nNot exactly. <br>\nFor the 2d model, during the training, I performed 512x512 non-empty mask random cropping of images. During the epoch, I did 2 such crops for each image from all views. In a such setup, one epoch was approximately 1000 batches with batch size 32 and took 20 minutes on 2 RTX 4090.</p>\n<p>For the 3d model, the cropping was balanced 50/50 of empty/non-empty 192x192x192 crops. One epoch here was 2000 samples with batch size 4. In this case epoch took 6 minutes. </p>",
      "rawMarkdown": "Thanks!\nNot exactly. \nFor the 2d model, during the training, I performed 512x512 non-empty mask random cropping of images. During the epoch, I did 2 such crops for each image from all views. In a such setup, one epoch was approximately 1000 batches with batch size 32 and took 20 minutes on 2 RTX 4090.\n\nFor the 3d model, the cropping was balanced 50/50 of empty/non-empty 192x192x192 crops. One epoch here was 2000 samples with batch size 4. In this case epoch took 6 minutes.",
      "votes": null
    },
    {
      "id": "2641307",
      "postDate": "02/07/2024 12:12:56",
      "content": "<p>thank you!</p>",
      "rawMarkdown": "thank you!",
      "votes": null
    },
    {
      "id": "2641743",
      "postDate": "02/07/2024 16:35:36",
      "content": "<p>congratz. And thanks for the insightful write up!</p>",
      "rawMarkdown": "congratz. And thanks for the insightful write up!",
      "votes": null
    },
    {
      "id": "2641946",
      "postDate": "02/07/2024 19:16:05",
      "content": "<p>\"P.S. If you were able to read all of this, the top score on the private LB was a simple mix of 2d and 2.5d models with 1 and 3 channels :) \"</p>\n<p>can i conclude that high or low private score is a bit of luck in model selection, rather than apply some \"crucial\"  methods/steps?</p>",
      "rawMarkdown": "\"P.S. If you were able to read all of this, the top score on the private LB was a simple mix of 2d and 2.5d models with 1 and 3 channels :) \"\n\ncan i conclude that high or low private score is a bit of luck in model selection, rather than apply some \"crucial\"  methods/steps?",
      "votes": null
    },
    {
      "id": "2641957",
      "postDate": "02/07/2024 19:27:42",
      "content": "<p>Luck + some tricks in the training and inference process. </p>\n<p>I was not able to select the top model based on only two data points (local CV and public LB), but at some point the shakeup didn’t impact me much. </p>",
      "rawMarkdown": "Luck + some tricks in the training and inference process. \n\nI was not able to select the top model based on only two data points (local CV and public LB), but at some point the shakeup didn’t impact me much.",
      "votes": null
    },
    {
      "id": "2642023",
      "postDate": "02/07/2024 20:57:31",
      "content": "<p>Congratulations on the solo gold! I and my friends want to become kaggle master one day, so we figured reproducing your results is a great start :). In doing so, we're curious about your methodologies of arriving at your solutions: </p>\n<ul>\n<li>what's the motivation behind training the model with 512x512 crops then running the inference at 800x800 crops? Is it purely for inference speed? or is there any accuracy benefit to it?</li>\n<li>how do you come up with the augmentation pipeline? i.e. is this the first and only version of the augmentation pipeline? or do you have some methodology on how you iterate on it?</li>\n<li>same question with the 30 epochs and the cosine scheduler's start and end LR, do you iterate on that much? and do you have some ways you go about iterating on it?</li>\n</ul>\n<p>Thank you for the solution, and once again, congrats!</p>",
      "rawMarkdown": "Congratulations on the solo gold! I and my friends want to become kaggle master one day, so we figured reproducing your results is a great start :). In doing so, we're curious about your methodologies of arriving at your solutions: \n\n- what's the motivation behind training the model with 512x512 crops then running the inference at 800x800 crops? Is it purely for inference speed? or is there any accuracy benefit to it?\n- how do you come up with the augmentation pipeline? i.e. is this the first and only version of the augmentation pipeline? or do you have some methodology on how you iterate on it?\n- same question with the 30 epochs and the cosine scheduler's start and end LR, do you iterate on that much? and do you have some ways you go about iterating on it?\n\nThank you for the solution, and once again, congrats!",
      "votes": null
    },
    {
      "id": "2642172",
      "postDate": "02/08/2024 01:46:27",
      "content": "<p>This is amazing! Great work! <a href=\"https://www.kaggle.com/igorkrashenyi\" target=\"_blank\">@igorkrashenyi</a> </p>",
      "rawMarkdown": "This is amazing! Great work! @igorkrashenyi",
      "votes": null
    },
    {
      "id": "2642205",
      "postDate": "02/08/2024 02:44:12",
      "content": "<p>very helpful, thanks</p>",
      "rawMarkdown": "very helpful, thanks",
      "votes": null
    },
    {
      "id": "2642226",
      "postDate": "02/08/2024 03:19:41",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/igorkrashenyi\" target=\"_blank\">@igorkrashenyi</a> , how large the improvements brought by BoundaryDOULoss, comparing to DiceLoss, on cv, public and private, respectively?</p>",
      "rawMarkdown": "Hi @igorkrashenyi , how large the improvements brought by BoundaryDOULoss, comparing to DiceLoss, on cv, public and private, respectively?",
      "votes": null
    },
    {
      "id": "2642345",
      "postDate": "02/08/2024 05:19:45",
      "content": "<p>Great question! <br>\nI discovered this loss early in the competition and switched after a few weeks. By that time I did all my experiments with the efficientnet_b3.<br>\nSo on my CV I got +2%, on the public LB the boost was ~1.5%, while on the private leaderboard I’ve got +5%</p>",
      "rawMarkdown": "Great question! \nI discovered this loss early in the competition and switched after a few weeks. By that time I did all my experiments with the efficientnet_b3.\nSo on my CV I got +2%, on the public LB the boost was ~1.5%, while on the private leaderboard I’ve got +5%",
      "votes": null
    },
    {
      "id": "2642411",
      "postDate": "02/08/2024 06:25:16",
      "content": "<p>Amazing! I can hardly witness a loss cause such a huge impact. if the effect of BoundaryDOULoss can be validated on several more datasets, maybe it will become a priority loss in image segmentation task.</p>",
      "rawMarkdown": "Amazing! I can hardly witness a loss cause such a huge impact. if the effect of BoundaryDOULoss can be validated on several more datasets, maybe it will become a priority loss in image segmentation task.",
      "votes": null
    },
    {
      "id": "2642415",
      "postDate": "02/08/2024 06:26:40",
      "content": "<p>Or maybe BoundaryDOULoss match the surface dice(metrics) better? </p>",
      "rawMarkdown": "Or maybe BoundaryDOULoss match the surface dice(metrics) better?",
      "votes": null
    },
    {
      "id": "2642482",
      "postDate": "02/08/2024 07:10:36",
      "content": "<p>Thanks a lot!<br>\nAnd thank you for your questions. </p>\n<ul>\n<li>Regarding the inference. It’s just an empirical finding. It boosted the score on both CV and LB by ~0.003.</li>\n<li>iterative process. I’ve started with the severe augmentation scheme, but investigating the learning curves I’ve noticed a heavy underfitting. So I dropped everything and stated to add augmentations one by one. </li>\n<li>same here. It’s not about tuning the number of epochs, it’s more about tuning the amount of batches till plateau. Usually, I start with 100k-250k batches. In this approach “the epoch” is just how often you want to check the validation score. </li>\n</ul>",
      "rawMarkdown": "Thanks a lot!\nAnd thank you for your questions. \n- Regarding the inference. It’s just an empirical finding. It boosted the score on both CV and LB by ~0.003.\n- iterative process. I’ve started with the severe augmentation scheme, but investigating the learning curves I’ve noticed a heavy underfitting. So I dropped everything and stated to add augmentations one by one. \n- same here. It’s not about tuning the number of epochs, it’s more about tuning the amount of batches till plateau. Usually, I start with 100k-250k batches. In this approach “the epoch” is just how often you want to check the validation score.",
      "votes": null
    },
    {
      "id": "2642485",
      "postDate": "02/08/2024 07:13:58",
      "content": "<p>I guess both. But for this specific task and surface Dice metric it fit perfectly, so the boost is so huge. </p>",
      "rawMarkdown": "I guess both. But for this specific task and surface Dice metric it fit perfectly, so the boost is so huge.",
      "votes": null
    },
    {
      "id": "2642566",
      "postDate": "02/08/2024 08:59:38",
      "content": "<p>Yes, it does fit this metric perfectly. Thanks for introduce this loss.</p>",
      "rawMarkdown": "Yes, it does fit this metric perfectly. Thanks for introduce this loss.",
      "votes": null
    },
    {
      "id": "2644742",
      "postDate": "02/09/2024 16:37:41",
      "content": "<p>added the source code to the post </p>",
      "rawMarkdown": "added the source code to the post",
      "votes": null
    },
    {
      "id": "2644772",
      "postDate": "02/09/2024 17:07:17",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/igorkrashenyi\" target=\"_blank\">@igorkrashenyi</a> this is great! Thanks for writing this :). Congrats on your solution!</p>\n<p>If I may ask, what do you mean stackwise normalization based on percentiles?<br>\nThis is amazing! Great work!</p>",
      "rawMarkdown": "Hey @igorkrashenyi this is great! Thanks for writing this :). Congrats on your solution!\n\nIf I may ask, what do you mean stackwise normalization based on percentiles?\nThis is amazing! Great work!",
      "votes": null
    },
    {
      "id": "2644870",
      "postDate": "02/09/2024 18:15:32",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/475052#2640504\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/475052#2640504</a></p>",
      "rawMarkdown": "https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/475052#2640504",
      "votes": null
    },
    {
      "id": "2649561",
      "postDate": "02/13/2024 00:59:58",
      "content": "<p>thank you for the answer and sorry for the late reply. We want to thank you that the codebase has been insightful.</p>\n<p>In the mean time we did make a meme…</p>\n<p><a href=\"https://www.youtube.com/watch?v=tKM4xxBoLJI&amp;t=3s\" target=\"_blank\">https://www.youtube.com/watch?v=tKM4xxBoLJI&amp;t=3s</a></p>",
      "rawMarkdown": "thank you for the answer and sorry for the late reply. We want to thank you that the codebase has been insightful.\n\nIn the mean time we did make a meme...\n\nhttps://www.youtube.com/watch?v=tKM4xxBoLJI&t=3s",
      "votes": null
    },
    {
      "id": "3087081",
      "postDate": "01/03/2025 04:33:34",
      "content": "<p>\"Boss, you're awesome!\"</p>",
      "rawMarkdown": "\"Boss, you're awesome!\"",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2640481,
      "author_name": "programmaticart",
      "author_url": "",
      "post_date": "02/07/2024 00:37:46",
      "content": "<p>Congratulations! 🎉</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2640483,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "02/07/2024 00:38:35",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/igorkrashenyi\" target=\"_blank\">@igorkrashenyi</a> this is great! Thanks for writing this :). Congrats on your solution!</p>\n<p>If I may ask, what do you mean stackwise normalization based on percentiles?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2640504,
          "author_name": "igorkrashenyi",
          "author_url": "",
          "post_date": "02/07/2024 00:54:37",
          "content": "<p>This normalization is based not on a single image but on a stack of images. <br>\nFirst, you calculate the stats for the stack (kidney_1, kidney_2, etc), each stack will have its own stats.</p>\n<p><code>xmin = np.percentile(volume, low)</code><br>\n<code>xmax = np.max([np.percentile(volume, high), 1])</code></p>\n<p>and afterward apply:<br>\n<code>\nimage = (image - xmin) / (xmax - xmin)</code><br>\n<code>image = np.clip(image, 0, 1)\n</code></p>",
          "votes": null,
          "replies": [
            {
              "id": 2640772,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "02/07/2024 04:46:45",
              "content": "<p>Ah, I see. Makes sense! Thank you </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2640484,
      "author_name": "cody11null",
      "author_url": "",
      "post_date": "02/07/2024 00:38:38",
      "content": "<p>Wow incredible solution! Congrats on your placement! I thought about using 2.5d and 3d ensemble but never made the try. I am feeling like I had a lot of room for improvement on my augmentations as I was only using flips and contrast because some of my images were naturally different sizes! I will learn a lot from this one! Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2640485,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/07/2024 00:38:52",
      "content": "<p>thanks for the write up and congrats for your top ranking and results!!!</p>\n<p>is there any comparison results with and without additional data?<br>\n(you mentioned \"I downloaded the additional data from <a href=\"http://human-organ-atlas.esrf.eu\" target=\"_blank\">http://human-organ-atlas.esrf.eu</a> \")</p>",
      "votes": null,
      "replies": [
        {
          "id": 2640496,
          "author_name": "igorkrashenyi",
          "author_url": "",
          "post_date": "02/07/2024 00:51:25",
          "content": "<p>Yeah, so with the usage of pseudos  I was able to get around 2% on CV, 1% on the public LB, and 4% on private LB. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2640528,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "02/07/2024 01:13:22",
              "content": "<p>thanks for the reply.</p>\n<p>i use boundary loss as well. This is for your reference.<br>\ni  implemented by BCE and interior weights and boundary weights .<br>\nfor boundary weights, it uses  both foreground and background pixels near the object boundary.<br>\n boundary loss produces my top performing models and stablised training as well.</p>\n<p>i note that towards the end of training, interior and boundary BCE completes against each other.<br>\nboundary BCE  prevents the collapse of the boundary predict. hence my validation boundary dice is very stable even after very very long training.   </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2640487,
      "author_name": "timothyzero",
      "author_url": "",
      "post_date": "02/07/2024 00:42:27",
      "content": "<p>Congratulations!  Thx for sharing.😀</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2640492,
      "author_name": "",
      "author_url": "",
      "post_date": "02/07/2024 00:48:08",
      "content": "<p>That's impressive, I've always wanted to try BD Loss but haven't succeeded yet.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2640649,
      "author_name": "mornicen",
      "author_url": "",
      "post_date": "02/07/2024 03:22:21",
      "content": "<p>Congratulations on your good results. Thank you for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2640693,
      "author_name": "shenyangping",
      "author_url": "",
      "post_date": "02/07/2024 03:49:42",
      "content": "<p>Congratulations! Thank you for sharing the results!❤️</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2641130,
      "author_name": "humaperveen",
      "author_url": "",
      "post_date": "02/07/2024 10:00:43",
      "content": "<p>Congratulations, Thank you for sharing and explaining your approach. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2641199,
      "author_name": "sergiosaharovskiy",
      "author_url": "",
      "post_date": "02/07/2024 10:45:13",
      "content": "<p><a href=\"https://www.kaggle.com/igorkrashenyi\" target=\"_blank\">@igorkrashenyi</a> congratz with solo gold and finally gm. By cropping technique you mean you did tiling 512x512 with overlap during inference, is that correct? What was the training time?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2641298,
          "author_name": "igorkrashenyi",
          "author_url": "",
          "post_date": "02/07/2024 12:08:33",
          "content": "<p>Thanks!<br>\nNot exactly. <br>\nFor the 2d model, during the training, I performed 512x512 non-empty mask random cropping of images. During the epoch, I did 2 such crops for each image from all views. In a such setup, one epoch was approximately 1000 batches with batch size 32 and took 20 minutes on 2 RTX 4090.</p>\n<p>For the 3d model, the cropping was balanced 50/50 of empty/non-empty 192x192x192 crops. One epoch here was 2000 samples with batch size 4. In this case epoch took 6 minutes. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2641260,
      "author_name": "forcewithme",
      "author_url": "",
      "post_date": "02/07/2024 11:27:40",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/igorkrashenyi\" target=\"_blank\">@igorkrashenyi</a> , solid and creative work. Congratulations for the solo gold and grandmaster! </p>",
      "votes": null,
      "replies": [
        {
          "id": 2641307,
          "author_name": "igorkrashenyi",
          "author_url": "",
          "post_date": "02/07/2024 12:12:56",
          "content": "<p>thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2641743,
      "author_name": "bitorqubitt",
      "author_url": "",
      "post_date": "02/07/2024 16:35:36",
      "content": "<p>congratz. And thanks for the insightful write up!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2641946,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/07/2024 19:16:05",
      "content": "<p>\"P.S. If you were able to read all of this, the top score on the private LB was a simple mix of 2d and 2.5d models with 1 and 3 channels :) \"</p>\n<p>can i conclude that high or low private score is a bit of luck in model selection, rather than apply some \"crucial\"  methods/steps?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2641957,
          "author_name": "igorkrashenyi",
          "author_url": "",
          "post_date": "02/07/2024 19:27:42",
          "content": "<p>Luck + some tricks in the training and inference process. </p>\n<p>I was not able to select the top model based on only two data points (local CV and public LB), but at some point the shakeup didn’t impact me much. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2642023,
      "author_name": "sirapoabchaikunsaeng",
      "author_url": "",
      "post_date": "02/07/2024 20:57:31",
      "content": "<p>Congratulations on the solo gold! I and my friends want to become kaggle master one day, so we figured reproducing your results is a great start :). In doing so, we're curious about your methodologies of arriving at your solutions: </p>\n<ul>\n<li>what's the motivation behind training the model with 512x512 crops then running the inference at 800x800 crops? Is it purely for inference speed? or is there any accuracy benefit to it?</li>\n<li>how do you come up with the augmentation pipeline? i.e. is this the first and only version of the augmentation pipeline? or do you have some methodology on how you iterate on it?</li>\n<li>same question with the 30 epochs and the cosine scheduler's start and end LR, do you iterate on that much? and do you have some ways you go about iterating on it?</li>\n</ul>\n<p>Thank you for the solution, and once again, congrats!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2642482,
          "author_name": "igorkrashenyi",
          "author_url": "",
          "post_date": "02/08/2024 07:10:36",
          "content": "<p>Thanks a lot!<br>\nAnd thank you for your questions. </p>\n<ul>\n<li>Regarding the inference. It’s just an empirical finding. It boosted the score on both CV and LB by ~0.003.</li>\n<li>iterative process. I’ve started with the severe augmentation scheme, but investigating the learning curves I’ve noticed a heavy underfitting. So I dropped everything and stated to add augmentations one by one. </li>\n<li>same here. It’s not about tuning the number of epochs, it’s more about tuning the amount of batches till plateau. Usually, I start with 100k-250k batches. In this approach “the epoch” is just how often you want to check the validation score. </li>\n</ul>",
          "votes": null,
          "replies": [
            {
              "id": 2649561,
              "author_name": "sirapoabchaikunsaeng",
              "author_url": "",
              "post_date": "02/13/2024 00:59:58",
              "content": "<p>thank you for the answer and sorry for the late reply. We want to thank you that the codebase has been insightful.</p>\n<p>In the mean time we did make a meme…</p>\n<p><a href=\"https://www.youtube.com/watch?v=tKM4xxBoLJI&amp;t=3s\" target=\"_blank\">https://www.youtube.com/watch?v=tKM4xxBoLJI&amp;t=3s</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2642172,
      "author_name": "seungwanhong",
      "author_url": "",
      "post_date": "02/08/2024 01:46:27",
      "content": "<p>This is amazing! Great work! <a href=\"https://www.kaggle.com/igorkrashenyi\" target=\"_blank\">@igorkrashenyi</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2642205,
      "author_name": "jovierrmatthew",
      "author_url": "",
      "post_date": "02/08/2024 02:44:12",
      "content": "<p>very helpful, thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2642226,
      "author_name": "forcewithme",
      "author_url": "",
      "post_date": "02/08/2024 03:19:41",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/igorkrashenyi\" target=\"_blank\">@igorkrashenyi</a> , how large the improvements brought by BoundaryDOULoss, comparing to DiceLoss, on cv, public and private, respectively?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2642345,
          "author_name": "igorkrashenyi",
          "author_url": "",
          "post_date": "02/08/2024 05:19:45",
          "content": "<p>Great question! <br>\nI discovered this loss early in the competition and switched after a few weeks. By that time I did all my experiments with the efficientnet_b3.<br>\nSo on my CV I got +2%, on the public LB the boost was ~1.5%, while on the private leaderboard I’ve got +5%</p>",
          "votes": null,
          "replies": [
            {
              "id": 2642411,
              "author_name": "forcewithme",
              "author_url": "",
              "post_date": "02/08/2024 06:25:16",
              "content": "<p>Amazing! I can hardly witness a loss cause such a huge impact. if the effect of BoundaryDOULoss can be validated on several more datasets, maybe it will become a priority loss in image segmentation task.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2642415,
              "author_name": "forcewithme",
              "author_url": "",
              "post_date": "02/08/2024 06:26:40",
              "content": "<p>Or maybe BoundaryDOULoss match the surface dice(metrics) better? </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2642485,
                  "author_name": "igorkrashenyi",
                  "author_url": "",
                  "post_date": "02/08/2024 07:13:58",
                  "content": "<p>I guess both. But for this specific task and surface Dice metric it fit perfectly, so the boost is so huge. </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2642566,
                      "author_name": "forcewithme",
                      "author_url": "",
                      "post_date": "02/08/2024 08:59:38",
                      "content": "<p>Yes, it does fit this metric perfectly. Thanks for introduce this loss.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2644742,
      "author_name": "igorkrashenyi",
      "author_url": "",
      "post_date": "02/09/2024 16:37:41",
      "content": "<p>added the source code to the post </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2644772,
      "author_name": "tanishqdublish",
      "author_url": "",
      "post_date": "02/09/2024 17:07:17",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/igorkrashenyi\" target=\"_blank\">@igorkrashenyi</a> this is great! Thanks for writing this :). Congrats on your solution!</p>\n<p>If I may ask, what do you mean stackwise normalization based on percentiles?<br>\nThis is amazing! Great work!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2644870,
          "author_name": "igorkrashenyi",
          "author_url": "",
          "post_date": "02/09/2024 18:15:32",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/475052#2640504\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/475052#2640504</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3087081,
      "author_name": "yolov5ssd",
      "author_url": "",
      "post_date": "01/03/2025 04:33:34",
      "content": "<p>\"Boss, you're awesome!\"</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2640468": "First of all, I would like to start my solution description with a few important words:\n\n*I would like to thank the Armed Forces of Ukraine, the Security Service of Ukraine, Defence Intelligence of Ukraine, and the State Emergency Service of Ukraine for providing safety and security to participate in this great competition, complete this work, and help science, technology, and business not to stop but to move forward.*\n\n# Context\n- Business context: https://www.kaggle.com/competitions/blood-vessel-segmentation \n- Data context: https://www.kaggle.com/competitions/blood-vessel-segmentation/data\n\n# Overview of the approach:\nMy final model is a mixture of 2d and 3d models with d4 tta. For the 2d model, the multiview tta was applied. All models were trained in a 2-fold setup with kidney_2 and kidney_3_dense selected as validation sets. The ensembling was performed with equal weights for both 2d and 3d models.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F207760%2F0e15deeafd39981c337888fdaf27e2c9%2FScreenshot%202024-02-07%20at%2002.40.31.png?generation=1707266453616337&alt=media)\n\n# Details of the submission\n\n## Data preparation and training data and validation scheme \n\nAll final (3d and 2d) models were trained on kidney_1_dense, kidney_2, kidney_3_dense, kidney_3_sparse and pseudo labels [50um_LADAF-2020-31_kidney_pag-0.01_0.02_jp2_](http://human-organ-atlas.esrf.eu). Initially, I used slice-wise normalization to normalize images but later switched to stack-wise normalization based on percentiles.\n\nThe 2D model was trained in a multiview setup: all images were stacked in a tensor and sliced in different axes afterward. During the training, the set of augmentations and sampling strategy was crucial. The weighted sampling was based on sparsity percentage: dense samples had a weight of 1, while sparse samples had a weight equal to their sparsity. For pseudo labels,  I chose the same weight as for kidney_2, e.g.: \n\n```python\nkidney_1_dense: 1, \nkidney_2: 0.65, \nkidney_3_dense: 1, \nkidney_3_sparse: 0.85, \n50um_LADAF-2020-31_kidney_pag-0.01_0.02_jp2_: 0.65. \n```\n\nThe augmentation scheme was the next one, with a chance of 0.5 CutMix augmentation being applied. The cropping was performed from the same organ and the same projection axis. Afterward, on top of CutMix, the next augmentation pipeline was applied:\n\n```python\nA.Compose(\n    [\n          A.PadIfNeeded(*crop_size),\n          A.CropNonEmptyMaskIfExists(*crop_size, p=1.0),\n          A.ShiftScaleRotate(scale_limit=0.2),\n          A.HorizontalFlip(p=0.5),\n          A.VerticalFlip(p=0.5),\n          A.RandomRotate90(p=0.5),\n          A.OneOf([\n                A.RandomBrightnessContrast(), \n                A.RandomBrightness(), \n                A.RandomGamma(),\n          ],p=1.0,),\n    ],p=1.0,)\n```\n\nThe crop size was set to 512. I’ve also tried higher resolution, but it performs +- the same result. \n\nI did some experiments with 2.5d approaches (3 and 5 channels), but it produced the same result or worse. \n\nThe 3d model augmentation scheme contained only d4 augmentations and random crops. The cropping was performed with a 0.5 probability of an empty mask. This was motivated by false positives that appeared outside the kidney volume. This could be improved by incorporating the two-class 3D segmentation, but I didn’t have much time and resources to perform such an experiment. Thus, I decided to create a post-processing that would handle this. \nThe crop size for the 3d model was 192x192x192.\n\nBoth models were trained in a 2-fold setup where as validation, I used kidney_2 (fold_1) and kidney_3_dense (fold_0). Removal of kidney_1 from the training set caused performance degradation in performance in both CV and LB, so I dropped the fold_2 and didn't perform training in that setup.\n\n## Model setup\nThe best results I was able to get using the efficientnet family models with UnetPlusPlus decoder and SCSE attention from the segmentation_models_pytorch library. I’ve tried the resnet50 model, like it was mentioned in the discussion section, different transformers and seresnext models, but could overcome the performance of efficientnet-b5 (which performed the best on both CV and LB). On my local validation, the score I was able to get with efficientnet_b7 encoder and mit_b5 encoder, but on the LB the score was significantly lower.\nThe training was performed for 30 epochs with a Cosine LR scheduler starting from 3e-4 to 1e-6. I saved the top 3 checkpoints and used the best-last checkpoint for the submission.\nThe model efficientnet_b5_UnetPlusPlus trained in such a setup was able to score 0.878 on the public LB and 0.714 on the private LB at a 0.05 threshold. \n\nHere yellow is TP, green is FP, and red is FN.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F207760%2F0484711261976f944aca4341305752f2%2FindividualImage.png?generation=1707265123800657&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F207760%2F157fa4895ef81f6e1fc0dbb8e685c17d%2FindividualImage-2.png?generation=1707265135912886&alt=media)\n\nThe 3d model was heavily inspired by the nnUnet model architecture and was pretty much the same. Instead of the native nnUnet model, I used DynUnet from the monai library with almost default configuration and trained in almost the same setup as for nnUnet. As the optimizer, I used SGD with initial LR 0.01 and the Cosine Annealing LR scheme instead of LinearLR and trained for 500 epochs with 2000 samples per epoch. \nThis model scored 0.869 (0.868 and 0.866 -- 0 and 1 folds respectively) on the public LB and 0.694 on the private LB (0.758 and 0.663 -- 0 and 1 folds respectively). \n\nBoth models were trained using the BoundaryDOULoss (https://arxiv.org/pdf/2308.00220.pdf), which performed the best. I’ve tried to modify it to perform better on sparse data but failed. \n\n## Pseudo labeling \n\nBased on the preprint, I downloaded the additional data from http://human-organ-atlas.esrf.eu site (2 datasets). It appeared, that one of the datasets overlaps with the kidney_3, so I dropped it to prevent leakage. I used the other one to generate pseudo labels. For pseudo labeling, I used an ensemble of 2d models (efficientnet-b5 and efficientnet-b6 with UnetPlusPlus) trained with the same setup but without CutMix. The correct setup of CutMix as well as the 3d model I was able to discover close to the competition deadline, so I didn’t retrain the original ensemble and stick to the first version of pseudos. \n\n## Inference setup and Post-processing \n\nThe inference for both models was performed using sliding_window_inference from monai library. Additionally, for 2d model I performed multi-view tta, which helped to detect small vessels and improve overall performance. \n\nFor the 2d model, the crop size was 800 pix, while for the 3d – 256 pix with 0.25 overlap and Gaussian merging. All models used d4_transform from ttach library. I’ve forked the ttach repository and implemented the logic for 3d images, but the inference time increased significantly, and there was no major boost in performance, so I’ve sticked with 2d d4_transform for both 2d and 3d models :)\n\nAs I mentioned before, the 3d model had decent performance on the non-empty cubes, while empty ones were confusing the model. To handle this issue, decided to experiment with post-processing. The idea was the next one: let's try to find ROI where the vessels were presented. Since the 2d model didn’t have such a problem I’ve decided to find a bounding polygon for vessels for each 2d slice. Having a mask of ROI, I multiplied it with 3d model predictions and got a boost from 0.869 to 0.881 public LB and 0.701 private LB for a single 3d model.\n\nEnsembling the 2d model and 3d model predictions with weights 1 and 1, I was able to improve the score from 0.881 to 0.884 on the public LB and 0.712 on the private LB.\n\nAnother post-processing approach that I’ve tried is to use Canny filters from cv2 to segment the kidney. This segmentation algorithm was not perfect, but applying such post-processing boosted my score from 0.884 to 0.892 on the public LB while failing on the private LB, scoring just 0.313.\n\n## What didn’t work\n- nnUnet out of the box. At the beginning of the challenge and after the pre-print reading, I tried to reproduce the result with nnUnet. The local score was promising, but the LB was 0. My intuition behind this issue related to the data normalization and spacing (scale), but I didn’t try to fix it and decided to build my own solution.\n- BCE and Focal Loss.\n- Transformers in both 2d and 3d model\n- Zoom and brightness augmentation for 3d images\n- Pseudo on top of sparse datasets. I’ve tried to fulfill the sparsity of the dataset by pseudo labeling and aggregation, but it didn’t improve the score.\n- Additional projections. I’ve performed experiments with 2d models and additional slices generated from the 3d stack, but LB performance dropped by 20% while CV was about the same. \n- Auxiliary outputs such as distance transform or center of mass. \n- **and the most important: validation**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F207760%2Fcd379db59f3575778ebc4eca63ef0189%2FScreenshot%202024-02-07%20at%2002.29.01.png?generation=1707265761683297&alt=media)\n\nP.S. If you were able to read all of this, the top score on the private LB was a simple mix of 2d and 2.5d models with 1 and 3 channels :) \n\nP.P.S. Thank you for reading!\n\n## Links\n- [Pseudo labels ](https://www.kaggle.com/datasets/igorkrashenyi/50um-ladaf-2020-31-kidney-pag-0-01-0-02-jp2 )\n- [Source code](https://github.com/burnmyletters/blood-vessel-segmentation-public)\n- Inference code https://www.kaggle.com/code/igorkrashenyi/4th-place-solution/notebook + https://www.kaggle.com/code/igorkrashenyi/fork-of-multiview-2-5-sennet-hoa-inference-v3",
    "2640481": "Congratulations! 🎉",
    "2640483": "Hey @igorkrashenyi this is great! Thanks for writing this :). Congrats on your solution!\n\nIf I may ask, what do you mean stackwise normalization based on percentiles?",
    "2640484": "Wow incredible solution! Congrats on your placement! I thought about using 2.5d and 3d ensemble but never made the try. I am feeling like I had a lot of room for improvement on my augmentations as I was only using flips and contrast because some of my images were naturally different sizes! I will learn a lot from this one! Thank you for sharing!",
    "2640485": "thanks for the write up and congrats for your top ranking and results!!!\n\nis there any comparison results with and without additional data?\n(you mentioned \"I downloaded the additional data from http://human-organ-atlas.esrf.eu \")",
    "2640487": "Congratulations!  Thx for sharing.😀",
    "2640492": "That's impressive, I've always wanted to try BD Loss but haven't succeeded yet.",
    "2640496": "Yeah, so with the usage of pseudos  I was able to get around 2% on CV, 1% on the public LB, and 4% on private LB.",
    "2640504": "This normalization is based not on a single image but on a stack of images. \nFirst, you calculate the stats for the stack (kidney_1, kidney_2, etc), each stack will have its own stats.\n\n`xmin = np.percentile(volume, low)`\n`xmax = np.max([np.percentile(volume, high), 1])`\n\nand afterward apply:\n`\nimage = (image - xmin) / (xmax - xmin)`\n`image = np.clip(image, 0, 1)\n`",
    "2640528": "thanks for the reply.\n\ni use boundary loss as well. This is for your reference.\ni  implemented by BCE and interior weights and boundary weights .\nfor boundary weights, it uses  both foreground and background pixels near the object boundary.\n boundary loss produces my top performing models and stablised training as well.\n\ni note that towards the end of training, interior and boundary BCE completes against each other.\nboundary BCE  prevents the collapse of the boundary predict. hence my validation boundary dice is very stable even after very very long training.",
    "2640649": "Congratulations on your good results. Thank you for sharing.",
    "2640693": "Congratulations! Thank you for sharing the results!❤️",
    "2640772": "Ah, I see. Makes sense! Thank you",
    "2641130": "Congratulations, Thank you for sharing and explaining your approach.",
    "2641199": "igorkrashenyi congratz with solo gold and finally gm. By cropping technique you mean you did tiling 512x512 with overlap during inference, is that correct? What was the training time?",
    "2641260": "Hi @igorkrashenyi , solid and creative work. Congratulations for the solo gold and grandmaster!",
    "2641298": "Thanks!\nNot exactly. \nFor the 2d model, during the training, I performed 512x512 non-empty mask random cropping of images. During the epoch, I did 2 such crops for each image from all views. In a such setup, one epoch was approximately 1000 batches with batch size 32 and took 20 minutes on 2 RTX 4090.\n\nFor the 3d model, the cropping was balanced 50/50 of empty/non-empty 192x192x192 crops. One epoch here was 2000 samples with batch size 4. In this case epoch took 6 minutes.",
    "2641307": "thank you!",
    "2641743": "congratz. And thanks for the insightful write up!",
    "2641946": "\"P.S. If you were able to read all of this, the top score on the private LB was a simple mix of 2d and 2.5d models with 1 and 3 channels :) \"\n\ncan i conclude that high or low private score is a bit of luck in model selection, rather than apply some \"crucial\"  methods/steps?",
    "2641957": "Luck + some tricks in the training and inference process. \n\nI was not able to select the top model based on only two data points (local CV and public LB), but at some point the shakeup didn’t impact me much.",
    "2642023": "Congratulations on the solo gold! I and my friends want to become kaggle master one day, so we figured reproducing your results is a great start :). In doing so, we're curious about your methodologies of arriving at your solutions: \n\n- what's the motivation behind training the model with 512x512 crops then running the inference at 800x800 crops? Is it purely for inference speed? or is there any accuracy benefit to it?\n- how do you come up with the augmentation pipeline? i.e. is this the first and only version of the augmentation pipeline? or do you have some methodology on how you iterate on it?\n- same question with the 30 epochs and the cosine scheduler's start and end LR, do you iterate on that much? and do you have some ways you go about iterating on it?\n\nThank you for the solution, and once again, congrats!",
    "2642172": "This is amazing! Great work! @igorkrashenyi",
    "2642205": "very helpful, thanks",
    "2642226": "Hi @igorkrashenyi , how large the improvements brought by BoundaryDOULoss, comparing to DiceLoss, on cv, public and private, respectively?",
    "2642345": "Great question! \nI discovered this loss early in the competition and switched after a few weeks. By that time I did all my experiments with the efficientnet_b3.\nSo on my CV I got +2%, on the public LB the boost was ~1.5%, while on the private leaderboard I’ve got +5%",
    "2642411": "Amazing! I can hardly witness a loss cause such a huge impact. if the effect of BoundaryDOULoss can be validated on several more datasets, maybe it will become a priority loss in image segmentation task.",
    "2642415": "Or maybe BoundaryDOULoss match the surface dice(metrics) better?",
    "2642482": "Thanks a lot!\nAnd thank you for your questions. \n- Regarding the inference. It’s just an empirical finding. It boosted the score on both CV and LB by ~0.003.\n- iterative process. I’ve started with the severe augmentation scheme, but investigating the learning curves I’ve noticed a heavy underfitting. So I dropped everything and stated to add augmentations one by one. \n- same here. It’s not about tuning the number of epochs, it’s more about tuning the amount of batches till plateau. Usually, I start with 100k-250k batches. In this approach “the epoch” is just how often you want to check the validation score.",
    "2642485": "I guess both. But for this specific task and surface Dice metric it fit perfectly, so the boost is so huge.",
    "2642566": "Yes, it does fit this metric perfectly. Thanks for introduce this loss.",
    "2644742": "added the source code to the post",
    "2644772": "Hey @igorkrashenyi this is great! Thanks for writing this :). Congrats on your solution!\n\nIf I may ask, what do you mean stackwise normalization based on percentiles?\nThis is amazing! Great work!",
    "2644870": "https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/475052#2640504",
    "2649561": "thank you for the answer and sorry for the late reply. We want to thank you that the codebase has been insightful.\n\nIn the mean time we did make a meme...\n\nhttps://www.youtube.com/watch?v=tKM4xxBoLJI&t=3s",
    "3087081": "\"Boss, you're awesome!\""
  },
  "source": "meta"
}