{
  "id": 475311,
  "title": "93rd Place Solution for the SenNet + HOA - Hacking the Human Vasculature in 3D",
  "url": "/competitions/blood-vessel-segmentation/discussion/475311",
  "author_name": "velangovan",
  "post_date": "2024-02-07T22:02:08.569000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I started late in the beginning of January. This is only my second kaggle competition and I am happy with the enormous learning and the outcome of placing in the top 10% with bronze. Thank you to the fellow teams for sharing your knowledge and to the organizers for a well-run competition.</p>\n<h1>Context</h1>\n<p>Business context: <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/overview\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/overview</a><br>\nData context: <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/data\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/data</a></p>\n<h1>Overview of the approach</h1>\n<p>My final model was an ensemble of two UNet 2D models trained with 1024x1024 and 512x512 sizes with equal weighting in inference. I used slices from X,Y and Z projections for both training and inference, a simplification inspired by the 2.5D Unet paper. </p>\n<p>For training, I used segmentation_models_pytorch smp.Unet architecture with resnext50_32x4d backbone and started with imagenet weights. Trained with kidney_1_dense and kidney_3_dense and validated with kidney_2. For preprocessing, used histogram equalization followed by minmax normalization. Adam optimizer, CosineAnnealingLR scheduler and smp.losses.DiceLoss were used for training. Following augmentations were used in training.</p>\n<pre><code>aug_prob     =  \ntrain_aug_list = [\n        A.Rotate(limit=, p=aug_prob),\n        A.RandomScale(scale_limit=(,),interpolation=cv2.INTER_CUBIC,p=aug_prob),\n        A.PadIfNeeded(min_height=img_size[], min_width=img_size[], p=),\n        A.RandomCrop(img_size[], img_size[], p=),\n        A.RandomBrightnessContrast(p=aug_prob),\n        A.GaussianBlur(p=aug_prob),\n        A.MotionBlur(p=aug_prob),\n        A.GridDistortion(num_steps=, distort_limit=, p=aug_prob),\n        ToTensorV2(transpose_mask=),\n    ]\n</code></pre>\n<p>For inference, I padded the image to 3072x3072 and used a 3x3 grid of 1024x1024 size tiles to run the model on. 4 rotations used on each tile for inference and averaged for Test Time Augmentation aka TTA. Used sigmoid activation layer in Unet. Also, ran inference on X, Y and Z projections and created three prediction volumes. On each projection’s prediction volume, applied sigmoid threshold of 0.0001 to get binary mask volumes, then transposed and added binary masks from the three projections, then used a majority voting to get final predictions.</p>\n<h1>Details of the submission</h1>\n<h2>What was special about the submission</h2>\n<p>-I did no resizing in training or inference to reduce errors and artifacts from downsizing and aspect ratio changes. In training, I used a random crop. In inference, I used padding and tiling.<br>\n-Fixed sigmoid threshold independent of dataset, as opposed to using top Nth percentile for thresholding produced more stable result on private LB. I was bumped up by 616 in ranking.<br>\n-Ensemble of 1024+512 gave higher score than each model applied separately.<br>\n-4x rotation Test Time Augmentation in inference boosted the score by 0.007 in public LB.<br>\n-During local validation and spot checking, my models were achieving very high 2D dice scores on the middle slices in the volumes. Most of the FP and FN errors were on the edge slices in the tiny vessels (1 or 2 pixel errors).<br>\n-While I thought histogram equalization was a secret sauce (since it improved public LB score over mean/std/clip normalization), it turns out my mean/std/clip normalization model produced much better private LB score.<br>\n-There were many high scoring inference notebooks publicly shared in this competition. While I studied them to understand what other teams are doing and did in fact get many great ideas, I chose not to use large sections of code directly since there were many questionable choices and I didn’t understand why those notebooks were producing the high public LB scores. This approach kept my solution unique and generalized enough to move up in the private LB.</p>\n<h2>What was tried and didn’t work</h2>\n<p>-The striding on 3 or 5 consecutive slices to create a multi-channel image for training and inference did not work well for me, both score wise and CPU/GPU/Memory resource wise.<br>\n-Tried median blur for preprocessing which did not work well.<br>\n-Went from efficientnetb0 to resnext50_32x4d. Maybe somewhere in between would have been better.<br>\n-It would have been better to stick with mean/std/clip norm as opposed to histogram equalization for preprocessing<br>\n-Tried inference with Z projection only initially, after 0.04 improvement in LB score with XYZ projections and majority voting, and decided to use it going forward.<br>\n-There were suggestions in public high-scoring notebooks to reduce augmentation probability to 0.05, I tried low augmentation and although that produced higher validation dice scores during training, almost always produced lower score in public LB. So, I decided to increase to 0.10 probability for augmentation, now that I read the solution writeups from top scoring teams, I realize an even higher augmentation would have been better<br>\n-I had tried training with 90% of slices of all three kidneys dense 1, 2 and 3 and validating with 10% of the slices. Since I read in the discussion board that there is label shift in kidney 2, I switched to using only kidney1 and kidney 3 for training. The results got slightly better. Every incremental improvement counted.<br>\n-Since I realized most of segmentation errors were in the first 100 slices or so with the tiny vessels, I tried to train a separate model with first 200 slices of all three kidneys. Model didn’t work at all on both early slices and middle slices, probably because there was not enough data to train.<br>\n-Tried morphological opening for post processing the predicted mask. I expected that it will reduce some false positives without affecting true positives that much. But nope. It made both false positives and false negatives significantly worse. Then, tried removing all one pixel blobs for post processing. It was a disaster with the score. Also, tried majority voting across three slices to retain positives, it was a disaster. After that, I decided that no post processing is best. Any improvement had to come from better model, not predicted mask cleanup.<br>\n-Tried 6x and 8x TTA with flips and rotations, but score was worse than 4x TTA.</p>\n<h2>What I didn’t try which I would consider next time</h2>\n<ol>\n<li><p>A loss function that considers 2D boundary loss since the “surface dice score” used for test evaluation uses only the 3D surface boundaries. 5th place solution used a clever loss function of CE_boundaries + Dice + Focal. 4th place solution used a BoundaryDOULoss (<a href=\"https://arxiv.org/pdf/2308.00220.pdf\" target=\"_blank\">https://arxiv.org/pdf/2308.00220.pdf</a>)</p></li>\n<li><p>Take into account that the native scan resolution of private test data is different (63um/voxel) from train (50 and 50.16um) and public test data (50.28um/voxel). This can be accomplished with 2D or 3D resize during inference or equivalent scaling augmentations in training. Top 5 solution write ups talk about this.</p></li>\n<li><p>Somehow incorporate the sparse annotations in training. I completely ignored the kidney_3_sparse and kidney_2 (also sparse and had shifted labels) in training which limited my training data to kidney_1_dense and the small number of slices from kidney_3_dense. Pseudo-annotating the sparse data is an approach I saw that many of the top scoring solutions used.</p></li>\n<li><p>Experiment more with different model architectures and ensembles? I was bogged down with the basics and getting the foundations working. Also, with the Kaggle 30 hours per week GPU limit, didn’t have time and resources for this.</p></li>\n<li><p>Increase the augmentation probabilities in intensity and scaling. Many of the top solutions have done this.</p></li>\n<li><p>2.5D approach with 5 consecutive slices concatenated as a 5-channel image. 6th place solution has successfully used this. I was running into CPU/GPU/Mem resource issues, but those could have been potentially resolved with more time and effort. </p></li>\n</ol>\n<h2>Preprocessing examples</h2>\n<p>Histogram equalization followed by Minmax normalization:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2Fe119a342eda6f290c058e94c97728e78%2FNorm-HistogramEqualization.png?generation=1707343051563429&amp;alt=media\"></p>\n<p>Mean/Std normalization with clipping followed by Minmax normalization:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2F9e72cb4f6f49eee6aae6213df2833e90%2FNorm-MeanStdClip.png?generation=1707343078883718&amp;alt=media\"></p>\n<h2>Padding and tiling kidney slice for inference</h2>\n<p>This is showing 2048x2048 padding and 512x512 tiles which I later changed to 3072x3072 padding and 1024x1024 tiles.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2F4acb3f75490749342dab0c8e3f7f91de%2FTiledKidneySlice.png?generation=1707342738234654&amp;alt=media\"></p>\n<h2>Sample segmentation results on slices for illustration</h2>\n<p>Slice 1000 from kidney_1_dense Label Vs Prediction. Prediction is color coded as Green for True Positives, Red for false positives and Blue for false negatives. Also all blobs are dilated 5x5 to observe the tiny blobs visually. Most of the false positives and false negatives are 1 pixel area blobs. However, there are also some 1 pixel area true positive blobs in this slice and many more of those in early slices for example 0100. So, cannot blindly remove 1 pixel blobs.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2F3d16023f25a7f35a225c15c2d88cc107%2FSegmentedKidney1_1000.png?generation=1707342563232471&amp;alt=media\"></p>\n<p>Slice 0100 from kidney_1_dense Label Vs Prediction. 119 pixels in label, 133 pixels found, 93 pixels true positives, 40 pixels false positives, 26 pixels false negatives. Many of the blob sizes are tiny 1 to 4 pixels area.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2F632bff25ba557b49017796618251f092%2FSegmentedKidney1_0100.png?generation=1707342598764919&amp;alt=media\"></p>\n<h1>Updates After the competition ended</h1>\n<p>After the competition ended, some new opportunities opened up for learning. First, I read the top solution write-ups looking for inspiration, especially low-hanging fruit ideas for adding to my existing implementation. Second, the private and public scores are now visible, enabling us to know how we do on the two test sets. Third, the 5 per day submission limit is lifted, so can do a lot more experiments more quickly. Given these, I was able to improve my private score to 0.682 which would have been 6th place (however not genuinely since I would never have chosen this submission due to the public score being so low). Well, anyway, here it is, my best improved score.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2F574ddf3fa25f91bab4dad90f81efd542%2FSenNetHOA-BestPrivateScore.png?generation=1707777659990418&amp;alt=media\"></p>\n<p>Here are the ideas I added to my implementation to achieve the above score.</p>\n<ol>\n<li>Changed histogram equalization preprocessing to mean/std/clip normalization (I learned from my own prev submissions scoring higher in private score)</li>\n<li>Changed DiceLoss to custom loss that is a combination of boundary loss, dice loss and focal loss (shared by 5th place solution)</li>\n<li>Changed the unet architecture to include an upscale layer (shared by 5th place solution)</li>\n<li>Changed to heavy augmentations in scale and intensity (strategy used by several top solutions)</li>\n<li>Used -0.45, 0.05 for scale limit to downsize more than upsize, to account for the private test set resolution. (shared by 3rd place solution)</li>\n</ol>\n<h1>Sources</h1>\n<p>Following kaggle sources were hugely helpful, many thanks to the contributors.</p>\n<p><a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456118\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456118</a><br>\n<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/464768\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/464768</a><br>\n<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/468525\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/468525</a><br>\n<a href=\"https://www.kaggle.com/code/junkoda/fast-surface-dice-computation\" target=\"_blank\">https://www.kaggle.com/code/junkoda/fast-surface-dice-computation</a><br>\n<a href=\"https://www.kaggle.com/code/hengck23/lb0-808-resnet50-2d-unet-xy-zy-zx-cc3d\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb0-808-resnet50-2d-unet-xy-zy-zx-cc3d</a><br>\n<a href=\"https://www.kaggle.com/code/yoyobar/2-5d-cutting-model-baseline-inference\" target=\"_blank\">https://www.kaggle.com/code/yoyobar/2-5d-cutting-model-baseline-inference</a><br>\n<a href=\"https://www.kaggle.com/code/yoyobar/2-5d-cutting-model-baseline-training\" target=\"_blank\">https://www.kaggle.com/code/yoyobar/2-5d-cutting-model-baseline-training</a><br>\n<a href=\"https://www.kaggle.com/code/misakimatsutomo/inference-1024-should-have-a-percentile-of-0-00149\" target=\"_blank\">https://www.kaggle.com/code/misakimatsutomo/inference-1024-should-have-a-percentile-of-0-00149</a></p>\n<p>Other sources I used for background and inspiration:<br>\n<a href=\"https://doi.org/10.48550/arXiv.2311.13319\" target=\"_blank\">https://doi.org/10.48550/arXiv.2311.13319</a><br>\n<a href=\"https://arxiv.org/abs/1902.00347\" target=\"_blank\">https://arxiv.org/abs/1902.00347</a><br>\n<a href=\"https://arxiv.org/abs/2010.0616\" target=\"_blank\">https://arxiv.org/abs/2010.0616</a></p>",
  "messages": [
    {
      "id": 2642074,
      "postDate": "2024-02-07T22:02:08.570Z",
      "content": "<p>I started late in the beginning of January. This is only my second kaggle competition and I am happy with the enormous learning and the outcome of placing in the top 10% with bronze. Thank you to the fellow teams for sharing your knowledge and to the organizers for a well-run competition.</p>\n<h1>Context</h1>\n<p>Business context: <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/overview\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/overview</a><br>\nData context: <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/data\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/data</a></p>\n<h1>Overview of the approach</h1>\n<p>My final model was an ensemble of two UNet 2D models trained with 1024x1024 and 512x512 sizes with equal weighting in inference. I used slices from X,Y and Z projections for both training and inference, a simplification inspired by the 2.5D Unet paper. </p>\n<p>For training, I used segmentation_models_pytorch smp.Unet architecture with resnext50_32x4d backbone and started with imagenet weights. Trained with kidney_1_dense and kidney_3_dense and validated with kidney_2. For preprocessing, used histogram equalization followed by minmax normalization. Adam optimizer, CosineAnnealingLR scheduler and smp.losses.DiceLoss were used for training. Following augmentations were used in training.</p>\n<pre><code>aug_prob     =  \ntrain_aug_list = [\n        A.Rotate(limit=, p=aug_prob),\n        A.RandomScale(scale_limit=(,),interpolation=cv2.INTER_CUBIC,p=aug_prob),\n        A.PadIfNeeded(min_height=img_size[], min_width=img_size[], p=),\n        A.RandomCrop(img_size[], img_size[], p=),\n        A.RandomBrightnessContrast(p=aug_prob),\n        A.GaussianBlur(p=aug_prob),\n        A.MotionBlur(p=aug_prob),\n        A.GridDistortion(num_steps=, distort_limit=, p=aug_prob),\n        ToTensorV2(transpose_mask=),\n    ]\n</code></pre>\n<p>For inference, I padded the image to 3072x3072 and used a 3x3 grid of 1024x1024 size tiles to run the model on. 4 rotations used on each tile for inference and averaged for Test Time Augmentation aka TTA. Used sigmoid activation layer in Unet. Also, ran inference on X, Y and Z projections and created three prediction volumes. On each projection’s prediction volume, applied sigmoid threshold of 0.0001 to get binary mask volumes, then transposed and added binary masks from the three projections, then used a majority voting to get final predictions.</p>\n<h1>Details of the submission</h1>\n<h2>What was special about the submission</h2>\n<p>-I did no resizing in training or inference to reduce errors and artifacts from downsizing and aspect ratio changes. In training, I used a random crop. In inference, I used padding and tiling.<br>\n-Fixed sigmoid threshold independent of dataset, as opposed to using top Nth percentile for thresholding produced more stable result on private LB. I was bumped up by 616 in ranking.<br>\n-Ensemble of 1024+512 gave higher score than each model applied separately.<br>\n-4x rotation Test Time Augmentation in inference boosted the score by 0.007 in public LB.<br>\n-During local validation and spot checking, my models were achieving very high 2D dice scores on the middle slices in the volumes. Most of the FP and FN errors were on the edge slices in the tiny vessels (1 or 2 pixel errors).<br>\n-While I thought histogram equalization was a secret sauce (since it improved public LB score over mean/std/clip normalization), it turns out my mean/std/clip normalization model produced much better private LB score.<br>\n-There were many high scoring inference notebooks publicly shared in this competition. While I studied them to understand what other teams are doing and did in fact get many great ideas, I chose not to use large sections of code directly since there were many questionable choices and I didn’t understand why those notebooks were producing the high public LB scores. This approach kept my solution unique and generalized enough to move up in the private LB.</p>\n<h2>What was tried and didn’t work</h2>\n<p>-The striding on 3 or 5 consecutive slices to create a multi-channel image for training and inference did not work well for me, both score wise and CPU/GPU/Memory resource wise.<br>\n-Tried median blur for preprocessing which did not work well.<br>\n-Went from efficientnetb0 to resnext50_32x4d. Maybe somewhere in between would have been better.<br>\n-It would have been better to stick with mean/std/clip norm as opposed to histogram equalization for preprocessing<br>\n-Tried inference with Z projection only initially, after 0.04 improvement in LB score with XYZ projections and majority voting, and decided to use it going forward.<br>\n-There were suggestions in public high-scoring notebooks to reduce augmentation probability to 0.05, I tried low augmentation and although that produced higher validation dice scores during training, almost always produced lower score in public LB. So, I decided to increase to 0.10 probability for augmentation, now that I read the solution writeups from top scoring teams, I realize an even higher augmentation would have been better<br>\n-I had tried training with 90% of slices of all three kidneys dense 1, 2 and 3 and validating with 10% of the slices. Since I read in the discussion board that there is label shift in kidney 2, I switched to using only kidney1 and kidney 3 for training. The results got slightly better. Every incremental improvement counted.<br>\n-Since I realized most of segmentation errors were in the first 100 slices or so with the tiny vessels, I tried to train a separate model with first 200 slices of all three kidneys. Model didn’t work at all on both early slices and middle slices, probably because there was not enough data to train.<br>\n-Tried morphological opening for post processing the predicted mask. I expected that it will reduce some false positives without affecting true positives that much. But nope. It made both false positives and false negatives significantly worse. Then, tried removing all one pixel blobs for post processing. It was a disaster with the score. Also, tried majority voting across three slices to retain positives, it was a disaster. After that, I decided that no post processing is best. Any improvement had to come from better model, not predicted mask cleanup.<br>\n-Tried 6x and 8x TTA with flips and rotations, but score was worse than 4x TTA.</p>\n<h2>What I didn’t try which I would consider next time</h2>\n<ol>\n<li><p>A loss function that considers 2D boundary loss since the “surface dice score” used for test evaluation uses only the 3D surface boundaries. 5th place solution used a clever loss function of CE_boundaries + Dice + Focal. 4th place solution used a BoundaryDOULoss (<a href=\"https://arxiv.org/pdf/2308.00220.pdf\" target=\"_blank\">https://arxiv.org/pdf/2308.00220.pdf</a>)</p></li>\n<li><p>Take into account that the native scan resolution of private test data is different (63um/voxel) from train (50 and 50.16um) and public test data (50.28um/voxel). This can be accomplished with 2D or 3D resize during inference or equivalent scaling augmentations in training. Top 5 solution write ups talk about this.</p></li>\n<li><p>Somehow incorporate the sparse annotations in training. I completely ignored the kidney_3_sparse and kidney_2 (also sparse and had shifted labels) in training which limited my training data to kidney_1_dense and the small number of slices from kidney_3_dense. Pseudo-annotating the sparse data is an approach I saw that many of the top scoring solutions used.</p></li>\n<li><p>Experiment more with different model architectures and ensembles? I was bogged down with the basics and getting the foundations working. Also, with the Kaggle 30 hours per week GPU limit, didn’t have time and resources for this.</p></li>\n<li><p>Increase the augmentation probabilities in intensity and scaling. Many of the top solutions have done this.</p></li>\n<li><p>2.5D approach with 5 consecutive slices concatenated as a 5-channel image. 6th place solution has successfully used this. I was running into CPU/GPU/Mem resource issues, but those could have been potentially resolved with more time and effort. </p></li>\n</ol>\n<h2>Preprocessing examples</h2>\n<p>Histogram equalization followed by Minmax normalization:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2Fe119a342eda6f290c058e94c97728e78%2FNorm-HistogramEqualization.png?generation=1707343051563429&amp;alt=media\"></p>\n<p>Mean/Std normalization with clipping followed by Minmax normalization:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2F9e72cb4f6f49eee6aae6213df2833e90%2FNorm-MeanStdClip.png?generation=1707343078883718&amp;alt=media\"></p>\n<h2>Padding and tiling kidney slice for inference</h2>\n<p>This is showing 2048x2048 padding and 512x512 tiles which I later changed to 3072x3072 padding and 1024x1024 tiles.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2F4acb3f75490749342dab0c8e3f7f91de%2FTiledKidneySlice.png?generation=1707342738234654&amp;alt=media\"></p>\n<h2>Sample segmentation results on slices for illustration</h2>\n<p>Slice 1000 from kidney_1_dense Label Vs Prediction. Prediction is color coded as Green for True Positives, Red for false positives and Blue for false negatives. Also all blobs are dilated 5x5 to observe the tiny blobs visually. Most of the false positives and false negatives are 1 pixel area blobs. However, there are also some 1 pixel area true positive blobs in this slice and many more of those in early slices for example 0100. So, cannot blindly remove 1 pixel blobs.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2F3d16023f25a7f35a225c15c2d88cc107%2FSegmentedKidney1_1000.png?generation=1707342563232471&amp;alt=media\"></p>\n<p>Slice 0100 from kidney_1_dense Label Vs Prediction. 119 pixels in label, 133 pixels found, 93 pixels true positives, 40 pixels false positives, 26 pixels false negatives. Many of the blob sizes are tiny 1 to 4 pixels area.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2F632bff25ba557b49017796618251f092%2FSegmentedKidney1_0100.png?generation=1707342598764919&amp;alt=media\"></p>\n<h1>Updates After the competition ended</h1>\n<p>After the competition ended, some new opportunities opened up for learning. First, I read the top solution write-ups looking for inspiration, especially low-hanging fruit ideas for adding to my existing implementation. Second, the private and public scores are now visible, enabling us to know how we do on the two test sets. Third, the 5 per day submission limit is lifted, so can do a lot more experiments more quickly. Given these, I was able to improve my private score to 0.682 which would have been 6th place (however not genuinely since I would never have chosen this submission due to the public score being so low). Well, anyway, here it is, my best improved score.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2F574ddf3fa25f91bab4dad90f81efd542%2FSenNetHOA-BestPrivateScore.png?generation=1707777659990418&amp;alt=media\"></p>\n<p>Here are the ideas I added to my implementation to achieve the above score.</p>\n<ol>\n<li>Changed histogram equalization preprocessing to mean/std/clip normalization (I learned from my own prev submissions scoring higher in private score)</li>\n<li>Changed DiceLoss to custom loss that is a combination of boundary loss, dice loss and focal loss (shared by 5th place solution)</li>\n<li>Changed the unet architecture to include an upscale layer (shared by 5th place solution)</li>\n<li>Changed to heavy augmentations in scale and intensity (strategy used by several top solutions)</li>\n<li>Used -0.45, 0.05 for scale limit to downsize more than upsize, to account for the private test set resolution. (shared by 3rd place solution)</li>\n</ol>\n<h1>Sources</h1>\n<p>Following kaggle sources were hugely helpful, many thanks to the contributors.</p>\n<p><a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456118\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456118</a><br>\n<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/464768\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/464768</a><br>\n<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/468525\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/468525</a><br>\n<a href=\"https://www.kaggle.com/code/junkoda/fast-surface-dice-computation\" target=\"_blank\">https://www.kaggle.com/code/junkoda/fast-surface-dice-computation</a><br>\n<a href=\"https://www.kaggle.com/code/hengck23/lb0-808-resnet50-2d-unet-xy-zy-zx-cc3d\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb0-808-resnet50-2d-unet-xy-zy-zx-cc3d</a><br>\n<a href=\"https://www.kaggle.com/code/yoyobar/2-5d-cutting-model-baseline-inference\" target=\"_blank\">https://www.kaggle.com/code/yoyobar/2-5d-cutting-model-baseline-inference</a><br>\n<a href=\"https://www.kaggle.com/code/yoyobar/2-5d-cutting-model-baseline-training\" target=\"_blank\">https://www.kaggle.com/code/yoyobar/2-5d-cutting-model-baseline-training</a><br>\n<a href=\"https://www.kaggle.com/code/misakimatsutomo/inference-1024-should-have-a-percentile-of-0-00149\" target=\"_blank\">https://www.kaggle.com/code/misakimatsutomo/inference-1024-should-have-a-percentile-of-0-00149</a></p>\n<p>Other sources I used for background and inspiration:<br>\n<a href=\"https://doi.org/10.48550/arXiv.2311.13319\" target=\"_blank\">https://doi.org/10.48550/arXiv.2311.13319</a><br>\n<a href=\"https://arxiv.org/abs/1902.00347\" target=\"_blank\">https://arxiv.org/abs/1902.00347</a><br>\n<a href=\"https://arxiv.org/abs/2010.0616\" target=\"_blank\">https://arxiv.org/abs/2010.0616</a></p>",
      "rawMarkdown": "I started late in the beginning of January. This is only my second kaggle competition and I am happy with the enormous learning and the outcome of placing in the top 10% with bronze. Thank you to the fellow teams for sharing your knowledge and to the organizers for a well-run competition.\n\n# Context\nBusiness context: https://www.kaggle.com/competitions/blood-vessel-segmentation/overview\nData context: https://www.kaggle.com/competitions/blood-vessel-segmentation/data\n\n\n# Overview of the approach\nMy final model was an ensemble of two UNet 2D models trained with 1024x1024 and 512x512 sizes with equal weighting in inference. I used slices from X,Y and Z projections for both training and inference, a simplification inspired by the 2.5D Unet paper. \n\nFor training, I used segmentation_models_pytorch smp.Unet architecture with resnext50_32x4d backbone and started with imagenet weights. Trained with kidney_1_dense and kidney_3_dense and validated with kidney_2. For preprocessing, used histogram equalization followed by minmax normalization. Adam optimizer, CosineAnnealingLR scheduler and smp.losses.DiceLoss were used for training. Following augmentations were used in training.\n\n```python\naug_prob     = 0.1 # Augmentation probability\ntrain_aug_list = [\n        A.Rotate(limit=45, p=aug_prob),\n        A.RandomScale(scale_limit=(1.0,1.25),interpolation=cv2.INTER_CUBIC,p=aug_prob),\n        A.PadIfNeeded(min_height=img_size[0], min_width=img_size[1], p=1),\n        A.RandomCrop(img_size[0], img_size[1], p=1),\n        A.RandomBrightnessContrast(p=aug_prob),\n        A.GaussianBlur(p=aug_prob),\n        A.MotionBlur(p=aug_prob),\n        A.GridDistortion(num_steps=5, distort_limit=0.3, p=aug_prob),\n        ToTensorV2(transpose_mask=True),\n    ]\n```\n\nFor inference, I padded the image to 3072x3072 and used a 3x3 grid of 1024x1024 size tiles to run the model on. 4 rotations used on each tile for inference and averaged for Test Time Augmentation aka TTA. Used sigmoid activation layer in Unet. Also, ran inference on X, Y and Z projections and created three prediction volumes. On each projection’s prediction volume, applied sigmoid threshold of 0.0001 to get binary mask volumes, then transposed and added binary masks from the three projections, then used a majority voting to get final predictions.\n\n# Details of the submission\n## What was special about the submission\n-I did no resizing in training or inference to reduce errors and artifacts from downsizing and aspect ratio changes. In training, I used a random crop. In inference, I used padding and tiling.\n-Fixed sigmoid threshold independent of dataset, as opposed to using top Nth percentile for thresholding produced more stable result on private LB. I was bumped up by 616 in ranking.\n-Ensemble of 1024+512 gave higher score than each model applied separately.\n-4x rotation Test Time Augmentation in inference boosted the score by 0.007 in public LB.\n-During local validation and spot checking, my models were achieving very high 2D dice scores on the middle slices in the volumes. Most of the FP and FN errors were on the edge slices in the tiny vessels (1 or 2 pixel errors).\n-While I thought histogram equalization was a secret sauce (since it improved public LB score over mean/std/clip normalization), it turns out my mean/std/clip normalization model produced much better private LB score.\n-There were many high scoring inference notebooks publicly shared in this competition. While I studied them to understand what other teams are doing and did in fact get many great ideas, I chose not to use large sections of code directly since there were many questionable choices and I didn’t understand why those notebooks were producing the high public LB scores. This approach kept my solution unique and generalized enough to move up in the private LB.\n\n## What was tried and didn’t work\n-The striding on 3 or 5 consecutive slices to create a multi-channel image for training and inference did not work well for me, both score wise and CPU/GPU/Memory resource wise.\n-Tried median blur for preprocessing which did not work well.\n-Went from efficientnetb0 to resnext50_32x4d. Maybe somewhere in between would have been better.\n-It would have been better to stick with mean/std/clip norm as opposed to histogram equalization for preprocessing\n-Tried inference with Z projection only initially, after 0.04 improvement in LB score with XYZ projections and majority voting, and decided to use it going forward.\n-There were suggestions in public high-scoring notebooks to reduce augmentation probability to 0.05, I tried low augmentation and although that produced higher validation dice scores during training, almost always produced lower score in public LB. So, I decided to increase to 0.10 probability for augmentation, now that I read the solution writeups from top scoring teams, I realize an even higher augmentation would have been better\n-I had tried training with 90% of slices of all three kidneys dense 1, 2 and 3 and validating with 10% of the slices. Since I read in the discussion board that there is label shift in kidney 2, I switched to using only kidney1 and kidney 3 for training. The results got slightly better. Every incremental improvement counted.\n-Since I realized most of segmentation errors were in the first 100 slices or so with the tiny vessels, I tried to train a separate model with first 200 slices of all three kidneys. Model didn’t work at all on both early slices and middle slices, probably because there was not enough data to train.\n-Tried morphological opening for post processing the predicted mask. I expected that it will reduce some false positives without affecting true positives that much. But nope. It made both false positives and false negatives significantly worse. Then, tried removing all one pixel blobs for post processing. It was a disaster with the score. Also, tried majority voting across three slices to retain positives, it was a disaster. After that, I decided that no post processing is best. Any improvement had to come from better model, not predicted mask cleanup.\n-Tried 6x and 8x TTA with flips and rotations, but score was worse than 4x TTA.\n\n## What I didn’t try which I would consider next time\n1. A loss function that considers 2D boundary loss since the “surface dice score” used for test evaluation uses only the 3D surface boundaries. 5th place solution used a clever loss function of CE_boundaries + Dice + Focal. 4th place solution used a BoundaryDOULoss (https://arxiv.org/pdf/2308.00220.pdf)\n\n2. Take into account that the native scan resolution of private test data is different (63um/voxel) from train (50 and 50.16um) and public test data (50.28um/voxel). This can be accomplished with 2D or 3D resize during inference or equivalent scaling augmentations in training. Top 5 solution write ups talk about this.\n\n3. Somehow incorporate the sparse annotations in training. I completely ignored the kidney_3_sparse and kidney_2 (also sparse and had shifted labels) in training which limited my training data to kidney_1_dense and the small number of slices from kidney_3_dense. Pseudo-annotating the sparse data is an approach I saw that many of the top scoring solutions used.\n\n4. Experiment more with different model architectures and ensembles? I was bogged down with the basics and getting the foundations working. Also, with the Kaggle 30 hours per week GPU limit, didn’t have time and resources for this.\n\n5. Increase the augmentation probabilities in intensity and scaling. Many of the top solutions have done this.\n\n6. 2.5D approach with 5 consecutive slices concatenated as a 5-channel image. 6th place solution has successfully used this. I was running into CPU/GPU/Mem resource issues, but those could have been potentially resolved with more time and effort. \n\n## Preprocessing examples\n\n\nHistogram equalization followed by Minmax normalization:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2Fe119a342eda6f290c058e94c97728e78%2FNorm-HistogramEqualization.png?generation=1707343051563429&alt=media)\n\nMean/Std normalization with clipping followed by Minmax normalization:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2F9e72cb4f6f49eee6aae6213df2833e90%2FNorm-MeanStdClip.png?generation=1707343078883718&alt=media)\n\n## Padding and tiling kidney slice for inference\nThis is showing 2048x2048 padding and 512x512 tiles which I later changed to 3072x3072 padding and 1024x1024 tiles.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2F4acb3f75490749342dab0c8e3f7f91de%2FTiledKidneySlice.png?generation=1707342738234654&alt=media)\n\n\n## Sample segmentation results on slices for illustration\n\nSlice 1000 from kidney_1_dense Label Vs Prediction. Prediction is color coded as Green for True Positives, Red for false positives and Blue for false negatives. Also all blobs are dilated 5x5 to observe the tiny blobs visually. Most of the false positives and false negatives are 1 pixel area blobs. However, there are also some 1 pixel area true positive blobs in this slice and many more of those in early slices for example 0100. So, cannot blindly remove 1 pixel blobs.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2F3d16023f25a7f35a225c15c2d88cc107%2FSegmentedKidney1_1000.png?generation=1707342563232471&alt=media)\n\nSlice 0100 from kidney_1_dense Label Vs Prediction. 119 pixels in label, 133 pixels found, 93 pixels true positives, 40 pixels false positives, 26 pixels false negatives. Many of the blob sizes are tiny 1 to 4 pixels area.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2F632bff25ba557b49017796618251f092%2FSegmentedKidney1_0100.png?generation=1707342598764919&alt=media)\n\n# Updates After the competition ended\nAfter the competition ended, some new opportunities opened up for learning. First, I read the top solution write-ups looking for inspiration, especially low-hanging fruit ideas for adding to my existing implementation. Second, the private and public scores are now visible, enabling us to know how we do on the two test sets. Third, the 5 per day submission limit is lifted, so can do a lot more experiments more quickly. Given these, I was able to improve my private score to 0.682 which would have been 6th place (however not genuinely since I would never have chosen this submission due to the public score being so low). Well, anyway, here it is, my best improved score.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2F574ddf3fa25f91bab4dad90f81efd542%2FSenNetHOA-BestPrivateScore.png?generation=1707777659990418&alt=media)\n\nHere are the ideas I added to my implementation to achieve the above score.\n1. Changed histogram equalization preprocessing to mean/std/clip normalization (I learned from my own prev submissions scoring higher in private score)\n2. Changed DiceLoss to custom loss that is a combination of boundary loss, dice loss and focal loss (shared by 5th place solution)\n3. Changed the unet architecture to include an upscale layer (shared by 5th place solution)\n4. Changed to heavy augmentations in scale and intensity (strategy used by several top solutions)\n5. Used -0.45, 0.05 for scale limit to downsize more than upsize, to account for the private test set resolution. (shared by 3rd place solution)\n\n# Sources\nFollowing kaggle sources were hugely helpful, many thanks to the contributors.\n\nhttps://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456118\nhttps://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/464768\nhttps://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/468525\nhttps://www.kaggle.com/code/junkoda/fast-surface-dice-computation\nhttps://www.kaggle.com/code/hengck23/lb0-808-resnet50-2d-unet-xy-zy-zx-cc3d\nhttps://www.kaggle.com/code/yoyobar/2-5d-cutting-model-baseline-inference\nhttps://www.kaggle.com/code/yoyobar/2-5d-cutting-model-baseline-training\nhttps://www.kaggle.com/code/misakimatsutomo/inference-1024-should-have-a-percentile-of-0-00149\n\nOther sources I used for background and inspiration:\nhttps://doi.org/10.48550/arXiv.2311.13319\nhttps://arxiv.org/abs/1902.00347\nhttps://arxiv.org/abs/2010.0616",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2642074": "I started late in the beginning of January. This is only my second kaggle competition and I am happy with the enormous learning and the outcome of placing in the top 10% with bronze. Thank you to the fellow teams for sharing your knowledge and to the organizers for a well-run competition.\n\n# Context\nBusiness context: https://www.kaggle.com/competitions/blood-vessel-segmentation/overview\nData context: https://www.kaggle.com/competitions/blood-vessel-segmentation/data\n\n\n# Overview of the approach\nMy final model was an ensemble of two UNet 2D models trained with 1024x1024 and 512x512 sizes with equal weighting in inference. I used slices from X,Y and Z projections for both training and inference, a simplification inspired by the 2.5D Unet paper. \n\nFor training, I used segmentation_models_pytorch smp.Unet architecture with resnext50_32x4d backbone and started with imagenet weights. Trained with kidney_1_dense and kidney_3_dense and validated with kidney_2. For preprocessing, used histogram equalization followed by minmax normalization. Adam optimizer, CosineAnnealingLR scheduler and smp.losses.DiceLoss were used for training. Following augmentations were used in training.\n\n```python\naug_prob     = 0.1 # Augmentation probability\ntrain_aug_list = [\n        A.Rotate(limit=45, p=aug_prob),\n        A.RandomScale(scale_limit=(1.0,1.25),interpolation=cv2.INTER_CUBIC,p=aug_prob),\n        A.PadIfNeeded(min_height=img_size[0], min_width=img_size[1], p=1),\n        A.RandomCrop(img_size[0], img_size[1], p=1),\n        A.RandomBrightnessContrast(p=aug_prob),\n        A.GaussianBlur(p=aug_prob),\n        A.MotionBlur(p=aug_prob),\n        A.GridDistortion(num_steps=5, distort_limit=0.3, p=aug_prob),\n        ToTensorV2(transpose_mask=True),\n    ]\n```\n\nFor inference, I padded the image to 3072x3072 and used a 3x3 grid of 1024x1024 size tiles to run the model on. 4 rotations used on each tile for inference and averaged for Test Time Augmentation aka TTA. Used sigmoid activation layer in Unet. Also, ran inference on X, Y and Z projections and created three prediction volumes. On each projection’s prediction volume, applied sigmoid threshold of 0.0001 to get binary mask volumes, then transposed and added binary masks from the three projections, then used a majority voting to get final predictions.\n\n# Details of the submission\n## What was special about the submission\n-I did no resizing in training or inference to reduce errors and artifacts from downsizing and aspect ratio changes. In training, I used a random crop. In inference, I used padding and tiling.\n-Fixed sigmoid threshold independent of dataset, as opposed to using top Nth percentile for thresholding produced more stable result on private LB. I was bumped up by 616 in ranking.\n-Ensemble of 1024+512 gave higher score than each model applied separately.\n-4x rotation Test Time Augmentation in inference boosted the score by 0.007 in public LB.\n-During local validation and spot checking, my models were achieving very high 2D dice scores on the middle slices in the volumes. Most of the FP and FN errors were on the edge slices in the tiny vessels (1 or 2 pixel errors).\n-While I thought histogram equalization was a secret sauce (since it improved public LB score over mean/std/clip normalization), it turns out my mean/std/clip normalization model produced much better private LB score.\n-There were many high scoring inference notebooks publicly shared in this competition. While I studied them to understand what other teams are doing and did in fact get many great ideas, I chose not to use large sections of code directly since there were many questionable choices and I didn’t understand why those notebooks were producing the high public LB scores. This approach kept my solution unique and generalized enough to move up in the private LB.\n\n## What was tried and didn’t work\n-The striding on 3 or 5 consecutive slices to create a multi-channel image for training and inference did not work well for me, both score wise and CPU/GPU/Memory resource wise.\n-Tried median blur for preprocessing which did not work well.\n-Went from efficientnetb0 to resnext50_32x4d. Maybe somewhere in between would have been better.\n-It would have been better to stick with mean/std/clip norm as opposed to histogram equalization for preprocessing\n-Tried inference with Z projection only initially, after 0.04 improvement in LB score with XYZ projections and majority voting, and decided to use it going forward.\n-There were suggestions in public high-scoring notebooks to reduce augmentation probability to 0.05, I tried low augmentation and although that produced higher validation dice scores during training, almost always produced lower score in public LB. So, I decided to increase to 0.10 probability for augmentation, now that I read the solution writeups from top scoring teams, I realize an even higher augmentation would have been better\n-I had tried training with 90% of slices of all three kidneys dense 1, 2 and 3 and validating with 10% of the slices. Since I read in the discussion board that there is label shift in kidney 2, I switched to using only kidney1 and kidney 3 for training. The results got slightly better. Every incremental improvement counted.\n-Since I realized most of segmentation errors were in the first 100 slices or so with the tiny vessels, I tried to train a separate model with first 200 slices of all three kidneys. Model didn’t work at all on both early slices and middle slices, probably because there was not enough data to train.\n-Tried morphological opening for post processing the predicted mask. I expected that it will reduce some false positives without affecting true positives that much. But nope. It made both false positives and false negatives significantly worse. Then, tried removing all one pixel blobs for post processing. It was a disaster with the score. Also, tried majority voting across three slices to retain positives, it was a disaster. After that, I decided that no post processing is best. Any improvement had to come from better model, not predicted mask cleanup.\n-Tried 6x and 8x TTA with flips and rotations, but score was worse than 4x TTA.\n\n## What I didn’t try which I would consider next time\n1. A loss function that considers 2D boundary loss since the “surface dice score” used for test evaluation uses only the 3D surface boundaries. 5th place solution used a clever loss function of CE_boundaries + Dice + Focal. 4th place solution used a BoundaryDOULoss (https://arxiv.org/pdf/2308.00220.pdf)\n\n2. Take into account that the native scan resolution of private test data is different (63um/voxel) from train (50 and 50.16um) and public test data (50.28um/voxel). This can be accomplished with 2D or 3D resize during inference or equivalent scaling augmentations in training. Top 5 solution write ups talk about this.\n\n3. Somehow incorporate the sparse annotations in training. I completely ignored the kidney_3_sparse and kidney_2 (also sparse and had shifted labels) in training which limited my training data to kidney_1_dense and the small number of slices from kidney_3_dense. Pseudo-annotating the sparse data is an approach I saw that many of the top scoring solutions used.\n\n4. Experiment more with different model architectures and ensembles? I was bogged down with the basics and getting the foundations working. Also, with the Kaggle 30 hours per week GPU limit, didn’t have time and resources for this.\n\n5. Increase the augmentation probabilities in intensity and scaling. Many of the top solutions have done this.\n\n6. 2.5D approach with 5 consecutive slices concatenated as a 5-channel image. 6th place solution has successfully used this. I was running into CPU/GPU/Mem resource issues, but those could have been potentially resolved with more time and effort. \n\n## Preprocessing examples\n\n\nHistogram equalization followed by Minmax normalization:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2Fe119a342eda6f290c058e94c97728e78%2FNorm-HistogramEqualization.png?generation=1707343051563429&alt=media)\n\nMean/Std normalization with clipping followed by Minmax normalization:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2F9e72cb4f6f49eee6aae6213df2833e90%2FNorm-MeanStdClip.png?generation=1707343078883718&alt=media)\n\n## Padding and tiling kidney slice for inference\nThis is showing 2048x2048 padding and 512x512 tiles which I later changed to 3072x3072 padding and 1024x1024 tiles.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2F4acb3f75490749342dab0c8e3f7f91de%2FTiledKidneySlice.png?generation=1707342738234654&alt=media)\n\n\n## Sample segmentation results on slices for illustration\n\nSlice 1000 from kidney_1_dense Label Vs Prediction. Prediction is color coded as Green for True Positives, Red for false positives and Blue for false negatives. Also all blobs are dilated 5x5 to observe the tiny blobs visually. Most of the false positives and false negatives are 1 pixel area blobs. However, there are also some 1 pixel area true positive blobs in this slice and many more of those in early slices for example 0100. So, cannot blindly remove 1 pixel blobs.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2F3d16023f25a7f35a225c15c2d88cc107%2FSegmentedKidney1_1000.png?generation=1707342563232471&alt=media)\n\nSlice 0100 from kidney_1_dense Label Vs Prediction. 119 pixels in label, 133 pixels found, 93 pixels true positives, 40 pixels false positives, 26 pixels false negatives. Many of the blob sizes are tiny 1 to 4 pixels area.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2F632bff25ba557b49017796618251f092%2FSegmentedKidney1_0100.png?generation=1707342598764919&alt=media)\n\n# Updates After the competition ended\nAfter the competition ended, some new opportunities opened up for learning. First, I read the top solution write-ups looking for inspiration, especially low-hanging fruit ideas for adding to my existing implementation. Second, the private and public scores are now visible, enabling us to know how we do on the two test sets. Third, the 5 per day submission limit is lifted, so can do a lot more experiments more quickly. Given these, I was able to improve my private score to 0.682 which would have been 6th place (however not genuinely since I would never have chosen this submission due to the public score being so low). Well, anyway, here it is, my best improved score.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18199244%2F574ddf3fa25f91bab4dad90f81efd542%2FSenNetHOA-BestPrivateScore.png?generation=1707777659990418&alt=media)\n\nHere are the ideas I added to my implementation to achieve the above score.\n1. Changed histogram equalization preprocessing to mean/std/clip normalization (I learned from my own prev submissions scoring higher in private score)\n2. Changed DiceLoss to custom loss that is a combination of boundary loss, dice loss and focal loss (shared by 5th place solution)\n3. Changed the unet architecture to include an upscale layer (shared by 5th place solution)\n4. Changed to heavy augmentations in scale and intensity (strategy used by several top solutions)\n5. Used -0.45, 0.05 for scale limit to downsize more than upsize, to account for the private test set resolution. (shared by 3rd place solution)\n\n# Sources\nFollowing kaggle sources were hugely helpful, many thanks to the contributors.\n\nhttps://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456118\nhttps://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/464768\nhttps://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/468525\nhttps://www.kaggle.com/code/junkoda/fast-surface-dice-computation\nhttps://www.kaggle.com/code/hengck23/lb0-808-resnet50-2d-unet-xy-zy-zx-cc3d\nhttps://www.kaggle.com/code/yoyobar/2-5d-cutting-model-baseline-inference\nhttps://www.kaggle.com/code/yoyobar/2-5d-cutting-model-baseline-training\nhttps://www.kaggle.com/code/misakimatsutomo/inference-1024-should-have-a-percentile-of-0-00149\n\nOther sources I used for background and inspiration:\nhttps://doi.org/10.48550/arXiv.2311.13319\nhttps://arxiv.org/abs/1902.00347\nhttps://arxiv.org/abs/2010.0616"
  }
}