{
  "id": 475252,
  "title": "6th place solution: Luck is All You Need",
  "url": "/competitions/blood-vessel-segmentation/discussion/475252",
  "author_name": "Volodymyr",
  "post_date": "2024-02-07T16:30:54.103000",
  "votes": 28,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi, Kagglers!</p>\n<p>After finding yourself on the Leaderboard after Grand Shakeup and restoring your mental health, we can dive deep into 6th place solution, but before this, a few very important words:<br>\n<em>I would like to thank the Armed Forces of Ukraine, Security Service of Ukraine, Defence Intelligence of Ukraine, and the State Emergency Service of Ukraine for providing safety and security to participate in this great competition, complete this work, and help science, technology, and business not to stop but to move forward.</em></p>\n<h1>Validation. Not really</h1>\n<p>I was using for validation 2 organs (in different folds): kidney_3_dense and kidney_2. After releasing a <a href=\"https://www.kaggle.com/code/junkoda/fast-surface-dice-computation\" target=\"_blank\">fast version of 3D Surface Dice</a>, I was able to compute validation scores while training, and I received the next insights:</p>\n<ul>\n<li>Tracking the score on kidney_2 was useless for me. The validation score decreased from epochs 2-3 on kidney_2</li>\n<li>Scores on kidney_3_dense were meaningful for checking “radical” features, like additional data and new losses. But then optimal score fluctuated between 0.9-0.925 dice without any reasonable correlation with Public or Private score</li>\n<li>The optimal threshold on kidney_3_dense was optimal for Private, Public, and kidney_3_dense scores - 0.1 and lower <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2Fa4050c6e4e6d27662cfddc62201e2f71%2Fkidney_3_dense_val.png?generation=1707321782497972&amp;alt=media\"></li>\n<li>Resize to constant um/voxel (I have picked 50 um/voxel) for prediction increased optimal threshold both on CV and Public but decreased optimal score dramatically. But on Private, it became one of the most robust approaches<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2F6f0a8a87cf758aa4920353ace23c11cf%2Fkidney_3_dense_resize_val.png?generation=1707323076488076&amp;alt=media\"><br>\nIn summary, Validation did not work (at least mine). It is not strange because of solo data point in CV, Public, and Private </li>\n</ul>\n<h1>Data</h1>\n<p>I was using all train data except kidney_1_voi sample<br>\nIn order to enlarge training data I have used 50um_LADAF_2020_31_kidney_pag from <a href=\"https://human-organ-atlas.esrf.eu/search?organ=kidney\" target=\"_blank\">Human Organ Atlas</a> <br>\nFor data normalization, I was using the approach proposed by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> - <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456118#2552053\" target=\"_blank\">percentile normalization</a></p>\n<h1>Training setup</h1>\n<p>I mostly stick to the 2.5D approach with 5 slices.<br>\nI started from one view model and iterated along the last axis, but then I switched to a multiview and used slices by all three axes in the training<br>\nI have used 512 square crops with Non Empty probability of 0.5, pretty much standard augmentations and CutMix within one organ and one view with 0.5 probability and 1.0 alpha :</p>\n<pre><code>: : [\n    A.PadIfNeeded(\n        min_height=crop_size,\n        min_width=crop_size,\n        always_apply=,\n    ),\n    \n    \n    A.OneOrOther(\n        first=A.CropNonEmptyMaskIfExists(crop_size, crop_size), \n        second=A.RandomCrop(crop_size, crop_size), \n        p=\n    ),\n],\n: A.Compose(\n    [\n        A.ShiftScaleRotate(\n            scale_limit=,\n        ),\n        \n        A.RandomRotate90(p=),\n        A.HorizontalFlip(p=),\n        A.VerticalFlip(p=),\n        A.Transpose(p=),\n        A.OneOf(\n            [\n                A.RandomBrightnessContrast(),\n                A.RandomBrightness(),\n                A.RandomGamma(),\n            ],\n            p=,\n        ),\n        ToTensorV2(transpose_mask=),\n    ]\n),\n\n: ,\n: {: , : },\n</code></pre>\n<p>I was using Adam optimizer and reduced learning rate with CosineAnnealingLR starting from 1e-3 and ending with 1e-6<br>\nRegarding loss function choice, I started with classical BCE+Dice loss and then tried to implement the loss function, which will directly optimize the metric, but unfortunately, it did not work well. Luckily, I have come across BoundaryLoss, which worked firstly comparable to BCE+Dice loss and then better. Interesting fact that the best (not selected) model was trained on BoundaryLoss + 0.5 Focal Symmetric Loss and scored 0.756 on Private, 0.867 on Public, and 0.916 on kidney_3_dense, which is a pretty much balanced score (of course, comparing to all other models score distribution 🙂)<br>\nI was training for 30 epochs. repeating the original train set 3 times and the pseudo train set 2 times<br>\nI have used a batch size of 14 samples and trained with DDP on 2 GPUs, so the final batch size was 28</p>\n<p>After I saw <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461213#2594751\" target=\"_blank\">the post</a> about promising results from 3D models, I started exploring 3D approaches, and they worked pretty well. In order to make it train without NaNs I have changed the optimization strategy and switched to SGD, with momentum=0.99, weight_decay=3e-5, nesterov=True and also changed starting learning rate to 1e-6 - taken from <a href=\"https://github.com/Project-MONAI/tutorials/tree/main/modules/dynunet_pipeline\" target=\"_blank\">monai example</a><br>\nAs the overall image resolution of the image was increased dramatically, I had to reduce the batch to 3 on one GPU, so the aggregated batch size was 6. I was training in total for ~120K iterations<br>\nAs for augmentations - they were pretty much the same as in the 2.5D setup, except from Zoom. </p>\n<pre><code>:  : mt.Compose(\n    [\n        mt.OneOf([\n            mt.RandRotate90d(keys=(, ), prob=, spatial_axes=(-,-)),\n            mt.RandRotate90d(keys=(, ), prob=, spatial_axes=(-,-)),\n            mt.RandRotate90d(keys=(, ), prob=, spatial_axes=(-,-))\n        ]),\n        mt.RandFlipd(keys=(, ), prob=, spatial_axis=-),\n        mt.RandFlipd(keys=(, ), prob=, spatial_axis=-),\n        mt.RandFlipd(keys=(, ), prob=, spatial_axis=-),\n\n        mt.RandScaleIntensityd(keys=(), prob=, factors=)\n    ]\n),\n</code></pre>\n<p>Zoom worked better on CV but worse on Public and also on Private (Why? - who knows …)</p>\n<h1>Neural Networks</h1>\n<p>I was mostly using <a href=\"https://arxiv.org/abs/1905.11946\" target=\"_blank\">EfficientNet family</a> as an Encoder (from noisy student weights), started from B3, then switched to B5, and unfortunately, B7 did not work well for me both on CV and Public </p>\n<p>Interestingly se_resnext50_32x4d performed not well on Public LB (0.852) and CV (0.909) but really well on Private (0.702)</p>\n<p>As for Decoder I was mostly using Unet++. I have tried <a href=\"https://arxiv.org/abs/2004.08790\" target=\"_blank\">Unet3+</a> but it showed considerably worse results</p>\n<p>As for 3D Nets, I was using <a href=\"https://monai-dev.readthedocs.io/en/fixes-sphinx/networks.html#dynunet\" target=\"_blank\">DynUNet</a> and adopted model architecture according to <a href=\"https://github.com/Project-MONAI/tutorials/blob/main/modules/dynunet_pipeline/create_network.py#L19\" target=\"_blank\">next script</a>. I have tried to use pretrained Unet from  <a href=\"https://monai.io/model-zoo.html\" target=\"_blank\">MONAI Model Zoo</a> but it performed badly on all sets </p>\n<h1>Inference and Post Processing</h1>\n<p>I was using 512 sliding window with 0.5 overlap, flip TTA, and last checkpoint from 2 folds.<br>\nAfter switching to multi view model, I have also added multi view TTA</p>\n<p>The next step was the creation of a kidney mask. I have tried several approaches </p>\n<ol>\n<li>Using segmentation net, trained on <a href=\"https://www.kaggle.com/datasets/squidinator/sennet-hoa-kidney-13-dense-full-kidney-masks\" target=\"_blank\">this dataset</a> + slight post-processing for removing binary holes and small connected regions </li>\n<li>Using an algorithmic approach based on intensity thresholding, erosion, and dilation</li>\n</ol>\n<p>The first one had a pretty high FP rate but nearly zero FN rate, while the second one had a pretty high FN rate. Both of them performed nearly ideal on kidney_3, so did not really influence the fold 0 scores, but an algorithmic approach cut out kidney regions for kidney_2 and dramatically reduced the fold 1 score. BUT at the same time, the second approach improved Public score (0.874-&gt;0.882). I understood that it was 90% overfit to Public LB, but I have decided to take the risk</p>\n<h1>Final Model</h1>\n<p>For final submission, I have selected the following ones:</p>\n<ul>\n<li>Pure 2.5D -&gt; algorithmic post-processing<ul>\n<li>Public: 0.886</li>\n<li>Private: 0.681</li>\n<li>Kidney 3 dense score: 0.917   </li></ul></li>\n<li>2.5D (weight 3.0) blended with 3D (weight 1.0) <ul>\n<li>Public: 0.871</li>\n<li>Private: 0.676</li>\n<li>Kidney 3 dense score: ~0.918<br>\nFor both models, I used 0.05 threshold </li></ul></li>\n</ul>\n<h1>The most popular rubric of this competition: Not Selected Best Submission</h1>\n<p>Here, I want to point out several of the most exciting approaches for me </p>\n<ul>\n<li>Pure 2.5D but add Symmetric Focal loss with 0.5 coefficient <ul>\n<li>Public: 0.867</li>\n<li>Private: 0.756</li>\n<li>Kidney 3 dense score: 0.916</li></ul></li>\n<li>Resize 2d slices to 50 um/voxel for prediction and then resize back <ul>\n<li>Public: 0.799</li>\n<li>Private: 0.753</li>\n<li>Kidney 3 dense score: 0.907</li></ul></li>\n<li>Resize the whole volume with scipy.zoom to 50 um/voxel for prediction and than resize back <ul>\n<li>Public: 0.7 resize back </li>\n<li>Public: 0.726</li>\n<li>Private: 0.745</li>\n<li>Kidney 3 dense score: Have not checked </li></ul></li>\n<li>Solo 3D model <ul>\n<li>Public: 0.849</li>\n<li>Private: 0.723</li>\n<li>Kidney 3 dense score: 0.915<br>\nFor me, it was logical to pick first or second, but as for all other better submissions, it sounds to me like pure random.</li></ul></li>\n</ul>\n<h1>Conclusions</h1>\n<p>Computing metrics on one data sample leads to severe shakeups 🙂</p>\n<h1>Closing words</h1>\n<p>I hope you have not fallen asleep while reading. Finally, I want to thank the entire Kaggle community, congratulate all participants and winners. Special thanks to Indian University Bloomington, University College London, Yashvardhan Jain (@yashvrdnjain), Claire Walsh (@clairewalsh), the Kaggle Team, and other organizers.</p>",
  "messages": [
    {
      "id": 2641732,
      "postDate": "2024-02-07T16:30:54.103Z",
      "content": "<p>Hi, Kagglers!</p>\n<p>After finding yourself on the Leaderboard after Grand Shakeup and restoring your mental health, we can dive deep into 6th place solution, but before this, a few very important words:<br>\n<em>I would like to thank the Armed Forces of Ukraine, Security Service of Ukraine, Defence Intelligence of Ukraine, and the State Emergency Service of Ukraine for providing safety and security to participate in this great competition, complete this work, and help science, technology, and business not to stop but to move forward.</em></p>\n<h1>Validation. Not really</h1>\n<p>I was using for validation 2 organs (in different folds): kidney_3_dense and kidney_2. After releasing a <a href=\"https://www.kaggle.com/code/junkoda/fast-surface-dice-computation\" target=\"_blank\">fast version of 3D Surface Dice</a>, I was able to compute validation scores while training, and I received the next insights:</p>\n<ul>\n<li>Tracking the score on kidney_2 was useless for me. The validation score decreased from epochs 2-3 on kidney_2</li>\n<li>Scores on kidney_3_dense were meaningful for checking “radical” features, like additional data and new losses. But then optimal score fluctuated between 0.9-0.925 dice without any reasonable correlation with Public or Private score</li>\n<li>The optimal threshold on kidney_3_dense was optimal for Private, Public, and kidney_3_dense scores - 0.1 and lower <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2Fa4050c6e4e6d27662cfddc62201e2f71%2Fkidney_3_dense_val.png?generation=1707321782497972&amp;alt=media\"></li>\n<li>Resize to constant um/voxel (I have picked 50 um/voxel) for prediction increased optimal threshold both on CV and Public but decreased optimal score dramatically. But on Private, it became one of the most robust approaches<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2F6f0a8a87cf758aa4920353ace23c11cf%2Fkidney_3_dense_resize_val.png?generation=1707323076488076&amp;alt=media\"><br>\nIn summary, Validation did not work (at least mine). It is not strange because of solo data point in CV, Public, and Private </li>\n</ul>\n<h1>Data</h1>\n<p>I was using all train data except kidney_1_voi sample<br>\nIn order to enlarge training data I have used 50um_LADAF_2020_31_kidney_pag from <a href=\"https://human-organ-atlas.esrf.eu/search?organ=kidney\" target=\"_blank\">Human Organ Atlas</a> <br>\nFor data normalization, I was using the approach proposed by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> - <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456118#2552053\" target=\"_blank\">percentile normalization</a></p>\n<h1>Training setup</h1>\n<p>I mostly stick to the 2.5D approach with 5 slices.<br>\nI started from one view model and iterated along the last axis, but then I switched to a multiview and used slices by all three axes in the training<br>\nI have used 512 square crops with Non Empty probability of 0.5, pretty much standard augmentations and CutMix within one organ and one view with 0.5 probability and 1.0 alpha :</p>\n<pre><code>: : [\n    A.PadIfNeeded(\n        min_height=crop_size,\n        min_width=crop_size,\n        always_apply=,\n    ),\n    \n    \n    A.OneOrOther(\n        first=A.CropNonEmptyMaskIfExists(crop_size, crop_size), \n        second=A.RandomCrop(crop_size, crop_size), \n        p=\n    ),\n],\n: A.Compose(\n    [\n        A.ShiftScaleRotate(\n            scale_limit=,\n        ),\n        \n        A.RandomRotate90(p=),\n        A.HorizontalFlip(p=),\n        A.VerticalFlip(p=),\n        A.Transpose(p=),\n        A.OneOf(\n            [\n                A.RandomBrightnessContrast(),\n                A.RandomBrightness(),\n                A.RandomGamma(),\n            ],\n            p=,\n        ),\n        ToTensorV2(transpose_mask=),\n    ]\n),\n\n: ,\n: {: , : },\n</code></pre>\n<p>I was using Adam optimizer and reduced learning rate with CosineAnnealingLR starting from 1e-3 and ending with 1e-6<br>\nRegarding loss function choice, I started with classical BCE+Dice loss and then tried to implement the loss function, which will directly optimize the metric, but unfortunately, it did not work well. Luckily, I have come across BoundaryLoss, which worked firstly comparable to BCE+Dice loss and then better. Interesting fact that the best (not selected) model was trained on BoundaryLoss + 0.5 Focal Symmetric Loss and scored 0.756 on Private, 0.867 on Public, and 0.916 on kidney_3_dense, which is a pretty much balanced score (of course, comparing to all other models score distribution 🙂)<br>\nI was training for 30 epochs. repeating the original train set 3 times and the pseudo train set 2 times<br>\nI have used a batch size of 14 samples and trained with DDP on 2 GPUs, so the final batch size was 28</p>\n<p>After I saw <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461213#2594751\" target=\"_blank\">the post</a> about promising results from 3D models, I started exploring 3D approaches, and they worked pretty well. In order to make it train without NaNs I have changed the optimization strategy and switched to SGD, with momentum=0.99, weight_decay=3e-5, nesterov=True and also changed starting learning rate to 1e-6 - taken from <a href=\"https://github.com/Project-MONAI/tutorials/tree/main/modules/dynunet_pipeline\" target=\"_blank\">monai example</a><br>\nAs the overall image resolution of the image was increased dramatically, I had to reduce the batch to 3 on one GPU, so the aggregated batch size was 6. I was training in total for ~120K iterations<br>\nAs for augmentations - they were pretty much the same as in the 2.5D setup, except from Zoom. </p>\n<pre><code>:  : mt.Compose(\n    [\n        mt.OneOf([\n            mt.RandRotate90d(keys=(, ), prob=, spatial_axes=(-,-)),\n            mt.RandRotate90d(keys=(, ), prob=, spatial_axes=(-,-)),\n            mt.RandRotate90d(keys=(, ), prob=, spatial_axes=(-,-))\n        ]),\n        mt.RandFlipd(keys=(, ), prob=, spatial_axis=-),\n        mt.RandFlipd(keys=(, ), prob=, spatial_axis=-),\n        mt.RandFlipd(keys=(, ), prob=, spatial_axis=-),\n\n        mt.RandScaleIntensityd(keys=(), prob=, factors=)\n    ]\n),\n</code></pre>\n<p>Zoom worked better on CV but worse on Public and also on Private (Why? - who knows …)</p>\n<h1>Neural Networks</h1>\n<p>I was mostly using <a href=\"https://arxiv.org/abs/1905.11946\" target=\"_blank\">EfficientNet family</a> as an Encoder (from noisy student weights), started from B3, then switched to B5, and unfortunately, B7 did not work well for me both on CV and Public </p>\n<p>Interestingly se_resnext50_32x4d performed not well on Public LB (0.852) and CV (0.909) but really well on Private (0.702)</p>\n<p>As for Decoder I was mostly using Unet++. I have tried <a href=\"https://arxiv.org/abs/2004.08790\" target=\"_blank\">Unet3+</a> but it showed considerably worse results</p>\n<p>As for 3D Nets, I was using <a href=\"https://monai-dev.readthedocs.io/en/fixes-sphinx/networks.html#dynunet\" target=\"_blank\">DynUNet</a> and adopted model architecture according to <a href=\"https://github.com/Project-MONAI/tutorials/blob/main/modules/dynunet_pipeline/create_network.py#L19\" target=\"_blank\">next script</a>. I have tried to use pretrained Unet from  <a href=\"https://monai.io/model-zoo.html\" target=\"_blank\">MONAI Model Zoo</a> but it performed badly on all sets </p>\n<h1>Inference and Post Processing</h1>\n<p>I was using 512 sliding window with 0.5 overlap, flip TTA, and last checkpoint from 2 folds.<br>\nAfter switching to multi view model, I have also added multi view TTA</p>\n<p>The next step was the creation of a kidney mask. I have tried several approaches </p>\n<ol>\n<li>Using segmentation net, trained on <a href=\"https://www.kaggle.com/datasets/squidinator/sennet-hoa-kidney-13-dense-full-kidney-masks\" target=\"_blank\">this dataset</a> + slight post-processing for removing binary holes and small connected regions </li>\n<li>Using an algorithmic approach based on intensity thresholding, erosion, and dilation</li>\n</ol>\n<p>The first one had a pretty high FP rate but nearly zero FN rate, while the second one had a pretty high FN rate. Both of them performed nearly ideal on kidney_3, so did not really influence the fold 0 scores, but an algorithmic approach cut out kidney regions for kidney_2 and dramatically reduced the fold 1 score. BUT at the same time, the second approach improved Public score (0.874-&gt;0.882). I understood that it was 90% overfit to Public LB, but I have decided to take the risk</p>\n<h1>Final Model</h1>\n<p>For final submission, I have selected the following ones:</p>\n<ul>\n<li>Pure 2.5D -&gt; algorithmic post-processing<ul>\n<li>Public: 0.886</li>\n<li>Private: 0.681</li>\n<li>Kidney 3 dense score: 0.917   </li></ul></li>\n<li>2.5D (weight 3.0) blended with 3D (weight 1.0) <ul>\n<li>Public: 0.871</li>\n<li>Private: 0.676</li>\n<li>Kidney 3 dense score: ~0.918<br>\nFor both models, I used 0.05 threshold </li></ul></li>\n</ul>\n<h1>The most popular rubric of this competition: Not Selected Best Submission</h1>\n<p>Here, I want to point out several of the most exciting approaches for me </p>\n<ul>\n<li>Pure 2.5D but add Symmetric Focal loss with 0.5 coefficient <ul>\n<li>Public: 0.867</li>\n<li>Private: 0.756</li>\n<li>Kidney 3 dense score: 0.916</li></ul></li>\n<li>Resize 2d slices to 50 um/voxel for prediction and then resize back <ul>\n<li>Public: 0.799</li>\n<li>Private: 0.753</li>\n<li>Kidney 3 dense score: 0.907</li></ul></li>\n<li>Resize the whole volume with scipy.zoom to 50 um/voxel for prediction and than resize back <ul>\n<li>Public: 0.7 resize back </li>\n<li>Public: 0.726</li>\n<li>Private: 0.745</li>\n<li>Kidney 3 dense score: Have not checked </li></ul></li>\n<li>Solo 3D model <ul>\n<li>Public: 0.849</li>\n<li>Private: 0.723</li>\n<li>Kidney 3 dense score: 0.915<br>\nFor me, it was logical to pick first or second, but as for all other better submissions, it sounds to me like pure random.</li></ul></li>\n</ul>\n<h1>Conclusions</h1>\n<p>Computing metrics on one data sample leads to severe shakeups 🙂</p>\n<h1>Closing words</h1>\n<p>I hope you have not fallen asleep while reading. Finally, I want to thank the entire Kaggle community, congratulate all participants and winners. Special thanks to Indian University Bloomington, University College London, Yashvardhan Jain (@yashvrdnjain), Claire Walsh (@clairewalsh), the Kaggle Team, and other organizers.</p>",
      "rawMarkdown": "Hi, Kagglers!\n\nAfter finding yourself on the Leaderboard after Grand Shakeup and restoring your mental health, we can dive deep into 6th place solution, but before this, a few very important words:\n*I would like to thank the Armed Forces of Ukraine, Security Service of Ukraine, Defence Intelligence of Ukraine, and the State Emergency Service of Ukraine for providing safety and security to participate in this great competition, complete this work, and help science, technology, and business not to stop but to move forward.*\n\n# Validation. Not really \n\nI was using for validation 2 organs (in different folds): kidney_3_dense and kidney_2. After releasing a [fast version of 3D Surface Dice](https://www.kaggle.com/code/junkoda/fast-surface-dice-computation), I was able to compute validation scores while training, and I received the next insights:\n- Tracking the score on kidney_2 was useless for me. The validation score decreased from epochs 2-3 on kidney_2\n- Scores on kidney_3_dense were meaningful for checking “radical” features, like additional data and new losses. But then optimal score fluctuated between 0.9-0.925 dice without any reasonable correlation with Public or Private score\n- The optimal threshold on kidney_3_dense was optimal for Private, Public, and kidney_3_dense scores - 0.1 and lower \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2Fa4050c6e4e6d27662cfddc62201e2f71%2Fkidney_3_dense_val.png?generation=1707321782497972&alt=media)\n- Resize to constant um/voxel (I have picked 50 um/voxel) for prediction increased optimal threshold both on CV and Public but decreased optimal score dramatically. But on Private, it became one of the most robust approaches\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2F6f0a8a87cf758aa4920353ace23c11cf%2Fkidney_3_dense_resize_val.png?generation=1707323076488076&alt=media)\nIn summary, Validation did not work (at least mine). It is not strange because of solo data point in CV, Public, and Private \n\n# Data \n\nI was using all train data except kidney_1_voi sample\nIn order to enlarge training data I have used 50um_LADAF_2020_31_kidney_pag from [Human Organ Atlas](https://human-organ-atlas.esrf.eu/search?organ=kidney) \nFor data normalization, I was using the approach proposed by @hengck23 - [percentile normalization](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456118#2552053)\n\n# Training setup\n\nI mostly stick to the 2.5D approach with 5 slices.\nI started from one view model and iterated along the last axis, but then I switched to a multiview and used slices by all three axes in the training\nI have used 512 square crops with Non Empty probability of 0.5, pretty much standard augmentations and CutMix within one organ and one view with 0.5 probability and 1.0 alpha :\n```python\n\"cutmix_transform\":lambda : [\n    A.PadIfNeeded(\n        min_height=crop_size,\n        min_width=crop_size,\n        always_apply=True,\n    ),\n    # Sample Non-Empty mask with prob 0.5\n    # Otherwise empty OR Non-Empty mask will be sampled\n    A.OneOrOther(\n        first=A.CropNonEmptyMaskIfExists(crop_size, crop_size), \n        second=A.RandomCrop(crop_size, crop_size), \n        p=0.5\n    ),\n],\n\"transform\": A.Compose(\n    [\n        A.ShiftScaleRotate(\n            scale_limit=0.2,\n        ),\n        # dihedral_aug\n        A.RandomRotate90(p=0.5),\n        A.HorizontalFlip(p=0.5),\n        A.VerticalFlip(p=0.5),\n        A.Transpose(p=0.5),\n        A.OneOf(\n            [\n                A.RandomBrightnessContrast(),\n                A.RandomBrightness(),\n                A.RandomGamma(),\n            ],\n            p=1.0,\n        ),\n        ToTensorV2(transpose_mask=True),\n    ]\n),\n\n\"do_cutmix\": True,\n\"cutmix_params\": {\"prob\": 0.5, \"alpha\": 1.0},\n```\nI was using Adam optimizer and reduced learning rate with CosineAnnealingLR starting from 1e-3 and ending with 1e-6\nRegarding loss function choice, I started with classical BCE+Dice loss and then tried to implement the loss function, which will directly optimize the metric, but unfortunately, it did not work well. Luckily, I have come across BoundaryLoss, which worked firstly comparable to BCE+Dice loss and then better. Interesting fact that the best (not selected) model was trained on BoundaryLoss + 0.5 Focal Symmetric Loss and scored 0.756 on Private, 0.867 on Public, and 0.916 on kidney_3_dense, which is a pretty much balanced score (of course, comparing to all other models score distribution 🙂)\nI was training for 30 epochs. repeating the original train set 3 times and the pseudo train set 2 times\nI have used a batch size of 14 samples and trained with DDP on 2 GPUs, so the final batch size was 28\n\nAfter I saw [the post](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461213#2594751) about promising results from 3D models, I started exploring 3D approaches, and they worked pretty well. In order to make it train without NaNs I have changed the optimization strategy and switched to SGD, with momentum=0.99, weight_decay=3e-5, nesterov=True and also changed starting learning rate to 1e-6 - taken from [monai example](https://github.com/Project-MONAI/tutorials/tree/main/modules/dynunet_pipeline)\nAs the overall image resolution of the image was increased dramatically, I had to reduce the batch to 3 on one GPU, so the aggregated batch size was 6. I was training in total for ~120K iterations\nAs for augmentations - they were pretty much the same as in the 2.5D setup, except from Zoom. \n```python\n\"transform_init\": lambda : mt.Compose(\n    [\n        mt.OneOf([\n            mt.RandRotate90d(keys=('image', 'mask'), prob=0.5, spatial_axes=(-3,-1)),\n            mt.RandRotate90d(keys=('image', 'mask'), prob=0.5, spatial_axes=(-2,-1)),\n            mt.RandRotate90d(keys=('image', 'mask'), prob=0.5, spatial_axes=(-3,-2))\n        ]),\n        mt.RandFlipd(keys=('image', 'mask'), prob=0.5, spatial_axis=-1),\n        mt.RandFlipd(keys=('image', 'mask'), prob=0.5, spatial_axis=-2),\n        mt.RandFlipd(keys=('image', 'mask'), prob=0.5, spatial_axis=-3),\n\n        mt.RandScaleIntensityd(keys=('image'), prob=0.5, factors=0.2)\n    ]\n),\n```\nZoom worked better on CV but worse on Public and also on Private (Why? - who knows …)\n\n# Neural Networks \n\nI was mostly using [EfficientNet family](https://arxiv.org/abs/1905.11946) as an Encoder (from noisy student weights), started from B3, then switched to B5, and unfortunately, B7 did not work well for me both on CV and Public \n\nInterestingly se_resnext50_32x4d performed not well on Public LB (0.852) and CV (0.909) but really well on Private (0.702)\n\nAs for Decoder I was mostly using Unet++. I have tried [Unet3+](https://arxiv.org/abs/2004.08790) but it showed considerably worse results\n\nAs for 3D Nets, I was using [DynUNet](https://monai-dev.readthedocs.io/en/fixes-sphinx/networks.html#dynunet) and adopted model architecture according to [next script](https://github.com/Project-MONAI/tutorials/blob/main/modules/dynunet_pipeline/create_network.py#L19). I have tried to use pretrained Unet from  [MONAI Model Zoo](https://monai.io/model-zoo.html) but it performed badly on all sets \n\n# Inference and Post Processing \n\nI was using 512 sliding window with 0.5 overlap, flip TTA, and last checkpoint from 2 folds.\nAfter switching to multi view model, I have also added multi view TTA\n\nThe next step was the creation of a kidney mask. I have tried several approaches \n1. Using segmentation net, trained on [this dataset](https://www.kaggle.com/datasets/squidinator/sennet-hoa-kidney-13-dense-full-kidney-masks) + slight post-processing for removing binary holes and small connected regions \n2. Using an algorithmic approach based on intensity thresholding, erosion, and dilation\n\nThe first one had a pretty high FP rate but nearly zero FN rate, while the second one had a pretty high FN rate. Both of them performed nearly ideal on kidney_3, so did not really influence the fold 0 scores, but an algorithmic approach cut out kidney regions for kidney_2 and dramatically reduced the fold 1 score. BUT at the same time, the second approach improved Public score (0.874->0.882). I understood that it was 90% overfit to Public LB, but I have decided to take the risk\n\n# Final Model\n\nFor final submission, I have selected the following ones:\n- Pure 2.5D -> algorithmic post-processing\n - Public: 0.886\n - Private: 0.681\n - Kidney 3 dense score: 0.917   \n- 2.5D (weight 3.0) blended with 3D (weight 1.0) \n - Public: 0.871\n - Private: 0.676\n - Kidney 3 dense score: ~0.918\nFor both models, I used 0.05 threshold \n\n# The most popular rubric of this competition: Not Selected Best Submission\n\nHere, I want to point out several of the most exciting approaches for me \n- Pure 2.5D but add Symmetric Focal loss with 0.5 coefficient \n - Public: 0.867\n - Private: 0.756\n - Kidney 3 dense score: 0.916\n- Resize 2d slices to 50 um/voxel for prediction and then resize back \n - Public: 0.799\n - Private: 0.753\n - Kidney 3 dense score: 0.907\n- Resize the whole volume with scipy.zoom to 50 um/voxel for prediction and than resize back \n - Public: 0.7 resize back \n - Public: 0.726\n - Private: 0.745\n - Kidney 3 dense score: Have not checked \n- Solo 3D model \n - Public: 0.849\n - Private: 0.723\n - Kidney 3 dense score: 0.915\nFor me, it was logical to pick first or second, but as for all other better submissions, it sounds to me like pure random.\n\n# Conclusions\n\nComputing metrics on one data sample leads to severe shakeups 🙂\n\n# Closing words\n\nI hope you have not fallen asleep while reading. Finally, I want to thank the entire Kaggle community, congratulate all participants and winners. Special thanks to Indian University Bloomington, University College London, Yashvardhan Jain (@yashvrdnjain), Claire Walsh (@clairewalsh), the Kaggle Team, and other organizers.",
      "votes": 28
    },
    {
      "id": 2642628,
      "postDate": "2024-02-08T09:49:07.647Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a> . Congratulations on another solo gold.  I have two questions.</p>\n<ol>\n<li><p><code>Resize to constant um/voxel(50um/voxel) makes the prediction more robust.</code> How do you resize them to 50um/voxel? And what is the reason you think they work well?</p></li>\n<li><p>Does the external data have annotation masks? If no, did you use pseudo label?</p></li>\n</ol>\n<p>Thank you for this great post again!</p>",
      "rawMarkdown": "Hi @vladimirsydor . Congratulations on another solo gold.  I have two questions.\n\n1. `Resize to constant um/voxel(50um/voxel) makes the prediction more robust. ` How do you resize them to 50um/voxel? And what is the reason you think they work well?\n\n2. Does the external data have annotation masks? If no, did you use pseudo label?\n\nThank you for this great post again!",
      "votes": 1,
      "replies": [
        {
          "id": 2642861,
          "postDate": "2024-02-08T13:22:04.820Z",
          "content": "<ol>\n<li>I have used 2 options: resize the whole volume with scipy.zoom and resize each slice with albumnetations. Second worked better. Intuition is that you get rid of additional domain shift caused by resolution difference </li>\n<li>I have used pseudo labels </li>\n</ol>\n<p>Congratulations on your 3d place!</p>",
          "rawMarkdown": "1.  I have used 2 options: resize the whole volume with scipy.zoom and resize each slice with albumnetations. Second worked better. Intuition is that you get rid of additional domain shift caused by resolution difference \n2. I have used pseudo labels \n\nCongratulations on your 3d place!",
          "votes": 2,
          "replies": [
            {
              "id": 2642963,
              "postDate": "2024-02-08T14:23:12.540Z",
              "content": "<blockquote>\n  <p>Intuition is that you get rid of additional domain shift caused by resolution difference</p>\n</blockquote>\n<p>Make sense. In my solution, I made the scaling center in training as 0.8, instead of 1.0. I can observe improvements by ablation experiments. Even our implemented in different methods, they share a similar idea, that is,  making the model more robust to the real word size in private to prevent domain shift.</p>\n<p>Thanks for your reply and congrats👍. </p>",
              "rawMarkdown": "> Intuition is that you get rid of additional domain shift caused by resolution difference\n\nMake sense. In my solution, I made the scaling center in training as 0.8, instead of 1.0. I can observe improvements by ablation experiments. Even our implemented in different methods, they share a similar idea, that is,  making the model more robust to the real word size in private to prevent domain shift.\n\nThanks for your reply and congrats👍. ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2641805,
      "postDate": "2024-02-07T17:14:35.543Z",
      "content": "<p>Best write-up so far!</p>",
      "rawMarkdown": "Best write-up so far!",
      "votes": 1
    },
    {
      "id": 2645451,
      "postDate": "2024-02-10T08:44:32.567Z",
      "content": "<p>Great and insightful work. Thanks, Volodymyr!</p>",
      "rawMarkdown": "Great and insightful work. Thanks, Volodymyr!"
    },
    {
      "id": 2643083,
      "postDate": "2024-02-08T15:51:16.237Z",
      "content": "<p>Congratulations!!!!</p>",
      "rawMarkdown": "Congratulations!!!!"
    },
    {
      "id": 2642744,
      "postDate": "2024-02-08T11:30:36.627Z",
      "content": "<p>Congratulations on topping this competition with 6th place. Thanks for sharing the details of your solution with code examples. </p>",
      "rawMarkdown": "Congratulations on topping this competition with 6th place. Thanks for sharing the details of your solution with code examples. "
    },
    {
      "id": 2642209,
      "postDate": "2024-02-08T02:59:01.677Z",
      "content": "<p>stability after shakeup is so crucial and difficult in this competition, and so rare. Your solution is the best one so far.</p>",
      "rawMarkdown": "stability after shakeup is so crucial and difficult in this competition, and so rare. Your solution is the best one so far."
    },
    {
      "id": 2642079,
      "postDate": "2024-02-07T22:02:59.600Z",
      "content": "<p>Very-very interesting!</p>",
      "rawMarkdown": "Very-very interesting!"
    },
    {
      "id": 2642902,
      "postDate": "2024-02-08T13:49:30.760Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2642628,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2024-02-08T09:49:07.647000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a> . Congratulations on another solo gold.  I have two questions.</p>\n<ol>\n<li><p><code>Resize to constant um/voxel(50um/voxel) makes the prediction more robust.</code> How do you resize them to 50um/voxel? And what is the reason you think they work well?</p></li>\n<li><p>Does the external data have annotation masks? If no, did you use pseudo label?</p></li>\n</ol>\n<p>Thank you for this great post again!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2642861,
          "author_name": "Volodymyr",
          "author_url": "",
          "post_date": "2024-02-08T13:22:04.820000",
          "content": "<ol>\n<li>I have used 2 options: resize the whole volume with scipy.zoom and resize each slice with albumnetations. Second worked better. Intuition is that you get rid of additional domain shift caused by resolution difference </li>\n<li>I have used pseudo labels </li>\n</ol>\n<p>Congratulations on your 3d place!</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2642963,
              "author_name": "ForcewithMe",
              "author_url": "",
              "post_date": "2024-02-08T14:23:12.540000",
              "content": "<blockquote>\n  <p>Intuition is that you get rid of additional domain shift caused by resolution difference</p>\n</blockquote>\n<p>Make sense. In my solution, I made the scaling center in training as 0.8, instead of 1.0. I can observe improvements by ablation experiments. Even our implemented in different methods, they share a similar idea, that is,  making the model more robust to the real word size in private to prevent domain shift.</p>\n<p>Thanks for your reply and congrats👍. </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2641805,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2024-02-07T17:14:35.543000",
      "content": "<p>Best write-up so far!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2645451,
      "author_name": "Andrii Shevtsov",
      "author_url": "",
      "post_date": "2024-02-10T08:44:32.567000",
      "content": "<p>Great and insightful work. Thanks, Volodymyr!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2643083,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-08T15:51:16.237000",
      "content": "<p>Congratulations!!!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2642744,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2024-02-08T11:30:36.627000",
      "content": "<p>Congratulations on topping this competition with 6th place. Thanks for sharing the details of your solution with code examples. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2642209,
      "author_name": "LLLEEEOOOH",
      "author_url": "",
      "post_date": "2024-02-08T02:59:01.677000",
      "content": "<p>stability after shakeup is so crucial and difficult in this competition, and so rare. Your solution is the best one so far.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2642079,
      "author_name": "Volodymyr Pivoshenko 🇺🇦",
      "author_url": "",
      "post_date": "2024-02-07T22:02:59.600000",
      "content": "<p>Very-very interesting!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2642902,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-08T13:49:30.760000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2641732": "Hi, Kagglers!\n\nAfter finding yourself on the Leaderboard after Grand Shakeup and restoring your mental health, we can dive deep into 6th place solution, but before this, a few very important words:\n*I would like to thank the Armed Forces of Ukraine, Security Service of Ukraine, Defence Intelligence of Ukraine, and the State Emergency Service of Ukraine for providing safety and security to participate in this great competition, complete this work, and help science, technology, and business not to stop but to move forward.*\n\n# Validation. Not really \n\nI was using for validation 2 organs (in different folds): kidney_3_dense and kidney_2. After releasing a [fast version of 3D Surface Dice](https://www.kaggle.com/code/junkoda/fast-surface-dice-computation), I was able to compute validation scores while training, and I received the next insights:\n- Tracking the score on kidney_2 was useless for me. The validation score decreased from epochs 2-3 on kidney_2\n- Scores on kidney_3_dense were meaningful for checking “radical” features, like additional data and new losses. But then optimal score fluctuated between 0.9-0.925 dice without any reasonable correlation with Public or Private score\n- The optimal threshold on kidney_3_dense was optimal for Private, Public, and kidney_3_dense scores - 0.1 and lower \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2Fa4050c6e4e6d27662cfddc62201e2f71%2Fkidney_3_dense_val.png?generation=1707321782497972&alt=media)\n- Resize to constant um/voxel (I have picked 50 um/voxel) for prediction increased optimal threshold both on CV and Public but decreased optimal score dramatically. But on Private, it became one of the most robust approaches\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2F6f0a8a87cf758aa4920353ace23c11cf%2Fkidney_3_dense_resize_val.png?generation=1707323076488076&alt=media)\nIn summary, Validation did not work (at least mine). It is not strange because of solo data point in CV, Public, and Private \n\n# Data \n\nI was using all train data except kidney_1_voi sample\nIn order to enlarge training data I have used 50um_LADAF_2020_31_kidney_pag from [Human Organ Atlas](https://human-organ-atlas.esrf.eu/search?organ=kidney) \nFor data normalization, I was using the approach proposed by @hengck23 - [percentile normalization](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456118#2552053)\n\n# Training setup\n\nI mostly stick to the 2.5D approach with 5 slices.\nI started from one view model and iterated along the last axis, but then I switched to a multiview and used slices by all three axes in the training\nI have used 512 square crops with Non Empty probability of 0.5, pretty much standard augmentations and CutMix within one organ and one view with 0.5 probability and 1.0 alpha :\n```python\n\"cutmix_transform\":lambda : [\n    A.PadIfNeeded(\n        min_height=crop_size,\n        min_width=crop_size,\n        always_apply=True,\n    ),\n    # Sample Non-Empty mask with prob 0.5\n    # Otherwise empty OR Non-Empty mask will be sampled\n    A.OneOrOther(\n        first=A.CropNonEmptyMaskIfExists(crop_size, crop_size), \n        second=A.RandomCrop(crop_size, crop_size), \n        p=0.5\n    ),\n],\n\"transform\": A.Compose(\n    [\n        A.ShiftScaleRotate(\n            scale_limit=0.2,\n        ),\n        # dihedral_aug\n        A.RandomRotate90(p=0.5),\n        A.HorizontalFlip(p=0.5),\n        A.VerticalFlip(p=0.5),\n        A.Transpose(p=0.5),\n        A.OneOf(\n            [\n                A.RandomBrightnessContrast(),\n                A.RandomBrightness(),\n                A.RandomGamma(),\n            ],\n            p=1.0,\n        ),\n        ToTensorV2(transpose_mask=True),\n    ]\n),\n\n\"do_cutmix\": True,\n\"cutmix_params\": {\"prob\": 0.5, \"alpha\": 1.0},\n```\nI was using Adam optimizer and reduced learning rate with CosineAnnealingLR starting from 1e-3 and ending with 1e-6\nRegarding loss function choice, I started with classical BCE+Dice loss and then tried to implement the loss function, which will directly optimize the metric, but unfortunately, it did not work well. Luckily, I have come across BoundaryLoss, which worked firstly comparable to BCE+Dice loss and then better. Interesting fact that the best (not selected) model was trained on BoundaryLoss + 0.5 Focal Symmetric Loss and scored 0.756 on Private, 0.867 on Public, and 0.916 on kidney_3_dense, which is a pretty much balanced score (of course, comparing to all other models score distribution 🙂)\nI was training for 30 epochs. repeating the original train set 3 times and the pseudo train set 2 times\nI have used a batch size of 14 samples and trained with DDP on 2 GPUs, so the final batch size was 28\n\nAfter I saw [the post](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461213#2594751) about promising results from 3D models, I started exploring 3D approaches, and they worked pretty well. In order to make it train without NaNs I have changed the optimization strategy and switched to SGD, with momentum=0.99, weight_decay=3e-5, nesterov=True and also changed starting learning rate to 1e-6 - taken from [monai example](https://github.com/Project-MONAI/tutorials/tree/main/modules/dynunet_pipeline)\nAs the overall image resolution of the image was increased dramatically, I had to reduce the batch to 3 on one GPU, so the aggregated batch size was 6. I was training in total for ~120K iterations\nAs for augmentations - they were pretty much the same as in the 2.5D setup, except from Zoom. \n```python\n\"transform_init\": lambda : mt.Compose(\n    [\n        mt.OneOf([\n            mt.RandRotate90d(keys=('image', 'mask'), prob=0.5, spatial_axes=(-3,-1)),\n            mt.RandRotate90d(keys=('image', 'mask'), prob=0.5, spatial_axes=(-2,-1)),\n            mt.RandRotate90d(keys=('image', 'mask'), prob=0.5, spatial_axes=(-3,-2))\n        ]),\n        mt.RandFlipd(keys=('image', 'mask'), prob=0.5, spatial_axis=-1),\n        mt.RandFlipd(keys=('image', 'mask'), prob=0.5, spatial_axis=-2),\n        mt.RandFlipd(keys=('image', 'mask'), prob=0.5, spatial_axis=-3),\n\n        mt.RandScaleIntensityd(keys=('image'), prob=0.5, factors=0.2)\n    ]\n),\n```\nZoom worked better on CV but worse on Public and also on Private (Why? - who knows …)\n\n# Neural Networks \n\nI was mostly using [EfficientNet family](https://arxiv.org/abs/1905.11946) as an Encoder (from noisy student weights), started from B3, then switched to B5, and unfortunately, B7 did not work well for me both on CV and Public \n\nInterestingly se_resnext50_32x4d performed not well on Public LB (0.852) and CV (0.909) but really well on Private (0.702)\n\nAs for Decoder I was mostly using Unet++. I have tried [Unet3+](https://arxiv.org/abs/2004.08790) but it showed considerably worse results\n\nAs for 3D Nets, I was using [DynUNet](https://monai-dev.readthedocs.io/en/fixes-sphinx/networks.html#dynunet) and adopted model architecture according to [next script](https://github.com/Project-MONAI/tutorials/blob/main/modules/dynunet_pipeline/create_network.py#L19). I have tried to use pretrained Unet from  [MONAI Model Zoo](https://monai.io/model-zoo.html) but it performed badly on all sets \n\n# Inference and Post Processing \n\nI was using 512 sliding window with 0.5 overlap, flip TTA, and last checkpoint from 2 folds.\nAfter switching to multi view model, I have also added multi view TTA\n\nThe next step was the creation of a kidney mask. I have tried several approaches \n1. Using segmentation net, trained on [this dataset](https://www.kaggle.com/datasets/squidinator/sennet-hoa-kidney-13-dense-full-kidney-masks) + slight post-processing for removing binary holes and small connected regions \n2. Using an algorithmic approach based on intensity thresholding, erosion, and dilation\n\nThe first one had a pretty high FP rate but nearly zero FN rate, while the second one had a pretty high FN rate. Both of them performed nearly ideal on kidney_3, so did not really influence the fold 0 scores, but an algorithmic approach cut out kidney regions for kidney_2 and dramatically reduced the fold 1 score. BUT at the same time, the second approach improved Public score (0.874->0.882). I understood that it was 90% overfit to Public LB, but I have decided to take the risk\n\n# Final Model\n\nFor final submission, I have selected the following ones:\n- Pure 2.5D -> algorithmic post-processing\n - Public: 0.886\n - Private: 0.681\n - Kidney 3 dense score: 0.917   \n- 2.5D (weight 3.0) blended with 3D (weight 1.0) \n - Public: 0.871\n - Private: 0.676\n - Kidney 3 dense score: ~0.918\nFor both models, I used 0.05 threshold \n\n# The most popular rubric of this competition: Not Selected Best Submission\n\nHere, I want to point out several of the most exciting approaches for me \n- Pure 2.5D but add Symmetric Focal loss with 0.5 coefficient \n - Public: 0.867\n - Private: 0.756\n - Kidney 3 dense score: 0.916\n- Resize 2d slices to 50 um/voxel for prediction and then resize back \n - Public: 0.799\n - Private: 0.753\n - Kidney 3 dense score: 0.907\n- Resize the whole volume with scipy.zoom to 50 um/voxel for prediction and than resize back \n - Public: 0.7 resize back \n - Public: 0.726\n - Private: 0.745\n - Kidney 3 dense score: Have not checked \n- Solo 3D model \n - Public: 0.849\n - Private: 0.723\n - Kidney 3 dense score: 0.915\nFor me, it was logical to pick first or second, but as for all other better submissions, it sounds to me like pure random.\n\n# Conclusions\n\nComputing metrics on one data sample leads to severe shakeups 🙂\n\n# Closing words\n\nI hope you have not fallen asleep while reading. Finally, I want to thank the entire Kaggle community, congratulate all participants and winners. Special thanks to Indian University Bloomington, University College London, Yashvardhan Jain (@yashvrdnjain), Claire Walsh (@clairewalsh), the Kaggle Team, and other organizers.",
    "2642628": "Hi @vladimirsydor . Congratulations on another solo gold.  I have two questions.\n\n1. `Resize to constant um/voxel(50um/voxel) makes the prediction more robust. ` How do you resize them to 50um/voxel? And what is the reason you think they work well?\n\n2. Does the external data have annotation masks? If no, did you use pseudo label?\n\nThank you for this great post again!",
    "2641805": "Best write-up so far!",
    "2645451": "Great and insightful work. Thanks, Volodymyr!",
    "2643083": "Congratulations!!!!",
    "2642744": "Congratulations on topping this competition with 6th place. Thanks for sharing the details of your solution with code examples. ",
    "2642209": "stability after shakeup is so crucial and difficult in this competition, and so rare. Your solution is the best one so far.",
    "2642079": "Very-very interesting!",
    "2642902": ""
  }
}