{
  "id": 475074,
  "title": "[3rd Place solution] Refine from Sparse to Dense",
  "url": "/competitions/blood-vessel-segmentation/writeups/forcewithme-3rd-place-solution-refine-from-sparse-",
  "author_name": "",
  "post_date": "2024-02-09T15:07:28.710Z",
  "votes": 69,
  "comment_count": 19,
  "views": 0,
  "content": "<p>First and foremost, I would like to extend my gratitude to the organizers and the official Kaggle team for orchestrating such an outstanding competition. I joined the contest at a very late stage. Despite having some experience with segmentation competitions, I must express my appreciation to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , <a href=\"https://www.kaggle.com/yoyobar\" target=\"_blank\">@yoyobar</a> , and <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> (implementation of metric), as well as the other community participants for their open-source contributions and discussions, which allowed me to quickly get up to speed with this contest.</p>\n<p>My approach was strikingly straightforward, relying solely on <strong>2D models</strong> and only utilizing <strong>smp</strong> (segmentation models pytorch) and <strong>timm</strong> (pytorch image models) in the whole training and inference pipeline. </p>\n<h2>Global key points</h2>\n<ol>\n<li><strong>Refining labels from sparse to dense.</strong></li>\n<li><strong>Emulating the magnification factor of the test set.</strong></li>\n<li><strong>Maintaining an appropriate resolution.</strong></li>\n</ol>\n<h2>1. From Sparse to Dense</h2>\n<p>Given that half of the training set has dense labels (kidney 1, kidney 3 dense), and the other half was sparse, utilizing dense labels to refine sparse ones was a crucial step. The overall process entailed:</p>\n<ol>\n<li>Training UNet(maxViT512) and UNet(EfficientNetv2s) using kidney 1 and kidney 3 dense.</li>\n<li>Generating supplemental labels for kidney 3 sparse using the trained UNet maxViT512 and UNet EfficientNetv2s models.</li>\n<li>Resuming the training of UNet maxViT512 and UNet EfficientNetv2s for a few epochs with kidney 1, kidney 3 (dense, sparse plus supplemental labels).</li>\n<li>Repeating the step2 on kidney 2.</li>\n<li>Training three UNet models (with EfficientNetv2s, SeResNext101, MaxViT512) and one UNet++ using all real labels from all kidney plus pseudo labels.</li>\n</ol>\n<p>Note: As the organizers disclosed the proportion of annotations within kidney 3 and kidney 2, I endeavored to select thresholds based on pixel quantity as close as possible to the official proportion when choosing threshold values for pseudo label.</p>\n<h2>2. Emulating the Magnification of the Private Test Set</h2>\n<p>A pivotal reason for my decision to participate in this competition was the disclosure of the magnification factors for the training and test sets by the hosts. The training set had a magnification of 50um/voxel, the public test set was the same at 50um/voxel, while the private test set was at 63um/voxel. A larger magnification factor implies a lower resolution. For instance, a 600um object would occupy 12 pixels in both the training and public sets, but only 10 pixels in the validation set. Hence, during training, <strong>I set the scaling center to 0.8</strong>, rather than 1, with a scaling range of 0.55 to 1.05, to simulate the private test set.</p>\n<pre><code>.ShiftScaleRotate(shift_limit=.,\n                    =(-., .),  \n                    =,\n                    \n                    =,\n                    =.),\n</code></pre>\n<h2>3. Maintaining an Appropriate Resolution</h2>\n<p>In this competition, training and inference along the x-axis, y-axis, and z-axis separately was a very important trick. However, this introduced a significant risk. The entire test set contained 1500 slices, with the public test set accounting for 67% and the private test set for 33%. This means that the private test set comprised only about <strong>500 slices</strong>. Inferring along the z-axis with a higher resolution (e.g., 1024) was feasible. But if inferring along the y-axis or x-axis, it would mean that one of the edges would only be 500 pixels long. At that point, if the model and code were configured for a larger resolution (say 1024), there would be a substantial risk of a huge shake down.</p>\n<p>My models primarily operated at a resolution of 512, with one model switching to higher resolution weights for larger resolution slices when the slice have appropriate resolution.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Backbone</th>\n<th>Resolution</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>UNet</td>\n<td>MaxViT-Large 512</td>\n<td>512</td>\n<td>0.846</td>\n<td>0.727(submission1)</td>\n</tr>\n<tr>\n<td>UNet</td>\n<td>SeResNext</td>\n<td>512</td>\n<td>0.819</td>\n<td>0.753</td>\n</tr>\n<tr>\n<td>UNet</td>\n<td>Efficiennet_v2_s</td>\n<td>448, 832</td>\n<td>0.799</td>\n<td>0.703</td>\n</tr>\n<tr>\n<td>UNet++</td>\n<td>Efficiennet_v2_l</td>\n<td>512</td>\n<td>0.817</td>\n<td>0.692</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td>-</td>\n<td>-</td>\n<td>0.846</td>\n<td>0.727(submission2)</td>\n</tr>\n</tbody>\n</table>\n<h2>4. Train on all data if convergence is Stable</h2>\n<p>During the early stages of the competition, whether validating on kidney 2 or kidney 3, I observed that if I trained for 20 epochs, after the initial few epochs, the dice coefficient (not surface dice) variation on the validation set was very minimal, with the MaxVit512 large model exhibiting the least fluctuation. Considering that we only had three kidneys, I decided to train on all kidneys directly after completing the pseudo labeling process, given the stability in convergence.</p>\n<h2>5. Minimizing the Impact of Threshold Values</h2>\n<p>I am grateful for the method provided by <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> for calculating metrics. My most stable single model was able to maintain very minor fluctuations in the surface dice score (less than 1) within a threshold range of 0.2. After model fusion, the stable threshold range could be potentially in 0.3~ 0.4. A stable threshold is extremely crucial in segmentation competitions. In this competition, as my final model lacked a validation set, I had to utilize thresholds searched with earlier trained models that included a validation set and apply them to the final version of the model. Fortunately, the models trained on the full dataset appeared to possess threshold values very close to those from the earlier models trained with k1+k2 (sparse), and validate on k3. At the same time, the fluctuation of threshold values across kidney 3 dense, public, and private was very small.</p>\n<h2>6. Heavy augmentation on intensity.</h2>\n<p>As mentioned by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , difference kidneys has large variance on intensity. So I used a heavy intensity augmentation.</p>\n<pre><code>.RandomBrightnessContrast(p=.),\n.RandomGamma(p=.),\n</code></pre>\n<h2>7. Quick ablation</h2>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Backbone</th>\n<th>points mentioned above</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>UNet</td>\n<td>MaxViT-Large 512</td>\n<td>3, 5, 6</td>\n<td>0.818</td>\n<td>0.586</td>\n</tr>\n<tr>\n<td>UNet</td>\n<td>MaxViT-Large 512</td>\n<td>3, 4, 5, 6</td>\n<td>0.857</td>\n<td>0.633</td>\n</tr>\n<tr>\n<td>UNet</td>\n<td>MaxViT-Large 512</td>\n<td>2, 3, 4, 5, 6</td>\n<td>0.849</td>\n<td>0.652</td>\n</tr>\n<tr>\n<td>UNet</td>\n<td>MaxViT-Large 512</td>\n<td>1, 2, 3, 4, 5, 6</td>\n<td>0.846</td>\n<td>0.727</td>\n</tr>\n</tbody>\n</table>\n<h2>8. Final Submission</h2>\n<p>My final submissions were a single model of MaxViT and a ensemble of the four models. Surprisingly, both submissions scored same at 0.727.  I did not use any form of weighting and MaxViT only constituted a quarter of the ensemble submission, but their scores were totally the same on private LB. Even more astonishing was that the single-model score of SeResNext on private LB turned out to be the highest. Its cv was nothing extraordinary, its convergence was not more stable than MaxViT's, and its public leaderboard score was not high, so I had no reason to choose it.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7285387%2F8fa373c75008e7cf3064b7f4c4089175%2F1.png?generation=1707272539926386&amp;alt=media\" alt=\"1\"></p>\n<p>Finally, I would like to extend my gratitude once again to the organizers, Kaggle, and all the participants again!</p>\n<hr>\n<p><strong>Inference(Submission)</strong> Code is published:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/forcewithme/sennet-final-submission2\" target=\"_blank\">MaxVit512 scored 0.727</a></li>\n<li><a href=\"https://www.kaggle.com/code/forcewithme/sennet-top3-final-submission?scriptVersionId=162311084\" target=\"_blank\">Ensemble submission scored 0.727</a></li>\n</ol>\n<p><strong>Training</strong> Code is published in the <a href=\"https://www.kaggle.com/datasets/forcewithme/sennettop3-training-code/data\" target=\"_blank\">kaggle dataset</a>. </p>",
  "messages": [
    {
      "id": "2640599",
      "postDate": "02/07/2024 02:32:28",
      "content": "<p>First and foremost, I would like to extend my gratitude to the organizers and the official Kaggle team for orchestrating such an outstanding competition. I joined the contest at a very late stage. Despite having some experience with segmentation competitions, I must express my appreciation to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , <a href=\"https://www.kaggle.com/yoyobar\" target=\"_blank\">@yoyobar</a> , and <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> (implementation of metric), as well as the other community participants for their open-source contributions and discussions, which allowed me to quickly get up to speed with this contest.</p>\n<p>My approach was strikingly straightforward, relying solely on <strong>2D models</strong> and only utilizing <strong>smp</strong> (segmentation models pytorch) and <strong>timm</strong> (pytorch image models) in the whole training and inference pipeline. </p>\n<h2>Global key points</h2>\n<ol>\n<li><strong>Refining labels from sparse to dense.</strong></li>\n<li><strong>Emulating the magnification factor of the test set.</strong></li>\n<li><strong>Maintaining an appropriate resolution.</strong></li>\n</ol>\n<h2>1. From Sparse to Dense</h2>\n<p>Given that half of the training set has dense labels (kidney 1, kidney 3 dense), and the other half was sparse, utilizing dense labels to refine sparse ones was a crucial step. The overall process entailed:</p>\n<ol>\n<li>Training UNet(maxViT512) and UNet(EfficientNetv2s) using kidney 1 and kidney 3 dense.</li>\n<li>Generating supplemental labels for kidney 3 sparse using the trained UNet maxViT512 and UNet EfficientNetv2s models.</li>\n<li>Resuming the training of UNet maxViT512 and UNet EfficientNetv2s for a few epochs with kidney 1, kidney 3 (dense, sparse plus supplemental labels).</li>\n<li>Repeating the step2 on kidney 2.</li>\n<li>Training three UNet models (with EfficientNetv2s, SeResNext101, MaxViT512) and one UNet++ using all real labels from all kidney plus pseudo labels.</li>\n</ol>\n<p>Note: As the organizers disclosed the proportion of annotations within kidney 3 and kidney 2, I endeavored to select thresholds based on pixel quantity as close as possible to the official proportion when choosing threshold values for pseudo label.</p>\n<h2>2. Emulating the Magnification of the Private Test Set</h2>\n<p>A pivotal reason for my decision to participate in this competition was the disclosure of the magnification factors for the training and test sets by the hosts. The training set had a magnification of 50um/voxel, the public test set was the same at 50um/voxel, while the private test set was at 63um/voxel. A larger magnification factor implies a lower resolution. For instance, a 600um object would occupy 12 pixels in both the training and public sets, but only 10 pixels in the validation set. Hence, during training, <strong>I set the scaling center to 0.8</strong>, rather than 1, with a scaling range of 0.55 to 1.05, to simulate the private test set.</p>\n<pre><code>.ShiftScaleRotate(shift_limit=.,\n                    =(-., .),  \n                    =,\n                    \n                    =,\n                    =.),\n</code></pre>\n<h2>3. Maintaining an Appropriate Resolution</h2>\n<p>In this competition, training and inference along the x-axis, y-axis, and z-axis separately was a very important trick. However, this introduced a significant risk. The entire test set contained 1500 slices, with the public test set accounting for 67% and the private test set for 33%. This means that the private test set comprised only about <strong>500 slices</strong>. Inferring along the z-axis with a higher resolution (e.g., 1024) was feasible. But if inferring along the y-axis or x-axis, it would mean that one of the edges would only be 500 pixels long. At that point, if the model and code were configured for a larger resolution (say 1024), there would be a substantial risk of a huge shake down.</p>\n<p>My models primarily operated at a resolution of 512, with one model switching to higher resolution weights for larger resolution slices when the slice have appropriate resolution.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Backbone</th>\n<th>Resolution</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>UNet</td>\n<td>MaxViT-Large 512</td>\n<td>512</td>\n<td>0.846</td>\n<td>0.727(submission1)</td>\n</tr>\n<tr>\n<td>UNet</td>\n<td>SeResNext</td>\n<td>512</td>\n<td>0.819</td>\n<td>0.753</td>\n</tr>\n<tr>\n<td>UNet</td>\n<td>Efficiennet_v2_s</td>\n<td>448, 832</td>\n<td>0.799</td>\n<td>0.703</td>\n</tr>\n<tr>\n<td>UNet++</td>\n<td>Efficiennet_v2_l</td>\n<td>512</td>\n<td>0.817</td>\n<td>0.692</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td>-</td>\n<td>-</td>\n<td>0.846</td>\n<td>0.727(submission2)</td>\n</tr>\n</tbody>\n</table>\n<h2>4. Train on all data if convergence is Stable</h2>\n<p>During the early stages of the competition, whether validating on kidney 2 or kidney 3, I observed that if I trained for 20 epochs, after the initial few epochs, the dice coefficient (not surface dice) variation on the validation set was very minimal, with the MaxVit512 large model exhibiting the least fluctuation. Considering that we only had three kidneys, I decided to train on all kidneys directly after completing the pseudo labeling process, given the stability in convergence.</p>\n<h2>5. Minimizing the Impact of Threshold Values</h2>\n<p>I am grateful for the method provided by <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> for calculating metrics. My most stable single model was able to maintain very minor fluctuations in the surface dice score (less than 1) within a threshold range of 0.2. After model fusion, the stable threshold range could be potentially in 0.3~ 0.4. A stable threshold is extremely crucial in segmentation competitions. In this competition, as my final model lacked a validation set, I had to utilize thresholds searched with earlier trained models that included a validation set and apply them to the final version of the model. Fortunately, the models trained on the full dataset appeared to possess threshold values very close to those from the earlier models trained with k1+k2 (sparse), and validate on k3. At the same time, the fluctuation of threshold values across kidney 3 dense, public, and private was very small.</p>\n<h2>6. Heavy augmentation on intensity.</h2>\n<p>As mentioned by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , difference kidneys has large variance on intensity. So I used a heavy intensity augmentation.</p>\n<pre><code>.RandomBrightnessContrast(p=.),\n.RandomGamma(p=.),\n</code></pre>\n<h2>7. Quick ablation</h2>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Backbone</th>\n<th>points mentioned above</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>UNet</td>\n<td>MaxViT-Large 512</td>\n<td>3, 5, 6</td>\n<td>0.818</td>\n<td>0.586</td>\n</tr>\n<tr>\n<td>UNet</td>\n<td>MaxViT-Large 512</td>\n<td>3, 4, 5, 6</td>\n<td>0.857</td>\n<td>0.633</td>\n</tr>\n<tr>\n<td>UNet</td>\n<td>MaxViT-Large 512</td>\n<td>2, 3, 4, 5, 6</td>\n<td>0.849</td>\n<td>0.652</td>\n</tr>\n<tr>\n<td>UNet</td>\n<td>MaxViT-Large 512</td>\n<td>1, 2, 3, 4, 5, 6</td>\n<td>0.846</td>\n<td>0.727</td>\n</tr>\n</tbody>\n</table>\n<h2>8. Final Submission</h2>\n<p>My final submissions were a single model of MaxViT and a ensemble of the four models. Surprisingly, both submissions scored same at 0.727.  I did not use any form of weighting and MaxViT only constituted a quarter of the ensemble submission, but their scores were totally the same on private LB. Even more astonishing was that the single-model score of SeResNext on private LB turned out to be the highest. Its cv was nothing extraordinary, its convergence was not more stable than MaxViT's, and its public leaderboard score was not high, so I had no reason to choose it.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7285387%2F8fa373c75008e7cf3064b7f4c4089175%2F1.png?generation=1707272539926386&amp;alt=media\" alt=\"1\"></p>\n<p>Finally, I would like to extend my gratitude once again to the organizers, Kaggle, and all the participants again!</p>\n<hr>\n<p><strong>Inference(Submission)</strong> Code is published:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/forcewithme/sennet-final-submission2\" target=\"_blank\">MaxVit512 scored 0.727</a></li>\n<li><a href=\"https://www.kaggle.com/code/forcewithme/sennet-top3-final-submission?scriptVersionId=162311084\" target=\"_blank\">Ensemble submission scored 0.727</a></li>\n</ol>\n<p><strong>Training</strong> Code is published in the <a href=\"https://www.kaggle.com/datasets/forcewithme/sennettop3-training-code/data\" target=\"_blank\">kaggle dataset</a>. </p>",
      "rawMarkdown": "First and foremost, I would like to extend my gratitude to the organizers and the official Kaggle team for orchestrating such an outstanding competition. I joined the contest at a very late stage. Despite having some experience with segmentation competitions, I must express my appreciation to @hengck23 , @yoyobar , and @junkoda (implementation of metric), as well as the other community participants for their open-source contributions and discussions, which allowed me to quickly get up to speed with this contest.\n\nMy approach was strikingly straightforward, relying solely on **2D models** and only utilizing **smp** (segmentation models pytorch) and **timm** (pytorch image models) in the whole training and inference pipeline. \n\n## Global key points\n1.   **Refining labels from sparse to dense.**\n2.  **Emulating the magnification factor of the test set.**\n3.  **Maintaining an appropriate resolution.**\n\n## 1. From Sparse to Dense\nGiven that half of the training set has dense labels (kidney 1, kidney 3 dense), and the other half was sparse, utilizing dense labels to refine sparse ones was a crucial step. The overall process entailed:\n\n1. Training UNet(maxViT512) and UNet(EfficientNetv2s) using kidney 1 and kidney 3 dense.\n2. Generating supplemental labels for kidney 3 sparse using the trained UNet maxViT512 and UNet EfficientNetv2s models.\n3. Resuming the training of UNet maxViT512 and UNet EfficientNetv2s for a few epochs with kidney 1, kidney 3 (dense, sparse plus supplemental labels).\n4. Repeating the step2 on kidney 2.\n5. Training three UNet models (with EfficientNetv2s, SeResNext101, MaxViT512) and one UNet++ using all real labels from all kidney plus pseudo labels.\n\nNote: As the organizers disclosed the proportion of annotations within kidney 3 and kidney 2, I endeavored to select thresholds based on pixel quantity as close as possible to the official proportion when choosing threshold values for pseudo label.\n\n## 2. Emulating the Magnification of the Private Test Set\nA pivotal reason for my decision to participate in this competition was the disclosure of the magnification factors for the training and test sets by the hosts. The training set had a magnification of 50um/voxel, the public test set was the same at 50um/voxel, while the private test set was at 63um/voxel. A larger magnification factor implies a lower resolution. For instance, a 600um object would occupy 12 pixels in both the training and public sets, but only 10 pixels in the validation set. Hence, during training, **I set the scaling center to 0.8**, rather than 1, with a scaling range of 0.55 to 1.05, to simulate the private test set.\n\n```\nA.ShiftScaleRotate(shift_limit=0.3,\n                    scale_limit=(-0.45, 0.05),  \n                    rotate_limit=45,\n                    # value=0,\n                    border_mode=4,\n                    p=0.95),\n```\n\n## 3. Maintaining an Appropriate Resolution\nIn this competition, training and inference along the x-axis, y-axis, and z-axis separately was a very important trick. However, this introduced a significant risk. The entire test set contained 1500 slices, with the public test set accounting for 67% and the private test set for 33%. This means that the private test set comprised only about **500 slices**. Inferring along the z-axis with a higher resolution (e.g., 1024) was feasible. But if inferring along the y-axis or x-axis, it would mean that one of the edges would only be 500 pixels long. At that point, if the model and code were configured for a larger resolution (say 1024), there would be a substantial risk of a huge shake down.\n\nMy models primarily operated at a resolution of 512, with one model switching to higher resolution weights for larger resolution slices when the slice have appropriate resolution.\n\n| Model | Backbone | Resolution | public | private|\n| --- | --- | --- | --- | --- |\n| UNet | MaxViT-Large 512 | 512 | 0.846 | 0.727(submission1) |\n| UNet | SeResNext | 512 | 0.819 | 0.753 |\n| UNet | Efficiennet_v2_s | 448, 832 | 0.799 | 0.703 | \n| UNet++ | Efficiennet_v2_l | 512 | 0.817 | 0.692 |\n| ensemble | - | - | 0.846 | 0.727(submission2) | \n\n\n## 4. Train on all data if convergence is Stable\nDuring the early stages of the competition, whether validating on kidney 2 or kidney 3, I observed that if I trained for 20 epochs, after the initial few epochs, the dice coefficient (not surface dice) variation on the validation set was very minimal, with the MaxVit512 large model exhibiting the least fluctuation. Considering that we only had three kidneys, I decided to train on all kidneys directly after completing the pseudo labeling process, given the stability in convergence.\n\n## 5. Minimizing the Impact of Threshold Values\nI am grateful for the method provided by @junkoda for calculating metrics. My most stable single model was able to maintain very minor fluctuations in the surface dice score (less than 1) within a threshold range of 0.2. After model fusion, the stable threshold range could be potentially in 0.3~ 0.4. A stable threshold is extremely crucial in segmentation competitions. In this competition, as my final model lacked a validation set, I had to utilize thresholds searched with earlier trained models that included a validation set and apply them to the final version of the model. Fortunately, the models trained on the full dataset appeared to possess threshold values very close to those from the earlier models trained with k1+k2 (sparse), and validate on k3. At the same time, the fluctuation of threshold values across kidney 3 dense, public, and private was very small.\n\n## 6. Heavy augmentation on intensity.\nAs mentioned by @hengck23 , difference kidneys has large variance on intensity. So I used a heavy intensity augmentation.\n```\nA.RandomBrightnessContrast(p=1.0),\nA.RandomGamma(p=0.8),\n```\n## 7. Quick ablation\n| Model | Backbone | points mentioned above | public | private|\n| --- | --- | --- | --- | --- |\n| UNet | MaxViT-Large 512 | 3, 5, 6 | 0.818 | 0.586 |\n| UNet | MaxViT-Large 512 | 3, 4, 5, 6 | 0.857 | 0.633 |\n| UNet | MaxViT-Large 512 | 2, 3, 4, 5, 6 | 0.849 | 0.652 |\n| UNet | MaxViT-Large 512 | 1, 2, 3, 4, 5, 6 | 0.846 | 0.727 |\n\n## 8. Final Submission\nMy final submissions were a single model of MaxViT and a ensemble of the four models. Surprisingly, both submissions scored same at 0.727.  I did not use any form of weighting and MaxViT only constituted a quarter of the ensemble submission, but their scores were totally the same on private LB. Even more astonishing was that the single-model score of SeResNext on private LB turned out to be the highest. Its cv was nothing extraordinary, its convergence was not more stable than MaxViT's, and its public leaderboard score was not high, so I had no reason to choose it.\n\n![1](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7285387%2F8fa373c75008e7cf3064b7f4c4089175%2F1.png?generation=1707272539926386&alt=media)\n\nFinally, I would like to extend my gratitude once again to the organizers, Kaggle, and all the participants again!\n\n------------------------------------------\n**Inference(Submission)** Code is published:\n1. [MaxVit512 scored 0.727](https://www.kaggle.com/forcewithme/sennet-final-submission2)\n2. [Ensemble submission scored 0.727](https://www.kaggle.com/code/forcewithme/sennet-top3-final-submission?scriptVersionId=162311084)\n\n**Training** Code is published in the [kaggle dataset](https://www.kaggle.com/datasets/forcewithme/sennettop3-training-code/data).",
      "votes": null
    },
    {
      "id": "2640625",
      "postDate": "02/07/2024 02:50:33",
      "content": "<p>Congratulations on winning another solo gold!🥳</p>",
      "rawMarkdown": "Congratulations on winning another solo gold!🥳",
      "votes": null
    },
    {
      "id": "2640635",
      "postDate": "02/07/2024 03:08:52",
      "content": "<p>Congratulations!  GM deserves it!</p>",
      "rawMarkdown": "Congratulations!  GM deserves it!",
      "votes": null
    },
    {
      "id": "2640646",
      "postDate": "02/07/2024 03:21:33",
      "content": "<p>Wow a lot to learn from this one! I would love to hear more about how you do the pseudo labeling here. Do you mean you are making a mask or assigning an actual label to the image? </p>",
      "rawMarkdown": "Wow a lot to learn from this one! I would love to hear more about how you do the pseudo labeling here. Do you mean you are making a mask or assigning an actual label to the image?",
      "votes": null
    },
    {
      "id": "2640776",
      "postDate": "02/07/2024 04:52:20",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> , thank you for asking this. Take kidney 3 sparse as an example. My pseudo label process contains 3 steps: </p>\n<ol>\n<li>Inference on the sparse kidney, <strong>save the mask in float format</strong>. Don't use any threshold to binarize them now!</li>\n<li>Calculated the positive pixels. It's a simple <code>np.sum</code> operation.</li>\n<li>Search the threshold, get TP, FP and FN, and find the threshold that meets the hosts description (85% for kidney 3 sparse) most. Suppose the model can find all the target perfectly, all the FP on sparse label can be treated as pseudo labels</li>\n</ol>\n<p>The reason that I start with kidney 3 instead of kidney 2, is that half of kidney 3 has dense annotation. According to the host's paper, if  a model is trained and test on the same kidney, the dice is super high. So I think the models can predict very well on kidney 3.</p>\n<p>When making the pseudo label on kidney 2, I can't find a appropriate threshold that perfectly meets 65% sparsity. Hence, I trained difference models on 0.1, 0.15, 0.2 for better diversity in ensembeling, respectively. But according to my final few submissions, I think the pseudo threshold doesn't matter too much.</p>",
      "rawMarkdown": "Hi @cody11null , thank you for asking this. Take kidney 3 sparse as an example. My pseudo label process contains 3 steps: \n\n1.  Inference on the sparse kidney, **save the mask in float format**. Don't use any threshold to binarize them now!\n2. Calculated the positive pixels. It's a simple `np.sum` operation.\n3. Search the threshold, get TP, FP and FN, and find the threshold that meets the hosts description (85% for kidney 3 sparse) most. Suppose the model can find all the target perfectly, all the FP on sparse label can be treated as pseudo labels\n\nThe reason that I start with kidney 3 instead of kidney 2, is that half of kidney 3 has dense annotation. According to the host's paper, if  a model is trained and test on the same kidney, the dice is super high. So I think the models can predict very well on kidney 3.\n\nWhen making the pseudo label on kidney 2, I can't find a appropriate threshold that perfectly meets 65% sparsity. Hence, I trained difference models on 0.1, 0.15, 0.2 for better diversity in ensembeling, respectively. But according to my final few submissions, I think the pseudo threshold doesn't matter too much.",
      "votes": null
    },
    {
      "id": "2640865",
      "postDate": "02/07/2024 05:46:41",
      "content": "<p>I guess that the 2nd, 3rd, 5th, 6th points mentioned in my post is the potential reason for the huge shake. If a participant had missed any one of these aspects and had not prepared adequately, they might have been at risk of experiencing a super huge 'shake down'. </p>\n<p>Conversely, if participants were aware of these issues, or if their models happened to circumvent them, their performance on both the public and private LB would be more consistent, or they might experience a 'shake up'.</p>\n<p>This is a preliminary guess, and perhaps a more comprehensive conclusion will emerge once more participants have disclosed their strategies. </p>",
      "rawMarkdown": "I guess that the 2nd, 3rd, 5th, 6th points mentioned in my post is the potential reason for the huge shake. If a participant had missed any one of these aspects and had not prepared adequately, they might have been at risk of experiencing a super huge 'shake down'. \n\nConversely, if participants were aware of these issues, or if their models happened to circumvent them, their performance on both the public and private LB would be more consistent, or they might experience a 'shake up'.\n\nThis is a preliminary guess, and perhaps a more comprehensive conclusion will emerge once more participants have disclosed their strategies.",
      "votes": null
    },
    {
      "id": "2640890",
      "postDate": "02/07/2024 06:17:48",
      "content": "<p>Wow! Great work! Congrats on your placement, well deserved! </p>",
      "rawMarkdown": "Wow! Great work! Congrats on your placement, well deserved!",
      "votes": null
    },
    {
      "id": "2640898",
      "postDate": "02/07/2024 06:25:03",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> , I added some ablation in the post. If you are interested, any questions are welcome! </p>",
      "rawMarkdown": "Hi @cody11null , I added some ablation in the post. If you are interested, any questions are welcome!",
      "votes": null
    },
    {
      "id": "2640941",
      "postDate": "02/07/2024 07:04:59",
      "content": "<p>Congratulations on the gold medal!<br>\nThank you for sharing your solution!</p>",
      "rawMarkdown": "Congratulations on the gold medal!\nThank you for sharing your solution!",
      "votes": null
    },
    {
      "id": "2641126",
      "postDate": "02/07/2024 09:57:00",
      "content": "<p>Congratulations 🎊 well deserved 👏 <br>\nI just worked during the last week of the comp so didn't get good results, but regarding point 4, i just trained on all images except 1_voi. I used an increasing number of epochs (5,10,25,25) based on sparse--&gt; dense levels. That already makes score improves from 0.45 to 0.55 in private. </p>",
      "rawMarkdown": "Congratulations 🎊 well deserved 👏 \nI just worked during the last week of the comp so didn't get good results, but regarding point 4, i just trained on all images except 1_voi. I used an increasing number of epochs (5,10,25,25) based on sparse--> dense levels. That already makes score improves from 0.45 to 0.55 in private.",
      "votes": null
    },
    {
      "id": "2641379",
      "postDate": "02/07/2024 13:11:40",
      "content": "<p><a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> thanks for sharing!</p>\n<ol>\n<li>There is another posting mentioning pseudo labelling private dataset during inference, any thoguhts on that? First I suppose the thresholding can't be done the same way as you showed (matching the sparsity), I wonder what your public and private score might be if the pseudo labelling is done in inference.</li>\n<li>Why do you think pseudo labelling works? I have seen many times people used pseudo labelling, although intuitively it doesn't make complete sense to me, the pseudo label we generate with the model is produced by the model itself (so perhaps it might be put this way - the model knows these pixels/predictions already), why would using these labels for a re-training (fine-tuning followed by a re-training on all data in your case) help?</li>\n</ol>",
      "rawMarkdown": "forcewithme thanks for sharing!\n1. There is another posting mentioning pseudo labelling private dataset during inference, any thoguhts on that? First I suppose the thresholding can't be done the same way as you showed (matching the sparsity), I wonder what your public and private score might be if the pseudo labelling is done in inference.\n2. Why do you think pseudo labelling works? I have seen many times people used pseudo labelling, although intuitively it doesn't make complete sense to me, the pseudo label we generate with the model is produced by the model itself (so perhaps it might be put this way - the model knows these pixels/predictions already), why would using these labels for a re-training (fine-tuning followed by a re-training on all data in your case) help?",
      "votes": null
    },
    {
      "id": "2641656",
      "postDate": "02/07/2024 15:43:14",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/samshipengs\" target=\"_blank\">@samshipengs</a> . I am not sure if I think it right. But here is my thought:</p>\n<p><strong>It's all about data distribution</strong>. A crucial fact in this competition is that we only have 3 kidneys for train and 2 kidneys in test. The 3 training kidneys only cover a narrow feature space, while the public kidney and private kidney also cover another two narrow feature space. Their feature space may have little overlap, much less than most CV tasks. So the more feature space our model can cover, the more winning chance we have.</p>\n<p>Therefore, </p>\n<ol>\n<li><p>if we retrained for private. Then even the pseudo label is noisy, at least the model <code>has seen</code> these kinds of data. Much better than unseen.</p></li>\n<li><p>When we make pseudo for k3 sparse and k2 sparse, we are making the label less noisy. Then we can better learn the data distribution of kidney 2/3 with less risk to overfitting to the wrong annotations. The data distribution of kidney1+2+3 is much larger than two of them. So the more kidney we trained on, the more overlap the feature space of our model shares with the test set.</p></li>\n</ol>\n<p>I am not a naive English speaker. Please tell me if any expression made you confused.</p>",
      "rawMarkdown": "Hi @samshipengs . I am not sure if I think it right. But here is my thought:\n\n**It's all about data distribution**. A crucial fact in this competition is that we only have 3 kidneys for train and 2 kidneys in test. The 3 training kidneys only cover a narrow feature space, while the public kidney and private kidney also cover another two narrow feature space. Their feature space may have little overlap, much less than most CV tasks. So the more feature space our model can cover, the more winning chance we have.\n\nTherefore, \n\n1. if we retrained for private. Then even the pseudo label is noisy, at least the model `has seen` these kinds of data. Much better than unseen.\n\n2. When we make pseudo for k3 sparse and k2 sparse, we are making the label less noisy. Then we can better learn the data distribution of kidney 2/3 with less risk to overfitting to the wrong annotations. The data distribution of kidney1+2+3 is much larger than two of them. So the more kidney we trained on, the more overlap the feature space of our model shares with the test set.\n\nI am not a naive English speaker. Please tell me if any expression made you confused.",
      "votes": null
    },
    {
      "id": "2641858",
      "postDate": "02/07/2024 17:56:35",
      "content": "<p>thanks, your expression is clear and makes sense.</p>",
      "rawMarkdown": "thanks, your expression is clear and makes sense.",
      "votes": null
    },
    {
      "id": "2642173",
      "postDate": "02/08/2024 01:47:22",
      "content": "<p>Great approach!  proud of you! <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> </p>",
      "rawMarkdown": "Great approach!  proud of you! @forcewithme",
      "votes": null
    },
    {
      "id": "2642737",
      "postDate": "02/08/2024 11:25:35",
      "content": "<p>Congratulations on achieving 3rd place in this competition. Thanks for sharing details of your solution. </p>",
      "rawMarkdown": "Congratulations on achieving 3rd place in this competition. Thanks for sharing details of your solution.",
      "votes": null
    },
    {
      "id": "2644370",
      "postDate": "02/09/2024 12:27:41",
      "content": "<p>The UNet-SeResNext101 in my pipeline is somehow fitting(maybe overfitting) the private LB. Probing with different resolution and threshold several times, it can achieve 0.78 on private LB. </p>\n<table>\n<thead>\n<tr>\n<th>Resolution</th>\n<th>Threshold</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>512</td>\n<td>0.2</td>\n<td>0.820</td>\n<td>0.753</td>\n</tr>\n<tr>\n<td>512</td>\n<td>0.15</td>\n<td>0.816</td>\n<td>0.757</td>\n</tr>\n<tr>\n<td>512</td>\n<td>0.1</td>\n<td>0.811</td>\n<td>0.766</td>\n</tr>\n<tr>\n<td>384+512</td>\n<td>0.1</td>\n<td>0.807</td>\n<td>0.778</td>\n</tr>\n<tr>\n<td>384+512</td>\n<td>0.12</td>\n<td>0.811</td>\n<td>0.775</td>\n</tr>\n<tr>\n<td>384+512</td>\n<td>0.08</td>\n<td>0.801</td>\n<td>0.780</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "The UNet-SeResNext101 in my pipeline is somehow fitting(maybe overfitting) the private LB. Probing with different resolution and threshold several times, it can achieve 0.78 on private LB. \n\n| Resolution | Threshold | public | private |\n| --- | --- |--- | --- |\n| 512 | 0.2 | 0.820 | 0.753 |\n| 512 | 0.15 |0.816 | 0.757 |\n| 512 | 0.1 |0.811 | 0.766 |\n| 384+512 | 0.1 |0.807 | 0.778 |\n| 384+512 | 0.12 |0.811 | 0.775 |\n| 384+512 | 0.08 |0.801 | 0.780 |",
      "votes": null
    },
    {
      "id": "2645443",
      "postDate": "02/10/2024 08:31:49",
      "content": "<p>congratulations</p>",
      "rawMarkdown": "congratulations",
      "votes": null
    },
    {
      "id": "2645583",
      "postDate": "02/10/2024 10:39:36",
      "content": "<p>Wow! Great work! Great approach!</p>",
      "rawMarkdown": "Wow! Great work! Great approach!",
      "votes": null
    },
    {
      "id": "2645871",
      "postDate": "02/10/2024 14:16:10",
      "content": "<p>Great approach! proud of you! <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a></p>\n<p>Congratulations on winning another solo gold!🥳</p>",
      "rawMarkdown": "Great approach! proud of you! @forcewithme\n\nCongratulations on winning another solo gold!🥳",
      "votes": null
    },
    {
      "id": "2646935",
      "postDate": "02/11/2024 09:44:47",
      "content": "<p>great accuracy..congo</p>",
      "rawMarkdown": "great accuracy..congo",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2640625,
      "author_name": "kevin1742064161",
      "author_url": "",
      "post_date": "02/07/2024 02:50:33",
      "content": "<p>Congratulations on winning another solo gold!🥳</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2640635,
      "author_name": "wisley1024",
      "author_url": "",
      "post_date": "02/07/2024 03:08:52",
      "content": "<p>Congratulations!  GM deserves it!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2640646,
      "author_name": "cody11null",
      "author_url": "",
      "post_date": "02/07/2024 03:21:33",
      "content": "<p>Wow a lot to learn from this one! I would love to hear more about how you do the pseudo labeling here. Do you mean you are making a mask or assigning an actual label to the image? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2640776,
          "author_name": "forcewithme",
          "author_url": "",
          "post_date": "02/07/2024 04:52:20",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> , thank you for asking this. Take kidney 3 sparse as an example. My pseudo label process contains 3 steps: </p>\n<ol>\n<li>Inference on the sparse kidney, <strong>save the mask in float format</strong>. Don't use any threshold to binarize them now!</li>\n<li>Calculated the positive pixels. It's a simple <code>np.sum</code> operation.</li>\n<li>Search the threshold, get TP, FP and FN, and find the threshold that meets the hosts description (85% for kidney 3 sparse) most. Suppose the model can find all the target perfectly, all the FP on sparse label can be treated as pseudo labels</li>\n</ol>\n<p>The reason that I start with kidney 3 instead of kidney 2, is that half of kidney 3 has dense annotation. According to the host's paper, if  a model is trained and test on the same kidney, the dice is super high. So I think the models can predict very well on kidney 3.</p>\n<p>When making the pseudo label on kidney 2, I can't find a appropriate threshold that perfectly meets 65% sparsity. Hence, I trained difference models on 0.1, 0.15, 0.2 for better diversity in ensembeling, respectively. But according to my final few submissions, I think the pseudo threshold doesn't matter too much.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2640890,
              "author_name": "cody11null",
              "author_url": "",
              "post_date": "02/07/2024 06:17:48",
              "content": "<p>Wow! Great work! Congrats on your placement, well deserved! </p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2640898,
          "author_name": "forcewithme",
          "author_url": "",
          "post_date": "02/07/2024 06:25:03",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> , I added some ablation in the post. If you are interested, any questions are welcome! </p>",
          "votes": null,
          "replies": [
            {
              "id": 2641379,
              "author_name": "samshipengs",
              "author_url": "",
              "post_date": "02/07/2024 13:11:40",
              "content": "<p><a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> thanks for sharing!</p>\n<ol>\n<li>There is another posting mentioning pseudo labelling private dataset during inference, any thoguhts on that? First I suppose the thresholding can't be done the same way as you showed (matching the sparsity), I wonder what your public and private score might be if the pseudo labelling is done in inference.</li>\n<li>Why do you think pseudo labelling works? I have seen many times people used pseudo labelling, although intuitively it doesn't make complete sense to me, the pseudo label we generate with the model is produced by the model itself (so perhaps it might be put this way - the model knows these pixels/predictions already), why would using these labels for a re-training (fine-tuning followed by a re-training on all data in your case) help?</li>\n</ol>",
              "votes": null,
              "replies": [
                {
                  "id": 2641656,
                  "author_name": "forcewithme",
                  "author_url": "",
                  "post_date": "02/07/2024 15:43:14",
                  "content": "<p>Hi <a href=\"https://www.kaggle.com/samshipengs\" target=\"_blank\">@samshipengs</a> . I am not sure if I think it right. But here is my thought:</p>\n<p><strong>It's all about data distribution</strong>. A crucial fact in this competition is that we only have 3 kidneys for train and 2 kidneys in test. The 3 training kidneys only cover a narrow feature space, while the public kidney and private kidney also cover another two narrow feature space. Their feature space may have little overlap, much less than most CV tasks. So the more feature space our model can cover, the more winning chance we have.</p>\n<p>Therefore, </p>\n<ol>\n<li><p>if we retrained for private. Then even the pseudo label is noisy, at least the model <code>has seen</code> these kinds of data. Much better than unseen.</p></li>\n<li><p>When we make pseudo for k3 sparse and k2 sparse, we are making the label less noisy. Then we can better learn the data distribution of kidney 2/3 with less risk to overfitting to the wrong annotations. The data distribution of kidney1+2+3 is much larger than two of them. So the more kidney we trained on, the more overlap the feature space of our model shares with the test set.</p></li>\n</ol>\n<p>I am not a naive English speaker. Please tell me if any expression made you confused.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2641858,
                      "author_name": "samshipengs",
                      "author_url": "",
                      "post_date": "02/07/2024 17:56:35",
                      "content": "<p>thanks, your expression is clear and makes sense.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2640865,
      "author_name": "forcewithme",
      "author_url": "",
      "post_date": "02/07/2024 05:46:41",
      "content": "<p>I guess that the 2nd, 3rd, 5th, 6th points mentioned in my post is the potential reason for the huge shake. If a participant had missed any one of these aspects and had not prepared adequately, they might have been at risk of experiencing a super huge 'shake down'. </p>\n<p>Conversely, if participants were aware of these issues, or if their models happened to circumvent them, their performance on both the public and private LB would be more consistent, or they might experience a 'shake up'.</p>\n<p>This is a preliminary guess, and perhaps a more comprehensive conclusion will emerge once more participants have disclosed their strategies. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2640941,
      "author_name": "raytency",
      "author_url": "",
      "post_date": "02/07/2024 07:04:59",
      "content": "<p>Congratulations on the gold medal!<br>\nThank you for sharing your solution!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2641126,
      "author_name": "mohammad2012191",
      "author_url": "",
      "post_date": "02/07/2024 09:57:00",
      "content": "<p>Congratulations 🎊 well deserved 👏 <br>\nI just worked during the last week of the comp so didn't get good results, but regarding point 4, i just trained on all images except 1_voi. I used an increasing number of epochs (5,10,25,25) based on sparse--&gt; dense levels. That already makes score improves from 0.45 to 0.55 in private. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2642173,
      "author_name": "seungwanhong",
      "author_url": "",
      "post_date": "02/08/2024 01:47:22",
      "content": "<p>Great approach!  proud of you! <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2642737,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "02/08/2024 11:25:35",
      "content": "<p>Congratulations on achieving 3rd place in this competition. Thanks for sharing details of your solution. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2644370,
      "author_name": "forcewithme",
      "author_url": "",
      "post_date": "02/09/2024 12:27:41",
      "content": "<p>The UNet-SeResNext101 in my pipeline is somehow fitting(maybe overfitting) the private LB. Probing with different resolution and threshold several times, it can achieve 0.78 on private LB. </p>\n<table>\n<thead>\n<tr>\n<th>Resolution</th>\n<th>Threshold</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>512</td>\n<td>0.2</td>\n<td>0.820</td>\n<td>0.753</td>\n</tr>\n<tr>\n<td>512</td>\n<td>0.15</td>\n<td>0.816</td>\n<td>0.757</td>\n</tr>\n<tr>\n<td>512</td>\n<td>0.1</td>\n<td>0.811</td>\n<td>0.766</td>\n</tr>\n<tr>\n<td>384+512</td>\n<td>0.1</td>\n<td>0.807</td>\n<td>0.778</td>\n</tr>\n<tr>\n<td>384+512</td>\n<td>0.12</td>\n<td>0.811</td>\n<td>0.775</td>\n</tr>\n<tr>\n<td>384+512</td>\n<td>0.08</td>\n<td>0.801</td>\n<td>0.780</td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2645443,
      "author_name": "taniya819",
      "author_url": "",
      "post_date": "02/10/2024 08:31:49",
      "content": "<p>congratulations</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2645583,
      "author_name": "ericka42",
      "author_url": "",
      "post_date": "02/10/2024 10:39:36",
      "content": "<p>Wow! Great work! Great approach!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2645871,
      "author_name": "tanishqdublish",
      "author_url": "",
      "post_date": "02/10/2024 14:16:10",
      "content": "<p>Great approach! proud of you! <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a></p>\n<p>Congratulations on winning another solo gold!🥳</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2646935,
      "author_name": "chetalipushkarna",
      "author_url": "",
      "post_date": "02/11/2024 09:44:47",
      "content": "<p>great accuracy..congo</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2640599": "First and foremost, I would like to extend my gratitude to the organizers and the official Kaggle team for orchestrating such an outstanding competition. I joined the contest at a very late stage. Despite having some experience with segmentation competitions, I must express my appreciation to @hengck23 , @yoyobar , and @junkoda (implementation of metric), as well as the other community participants for their open-source contributions and discussions, which allowed me to quickly get up to speed with this contest.\n\nMy approach was strikingly straightforward, relying solely on **2D models** and only utilizing **smp** (segmentation models pytorch) and **timm** (pytorch image models) in the whole training and inference pipeline. \n\n## Global key points\n1.   **Refining labels from sparse to dense.**\n2.  **Emulating the magnification factor of the test set.**\n3.  **Maintaining an appropriate resolution.**\n\n## 1. From Sparse to Dense\nGiven that half of the training set has dense labels (kidney 1, kidney 3 dense), and the other half was sparse, utilizing dense labels to refine sparse ones was a crucial step. The overall process entailed:\n\n1. Training UNet(maxViT512) and UNet(EfficientNetv2s) using kidney 1 and kidney 3 dense.\n2. Generating supplemental labels for kidney 3 sparse using the trained UNet maxViT512 and UNet EfficientNetv2s models.\n3. Resuming the training of UNet maxViT512 and UNet EfficientNetv2s for a few epochs with kidney 1, kidney 3 (dense, sparse plus supplemental labels).\n4. Repeating the step2 on kidney 2.\n5. Training three UNet models (with EfficientNetv2s, SeResNext101, MaxViT512) and one UNet++ using all real labels from all kidney plus pseudo labels.\n\nNote: As the organizers disclosed the proportion of annotations within kidney 3 and kidney 2, I endeavored to select thresholds based on pixel quantity as close as possible to the official proportion when choosing threshold values for pseudo label.\n\n## 2. Emulating the Magnification of the Private Test Set\nA pivotal reason for my decision to participate in this competition was the disclosure of the magnification factors for the training and test sets by the hosts. The training set had a magnification of 50um/voxel, the public test set was the same at 50um/voxel, while the private test set was at 63um/voxel. A larger magnification factor implies a lower resolution. For instance, a 600um object would occupy 12 pixels in both the training and public sets, but only 10 pixels in the validation set. Hence, during training, **I set the scaling center to 0.8**, rather than 1, with a scaling range of 0.55 to 1.05, to simulate the private test set.\n\n```\nA.ShiftScaleRotate(shift_limit=0.3,\n                    scale_limit=(-0.45, 0.05),  \n                    rotate_limit=45,\n                    # value=0,\n                    border_mode=4,\n                    p=0.95),\n```\n\n## 3. Maintaining an Appropriate Resolution\nIn this competition, training and inference along the x-axis, y-axis, and z-axis separately was a very important trick. However, this introduced a significant risk. The entire test set contained 1500 slices, with the public test set accounting for 67% and the private test set for 33%. This means that the private test set comprised only about **500 slices**. Inferring along the z-axis with a higher resolution (e.g., 1024) was feasible. But if inferring along the y-axis or x-axis, it would mean that one of the edges would only be 500 pixels long. At that point, if the model and code were configured for a larger resolution (say 1024), there would be a substantial risk of a huge shake down.\n\nMy models primarily operated at a resolution of 512, with one model switching to higher resolution weights for larger resolution slices when the slice have appropriate resolution.\n\n| Model | Backbone | Resolution | public | private|\n| --- | --- | --- | --- | --- |\n| UNet | MaxViT-Large 512 | 512 | 0.846 | 0.727(submission1) |\n| UNet | SeResNext | 512 | 0.819 | 0.753 |\n| UNet | Efficiennet_v2_s | 448, 832 | 0.799 | 0.703 | \n| UNet++ | Efficiennet_v2_l | 512 | 0.817 | 0.692 |\n| ensemble | - | - | 0.846 | 0.727(submission2) | \n\n\n## 4. Train on all data if convergence is Stable\nDuring the early stages of the competition, whether validating on kidney 2 or kidney 3, I observed that if I trained for 20 epochs, after the initial few epochs, the dice coefficient (not surface dice) variation on the validation set was very minimal, with the MaxVit512 large model exhibiting the least fluctuation. Considering that we only had three kidneys, I decided to train on all kidneys directly after completing the pseudo labeling process, given the stability in convergence.\n\n## 5. Minimizing the Impact of Threshold Values\nI am grateful for the method provided by @junkoda for calculating metrics. My most stable single model was able to maintain very minor fluctuations in the surface dice score (less than 1) within a threshold range of 0.2. After model fusion, the stable threshold range could be potentially in 0.3~ 0.4. A stable threshold is extremely crucial in segmentation competitions. In this competition, as my final model lacked a validation set, I had to utilize thresholds searched with earlier trained models that included a validation set and apply them to the final version of the model. Fortunately, the models trained on the full dataset appeared to possess threshold values very close to those from the earlier models trained with k1+k2 (sparse), and validate on k3. At the same time, the fluctuation of threshold values across kidney 3 dense, public, and private was very small.\n\n## 6. Heavy augmentation on intensity.\nAs mentioned by @hengck23 , difference kidneys has large variance on intensity. So I used a heavy intensity augmentation.\n```\nA.RandomBrightnessContrast(p=1.0),\nA.RandomGamma(p=0.8),\n```\n## 7. Quick ablation\n| Model | Backbone | points mentioned above | public | private|\n| --- | --- | --- | --- | --- |\n| UNet | MaxViT-Large 512 | 3, 5, 6 | 0.818 | 0.586 |\n| UNet | MaxViT-Large 512 | 3, 4, 5, 6 | 0.857 | 0.633 |\n| UNet | MaxViT-Large 512 | 2, 3, 4, 5, 6 | 0.849 | 0.652 |\n| UNet | MaxViT-Large 512 | 1, 2, 3, 4, 5, 6 | 0.846 | 0.727 |\n\n## 8. Final Submission\nMy final submissions were a single model of MaxViT and a ensemble of the four models. Surprisingly, both submissions scored same at 0.727.  I did not use any form of weighting and MaxViT only constituted a quarter of the ensemble submission, but their scores were totally the same on private LB. Even more astonishing was that the single-model score of SeResNext on private LB turned out to be the highest. Its cv was nothing extraordinary, its convergence was not more stable than MaxViT's, and its public leaderboard score was not high, so I had no reason to choose it.\n\n![1](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7285387%2F8fa373c75008e7cf3064b7f4c4089175%2F1.png?generation=1707272539926386&alt=media)\n\nFinally, I would like to extend my gratitude once again to the organizers, Kaggle, and all the participants again!\n\n------------------------------------------\n**Inference(Submission)** Code is published:\n1. [MaxVit512 scored 0.727](https://www.kaggle.com/forcewithme/sennet-final-submission2)\n2. [Ensemble submission scored 0.727](https://www.kaggle.com/code/forcewithme/sennet-top3-final-submission?scriptVersionId=162311084)\n\n**Training** Code is published in the [kaggle dataset](https://www.kaggle.com/datasets/forcewithme/sennettop3-training-code/data).",
    "2640625": "Congratulations on winning another solo gold!🥳",
    "2640635": "Congratulations!  GM deserves it!",
    "2640646": "Wow a lot to learn from this one! I would love to hear more about how you do the pseudo labeling here. Do you mean you are making a mask or assigning an actual label to the image?",
    "2640776": "Hi @cody11null , thank you for asking this. Take kidney 3 sparse as an example. My pseudo label process contains 3 steps: \n\n1.  Inference on the sparse kidney, **save the mask in float format**. Don't use any threshold to binarize them now!\n2. Calculated the positive pixels. It's a simple `np.sum` operation.\n3. Search the threshold, get TP, FP and FN, and find the threshold that meets the hosts description (85% for kidney 3 sparse) most. Suppose the model can find all the target perfectly, all the FP on sparse label can be treated as pseudo labels\n\nThe reason that I start with kidney 3 instead of kidney 2, is that half of kidney 3 has dense annotation. According to the host's paper, if  a model is trained and test on the same kidney, the dice is super high. So I think the models can predict very well on kidney 3.\n\nWhen making the pseudo label on kidney 2, I can't find a appropriate threshold that perfectly meets 65% sparsity. Hence, I trained difference models on 0.1, 0.15, 0.2 for better diversity in ensembeling, respectively. But according to my final few submissions, I think the pseudo threshold doesn't matter too much.",
    "2640865": "I guess that the 2nd, 3rd, 5th, 6th points mentioned in my post is the potential reason for the huge shake. If a participant had missed any one of these aspects and had not prepared adequately, they might have been at risk of experiencing a super huge 'shake down'. \n\nConversely, if participants were aware of these issues, or if their models happened to circumvent them, their performance on both the public and private LB would be more consistent, or they might experience a 'shake up'.\n\nThis is a preliminary guess, and perhaps a more comprehensive conclusion will emerge once more participants have disclosed their strategies.",
    "2640890": "Wow! Great work! Congrats on your placement, well deserved!",
    "2640898": "Hi @cody11null , I added some ablation in the post. If you are interested, any questions are welcome!",
    "2640941": "Congratulations on the gold medal!\nThank you for sharing your solution!",
    "2641126": "Congratulations 🎊 well deserved 👏 \nI just worked during the last week of the comp so didn't get good results, but regarding point 4, i just trained on all images except 1_voi. I used an increasing number of epochs (5,10,25,25) based on sparse--> dense levels. That already makes score improves from 0.45 to 0.55 in private.",
    "2641379": "forcewithme thanks for sharing!\n1. There is another posting mentioning pseudo labelling private dataset during inference, any thoguhts on that? First I suppose the thresholding can't be done the same way as you showed (matching the sparsity), I wonder what your public and private score might be if the pseudo labelling is done in inference.\n2. Why do you think pseudo labelling works? I have seen many times people used pseudo labelling, although intuitively it doesn't make complete sense to me, the pseudo label we generate with the model is produced by the model itself (so perhaps it might be put this way - the model knows these pixels/predictions already), why would using these labels for a re-training (fine-tuning followed by a re-training on all data in your case) help?",
    "2641656": "Hi @samshipengs . I am not sure if I think it right. But here is my thought:\n\n**It's all about data distribution**. A crucial fact in this competition is that we only have 3 kidneys for train and 2 kidneys in test. The 3 training kidneys only cover a narrow feature space, while the public kidney and private kidney also cover another two narrow feature space. Their feature space may have little overlap, much less than most CV tasks. So the more feature space our model can cover, the more winning chance we have.\n\nTherefore, \n\n1. if we retrained for private. Then even the pseudo label is noisy, at least the model `has seen` these kinds of data. Much better than unseen.\n\n2. When we make pseudo for k3 sparse and k2 sparse, we are making the label less noisy. Then we can better learn the data distribution of kidney 2/3 with less risk to overfitting to the wrong annotations. The data distribution of kidney1+2+3 is much larger than two of them. So the more kidney we trained on, the more overlap the feature space of our model shares with the test set.\n\nI am not a naive English speaker. Please tell me if any expression made you confused.",
    "2641858": "thanks, your expression is clear and makes sense.",
    "2642173": "Great approach!  proud of you! @forcewithme",
    "2642737": "Congratulations on achieving 3rd place in this competition. Thanks for sharing details of your solution.",
    "2644370": "The UNet-SeResNext101 in my pipeline is somehow fitting(maybe overfitting) the private LB. Probing with different resolution and threshold several times, it can achieve 0.78 on private LB. \n\n| Resolution | Threshold | public | private |\n| --- | --- |--- | --- |\n| 512 | 0.2 | 0.820 | 0.753 |\n| 512 | 0.15 |0.816 | 0.757 |\n| 512 | 0.1 |0.811 | 0.766 |\n| 384+512 | 0.1 |0.807 | 0.778 |\n| 384+512 | 0.12 |0.811 | 0.775 |\n| 384+512 | 0.08 |0.801 | 0.780 |",
    "2645443": "congratulations",
    "2645583": "Wow! Great work! Great approach!",
    "2645871": "Great approach! proud of you! @forcewithme\n\nCongratulations on winning another solo gold!🥳",
    "2646935": "great accuracy..congo"
  },
  "source": "meta"
}