{
  "id": 561510,
  "title": "1st place solution [segmentation with partly U-NET and ensembling part]",
  "url": "/competitions/czii-cryo-et-object-identification/writeups/daddies-1st-place-solution-segmentation-with-partl",
  "author_name": "",
  "post_date": "2025-02-24T14:00:22.720Z",
  "votes": 103,
  "comment_count": 35,
  "views": 0,
  "content": "<p>Thanks to kaggle and everyone involved for hosting this exciting competition. It was a great learning experience and it was very interesting to see how much of our computer vision experience could also be applied to 3D imaging. Thanks to <a href=\"https://www.kaggle.com/bloodaxe\" target=\"_blank\">@bloodaxe</a> for this great team experience. </p>\n<h2>TLDR</h2>\n<p>The solution in an ensemble of segmentation (3D Unets with ResNet &amp; B3 encoders) and object detection models (SegResNet and DynUnet backbones) from <a href=\"https://github.com/Project-MONAI/MONAI\" target=\"_blank\">MONAI</a>. We also used MONAI for augmentations, and exported models via jit or TensorRT, which gave 200% speedup increase and enabled us to have a slightly larger ensemble. We did not use any external or simulated data!</p>\n<p>This post covers the segmentation based approach and ensembling. For object detection part see <a href=\"https://www.kaggle.com/bloodaxe\" target=\"_blank\">@bloodaxe</a> writeup: <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561440\" target=\"_blank\">1st place solution [Object Detection Part]</a></p>\n<h2>Cross validation</h2>\n<p>For segmentation approach 7 folds were used, simply by splitting by experiment. Using mean of f4-score of all 7 folds had some good correlation with LB. During model training I optimized individual class thresholds at the end of each epoch by simple grid-search on the validation experiment. After all 7 folds were trained we could re-calibrate thresholds, by taking OOF predictions. And fitting the threshold for one fold on the predictions of the other 6. Then we average the resulting f4-curves and take the best threshold.</p>\n<h2>Data preprocessing/ augmentations</h2>\n<p>3D images were normalized by standard normalization, i.e. for each 630x630x184 image, we substract mean and devide by standard deviation before splitting the images into patches.<br>\nSince models are trained from scratch, augmentations were essential to prevent overfitting. <br>\nWe used RandomCrop, Flip on each axis, and rotation, which all are available with MONAI. Additionally I used my own implementation of MixUp which was highly effective to train longer and prevent overfitting.</p>\n<h2>Model</h2>\n<p>Modelling was quite a ride in this competition. I started with simple UNET with having 3D gaussian balls as segmentation target. For diversity I also tried <a href=\"https://github.com/Project-MONAI/tutorials/tree/main/detection\" target=\"_blank\">object detection example from MONAI</a> and realized its working very well out of the box. But when analyzing the different output feature maps and trying to isolate the perfomance gain over gaussian heatmap based segmentation I realized where the advantage was and adjusted my segmentation model accordingly. I learned that </p>\n<ul>\n<li>the penultimate feature map has a higher accuracy than the last one, which is surprising at first</li>\n<li>gains from box regression are negligable, as particles from the same type have mostly same size anyways.</li>\n</ul>\n<p>Hence its sufficient to have a pixel-wise loss on the penultimate feature map output. I.e. use a <strong>partly</strong> UNET. The gaussian heatmap is not needed when suppressing background with a low class weight, and just use single pixels as targets. Using the same approach on lower level outputs (= deep supervision) is possible, but does not provide much gain. Input for the segmentation models are 96x96x96 image patches and the loss is calculated on the 48x48x48 output. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F961b07916f0b97850233aae461704c69%2FScreenshot%202025-02-06%20at%2014.05.42.png?generation=1738849991855202&amp;alt=media\" alt=\"\"></p>\n<p>In general, we observed that relatively small models work really and the design of loss is the most important aspect. We used MONAIs <a href=\"https://docs.monai.io/en/stable/networks.html#flexibleunet\" target=\"_blank\">FlexibleUnet</a> with backbones resnet34 and efficientnet-b3. Six checkpoints of this architecture would finish under 2h and score 7th place on LB.</p>\n<h2>Training procedure</h2>\n<p>We used 7 classes (incl background) and weighted CrossEntropy as loss. Notably keeping beta-amylase as a class although it is not scored is quite helpful as model learns to differentiate beta-galactosidase from it. To account for low number of positive pixels, positive pixels are weighted by 256 and background has weight 1. Models were trained with a cosine learning rate schedule with peak LR of 0.001, mixed precision and an effective batch size of 32 samples. Training is based on Random crops and for validation the single experiment image is divided into patches and stored in RAM. </p>\n<h2>Ensembling</h2>\n<p>Ensembling was very challenging, as our two approaches are quite diverse. While in theory predictions from the segmentation models can be ensembled with feature map outputs of the object detection model before runnning the object detection postprocessing, in practice those feature maps have a very different distribution due to difference architectures and loss functions. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F9d47503c1fd18872ef42b466f72c6732%2FScreenshot%202025-02-06%20at%2014.25.00.png?generation=1738850078673576&amp;alt=media\" alt=\"\"></p>\n<p>We were eager to find an elegant way to fix this scaling issue as we saw the potential of a possible ensemble. The scaling of our best submission works as following. To combine predictions A with predictions B, for each class sort all pixel values for A and B and replace values of B with the corresponding values of A of same rank. In code: </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F05670c1a2a658de03ad73522d1efe22e%2FScreenshot%202025-02-06%20at%2014.58.11.png?generation=1738850308432936&amp;alt=media\" alt=\"\"></p>\n<p>This results in both predictions having the same distribution, and hence we could simple blend the feature maps before performing object de<br>\ntection task.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fc5ccff4d072d09afe6b6c03273b65d40%2FScreenshot%202025-02-06%20at%2014.38.13.png?generation=1738850131187278&amp;alt=media\" alt=\"\"></p>\n<h2>What did not work</h2>\n<ul>\n<li>Using supplemental data (external or simulated)</li>\n<li>Other augs</li>\n<li>Other losses (Tversky, Dice)</li>\n</ul>\n<p>Thanks for reading. </p>\n<p>Edits:<br>\ntraining code: <a href=\"https://github.com/ChristofHenkel/kaggle-cryoet-1st-place-segmentation\" target=\"_blank\">https://github.com/ChristofHenkel/kaggle-cryoet-1st-place-segmentation</a><br>\ninference kernel: <a href=\"https://www.kaggle.com/code/christofhenkel/cryo-et-1st-place-solution?scriptVersionId=223259615\" target=\"_blank\">https://www.kaggle.com/code/christofhenkel/cryo-et-1st-place-solution?scriptVersionId=223259615</a></p>",
  "messages": [
    {
      "id": "3116985",
      "postDate": "02/06/2025 13:59:32",
      "content": "<p>Thanks to kaggle and everyone involved for hosting this exciting competition. It was a great learning experience and it was very interesting to see how much of our computer vision experience could also be applied to 3D imaging. Thanks to <a href=\"https://www.kaggle.com/bloodaxe\" target=\"_blank\">@bloodaxe</a> for this great team experience. </p>\n<h2>TLDR</h2>\n<p>The solution in an ensemble of segmentation (3D Unets with ResNet &amp; B3 encoders) and object detection models (SegResNet and DynUnet backbones) from <a href=\"https://github.com/Project-MONAI/MONAI\" target=\"_blank\">MONAI</a>. We also used MONAI for augmentations, and exported models via jit or TensorRT, which gave 200% speedup increase and enabled us to have a slightly larger ensemble. We did not use any external or simulated data!</p>\n<p>This post covers the segmentation based approach and ensembling. For object detection part see <a href=\"https://www.kaggle.com/bloodaxe\" target=\"_blank\">@bloodaxe</a> writeup: <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561440\" target=\"_blank\">1st place solution [Object Detection Part]</a></p>\n<h2>Cross validation</h2>\n<p>For segmentation approach 7 folds were used, simply by splitting by experiment. Using mean of f4-score of all 7 folds had some good correlation with LB. During model training I optimized individual class thresholds at the end of each epoch by simple grid-search on the validation experiment. After all 7 folds were trained we could re-calibrate thresholds, by taking OOF predictions. And fitting the threshold for one fold on the predictions of the other 6. Then we average the resulting f4-curves and take the best threshold.</p>\n<h2>Data preprocessing/ augmentations</h2>\n<p>3D images were normalized by standard normalization, i.e. for each 630x630x184 image, we substract mean and devide by standard deviation before splitting the images into patches.<br>\nSince models are trained from scratch, augmentations were essential to prevent overfitting. <br>\nWe used RandomCrop, Flip on each axis, and rotation, which all are available with MONAI. Additionally I used my own implementation of MixUp which was highly effective to train longer and prevent overfitting.</p>\n<h2>Model</h2>\n<p>Modelling was quite a ride in this competition. I started with simple UNET with having 3D gaussian balls as segmentation target. For diversity I also tried <a href=\"https://github.com/Project-MONAI/tutorials/tree/main/detection\" target=\"_blank\">object detection example from MONAI</a> and realized its working very well out of the box. But when analyzing the different output feature maps and trying to isolate the perfomance gain over gaussian heatmap based segmentation I realized where the advantage was and adjusted my segmentation model accordingly. I learned that </p>\n<ul>\n<li>the penultimate feature map has a higher accuracy than the last one, which is surprising at first</li>\n<li>gains from box regression are negligable, as particles from the same type have mostly same size anyways.</li>\n</ul>\n<p>Hence its sufficient to have a pixel-wise loss on the penultimate feature map output. I.e. use a <strong>partly</strong> UNET. The gaussian heatmap is not needed when suppressing background with a low class weight, and just use single pixels as targets. Using the same approach on lower level outputs (= deep supervision) is possible, but does not provide much gain. Input for the segmentation models are 96x96x96 image patches and the loss is calculated on the 48x48x48 output. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F961b07916f0b97850233aae461704c69%2FScreenshot%202025-02-06%20at%2014.05.42.png?generation=1738849991855202&amp;alt=media\" alt=\"\"></p>\n<p>In general, we observed that relatively small models work really and the design of loss is the most important aspect. We used MONAIs <a href=\"https://docs.monai.io/en/stable/networks.html#flexibleunet\" target=\"_blank\">FlexibleUnet</a> with backbones resnet34 and efficientnet-b3. Six checkpoints of this architecture would finish under 2h and score 7th place on LB.</p>\n<h2>Training procedure</h2>\n<p>We used 7 classes (incl background) and weighted CrossEntropy as loss. Notably keeping beta-amylase as a class although it is not scored is quite helpful as model learns to differentiate beta-galactosidase from it. To account for low number of positive pixels, positive pixels are weighted by 256 and background has weight 1. Models were trained with a cosine learning rate schedule with peak LR of 0.001, mixed precision and an effective batch size of 32 samples. Training is based on Random crops and for validation the single experiment image is divided into patches and stored in RAM. </p>\n<h2>Ensembling</h2>\n<p>Ensembling was very challenging, as our two approaches are quite diverse. While in theory predictions from the segmentation models can be ensembled with feature map outputs of the object detection model before runnning the object detection postprocessing, in practice those feature maps have a very different distribution due to difference architectures and loss functions. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F9d47503c1fd18872ef42b466f72c6732%2FScreenshot%202025-02-06%20at%2014.25.00.png?generation=1738850078673576&amp;alt=media\" alt=\"\"></p>\n<p>We were eager to find an elegant way to fix this scaling issue as we saw the potential of a possible ensemble. The scaling of our best submission works as following. To combine predictions A with predictions B, for each class sort all pixel values for A and B and replace values of B with the corresponding values of A of same rank. In code: </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F05670c1a2a658de03ad73522d1efe22e%2FScreenshot%202025-02-06%20at%2014.58.11.png?generation=1738850308432936&amp;alt=media\" alt=\"\"></p>\n<p>This results in both predictions having the same distribution, and hence we could simple blend the feature maps before performing object de<br>\ntection task.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fc5ccff4d072d09afe6b6c03273b65d40%2FScreenshot%202025-02-06%20at%2014.38.13.png?generation=1738850131187278&amp;alt=media\" alt=\"\"></p>\n<h2>What did not work</h2>\n<ul>\n<li>Using supplemental data (external or simulated)</li>\n<li>Other augs</li>\n<li>Other losses (Tversky, Dice)</li>\n</ul>\n<p>Thanks for reading. </p>\n<p>Edits:<br>\ntraining code: <a href=\"https://github.com/ChristofHenkel/kaggle-cryoet-1st-place-segmentation\" target=\"_blank\">https://github.com/ChristofHenkel/kaggle-cryoet-1st-place-segmentation</a><br>\ninference kernel: <a href=\"https://www.kaggle.com/code/christofhenkel/cryo-et-1st-place-solution?scriptVersionId=223259615\" target=\"_blank\">https://www.kaggle.com/code/christofhenkel/cryo-et-1st-place-solution?scriptVersionId=223259615</a></p>",
      "rawMarkdown": "Thanks to kaggle and everyone involved for hosting this exciting competition. It was a great learning experience and it was very interesting to see how much of our computer vision experience could also be applied to 3D imaging. Thanks to @bloodaxe for this great team experience. \n\n## TLDR\n\nThe solution in an ensemble of segmentation (3D Unets with ResNet & B3 encoders) and object detection models (SegResNet and DynUnet backbones) from [MONAI](https://github.com/Project-MONAI/MONAI). We also used MONAI for augmentations, and exported models via jit or TensorRT, which gave 200% speedup increase and enabled us to have a slightly larger ensemble. We did not use any external or simulated data!\n\nThis post covers the segmentation based approach and ensembling. For object detection part see @bloodaxe writeup: [1st place solution [Object Detection Part]](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561440)\n\n## Cross validation\n\nFor segmentation approach 7 folds were used, simply by splitting by experiment. Using mean of f4-score of all 7 folds had some good correlation with LB. During model training I optimized individual class thresholds at the end of each epoch by simple grid-search on the validation experiment. After all 7 folds were trained we could re-calibrate thresholds, by taking OOF predictions. And fitting the threshold for one fold on the predictions of the other 6. Then we average the resulting f4-curves and take the best threshold.\n\n## Data preprocessing/ augmentations\n\n3D images were normalized by standard normalization, i.e. for each 630x630x184 image, we substract mean and devide by standard deviation before splitting the images into patches.\nSince models are trained from scratch, augmentations were essential to prevent overfitting. \nWe used RandomCrop, Flip on each axis, and rotation, which all are available with MONAI. Additionally I used my own implementation of MixUp which was highly effective to train longer and prevent overfitting.\n\n## Model\n\nModelling was quite a ride in this competition. I started with simple UNET with having 3D gaussian balls as segmentation target. For diversity I also tried [object detection example from MONAI](https://github.com/Project-MONAI/tutorials/tree/main/detection) and realized its working very well out of the box. But when analyzing the different output feature maps and trying to isolate the perfomance gain over gaussian heatmap based segmentation I realized where the advantage was and adjusted my segmentation model accordingly. I learned that \n\n- the penultimate feature map has a higher accuracy than the last one, which is surprising at first\n- gains from box regression are negligable, as particles from the same type have mostly same size anyways.\n\nHence its sufficient to have a pixel-wise loss on the penultimate feature map output. I.e. use a **partly** UNET. The gaussian heatmap is not needed when suppressing background with a low class weight, and just use single pixels as targets. Using the same approach on lower level outputs (= deep supervision) is possible, but does not provide much gain. Input for the segmentation models are 96x96x96 image patches and the loss is calculated on the 48x48x48 output. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F961b07916f0b97850233aae461704c69%2FScreenshot%202025-02-06%20at%2014.05.42.png?generation=1738849991855202&alt=media)\n\nIn general, we observed that relatively small models work really and the design of loss is the most important aspect. We used MONAIs [FlexibleUnet](https://docs.monai.io/en/stable/networks.html#flexibleunet) with backbones resnet34 and efficientnet-b3. Six checkpoints of this architecture would finish under 2h and score 7th place on LB.\n\n## Training procedure\n\nWe used 7 classes (incl background) and weighted CrossEntropy as loss. Notably keeping beta-amylase as a class although it is not scored is quite helpful as model learns to differentiate beta-galactosidase from it. To account for low number of positive pixels, positive pixels are weighted by 256 and background has weight 1. Models were trained with a cosine learning rate schedule with peak LR of 0.001, mixed precision and an effective batch size of 32 samples. Training is based on Random crops and for validation the single experiment image is divided into patches and stored in RAM. \n\n## Ensembling\n\nEnsembling was very challenging, as our two approaches are quite diverse. While in theory predictions from the segmentation models can be ensembled with feature map outputs of the object detection model before runnning the object detection postprocessing, in practice those feature maps have a very different distribution due to difference architectures and loss functions. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F9d47503c1fd18872ef42b466f72c6732%2FScreenshot%202025-02-06%20at%2014.25.00.png?generation=1738850078673576&alt=media)\n\nWe were eager to find an elegant way to fix this scaling issue as we saw the potential of a possible ensemble. The scaling of our best submission works as following. To combine predictions A with predictions B, for each class sort all pixel values for A and B and replace values of B with the corresponding values of A of same rank. In code: \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F05670c1a2a658de03ad73522d1efe22e%2FScreenshot%202025-02-06%20at%2014.58.11.png?generation=1738850308432936&alt=media)\n\nThis results in both predictions having the same distribution, and hence we could simple blend the feature maps before performing object de\ntection task.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fc5ccff4d072d09afe6b6c03273b65d40%2FScreenshot%202025-02-06%20at%2014.38.13.png?generation=1738850131187278&alt=media)\n\n## What did not work\n- Using supplemental data (external or simulated)\n- Other augs\n- Other losses (Tversky, Dice)\n\nThanks for reading. \n\nEdits:\ntraining code: https://github.com/ChristofHenkel/kaggle-cryoet-1st-place-segmentation\ninference kernel: https://www.kaggle.com/code/christofhenkel/cryo-et-1st-place-solution?scriptVersionId=223259615",
      "votes": null
    },
    {
      "id": "3116995",
      "postDate": "02/06/2025 14:17:42",
      "content": "<p>Congratulations!!!<br>\nWould you mind sharing the single model performance in the final ensemble models?</p>",
      "rawMarkdown": "Congratulations!!!\nWould you mind sharing the single model performance in the final ensemble models?",
      "votes": null
    },
    {
      "id": "3117004",
      "postDate": "02/06/2025 14:25:51",
      "content": "<p>for segmentation:</p>\n<p>resnet34 backbone: 0.766<br>\nresnet34 backbone + deep supervision: 0.767<br>\neffnet-b3 backbone: 0.765<br>\ncombined: 0.778</p>",
      "rawMarkdown": "for segmentation:\n\nresnet34 backbone: 0.766\nresnet34 backbone + deep supervision: 0.767\neffnet-b3 backbone: 0.765\ncombined: 0.778",
      "votes": null
    },
    {
      "id": "3117011",
      "postDate": "02/06/2025 14:36:58",
      "content": "<p>Thanks for sharing 🙏</p>",
      "rawMarkdown": "Thanks for sharing 🙏",
      "votes": null
    },
    {
      "id": "3117024",
      "postDate": "02/06/2025 14:53:37",
      "content": "<p>Thank you so much for sharing!</p>",
      "rawMarkdown": "Thank you so much for sharing!",
      "votes": null
    },
    {
      "id": "3117050",
      "postDate": "02/06/2025 15:32:37",
      "content": "<p>Congratulations on the win! Curious to see if the number one spot will come back again shortly 👀</p>\n<p>Your solo scores are amazing with no external data. It would be greatly appreciated if you plan to open source your approach. I got even inspired from your solution on: <a href=\"https://github.com/Project-MONAI/tutorials/tree/main/competitions/kaggle/RANZCR/4th_place_solution\" target=\"_blank\">https://github.com/Project-MONAI/tutorials/tree/main/competitions/kaggle/RANZCR/4th_place_solution</a> while learning about MONAI.</p>",
      "rawMarkdown": "Congratulations on the win! Curious to see if the number one spot will come back again shortly 👀\n\nYour solo scores are amazing with no external data. It would be greatly appreciated if you plan to open source your approach. I got even inspired from your solution on: https://github.com/Project-MONAI/tutorials/tree/main/competitions/kaggle/RANZCR/4th_place_solution while learning about MONAI.",
      "votes": null
    },
    {
      "id": "3117191",
      "postDate": "02/06/2025 17:58:23",
      "content": "<p>Thanks for the detailed write-up and nice ensemble strategy.</p>\n<p>How was your version of MixUp different from the usual implementation? </p>",
      "rawMarkdown": "Thanks for the detailed write-up and nice ensemble strategy.\n\nHow was your version of MixUp different from the usual implementation?",
      "votes": null
    },
    {
      "id": "3117261",
      "postDate": "02/06/2025 19:18:25",
      "content": "<p>congratulations , not fully understand the solution , but i learn a lot from this competition</p>",
      "rawMarkdown": "congratulations , not fully understand the solution , but i learn a lot from this competition",
      "votes": null
    },
    {
      "id": "3117289",
      "postDate": "02/06/2025 20:02:15",
      "content": "<p>Very nice!!!!!!!!!!!!!!</p>",
      "rawMarkdown": "Very nice!!!!!!!!!!!!!!",
      "votes": null
    },
    {
      "id": "3117327",
      "postDate": "02/06/2025 21:12:25",
      "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> congratulations! very interesting solution!</p>",
      "rawMarkdown": "christofhenkel congratulations! very interesting solution!",
      "votes": null
    },
    {
      "id": "3117417",
      "postDate": "02/07/2025 01:26:00",
      "content": "<p>I mean…who thinks of measuring the accuracy of the penultimate feature map?  That or the way you managed to do the ensemble.  Pretty cool stuff!  Thanks for the great explanation!</p>",
      "rawMarkdown": "I mean...who thinks of measuring the accuracy of the penultimate feature map?  That or the way you managed to do the ensemble.  Pretty cool stuff!  Thanks for the great explanation!",
      "votes": null
    },
    {
      "id": "3117717",
      "postDate": "02/07/2025 07:19:16",
      "content": "<p>Congratulations and thank you so much for your write-up, as always!<br>\nThe CV scheme truly impressive. The ensemble approach seems like a form of <a href=\"https://en.wikipedia.org/wiki/Histogram_matching\" target=\"_blank\">histogram matching</a>, isn't it?<br>\nFor single model, it's remarkable how your team and the top solutions have made everything work so seamlessly, especially when I struggled to achieve just a poorer result with a \"sound-like similar\" setup. There are likely many implementation details that come naturally to you but still pose a bit of a challenge for me.<br>\nReally looking forward to seeing your train/infer code so I can better understanding of our problems and failures!</p>",
      "rawMarkdown": "Congratulations and thank you so much for your write-up, as always!\nThe CV scheme truly impressive. The ensemble approach seems like a form of [histogram matching](https://en.wikipedia.org/wiki/Histogram_matching), isn't it?\nFor single model, it's remarkable how your team and the top solutions have made everything work so seamlessly, especially when I struggled to achieve just a poorer result with a \"sound-like similar\" setup. There are likely many implementation details that come naturally to you but still pose a bit of a challenge for me.\nReally looking forward to seeing your train/infer code so I can better understanding of our problems and failures!",
      "votes": null
    },
    {
      "id": "3118259",
      "postDate": "02/07/2025 19:45:43",
      "content": "<p>Well done and congrats. Can you provide some details on the training infrastructure please (how many GPUs, how long it trained, any issues on the losses…)?</p>",
      "rawMarkdown": "Well done and congrats. Can you provide some details on the training infrastructure please (how many GPUs, how long it trained, any issues on the losses...)?",
      "votes": null
    },
    {
      "id": "3118292",
      "postDate": "02/07/2025 20:22:59",
      "content": "<p>Its not much different to other implementations. Some mix rather the losses, I like mixing the targets better. I like a simple pure torch implementation so you can have it as part of the model on GPU. And its flexible for 1D, 2D, 3D. </p>\n<pre><code> torch\n torch  nn\n torch.nn.functional  F\n torch.nn.parameter  Parameter\n torch.distributions  Beta\n\n (nn.Module):\n     ():\n\n        (Mixup, ).__init__()\n        .beta_distribution = Beta(mix_beta, mix_beta)\n        .mixadd = mixadd\n\n     ():\n\n        bs = X.shape[]\n        n_dims = (X.shape)\n        perm = torch.randperm(bs)\n        coeffs = .beta_distribution.rsample(torch.Size((bs,))).to(X.device)\n        X_coeffs = coeffs.view((-,) + (,)*(X.ndim-))\n        Y_coeffs = coeffs.view((-,) + (,)*(Y.ndim-))\n\n        X = X_coeffs * X + (-X_coeffs) * X[perm]\n\n         .mixadd:\n            Y = (Y + Y[perm]).clip(, )\n        :\n            Y = Y_coeffs * Y + ( - Y_coeffs) * Y[perm]\n\n         Z:\n             X, Y, Z\n\n         X, Y\n\n\n (nn.Module):\n\n     ():\n        (Net, ).__init__()\n\n        ...\n        .mixup = Mixup(cfg.mixup_beta)\n        ...\n\n     ():\n\n        x = batch[]\n        y = batch[]\n         .training:\n             torch.rand()[] &lt; .cfg.mixup_p:\n                x, y = .mixup(x,y)\n        ....\n</code></pre>",
      "rawMarkdown": "Its not much different to other implementations. Some mix rather the losses, I like mixing the targets better. I like a simple pure torch implementation so you can have it as part of the model on GPU. And its flexible for 1D, 2D, 3D. \n\n```python\n\nimport torch\nfrom torch import nn\nimport torch.nn.functional as F\nfrom torch.nn.parameter import Parameter\nfrom torch.distributions import Beta\n\nclass Mixup(nn.Module):\n    def __init__(self, mix_beta, mixadd=False):\n\n        super(Mixup, self).__init__()\n        self.beta_distribution = Beta(mix_beta, mix_beta)\n        self.mixadd = mixadd\n\n    def forward(self, X, Y, Z=None):\n\n        bs = X.shape[0]\n        n_dims = len(X.shape)\n        perm = torch.randperm(bs)\n        coeffs = self.beta_distribution.rsample(torch.Size((bs,))).to(X.device)\n        X_coeffs = coeffs.view((-1,) + (1,)*(X.ndim-1))\n        Y_coeffs = coeffs.view((-1,) + (1,)*(Y.ndim-1))\n        \n        X = X_coeffs * X + (1-X_coeffs) * X[perm]\n\n        if self.mixadd:\n            Y = (Y + Y[perm]).clip(0, 1)\n        else:\n            Y = Y_coeffs * Y + (1 - Y_coeffs) * Y[perm]\n                \n        if Z:\n            return X, Y, Z\n\n        return X, Y\n\n\nclass Net(nn.Module):\n\n    def __init__(self, cfg):\n        super(Net, self).__init__()\n\n        ...\n        self.mixup = Mixup(cfg.mixup_beta)\n        ...\n\n    def forward(self, batch):\n\n        x = batch['input']\n        y = batch[\"target\"]\n        if self.training:\n            if torch.rand(1)[0] < self.cfg.mixup_p:\n                x, y = self.mixup(x,y)\n        ....\n```",
      "votes": null
    },
    {
      "id": "3118945",
      "postDate": "02/08/2025 17:22:50",
      "content": "<p>I guess he adapted an existing implementation to work with 3D patches? As far as my knowledge goes, MixUp works with two images with weights alpha and (1-alpha), maybe he did something clever to have images over the temporal dimension?</p>",
      "rawMarkdown": "I guess he adapted an existing implementation to work with 3D patches? As far as my knowledge goes, MixUp works with two images with weights alpha and (1-alpha), maybe he did something clever to have images over the temporal dimension?",
      "votes": null
    },
    {
      "id": "3119445",
      "postDate": "02/09/2025 09:27:54",
      "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>  congratulations to be 1 ranking other time, can you share with us what environement that use to develop your solution ( IDE, HARDWARE…)</p>",
      "rawMarkdown": "christofhenkel  congratulations to be 1 ranking other time, can you share with us what environement that use to develop your solution ( IDE, HARDWARE...)",
      "votes": null
    },
    {
      "id": "3119513",
      "postDate": "02/09/2025 11:12:52",
      "content": "<p>If you have specific points where I should be clearer, I am happy to edit the post and elaborate more. My aim is that everybody can understand the solution.</p>",
      "rawMarkdown": "If you have specific points where I should be clearer, I am happy to edit the post and elaborate more. My aim is that everybody can understand the solution.",
      "votes": null
    },
    {
      "id": "3119522",
      "postDate": "02/09/2025 11:22:02",
      "content": "<p>I trained 6 checkpoints, each takes 2h on a A100 with bf16 mixed-precision. For my models mixup was crucial which more than doubles the number of epochs for training.</p>",
      "rawMarkdown": "I trained 6 checkpoints, each takes 2h on a A100 with bf16 mixed-precision. For my models mixup was crucial which more than doubles the number of epochs for training.",
      "votes": null
    },
    {
      "id": "3119527",
      "postDate": "02/09/2025 11:24:53",
      "content": "<p>For this competition I switched from jupyter notebook as IDE to using VSCODE, simply because it enables the use of Copilot. </p>",
      "rawMarkdown": "For this competition I switched from jupyter notebook as IDE to using VSCODE, simply because it enables the use of Copilot.",
      "votes": null
    },
    {
      "id": "3119537",
      "postDate": "02/09/2025 11:30:19",
      "content": "<p>I was already 1st place on public LB two months ago, with a simple 3D-Unet with gaussian balls as target. Score was around 0.74 back then. It took us 2 month of hard work to improve and refine so the final solution can look \"seamlessly\" as you say. </p>",
      "rawMarkdown": "I was already 1st place on public LB two months ago, with a simple 3D-Unet with gaussian balls as target. Score was around 0.74 back then. It took us 2 month of hard work to improve and refine so the final solution can look \"seamlessly\" as you say.",
      "votes": null
    },
    {
      "id": "3119545",
      "postDate": "02/09/2025 11:43:45",
      "content": "<p>thank you for your reply </p>",
      "rawMarkdown": "thank you for your reply",
      "votes": null
    },
    {
      "id": "3119549",
      "postDate": "02/09/2025 11:45:57",
      "content": "<p>Mixup really just is averaging inputs and targets of a batch with a weight drawn from a Beta-distribution.<br>\nActually we ( <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> and <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> ) coded this four years ago as part of the rainforest competition if I recall correctly. Mixup is really strong for audio. Our implementation was already flexible for arbitrary dimension input and was running on GPU, and I re-used in several competitions since then.</p>",
      "rawMarkdown": "Mixup really just is averaging inputs and targets of a batch with a weight drawn from a Beta-distribution.\nActually we ( @philippsinger and @ilu000 ) coded this four years ago as part of the rainforest competition if I recall correctly. Mixup is really strong for audio. Our implementation was already flexible for arbitrary dimension input and was running on GPU, and I re-used in several competitions since then.",
      "votes": null
    },
    {
      "id": "3119632",
      "postDate": "02/09/2025 13:45:54",
      "content": "<p>Thanks for the detailed write-up and nice ensemble strategy.</p>",
      "rawMarkdown": "Thanks for the detailed write-up and nice ensemble strategy.",
      "votes": null
    },
    {
      "id": "3120235",
      "postDate": "02/10/2025 10:16:56",
      "content": "<p>Thanks for the details. <code>bf16</code> didn't introduce any instabilities?</p>",
      "rawMarkdown": "Thanks for the details. `bf16` didn't introduce any instabilities?",
      "votes": null
    },
    {
      "id": "3120599",
      "postDate": "02/10/2025 17:16:43",
      "content": "<p>Very informative</p>",
      "rawMarkdown": "Very informative",
      "votes": null
    },
    {
      "id": "3120653",
      "postDate": "02/10/2025 18:31:12",
      "content": "<p>very informative </p>",
      "rawMarkdown": "very informative",
      "votes": null
    },
    {
      "id": "3121064",
      "postDate": "02/11/2025 08:17:15",
      "content": "<p>Thanks! You've really motivated us 😀<br>\n0.74 to 0.765+ for single model, i mean all things were improved, but beside <code>design of loss is the most important aspect</code> as you mentioned, could you share your thoughts on a few other key factors that contributed to this improvement? A short ranked list would be incredibly helpful.</p>",
      "rawMarkdown": "Thanks! You've really motivated us 😀\n0.74 to 0.765+ for single model, i mean all things were improved, but beside `design of loss is the most important aspect` as you mentioned, could you share your thoughts on a few other key factors that contributed to this improvement? A short ranked list would be incredibly helpful.",
      "votes": null
    },
    {
      "id": "3122896",
      "postDate": "02/13/2025 06:36:11",
      "content": "<p>Thanks for the detailed write-up and nice ensemble strategy.💪</p>",
      "rawMarkdown": "Thanks for the detailed write-up and nice ensemble strategy.💪",
      "votes": null
    },
    {
      "id": "3123759",
      "postDate": "02/14/2025 06:49:40",
      "content": "<p>I guess you will be happy to hear that we will add our solution to MONAI as a tutorial </p>",
      "rawMarkdown": "I guess you will be happy to hear that we will add our solution to MONAI as a tutorial",
      "votes": null
    },
    {
      "id": "3124038",
      "postDate": "02/14/2025 13:50:05",
      "content": "<p>Interesting🤔</p>",
      "rawMarkdown": "Interesting🤔",
      "votes": null
    },
    {
      "id": "3124063",
      "postDate": "02/14/2025 14:23:58",
      "content": "<p>Oh, that's such a great news!! I am looking forward to reading your tutorial when published. </p>",
      "rawMarkdown": "Oh, that's such a great news!! I am looking forward to reading your tutorial when published.",
      "votes": null
    },
    {
      "id": "3130650",
      "postDate": "02/21/2025 21:39:29",
      "content": "<p>Thanks for the detailed writeup. Using x2 downsampled output can speed up the model greatly. Could you share how you do postprocessing? If you use maxpooling, can you share the kernel size. I am concerned about downsampling might hurt the localization accuracy.</p>",
      "rawMarkdown": "Thanks for the detailed writeup. Using x2 downsampled output can speed up the model greatly. Could you share how you do postprocessing? If you use maxpooling, can you share the kernel size. I am concerned about downsampling might hurt the localization accuracy.",
      "votes": null
    },
    {
      "id": "3135094",
      "postDate": "02/27/2025 03:54:39",
      "content": "<p>Congratulations!!  Thank you, it was a fantastic learning opportunity, and your work is greatly appreciated! </p>",
      "rawMarkdown": "Congratulations!!  Thank you, it was a fantastic learning opportunity, and your work is greatly appreciated!",
      "votes": null
    },
    {
      "id": "3145735",
      "postDate": "03/10/2025 07:59:15",
      "content": "<p>At some point, we will get cursor/copilot within notebooks I believe. 👌</p>",
      "rawMarkdown": "At some point, we will get cursor/copilot within notebooks I believe. 👌",
      "votes": null
    },
    {
      "id": "3145976",
      "postDate": "03/10/2025 12:34:38",
      "content": "<p>for pycharm IDE , it is available in professional version not free version</p>",
      "rawMarkdown": "for pycharm IDE , it is available in professional version not free version",
      "votes": null
    },
    {
      "id": "3186349",
      "postDate": "04/24/2025 14:22:43",
      "content": "<p>Congratulations for the win! What did you mean by gaussian balls as target? Is it some sort of baseline?</p>",
      "rawMarkdown": "Congratulations for the win! What did you mean by gaussian balls as target? Is it some sort of baseline?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3116995,
      "author_name": "maedward",
      "author_url": "",
      "post_date": "02/06/2025 14:17:42",
      "content": "<p>Congratulations!!!<br>\nWould you mind sharing the single model performance in the final ensemble models?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3117004,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "02/06/2025 14:25:51",
          "content": "<p>for segmentation:</p>\n<p>resnet34 backbone: 0.766<br>\nresnet34 backbone + deep supervision: 0.767<br>\neffnet-b3 backbone: 0.765<br>\ncombined: 0.778</p>",
          "votes": null,
          "replies": [
            {
              "id": 3117011,
              "author_name": "nancyalaswad90",
              "author_url": "",
              "post_date": "02/06/2025 14:36:58",
              "content": "<p>Thanks for sharing 🙏</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3117024,
              "author_name": "maedward",
              "author_url": "",
              "post_date": "02/06/2025 14:53:37",
              "content": "<p>Thank you so much for sharing!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3117050,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "02/06/2025 15:32:37",
      "content": "<p>Congratulations on the win! Curious to see if the number one spot will come back again shortly 👀</p>\n<p>Your solo scores are amazing with no external data. It would be greatly appreciated if you plan to open source your approach. I got even inspired from your solution on: <a href=\"https://github.com/Project-MONAI/tutorials/tree/main/competitions/kaggle/RANZCR/4th_place_solution\" target=\"_blank\">https://github.com/Project-MONAI/tutorials/tree/main/competitions/kaggle/RANZCR/4th_place_solution</a> while learning about MONAI.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3123759,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "02/14/2025 06:49:40",
          "content": "<p>I guess you will be happy to hear that we will add our solution to MONAI as a tutorial </p>",
          "votes": null,
          "replies": [
            {
              "id": 3124063,
              "author_name": "snnclsr",
              "author_url": "",
              "post_date": "02/14/2025 14:23:58",
              "content": "<p>Oh, that's such a great news!! I am looking forward to reading your tutorial when published. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3117191,
      "author_name": "brendanartley",
      "author_url": "",
      "post_date": "02/06/2025 17:58:23",
      "content": "<p>Thanks for the detailed write-up and nice ensemble strategy.</p>\n<p>How was your version of MixUp different from the usual implementation? </p>",
      "votes": null,
      "replies": [
        {
          "id": 3118292,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "02/07/2025 20:22:59",
          "content": "<p>Its not much different to other implementations. Some mix rather the losses, I like mixing the targets better. I like a simple pure torch implementation so you can have it as part of the model on GPU. And its flexible for 1D, 2D, 3D. </p>\n<pre><code> torch\n torch  nn\n torch.nn.functional  F\n torch.nn.parameter  Parameter\n torch.distributions  Beta\n\n (nn.Module):\n     ():\n\n        (Mixup, ).__init__()\n        .beta_distribution = Beta(mix_beta, mix_beta)\n        .mixadd = mixadd\n\n     ():\n\n        bs = X.shape[]\n        n_dims = (X.shape)\n        perm = torch.randperm(bs)\n        coeffs = .beta_distribution.rsample(torch.Size((bs,))).to(X.device)\n        X_coeffs = coeffs.view((-,) + (,)*(X.ndim-))\n        Y_coeffs = coeffs.view((-,) + (,)*(Y.ndim-))\n\n        X = X_coeffs * X + (-X_coeffs) * X[perm]\n\n         .mixadd:\n            Y = (Y + Y[perm]).clip(, )\n        :\n            Y = Y_coeffs * Y + ( - Y_coeffs) * Y[perm]\n\n         Z:\n             X, Y, Z\n\n         X, Y\n\n\n (nn.Module):\n\n     ():\n        (Net, ).__init__()\n\n        ...\n        .mixup = Mixup(cfg.mixup_beta)\n        ...\n\n     ():\n\n        x = batch[]\n        y = batch[]\n         .training:\n             torch.rand()[] &lt; .cfg.mixup_p:\n                x, y = .mixup(x,y)\n        ....\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3118945,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "02/08/2025 17:22:50",
          "content": "<p>I guess he adapted an existing implementation to work with 3D patches? As far as my knowledge goes, MixUp works with two images with weights alpha and (1-alpha), maybe he did something clever to have images over the temporal dimension?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3119549,
              "author_name": "christofhenkel",
              "author_url": "",
              "post_date": "02/09/2025 11:45:57",
              "content": "<p>Mixup really just is averaging inputs and targets of a batch with a weight drawn from a Beta-distribution.<br>\nActually we ( <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> and <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> ) coded this four years ago as part of the rainforest competition if I recall correctly. Mixup is really strong for audio. Our implementation was already flexible for arbitrary dimension input and was running on GPU, and I re-used in several competitions since then.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3117261,
      "author_name": "saidkoussi",
      "author_url": "",
      "post_date": "02/06/2025 19:18:25",
      "content": "<p>congratulations , not fully understand the solution , but i learn a lot from this competition</p>",
      "votes": null,
      "replies": [
        {
          "id": 3119513,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "02/09/2025 11:12:52",
          "content": "<p>If you have specific points where I should be clearer, I am happy to edit the post and elaborate more. My aim is that everybody can understand the solution.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3117289,
      "author_name": "saaranshgupta19",
      "author_url": "",
      "post_date": "02/06/2025 20:02:15",
      "content": "<p>Very nice!!!!!!!!!!!!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3117327,
      "author_name": "volodymyrpivoshenko",
      "author_url": "",
      "post_date": "02/06/2025 21:12:25",
      "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> congratulations! very interesting solution!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3117417,
      "author_name": "davidlist",
      "author_url": "",
      "post_date": "02/07/2025 01:26:00",
      "content": "<p>I mean…who thinks of measuring the accuracy of the penultimate feature map?  That or the way you managed to do the ensemble.  Pretty cool stuff!  Thanks for the great explanation!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3117717,
      "author_name": "dangnh0611",
      "author_url": "",
      "post_date": "02/07/2025 07:19:16",
      "content": "<p>Congratulations and thank you so much for your write-up, as always!<br>\nThe CV scheme truly impressive. The ensemble approach seems like a form of <a href=\"https://en.wikipedia.org/wiki/Histogram_matching\" target=\"_blank\">histogram matching</a>, isn't it?<br>\nFor single model, it's remarkable how your team and the top solutions have made everything work so seamlessly, especially when I struggled to achieve just a poorer result with a \"sound-like similar\" setup. There are likely many implementation details that come naturally to you but still pose a bit of a challenge for me.<br>\nReally looking forward to seeing your train/infer code so I can better understanding of our problems and failures!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3119537,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "02/09/2025 11:30:19",
          "content": "<p>I was already 1st place on public LB two months ago, with a simple 3D-Unet with gaussian balls as target. Score was around 0.74 back then. It took us 2 month of hard work to improve and refine so the final solution can look \"seamlessly\" as you say. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3121064,
              "author_name": "dangnh0611",
              "author_url": "",
              "post_date": "02/11/2025 08:17:15",
              "content": "<p>Thanks! You've really motivated us 😀<br>\n0.74 to 0.765+ for single model, i mean all things were improved, but beside <code>design of loss is the most important aspect</code> as you mentioned, could you share your thoughts on a few other key factors that contributed to this improvement? A short ranked list would be incredibly helpful.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3186349,
              "author_name": "adrian774",
              "author_url": "",
              "post_date": "04/24/2025 14:22:43",
              "content": "<p>Congratulations for the win! What did you mean by gaussian balls as target? Is it some sort of baseline?</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3118259,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "02/07/2025 19:45:43",
      "content": "<p>Well done and congrats. Can you provide some details on the training infrastructure please (how many GPUs, how long it trained, any issues on the losses…)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3119522,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "02/09/2025 11:22:02",
          "content": "<p>I trained 6 checkpoints, each takes 2h on a A100 with bf16 mixed-precision. For my models mixup was crucial which more than doubles the number of epochs for training.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3120235,
              "author_name": "yassinealouini",
              "author_url": "",
              "post_date": "02/10/2025 10:16:56",
              "content": "<p>Thanks for the details. <code>bf16</code> didn't introduce any instabilities?</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3119445,
      "author_name": "saidkoussi",
      "author_url": "",
      "post_date": "02/09/2025 09:27:54",
      "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>  congratulations to be 1 ranking other time, can you share with us what environement that use to develop your solution ( IDE, HARDWARE…)</p>",
      "votes": null,
      "replies": [
        {
          "id": 3119527,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "02/09/2025 11:24:53",
          "content": "<p>For this competition I switched from jupyter notebook as IDE to using VSCODE, simply because it enables the use of Copilot. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3119545,
              "author_name": "saidkoussi",
              "author_url": "",
              "post_date": "02/09/2025 11:43:45",
              "content": "<p>thank you for your reply </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3145735,
                  "author_name": "yassinealouini",
                  "author_url": "",
                  "post_date": "03/10/2025 07:59:15",
                  "content": "<p>At some point, we will get cursor/copilot within notebooks I believe. 👌</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3145976,
                      "author_name": "saidkoussi",
                      "author_url": "",
                      "post_date": "03/10/2025 12:34:38",
                      "content": "<p>for pycharm IDE , it is available in professional version not free version</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3119632,
      "author_name": "lavanyabanga25",
      "author_url": "",
      "post_date": "02/09/2025 13:45:54",
      "content": "<p>Thanks for the detailed write-up and nice ensemble strategy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3120599,
      "author_name": "harsaihajsinghgill",
      "author_url": "",
      "post_date": "02/10/2025 17:16:43",
      "content": "<p>Very informative</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3120653,
      "author_name": "bhumikamahajan",
      "author_url": "",
      "post_date": "02/10/2025 18:31:12",
      "content": "<p>very informative </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3122896,
      "author_name": "tongkitmin",
      "author_url": "",
      "post_date": "02/13/2025 06:36:11",
      "content": "<p>Thanks for the detailed write-up and nice ensemble strategy.💪</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3124038,
      "author_name": "yashmaurya2025",
      "author_url": "",
      "post_date": "02/14/2025 13:50:05",
      "content": "<p>Interesting🤔</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3130650,
      "author_name": "linhanwang2",
      "author_url": "",
      "post_date": "02/21/2025 21:39:29",
      "content": "<p>Thanks for the detailed writeup. Using x2 downsampled output can speed up the model greatly. Could you share how you do postprocessing? If you use maxpooling, can you share the kernel size. I am concerned about downsampling might hurt the localization accuracy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3135094,
      "author_name": "sarkarsagar",
      "author_url": "",
      "post_date": "02/27/2025 03:54:39",
      "content": "<p>Congratulations!!  Thank you, it was a fantastic learning opportunity, and your work is greatly appreciated! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3116985": "Thanks to kaggle and everyone involved for hosting this exciting competition. It was a great learning experience and it was very interesting to see how much of our computer vision experience could also be applied to 3D imaging. Thanks to @bloodaxe for this great team experience. \n\n## TLDR\n\nThe solution in an ensemble of segmentation (3D Unets with ResNet & B3 encoders) and object detection models (SegResNet and DynUnet backbones) from [MONAI](https://github.com/Project-MONAI/MONAI). We also used MONAI for augmentations, and exported models via jit or TensorRT, which gave 200% speedup increase and enabled us to have a slightly larger ensemble. We did not use any external or simulated data!\n\nThis post covers the segmentation based approach and ensembling. For object detection part see @bloodaxe writeup: [1st place solution [Object Detection Part]](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561440)\n\n## Cross validation\n\nFor segmentation approach 7 folds were used, simply by splitting by experiment. Using mean of f4-score of all 7 folds had some good correlation with LB. During model training I optimized individual class thresholds at the end of each epoch by simple grid-search on the validation experiment. After all 7 folds were trained we could re-calibrate thresholds, by taking OOF predictions. And fitting the threshold for one fold on the predictions of the other 6. Then we average the resulting f4-curves and take the best threshold.\n\n## Data preprocessing/ augmentations\n\n3D images were normalized by standard normalization, i.e. for each 630x630x184 image, we substract mean and devide by standard deviation before splitting the images into patches.\nSince models are trained from scratch, augmentations were essential to prevent overfitting. \nWe used RandomCrop, Flip on each axis, and rotation, which all are available with MONAI. Additionally I used my own implementation of MixUp which was highly effective to train longer and prevent overfitting.\n\n## Model\n\nModelling was quite a ride in this competition. I started with simple UNET with having 3D gaussian balls as segmentation target. For diversity I also tried [object detection example from MONAI](https://github.com/Project-MONAI/tutorials/tree/main/detection) and realized its working very well out of the box. But when analyzing the different output feature maps and trying to isolate the perfomance gain over gaussian heatmap based segmentation I realized where the advantage was and adjusted my segmentation model accordingly. I learned that \n\n- the penultimate feature map has a higher accuracy than the last one, which is surprising at first\n- gains from box regression are negligable, as particles from the same type have mostly same size anyways.\n\nHence its sufficient to have a pixel-wise loss on the penultimate feature map output. I.e. use a **partly** UNET. The gaussian heatmap is not needed when suppressing background with a low class weight, and just use single pixels as targets. Using the same approach on lower level outputs (= deep supervision) is possible, but does not provide much gain. Input for the segmentation models are 96x96x96 image patches and the loss is calculated on the 48x48x48 output. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F961b07916f0b97850233aae461704c69%2FScreenshot%202025-02-06%20at%2014.05.42.png?generation=1738849991855202&alt=media)\n\nIn general, we observed that relatively small models work really and the design of loss is the most important aspect. We used MONAIs [FlexibleUnet](https://docs.monai.io/en/stable/networks.html#flexibleunet) with backbones resnet34 and efficientnet-b3. Six checkpoints of this architecture would finish under 2h and score 7th place on LB.\n\n## Training procedure\n\nWe used 7 classes (incl background) and weighted CrossEntropy as loss. Notably keeping beta-amylase as a class although it is not scored is quite helpful as model learns to differentiate beta-galactosidase from it. To account for low number of positive pixels, positive pixels are weighted by 256 and background has weight 1. Models were trained with a cosine learning rate schedule with peak LR of 0.001, mixed precision and an effective batch size of 32 samples. Training is based on Random crops and for validation the single experiment image is divided into patches and stored in RAM. \n\n## Ensembling\n\nEnsembling was very challenging, as our two approaches are quite diverse. While in theory predictions from the segmentation models can be ensembled with feature map outputs of the object detection model before runnning the object detection postprocessing, in practice those feature maps have a very different distribution due to difference architectures and loss functions. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F9d47503c1fd18872ef42b466f72c6732%2FScreenshot%202025-02-06%20at%2014.25.00.png?generation=1738850078673576&alt=media)\n\nWe were eager to find an elegant way to fix this scaling issue as we saw the potential of a possible ensemble. The scaling of our best submission works as following. To combine predictions A with predictions B, for each class sort all pixel values for A and B and replace values of B with the corresponding values of A of same rank. In code: \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F05670c1a2a658de03ad73522d1efe22e%2FScreenshot%202025-02-06%20at%2014.58.11.png?generation=1738850308432936&alt=media)\n\nThis results in both predictions having the same distribution, and hence we could simple blend the feature maps before performing object de\ntection task.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fc5ccff4d072d09afe6b6c03273b65d40%2FScreenshot%202025-02-06%20at%2014.38.13.png?generation=1738850131187278&alt=media)\n\n## What did not work\n- Using supplemental data (external or simulated)\n- Other augs\n- Other losses (Tversky, Dice)\n\nThanks for reading. \n\nEdits:\ntraining code: https://github.com/ChristofHenkel/kaggle-cryoet-1st-place-segmentation\ninference kernel: https://www.kaggle.com/code/christofhenkel/cryo-et-1st-place-solution?scriptVersionId=223259615",
    "3116995": "Congratulations!!!\nWould you mind sharing the single model performance in the final ensemble models?",
    "3117004": "for segmentation:\n\nresnet34 backbone: 0.766\nresnet34 backbone + deep supervision: 0.767\neffnet-b3 backbone: 0.765\ncombined: 0.778",
    "3117011": "Thanks for sharing 🙏",
    "3117024": "Thank you so much for sharing!",
    "3117050": "Congratulations on the win! Curious to see if the number one spot will come back again shortly 👀\n\nYour solo scores are amazing with no external data. It would be greatly appreciated if you plan to open source your approach. I got even inspired from your solution on: https://github.com/Project-MONAI/tutorials/tree/main/competitions/kaggle/RANZCR/4th_place_solution while learning about MONAI.",
    "3117191": "Thanks for the detailed write-up and nice ensemble strategy.\n\nHow was your version of MixUp different from the usual implementation?",
    "3117261": "congratulations , not fully understand the solution , but i learn a lot from this competition",
    "3117289": "Very nice!!!!!!!!!!!!!!",
    "3117327": "christofhenkel congratulations! very interesting solution!",
    "3117417": "I mean...who thinks of measuring the accuracy of the penultimate feature map?  That or the way you managed to do the ensemble.  Pretty cool stuff!  Thanks for the great explanation!",
    "3117717": "Congratulations and thank you so much for your write-up, as always!\nThe CV scheme truly impressive. The ensemble approach seems like a form of [histogram matching](https://en.wikipedia.org/wiki/Histogram_matching), isn't it?\nFor single model, it's remarkable how your team and the top solutions have made everything work so seamlessly, especially when I struggled to achieve just a poorer result with a \"sound-like similar\" setup. There are likely many implementation details that come naturally to you but still pose a bit of a challenge for me.\nReally looking forward to seeing your train/infer code so I can better understanding of our problems and failures!",
    "3118259": "Well done and congrats. Can you provide some details on the training infrastructure please (how many GPUs, how long it trained, any issues on the losses...)?",
    "3118292": "Its not much different to other implementations. Some mix rather the losses, I like mixing the targets better. I like a simple pure torch implementation so you can have it as part of the model on GPU. And its flexible for 1D, 2D, 3D. \n\n```python\n\nimport torch\nfrom torch import nn\nimport torch.nn.functional as F\nfrom torch.nn.parameter import Parameter\nfrom torch.distributions import Beta\n\nclass Mixup(nn.Module):\n    def __init__(self, mix_beta, mixadd=False):\n\n        super(Mixup, self).__init__()\n        self.beta_distribution = Beta(mix_beta, mix_beta)\n        self.mixadd = mixadd\n\n    def forward(self, X, Y, Z=None):\n\n        bs = X.shape[0]\n        n_dims = len(X.shape)\n        perm = torch.randperm(bs)\n        coeffs = self.beta_distribution.rsample(torch.Size((bs,))).to(X.device)\n        X_coeffs = coeffs.view((-1,) + (1,)*(X.ndim-1))\n        Y_coeffs = coeffs.view((-1,) + (1,)*(Y.ndim-1))\n        \n        X = X_coeffs * X + (1-X_coeffs) * X[perm]\n\n        if self.mixadd:\n            Y = (Y + Y[perm]).clip(0, 1)\n        else:\n            Y = Y_coeffs * Y + (1 - Y_coeffs) * Y[perm]\n                \n        if Z:\n            return X, Y, Z\n\n        return X, Y\n\n\nclass Net(nn.Module):\n\n    def __init__(self, cfg):\n        super(Net, self).__init__()\n\n        ...\n        self.mixup = Mixup(cfg.mixup_beta)\n        ...\n\n    def forward(self, batch):\n\n        x = batch['input']\n        y = batch[\"target\"]\n        if self.training:\n            if torch.rand(1)[0] < self.cfg.mixup_p:\n                x, y = self.mixup(x,y)\n        ....\n```",
    "3118945": "I guess he adapted an existing implementation to work with 3D patches? As far as my knowledge goes, MixUp works with two images with weights alpha and (1-alpha), maybe he did something clever to have images over the temporal dimension?",
    "3119445": "christofhenkel  congratulations to be 1 ranking other time, can you share with us what environement that use to develop your solution ( IDE, HARDWARE...)",
    "3119513": "If you have specific points where I should be clearer, I am happy to edit the post and elaborate more. My aim is that everybody can understand the solution.",
    "3119522": "I trained 6 checkpoints, each takes 2h on a A100 with bf16 mixed-precision. For my models mixup was crucial which more than doubles the number of epochs for training.",
    "3119527": "For this competition I switched from jupyter notebook as IDE to using VSCODE, simply because it enables the use of Copilot.",
    "3119537": "I was already 1st place on public LB two months ago, with a simple 3D-Unet with gaussian balls as target. Score was around 0.74 back then. It took us 2 month of hard work to improve and refine so the final solution can look \"seamlessly\" as you say.",
    "3119545": "thank you for your reply",
    "3119549": "Mixup really just is averaging inputs and targets of a batch with a weight drawn from a Beta-distribution.\nActually we ( @philippsinger and @ilu000 ) coded this four years ago as part of the rainforest competition if I recall correctly. Mixup is really strong for audio. Our implementation was already flexible for arbitrary dimension input and was running on GPU, and I re-used in several competitions since then.",
    "3119632": "Thanks for the detailed write-up and nice ensemble strategy.",
    "3120235": "Thanks for the details. `bf16` didn't introduce any instabilities?",
    "3120599": "Very informative",
    "3120653": "very informative",
    "3121064": "Thanks! You've really motivated us 😀\n0.74 to 0.765+ for single model, i mean all things were improved, but beside `design of loss is the most important aspect` as you mentioned, could you share your thoughts on a few other key factors that contributed to this improvement? A short ranked list would be incredibly helpful.",
    "3122896": "Thanks for the detailed write-up and nice ensemble strategy.💪",
    "3123759": "I guess you will be happy to hear that we will add our solution to MONAI as a tutorial",
    "3124038": "Interesting🤔",
    "3124063": "Oh, that's such a great news!! I am looking forward to reading your tutorial when published.",
    "3130650": "Thanks for the detailed writeup. Using x2 downsampled output can speed up the model greatly. Could you share how you do postprocessing? If you use maxpooling, can you share the kernel size. I am concerned about downsampling might hurt the localization accuracy.",
    "3135094": "Congratulations!!  Thank you, it was a fantastic learning opportunity, and your work is greatly appreciated!",
    "3145735": "At some point, we will get cursor/copilot within notebooks I believe. 👌",
    "3145976": "for pycharm IDE , it is available in professional version not free version",
    "3186349": "Congratulations for the win! What did you mean by gaussian balls as target? Is it some sort of baseline?"
  },
  "source": "meta"
}