{
  "id": 583143,
  "title": "1st Place - 3D U-Net + Quantile Thresholding",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/583143",
  "author_name": "Bartley",
  "post_date": "2025-06-05T04:04:45.037000",
  "votes": 146,
  "comment_count": 60,
  "views": 0,
  "content": "<p>Thanks to BYU and Kaggle for hosting this competition. It was nice to have another well-run tomography competition and the hosts were awesome. I can't believe the result!</p>\n<h2>TLDR</h2>\n<p>My solution uses a 3D U-Net trained with heavy augmentations and auxiliary loss functions. During inference, I rank each tomogram based on the max predicted pixel value and use quantile thresholding to determine if a motor is present.</p>\n<h2>Cross Validation</h2>\n<p>For validating models, the competition data is split into 4 folds. Local CV strongly correlates with the LB up to about 0.93. Beyond that, I used the public LB for validation. It was important to use quantile thresholding to get reliable feedback from the LB. More on this in the post-processing section.</p>\n<h2>Preprocessing</h2>\n<p>Tomograms from the competition data and the CryoET Data Portal are used to create a training set. Each tomogram is resized to (128, 704, 704) using <code>scipy.ndimage.zoom()</code>, and tomograms without motors are discarded. As others noted, the competition data is quite noisy, so Napari was used to manually add missing motors. I will add the updated data <a href=\"https://www.kaggle.com/datasets/brendanartley/cryoet-flagellar-motors-dataset\" target=\"_blank\">here</a>.</p>\n<p>For the labels, I use a Gaussian heat map centered on each motor. Similar to <a href=\"https://www.kaggle.com/bloodaxe\" target=\"_blank\">@bloodaxe</a> and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>'s solution in the CZII competition, the resolution of the heatmap is reduced by 8x. This works especially well for this competition as there is a high tolerance for distance error in the metric. This means that predicting the exact pixel is not as important as predicting motor presence. If you are not convinced, the following plot shows roughly how much error is allowed around each motor when voxel spacing equals 10.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F406663dc00b116aa614b191abd71b37f%2Ftomograms.JPG?generation=1749094386913073&amp;alt=media\" alt=\"tomogram_image\"></p>\n<h2>Model</h2>\n<p>The model is a 3D U-Net (sort of). The encoder is a pre-trained ResNet200 from Kenoshara’s repository <a href=\"https://github.com/kenshohara/3D-ResNets-PyTorch\" target=\"_blank\">here</a>. For most experiments, I used the ResNet101 variant, but increasing the capacity of the encoder yields better performance. In addition, stochastic dropout is applied for regularization, and gradient checkpointing is used to reduce vRAM usage during training. The decoder uses a single deconvolution block before the segmentation head.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F9242bdbaf78089afc0842c382a4e40a2%2FBYU2025-Model%20(2).jpg?generation=1749095150932706&amp;alt=media\" alt=\"model_image\"></p>\n<h2>Loss</h2>\n<p>The model is trained using SmoothBCE loss with 3 contributions. The main segmentation head predicts the output logits, a deep supervision head is applied to the second last feature map, and a max pooled loss (kernel size and stride of 4) is applied on the main segmentation head. Moreover, the pooled loss encourages high probabilities around the motor region, while reducing the penalty for small localization errors.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fcedd954d5796ada89d6dd3749b00afa2%2FBYU2025-Loss%20(1).jpg?generation=1751317472394799&amp;alt=media\" alt=\"loss_image\"></p>\n<h2>Augmentations</h2>\n<p>Heavy augmentations enabled training for 400 epochs without overfitting. Although, I could probably have trained longer there was no change in the public LB scores beyond 250 epochs.</p>\n<ul>\n<li>Mixup (100%)</li>\n<li>Rescale/Zoom (100%)</li>\n<li>Rotate90/180/270 (100%)</li>\n<li>Axis Flips (100%)</li>\n<li>Axis Swap (100%)</li>\n<li>Coarse Dropout (50%)</li>\n<li>Color inversion (25%)</li>\n<li>Simple Cutmix (15%)</li>\n</ul>\n<p>Loading tomograms from disk is slow, which limits the time for augmentations on the CPU. To address this, all augmentations but rescaling are applied on the GPU. To keep rescaling as fast as possible <code>scipy.ndimage.zoom(..., order=0)</code> is used. </p>\n<h2>Inference</h2>\n<p>Initially, the same preprocessing pipeline was applied during inference. This worked well, but it was 4x faster to match the patch height and width, and only slide over the depth. This allows more time for TTA and a very high overlap (0.875). Both approaches scored about the same, but my final solution uses the latter.</p>\n<p>All edge predictions are down-weighted using the <code>roi_weight_map</code> parameter. The middle 40% of the logits are weighted as 1.0 and other logits are weighted as 0.001 when aggregating the sliding window.</p>\n<h2>Ensembling</h2>\n<p>The final submission uses an 8-seed ensemble. Sigmoid is applied to each model output and the logits are summed. Inference takes ~10 hrs.</p>\n<h2>Postprocessing</h2>\n<p>Like many others, I found that fixed thresholds were unstable. Instead, I use quantile thresholding to determine motor presence.</p>\n<p>To apply this, all tomograms are ranked based on their max predicted pixel value. Then, predictions for the lowest quantile are removed. I tuned the quantile on the public LB and then prayed to the Kaggle gods that the private LB was similar. On the public LB the optimal threshold was 0.565 and on private it was 0.560. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F6a05822baf7e6d70c6f8d932aabb51b0%2Fscores.JPG?generation=1749094657707643&amp;alt=media\" alt=\"LB_image\"></p>\n<h2>Final Note</h2>\n<p>Thanks for reading, and thanks to everyone who showed their appreciation for the external dataset. </p>\n<p>External data <a href=\"https://www.kaggle.com/datasets/brendanartley/cryoet-flagellar-motors-dataset\" target=\"_blank\">here</a><br>\nGithub repository <a href=\"https://github.com/brendanartley/BYU-competition\" target=\"_blank\">here</a><br>\nMetadata <a href=\"https://www.kaggle.com/datasets/brendanartley/solution-ds-byu-1st-place-metadata/data\" target=\"_blank\">here</a></p>\n<p>Happy Kaggling!</p>",
  "messages": [
    {
      "id": 3217480,
      "postDate": "2025-06-05T04:04:45.037Z",
      "content": "<p>Thanks to BYU and Kaggle for hosting this competition. It was nice to have another well-run tomography competition and the hosts were awesome. I can't believe the result!</p>\n<h2>TLDR</h2>\n<p>My solution uses a 3D U-Net trained with heavy augmentations and auxiliary loss functions. During inference, I rank each tomogram based on the max predicted pixel value and use quantile thresholding to determine if a motor is present.</p>\n<h2>Cross Validation</h2>\n<p>For validating models, the competition data is split into 4 folds. Local CV strongly correlates with the LB up to about 0.93. Beyond that, I used the public LB for validation. It was important to use quantile thresholding to get reliable feedback from the LB. More on this in the post-processing section.</p>\n<h2>Preprocessing</h2>\n<p>Tomograms from the competition data and the CryoET Data Portal are used to create a training set. Each tomogram is resized to (128, 704, 704) using <code>scipy.ndimage.zoom()</code>, and tomograms without motors are discarded. As others noted, the competition data is quite noisy, so Napari was used to manually add missing motors. I will add the updated data <a href=\"https://www.kaggle.com/datasets/brendanartley/cryoet-flagellar-motors-dataset\" target=\"_blank\">here</a>.</p>\n<p>For the labels, I use a Gaussian heat map centered on each motor. Similar to <a href=\"https://www.kaggle.com/bloodaxe\" target=\"_blank\">@bloodaxe</a> and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>'s solution in the CZII competition, the resolution of the heatmap is reduced by 8x. This works especially well for this competition as there is a high tolerance for distance error in the metric. This means that predicting the exact pixel is not as important as predicting motor presence. If you are not convinced, the following plot shows roughly how much error is allowed around each motor when voxel spacing equals 10.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F406663dc00b116aa614b191abd71b37f%2Ftomograms.JPG?generation=1749094386913073&amp;alt=media\" alt=\"tomogram_image\"></p>\n<h2>Model</h2>\n<p>The model is a 3D U-Net (sort of). The encoder is a pre-trained ResNet200 from Kenoshara’s repository <a href=\"https://github.com/kenshohara/3D-ResNets-PyTorch\" target=\"_blank\">here</a>. For most experiments, I used the ResNet101 variant, but increasing the capacity of the encoder yields better performance. In addition, stochastic dropout is applied for regularization, and gradient checkpointing is used to reduce vRAM usage during training. The decoder uses a single deconvolution block before the segmentation head.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F9242bdbaf78089afc0842c382a4e40a2%2FBYU2025-Model%20(2).jpg?generation=1749095150932706&amp;alt=media\" alt=\"model_image\"></p>\n<h2>Loss</h2>\n<p>The model is trained using SmoothBCE loss with 3 contributions. The main segmentation head predicts the output logits, a deep supervision head is applied to the second last feature map, and a max pooled loss (kernel size and stride of 4) is applied on the main segmentation head. Moreover, the pooled loss encourages high probabilities around the motor region, while reducing the penalty for small localization errors.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fcedd954d5796ada89d6dd3749b00afa2%2FBYU2025-Loss%20(1).jpg?generation=1751317472394799&amp;alt=media\" alt=\"loss_image\"></p>\n<h2>Augmentations</h2>\n<p>Heavy augmentations enabled training for 400 epochs without overfitting. Although, I could probably have trained longer there was no change in the public LB scores beyond 250 epochs.</p>\n<ul>\n<li>Mixup (100%)</li>\n<li>Rescale/Zoom (100%)</li>\n<li>Rotate90/180/270 (100%)</li>\n<li>Axis Flips (100%)</li>\n<li>Axis Swap (100%)</li>\n<li>Coarse Dropout (50%)</li>\n<li>Color inversion (25%)</li>\n<li>Simple Cutmix (15%)</li>\n</ul>\n<p>Loading tomograms from disk is slow, which limits the time for augmentations on the CPU. To address this, all augmentations but rescaling are applied on the GPU. To keep rescaling as fast as possible <code>scipy.ndimage.zoom(..., order=0)</code> is used. </p>\n<h2>Inference</h2>\n<p>Initially, the same preprocessing pipeline was applied during inference. This worked well, but it was 4x faster to match the patch height and width, and only slide over the depth. This allows more time for TTA and a very high overlap (0.875). Both approaches scored about the same, but my final solution uses the latter.</p>\n<p>All edge predictions are down-weighted using the <code>roi_weight_map</code> parameter. The middle 40% of the logits are weighted as 1.0 and other logits are weighted as 0.001 when aggregating the sliding window.</p>\n<h2>Ensembling</h2>\n<p>The final submission uses an 8-seed ensemble. Sigmoid is applied to each model output and the logits are summed. Inference takes ~10 hrs.</p>\n<h2>Postprocessing</h2>\n<p>Like many others, I found that fixed thresholds were unstable. Instead, I use quantile thresholding to determine motor presence.</p>\n<p>To apply this, all tomograms are ranked based on their max predicted pixel value. Then, predictions for the lowest quantile are removed. I tuned the quantile on the public LB and then prayed to the Kaggle gods that the private LB was similar. On the public LB the optimal threshold was 0.565 and on private it was 0.560. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F6a05822baf7e6d70c6f8d932aabb51b0%2Fscores.JPG?generation=1749094657707643&amp;alt=media\" alt=\"LB_image\"></p>\n<h2>Final Note</h2>\n<p>Thanks for reading, and thanks to everyone who showed their appreciation for the external dataset. </p>\n<p>External data <a href=\"https://www.kaggle.com/datasets/brendanartley/cryoet-flagellar-motors-dataset\" target=\"_blank\">here</a><br>\nGithub repository <a href=\"https://github.com/brendanartley/BYU-competition\" target=\"_blank\">here</a><br>\nMetadata <a href=\"https://www.kaggle.com/datasets/brendanartley/solution-ds-byu-1st-place-metadata/data\" target=\"_blank\">here</a></p>\n<p>Happy Kaggling!</p>",
      "rawMarkdown": "Thanks to BYU and Kaggle for hosting this competition. It was nice to have another well-run tomography competition and the hosts were awesome. I can't believe the result!\n\n## TLDR\n\nMy solution uses a 3D U-Net trained with heavy augmentations and auxiliary loss functions. During inference, I rank each tomogram based on the max predicted pixel value and use quantile thresholding to determine if a motor is present.\n\n## Cross Validation\n\nFor validating models, the competition data is split into 4 folds. Local CV strongly correlates with the LB up to about 0.93. Beyond that, I used the public LB for validation. It was important to use quantile thresholding to get reliable feedback from the LB. More on this in the post-processing section.\n\n## Preprocessing\n\nTomograms from the competition data and the CryoET Data Portal are used to create a training set. Each tomogram is resized to (128, 704, 704) using `scipy.ndimage.zoom()`, and tomograms without motors are discarded. As others noted, the competition data is quite noisy, so Napari was used to manually add missing motors. I will add the updated data [here](https://www.kaggle.com/datasets/brendanartley/cryoet-flagellar-motors-dataset).\n\nFor the labels, I use a Gaussian heat map centered on each motor. Similar to @bloodaxe and @christofhenkel's solution in the CZII competition, the resolution of the heatmap is reduced by 8x. This works especially well for this competition as there is a high tolerance for distance error in the metric. This means that predicting the exact pixel is not as important as predicting motor presence. If you are not convinced, the following plot shows roughly how much error is allowed around each motor when voxel spacing equals 10.\n\n![tomogram_image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F406663dc00b116aa614b191abd71b37f%2Ftomograms.JPG?generation=1749094386913073&alt=media)\n\n## Model\n\nThe model is a 3D U-Net (sort of). The encoder is a pre-trained ResNet200 from Kenoshara’s repository [here](https://github.com/kenshohara/3D-ResNets-PyTorch). For most experiments, I used the ResNet101 variant, but increasing the capacity of the encoder yields better performance. In addition, stochastic dropout is applied for regularization, and gradient checkpointing is used to reduce vRAM usage during training. The decoder uses a single deconvolution block before the segmentation head.\n\n![model_image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F9242bdbaf78089afc0842c382a4e40a2%2FBYU2025-Model%20(2).jpg?generation=1749095150932706&alt=media)\n\n## Loss\n\nThe model is trained using SmoothBCE loss with 3 contributions. The main segmentation head predicts the output logits, a deep supervision head is applied to the second last feature map, and a max pooled loss (kernel size and stride of 4) is applied on the main segmentation head. Moreover, the pooled loss encourages high probabilities around the motor region, while reducing the penalty for small localization errors.\n\n![loss_image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fcedd954d5796ada89d6dd3749b00afa2%2FBYU2025-Loss%20(1).jpg?generation=1751317472394799&alt=media)\n\n## Augmentations\n\nHeavy augmentations enabled training for 400 epochs without overfitting. Although, I could probably have trained longer there was no change in the public LB scores beyond 250 epochs.\n\n- Mixup (100%)\n- Rescale/Zoom (100%)\n- Rotate90/180/270 (100%)\n- Axis Flips (100%)\n- Axis Swap (100%)\n- Coarse Dropout (50%)\n- Color inversion (25%)\n- Simple Cutmix (15%)\n\nLoading tomograms from disk is slow, which limits the time for augmentations on the CPU. To address this, all augmentations but rescaling are applied on the GPU. To keep rescaling as fast as possible `scipy.ndimage.zoom(..., order=0)` is used. \n\n## Inference\n\nInitially, the same preprocessing pipeline was applied during inference. This worked well, but it was 4x faster to match the patch height and width, and only slide over the depth. This allows more time for TTA and a very high overlap (0.875). Both approaches scored about the same, but my final solution uses the latter.\n\nAll edge predictions are down-weighted using the `roi_weight_map` parameter. The middle 40% of the logits are weighted as 1.0 and other logits are weighted as 0.001 when aggregating the sliding window.\n\n## Ensembling\n\nThe final submission uses an 8-seed ensemble. Sigmoid is applied to each model output and the logits are summed. Inference takes ~10 hrs.\n\n## Postprocessing\n\nLike many others, I found that fixed thresholds were unstable. Instead, I use quantile thresholding to determine motor presence.\n\nTo apply this, all tomograms are ranked based on their max predicted pixel value. Then, predictions for the lowest quantile are removed. I tuned the quantile on the public LB and then prayed to the Kaggle gods that the private LB was similar. On the public LB the optimal threshold was 0.565 and on private it was 0.560. \n\n![LB_image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F6a05822baf7e6d70c6f8d932aabb51b0%2Fscores.JPG?generation=1749094657707643&alt=media)\n\n## Final Note\n\nThanks for reading, and thanks to everyone who showed their appreciation for the external dataset. \n\nExternal data [here](https://www.kaggle.com/datasets/brendanartley/cryoet-flagellar-motors-dataset)\nGithub repository [here](https://github.com/brendanartley/BYU-competition)\nMetadata [here](https://www.kaggle.com/datasets/brendanartley/solution-ds-byu-1st-place-metadata/data)\n\nHappy Kaggling!",
      "votes": 146
    },
    {
      "id": 3217533,
      "postDate": "2025-06-05T05:39:20.967Z",
      "content": "<p>Congratulations, and thanks for the detailed writeup!</p>\n<p>I actually spent my first month on this competition on a very similar 3D UNet approach, building from my CZII competition solution. A very simple first version already got to ~0.85 CV in the first few days, but never scored above 0.2 on the leaderboard (and usually 0.0). I spent a month trying to figure out what was going on, but failed and fell back to YOLO.</p>\n<p>Did you face anything similar at any point? I really want to know what was going on…</p>",
      "rawMarkdown": "Congratulations, and thanks for the detailed writeup!\n\nI actually spent my first month on this competition on a very similar 3D UNet approach, building from my CZII competition solution. A very simple first version already got to ~0.85 CV in the first few days, but never scored above 0.2 on the leaderboard (and usually 0.0). I spent a month trying to figure out what was going on, but failed and fell back to YOLO.\n\nDid you face anything similar at any point? I really want to know what was going on...",
      "votes": 7,
      "replies": [
        {
          "id": 3217663,
          "postDate": "2025-06-05T09:16:29.337Z",
          "content": "<p>Yup, myself and <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> as well faced similar issues independently while trying to work with 3D-UNet. The best I could get on the public LB was around ~0.22 and for sersasj it was around ~0.5. </p>",
          "rawMarkdown": "Yup, myself and @sersasj as well faced similar issues independently while trying to work with 3D-UNet. The best I could get on the public LB was around ~0.22 and for sersasj it was around ~0.5. ",
          "votes": 1
        },
        {
          "id": 3217683,
          "postDate": "2025-06-05T09:55:16.140Z",
          "content": "<p>Coincidentally, I encountered this similar problem in YOLOv10. I found that parameter adjustments such as conf_thr, iou have a great impact on LB. For example, I reduced conf_thr from 0.7 to 0.4, and lb dropped 4pp, which also confused me very much. In theory, it is unlikely that the parameter adjustment will have such a large fluctuation.<br>\nI attribute it to the sparse data, and I did not check the specific reasons.</p>",
          "rawMarkdown": "Coincidentally, I encountered this similar problem in YOLOv10. I found that parameter adjustments such as conf_thr, iou have a great impact on LB. For example, I reduced conf_thr from 0.7 to 0.4, and lb dropped 4pp, which also confused me very much. In theory, it is unlikely that the parameter adjustment will have such a large fluctuation.\nI attribute it to the sparse data, and I did not check the specific reasons.",
          "votes": 1
        },
        {
          "id": 3217730,
          "postDate": "2025-06-05T11:06:17.570Z",
          "content": "<p>I have the same problem, local cv scores more than 0.95 on sliced inference with 3d unet but max public score is 0.68</p>",
          "rawMarkdown": "I have the same problem, local cv scores more than 0.95 on sliced inference with 3d unet but max public score is 0.68",
          "votes": 1
        },
        {
          "id": 3217810,
          "postDate": "2025-06-05T12:21:24.893Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/Jeroen\" target=\"_blank\">@Jeroen</a>,</p>\n<p>Its hard to say - my first submissions with a resnet18 encoder scored in the 0.5 - 0.7 range.  Some key things that helped boost performance from there were:</p>\n<ol>\n<li>Deeper/larger encoders</li>\n<li>Large patch sizes (that covered &gt;85% of the height and width)</li>\n<li>High resolution tomograms</li>\n<li>&gt;50% depth overlap during inference</li>\n</ol>\n<p>Were you using a pretrained encoder?</p>",
          "rawMarkdown": "Hi @Jeroen,\n\nIts hard to say - my first submissions with a resnet18 encoder scored in the 0.5 - 0.7 range.  Some key things that helped boost performance from there were:\n\n1. Deeper/larger encoders\n2. Large patch sizes (that covered >85% of the height and width)\n3. High resolution tomograms\n4. >50% depth overlap during inference\n\nWere you using a pretrained encoder?\n",
          "votes": 2
        }
      ]
    },
    {
      "id": 3218995,
      "postDate": "2025-06-07T03:30:36.697Z",
      "content": "<p>Congratulations!!, and particularly thanks for your effort in providing the external dataset—made a significant impact on the entire community. This recognition is well-deserved, and your contribution truly set the stage for success.❕</p>",
      "rawMarkdown": "Congratulations!!, and particularly thanks for your effort in providing the external dataset—made a significant impact on the entire community. This recognition is well-deserved, and your contribution truly set the stage for success.❕",
      "votes": 3
    },
    {
      "id": 3271444,
      "postDate": "2025-08-18T23:23:26.723Z",
      "content": "<p>Thanks for the write up and congratulations on the win. As someone fairly new to data science I appreciate people like you who share their techniques to help me learn faster.</p>",
      "rawMarkdown": "Thanks for the write up and congratulations on the win. As someone fairly new to data science I appreciate people like you who share their techniques to help me learn faster.",
      "votes": 2
    },
    {
      "id": 3229334,
      "postDate": "2025-06-21T10:28:31.383Z",
      "content": "<p>Congratulations!! A lot to learn from you🙂</p>",
      "rawMarkdown": "Congratulations!! A lot to learn from you🙂",
      "votes": 1
    },
    {
      "id": 3223244,
      "postDate": "2025-06-13T04:07:11.933Z",
      "content": "<p>Congratulations! Your merit is greater for competing alone, and with such close scores. I wonder if it's possible to experiment with your work in a Jupyter notebook and Anaconda?  <br>\nAll the best</p>",
      "rawMarkdown": "Congratulations! Your merit is greater for competing alone, and with such close scores. I wonder if it's possible to experiment with your work in a Jupyter notebook and Anaconda?  \nAll the best",
      "votes": 1
    },
    {
      "id": 3223200,
      "postDate": "2025-06-13T02:26:35.490Z",
      "content": "<p>Massive congrats on the 1st place finish—what a brilliant piece of work. The 3D U-Net, all that heavy augmentation, and the smart use of quantile thresholding really paid off. well deserved and super inspiring.</p>",
      "rawMarkdown": "Massive congrats on the 1st place finish—what a brilliant piece of work. The 3D U-Net, all that heavy augmentation, and the smart use of quantile thresholding really paid off. well deserved and super inspiring.",
      "votes": 1
    },
    {
      "id": 3222828,
      "postDate": "2025-06-12T14:57:19.640Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> on winning and thanks for the writeup!<br>\nI have a question regarding the quantile thresholding: What exactly is the advantage in your opinion? Because I am not sure I have understood it completly and as far as I understand it, you still have to tune the threshold and the threshold would still change depending on the used model right?<br>\nIs it because a quantile threshold gives somewhat of a more smooth threshold on how many predictions are counted?</p>\n<p>It will probably be clear, when your inference code is released but still wanted to ask.</p>",
      "rawMarkdown": "Congrats @brendanartley on winning and thanks for the writeup!\nI have a question regarding the quantile thresholding: What exactly is the advantage in your opinion? Because I am not sure I have understood it completly and as far as I understand it, you still have to tune the threshold and the threshold would still change depending on the used model right?\nIs it because a quantile threshold gives somewhat of a more smooth threshold on how many predictions are counted?\n\nIt will probably be clear, when your inference code is released but still wanted to ask.",
      "votes": 1
    },
    {
      "id": 3222334,
      "postDate": "2025-06-12T06:16:17.510Z",
      "content": "<p>Thanks for this solution,its really helpful.</p>",
      "rawMarkdown": "Thanks for this solution,its really helpful.",
      "votes": 1
    },
    {
      "id": 3222121,
      "postDate": "2025-06-11T21:10:19.610Z",
      "content": "<p>Hi,<br>\nThe aux head shouldn't be applied on the last feature map of the encoder as you only use one decoder block?<br>\nIn the <strong>Loss</strong> image you shared, it is applied on the decoder block instead.</p>",
      "rawMarkdown": "Hi,\nThe aux head shouldn't be applied on the last feature map of the encoder as you only use one decoder block?\nIn the **Loss** image you shared, it is applied on the decoder block instead.",
      "votes": 1
    },
    {
      "id": 3220587,
      "postDate": "2025-06-09T14:56:57.667Z",
      "content": "<p>Congrutulations,thanks for your solution.It's amazing.</p>",
      "rawMarkdown": "Congrutulations,thanks for your solution.It's amazing.",
      "votes": 1
    },
    {
      "id": 3220050,
      "postDate": "2025-06-08T17:53:04.080Z",
      "content": "<p>Congratulations, and thanks for detailed writeup…!</p>",
      "rawMarkdown": "Congratulations, and thanks for detailed writeup…!",
      "votes": 1
    },
    {
      "id": 3218727,
      "postDate": "2025-06-06T16:12:07.793Z",
      "content": "<p>Congratulations, and thanks for detailed writeup…!</p>",
      "rawMarkdown": "Congratulations, and thanks for detailed writeup...!",
      "votes": 1
    },
    {
      "id": 3218188,
      "postDate": "2025-06-06T00:50:25.983Z",
      "content": "<p>Congrats! Is it possible to share this solution's training and inference source code?</p>",
      "rawMarkdown": "Congrats! Is it possible to share this solution's training and inference source code?",
      "votes": 1,
      "replies": [
        {
          "id": 3218255,
          "postDate": "2025-06-06T03:06:10.907Z",
          "content": "<p>Cheers <a href=\"https://www.kaggle.com/calebyenusah\" target=\"_blank\">@calebyenusah</a>. </p>\n<p>I just released the training code <a href=\"https://github.com/brendanartley/BYU-competition\" target=\"_blank\">here</a>. The inference pipeline is coming soon.</p>",
          "rawMarkdown": "Cheers @calebyenusah. \n\nI just released the training code [here](https://github.com/brendanartley/BYU-competition). The inference pipeline is coming soon.",
          "votes": 1,
          "replies": [
            {
              "id": 3301015,
              "postDate": "2025-10-12T09:05:06.603Z",
              "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> would love to get the inference pipeline to understand your approach to the problem</p>",
              "rawMarkdown": "@brendanartley would love to get the inference pipeline to understand your approach to the problem"
            }
          ]
        }
      ]
    },
    {
      "id": 3217850,
      "postDate": "2025-06-05T13:00:04.373Z",
      "content": "<p>Congrats! I think the thresholding trick you have done has a lot of weight in the final results, you did a great job nailing that! </p>",
      "rawMarkdown": "Congrats! I think the thresholding trick you have done has a lot of weight in the final results, you did a great job nailing that! ",
      "votes": 1
    },
    {
      "id": 3217717,
      "postDate": "2025-06-05T10:54:29.667Z",
      "content": "<p>Hi and congrats, I have some questions seeing as I used a similar approach but didn't get past .7 LB:</p>\n<ul>\n<li>Did you use all of the external data? I found that the tomograms with the smallest voxel spacings hindered my models significantly</li>\n<li>What crop size did you use for training? I tried to go as large as possible but I think it was a mistake.</li>\n<li>How long did training and especially validation take? I had to reduce the size of my validation sets because it was difficult to train for a lot of epochs</li>\n</ul>\n<p>Anectodally I estimated that around 35% of the test samples were positive, and settled for a threshold around the 58th percentile.</p>",
      "rawMarkdown": "Hi and congrats, I have some questions seeing as I used a similar approach but didn't get past .7 LB:\n\n- Did you use all of the external data? I found that the tomograms with the smallest voxel spacings hindered my models significantly\n- What crop size did you use for training? I tried to go as large as possible but I think it was a mistake.\n- How long did training and especially validation take? I had to reduce the size of my validation sets because it was difficult to train for a lot of epochs\n\nAnectodally I estimated that around 35% of the test samples were positive, and settled for a threshold around the 58th percentile.",
      "votes": 1,
      "replies": [
        {
          "id": 3217820,
          "postDate": "2025-06-05T12:39:46.897Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tennogh\" target=\"_blank\">@tennogh</a>,</p>\n<ol>\n<li><p>I used all tomograms.</p></li>\n<li><p>The crop size was <code>(64, 674, 674)</code> - I think you were on the right track with this.</p></li>\n<li><p>Small models trained in 2 - 8 hours, while the final ensemble models took 35 - 40 hours each.</p></li>\n</ol>",
          "rawMarkdown": "Hi @tennogh,\n\n1. I used all tomograms.\n\n2. The crop size was `(64, 674, 674)` - I think you were on the right track with this.\n\n3. Small models trained in 2 - 8 hours, while the final ensemble models took 35 - 40 hours each.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3217708,
      "postDate": "2025-06-05T10:36:39.920Z",
      "content": "<p>Congratulations. Thanks for share.</p>\n<blockquote>\n  <p>Rotate90/180/270 (100%)</p>\n</blockquote>\n<p>So no free rotations at all. Do you think apply free rotations as I did could been perjudicial because the small size of the actual target, the motor?</p>",
      "rawMarkdown": "Congratulations. Thanks for share.\n>Rotate90/180/270 (100%)\n\nSo no free rotations at all. Do you think apply free rotations as I did could been perjudicial because the small size of the actual target, the motor?",
      "votes": 1,
      "replies": [
        {
          "id": 3217788,
          "postDate": "2025-06-05T11:59:13.693Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sacuscreed\" target=\"_blank\">@sacuscreed</a>, thanks for your comment. </p>\n<p>Its hard to say without validation, but some pipelines will be more susceptible to noise than others and maybe free rotations were too much for yours. I did not try them myself, but maybe I should have!</p>",
          "rawMarkdown": "Hi @sacuscreed, thanks for your comment. \n\nIts hard to say without validation, but some pipelines will be more susceptible to noise than others and maybe free rotations were too much for yours. I did not try them myself, but maybe I should have!\n\n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 3217707,
      "postDate": "2025-06-05T10:35:30.077Z",
      "content": "<p>Congratulations! Great solution.</p>",
      "rawMarkdown": "Congratulations! Great solution.",
      "votes": 1
    },
    {
      "id": 3217687,
      "postDate": "2025-06-05T10:02:04.257Z",
      "content": "<p>Hello Brendanartley, your work is awesome! I have some questions, I want to ask you, as follows:<br>\nDo you have prior knowledge of medical images, so can you correct the label manually?<br>\nDid you mention discarding negative samples because you think the results should focus more on recalls or because of experience?<br>\nThank you for your reply!</p>",
      "rawMarkdown": "Hello Brendanartley, your work is awesome! I have some questions, I want to ask you, as follows:\nDo you have prior knowledge of medical images, so can you correct the label manually?\nDid you mention discarding negative samples because you think the results should focus more on recalls or because of experience?\nThank you for your reply!",
      "votes": 1,
      "replies": [
        {
          "id": 3217815,
          "postDate": "2025-06-05T12:32:54.487Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/wym2024\" target=\"_blank\">@wym2024</a>, thanks for the comment.</p>\n<p>My only experience with tomograms or medical images comes from ML competitions like the CZII one <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification\" target=\"_blank\">here</a>. I discarded negative samples due to concerns about missed annotations, but in hindsight this had no impact on the private LB.</p>",
          "rawMarkdown": "Hi @wym2024, thanks for the comment.\n\nMy only experience with tomograms or medical images comes from ML competitions like the CZII one [here](https://www.kaggle.com/competitions/czii-cryo-et-object-identification). I discarded negative samples due to concerns about missed annotations, but in hindsight this had no impact on the private LB.",
          "replies": [
            {
              "id": 3218226,
              "postDate": "2025-06-06T02:09:40.787Z",
              "content": "<p>Thank you for your reply!</p>",
              "rawMarkdown": "Thank you for your reply!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3217642,
      "postDate": "2025-06-05T08:50:13.363Z",
      "content": "<p>Congratulations for the rank 1 solution and extra congratulations for sharing the extra data! I saw you get all tomographs at 128, 704, 704. How do you downsample the Z dim? if they are like 300 or 500 slices do you keep one every 3 for example (and i guess a bit more around the mottor?</p>",
      "rawMarkdown": "Congratulations for the rank 1 solution and extra congratulations for sharing the extra data! I saw you get all tomographs at 128, 704, 704. How do you downsample the Z dim? if they are like 300 or 500 slices do you keep one every 3 for example (and i guess a bit more around the mottor?",
      "votes": 1,
      "replies": [
        {
          "id": 3217903,
          "postDate": "2025-06-05T14:01:17.167Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/vasileioscharatsidis\" target=\"_blank\">@vasileioscharatsidis</a>. All tomograms were resized with <code>scipy.ndimage.zoom()</code>. I have updated the preprocessing code <a href=\"https://www.kaggle.com/code/brendanartley/flagellar-motors-dataset-code\" target=\"_blank\">here</a>.</p>",
          "rawMarkdown": "Thanks @vasileioscharatsidis. All tomograms were resized with `scipy.ndimage.zoom()`. I have updated the preprocessing code [here](https://www.kaggle.com/code/brendanartley/flagellar-motors-dataset-code).",
          "votes": 1
        }
      ]
    },
    {
      "id": 3217601,
      "postDate": "2025-06-05T07:28:23.803Z",
      "content": "<p>Congrats in your solo gold medal! I've seen you in past competitions and I knew you were going to make it soon! Happy Kaggling! </p>",
      "rawMarkdown": "Congrats in your solo gold medal! I've seen you in past competitions and I knew you were going to make it soon! Happy Kaggling! ",
      "votes": 1
    },
    {
      "id": 3217528,
      "postDate": "2025-06-05T05:27:33.027Z",
      "content": "<p>Congratulations. I'm so glad to see the winning solution isnt YOLO. Also, Hats off to you for the dataset you provided. Kudos to you again</p>",
      "rawMarkdown": "Congratulations. I'm so glad to see the winning solution isnt YOLO. Also, Hats off to you for the dataset you provided. Kudos to you again",
      "votes": 1
    },
    {
      "id": 3217527,
      "postDate": "2025-06-05T05:26:43.400Z",
      "content": "<p>Congratulations! <br>\nWhat patch size, label radius did you use?</p>",
      "rawMarkdown": "Congratulations! \nWhat patch size, label radius did you use?",
      "votes": 1,
      "replies": [
        {
          "id": 3217879,
          "postDate": "2025-06-05T13:30:26.997Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/siwooyong\" target=\"_blank\">@siwooyong</a>! </p>\n<p>The patch size was <code>(64, 704, 704)</code> and the kernel_size was 7 (on the 8x downsampled label). Here is a label overlaid on a tomogram to show the scale.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F752991d5f5f47bc8840420d915e55356%2Fsiwoo.jpg?generation=1749130192802729&amp;alt=media\" alt=\"Cropper\" style=\"max-width: 50%\"></p>",
          "rawMarkdown": "Thanks @siwooyong! \n\nThe patch size was `(64, 704, 704)` and the kernel_size was 7 (on the 8x downsampled label). Here is a label overlaid on a tomogram to show the scale.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F752991d5f5f47bc8840420d915e55356%2Fsiwoo.jpg?generation=1749130192802729&alt=media\" alt=\"Cropper\" style=\"max-width: 50%;\">",
          "votes": 1
        }
      ]
    },
    {
      "id": 3217504,
      "postDate": "2025-06-05T04:46:12.397Z",
      "content": "<p>Congratulations, PrivateLB has a slight improvement over PublicLB, I think it is due to your data processing and 3D model. Can the model run on Kaggle?</p>",
      "rawMarkdown": "Congratulations, PrivateLB has a slight improvement over PublicLB, I think it is due to your data processing and 3D model. Can the model run on Kaggle?",
      "votes": 1,
      "replies": [
        {
          "id": 3217759,
          "postDate": "2025-06-05T11:33:14.577Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ynhuhu\" target=\"_blank\">@ynhuhu</a>, thanks for the comment. </p>\n<p>The final model and patch size are too large to train on Kaggle, but there is more than enough time for inference.</p>",
          "rawMarkdown": "Hi @ynhuhu, thanks for the comment. \n\nThe final model and patch size are too large to train on Kaggle, but there is more than enough time for inference."
        }
      ]
    },
    {
      "id": 3217503,
      "postDate": "2025-06-05T04:43:36.137Z",
      "content": "<p>Thank you for a very good solution and having excellent generalizability, learning a lot!🥰</p>",
      "rawMarkdown": "Thank you for a very good solution and having excellent generalizability, learning a lot!🥰",
      "votes": 1
    },
    {
      "id": 3219746,
      "postDate": "2025-06-08T09:09:51.300Z",
      "content": "<p>Congratulations! Thank you for this external dataset, as a beginner in Data Science and Machine Learning, there was a lot of learning when reading this. </p>",
      "rawMarkdown": "Congratulations! Thank you for this external dataset, as a beginner in Data Science and Machine Learning, there was a lot of learning when reading this. ",
      "votes": 2
    },
    {
      "id": 3217704,
      "postDate": "2025-06-05T10:30:29.147Z",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> Congratulations on the 1st place!  <br>\nYour early contribution—especially sharing the external dataset—was incredibly helpful for the whole community. You truly deserved this win.</p>",
      "rawMarkdown": "@brendanartley Congratulations on the 1st place!  \nYour early contribution—especially sharing the external dataset—was incredibly helpful for the whole community. You truly deserved this win.\n",
      "votes": 2
    },
    {
      "id": 3217667,
      "postDate": "2025-06-05T09:26:48.030Z",
      "content": "<p>Nice! Simple and elegant solution) Well deserved medal 💪<br>\nSeparate kudos for the external data! </p>",
      "rawMarkdown": "Nice! Simple and elegant solution) Well deserved medal 💪\nSeparate kudos for the external data! ",
      "votes": 2,
      "replies": [
        {
          "id": 3217770,
          "postDate": "2025-06-05T11:46:48.097Z",
          "content": "<p>Cheers <a href=\"https://www.kaggle.com/bloodaxe\" target=\"_blank\">@bloodaxe</a>. </p>\n<p>I got some great tips from your solution in CZII. Thanks for all that you shared in that competition!</p>",
          "rawMarkdown": "Cheers @bloodaxe. \n\nI got some great tips from your solution in CZII. Thanks for all that you shared in that competition!",
          "votes": 1
        }
      ]
    },
    {
      "id": 3217598,
      "postDate": "2025-06-05T07:25:02.410Z",
      "content": "<p>Congrats !<br>\nI have just 1 question what software are you using to generate those images to explain the approach?</p>",
      "rawMarkdown": "Congrats !\nI have just 1 question what software are you using to generate those images to explain the approach?",
      "votes": 2,
      "replies": [
        {
          "id": 3217826,
          "postDate": "2025-06-05T12:41:56.097Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/justforfun44\" target=\"_blank\">@justforfun44</a>! I used <a href=\"https://draw.io/\" target=\"_blank\">draw.io</a> for the figures.</p>",
          "rawMarkdown": "Thanks @justforfun44! I used [draw.io](https://draw.io/) for the figures.",
          "votes": 2
        }
      ]
    },
    {
      "id": 3217516,
      "postDate": "2025-06-05T05:08:35.420Z",
      "content": "<p>Can’t wait to dive into this further! So glad that the winning solution is more unique than what feels like a random mix of data inputted to a yolo model! Well deserved! </p>",
      "rawMarkdown": "Can’t wait to dive into this further! So glad that the winning solution is more unique than what feels like a random mix of data inputted to a yolo model! Well deserved! ",
      "votes": 2
    },
    {
      "id": 3217505,
      "postDate": "2025-06-05T04:47:40.367Z",
      "content": "<p>Congratulations! I am new to this field and have a question about training. How did you train the 3D U-Net with this data? I have also started training, but, it exceeded the allocated time.</p>",
      "rawMarkdown": "Congratulations! I am new to this field and have a question about training. How did you train the 3D U-Net with this data? I have also started training, but, it exceeded the allocated time.",
      "votes": 2,
      "replies": [
        {
          "id": 3217765,
          "postDate": "2025-06-05T11:39:30.063Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/rustambazarbayev\" target=\"_blank\">@rustambazarbayev</a>, thanks for your comment. I ran experiments on a machine outside of the Kaggle environment.</p>",
          "rawMarkdown": "Hi @rustambazarbayev, thanks for your comment. I ran experiments on a machine outside of the Kaggle environment."
        }
      ]
    },
    {
      "id": 3217502,
      "postDate": "2025-06-05T04:36:35.447Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>! Thank you for defeating the devil.</p>",
      "rawMarkdown": "Congrats @brendanartley! Thank you for defeating the devil.",
      "votes": 2
    },
    {
      "id": 3217491,
      "postDate": "2025-06-05T04:24:40.327Z",
      "content": "<p>Congrats and really happy to see you at the 1st place! <br>\nI'm also using your external dataset with some modifications. It was very kind of you for sharing the dataset and download notebook at the early stage of the competition. Started download my own version and use it a week ago, now I regret I did not using it earlier :3</p>",
      "rawMarkdown": "Congrats and really happy to see you at the 1st place! \nI'm also using your external dataset with some modifications. It was very kind of you for sharing the dataset and download notebook at the early stage of the competition. Started download my own version and use it a week ago, now I regret I did not using it earlier :3",
      "votes": 2,
      "replies": [
        {
          "id": 3217756,
          "postDate": "2025-06-05T11:29:23.207Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/dangnh0611\" target=\"_blank\">@dangnh0611</a>!</p>",
          "rawMarkdown": "Thanks @dangnh0611!"
        }
      ]
    },
    {
      "id": 3247592,
      "postDate": "2025-07-13T04:15:40.363Z",
      "content": "<p>Hello, I would like to know whether the division in cross-validation is based on certain criteria or is random.</p>",
      "rawMarkdown": "Hello, I would like to know whether the division in cross-validation is based on certain criteria or is random."
    },
    {
      "id": 3222460,
      "postDate": "2025-06-12T07:55:30.667Z",
      "content": "<p>Intersted in the same</p>",
      "rawMarkdown": "Intersted in the same"
    },
    {
      "id": 3222458,
      "postDate": "2025-06-12T07:54:22.457Z",
      "content": "<p>3D U-Net + Quantile Thresholding intersted</p>",
      "rawMarkdown": "3D U-Net + Quantile Thresholding intersted"
    },
    {
      "id": 3247582,
      "postDate": "2025-07-13T03:35:11.587Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3217686,
      "postDate": "2025-06-05T09:59:23.937Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3224759,
      "postDate": "2025-06-15T12:24:56.563Z",
      "content": "<p>Congratulations,👏 Thanks for sharing✨</p>",
      "rawMarkdown": "Congratulations,👏 Thanks for sharing✨",
      "votes": 1
    },
    {
      "id": 3223172,
      "postDate": "2025-06-13T00:56:57.977Z",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": 1
    },
    {
      "id": 3219718,
      "postDate": "2025-06-08T08:35:57.427Z",
      "content": "<p>Thank you for the nice visualizations!</p>",
      "rawMarkdown": "Thank you for the nice visualizations!",
      "votes": 1
    },
    {
      "id": 3219081,
      "postDate": "2025-06-07T06:32:52.233Z",
      "content": "<p>Congratulations !</p>",
      "rawMarkdown": "Congratulations !",
      "votes": 1
    },
    {
      "id": 3218718,
      "postDate": "2025-06-06T16:01:37.093Z",
      "content": "<p>👏🏻 Congrats! Thanks for sharing.</p>",
      "rawMarkdown": "👏🏻 Congrats! Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 3218379,
      "postDate": "2025-06-06T05:51:18.380Z",
      "content": "<p>good, thanks you</p>",
      "rawMarkdown": "good, thanks you\n",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 3217533,
      "author_name": "Jeroen Cottaar",
      "author_url": "",
      "post_date": "2025-06-05T05:39:20.967000",
      "content": "<p>Congratulations, and thanks for the detailed writeup!</p>\n<p>I actually spent my first month on this competition on a very similar 3D UNet approach, building from my CZII competition solution. A very simple first version already got to ~0.85 CV in the first few days, but never scored above 0.2 on the leaderboard (and usually 0.0). I spent a month trying to figure out what was going on, but failed and fell back to YOLO.</p>\n<p>Did you face anything similar at any point? I really want to know what was going on…</p>",
      "votes": 7,
      "replies": [
        {
          "id": 3217663,
          "author_name": "IAmParadox",
          "author_url": "",
          "post_date": "2025-06-05T09:16:29.337000",
          "content": "<p>Yup, myself and <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> as well faced similar issues independently while trying to work with 3D-UNet. The best I could get on the public LB was around ~0.22 and for sersasj it was around ~0.5. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3217683,
          "author_name": "wym2024",
          "author_url": "",
          "post_date": "2025-06-05T09:55:16.140000",
          "content": "<p>Coincidentally, I encountered this similar problem in YOLOv10. I found that parameter adjustments such as conf_thr, iou have a great impact on LB. For example, I reduced conf_thr from 0.7 to 0.4, and lb dropped 4pp, which also confused me very much. In theory, it is unlikely that the parameter adjustment will have such a large fluctuation.<br>\nI attribute it to the sparse data, and I did not check the specific reasons.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3217730,
          "author_name": "MLArt",
          "author_url": "",
          "post_date": "2025-06-05T11:06:17.570000",
          "content": "<p>I have the same problem, local cv scores more than 0.95 on sliced inference with 3d unet but max public score is 0.68</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3217810,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-06-05T12:21:24.893000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/Jeroen\" target=\"_blank\">@Jeroen</a>,</p>\n<p>Its hard to say - my first submissions with a resnet18 encoder scored in the 0.5 - 0.7 range.  Some key things that helped boost performance from there were:</p>\n<ol>\n<li>Deeper/larger encoders</li>\n<li>Large patch sizes (that covered &gt;85% of the height and width)</li>\n<li>High resolution tomograms</li>\n<li>&gt;50% depth overlap during inference</li>\n</ol>\n<p>Were you using a pretrained encoder?</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3218995,
      "author_name": "Sadeep Dilshan Kasthuriarachchi",
      "author_url": "",
      "post_date": "2025-06-07T03:30:36.697000",
      "content": "<p>Congratulations!!, and particularly thanks for your effort in providing the external dataset—made a significant impact on the entire community. This recognition is well-deserved, and your contribution truly set the stage for success.❕</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3271444,
      "author_name": "Lonnie Wibberding",
      "author_url": "",
      "post_date": "2025-08-18T23:23:26.723000",
      "content": "<p>Thanks for the write up and congratulations on the win. As someone fairly new to data science I appreciate people like you who share their techniques to help me learn faster.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3229334,
      "author_name": "Bhavya Garg",
      "author_url": "",
      "post_date": "2025-06-21T10:28:31.383000",
      "content": "<p>Congratulations!! A lot to learn from you🙂</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3223244,
      "author_name": "Guillermo Perez G",
      "author_url": "",
      "post_date": "2025-06-13T04:07:11.933000",
      "content": "<p>Congratulations! Your merit is greater for competing alone, and with such close scores. I wonder if it's possible to experiment with your work in a Jupyter notebook and Anaconda?  <br>\nAll the best</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3223200,
      "author_name": "Charvak Upadhyay",
      "author_url": "",
      "post_date": "2025-06-13T02:26:35.490000",
      "content": "<p>Massive congrats on the 1st place finish—what a brilliant piece of work. The 3D U-Net, all that heavy augmentation, and the smart use of quantile thresholding really paid off. well deserved and super inspiring.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3222828,
      "author_name": "Champ",
      "author_url": "",
      "post_date": "2025-06-12T14:57:19.640000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> on winning and thanks for the writeup!<br>\nI have a question regarding the quantile thresholding: What exactly is the advantage in your opinion? Because I am not sure I have understood it completly and as far as I understand it, you still have to tune the threshold and the threshold would still change depending on the used model right?<br>\nIs it because a quantile threshold gives somewhat of a more smooth threshold on how many predictions are counted?</p>\n<p>It will probably be clear, when your inference code is released but still wanted to ask.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3222334,
      "author_name": "tenzy123",
      "author_url": "",
      "post_date": "2025-06-12T06:16:17.510000",
      "content": "<p>Thanks for this solution,its really helpful.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3222121,
      "author_name": "Victor",
      "author_url": "",
      "post_date": "2025-06-11T21:10:19.610000",
      "content": "<p>Hi,<br>\nThe aux head shouldn't be applied on the last feature map of the encoder as you only use one decoder block?<br>\nIn the <strong>Loss</strong> image you shared, it is applied on the decoder block instead.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3220587,
      "author_name": "alex5051",
      "author_url": "",
      "post_date": "2025-06-09T14:56:57.667000",
      "content": "<p>Congrutulations,thanks for your solution.It's amazing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3220050,
      "author_name": "Jiahao Liu",
      "author_url": "",
      "post_date": "2025-06-08T17:53:04.080000",
      "content": "<p>Congratulations, and thanks for detailed writeup…!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3218727,
      "author_name": "Sarah Arshad",
      "author_url": "",
      "post_date": "2025-06-06T16:12:07.793000",
      "content": "<p>Congratulations, and thanks for detailed writeup…!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3218188,
      "author_name": "Caleb Yenusah",
      "author_url": "",
      "post_date": "2025-06-06T00:50:25.983000",
      "content": "<p>Congrats! Is it possible to share this solution's training and inference source code?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3218255,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-06-06T03:06:10.907000",
          "content": "<p>Cheers <a href=\"https://www.kaggle.com/calebyenusah\" target=\"_blank\">@calebyenusah</a>. </p>\n<p>I just released the training code <a href=\"https://github.com/brendanartley/BYU-competition\" target=\"_blank\">here</a>. The inference pipeline is coming soon.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3301015,
              "author_name": "Sweksha Sinha",
              "author_url": "",
              "post_date": "2025-10-12T09:05:06.603000",
              "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> would love to get the inference pipeline to understand your approach to the problem</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3217850,
      "author_name": "Cyrus",
      "author_url": "",
      "post_date": "2025-06-05T13:00:04.373000",
      "content": "<p>Congrats! I think the thresholding trick you have done has a lot of weight in the final results, you did a great job nailing that! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3217717,
      "author_name": "tennogh",
      "author_url": "",
      "post_date": "2025-06-05T10:54:29.667000",
      "content": "<p>Hi and congrats, I have some questions seeing as I used a similar approach but didn't get past .7 LB:</p>\n<ul>\n<li>Did you use all of the external data? I found that the tomograms with the smallest voxel spacings hindered my models significantly</li>\n<li>What crop size did you use for training? I tried to go as large as possible but I think it was a mistake.</li>\n<li>How long did training and especially validation take? I had to reduce the size of my validation sets because it was difficult to train for a lot of epochs</li>\n</ul>\n<p>Anectodally I estimated that around 35% of the test samples were positive, and settled for a threshold around the 58th percentile.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3217820,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-06-05T12:39:46.897000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tennogh\" target=\"_blank\">@tennogh</a>,</p>\n<ol>\n<li><p>I used all tomograms.</p></li>\n<li><p>The crop size was <code>(64, 674, 674)</code> - I think you were on the right track with this.</p></li>\n<li><p>Small models trained in 2 - 8 hours, while the final ensemble models took 35 - 40 hours each.</p></li>\n</ol>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3217708,
      "author_name": "Ángel Jacinto Sánchez Ruiz",
      "author_url": "",
      "post_date": "2025-06-05T10:36:39.920000",
      "content": "<p>Congratulations. Thanks for share.</p>\n<blockquote>\n  <p>Rotate90/180/270 (100%)</p>\n</blockquote>\n<p>So no free rotations at all. Do you think apply free rotations as I did could been perjudicial because the small size of the actual target, the motor?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3217788,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-06-05T11:59:13.693000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sacuscreed\" target=\"_blank\">@sacuscreed</a>, thanks for your comment. </p>\n<p>Its hard to say without validation, but some pipelines will be more susceptible to noise than others and maybe free rotations were too much for yours. I did not try them myself, but maybe I should have!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3217707,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2025-06-05T10:35:30.077000",
      "content": "<p>Congratulations! Great solution.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3217687,
      "author_name": "wym2024",
      "author_url": "",
      "post_date": "2025-06-05T10:02:04.257000",
      "content": "<p>Hello Brendanartley, your work is awesome! I have some questions, I want to ask you, as follows:<br>\nDo you have prior knowledge of medical images, so can you correct the label manually?<br>\nDid you mention discarding negative samples because you think the results should focus more on recalls or because of experience?<br>\nThank you for your reply!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3217815,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-06-05T12:32:54.487000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/wym2024\" target=\"_blank\">@wym2024</a>, thanks for the comment.</p>\n<p>My only experience with tomograms or medical images comes from ML competitions like the CZII one <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification\" target=\"_blank\">here</a>. I discarded negative samples due to concerns about missed annotations, but in hindsight this had no impact on the private LB.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3218226,
              "author_name": "wym2024",
              "author_url": "",
              "post_date": "2025-06-06T02:09:40.787000",
              "content": "<p>Thank you for your reply!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3217642,
      "author_name": "Vasilis",
      "author_url": "",
      "post_date": "2025-06-05T08:50:13.363000",
      "content": "<p>Congratulations for the rank 1 solution and extra congratulations for sharing the extra data! I saw you get all tomographs at 128, 704, 704. How do you downsample the Z dim? if they are like 300 or 500 slices do you keep one every 3 for example (and i guess a bit more around the mottor?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3217903,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-06-05T14:01:17.167000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/vasileioscharatsidis\" target=\"_blank\">@vasileioscharatsidis</a>. All tomograms were resized with <code>scipy.ndimage.zoom()</code>. I have updated the preprocessing code <a href=\"https://www.kaggle.com/code/brendanartley/flagellar-motors-dataset-code\" target=\"_blank\">here</a>.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3217601,
      "author_name": "Carlos Pérez Ricardo",
      "author_url": "",
      "post_date": "2025-06-05T07:28:23.803000",
      "content": "<p>Congrats in your solo gold medal! I've seen you in past competitions and I knew you were going to make it soon! Happy Kaggling! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3217528,
      "author_name": "NarayanNarayan",
      "author_url": "",
      "post_date": "2025-06-05T05:27:33.027000",
      "content": "<p>Congratulations. I'm so glad to see the winning solution isnt YOLO. Also, Hats off to you for the dataset you provided. Kudos to you again</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3217527,
      "author_name": "siwooyong",
      "author_url": "",
      "post_date": "2025-06-05T05:26:43.400000",
      "content": "<p>Congratulations! <br>\nWhat patch size, label radius did you use?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3217879,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-06-05T13:30:26.997000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/siwooyong\" target=\"_blank\">@siwooyong</a>! </p>\n<p>The patch size was <code>(64, 704, 704)</code> and the kernel_size was 7 (on the 8x downsampled label). Here is a label overlaid on a tomogram to show the scale.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F752991d5f5f47bc8840420d915e55356%2Fsiwoo.jpg?generation=1749130192802729&amp;alt=media\" alt=\"Cropper\" style=\"max-width: 50%\"></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3217504,
      "author_name": "ynhuhu",
      "author_url": "",
      "post_date": "2025-06-05T04:46:12.397000",
      "content": "<p>Congratulations, PrivateLB has a slight improvement over PublicLB, I think it is due to your data processing and 3D model. Can the model run on Kaggle?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3217759,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-06-05T11:33:14.577000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ynhuhu\" target=\"_blank\">@ynhuhu</a>, thanks for the comment. </p>\n<p>The final model and patch size are too large to train on Kaggle, but there is more than enough time for inference.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3217503,
      "author_name": "yyyy0201",
      "author_url": "",
      "post_date": "2025-06-05T04:43:36.137000",
      "content": "<p>Thank you for a very good solution and having excellent generalizability, learning a lot!🥰</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3219746,
      "author_name": "aaaditya56",
      "author_url": "",
      "post_date": "2025-06-08T09:09:51.300000",
      "content": "<p>Congratulations! Thank you for this external dataset, as a beginner in Data Science and Machine Learning, there was a lot of learning when reading this. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3217704,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2025-06-05T10:30:29.147000",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> Congratulations on the 1st place!  <br>\nYour early contribution—especially sharing the external dataset—was incredibly helpful for the whole community. You truly deserved this win.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3217667,
      "author_name": "Eugene Khvedchenya",
      "author_url": "",
      "post_date": "2025-06-05T09:26:48.030000",
      "content": "<p>Nice! Simple and elegant solution) Well deserved medal 💪<br>\nSeparate kudos for the external data! </p>",
      "votes": 2,
      "replies": [
        {
          "id": 3217770,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-06-05T11:46:48.097000",
          "content": "<p>Cheers <a href=\"https://www.kaggle.com/bloodaxe\" target=\"_blank\">@bloodaxe</a>. </p>\n<p>I got some great tips from your solution in CZII. Thanks for all that you shared in that competition!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3217598,
      "author_name": "Shapu",
      "author_url": "",
      "post_date": "2025-06-05T07:25:02.410000",
      "content": "<p>Congrats !<br>\nI have just 1 question what software are you using to generate those images to explain the approach?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3217826,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-06-05T12:41:56.097000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/justforfun44\" target=\"_blank\">@justforfun44</a>! I used <a href=\"https://draw.io/\" target=\"_blank\">draw.io</a> for the figures.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3217516,
      "author_name": "Cody_Null",
      "author_url": "",
      "post_date": "2025-06-05T05:08:35.420000",
      "content": "<p>Can’t wait to dive into this further! So glad that the winning solution is more unique than what feels like a random mix of data inputted to a yolo model! Well deserved! </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3217505,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-05T04:47:40.367000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 3217765,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-06-05T11:39:30.063000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3217502,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-05T04:36:35.447000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3217491,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-05T04:24:40.327000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 3217756,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-06-05T11:29:23.207000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3247592,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-07-13T04:15:40.363000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3222460,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-12T07:55:30.667000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3222458,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-12T07:54:22.457000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3247582,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-07-13T03:35:11.587000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3217686,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-05T09:59:23.937000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3224759,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-15T12:24:56.563000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3223172,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-13T00:56:57.977000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3219718,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-08T08:35:57.427000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3219081,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-07T06:32:52.233000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3218718,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-06T16:01:37.093000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3218379,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-06T05:51:18.380000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3217480": "Thanks to BYU and Kaggle for hosting this competition. It was nice to have another well-run tomography competition and the hosts were awesome. I can't believe the result!\n\n## TLDR\n\nMy solution uses a 3D U-Net trained with heavy augmentations and auxiliary loss functions. During inference, I rank each tomogram based on the max predicted pixel value and use quantile thresholding to determine if a motor is present.\n\n## Cross Validation\n\nFor validating models, the competition data is split into 4 folds. Local CV strongly correlates with the LB up to about 0.93. Beyond that, I used the public LB for validation. It was important to use quantile thresholding to get reliable feedback from the LB. More on this in the post-processing section.\n\n## Preprocessing\n\nTomograms from the competition data and the CryoET Data Portal are used to create a training set. Each tomogram is resized to (128, 704, 704) using `scipy.ndimage.zoom()`, and tomograms without motors are discarded. As others noted, the competition data is quite noisy, so Napari was used to manually add missing motors. I will add the updated data [here](https://www.kaggle.com/datasets/brendanartley/cryoet-flagellar-motors-dataset).\n\nFor the labels, I use a Gaussian heat map centered on each motor. Similar to @bloodaxe and @christofhenkel's solution in the CZII competition, the resolution of the heatmap is reduced by 8x. This works especially well for this competition as there is a high tolerance for distance error in the metric. This means that predicting the exact pixel is not as important as predicting motor presence. If you are not convinced, the following plot shows roughly how much error is allowed around each motor when voxel spacing equals 10.\n\n![tomogram_image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F406663dc00b116aa614b191abd71b37f%2Ftomograms.JPG?generation=1749094386913073&alt=media)\n\n## Model\n\nThe model is a 3D U-Net (sort of). The encoder is a pre-trained ResNet200 from Kenoshara’s repository [here](https://github.com/kenshohara/3D-ResNets-PyTorch). For most experiments, I used the ResNet101 variant, but increasing the capacity of the encoder yields better performance. In addition, stochastic dropout is applied for regularization, and gradient checkpointing is used to reduce vRAM usage during training. The decoder uses a single deconvolution block before the segmentation head.\n\n![model_image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F9242bdbaf78089afc0842c382a4e40a2%2FBYU2025-Model%20(2).jpg?generation=1749095150932706&alt=media)\n\n## Loss\n\nThe model is trained using SmoothBCE loss with 3 contributions. The main segmentation head predicts the output logits, a deep supervision head is applied to the second last feature map, and a max pooled loss (kernel size and stride of 4) is applied on the main segmentation head. Moreover, the pooled loss encourages high probabilities around the motor region, while reducing the penalty for small localization errors.\n\n![loss_image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fcedd954d5796ada89d6dd3749b00afa2%2FBYU2025-Loss%20(1).jpg?generation=1751317472394799&alt=media)\n\n## Augmentations\n\nHeavy augmentations enabled training for 400 epochs without overfitting. Although, I could probably have trained longer there was no change in the public LB scores beyond 250 epochs.\n\n- Mixup (100%)\n- Rescale/Zoom (100%)\n- Rotate90/180/270 (100%)\n- Axis Flips (100%)\n- Axis Swap (100%)\n- Coarse Dropout (50%)\n- Color inversion (25%)\n- Simple Cutmix (15%)\n\nLoading tomograms from disk is slow, which limits the time for augmentations on the CPU. To address this, all augmentations but rescaling are applied on the GPU. To keep rescaling as fast as possible `scipy.ndimage.zoom(..., order=0)` is used. \n\n## Inference\n\nInitially, the same preprocessing pipeline was applied during inference. This worked well, but it was 4x faster to match the patch height and width, and only slide over the depth. This allows more time for TTA and a very high overlap (0.875). Both approaches scored about the same, but my final solution uses the latter.\n\nAll edge predictions are down-weighted using the `roi_weight_map` parameter. The middle 40% of the logits are weighted as 1.0 and other logits are weighted as 0.001 when aggregating the sliding window.\n\n## Ensembling\n\nThe final submission uses an 8-seed ensemble. Sigmoid is applied to each model output and the logits are summed. Inference takes ~10 hrs.\n\n## Postprocessing\n\nLike many others, I found that fixed thresholds were unstable. Instead, I use quantile thresholding to determine motor presence.\n\nTo apply this, all tomograms are ranked based on their max predicted pixel value. Then, predictions for the lowest quantile are removed. I tuned the quantile on the public LB and then prayed to the Kaggle gods that the private LB was similar. On the public LB the optimal threshold was 0.565 and on private it was 0.560. \n\n![LB_image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F6a05822baf7e6d70c6f8d932aabb51b0%2Fscores.JPG?generation=1749094657707643&alt=media)\n\n## Final Note\n\nThanks for reading, and thanks to everyone who showed their appreciation for the external dataset. \n\nExternal data [here](https://www.kaggle.com/datasets/brendanartley/cryoet-flagellar-motors-dataset)\nGithub repository [here](https://github.com/brendanartley/BYU-competition)\nMetadata [here](https://www.kaggle.com/datasets/brendanartley/solution-ds-byu-1st-place-metadata/data)\n\nHappy Kaggling!",
    "3217533": "Congratulations, and thanks for the detailed writeup!\n\nI actually spent my first month on this competition on a very similar 3D UNet approach, building from my CZII competition solution. A very simple first version already got to ~0.85 CV in the first few days, but never scored above 0.2 on the leaderboard (and usually 0.0). I spent a month trying to figure out what was going on, but failed and fell back to YOLO.\n\nDid you face anything similar at any point? I really want to know what was going on...",
    "3218995": "Congratulations!!, and particularly thanks for your effort in providing the external dataset—made a significant impact on the entire community. This recognition is well-deserved, and your contribution truly set the stage for success.❕",
    "3271444": "Thanks for the write up and congratulations on the win. As someone fairly new to data science I appreciate people like you who share their techniques to help me learn faster.",
    "3229334": "Congratulations!! A lot to learn from you🙂",
    "3223244": "Congratulations! Your merit is greater for competing alone, and with such close scores. I wonder if it's possible to experiment with your work in a Jupyter notebook and Anaconda?  \nAll the best",
    "3223200": "Massive congrats on the 1st place finish—what a brilliant piece of work. The 3D U-Net, all that heavy augmentation, and the smart use of quantile thresholding really paid off. well deserved and super inspiring.",
    "3222828": "Congrats @brendanartley on winning and thanks for the writeup!\nI have a question regarding the quantile thresholding: What exactly is the advantage in your opinion? Because I am not sure I have understood it completly and as far as I understand it, you still have to tune the threshold and the threshold would still change depending on the used model right?\nIs it because a quantile threshold gives somewhat of a more smooth threshold on how many predictions are counted?\n\nIt will probably be clear, when your inference code is released but still wanted to ask.",
    "3222334": "Thanks for this solution,its really helpful.",
    "3222121": "Hi,\nThe aux head shouldn't be applied on the last feature map of the encoder as you only use one decoder block?\nIn the **Loss** image you shared, it is applied on the decoder block instead.",
    "3220587": "Congrutulations,thanks for your solution.It's amazing.",
    "3220050": "Congratulations, and thanks for detailed writeup…!",
    "3218727": "Congratulations, and thanks for detailed writeup...!",
    "3218188": "Congrats! Is it possible to share this solution's training and inference source code?",
    "3217850": "Congrats! I think the thresholding trick you have done has a lot of weight in the final results, you did a great job nailing that! ",
    "3217717": "Hi and congrats, I have some questions seeing as I used a similar approach but didn't get past .7 LB:\n\n- Did you use all of the external data? I found that the tomograms with the smallest voxel spacings hindered my models significantly\n- What crop size did you use for training? I tried to go as large as possible but I think it was a mistake.\n- How long did training and especially validation take? I had to reduce the size of my validation sets because it was difficult to train for a lot of epochs\n\nAnectodally I estimated that around 35% of the test samples were positive, and settled for a threshold around the 58th percentile.",
    "3217708": "Congratulations. Thanks for share.\n>Rotate90/180/270 (100%)\n\nSo no free rotations at all. Do you think apply free rotations as I did could been perjudicial because the small size of the actual target, the motor?",
    "3217707": "Congratulations! Great solution.",
    "3217687": "Hello Brendanartley, your work is awesome! I have some questions, I want to ask you, as follows:\nDo you have prior knowledge of medical images, so can you correct the label manually?\nDid you mention discarding negative samples because you think the results should focus more on recalls or because of experience?\nThank you for your reply!",
    "3217642": "Congratulations for the rank 1 solution and extra congratulations for sharing the extra data! I saw you get all tomographs at 128, 704, 704. How do you downsample the Z dim? if they are like 300 or 500 slices do you keep one every 3 for example (and i guess a bit more around the mottor?",
    "3217601": "Congrats in your solo gold medal! I've seen you in past competitions and I knew you were going to make it soon! Happy Kaggling! ",
    "3217528": "Congratulations. I'm so glad to see the winning solution isnt YOLO. Also, Hats off to you for the dataset you provided. Kudos to you again",
    "3217527": "Congratulations! \nWhat patch size, label radius did you use?",
    "3217504": "Congratulations, PrivateLB has a slight improvement over PublicLB, I think it is due to your data processing and 3D model. Can the model run on Kaggle?",
    "3217503": "Thank you for a very good solution and having excellent generalizability, learning a lot!🥰",
    "3219746": "Congratulations! Thank you for this external dataset, as a beginner in Data Science and Machine Learning, there was a lot of learning when reading this. ",
    "3217704": "@brendanartley Congratulations on the 1st place!  \nYour early contribution—especially sharing the external dataset—was incredibly helpful for the whole community. You truly deserved this win.\n",
    "3217667": "Nice! Simple and elegant solution) Well deserved medal 💪\nSeparate kudos for the external data! ",
    "3217598": "Congrats !\nI have just 1 question what software are you using to generate those images to explain the approach?",
    "3217516": "Can’t wait to dive into this further! So glad that the winning solution is more unique than what feels like a random mix of data inputted to a yolo model! Well deserved! ",
    "3217505": "Congratulations! I am new to this field and have a question about training. How did you train the 3D U-Net with this data? I have also started training, but, it exceeded the allocated time.",
    "3217502": "Congrats @brendanartley! Thank you for defeating the devil.",
    "3217491": "Congrats and really happy to see you at the 1st place! \nI'm also using your external dataset with some modifications. It was very kind of you for sharing the dataset and download notebook at the early stage of the competition. Started download my own version and use it a week ago, now I regret I did not using it earlier :3",
    "3247592": "Hello, I would like to know whether the division in cross-validation is based on certain criteria or is random.",
    "3222460": "Intersted in the same",
    "3222458": "3D U-Net + Quantile Thresholding intersted",
    "3247582": "",
    "3217686": "",
    "3224759": "Congratulations,👏 Thanks for sharing✨",
    "3223172": "Thank you!",
    "3219718": "Thank you for the nice visualizations!",
    "3219081": "Congratulations !",
    "3218718": "👏🏻 Congrats! Thanks for sharing.",
    "3218379": "good, thanks you\n"
  }
}