{
  "id": 561401,
  "title": "4th Place Solution [Source Codes & Submission Notebook Released!]",
  "url": "/competitions/czii-cryo-et-object-identification/writeups/yu4u-tattaka-4th-place-solution-source-codes-submi",
  "author_name": "",
  "post_date": "2025-02-09T23:36:48.063Z",
  "votes": 68,
  "comment_count": 16,
  "views": 0,
  "content": "<p>We would first like to express our gratitude to the competition host and the Kaggle staff for organizing this outstanding competition. Below, we introduce the solution of Team yu4u &amp; tattaka.</p>\n<h1>Summary</h1>\n<p>We adopted an approach to detect particle points using a heatmap-based method, which is the most commonly employed technique in pose estimation and facial keypoint detection.<br>\nSince this competition deals with 3D images rather than 2D images, we utilized two types of UNet-like models (yu4u's model and tattaka's model) that take 3D voxels as input and outputs 3D heatmaps.</p>\n<h1>Our Approach to This Competition</h1>\n<p>First, we will explain our approach to this competition, specifically how we addressed the issue of CV and LB not correlating. We used CV only to confirm that the metric produced was somewhat reasonable and for selecting checkpoints. For models with potential for improvement, we simply submitted them and relied on the LB to make decisions about which methods to adopt or discard.</p>\n<h1>Creating the Ground Truth Heatmap</h1>\n<p>We generate the ground truth heatmap necessary for model training. This involves converting the ground truth particle coordinates into the pixel coordinate system and creating a mask using a Gaussian function, where the particle center is set to 1.0 and sigma is 6 pixels for yu4u's model. For tattaka's model, different sigma values were used for different particles based on their sizes.<br>\nWe believe that an offset of 1.0 should be added when converting particle coordinates into the pixel coordinate system. While <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/553126\" target=\"_blank\">this discussion</a> suggests adding 0.5, <a href=\"https://www.kaggle.com/code/ren4yu/czii-coordinate-eda\" target=\"_blank\">our notebook</a> demonstrates that 1.0 is the correct value. The main difference is that the previous discussion assumes the particle center is at the top-left of a pixel, whereas we argue that, on average, the circle should be drawn from the pixel center (0.5, 0.5).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2F01f088d7b84cf5710d322fe17e7aaf5a%2F2025-02-06%2011.30.19.png?generation=1738809090496377&amp;alt=media\" alt=\"\"></p>\n<h1>yu4u's Model</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2F7fad2afcd6dfd4b83065d9692ecb8479%2Fmodel.png?generation=1738807875765144&amp;alt=media\" alt=\"\"></p>\n<p>We adopted a 2.5D-UNet, which utilizes a 2D image-based model as the backbone. The outputs from each stage of this backbone are pooled along the depth direction, enabling hierarchical feature extraction in the depth dimension as well. This idea was borrowed from <a href=\"https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder\" target=\"_blank\">the excellent notebook</a>. An interesting observation is that replacing this pooling operation with strided 3D convolutions degrades performance. This would be because the pooling method effectively aggregates depth features while preserving the original 2D backbone’s feature maps as much as possible. Similar to many other Kaggle competitions dealing with 3D data, a UNet utilizing a 2D backbone outperformed a straightforward UNet with a 3D backbone.</p>\n<p>We also applied 3D convolution between the encoder and decoder, inspired by <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430685\" target=\"_blank\">the 3rd Place Solution of the contrails competition</a>.</p>\n<p>Initially, we used a plain UNet architecture, but processing high-resolution feature maps required significant memory and computation. To address this, we adopted a model that outputs the final heatmap using pixel shuffle from a feature map with a stride of 4. Pixel shuffle, also known as depth_to_space in TensorFlow, is an operation that redistributes information from the channel dimension to the spatial dimensions. Compared to deconvolution, it offers advantages in computational efficiency and reducing artifacts.</p>\n<p>For the final submission, we used four models with different folds of a ConvNeXt Nano model as the backbone.</p>\n<h1>tattaka's Model</h1>\n<h2>Model Architecture</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2Fe15184020a0b9ea5e2ff58bf5d17ce25%2F2025-02-05%2023.11.47.png?generation=1738799891700544&amp;alt=media\" alt=\"\"></p>\n<p>This model is a lightweight 2.5D UNet with ResNetRS-50 as the backbone.  </p>\n<p>The input to the model is a volume of size 32×128×128 (D×H×W), and it outputs a 3D heatmap of the same size. Within the backbone, the depth is progressively reduced by half using average pooling for the first two stages. After that, average pooling with kernel=3, stride=1, padding=1 is used to maintain the depth while computing feature maps. As a result, the feature map shapes at each stage of the backbone are as follows:  <br>\n<code>(bs, ch, 16, 64, 64)</code>, <code>(bs, ch, 8, 32, 32)</code>, <code>(bs, ch, 8, 16, 16)</code>, <code>(bs, ch, 8, 8, 8)</code>, <code>(bs, ch, 8, 4, 4)</code>.  </p>\n<p>In the decoder, the three lowest-resolution feature maps are fed into <a href=\"https://arxiv.org/abs/1903.11816\" target=\"_blank\">Joint Pyramid Upsampling</a>. These maps are then progressively upsampled using 3D CNNs, SESC attention, and upsampling layers until they reach the same size as the input volume.</p>\n<h2>Loss Function</h2>\n<p>Since the number of particles within the volume is relatively small, there is a significant class imbalance between positive and negative samples during training. We attempted to adjust the parameters for generating ground truth heatmaps, but this did not lead to any improvement in cross-validation performance.  </p>\n<p>Ultimately, we implemented a simple MSE-based loss function to balance positive and negative samples, which allowed for faster convergence:  </p>\n<pre><code>loss = MeanSquaredError(pred, true)\n\npos_loss = (loss * true).() / (true.() + )\nneg_loss = (loss * ( - true)).() / (( - true).() + )\n\nbalanced_loss = pos_loss + neg_loss\n</code></pre>\n<h1>Inference Tips</h1>\n<p>Finally, we used four yu4u's models and three tattaka's models in the final submission.<br>\nTo stay within the time limit, we optimized our models by converting them to TensorRT format for faster inference. The conversion process was based on <a href=\"https://www.kaggle.com/code/sjtuwangshuo/converting-pytorch-checkpoints-to-tensorrt-models\" target=\"_blank\">this notebook</a>.<br>\nAdditionally, we selected a Kaggle Notebook instance with dual T4 GPUs and leveraged multiprocessing to parallelize inference.</p>\n<h1>Post Processing</h1>\n<p>For the final heatmap, we first detect local maxima using non-maximum suppression, which is implemented via max pooling with a kernel size of 7. Next, the detected points are filtered using different thresholds for each particle type.</p>\n<p>Since the detected points are in the pixel coordinate system, we need to convert them into the particle coordinate system. To do so, we proceed as follows:</p>\n<ol>\n<li><strong>Centering</strong>: Add 0.5 to the pixel coordinates to shift from the pixel’s top-left to its center.</li>\n<li><strong>Offset Correction</strong>: Subtract the 1.0 offset that was added during heatmap generation.</li>\n<li><strong>Scaling</strong>: Multiply by 10.012 to convert the adjusted pixel coordinates to the particle coordinate system.</li>\n</ol>\n<h1>Does Not Work for Us</h1>\n<ul>\n<li>Two-stage model: We built a model that refines the scores by cropping regions around the points detected using a heatmap approach and then applying a classification model to those cropped regions. Although it worked well in terms of CV scores, it did not improve the LB performance.</li>\n</ul>\n<h1>Source Code and Notebooks</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/ren4yu/czii-ensemble-tensorrt-xy-stride-th/notebook?scriptVersionId=220758003\" target=\"_blank\">Final submission</a></li>\n<li><a href=\"https://github.com/tattaka/czii-cryo-et-object-identification-public\" target=\"_blank\">Training tattaka's model</a></li>\n<li><a href=\"https://github.com/yu4u/kaggle-czii-4th\" target=\"_blank\">Training yu4u's model</a></li>\n</ul>",
  "messages": [
    {
      "id": "3116410",
      "postDate": "02/06/2025 00:01:01",
      "content": "<p>We would first like to express our gratitude to the competition host and the Kaggle staff for organizing this outstanding competition. Below, we introduce the solution of Team yu4u &amp; tattaka.</p>\n<h1>Summary</h1>\n<p>We adopted an approach to detect particle points using a heatmap-based method, which is the most commonly employed technique in pose estimation and facial keypoint detection.<br>\nSince this competition deals with 3D images rather than 2D images, we utilized two types of UNet-like models (yu4u's model and tattaka's model) that take 3D voxels as input and outputs 3D heatmaps.</p>\n<h1>Our Approach to This Competition</h1>\n<p>First, we will explain our approach to this competition, specifically how we addressed the issue of CV and LB not correlating. We used CV only to confirm that the metric produced was somewhat reasonable and for selecting checkpoints. For models with potential for improvement, we simply submitted them and relied on the LB to make decisions about which methods to adopt or discard.</p>\n<h1>Creating the Ground Truth Heatmap</h1>\n<p>We generate the ground truth heatmap necessary for model training. This involves converting the ground truth particle coordinates into the pixel coordinate system and creating a mask using a Gaussian function, where the particle center is set to 1.0 and sigma is 6 pixels for yu4u's model. For tattaka's model, different sigma values were used for different particles based on their sizes.<br>\nWe believe that an offset of 1.0 should be added when converting particle coordinates into the pixel coordinate system. While <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/553126\" target=\"_blank\">this discussion</a> suggests adding 0.5, <a href=\"https://www.kaggle.com/code/ren4yu/czii-coordinate-eda\" target=\"_blank\">our notebook</a> demonstrates that 1.0 is the correct value. The main difference is that the previous discussion assumes the particle center is at the top-left of a pixel, whereas we argue that, on average, the circle should be drawn from the pixel center (0.5, 0.5).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2F01f088d7b84cf5710d322fe17e7aaf5a%2F2025-02-06%2011.30.19.png?generation=1738809090496377&amp;alt=media\" alt=\"\"></p>\n<h1>yu4u's Model</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2F7fad2afcd6dfd4b83065d9692ecb8479%2Fmodel.png?generation=1738807875765144&amp;alt=media\" alt=\"\"></p>\n<p>We adopted a 2.5D-UNet, which utilizes a 2D image-based model as the backbone. The outputs from each stage of this backbone are pooled along the depth direction, enabling hierarchical feature extraction in the depth dimension as well. This idea was borrowed from <a href=\"https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder\" target=\"_blank\">the excellent notebook</a>. An interesting observation is that replacing this pooling operation with strided 3D convolutions degrades performance. This would be because the pooling method effectively aggregates depth features while preserving the original 2D backbone’s feature maps as much as possible. Similar to many other Kaggle competitions dealing with 3D data, a UNet utilizing a 2D backbone outperformed a straightforward UNet with a 3D backbone.</p>\n<p>We also applied 3D convolution between the encoder and decoder, inspired by <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430685\" target=\"_blank\">the 3rd Place Solution of the contrails competition</a>.</p>\n<p>Initially, we used a plain UNet architecture, but processing high-resolution feature maps required significant memory and computation. To address this, we adopted a model that outputs the final heatmap using pixel shuffle from a feature map with a stride of 4. Pixel shuffle, also known as depth_to_space in TensorFlow, is an operation that redistributes information from the channel dimension to the spatial dimensions. Compared to deconvolution, it offers advantages in computational efficiency and reducing artifacts.</p>\n<p>For the final submission, we used four models with different folds of a ConvNeXt Nano model as the backbone.</p>\n<h1>tattaka's Model</h1>\n<h2>Model Architecture</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2Fe15184020a0b9ea5e2ff58bf5d17ce25%2F2025-02-05%2023.11.47.png?generation=1738799891700544&amp;alt=media\" alt=\"\"></p>\n<p>This model is a lightweight 2.5D UNet with ResNetRS-50 as the backbone.  </p>\n<p>The input to the model is a volume of size 32×128×128 (D×H×W), and it outputs a 3D heatmap of the same size. Within the backbone, the depth is progressively reduced by half using average pooling for the first two stages. After that, average pooling with kernel=3, stride=1, padding=1 is used to maintain the depth while computing feature maps. As a result, the feature map shapes at each stage of the backbone are as follows:  <br>\n<code>(bs, ch, 16, 64, 64)</code>, <code>(bs, ch, 8, 32, 32)</code>, <code>(bs, ch, 8, 16, 16)</code>, <code>(bs, ch, 8, 8, 8)</code>, <code>(bs, ch, 8, 4, 4)</code>.  </p>\n<p>In the decoder, the three lowest-resolution feature maps are fed into <a href=\"https://arxiv.org/abs/1903.11816\" target=\"_blank\">Joint Pyramid Upsampling</a>. These maps are then progressively upsampled using 3D CNNs, SESC attention, and upsampling layers until they reach the same size as the input volume.</p>\n<h2>Loss Function</h2>\n<p>Since the number of particles within the volume is relatively small, there is a significant class imbalance between positive and negative samples during training. We attempted to adjust the parameters for generating ground truth heatmaps, but this did not lead to any improvement in cross-validation performance.  </p>\n<p>Ultimately, we implemented a simple MSE-based loss function to balance positive and negative samples, which allowed for faster convergence:  </p>\n<pre><code>loss = MeanSquaredError(pred, true)\n\npos_loss = (loss * true).() / (true.() + )\nneg_loss = (loss * ( - true)).() / (( - true).() + )\n\nbalanced_loss = pos_loss + neg_loss\n</code></pre>\n<h1>Inference Tips</h1>\n<p>Finally, we used four yu4u's models and three tattaka's models in the final submission.<br>\nTo stay within the time limit, we optimized our models by converting them to TensorRT format for faster inference. The conversion process was based on <a href=\"https://www.kaggle.com/code/sjtuwangshuo/converting-pytorch-checkpoints-to-tensorrt-models\" target=\"_blank\">this notebook</a>.<br>\nAdditionally, we selected a Kaggle Notebook instance with dual T4 GPUs and leveraged multiprocessing to parallelize inference.</p>\n<h1>Post Processing</h1>\n<p>For the final heatmap, we first detect local maxima using non-maximum suppression, which is implemented via max pooling with a kernel size of 7. Next, the detected points are filtered using different thresholds for each particle type.</p>\n<p>Since the detected points are in the pixel coordinate system, we need to convert them into the particle coordinate system. To do so, we proceed as follows:</p>\n<ol>\n<li><strong>Centering</strong>: Add 0.5 to the pixel coordinates to shift from the pixel’s top-left to its center.</li>\n<li><strong>Offset Correction</strong>: Subtract the 1.0 offset that was added during heatmap generation.</li>\n<li><strong>Scaling</strong>: Multiply by 10.012 to convert the adjusted pixel coordinates to the particle coordinate system.</li>\n</ol>\n<h1>Does Not Work for Us</h1>\n<ul>\n<li>Two-stage model: We built a model that refines the scores by cropping regions around the points detected using a heatmap approach and then applying a classification model to those cropped regions. Although it worked well in terms of CV scores, it did not improve the LB performance.</li>\n</ul>\n<h1>Source Code and Notebooks</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/ren4yu/czii-ensemble-tensorrt-xy-stride-th/notebook?scriptVersionId=220758003\" target=\"_blank\">Final submission</a></li>\n<li><a href=\"https://github.com/tattaka/czii-cryo-et-object-identification-public\" target=\"_blank\">Training tattaka's model</a></li>\n<li><a href=\"https://github.com/yu4u/kaggle-czii-4th\" target=\"_blank\">Training yu4u's model</a></li>\n</ul>",
      "rawMarkdown": "We would first like to express our gratitude to the competition host and the Kaggle staff for organizing this outstanding competition. Below, we introduce the solution of Team yu4u & tattaka.\n\n# Summary\n\nWe adopted an approach to detect particle points using a heatmap-based method, which is the most commonly employed technique in pose estimation and facial keypoint detection.\nSince this competition deals with 3D images rather than 2D images, we utilized two types of UNet-like models (yu4u's model and tattaka's model) that take 3D voxels as input and outputs 3D heatmaps.\n\n\n# Our Approach to This Competition\n\nFirst, we will explain our approach to this competition, specifically how we addressed the issue of CV and LB not correlating. We used CV only to confirm that the metric produced was somewhat reasonable and for selecting checkpoints. For models with potential for improvement, we simply submitted them and relied on the LB to make decisions about which methods to adopt or discard.\n\n# Creating the Ground Truth Heatmap\n\nWe generate the ground truth heatmap necessary for model training. This involves converting the ground truth particle coordinates into the pixel coordinate system and creating a mask using a Gaussian function, where the particle center is set to 1.0 and sigma is 6 pixels for yu4u's model. For tattaka's model, different sigma values were used for different particles based on their sizes.\nWe believe that an offset of 1.0 should be added when converting particle coordinates into the pixel coordinate system. While [this discussion](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/553126) suggests adding 0.5, [our notebook](https://www.kaggle.com/code/ren4yu/czii-coordinate-eda) demonstrates that 1.0 is the correct value. The main difference is that the previous discussion assumes the particle center is at the top-left of a pixel, whereas we argue that, on average, the circle should be drawn from the pixel center (0.5, 0.5).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2F01f088d7b84cf5710d322fe17e7aaf5a%2F2025-02-06%2011.30.19.png?generation=1738809090496377&alt=media)\n\n# yu4u's Model\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2F7fad2afcd6dfd4b83065d9692ecb8479%2Fmodel.png?generation=1738807875765144&alt=media)\n\nWe adopted a 2.5D-UNet, which utilizes a 2D image-based model as the backbone. The outputs from each stage of this backbone are pooled along the depth direction, enabling hierarchical feature extraction in the depth dimension as well. This idea was borrowed from [the excellent notebook](https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder). An interesting observation is that replacing this pooling operation with strided 3D convolutions degrades performance. This would be because the pooling method effectively aggregates depth features while preserving the original 2D backbone’s feature maps as much as possible. Similar to many other Kaggle competitions dealing with 3D data, a UNet utilizing a 2D backbone outperformed a straightforward UNet with a 3D backbone.\n\nWe also applied 3D convolution between the encoder and decoder, inspired by [the 3rd Place Solution of the contrails competition](https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430685).\n\nInitially, we used a plain UNet architecture, but processing high-resolution feature maps required significant memory and computation. To address this, we adopted a model that outputs the final heatmap using pixel shuffle from a feature map with a stride of 4. Pixel shuffle, also known as depth_to_space in TensorFlow, is an operation that redistributes information from the channel dimension to the spatial dimensions. Compared to deconvolution, it offers advantages in computational efficiency and reducing artifacts.\n\nFor the final submission, we used four models with different folds of a ConvNeXt Nano model as the backbone.\n\n# tattaka's Model\n\n## Model Architecture\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2Fe15184020a0b9ea5e2ff58bf5d17ce25%2F2025-02-05%2023.11.47.png?generation=1738799891700544&alt=media)\n\nThis model is a lightweight 2.5D UNet with ResNetRS-50 as the backbone.  \n\nThe input to the model is a volume of size 32×128×128 (D×H×W), and it outputs a 3D heatmap of the same size. Within the backbone, the depth is progressively reduced by half using average pooling for the first two stages. After that, average pooling with kernel=3, stride=1, padding=1 is used to maintain the depth while computing feature maps. As a result, the feature map shapes at each stage of the backbone are as follows:  \n`(bs, ch, 16, 64, 64)`, `(bs, ch, 8, 32, 32)`, `(bs, ch, 8, 16, 16)`, `(bs, ch, 8, 8, 8)`, `(bs, ch, 8, 4, 4)`.  \n\nIn the decoder, the three lowest-resolution feature maps are fed into [Joint Pyramid Upsampling](https://arxiv.org/abs/1903.11816). These maps are then progressively upsampled using 3D CNNs, SESC attention, and upsampling layers until they reach the same size as the input volume.\n\n## Loss Function\n\nSince the number of particles within the volume is relatively small, there is a significant class imbalance between positive and negative samples during training. We attempted to adjust the parameters for generating ground truth heatmaps, but this did not lead to any improvement in cross-validation performance.  \n\nUltimately, we implemented a simple MSE-based loss function to balance positive and negative samples, which allowed for faster convergence:  \n\n```python\nloss = MeanSquaredError(pred, true)\n\npos_loss = (loss * true).sum() / (true.sum() + 1e-6)\nneg_loss = (loss * (1 - true)).sum() / ((1 - true).sum() + 1e-6)\n\nbalanced_loss = pos_loss + neg_loss\n```\n\n# Inference Tips\n\nFinally, we used four yu4u's models and three tattaka's models in the final submission.\nTo stay within the time limit, we optimized our models by converting them to TensorRT format for faster inference. The conversion process was based on [this notebook](https://www.kaggle.com/code/sjtuwangshuo/converting-pytorch-checkpoints-to-tensorrt-models).\nAdditionally, we selected a Kaggle Notebook instance with dual T4 GPUs and leveraged multiprocessing to parallelize inference.\n\n# Post Processing\n\nFor the final heatmap, we first detect local maxima using non-maximum suppression, which is implemented via max pooling with a kernel size of 7. Next, the detected points are filtered using different thresholds for each particle type.\n\nSince the detected points are in the pixel coordinate system, we need to convert them into the particle coordinate system. To do so, we proceed as follows:\n\n1. **Centering**: Add 0.5 to the pixel coordinates to shift from the pixel’s top-left to its center.\n2. **Offset Correction**: Subtract the 1.0 offset that was added during heatmap generation.\n3. **Scaling**: Multiply by 10.012 to convert the adjusted pixel coordinates to the particle coordinate system.\n\n\n# Does Not Work for Us\n\n* Two-stage model: We built a model that refines the scores by cropping regions around the points detected using a heatmap approach and then applying a classification model to those cropped regions. Although it worked well in terms of CV scores, it did not improve the LB performance.\n\n# Source Code and Notebooks\n\n- [Final submission](https://www.kaggle.com/code/ren4yu/czii-ensemble-tensorrt-xy-stride-th/notebook?scriptVersionId=220758003)\n- [Training tattaka's model](https://github.com/tattaka/czii-cryo-et-object-identification-public)\n- [Training yu4u's model](https://github.com/yu4u/kaggle-czii-4th)",
      "votes": null
    },
    {
      "id": "3116412",
      "postDate": "02/06/2025 00:04:16",
      "content": "<p>I’m very pleased to see a 2.5D UNet architecture as a winning solution.  I read through the same notebook you referenced early on and thought it was a smart architecture given the size of the architecture having larger height and width dimensions.  The use of a 3D convolution layer between the encoder and decoder was not something I thought of and it’s very interesting.  Thank you for sharing your solution and congratulations!</p>",
      "rawMarkdown": "I’m very pleased to see a 2.5D UNet architecture as a winning solution.  I read through the same notebook you referenced early on and thought it was a smart architecture given the size of the architecture having larger height and width dimensions.  The use of a 3D convolution layer between the encoder and decoder was not something I thought of and it’s very interesting.  Thank you for sharing your solution and congratulations!",
      "votes": null
    },
    {
      "id": "3116417",
      "postDate": "02/06/2025 00:15:00",
      "content": "<p>Huge congrats on your strong finish and the prize! This is also a solid write-up with all the details. I look forward to seeing your code. I am sure there is a lot to study.</p>",
      "rawMarkdown": "Huge congrats on your strong finish and the prize! This is also a solid write-up with all the details. I look forward to seeing your code. I am sure there is a lot to study.",
      "votes": null
    },
    {
      "id": "3116418",
      "postDate": "02/06/2025 00:17:37",
      "content": "<p>Congratulations!!!!!!</p>\n<p>I have three questions:</p>\n<ol>\n<li>Compared to the standard MSE, Focal loss, BCE, did the loss function you used have a different effect to performance?(How much difference?)</li>\n<li>You set the kernel size to 7. Is there any reason you didn't make it variable based on the size of the evaluation particles?</li>\n<li>You are using max pooling. Is there any reason for not using average pooling?</li>\n</ol>",
      "rawMarkdown": "Congratulations!!!!!!\n\nI have three questions:\n1. Compared to the standard MSE, Focal loss, BCE, did the loss function you used have a different effect to performance?(How much difference?)\n2. You set the kernel size to 7. Is there any reason you didn't make it variable based on the size of the evaluation particles?\n3. You are using max pooling. Is there any reason for not using average pooling?",
      "votes": null
    },
    {
      "id": "3116426",
      "postDate": "02/06/2025 00:30:09",
      "content": "<p>Thanks for sharing your solution, and congratulations on achieving 4th place! 🎉.<br>\nYour approach using heatmap-based detection and 2.5D UNet architectures is fascinating, especially the combination of pooling along the depth direction and the use of pixel shuffle for efficient upsampling. The insights on CV-LB correlation and post-processing steps are particularly valuable.</p>\n<p>I also found that discussion on offset corrections and coordinate conversions very insightful. these small details can make a big difference in final accuracy. Great job, and thanks again for sharing your methodology. </p>",
      "rawMarkdown": "Thanks for sharing your solution, and congratulations on achieving 4th place! 🎉.\nYour approach using heatmap-based detection and 2.5D UNet architectures is fascinating, especially the combination of pooling along the depth direction and the use of pixel shuffle for efficient upsampling. The insights on CV-LB correlation and post-processing steps are particularly valuable.\n\nI also found that discussion on offset corrections and coordinate conversions very insightful. these small details can make a big difference in final accuracy. Great job, and thanks again for sharing your methodology.",
      "votes": null
    },
    {
      "id": "3116432",
      "postDate": "02/06/2025 00:43:31",
      "content": "<p>I'm also curious about the 1st question. I don't know whether the loss function is the key to the competition, because I tried many backbone with different combination like focal loss, BCE and Tversky loss, but no good effects. </p>",
      "rawMarkdown": "I'm also curious about the 1st question. I don't know whether the loss function is the key to the competition, because I tried many backbone with different combination like focal loss, BCE and Tversky loss, but no good effects.",
      "votes": null
    },
    {
      "id": "3116433",
      "postDate": "02/06/2025 00:44:03",
      "content": "<ol>\n<li><p>Initially, I experimented with Focal Loss, BCE, and FBeta Loss for training, but the MSE-based loss ultimately yielded better final accuracy. Additionally, when using EMA for smoothing, it is crucial to balance positive and negative samples; otherwise, convergence slows down significantly.</p></li>\n<li><p><a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/blob_detector_inference.ipynb\" target=\"_blank\">In the host notebook</a>, the kernel size was dynamically adjusted for each particle. We considered implementing the same approach, but due to its lower priority and limited submission attempts, I was unable to test it.</p></li>\n<li><p>Max pooling in post-processing serves as a peak detection mechanism.</p></li>\n</ol>",
      "rawMarkdown": "1. Initially, I experimented with Focal Loss, BCE, and FBeta Loss for training, but the MSE-based loss ultimately yielded better final accuracy. Additionally, when using EMA for smoothing, it is crucial to balance positive and negative samples; otherwise, convergence slows down significantly.\n\n2. [In the host notebook](https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/blob_detector_inference.ipynb), the kernel size was dynamically adjusted for each particle. We considered implementing the same approach, but due to its lower priority and limited submission attempts, I was unable to test it.\n\n3. Max pooling in post-processing serves as a peak detection mechanism.",
      "votes": null
    },
    {
      "id": "3116438",
      "postDate": "02/06/2025 00:54:59",
      "content": "<p>The loss section is fascinating and carries many implications. Thank you for sharing.</p>",
      "rawMarkdown": "The loss section is fascinating and carries many implications. Thank you for sharing.",
      "votes": null
    },
    {
      "id": "3116440",
      "postDate": "02/06/2025 00:55:44",
      "content": "<p>Congratulations! We found that the norm layer and activation layer may be the key. However, the 3D convolution layer between the encoder and decoder is really interesting. Thank you for coming up with the solution so quickly after the close of the competition.</p>",
      "rawMarkdown": "Congratulations! We found that the norm layer and activation layer may be the key. However, the 3D convolution layer between the encoder and decoder is really interesting. Thank you for coming up with the solution so quickly after the close of the competition.",
      "votes": null
    },
    {
      "id": "3116447",
      "postDate": "02/06/2025 00:58:51",
      "content": "<p>Congratulation to your team.  What's the norm layer and activation layer you use? softmax?</p>",
      "rawMarkdown": "Congratulation to your team.  What's the norm layer and activation layer you use? softmax?",
      "votes": null
    },
    {
      "id": "3116454",
      "postDate": "02/06/2025 01:05:51",
      "content": "<p>We tried replacing the norm layer with InstanceNorm3d and the activation layer with PReLU due to the high score of the MONAI UNet.</p>",
      "rawMarkdown": "We tried replacing the norm layer with InstanceNorm3d and the activation layer with PReLU due to the high score of the MONAI UNet.",
      "votes": null
    },
    {
      "id": "3116502",
      "postDate": "02/06/2025 03:14:24",
      "content": "<p>Training code of tattaka's part: <a href=\"https://github.com/tattaka/czii-cryo-et-object-identification-public\" target=\"_blank\">https://github.com/tattaka/czii-cryo-et-object-identification-public</a></p>",
      "rawMarkdown": "Training code of tattaka's part: https://github.com/tattaka/czii-cryo-et-object-identification-public",
      "votes": null
    },
    {
      "id": "3116699",
      "postDate": "02/06/2025 08:20:24",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> ! May I ask which where you specific thresholds in post processing? <code>Next, the detected points are filtered using different thresholds for each particle type.</code></p>",
      "rawMarkdown": "Congratulations @ren4yu ! May I ask which where you specific thresholds in post processing? `Next, the detected points are filtered using different thresholds for each particle type.`",
      "votes": null
    },
    {
      "id": "3116724",
      "postDate": "02/06/2025 08:53:41",
      "content": "<p>Your approach using heatmap-based detection and 2.5D UNet architectures is fascinating, especially the combination of pooling along the depth direction and the use of pixel shuffle for efficient upsampling. The insights on CV-LB correlation and post-processing steps are particularly valuable.</p>",
      "rawMarkdown": "Your approach using heatmap-based detection and 2.5D UNet architectures is fascinating, especially the combination of pooling along the depth direction and the use of pixel shuffle for efficient upsampling. The insights on CV-LB correlation and post-processing steps are particularly valuable.",
      "votes": null
    },
    {
      "id": "3116751",
      "postDate": "02/06/2025 09:13:48",
      "content": "<pre><code>thresholds = [, , , , ] \n</code></pre>\n<pre><code>threshold = thresholds[i]\ncoordinates = find_local_maxima(preds[i], threshold, nms_filter_size)\n</code></pre>\n<p>used <a href=\"https://www.kaggle.com/code/ren4yu/czii-ensemble-tensorrt-xy-stride-th?scriptVersionId=220758003&amp;cellId=12\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "```python\nthresholds = [0.5, 0.5, 0.5, 0.5, 0.4] # ensemble\n```\n\n```python\nthreshold = thresholds[i]\ncoordinates = find_local_maxima(preds[i], threshold, nms_filter_size)\n```\n\n\nused [here](https://www.kaggle.com/code/ren4yu/czii-ensemble-tensorrt-xy-stride-th?scriptVersionId=220758003&cellId=12)",
      "votes": null
    },
    {
      "id": "3116793",
      "postDate": "02/06/2025 10:14:45",
      "content": "<p>Thank you </p>",
      "rawMarkdown": "Thank you",
      "votes": null
    },
    {
      "id": "3116794",
      "postDate": "02/06/2025 10:15:07",
      "content": "<p>Good work …</p>",
      "rawMarkdown": "Good work ...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3116412,
      "author_name": "connorjd",
      "author_url": "",
      "post_date": "02/06/2025 00:04:16",
      "content": "<p>I’m very pleased to see a 2.5D UNet architecture as a winning solution.  I read through the same notebook you referenced early on and thought it was a smart architecture given the size of the architecture having larger height and width dimensions.  The use of a 3D convolution layer between the encoder and decoder was not something I thought of and it’s very interesting.  Thank you for sharing your solution and congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3116417,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "02/06/2025 00:15:00",
      "content": "<p>Huge congrats on your strong finish and the prize! This is also a solid write-up with all the details. I look forward to seeing your code. I am sure there is a lot to study.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3116418,
      "author_name": "sugupoko",
      "author_url": "",
      "post_date": "02/06/2025 00:17:37",
      "content": "<p>Congratulations!!!!!!</p>\n<p>I have three questions:</p>\n<ol>\n<li>Compared to the standard MSE, Focal loss, BCE, did the loss function you used have a different effect to performance?(How much difference?)</li>\n<li>You set the kernel size to 7. Is there any reason you didn't make it variable based on the size of the evaluation particles?</li>\n<li>You are using max pooling. Is there any reason for not using average pooling?</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 3116432,
          "author_name": "sweetyheehee",
          "author_url": "",
          "post_date": "02/06/2025 00:43:31",
          "content": "<p>I'm also curious about the 1st question. I don't know whether the loss function is the key to the competition, because I tried many backbone with different combination like focal loss, BCE and Tversky loss, but no good effects. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3116433,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "02/06/2025 00:44:03",
          "content": "<ol>\n<li><p>Initially, I experimented with Focal Loss, BCE, and FBeta Loss for training, but the MSE-based loss ultimately yielded better final accuracy. Additionally, when using EMA for smoothing, it is crucial to balance positive and negative samples; otherwise, convergence slows down significantly.</p></li>\n<li><p><a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/blob_detector_inference.ipynb\" target=\"_blank\">In the host notebook</a>, the kernel size was dynamically adjusted for each particle. We considered implementing the same approach, but due to its lower priority and limited submission attempts, I was unable to test it.</p></li>\n<li><p>Max pooling in post-processing serves as a peak detection mechanism.</p></li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3116426,
      "author_name": "engadamalmohammedi",
      "author_url": "",
      "post_date": "02/06/2025 00:30:09",
      "content": "<p>Thanks for sharing your solution, and congratulations on achieving 4th place! 🎉.<br>\nYour approach using heatmap-based detection and 2.5D UNet architectures is fascinating, especially the combination of pooling along the depth direction and the use of pixel shuffle for efficient upsampling. The insights on CV-LB correlation and post-processing steps are particularly valuable.</p>\n<p>I also found that discussion on offset corrections and coordinate conversions very insightful. these small details can make a big difference in final accuracy. Great job, and thanks again for sharing your methodology. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3116438,
      "author_name": "junhanzangai",
      "author_url": "",
      "post_date": "02/06/2025 00:54:59",
      "content": "<p>The loss section is fascinating and carries many implications. Thank you for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3116440,
      "author_name": "luoziqian",
      "author_url": "",
      "post_date": "02/06/2025 00:55:44",
      "content": "<p>Congratulations! We found that the norm layer and activation layer may be the key. However, the 3D convolution layer between the encoder and decoder is really interesting. Thank you for coming up with the solution so quickly after the close of the competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3116447,
          "author_name": "sweetyheehee",
          "author_url": "",
          "post_date": "02/06/2025 00:58:51",
          "content": "<p>Congratulation to your team.  What's the norm layer and activation layer you use? softmax?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3116454,
              "author_name": "luoziqian",
              "author_url": "",
              "post_date": "02/06/2025 01:05:51",
              "content": "<p>We tried replacing the norm layer with InstanceNorm3d and the activation layer with PReLU due to the high score of the MONAI UNet.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3116793,
          "author_name": "rahul1845",
          "author_url": "",
          "post_date": "02/06/2025 10:14:45",
          "content": "<p>Thank you </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3116502,
      "author_name": "tattaka",
      "author_url": "",
      "post_date": "02/06/2025 03:14:24",
      "content": "<p>Training code of tattaka's part: <a href=\"https://github.com/tattaka/czii-cryo-et-object-identification-public\" target=\"_blank\">https://github.com/tattaka/czii-cryo-et-object-identification-public</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3116699,
      "author_name": "octaviograu",
      "author_url": "",
      "post_date": "02/06/2025 08:20:24",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> ! May I ask which where you specific thresholds in post processing? <code>Next, the detected points are filtered using different thresholds for each particle type.</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 3116751,
          "author_name": "ren4yu",
          "author_url": "",
          "post_date": "02/06/2025 09:13:48",
          "content": "<pre><code>thresholds = [, , , , ] \n</code></pre>\n<pre><code>threshold = thresholds[i]\ncoordinates = find_local_maxima(preds[i], threshold, nms_filter_size)\n</code></pre>\n<p>used <a href=\"https://www.kaggle.com/code/ren4yu/czii-ensemble-tensorrt-xy-stride-th?scriptVersionId=220758003&amp;cellId=12\" target=\"_blank\">here</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3116724,
      "author_name": "yuvikagogarr",
      "author_url": "",
      "post_date": "02/06/2025 08:53:41",
      "content": "<p>Your approach using heatmap-based detection and 2.5D UNet architectures is fascinating, especially the combination of pooling along the depth direction and the use of pixel shuffle for efficient upsampling. The insights on CV-LB correlation and post-processing steps are particularly valuable.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3116794,
      "author_name": "rahul1845",
      "author_url": "",
      "post_date": "02/06/2025 10:15:07",
      "content": "<p>Good work …</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3116410": "We would first like to express our gratitude to the competition host and the Kaggle staff for organizing this outstanding competition. Below, we introduce the solution of Team yu4u & tattaka.\n\n# Summary\n\nWe adopted an approach to detect particle points using a heatmap-based method, which is the most commonly employed technique in pose estimation and facial keypoint detection.\nSince this competition deals with 3D images rather than 2D images, we utilized two types of UNet-like models (yu4u's model and tattaka's model) that take 3D voxels as input and outputs 3D heatmaps.\n\n\n# Our Approach to This Competition\n\nFirst, we will explain our approach to this competition, specifically how we addressed the issue of CV and LB not correlating. We used CV only to confirm that the metric produced was somewhat reasonable and for selecting checkpoints. For models with potential for improvement, we simply submitted them and relied on the LB to make decisions about which methods to adopt or discard.\n\n# Creating the Ground Truth Heatmap\n\nWe generate the ground truth heatmap necessary for model training. This involves converting the ground truth particle coordinates into the pixel coordinate system and creating a mask using a Gaussian function, where the particle center is set to 1.0 and sigma is 6 pixels for yu4u's model. For tattaka's model, different sigma values were used for different particles based on their sizes.\nWe believe that an offset of 1.0 should be added when converting particle coordinates into the pixel coordinate system. While [this discussion](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/553126) suggests adding 0.5, [our notebook](https://www.kaggle.com/code/ren4yu/czii-coordinate-eda) demonstrates that 1.0 is the correct value. The main difference is that the previous discussion assumes the particle center is at the top-left of a pixel, whereas we argue that, on average, the circle should be drawn from the pixel center (0.5, 0.5).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2F01f088d7b84cf5710d322fe17e7aaf5a%2F2025-02-06%2011.30.19.png?generation=1738809090496377&alt=media)\n\n# yu4u's Model\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2F7fad2afcd6dfd4b83065d9692ecb8479%2Fmodel.png?generation=1738807875765144&alt=media)\n\nWe adopted a 2.5D-UNet, which utilizes a 2D image-based model as the backbone. The outputs from each stage of this backbone are pooled along the depth direction, enabling hierarchical feature extraction in the depth dimension as well. This idea was borrowed from [the excellent notebook](https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder). An interesting observation is that replacing this pooling operation with strided 3D convolutions degrades performance. This would be because the pooling method effectively aggregates depth features while preserving the original 2D backbone’s feature maps as much as possible. Similar to many other Kaggle competitions dealing with 3D data, a UNet utilizing a 2D backbone outperformed a straightforward UNet with a 3D backbone.\n\nWe also applied 3D convolution between the encoder and decoder, inspired by [the 3rd Place Solution of the contrails competition](https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430685).\n\nInitially, we used a plain UNet architecture, but processing high-resolution feature maps required significant memory and computation. To address this, we adopted a model that outputs the final heatmap using pixel shuffle from a feature map with a stride of 4. Pixel shuffle, also known as depth_to_space in TensorFlow, is an operation that redistributes information from the channel dimension to the spatial dimensions. Compared to deconvolution, it offers advantages in computational efficiency and reducing artifacts.\n\nFor the final submission, we used four models with different folds of a ConvNeXt Nano model as the backbone.\n\n# tattaka's Model\n\n## Model Architecture\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2Fe15184020a0b9ea5e2ff58bf5d17ce25%2F2025-02-05%2023.11.47.png?generation=1738799891700544&alt=media)\n\nThis model is a lightweight 2.5D UNet with ResNetRS-50 as the backbone.  \n\nThe input to the model is a volume of size 32×128×128 (D×H×W), and it outputs a 3D heatmap of the same size. Within the backbone, the depth is progressively reduced by half using average pooling for the first two stages. After that, average pooling with kernel=3, stride=1, padding=1 is used to maintain the depth while computing feature maps. As a result, the feature map shapes at each stage of the backbone are as follows:  \n`(bs, ch, 16, 64, 64)`, `(bs, ch, 8, 32, 32)`, `(bs, ch, 8, 16, 16)`, `(bs, ch, 8, 8, 8)`, `(bs, ch, 8, 4, 4)`.  \n\nIn the decoder, the three lowest-resolution feature maps are fed into [Joint Pyramid Upsampling](https://arxiv.org/abs/1903.11816). These maps are then progressively upsampled using 3D CNNs, SESC attention, and upsampling layers until they reach the same size as the input volume.\n\n## Loss Function\n\nSince the number of particles within the volume is relatively small, there is a significant class imbalance between positive and negative samples during training. We attempted to adjust the parameters for generating ground truth heatmaps, but this did not lead to any improvement in cross-validation performance.  \n\nUltimately, we implemented a simple MSE-based loss function to balance positive and negative samples, which allowed for faster convergence:  \n\n```python\nloss = MeanSquaredError(pred, true)\n\npos_loss = (loss * true).sum() / (true.sum() + 1e-6)\nneg_loss = (loss * (1 - true)).sum() / ((1 - true).sum() + 1e-6)\n\nbalanced_loss = pos_loss + neg_loss\n```\n\n# Inference Tips\n\nFinally, we used four yu4u's models and three tattaka's models in the final submission.\nTo stay within the time limit, we optimized our models by converting them to TensorRT format for faster inference. The conversion process was based on [this notebook](https://www.kaggle.com/code/sjtuwangshuo/converting-pytorch-checkpoints-to-tensorrt-models).\nAdditionally, we selected a Kaggle Notebook instance with dual T4 GPUs and leveraged multiprocessing to parallelize inference.\n\n# Post Processing\n\nFor the final heatmap, we first detect local maxima using non-maximum suppression, which is implemented via max pooling with a kernel size of 7. Next, the detected points are filtered using different thresholds for each particle type.\n\nSince the detected points are in the pixel coordinate system, we need to convert them into the particle coordinate system. To do so, we proceed as follows:\n\n1. **Centering**: Add 0.5 to the pixel coordinates to shift from the pixel’s top-left to its center.\n2. **Offset Correction**: Subtract the 1.0 offset that was added during heatmap generation.\n3. **Scaling**: Multiply by 10.012 to convert the adjusted pixel coordinates to the particle coordinate system.\n\n\n# Does Not Work for Us\n\n* Two-stage model: We built a model that refines the scores by cropping regions around the points detected using a heatmap approach and then applying a classification model to those cropped regions. Although it worked well in terms of CV scores, it did not improve the LB performance.\n\n# Source Code and Notebooks\n\n- [Final submission](https://www.kaggle.com/code/ren4yu/czii-ensemble-tensorrt-xy-stride-th/notebook?scriptVersionId=220758003)\n- [Training tattaka's model](https://github.com/tattaka/czii-cryo-et-object-identification-public)\n- [Training yu4u's model](https://github.com/yu4u/kaggle-czii-4th)",
    "3116412": "I’m very pleased to see a 2.5D UNet architecture as a winning solution.  I read through the same notebook you referenced early on and thought it was a smart architecture given the size of the architecture having larger height and width dimensions.  The use of a 3D convolution layer between the encoder and decoder was not something I thought of and it’s very interesting.  Thank you for sharing your solution and congratulations!",
    "3116417": "Huge congrats on your strong finish and the prize! This is also a solid write-up with all the details. I look forward to seeing your code. I am sure there is a lot to study.",
    "3116418": "Congratulations!!!!!!\n\nI have three questions:\n1. Compared to the standard MSE, Focal loss, BCE, did the loss function you used have a different effect to performance?(How much difference?)\n2. You set the kernel size to 7. Is there any reason you didn't make it variable based on the size of the evaluation particles?\n3. You are using max pooling. Is there any reason for not using average pooling?",
    "3116426": "Thanks for sharing your solution, and congratulations on achieving 4th place! 🎉.\nYour approach using heatmap-based detection and 2.5D UNet architectures is fascinating, especially the combination of pooling along the depth direction and the use of pixel shuffle for efficient upsampling. The insights on CV-LB correlation and post-processing steps are particularly valuable.\n\nI also found that discussion on offset corrections and coordinate conversions very insightful. these small details can make a big difference in final accuracy. Great job, and thanks again for sharing your methodology.",
    "3116432": "I'm also curious about the 1st question. I don't know whether the loss function is the key to the competition, because I tried many backbone with different combination like focal loss, BCE and Tversky loss, but no good effects.",
    "3116433": "1. Initially, I experimented with Focal Loss, BCE, and FBeta Loss for training, but the MSE-based loss ultimately yielded better final accuracy. Additionally, when using EMA for smoothing, it is crucial to balance positive and negative samples; otherwise, convergence slows down significantly.\n\n2. [In the host notebook](https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/blob_detector_inference.ipynb), the kernel size was dynamically adjusted for each particle. We considered implementing the same approach, but due to its lower priority and limited submission attempts, I was unable to test it.\n\n3. Max pooling in post-processing serves as a peak detection mechanism.",
    "3116438": "The loss section is fascinating and carries many implications. Thank you for sharing.",
    "3116440": "Congratulations! We found that the norm layer and activation layer may be the key. However, the 3D convolution layer between the encoder and decoder is really interesting. Thank you for coming up with the solution so quickly after the close of the competition.",
    "3116447": "Congratulation to your team.  What's the norm layer and activation layer you use? softmax?",
    "3116454": "We tried replacing the norm layer with InstanceNorm3d and the activation layer with PReLU due to the high score of the MONAI UNet.",
    "3116502": "Training code of tattaka's part: https://github.com/tattaka/czii-cryo-et-object-identification-public",
    "3116699": "Congratulations @ren4yu ! May I ask which where you specific thresholds in post processing? `Next, the detected points are filtered using different thresholds for each particle type.`",
    "3116724": "Your approach using heatmap-based detection and 2.5D UNet architectures is fascinating, especially the combination of pooling along the depth direction and the use of pixel shuffle for efficient upsampling. The insights on CV-LB correlation and post-processing steps are particularly valuable.",
    "3116751": "```python\nthresholds = [0.5, 0.5, 0.5, 0.5, 0.4] # ensemble\n```\n\n```python\nthreshold = thresholds[i]\ncoordinates = find_local_maxima(preds[i], threshold, nms_filter_size)\n```\n\n\nused [here](https://www.kaggle.com/code/ren4yu/czii-ensemble-tensorrt-xy-stride-th?scriptVersionId=220758003&cellId=12)",
    "3116793": "Thank you",
    "3116794": "Good work ..."
  },
  "source": "meta"
}