{
  "id": 561515,
  "title": "8th Place Solution for the CZII - CryoET Object Identification Competition + Code Released",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/561515",
  "author_name": "Sergio Alvarez",
  "post_date": "2025-02-06T14:40:24.903000",
  "votes": 27,
  "comment_count": 31,
  "views": 0,
  "content": "<p>First, we thank the competition host and Kaggle staff for organizing this competition. Below, we introduce the solution of the team I Cryo Everyteim -- <a href=\"https://www.kaggle.com/sirapoabchaikunsaeng\" target=\"_blank\">@sirapoabchaikunsaeng</a>, <a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a>, <a href=\"https://www.kaggle.com/itsuki9180\" target=\"_blank\">@itsuki9180</a>, <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> -- </p>\n<h2>Context</h2>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/overview\" target=\"_blank\">competition overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/data\" target=\"_blank\">competition data</a></li>\n</ul>\n<h2>Overview of the approach</h2>\n<p>Our final submissions consisted of four 3D U-Net model soups, trained with different model sizes, parameters, and training data to ensure strong model complementarity. All 3D U-Nets were trained using patch sizes of (128, 128, 128) but were inferred with patch sizes of (160, 384, 384), with a 25% overlap using gaussian reconstruction to handle border artifacts. Additionally, we employed geometric test-time augmentation (TTA), including flipping and transpose.</p>\n<h3>Models</h3>\n<p>Our models were originally based on the host's example notebook and <a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a> notebook . We utilized 3D U-Net architectures from the MONAI library, trained with patch sizes of (128, 128, 128).<br>\nThe models were 3 levels deep with strides of (2, 2, 1). We started with a simple model:</p>\n<pre><code>: \n: \n: \n</code></pre>\n<p>This single model with our inference strategy reached 0.744 in public leaderboard score.<br>\nLater, we trained more complex U-Nets with configs:</p>\n<pre><code>: ,\n: ,\n: ,\n</code></pre>\n<p>The model with above config alone achieved a score of 0.759 LB.</p>\n<pre><code>: ,\n: ,\n: ,\n</code></pre>\n<p>We also applied a dropout of 0.2 or 0.3 in the model.</p>\n<h3>Training</h3>\n<p>The models were pre-trained on six synthetic tomograms denoised with Gaussian denoising—specifically, the 'TS_0', 'TS_1', 'TS_10', 'TS_11', 'TS_12', and 'TS_13' tomograms.<br>\n<em>@sersasj note: Gaussian denoising was applied because it visually improved the WBP tomograms particle visualization and was easy to implement. I hypothesized that better results could be achieved with a more advanced denoiser, but attempting to code a model for denoising used too much of my Kaggle quota, so I gave up.</em><br>\nPretraining not only reduced the time needed for the models to learn the particles but also increased the LB score by roughly 0.01.<br>\nBoth pretraining on synthetic data and fine-tuning uses the following transformations:</p>\n<pre><code>Compose([\n   RandCropByLabelClassesd(\n       keys=[, ],\n       =,\n       spatial_size=[128, 128, 128],\n       =7,  # background,  all 6 classes\n       =16,\n   ),\n   RandFlipd(keys=[, ], =0.5, =0),\n   RandFlipd(keys=[, ], =0.5, =1),\n   RandFlipd(keys=[, ], =0.5, =2),\n])\n</code></pre>\n<p>We used MONAI's DiceCELoss and optimized the models with AdamW, employing a learning rate reduction on plateau and an initial learning rate of 1e-3.<br>\nWe also experimented with Exponential Moving Average (EMA), which showed good results. In the last 3 days of the competition, we trained models using other tomo types: \"denoised\", \"ctfdeconvolved\", \"isonetcorrected\". For almost the entire competition, we hadn't found any increase in LB strategies other than geometric data augmentation in training (flip and rotate); nevertheless, training with other tomo types yielded good results.</p>\n<h3>Validation strategy</h3>\n<p>We primarily relied on out-of-fold predictions from our k-fold models. Specifically, we implemented a 7-fold cross-validation approach where we trained on all tomographies except one, which was used as a validation set. This process was repeated seven times, with each fold serving as the validation set once. This helped us obtain a score that correlated well with the public leaderboard score, typically 0.01 to 0.02 points higher.<br>\nOne limitation we encountered was our inability to apply this validation strategy when testing ensembles of our model soups. In these cases, we had to rely on leaked data scores. While this approach didn't correlate as strongly with leaderboard scores, most of the time it allowed us to rank the performance of different ensembles, and compare TTA strategies. </p>\n<h2>Details of the submission</h2>\n<h3>Inference</h3>\n<p>During inference, we use MONAI's sliding_window_inference with 25% overlap and Gaussian reconstruction to handle border predictions more accurately; this boosted the public LB by around 0.002. <strong>One of the greatest findings was making the inference in patches as large as possible; we used a window size of (160, 384, 384), boosting our score by around 0.01 compared to a size of (128, 128, 128).</strong> (Thanks <a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a> for finding this)<br>\nWe also experimented with combining different sizes of predictions, for example (160, 384, 384) + (136, 248, 248), but inference took much longer while yielding inconclusive gains. Therefore, we abandoned this approach in favor of creating more ensembles and employing more TTA.<br>\nFor post-processing, we improved particle detection by combining model predictions at the logits level rather than averaging probabilities directly.<br>\nWe then refined particle identification using watershed segmentation, which helped separate touching or overlapping particles. The watershed process created binary masks using particle-specific certainty thresholds, computed distance transforms to measure particle separation, identified distinct particles using local maxima as markers, and finally applied watershed segmentation to determine particle boundaries. We switched from using skimage to cucim and cupy to decrease the inference time. <br>\nWe also used specific blob threshold sizes for each particle to decrease the number of false positives in our predictions.<br>\nMoreover, we used flip along the X and Y axes and transpose TTA.</p>\n<h3>Ensembles</h3>\n<p>Ensembling was always one of the main ideas in our team. When the team was fully formed 2 weeks before the deadline, we were all around a 0.73 LB score, around 50th place on the leaderboard. At the time, I (@sersasj) had a YOLO scoring 0.682 and a U-Net scoring 0.692 that, when ensembled, reached 0.732 in LB. The rest of the team had only a 3D U-Net that reached 0.739. So it was obvious that we would ensemble everything and achieve amazing results; sadly, it didn't work.<br>\nNevertheless, we learned valuable lessons from this experience. My 3D U-Net was ~30 times smaller than the 0.739 scoring U-Net, but it proved quite effective when used with the team's amazing inference techniques: <em>larger patches, blob thresholds, ensemble at logits level, and watershed processing</em>. After implementing pretraining and additional training, it achieved a 0.744 leaderboard score. <br>\nThanks to our group work (especially <a href=\"https://www.kaggle.com/sirapoabchaikunsaeng\" target=\"_blank\">@sirapoabchaikunsaeng</a> and <a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a>), we successfully trained several U-Net variants derived from the 0.744 model by modifying parameters, implementing EMA (Exponential Moving Average), incorporating more data, and applying the techniques mentioned above.<br>\nBy doing so, we had various U-Nets that were complementary to each other, but we couldn't use them all. To manage \"using\" all we had, <a href=\"https://www.kaggle.com/itsuki9180\" target=\"_blank\">@itsuki9180</a> presented the idea and code for a technique called model soup, which essentially averages the weights of multiple models from the same pre-trained weights.<br>\nWe applied this to the models from our k-fold training.<br>\nOur final ensemble consisted of:</p>\n<ol>\n<li>A soup of 7 folds trained with parameters:<ul>\n<li>Channels: (32, 64, 128, 128)</li>\n<li>Number of residual units: 1</li>\n<li>Pretrained on synthetic data</li>\n<li>Trained on denoised images without EMA</li></ul></li>\n<li>A soup of 7 folds trained with parameters:<ul>\n<li>Channels: (32, 64, 128, 256)</li>\n<li>Number of residual units: 2</li>\n<li>Pretrained on synthetic data</li>\n<li>Trained on all denoised, IsoNet-corrected, and CTF-deconvolved images without EMA</li></ul></li>\n<li>A soup of 3 folds (TS_69_2, TS_86_3, TS_99_9) trained with parameters:<ul>\n<li>Channels: (32, 96, 256, 384)</li>\n<li>Number of residual units: 2</li>\n<li>Pretrained on synthetic data</li>\n<li>Trained on denoised images with EMA</li></ul></li>\n<li>A soup of 7 folds trained with parameters:<ul>\n<li>Channels: (32, 96, 256, 384)</li>\n<li>Number of residual units: 2</li>\n<li>Pretrained on synthetic data</li>\n<li>Trained on all denoised, IsoNet-corrected, and CTF-deconvolved images without EMA</li></ul></li>\n</ol>\n<h3>What didn't work:</h3>\n<ul>\n<li>Increased Augmentations</li>\n<li>CutMix, MixUp</li>\n<li>2D U-Net approaches (reached 0.636 lb only)</li>\n<li>Ensembling with YOLO (for final solution, works well to reach ~0.73 lb)</li>\n<li>Bigger and deeper U-Nets</li>\n<li>Cross entropy only loss functions</li>\n</ul>\n<h2>Sources section</h2>\n<ul>\n<li><a href=\"https://arxiv.org/abs/2203.05482\" target=\"_blank\">Model Soup paper</a></li>\n<li><a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/3d_unet_monai/train.ipynb\" target=\"_blank\">Host's example notebook</a></li>\n<li><a href=\"https://www.kaggle.com/code/fnands/baseline-unet-train-submit\" target=\"_blank\">@fnands notebook</a></li>\n</ul>\n<h2>Code</h2>\n<p><a href=\"https://github.com/IAmPara0x/czii-8th-solution\" target=\"_blank\">8th place solution of kaggle czii competition code github\n</a><br>\n<a href=\"https://www.kaggle.com/code/iamparadox/czii-final-sub-reproduce\" target=\"_blank\">Submission Notebook</a><br>\n<a href=\"https://www.kaggle.com/code/sirapoabchaikunsaeng/czii-final-sub-reproduce\" target=\"_blank\">Submission Notebook Clean Version</a></p>",
  "messages": [
    {
      "id": 3117017,
      "postDate": "2025-02-06T14:40:24.903Z",
      "content": "<p>First, we thank the competition host and Kaggle staff for organizing this competition. Below, we introduce the solution of the team I Cryo Everyteim -- <a href=\"https://www.kaggle.com/sirapoabchaikunsaeng\" target=\"_blank\">@sirapoabchaikunsaeng</a>, <a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a>, <a href=\"https://www.kaggle.com/itsuki9180\" target=\"_blank\">@itsuki9180</a>, <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> -- </p>\n<h2>Context</h2>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/overview\" target=\"_blank\">competition overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/data\" target=\"_blank\">competition data</a></li>\n</ul>\n<h2>Overview of the approach</h2>\n<p>Our final submissions consisted of four 3D U-Net model soups, trained with different model sizes, parameters, and training data to ensure strong model complementarity. All 3D U-Nets were trained using patch sizes of (128, 128, 128) but were inferred with patch sizes of (160, 384, 384), with a 25% overlap using gaussian reconstruction to handle border artifacts. Additionally, we employed geometric test-time augmentation (TTA), including flipping and transpose.</p>\n<h3>Models</h3>\n<p>Our models were originally based on the host's example notebook and <a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a> notebook . We utilized 3D U-Net architectures from the MONAI library, trained with patch sizes of (128, 128, 128).<br>\nThe models were 3 levels deep with strides of (2, 2, 1). We started with a simple model:</p>\n<pre><code>: \n: \n: \n</code></pre>\n<p>This single model with our inference strategy reached 0.744 in public leaderboard score.<br>\nLater, we trained more complex U-Nets with configs:</p>\n<pre><code>: ,\n: ,\n: ,\n</code></pre>\n<p>The model with above config alone achieved a score of 0.759 LB.</p>\n<pre><code>: ,\n: ,\n: ,\n</code></pre>\n<p>We also applied a dropout of 0.2 or 0.3 in the model.</p>\n<h3>Training</h3>\n<p>The models were pre-trained on six synthetic tomograms denoised with Gaussian denoising—specifically, the 'TS_0', 'TS_1', 'TS_10', 'TS_11', 'TS_12', and 'TS_13' tomograms.<br>\n<em>@sersasj note: Gaussian denoising was applied because it visually improved the WBP tomograms particle visualization and was easy to implement. I hypothesized that better results could be achieved with a more advanced denoiser, but attempting to code a model for denoising used too much of my Kaggle quota, so I gave up.</em><br>\nPretraining not only reduced the time needed for the models to learn the particles but also increased the LB score by roughly 0.01.<br>\nBoth pretraining on synthetic data and fine-tuning uses the following transformations:</p>\n<pre><code>Compose([\n   RandCropByLabelClassesd(\n       keys=[, ],\n       =,\n       spatial_size=[128, 128, 128],\n       =7,  # background,  all 6 classes\n       =16,\n   ),\n   RandFlipd(keys=[, ], =0.5, =0),\n   RandFlipd(keys=[, ], =0.5, =1),\n   RandFlipd(keys=[, ], =0.5, =2),\n])\n</code></pre>\n<p>We used MONAI's DiceCELoss and optimized the models with AdamW, employing a learning rate reduction on plateau and an initial learning rate of 1e-3.<br>\nWe also experimented with Exponential Moving Average (EMA), which showed good results. In the last 3 days of the competition, we trained models using other tomo types: \"denoised\", \"ctfdeconvolved\", \"isonetcorrected\". For almost the entire competition, we hadn't found any increase in LB strategies other than geometric data augmentation in training (flip and rotate); nevertheless, training with other tomo types yielded good results.</p>\n<h3>Validation strategy</h3>\n<p>We primarily relied on out-of-fold predictions from our k-fold models. Specifically, we implemented a 7-fold cross-validation approach where we trained on all tomographies except one, which was used as a validation set. This process was repeated seven times, with each fold serving as the validation set once. This helped us obtain a score that correlated well with the public leaderboard score, typically 0.01 to 0.02 points higher.<br>\nOne limitation we encountered was our inability to apply this validation strategy when testing ensembles of our model soups. In these cases, we had to rely on leaked data scores. While this approach didn't correlate as strongly with leaderboard scores, most of the time it allowed us to rank the performance of different ensembles, and compare TTA strategies. </p>\n<h2>Details of the submission</h2>\n<h3>Inference</h3>\n<p>During inference, we use MONAI's sliding_window_inference with 25% overlap and Gaussian reconstruction to handle border predictions more accurately; this boosted the public LB by around 0.002. <strong>One of the greatest findings was making the inference in patches as large as possible; we used a window size of (160, 384, 384), boosting our score by around 0.01 compared to a size of (128, 128, 128).</strong> (Thanks <a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a> for finding this)<br>\nWe also experimented with combining different sizes of predictions, for example (160, 384, 384) + (136, 248, 248), but inference took much longer while yielding inconclusive gains. Therefore, we abandoned this approach in favor of creating more ensembles and employing more TTA.<br>\nFor post-processing, we improved particle detection by combining model predictions at the logits level rather than averaging probabilities directly.<br>\nWe then refined particle identification using watershed segmentation, which helped separate touching or overlapping particles. The watershed process created binary masks using particle-specific certainty thresholds, computed distance transforms to measure particle separation, identified distinct particles using local maxima as markers, and finally applied watershed segmentation to determine particle boundaries. We switched from using skimage to cucim and cupy to decrease the inference time. <br>\nWe also used specific blob threshold sizes for each particle to decrease the number of false positives in our predictions.<br>\nMoreover, we used flip along the X and Y axes and transpose TTA.</p>\n<h3>Ensembles</h3>\n<p>Ensembling was always one of the main ideas in our team. When the team was fully formed 2 weeks before the deadline, we were all around a 0.73 LB score, around 50th place on the leaderboard. At the time, I (@sersasj) had a YOLO scoring 0.682 and a U-Net scoring 0.692 that, when ensembled, reached 0.732 in LB. The rest of the team had only a 3D U-Net that reached 0.739. So it was obvious that we would ensemble everything and achieve amazing results; sadly, it didn't work.<br>\nNevertheless, we learned valuable lessons from this experience. My 3D U-Net was ~30 times smaller than the 0.739 scoring U-Net, but it proved quite effective when used with the team's amazing inference techniques: <em>larger patches, blob thresholds, ensemble at logits level, and watershed processing</em>. After implementing pretraining and additional training, it achieved a 0.744 leaderboard score. <br>\nThanks to our group work (especially <a href=\"https://www.kaggle.com/sirapoabchaikunsaeng\" target=\"_blank\">@sirapoabchaikunsaeng</a> and <a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a>), we successfully trained several U-Net variants derived from the 0.744 model by modifying parameters, implementing EMA (Exponential Moving Average), incorporating more data, and applying the techniques mentioned above.<br>\nBy doing so, we had various U-Nets that were complementary to each other, but we couldn't use them all. To manage \"using\" all we had, <a href=\"https://www.kaggle.com/itsuki9180\" target=\"_blank\">@itsuki9180</a> presented the idea and code for a technique called model soup, which essentially averages the weights of multiple models from the same pre-trained weights.<br>\nWe applied this to the models from our k-fold training.<br>\nOur final ensemble consisted of:</p>\n<ol>\n<li>A soup of 7 folds trained with parameters:<ul>\n<li>Channels: (32, 64, 128, 128)</li>\n<li>Number of residual units: 1</li>\n<li>Pretrained on synthetic data</li>\n<li>Trained on denoised images without EMA</li></ul></li>\n<li>A soup of 7 folds trained with parameters:<ul>\n<li>Channels: (32, 64, 128, 256)</li>\n<li>Number of residual units: 2</li>\n<li>Pretrained on synthetic data</li>\n<li>Trained on all denoised, IsoNet-corrected, and CTF-deconvolved images without EMA</li></ul></li>\n<li>A soup of 3 folds (TS_69_2, TS_86_3, TS_99_9) trained with parameters:<ul>\n<li>Channels: (32, 96, 256, 384)</li>\n<li>Number of residual units: 2</li>\n<li>Pretrained on synthetic data</li>\n<li>Trained on denoised images with EMA</li></ul></li>\n<li>A soup of 7 folds trained with parameters:<ul>\n<li>Channels: (32, 96, 256, 384)</li>\n<li>Number of residual units: 2</li>\n<li>Pretrained on synthetic data</li>\n<li>Trained on all denoised, IsoNet-corrected, and CTF-deconvolved images without EMA</li></ul></li>\n</ol>\n<h3>What didn't work:</h3>\n<ul>\n<li>Increased Augmentations</li>\n<li>CutMix, MixUp</li>\n<li>2D U-Net approaches (reached 0.636 lb only)</li>\n<li>Ensembling with YOLO (for final solution, works well to reach ~0.73 lb)</li>\n<li>Bigger and deeper U-Nets</li>\n<li>Cross entropy only loss functions</li>\n</ul>\n<h2>Sources section</h2>\n<ul>\n<li><a href=\"https://arxiv.org/abs/2203.05482\" target=\"_blank\">Model Soup paper</a></li>\n<li><a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/3d_unet_monai/train.ipynb\" target=\"_blank\">Host's example notebook</a></li>\n<li><a href=\"https://www.kaggle.com/code/fnands/baseline-unet-train-submit\" target=\"_blank\">@fnands notebook</a></li>\n</ul>\n<h2>Code</h2>\n<p><a href=\"https://github.com/IAmPara0x/czii-8th-solution\" target=\"_blank\">8th place solution of kaggle czii competition code github\n</a><br>\n<a href=\"https://www.kaggle.com/code/iamparadox/czii-final-sub-reproduce\" target=\"_blank\">Submission Notebook</a><br>\n<a href=\"https://www.kaggle.com/code/sirapoabchaikunsaeng/czii-final-sub-reproduce\" target=\"_blank\">Submission Notebook Clean Version</a></p>",
      "rawMarkdown": "\nFirst, we thank the competition host and Kaggle staff for organizing this competition. Below, we introduce the solution of the team I Cryo Everyteim -- @sirapoabchaikunsaeng, @iamparadox, @itsuki9180, @sersasj -- \n\n## Context\n\n- Business context: [competition overview](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/overview)\n- Data context: [competition data](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/data)\n\n\n## Overview of the approach\n\nOur final submissions consisted of four 3D U-Net model soups, trained with different model sizes, parameters, and training data to ensure strong model complementarity. All 3D U-Nets were trained using patch sizes of (128, 128, 128) but were inferred with patch sizes of (160, 384, 384), with a 25% overlap using gaussian reconstruction to handle border artifacts. Additionally, we employed geometric test-time augmentation (TTA), including flipping and transpose.\n\n### Models \n\nOur models were originally based on the host's example notebook and @fnands notebook . We utilized 3D U-Net architectures from the MONAI library, trained with patch sizes of (128, 128, 128).\n\nThe models were 3 levels deep with strides of (2, 2, 1). We started with a simple model:\n\n```\n\"channels\": (32, 64, 128, 128)\n\"strides\": (2, 2, 1)\n\"num_res_units\": 1\n```\n\nThis single model with our inference strategy reached 0.744 in public leaderboard score.\n\nLater, we trained more complex U-Nets with configs:\n\n```\n\"channels\": (32, 64, 128, 256),\n\"strides\": (2, 2, 1),\n\"num_res_units\": 2,\n```\nThe model with above config alone achieved a score of 0.759 LB.\n\n```\n\"channels\": (32, 96, 256, 384),\n\"strides\": (2, 2, 1),\n\"num_res_units\": 2,\n```\n\nWe also applied a dropout of 0.2 or 0.3 in the model.\n\n### Training\n\nThe models were pre-trained on six synthetic tomograms denoised with Gaussian denoising—specifically, the 'TS_0', 'TS_1', 'TS_10', 'TS_11', 'TS_12', and 'TS_13' tomograms.\n\n*@sersasj note: Gaussian denoising was applied because it visually improved the WBP tomograms particle visualization and was easy to implement. I hypothesized that better results could be achieved with a more advanced denoiser, but attempting to code a model for denoising used too much of my Kaggle quota, so I gave up.*\n\nPretraining not only reduced the time needed for the models to learn the particles but also increased the LB score by roughly 0.01.\n\nBoth pretraining on synthetic data and fine-tuning uses the following transformations:\n```\nCompose([\n    RandCropByLabelClassesd(\n        keys=[\"image\", \"label\"],\n        label_key=\"label\",\n        spatial_size=[128, 128, 128],\n        num_classes=7,  # background, and all 6 classes\n        num_samples=16,\n    ),\n    RandFlipd(keys=[\"image\", \"label\"], prob=0.5, spatial_axis=0),\n    RandFlipd(keys=[\"image\", \"label\"], prob=0.5, spatial_axis=1),\n    RandFlipd(keys=[\"image\", \"label\"], prob=0.5, spatial_axis=2),\n])\n```\n\nWe used MONAI's DiceCELoss and optimized the models with AdamW, employing a learning rate reduction on plateau and an initial learning rate of 1e-3.\n\nWe also experimented with Exponential Moving Average (EMA), which showed good results. In the last 3 days of the competition, we trained models using other tomo types: \"denoised\", \"ctfdeconvolved\", \"isonetcorrected\". For almost the entire competition, we hadn't found any increase in LB strategies other than geometric data augmentation in training (flip and rotate); nevertheless, training with other tomo types yielded good results.\n\n### Validation strategy\n\nWe primarily relied on out-of-fold predictions from our k-fold models. Specifically, we implemented a 7-fold cross-validation approach where we trained on all tomographies except one, which was used as a validation set. This process was repeated seven times, with each fold serving as the validation set once. This helped us obtain a score that correlated well with the public leaderboard score, typically 0.01 to 0.02 points higher.\n\nOne limitation we encountered was our inability to apply this validation strategy when testing ensembles of our model soups. In these cases, we had to rely on leaked data scores. While this approach didn't correlate as strongly with leaderboard scores, most of the time it allowed us to rank the performance of different ensembles, and compare TTA strategies. \n\n## Details of the submission\n\n### Inference\n\nDuring inference, we use MONAI's sliding_window_inference with 25% overlap and Gaussian reconstruction to handle border predictions more accurately; this boosted the public LB by around 0.002. **One of the greatest findings was making the inference in patches as large as possible; we used a window size of (160, 384, 384), boosting our score by around 0.01 compared to a size of (128, 128, 128).** (Thanks @iamparadox for finding this)\n\nWe also experimented with combining different sizes of predictions, for example (160, 384, 384) + (136, 248, 248), but inference took much longer while yielding inconclusive gains. Therefore, we abandoned this approach in favor of creating more ensembles and employing more TTA.\n\nFor post-processing, we improved particle detection by combining model predictions at the logits level rather than averaging probabilities directly.\n\nWe then refined particle identification using watershed segmentation, which helped separate touching or overlapping particles. The watershed process created binary masks using particle-specific certainty thresholds, computed distance transforms to measure particle separation, identified distinct particles using local maxima as markers, and finally applied watershed segmentation to determine particle boundaries. We switched from using skimage to cucim and cupy to decrease the inference time. \n\nWe also used specific blob threshold sizes for each particle to decrease the number of false positives in our predictions.\n\nMoreover, we used flip along the X and Y axes and transpose TTA.\n   \n### Ensembles\n\nEnsembling was always one of the main ideas in our team. When the team was fully formed 2 weeks before the deadline, we were all around a 0.73 LB score, around 50th place on the leaderboard. At the time, I (@sersasj) had a YOLO scoring 0.682 and a U-Net scoring 0.692 that, when ensembled, reached 0.732 in LB. The rest of the team had only a 3D U-Net that reached 0.739. So it was obvious that we would ensemble everything and achieve amazing results; sadly, it didn't work.\n\nNevertheless, we learned valuable lessons from this experience. My 3D U-Net was ~30 times smaller than the 0.739 scoring U-Net, but it proved quite effective when used with the team's amazing inference techniques: *larger patches, blob thresholds, ensemble at logits level, and watershed processing*. After implementing pretraining and additional training, it achieved a 0.744 leaderboard score. \n\nThanks to our group work (especially @sirapoabchaikunsaeng and @iamparadox), we successfully trained several U-Net variants derived from the 0.744 model by modifying parameters, implementing EMA (Exponential Moving Average), incorporating more data, and applying the techniques mentioned above.\n\nBy doing so, we had various U-Nets that were complementary to each other, but we couldn't use them all. To manage \"using\" all we had, @itsuki9180 presented the idea and code for a technique called model soup, which essentially averages the weights of multiple models from the same pre-trained weights.\nWe applied this to the models from our k-fold training.\n\nOur final ensemble consisted of:\n1. A soup of 7 folds trained with parameters:\n   - Channels: (32, 64, 128, 128)\n   - Number of residual units: 1\n   - Pretrained on synthetic data\n   - Trained on denoised images without EMA\n\n2. A soup of 7 folds trained with parameters:\n   - Channels: (32, 64, 128, 256)\n   - Number of residual units: 2\n   - Pretrained on synthetic data\n   - Trained on all denoised, IsoNet-corrected, and CTF-deconvolved images without EMA\n\n3. A soup of 3 folds (TS_69_2, TS_86_3, TS_99_9) trained with parameters:\n   - Channels: (32, 96, 256, 384)\n   - Number of residual units: 2\n   - Pretrained on synthetic data\n   - Trained on denoised images with EMA\n\n4. A soup of 7 folds trained with parameters:\n   - Channels: (32, 96, 256, 384)\n   - Number of residual units: 2\n   - Pretrained on synthetic data\n   - Trained on all denoised, IsoNet-corrected, and CTF-deconvolved images without EMA\n\n### What didn't work:\n- Increased Augmentations\n- CutMix, MixUp\n- 2D U-Net approaches (reached 0.636 lb only)\n- Ensembling with YOLO (for final solution, works well to reach ~0.73 lb)\n- Bigger and deeper U-Nets\n- Cross entropy only loss functions\n\n\n## Sources section\n\n- [Model Soup paper](https://arxiv.org/abs/2203.05482)\n- [Host's example notebook](https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/3d_unet_monai/train.ipynb)\n- [@fnands notebook](https://www.kaggle.com/code/fnands/baseline-unet-train-submit)\n\n\n## Code\n\n[8th place solution of kaggle czii competition code github\n](https://github.com/IAmPara0x/czii-8th-solution)\n[Submission Notebook](https://www.kaggle.com/code/iamparadox/czii-final-sub-reproduce)\n[Submission Notebook Clean Version](https://www.kaggle.com/code/sirapoabchaikunsaeng/czii-final-sub-reproduce)",
      "votes": 27
    },
    {
      "id": 3118576,
      "postDate": "2025-02-08T08:14:54.210Z",
      "content": "<p>Congratulations to your team, and very glad to see you <a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a> winning a gold medal.</p>\n<p>I noticed a post on Xiaohongshu regarding accusations against one of your team members. Many Chinese kagglers in the discussion group echoed similar experiences. I don't know the private communications within your team regarding this matter. Here, I am posting the screenshot to remind everyone to be cautious in choosing teammates when teaming up. Of course, if the relevant accusations are proven to be untrue, I am willing to publicly apologize to you <a href=\"https://www.kaggle.com/yinhewang\" target=\"_blank\">@yinhewang</a> .</p>\n<p>This may affect the relationships among your team members, and for that I sincerely apologize.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18538441%2F29191241e615e6c202092157f3203a9f%2FWechatIMG20590.jpg?generation=1739001808068077&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Congratulations to your team, and very glad to see you @iamparadox winning a gold medal.\n\nI noticed a post on Xiaohongshu regarding accusations against one of your team members. Many Chinese kagglers in the discussion group echoed similar experiences. I don't know the private communications within your team regarding this matter. Here, I am posting the screenshot to remind everyone to be cautious in choosing teammates when teaming up. Of course, if the relevant accusations are proven to be untrue, I am willing to publicly apologize to you @yinhewang .\n\nThis may affect the relationships among your team members, and for that I sincerely apologize.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18538441%2F29191241e615e6c202092157f3203a9f%2FWechatIMG20590.jpg?generation=1739001808068077&alt=media)",
      "votes": 21,
      "replies": [
        {
          "id": 3118589,
          "postDate": "2025-02-08T08:38:07.860Z",
          "content": "<p>The translation:<br>\n\"Bro, you really managed to fool people. This person contacted me to team up on Kaggle two months ago. At that time, I was ranked relatively high on the leaderboard for a silver medal in one competition. I didn't think much about it. He said he had several L20s to run experiments. But after teaming up, he just disappeared, casting a wide net in various competitions. Then other people from different competitions approached me, asking if this person was just a freeloader doing no work. Now that CZII has ended, you really managed to get yourself a gold medal. I don't know if you disappeared in this team as well. I just know that in their team's solution, four people were mentioned, but not you. Foreigners really don’t understand just how cunning some Chinese people can be, huh?\"</p>",
          "rawMarkdown": "The translation:\n\"Bro, you really managed to fool people. This person contacted me to team up on Kaggle two months ago. At that time, I was ranked relatively high on the leaderboard for a silver medal in one competition. I didn't think much about it. He said he had several L20s to run experiments. But after teaming up, he just disappeared, casting a wide net in various competitions. Then other people from different competitions approached me, asking if this person was just a freeloader doing no work. Now that CZII has ended, you really managed to get yourself a gold medal. I don't know if you disappeared in this team as well. I just know that in their team's solution, four people were mentioned, but not you. Foreigners really don’t understand just how cunning some Chinese people can be, huh?\"",
          "votes": 12,
          "replies": [
            {
              "id": 3118693,
              "postDate": "2025-02-08T11:10:37.887Z",
              "content": "<p>I am sorry to hear about this. As a Chinese, I feel so ashamed of such inappropriate behavior. However, it is important to acknowledge that many Chinese individuals work diligently and collaborate effectively in teamworks. The key challenge is managing cunning and fool people who may ghost team members or steal your contributions. Therefore, it is essential to remain cautious and careful, and in certain situations, verification may be necessary.</p>",
              "rawMarkdown": "I am sorry to hear about this. As a Chinese, I feel so ashamed of such inappropriate behavior. However, it is important to acknowledge that many Chinese individuals work diligently and collaborate effectively in teamworks. The key challenge is managing cunning and fool people who may ghost team members or steal your contributions. Therefore, it is essential to remain cautious and careful, and in certain situations, verification may be necessary.",
              "votes": 3
            }
          ]
        },
        {
          "id": 3118663,
          "postDate": "2025-02-08T10:26:51.150Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sweetyheehee\" target=\"_blank\">@sweetyheehee</a>, we are aware of <a href=\"https://www.kaggle.com/yinhewang\" target=\"_blank\">@yinhewang</a>'s situation. He had no involvement in winning the gold medal. Essentially, he ghosted us from the beginning and only returned after we won.</p>\n<p>We tried to investigate, but things started getting very complicated and weird, many lies were appearing. We also reached out to other teams.</p>\n<p>We sent an email to Kaggle and were waiting for a response before going public and starting a discussion on Kaggle with all the proof we gathered.</p>\n<p>I entered right before merge so don't have all info to share. But my teammates will come here later.</p>",
          "rawMarkdown": "Hi @sweetyheehee, we are aware of @yinhewang's situation. He had no involvement in winning the gold medal. Essentially, he ghosted us from the beginning and only returned after we won.\n\nWe tried to investigate, but things started getting very complicated and weird, many lies were appearing. We also reached out to other teams.\n\nWe sent an email to Kaggle and were waiting for a response before going public and starting a discussion on Kaggle with all the proof we gathered.\n\nI entered right before merge so don't have all info to share. But my teammates will come here later.\n\n\n\n",
          "votes": 6,
          "replies": [
            {
              "id": 3118681,
              "postDate": "2025-02-08T10:54:41.173Z",
              "content": "<p>Thanks for your feedback <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> , </p>\n<p>The reason I made this post is that many Chinese kagglers are already discussing his alleged misconduct. We don't want more people to be deceived, and we also hope that he can clarify these matters. However, his account has already been deleted, which suggests that the previous accusations are true. </p>",
              "rawMarkdown": "Thanks for your feedback @sersasj , \n\nThe reason I made this post is that many Chinese kagglers are already discussing his alleged misconduct. We don't want more people to be deceived, and we also hope that he can clarify these matters. However, his account has already been deleted, which suggests that the previous accusations are true. ",
              "votes": 3
            }
          ]
        },
        {
          "id": 3118729,
          "postDate": "2025-02-08T12:24:37.533Z",
          "content": "<p>As a Chinese, I feel ashamed of this behavior, but more importantly, this behavior has seriously damaged the friendly and cooperative atmosphere of the kaggle community. Therefore, I would like to warn everyone to be careful when facing the application of unknown person to join the team, especially for such low-activity accounts. It is hoped that through the joint efforts of community members, the kaggle community can be maintained in a positive and healthy direction. Finally, I would like to apologize again for this happening in the kaggle community.</p>\n<p>There is a Chinese poem that I think is a powerful criticism of such people：<br>\n“小人者其未得也，则忧不得； 既已得之，又恐失之。是以有终身之忧，无一日之乐也。”</p>",
          "rawMarkdown": "As a Chinese, I feel ashamed of this behavior, but more importantly, this behavior has seriously damaged the friendly and cooperative atmosphere of the kaggle community. Therefore, I would like to warn everyone to be careful when facing the application of unknown person to join the team, especially for such low-activity accounts. It is hoped that through the joint efforts of community members, the kaggle community can be maintained in a positive and healthy direction. Finally, I would like to apologize again for this happening in the kaggle community.\n\nThere is a Chinese poem that I think is a powerful criticism of such people：\n“小人者其未得也，则忧不得； 既已得之，又恐失之。是以有终身之忧，无一日之乐也。”\n",
          "votes": 4
        },
        {
          "id": 3118734,
          "postDate": "2025-02-08T12:33:27.037Z",
          "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/sweetyheehee\" target=\"_blank\">@sweetyheehee</a>, as <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> mentioned, we have started our investigations ourselves and have found a lot of evidence that points towards <a href=\"https://www.kaggle.com/yinhewang\" target=\"_blank\">@yinhewang</a> deliberately leeching off of teams.</p>\n<p>From our side we have submitted an inquiry to kaggle, which they have forwarded to kaggle's compliance team. I'm not sure if this is correlated with <a href=\"https://www.kaggle.com/yinhewang\" target=\"_blank\">@yinhewang</a>'s account being deleted or not. But as of the time I post this comment, the account is gone <a href=\"https://www.kaggle.com/yinhewang\" target=\"_blank\">https://www.kaggle.com/yinhewang</a></p>\n<p>In addition to this, we have evidence that <a href=\"https://www.kaggle.com/yinhewang\" target=\"_blank\">@yinhewang</a> is also planning to even take parts of our prize money by submitting the winning team form. We are not sure what additional steps we can take to guarantee this doesn't happen. But we do have couple of written evidence from him that he is willing to not take any part in prize money.</p>",
          "rawMarkdown": "Thank you very much @sweetyheehee, as @sersasj mentioned, we have started our investigations ourselves and have found a lot of evidence that points towards @yinhewang deliberately leeching off of teams.\n\nFrom our side we have submitted an inquiry to kaggle, which they have forwarded to kaggle's compliance team. I'm not sure if this is correlated with @yinhewang's account being deleted or not. But as of the time I post this comment, the account is gone https://www.kaggle.com/yinhewang\n\nIn addition to this, we have evidence that @yinhewang is also planning to even take parts of our prize money by submitting the winning team form. We are not sure what additional steps we can take to guarantee this doesn't happen. But we do have couple of written evidence from him that he is willing to not take any part in prize money.",
          "votes": 3,
          "replies": [
            {
              "id": 3118757,
              "postDate": "2025-02-08T13:11:34.840Z",
              "content": "<p>I've made a post in the general thread to try bring more attention to this issue <a href=\"https://www.kaggle.com/discussions/general/561879\" target=\"_blank\">https://www.kaggle.com/discussions/general/561879</a></p>",
              "rawMarkdown": "I've made a post in the general thread to try bring more attention to this issue https://www.kaggle.com/discussions/general/561879",
              "votes": 4
            },
            {
              "id": 3118769,
              "postDate": "2025-02-08T13:29:42.593Z",
              "content": "<p>Thank you, Sumo. I will forward your post to my friends as well as his former teammates.</p>",
              "rawMarkdown": "Thank you, Sumo. I will forward your post to my friends as well as his former teammates.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3117418,
      "postDate": "2025-02-07T01:27:27.053Z",
      "content": "<p>I am super excited to team up with <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> , <a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a> and <a href=\"https://www.kaggle.com/sirapoabchaikunsaeng\" target=\"_blank\">@sirapoabchaikunsaeng</a>  to complete this competition! They are all wonderful mates who are very proactive in discussions and didn't hesitate to try out my crazy ideas.<br>\n<a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> picked out and summarized the ideas that worked for me, but in this post I'll cover the ones that didn't.</p>\n<h3>What I tried that didn't work or didn't used in submission</h3>\n<ol>\n<li>YOLO approach. I believe many competitors created their solutions based on my baseline, but our team ultimately chose a solution based primarily on 3D-UNet. Therefore, I believe that the solution for YOLO is not worthless, so I will describe it here. In this method, the following preprocessing was applied:<br>\n・For one z, a 3-channel image (z-1, z, z+1) is used instead of a 1-channel image (z, z, z).<br>\n・Scale the image to the range [-1, 1] and then raise the pixel values ​​to the p-th power (approximately 1&lt;p&lt;=3) <br>\n・For min-max scaling, used a 3-channel image with clip values ​​of (0.0001%, 2%), (2%, 98%), and (98%, 99.999%).<br>\nThe second and third ideas take into account the fact that pixels around the target have small values, as considered by <a href=\"https://www.kaggle.com/sirapoabchaikunsaeng\" target=\"_blank\">@sirapoabchaikunsaeng</a> . They improved LB 0.015~0.03.<br>\nThere were few changes to the training. As a tip, less strong data augmentation tends to produce better scores.<br>\nThere are various ensemble and postprocessing methods for object detection, but I used Weighted Box Fusion. Due to its characteristics, it is necessary to lower the threshold score a little. Applying the techniques above to the publicly available high-scoring notebooks would have been enough to win a bronze medal, but we weren't satisfied with that.</li>\n<li>Data generation using ControlNet w/ Stable Diffusion<br>\nAs mentioned in my other discussion posts, I thought the bad scores were due to a lack of training data, so I experimented with generating slice images from the GT of the segmentation map (i.e. the inverse of segmentation). This didn't work at all, and produced terrible results with mode collapse.</li>\n<li>Use a larger volume<br>\nI think we all know that a fundamental principle of image processing is that the larger the training image, the better the performance tends to be. So I scaled the (184, 630, 630) volume by a factor of 1.5 and trained the (276, 945, 945) one. This also didn't work well (around LB 0.7), Also, because the training took so long, there was no time to do any in-depth research.</li>\n</ol>\n<h3>Things I wanted to try but couldn't</h3>\n<ol>\n<li>Auxiliary task and aux loss<br>\nThe test dataset only has the denoising method, but the training dataset has other process methods. For example, we can take the denoised image as input and add an auxiliary task of reconstructing the ctfdeconvolved to the segmentation task. Multi-task learning may lead to faster training or better accuracy. I wanted to try this, but I didn't have time.</li>\n<li>Contrastive Learning<br>\nI considered contrastive learning with two input images. We suggested an auxiliary task that would give a positive label when the second image was an augmentation or identity of the first image, and a negative label when the first and second images were different, but we didn't have time to do this either.</li>\n</ol>\n<h3>Summary</h3>\n<p>Many of my suggestions went to waste, but I'm still happy to have contributed even a small amount. I ended up not being of much use to the team, so I would like to take this opportunity to thank the team once again.</p>\n<h3>Finally some words</h3>\n<p>In this post, I will first state what I want to say. <strong>Most of our team's experiments (especially mine) have been failures</strong>, Maybe, as have most of the other gold medal winners. <strong>Don't be afraid of failing an experiment or not winning a medal in a competition.</strong> Don't give up just because you lost one competition. Kaggle <strong>NEVER</strong> penalize you for losing. Your competitors are also your best companions to expand your knowledge. Don't stop learning from them, and one day you will become a great Kaggler.<br>\nThere is a Japanese-English word, \"no side\" It is a word used to praise both sides for their good fight after the game is over, regardless of whether they are allies or opponents. This is getting long, but my final word is…<br>\n<strong>NO SIDE!</strong></p>",
      "rawMarkdown": "I am super excited to team up with @sersasj , @iamparadox and @sirapoabchaikunsaeng  to complete this competition! They are all wonderful mates who are very proactive in discussions and didn't hesitate to try out my crazy ideas.\n@sersasj picked out and summarized the ideas that worked for me, but in this post I'll cover the ones that didn't.\n### What I tried that didn't work or didn't used in submission\n1. YOLO approach. I believe many competitors created their solutions based on my baseline, but our team ultimately chose a solution based primarily on 3D-UNet. Therefore, I believe that the solution for YOLO is not worthless, so I will describe it here. In this method, the following preprocessing was applied:\n・For one z, a 3-channel image (z-1, z, z+1) is used instead of a 1-channel image (z, z, z).\n・Scale the image to the range [-1, 1] and then raise the pixel values ​​to the p-th power (approximately 1<p<=3) \n・For min-max scaling, used a 3-channel image with clip values ​​of (0.0001%, 2%), (2%, 98%), and (98%, 99.999%).\nThe second and third ideas take into account the fact that pixels around the target have small values, as considered by @sirapoabchaikunsaeng . They improved LB 0.015~0.03.\nThere were few changes to the training. As a tip, less strong data augmentation tends to produce better scores.\nThere are various ensemble and postprocessing methods for object detection, but I used Weighted Box Fusion. Due to its characteristics, it is necessary to lower the threshold score a little. Applying the techniques above to the publicly available high-scoring notebooks would have been enough to win a bronze medal, but we weren't satisfied with that.\n2. Data generation using ControlNet w/ Stable Diffusion\nAs mentioned in my other discussion posts, I thought the bad scores were due to a lack of training data, so I experimented with generating slice images from the GT of the segmentation map (i.e. the inverse of segmentation). This didn't work at all, and produced terrible results with mode collapse.\n3. Use a larger volume\nI think we all know that a fundamental principle of image processing is that the larger the training image, the better the performance tends to be. So I scaled the (184, 630, 630) volume by a factor of 1.5 and trained the (276, 945, 945) one. This also didn't work well (around LB 0.7), Also, because the training took so long, there was no time to do any in-depth research.\n\n### Things I wanted to try but couldn't\n1. Auxiliary task and aux loss\nThe test dataset only has the denoising method, but the training dataset has other process methods. For example, we can take the denoised image as input and add an auxiliary task of reconstructing the ctfdeconvolved to the segmentation task. Multi-task learning may lead to faster training or better accuracy. I wanted to try this, but I didn't have time.\n2. Contrastive Learning\nI considered contrastive learning with two input images. We suggested an auxiliary task that would give a positive label when the second image was an augmentation or identity of the first image, and a negative label when the first and second images were different, but we didn't have time to do this either.\n\n### Summary\nMany of my suggestions went to waste, but I'm still happy to have contributed even a small amount. I ended up not being of much use to the team, so I would like to take this opportunity to thank the team once again.\n\n### Finally some words\nIn this post, I will first state what I want to say. **Most of our team's experiments (especially mine) have been failures**, Maybe, as have most of the other gold medal winners. **Don't be afraid of failing an experiment or not winning a medal in a competition.** Don't give up just because you lost one competition. Kaggle **NEVER** penalize you for losing. Your competitors are also your best companions to expand your knowledge. Don't stop learning from them, and one day you will become a great Kaggler.\nThere is a Japanese-English word, \"no side\" It is a word used to praise both sides for their good fight after the game is over, regardless of whether they are allies or opponents. This is getting long, but my final word is...\n**NO SIDE!**",
      "votes": 10,
      "replies": [
        {
          "id": 3117798,
          "postDate": "2025-02-07T08:58:29.360Z",
          "content": "<p><a href=\"https://www.kaggle.com/ITK8191\" target=\"_blank\">@ITK8191</a> <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a><br>\nThank you for your notebooks </p>\n<p>I had never used YOLO before, so I continued submitting in YOLO for practice. However, I did not do a team merge because I had doubts about my YOLO license and could have been expelled.</p>\n<p>One of my innovations was the calculation of Z coordinates.<br>\nWhen calculating the center coordinates in KDTree, only the Z coordinate is calculated with confidence as a weight. This improved the LB by about 0.005</p>\n<p>particle_confidences = pdf['confidence'].tolist()<br>\nparticle_xs = pdf['x'].tolist()<br>\nparticle_ys = pdf['y'].tolist()<br>\nparticle_zs = pdf['z'].tolist()<br>\nparticle_zs_conf = (pdf['z']*pdf['confidence']).tolist()<br>\n～<br>\ncenters_x = np.bincount(inverse_indices, weights=particle_xs) / counts<br>\ncenters_y = np.bincount(inverse_indices, weights=particle_ys) / counts<br>\n# centers_z = np.bincount(inverse_indices, weights=particle_yz) / counts<br>\ncenters_z = np.bincount(inverse_indices, weights=particle_zs_conf) / conf_sums</p>",
          "rawMarkdown": "@ITK8191 @sersasj\nThank you for your notebooks \n\nI had never used YOLO before, so I continued submitting in YOLO for practice. However, I did not do a team merge because I had doubts about my YOLO license and could have been expelled.\n\nOne of my innovations was the calculation of Z coordinates.\nWhen calculating the center coordinates in KDTree, only the Z coordinate is calculated with confidence as a weight. This improved the LB by about 0.005\n\nparticle_confidences = pdf['confidence'].tolist()\nparticle_xs = pdf['x'].tolist()\nparticle_ys = pdf['y'].tolist()\nparticle_zs = pdf['z'].tolist()\nparticle_zs_conf = (pdf['z']*pdf['confidence']).tolist()\n～\ncenters_x = np.bincount(inverse_indices, weights=particle_xs) / counts\ncenters_y = np.bincount(inverse_indices, weights=particle_ys) / counts\n\\# centers_z = np.bincount(inverse_indices, weights=particle_yz) / counts\ncenters_z = np.bincount(inverse_indices, weights=particle_zs_conf) / conf_sums",
          "votes": 3
        }
      ]
    },
    {
      "id": 3117540,
      "postDate": "2025-02-07T03:49:07.167Z",
      "content": "<p>Congrats all of you!  Very nice job!  Thanks for the great explanation!</p>",
      "rawMarkdown": "Congrats all of you!  Very nice job!  Thanks for the great explanation!",
      "votes": 4,
      "replies": [
        {
          "id": 3117574,
          "postDate": "2025-02-07T04:15:37.163Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a> congrats on your gold as well!!, You posts really helped us during the competition!</p>",
          "rawMarkdown": "Thanks @davidlist congrats on your gold as well!!, You posts really helped us during the competition!",
          "votes": 2,
          "replies": [
            {
              "id": 3117809,
              "postDate": "2025-02-07T09:06:46.490Z",
              "content": "<p>Yeah, I may have to tone that down a little next time.  Too many people ahead of me thanking me.  😀  (Kidding, of course.)</p>",
              "rawMarkdown": "Yeah, I may have to tone that down a little next time.  Too many people ahead of me thanking me.  😀  (Kidding, of course.)",
              "votes": 2
            },
            {
              "id": 3118039,
              "postDate": "2025-02-07T14:05:54.540Z",
              "content": "<p>People behind doing the same 😆 That's the Kaggle spirit, and it's what makes these competitions so fun!</p>",
              "rawMarkdown": "People behind doing the same 😆 That's the Kaggle spirit, and it's what makes these competitions so fun!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3119442,
      "postDate": "2025-02-09T09:22:03.693Z",
      "content": "<p>Congratulations to your team for winning this game, and I am also sorry for yinhe wang's situation.<br>\nIn addition, I would like to know how you use the EMA strategy. When I tried basic EMA, it had almost no effect on improving my score. So I adjusted to update once after each round and tried using the dynamic deck method, which can achieve slight improvements in some cases. I saw that you said EMA has a good effect, so I would like to consult your usage strategy</p>",
      "rawMarkdown": "Congratulations to your team for winning this game, and I am also sorry for yinhe wang's situation.\nIn addition, I would like to know how you use the EMA strategy. When I tried basic EMA, it had almost no effect on improving my score. So I adjusted to update once after each round and tried using the dynamic deck method, which can achieve slight improvements in some cases. I saw that you said EMA has a good effect, so I would like to consult your usage strategy",
      "votes": 1
    },
    {
      "id": 3118117,
      "postDate": "2025-02-07T16:22:55.983Z",
      "content": "<p>Congratulations!</p>\n<p><a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a> <br>\nThis is interesting for me.</p>\n<blockquote>\n  <p>One of the greatest findings was making the inference in patches as large as possible; we used a window size of (160, 384, 384), boosting our score by around 0.01 compared to a size of (128, 128, 128).</p>\n</blockquote>\n<p>Do you train model with same patch size or not ?<br>\nIf not, I cannot understand why this improve score.<br>\nI would appreciate it if you could share any theoretical background.</p>",
      "rawMarkdown": "Congratulations!\n\n@iamparadox \nThis is interesting for me.\n>One of the greatest findings was making the inference in patches as large as possible; we used a window size of (160, 384, 384), boosting our score by around 0.01 compared to a size of (128, 128, 128).\n\nDo you train model with same patch size or not ?\nIf not, I cannot understand why this improve score.\nI would appreciate it if you could share any theoretical background.\n",
      "votes": 1,
      "replies": [
        {
          "id": 3118821,
          "postDate": "2025-02-08T14:16:10.763Z",
          "content": "<p>The model was only trained on 128x128x128 patch size but it was inferred on 160x384x384 patch size. I honestly don't know a lot of theoretical background on this topic, but the idea here's that the contextual information doesn't change while we are doing the inference, in a sense that suppose, a 3x3x3 conv filter expects high value in the center pixel of the input image to activate then this doesn't change even if I increase the patch size at inference and the second reason is doing inference on larger patchsize reduces the border artifacts. </p>\n<p>The large patchsize trick will not work if we downsample the image into smaller dimension, and then infer them on higher dimension because in this case the contextual information will change, i.e it will see finer details that it has never seen before.</p>",
          "rawMarkdown": "The model was only trained on 128x128x128 patch size but it was inferred on 160x384x384 patch size. I honestly don't know a lot of theoretical background on this topic, but the idea here's that the contextual information doesn't change while we are doing the inference, in a sense that suppose, a 3x3x3 conv filter expects high value in the center pixel of the input image to activate then this doesn't change even if I increase the patch size at inference and the second reason is doing inference on larger patchsize reduces the border artifacts. \n\nThe large patchsize trick will not work if we downsample the image into smaller dimension, and then infer them on higher dimension because in this case the contextual information will change, i.e it will see finer details that it has never seen before.",
          "votes": 3
        },
        {
          "id": 3119101,
          "postDate": "2025-02-08T21:49:04.200Z",
          "content": "<p><a href=\"https://www.kaggle.com/clearwaterkzk\" target=\"_blank\">@clearwaterkzk</a> One pretty big issue I found with the <a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a> notebook was the borders of each patch were different from the rest of the model output so you could see a grid pattern in the reconstructed model output.  You can see some evidence of this here: <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549615\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549615</a> which shows single patches.  I wonder if this is the same issue.  Making the patches bigger would reduce the problem.  I likewise saw a big performance increase by deleting the outer 3 rows of pixels from each patch before reconstruction and averaging a small overlap.  slinding_window_inference() makes this pretty straightforward.</p>",
          "rawMarkdown": "@clearwaterkzk One pretty big issue I found with the @fnands notebook was the borders of each patch were different from the rest of the model output so you could see a grid pattern in the reconstructed model output.  You can see some evidence of this here: [https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549615](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549615) which shows single patches.  I wonder if this is the same issue.  Making the patches bigger would reduce the problem.  I likewise saw a big performance increase by deleting the outer 3 rows of pixels from each patch before reconstruction and averaging a small overlap.  slinding_window_inference() makes this pretty straightforward.",
          "votes": 1,
          "replies": [
            {
              "id": 3119145,
              "postDate": "2025-02-08T23:27:50.210Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 3119146,
          "postDate": "2025-02-08T23:29:04.317Z",
          "content": "<p>Thanks for replying.<br>\nI also got the answer from itsuki9180 directly in Twitter(X) regarding this.</p>\n<p>I found the border issue and deal with it by padding the voxel and sliding window with a lot of overlap, in which only the center part are predicted, then averaging probability like following figure.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F88c71d74e91423a3c61ec3136ccd3c4d%2Fsliding_window.png?generation=1739056705345010&amp;alt=media\" alt=\"\"><br>\nyellow: original exp voxel size, black: patch size for inference, red: probability to use for reconstruction</p>\n<p>It seems that this method is used in some segmentation past competition like Vesuvius Challenge, but it requires more inference time. Your method seems smarter and faster.</p>\n<p>I used to have the fixed idea that CNNs must be inferred with the same size as they were trained on. However, after seeing this solution, I realized for the first time that convolution-only calculations can accept any input size. I learned a lot from this. Thank you very much.</p>",
          "rawMarkdown": "Thanks for replying.\nI also got the answer from itsuki9180 directly in Twitter(X) regarding this.\n\nI found the border issue and deal with it by padding the voxel and sliding window with a lot of overlap, in which only the center part are predicted, then averaging probability like following figure.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F88c71d74e91423a3c61ec3136ccd3c4d%2Fsliding_window.png?generation=1739056705345010&alt=media)\nyellow: original exp voxel size, black: patch size for inference, red: probability to use for reconstruction\n\n\nIt seems that this method is used in some segmentation past competition like Vesuvius Challenge, but it requires more inference time. Your method seems smarter and faster.\n\n\nI used to have the fixed idea that CNNs must be inferred with the same size as they were trained on. However, after seeing this solution, I realized for the first time that convolution-only calculations can accept any input size. I learned a lot from this. Thank you very much.",
          "votes": 1,
          "replies": [
            {
              "id": 3119158,
              "postDate": "2025-02-09T00:18:31.917Z",
              "content": "<p><a href=\"https://www.kaggle.com/clearwaterkzk\" target=\"_blank\">@clearwaterkzk</a> Just to shown the difference bigger patches made, I submitted our best submission with patches 128.<br>\nBest submission: public lb: 0.78022, private lb: 0.77612<br>\nPatches 128 submission: public lb: 0.77464, private lb: 0.76604</p>",
              "rawMarkdown": "@clearwaterkzk Just to shown the difference bigger patches made, I submitted our best submission with patches 128.\nBest submission: public lb: 0.78022, private lb: 0.77612\nPatches 128 submission: public lb: 0.77464, private lb: 0.76604",
              "votes": 1
            },
            {
              "id": 3119859,
              "postDate": "2025-02-09T19:06:42.357Z",
              "content": "<p>That's interesting, I was also under the impression that I must infer on same subvolume size I trained on, until I accidentally discovered that's not the case. However, I thought that somehow, maintaining isotropic subvolume sizes would be advantageous (I worked on 184x184x184 for most of the competition, started with 128x320x320), and I remember checking an experiment where I had my best submission perform better on the 184x184x184. I have just sent a few experiments on 128x320x320 to see if what <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> found stands true for my solution as well. <br>\nI understand that with bigger patches, there's less border artifacts and less merges to be made, but assuming a sizable overlap and gaussian weights for reconstruction, I was inclined to believe that the edges would have a poor enough weight to be negligible.</p>",
              "rawMarkdown": "That's interesting, I was also under the impression that I must infer on same subvolume size I trained on, until I accidentally discovered that's not the case. However, I thought that somehow, maintaining isotropic subvolume sizes would be advantageous (I worked on 184x184x184 for most of the competition, started with 128x320x320), and I remember checking an experiment where I had my best submission perform better on the 184x184x184. I have just sent a few experiments on 128x320x320 to see if what @sersasj found stands true for my solution as well. \nI understand that with bigger patches, there's less border artifacts and less merges to be made, but assuming a sizable overlap and gaussian weights for reconstruction, I was inclined to believe that the edges would have a poor enough weight to be negligible.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3118096,
      "postDate": "2025-02-07T15:26:24.123Z",
      "content": "<p>brilliant, interesting that augmentations didn't work at all. did you try mirroring?</p>",
      "rawMarkdown": "brilliant, interesting that augmentations didn't work at all. did you try mirroring?"
    },
    {
      "id": 3117093,
      "postDate": "2025-02-06T16:19:45.010Z",
      "content": "<p><a href=\"https://www.kaggle.com/rezaparaan\" target=\"_blank\">@rezaparaan</a>,<br>\nthank you again for the competition! we're very grateful for all the learnings we've had from this, and also all the friendships we've made (this is the first time we all team up), as well as designing the public and private LB such that there's no massive shakeup!</p>\n<p>this is the first time our team won gold, may I ask you to outline what will happen next? As in what should we provide you with (code / model weights / reports). So that I can make sure that I don't miss anything in the process.</p>\n<p>Thank you!</p>",
      "rawMarkdown": "@rezaparaan,\nthank you again for the competition! we're very grateful for all the learnings we've had from this, and also all the friendships we've made (this is the first time we all team up), as well as designing the public and private LB such that there's no massive shakeup!\n\nthis is the first time our team won gold, may I ask you to outline what will happen next? As in what should we provide you with (code / model weights / reports). So that I can make sure that I don't miss anything in the process.\n\nThank you!",
      "replies": [
        {
          "id": 3117094,
          "postDate": "2025-02-06T16:22:20.337Z",
          "content": "<p>cc <a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a>, provided you're also this competition host, apologies for any mistakes in tagging anyone</p>",
          "rawMarkdown": "cc @kharrington, provided you're also this competition host, apologies for any mistakes in tagging anyone"
        },
        {
          "id": 3119183,
          "postDate": "2025-02-09T01:28:33.623Z",
          "content": "<p>Hi Sumo, thank you so much for participating. We are thrilled to see the Kaggle community in action tackling problems that are important to us. I have copy-pasted this from the Rules section:</p>\n<blockquote>\n  <ol>\n  <li>WINNER’S OBLIGATIONS.<br>\n  a. As a condition to being awarded a Prize, a Prize winner must fulfill the following obligations:</li>\n  <li>Deliver to the Competition Sponsor the final model's software code as used to generate the winning Submission and associated documentation. The delivered software code should follow these documentation guidelines, must be capable of generating the winning Submission, and contain a description of resources required to build and/or run the executable code successfully. For avoidance of doubt, delivered software code should include training code, inference code, and a description of the required computational environment. To the extent that the final model’s software code includes generally commercially available software that is not owned by you, but that can be procured by the Competition Sponsor without undue expense, then instead of delivering the code for that software to the Competition Sponsor, you must identify that software, method for procuring it, and any parameters or other information necessary to replicate the winning Submission;</li>\n  <li>Grant to the Competition Sponsor the license to the winning Submission stated in the Competition Specific Rules above, and represent that you have the unrestricted right to grant that license;</li>\n  <li>Sign and return all Prize acceptance documents as may be required by Competition Sponsor or Kaggle, including without limitation: (a) eligibility certifications; (b) licenses, releases and other agreements required under the Rules; and (c) U.S. tax forms (such as IRS Form W-9 if U.S. resident, IRS Form W-8BEN if foreign resident, or future equivalents).</li>\n  <li>Individual Participants and Teams who create a Submission using an AMLT may win a Prize. However, for clarity, the potential winner’s Submission must still meet the requirements of these Rules, including but not limited to Section 2.5 (Winner’s License and Non-confidentiality), Section 2.8 (Winner’s Obligations), and Section 3.14 (Warranty, Indemnity, and Release).&lt;</li>\n  </ol>\n</blockquote>",
          "rawMarkdown": "Hi Sumo, thank you so much for participating. We are thrilled to see the Kaggle community in action tackling problems that are important to us. I have copy-pasted this from the Rules section:\n>8. WINNER’S OBLIGATIONS.\na. As a condition to being awarded a Prize, a Prize winner must fulfill the following obligations:\n1. Deliver to the Competition Sponsor the final model's software code as used to generate the winning Submission and associated documentation. The delivered software code should follow these documentation guidelines, must be capable of generating the winning Submission, and contain a description of resources required to build and/or run the executable code successfully. For avoidance of doubt, delivered software code should include training code, inference code, and a description of the required computational environment. To the extent that the final model’s software code includes generally commercially available software that is not owned by you, but that can be procured by the Competition Sponsor without undue expense, then instead of delivering the code for that software to the Competition Sponsor, you must identify that software, method for procuring it, and any parameters or other information necessary to replicate the winning Submission;\n2. Grant to the Competition Sponsor the license to the winning Submission stated in the Competition Specific Rules above, and represent that you have the unrestricted right to grant that license;\n3. Sign and return all Prize acceptance documents as may be required by Competition Sponsor or Kaggle, including without limitation: (a) eligibility certifications; (b) licenses, releases and other agreements required under the Rules; and (c) U.S. tax forms (such as IRS Form W-9 if U.S. resident, IRS Form W-8BEN if foreign resident, or future equivalents).\n4. Individual Participants and Teams who create a Submission using an AMLT may win a Prize. However, for clarity, the potential winner’s Submission must still meet the requirements of these Rules, including but not limited to Section 2.5 (Winner’s License and Non-confidentiality), Section 2.8 (Winner’s Obligations), and Section 3.14 (Warranty, Indemnity, and Release).<",
          "votes": 2
        }
      ]
    },
    {
      "id": 3117027,
      "postDate": "2025-02-06T14:57:47.563Z",
      "content": "<p>Congratulations!!! Very insightful ensemble solutions. <br>\nWould you mind sharing the single model performance in the final ensemble models?</p>",
      "rawMarkdown": "Congratulations!!! Very insightful ensemble solutions. \nWould you mind sharing the single model performance in the final ensemble models?",
      "replies": [
        {
          "id": 3117031,
          "postDate": "2025-02-06T15:00:07.970Z",
          "content": "<p>one of the single model of final ensemble was of following config:</p>\n<p><code>\"channels\": (32, 64, 128, 256),\n\"strides\": (2, 2, 1),\n\"num_res_units\": 2,</code></p>\n<p>we created model soup from all the folds and it's public LB was 0.762</p>",
          "rawMarkdown": "one of the single model of final ensemble was of following config:\n\n`\"channels\": (32, 64, 128, 256),\n\"strides\": (2, 2, 1),\n\"num_res_units\": 2,`\n\nwe created model soup from all the folds and it's public LB was 0.762\n",
          "votes": 1,
          "replies": [
            {
              "id": 3117037,
              "postDate": "2025-02-06T15:10:26.230Z",
              "content": "<p>thank you very much for sharing!</p>",
              "rawMarkdown": "thank you very much for sharing!"
            }
          ]
        }
      ]
    },
    {
      "id": 3118080,
      "postDate": "2025-02-07T15:07:16.787Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3118576,
      "author_name": "Timmy Juicehouse",
      "author_url": "",
      "post_date": "2025-02-08T08:14:54.210000",
      "content": "<p>Congratulations to your team, and very glad to see you <a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a> winning a gold medal.</p>\n<p>I noticed a post on Xiaohongshu regarding accusations against one of your team members. Many Chinese kagglers in the discussion group echoed similar experiences. I don't know the private communications within your team regarding this matter. Here, I am posting the screenshot to remind everyone to be cautious in choosing teammates when teaming up. Of course, if the relevant accusations are proven to be untrue, I am willing to publicly apologize to you <a href=\"https://www.kaggle.com/yinhewang\" target=\"_blank\">@yinhewang</a> .</p>\n<p>This may affect the relationships among your team members, and for that I sincerely apologize.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18538441%2F29191241e615e6c202092157f3203a9f%2FWechatIMG20590.jpg?generation=1739001808068077&amp;alt=media\" alt=\"\"></p>",
      "votes": 21,
      "replies": [
        {
          "id": 3118589,
          "author_name": "Joseph Zhou",
          "author_url": "",
          "post_date": "2025-02-08T08:38:07.860000",
          "content": "<p>The translation:<br>\n\"Bro, you really managed to fool people. This person contacted me to team up on Kaggle two months ago. At that time, I was ranked relatively high on the leaderboard for a silver medal in one competition. I didn't think much about it. He said he had several L20s to run experiments. But after teaming up, he just disappeared, casting a wide net in various competitions. Then other people from different competitions approached me, asking if this person was just a freeloader doing no work. Now that CZII has ended, you really managed to get yourself a gold medal. I don't know if you disappeared in this team as well. I just know that in their team's solution, four people were mentioned, but not you. Foreigners really don’t understand just how cunning some Chinese people can be, huh?\"</p>",
          "votes": 12,
          "replies": [
            {
              "id": 3118693,
              "author_name": "FakeOrange",
              "author_url": "",
              "post_date": "2025-02-08T11:10:37.887000",
              "content": "<p>I am sorry to hear about this. As a Chinese, I feel so ashamed of such inappropriate behavior. However, it is important to acknowledge that many Chinese individuals work diligently and collaborate effectively in teamworks. The key challenge is managing cunning and fool people who may ghost team members or steal your contributions. Therefore, it is essential to remain cautious and careful, and in certain situations, verification may be necessary.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 3118663,
          "author_name": "Sergio Alvarez",
          "author_url": "",
          "post_date": "2025-02-08T10:26:51.150000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sweetyheehee\" target=\"_blank\">@sweetyheehee</a>, we are aware of <a href=\"https://www.kaggle.com/yinhewang\" target=\"_blank\">@yinhewang</a>'s situation. He had no involvement in winning the gold medal. Essentially, he ghosted us from the beginning and only returned after we won.</p>\n<p>We tried to investigate, but things started getting very complicated and weird, many lies were appearing. We also reached out to other teams.</p>\n<p>We sent an email to Kaggle and were waiting for a response before going public and starting a discussion on Kaggle with all the proof we gathered.</p>\n<p>I entered right before merge so don't have all info to share. But my teammates will come here later.</p>",
          "votes": 6,
          "replies": [
            {
              "id": 3118681,
              "author_name": "Timmy Juicehouse",
              "author_url": "",
              "post_date": "2025-02-08T10:54:41.173000",
              "content": "<p>Thanks for your feedback <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> , </p>\n<p>The reason I made this post is that many Chinese kagglers are already discussing his alleged misconduct. We don't want more people to be deceived, and we also hope that he can clarify these matters. However, his account has already been deleted, which suggests that the previous accusations are true. </p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 3118729,
          "author_name": "StarxSky",
          "author_url": "",
          "post_date": "2025-02-08T12:24:37.533000",
          "content": "<p>As a Chinese, I feel ashamed of this behavior, but more importantly, this behavior has seriously damaged the friendly and cooperative atmosphere of the kaggle community. Therefore, I would like to warn everyone to be careful when facing the application of unknown person to join the team, especially for such low-activity accounts. It is hoped that through the joint efforts of community members, the kaggle community can be maintained in a positive and healthy direction. Finally, I would like to apologize again for this happening in the kaggle community.</p>\n<p>There is a Chinese poem that I think is a powerful criticism of such people：<br>\n“小人者其未得也，则忧不得； 既已得之，又恐失之。是以有终身之忧，无一日之乐也。”</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 3118734,
          "author_name": "Sumo",
          "author_url": "",
          "post_date": "2025-02-08T12:33:27.037000",
          "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/sweetyheehee\" target=\"_blank\">@sweetyheehee</a>, as <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> mentioned, we have started our investigations ourselves and have found a lot of evidence that points towards <a href=\"https://www.kaggle.com/yinhewang\" target=\"_blank\">@yinhewang</a> deliberately leeching off of teams.</p>\n<p>From our side we have submitted an inquiry to kaggle, which they have forwarded to kaggle's compliance team. I'm not sure if this is correlated with <a href=\"https://www.kaggle.com/yinhewang\" target=\"_blank\">@yinhewang</a>'s account being deleted or not. But as of the time I post this comment, the account is gone <a href=\"https://www.kaggle.com/yinhewang\" target=\"_blank\">https://www.kaggle.com/yinhewang</a></p>\n<p>In addition to this, we have evidence that <a href=\"https://www.kaggle.com/yinhewang\" target=\"_blank\">@yinhewang</a> is also planning to even take parts of our prize money by submitting the winning team form. We are not sure what additional steps we can take to guarantee this doesn't happen. But we do have couple of written evidence from him that he is willing to not take any part in prize money.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 3118757,
              "author_name": "Sumo",
              "author_url": "",
              "post_date": "2025-02-08T13:11:34.840000",
              "content": "<p>I've made a post in the general thread to try bring more attention to this issue <a href=\"https://www.kaggle.com/discussions/general/561879\" target=\"_blank\">https://www.kaggle.com/discussions/general/561879</a></p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3118769,
              "author_name": "Timmy Juicehouse",
              "author_url": "",
              "post_date": "2025-02-08T13:29:42.593000",
              "content": "<p>Thank you, Sumo. I will forward your post to my friends as well as his former teammates.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3117418,
      "author_name": "ITK8191",
      "author_url": "",
      "post_date": "2025-02-07T01:27:27.053000",
      "content": "<p>I am super excited to team up with <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> , <a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a> and <a href=\"https://www.kaggle.com/sirapoabchaikunsaeng\" target=\"_blank\">@sirapoabchaikunsaeng</a>  to complete this competition! They are all wonderful mates who are very proactive in discussions and didn't hesitate to try out my crazy ideas.<br>\n<a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> picked out and summarized the ideas that worked for me, but in this post I'll cover the ones that didn't.</p>\n<h3>What I tried that didn't work or didn't used in submission</h3>\n<ol>\n<li>YOLO approach. I believe many competitors created their solutions based on my baseline, but our team ultimately chose a solution based primarily on 3D-UNet. Therefore, I believe that the solution for YOLO is not worthless, so I will describe it here. In this method, the following preprocessing was applied:<br>\n・For one z, a 3-channel image (z-1, z, z+1) is used instead of a 1-channel image (z, z, z).<br>\n・Scale the image to the range [-1, 1] and then raise the pixel values ​​to the p-th power (approximately 1&lt;p&lt;=3) <br>\n・For min-max scaling, used a 3-channel image with clip values ​​of (0.0001%, 2%), (2%, 98%), and (98%, 99.999%).<br>\nThe second and third ideas take into account the fact that pixels around the target have small values, as considered by <a href=\"https://www.kaggle.com/sirapoabchaikunsaeng\" target=\"_blank\">@sirapoabchaikunsaeng</a> . They improved LB 0.015~0.03.<br>\nThere were few changes to the training. As a tip, less strong data augmentation tends to produce better scores.<br>\nThere are various ensemble and postprocessing methods for object detection, but I used Weighted Box Fusion. Due to its characteristics, it is necessary to lower the threshold score a little. Applying the techniques above to the publicly available high-scoring notebooks would have been enough to win a bronze medal, but we weren't satisfied with that.</li>\n<li>Data generation using ControlNet w/ Stable Diffusion<br>\nAs mentioned in my other discussion posts, I thought the bad scores were due to a lack of training data, so I experimented with generating slice images from the GT of the segmentation map (i.e. the inverse of segmentation). This didn't work at all, and produced terrible results with mode collapse.</li>\n<li>Use a larger volume<br>\nI think we all know that a fundamental principle of image processing is that the larger the training image, the better the performance tends to be. So I scaled the (184, 630, 630) volume by a factor of 1.5 and trained the (276, 945, 945) one. This also didn't work well (around LB 0.7), Also, because the training took so long, there was no time to do any in-depth research.</li>\n</ol>\n<h3>Things I wanted to try but couldn't</h3>\n<ol>\n<li>Auxiliary task and aux loss<br>\nThe test dataset only has the denoising method, but the training dataset has other process methods. For example, we can take the denoised image as input and add an auxiliary task of reconstructing the ctfdeconvolved to the segmentation task. Multi-task learning may lead to faster training or better accuracy. I wanted to try this, but I didn't have time.</li>\n<li>Contrastive Learning<br>\nI considered contrastive learning with two input images. We suggested an auxiliary task that would give a positive label when the second image was an augmentation or identity of the first image, and a negative label when the first and second images were different, but we didn't have time to do this either.</li>\n</ol>\n<h3>Summary</h3>\n<p>Many of my suggestions went to waste, but I'm still happy to have contributed even a small amount. I ended up not being of much use to the team, so I would like to take this opportunity to thank the team once again.</p>\n<h3>Finally some words</h3>\n<p>In this post, I will first state what I want to say. <strong>Most of our team's experiments (especially mine) have been failures</strong>, Maybe, as have most of the other gold medal winners. <strong>Don't be afraid of failing an experiment or not winning a medal in a competition.</strong> Don't give up just because you lost one competition. Kaggle <strong>NEVER</strong> penalize you for losing. Your competitors are also your best companions to expand your knowledge. Don't stop learning from them, and one day you will become a great Kaggler.<br>\nThere is a Japanese-English word, \"no side\" It is a word used to praise both sides for their good fight after the game is over, regardless of whether they are allies or opponents. This is getting long, but my final word is…<br>\n<strong>NO SIDE!</strong></p>",
      "votes": 10,
      "replies": [
        {
          "id": 3117798,
          "author_name": "min fuka",
          "author_url": "",
          "post_date": "2025-02-07T08:58:29.360000",
          "content": "<p><a href=\"https://www.kaggle.com/ITK8191\" target=\"_blank\">@ITK8191</a> <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a><br>\nThank you for your notebooks </p>\n<p>I had never used YOLO before, so I continued submitting in YOLO for practice. However, I did not do a team merge because I had doubts about my YOLO license and could have been expelled.</p>\n<p>One of my innovations was the calculation of Z coordinates.<br>\nWhen calculating the center coordinates in KDTree, only the Z coordinate is calculated with confidence as a weight. This improved the LB by about 0.005</p>\n<p>particle_confidences = pdf['confidence'].tolist()<br>\nparticle_xs = pdf['x'].tolist()<br>\nparticle_ys = pdf['y'].tolist()<br>\nparticle_zs = pdf['z'].tolist()<br>\nparticle_zs_conf = (pdf['z']*pdf['confidence']).tolist()<br>\n～<br>\ncenters_x = np.bincount(inverse_indices, weights=particle_xs) / counts<br>\ncenters_y = np.bincount(inverse_indices, weights=particle_ys) / counts<br>\n# centers_z = np.bincount(inverse_indices, weights=particle_yz) / counts<br>\ncenters_z = np.bincount(inverse_indices, weights=particle_zs_conf) / conf_sums</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3117540,
      "author_name": "David List",
      "author_url": "",
      "post_date": "2025-02-07T03:49:07.167000",
      "content": "<p>Congrats all of you!  Very nice job!  Thanks for the great explanation!</p>",
      "votes": 4,
      "replies": [
        {
          "id": 3117574,
          "author_name": "IAmParadox",
          "author_url": "",
          "post_date": "2025-02-07T04:15:37.163000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a> congrats on your gold as well!!, You posts really helped us during the competition!</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3117809,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2025-02-07T09:06:46.490000",
              "content": "<p>Yeah, I may have to tone that down a little next time.  Too many people ahead of me thanking me.  😀  (Kidding, of course.)</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3118039,
              "author_name": "Andrei Zamfir",
              "author_url": "",
              "post_date": "2025-02-07T14:05:54.540000",
              "content": "<p>People behind doing the same 😆 That's the Kaggle spirit, and it's what makes these competitions so fun!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3119442,
      "author_name": "wpl",
      "author_url": "",
      "post_date": "2025-02-09T09:22:03.693000",
      "content": "<p>Congratulations to your team for winning this game, and I am also sorry for yinhe wang's situation.<br>\nIn addition, I would like to know how you use the EMA strategy. When I tried basic EMA, it had almost no effect on improving my score. So I adjusted to update once after each round and tried using the dynamic deck method, which can achieve slight improvements in some cases. I saw that you said EMA has a good effect, so I would like to consult your usage strategy</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3118117,
      "author_name": "Aurora_blue",
      "author_url": "",
      "post_date": "2025-02-07T16:22:55.983000",
      "content": "<p>Congratulations!</p>\n<p><a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a> <br>\nThis is interesting for me.</p>\n<blockquote>\n  <p>One of the greatest findings was making the inference in patches as large as possible; we used a window size of (160, 384, 384), boosting our score by around 0.01 compared to a size of (128, 128, 128).</p>\n</blockquote>\n<p>Do you train model with same patch size or not ?<br>\nIf not, I cannot understand why this improve score.<br>\nI would appreciate it if you could share any theoretical background.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3118821,
          "author_name": "IAmParadox",
          "author_url": "",
          "post_date": "2025-02-08T14:16:10.763000",
          "content": "<p>The model was only trained on 128x128x128 patch size but it was inferred on 160x384x384 patch size. I honestly don't know a lot of theoretical background on this topic, but the idea here's that the contextual information doesn't change while we are doing the inference, in a sense that suppose, a 3x3x3 conv filter expects high value in the center pixel of the input image to activate then this doesn't change even if I increase the patch size at inference and the second reason is doing inference on larger patchsize reduces the border artifacts. </p>\n<p>The large patchsize trick will not work if we downsample the image into smaller dimension, and then infer them on higher dimension because in this case the contextual information will change, i.e it will see finer details that it has never seen before.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 3119101,
          "author_name": "David List",
          "author_url": "",
          "post_date": "2025-02-08T21:49:04.200000",
          "content": "<p><a href=\"https://www.kaggle.com/clearwaterkzk\" target=\"_blank\">@clearwaterkzk</a> One pretty big issue I found with the <a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a> notebook was the borders of each patch were different from the rest of the model output so you could see a grid pattern in the reconstructed model output.  You can see some evidence of this here: <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549615\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549615</a> which shows single patches.  I wonder if this is the same issue.  Making the patches bigger would reduce the problem.  I likewise saw a big performance increase by deleting the outer 3 rows of pixels from each patch before reconstruction and averaging a small overlap.  slinding_window_inference() makes this pretty straightforward.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3119145,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-02-08T23:27:50.210000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3119146,
          "author_name": "Aurora_blue",
          "author_url": "",
          "post_date": "2025-02-08T23:29:04.317000",
          "content": "<p>Thanks for replying.<br>\nI also got the answer from itsuki9180 directly in Twitter(X) regarding this.</p>\n<p>I found the border issue and deal with it by padding the voxel and sliding window with a lot of overlap, in which only the center part are predicted, then averaging probability like following figure.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F88c71d74e91423a3c61ec3136ccd3c4d%2Fsliding_window.png?generation=1739056705345010&amp;alt=media\" alt=\"\"><br>\nyellow: original exp voxel size, black: patch size for inference, red: probability to use for reconstruction</p>\n<p>It seems that this method is used in some segmentation past competition like Vesuvius Challenge, but it requires more inference time. Your method seems smarter and faster.</p>\n<p>I used to have the fixed idea that CNNs must be inferred with the same size as they were trained on. However, after seeing this solution, I realized for the first time that convolution-only calculations can accept any input size. I learned a lot from this. Thank you very much.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3119158,
              "author_name": "Sergio Alvarez",
              "author_url": "",
              "post_date": "2025-02-09T00:18:31.917000",
              "content": "<p><a href=\"https://www.kaggle.com/clearwaterkzk\" target=\"_blank\">@clearwaterkzk</a> Just to shown the difference bigger patches made, I submitted our best submission with patches 128.<br>\nBest submission: public lb: 0.78022, private lb: 0.77612<br>\nPatches 128 submission: public lb: 0.77464, private lb: 0.76604</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3119859,
              "author_name": "Andrei Zamfir",
              "author_url": "",
              "post_date": "2025-02-09T19:06:42.357000",
              "content": "<p>That's interesting, I was also under the impression that I must infer on same subvolume size I trained on, until I accidentally discovered that's not the case. However, I thought that somehow, maintaining isotropic subvolume sizes would be advantageous (I worked on 184x184x184 for most of the competition, started with 128x320x320), and I remember checking an experiment where I had my best submission perform better on the 184x184x184. I have just sent a few experiments on 128x320x320 to see if what <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> found stands true for my solution as well. <br>\nI understand that with bigger patches, there's less border artifacts and less merges to be made, but assuming a sizable overlap and gaussian weights for reconstruction, I was inclined to believe that the edges would have a poor enough weight to be negligible.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3118096,
      "author_name": "MK245",
      "author_url": "",
      "post_date": "2025-02-07T15:26:24.123000",
      "content": "<p>brilliant, interesting that augmentations didn't work at all. did you try mirroring?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3117093,
      "author_name": "Sumo",
      "author_url": "",
      "post_date": "2025-02-06T16:19:45.010000",
      "content": "<p><a href=\"https://www.kaggle.com/rezaparaan\" target=\"_blank\">@rezaparaan</a>,<br>\nthank you again for the competition! we're very grateful for all the learnings we've had from this, and also all the friendships we've made (this is the first time we all team up), as well as designing the public and private LB such that there's no massive shakeup!</p>\n<p>this is the first time our team won gold, may I ask you to outline what will happen next? As in what should we provide you with (code / model weights / reports). So that I can make sure that I don't miss anything in the process.</p>\n<p>Thank you!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3117094,
          "author_name": "Sumo",
          "author_url": "",
          "post_date": "2025-02-06T16:22:20.337000",
          "content": "<p>cc <a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a>, provided you're also this competition host, apologies for any mistakes in tagging anyone</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3119183,
          "author_name": "Reza Paraan",
          "author_url": "",
          "post_date": "2025-02-09T01:28:33.623000",
          "content": "<p>Hi Sumo, thank you so much for participating. We are thrilled to see the Kaggle community in action tackling problems that are important to us. I have copy-pasted this from the Rules section:</p>\n<blockquote>\n  <ol>\n  <li>WINNER’S OBLIGATIONS.<br>\n  a. As a condition to being awarded a Prize, a Prize winner must fulfill the following obligations:</li>\n  <li>Deliver to the Competition Sponsor the final model's software code as used to generate the winning Submission and associated documentation. The delivered software code should follow these documentation guidelines, must be capable of generating the winning Submission, and contain a description of resources required to build and/or run the executable code successfully. For avoidance of doubt, delivered software code should include training code, inference code, and a description of the required computational environment. To the extent that the final model’s software code includes generally commercially available software that is not owned by you, but that can be procured by the Competition Sponsor without undue expense, then instead of delivering the code for that software to the Competition Sponsor, you must identify that software, method for procuring it, and any parameters or other information necessary to replicate the winning Submission;</li>\n  <li>Grant to the Competition Sponsor the license to the winning Submission stated in the Competition Specific Rules above, and represent that you have the unrestricted right to grant that license;</li>\n  <li>Sign and return all Prize acceptance documents as may be required by Competition Sponsor or Kaggle, including without limitation: (a) eligibility certifications; (b) licenses, releases and other agreements required under the Rules; and (c) U.S. tax forms (such as IRS Form W-9 if U.S. resident, IRS Form W-8BEN if foreign resident, or future equivalents).</li>\n  <li>Individual Participants and Teams who create a Submission using an AMLT may win a Prize. However, for clarity, the potential winner’s Submission must still meet the requirements of these Rules, including but not limited to Section 2.5 (Winner’s License and Non-confidentiality), Section 2.8 (Winner’s Obligations), and Section 3.14 (Warranty, Indemnity, and Release).&lt;</li>\n  </ol>\n</blockquote>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3117027,
      "author_name": "Ma Edward",
      "author_url": "",
      "post_date": "2025-02-06T14:57:47.563000",
      "content": "<p>Congratulations!!! Very insightful ensemble solutions. <br>\nWould you mind sharing the single model performance in the final ensemble models?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3117031,
          "author_name": "IAmParadox",
          "author_url": "",
          "post_date": "2025-02-06T15:00:07.970000",
          "content": "<p>one of the single model of final ensemble was of following config:</p>\n<p><code>\"channels\": (32, 64, 128, 256),\n\"strides\": (2, 2, 1),\n\"num_res_units\": 2,</code></p>\n<p>we created model soup from all the folds and it's public LB was 0.762</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3117037,
              "author_name": "Ma Edward",
              "author_url": "",
              "post_date": "2025-02-06T15:10:26.230000",
              "content": "<p>thank you very much for sharing!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3118080,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-02-07T15:07:16.787000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3117017": "\nFirst, we thank the competition host and Kaggle staff for organizing this competition. Below, we introduce the solution of the team I Cryo Everyteim -- @sirapoabchaikunsaeng, @iamparadox, @itsuki9180, @sersasj -- \n\n## Context\n\n- Business context: [competition overview](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/overview)\n- Data context: [competition data](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/data)\n\n\n## Overview of the approach\n\nOur final submissions consisted of four 3D U-Net model soups, trained with different model sizes, parameters, and training data to ensure strong model complementarity. All 3D U-Nets were trained using patch sizes of (128, 128, 128) but were inferred with patch sizes of (160, 384, 384), with a 25% overlap using gaussian reconstruction to handle border artifacts. Additionally, we employed geometric test-time augmentation (TTA), including flipping and transpose.\n\n### Models \n\nOur models were originally based on the host's example notebook and @fnands notebook . We utilized 3D U-Net architectures from the MONAI library, trained with patch sizes of (128, 128, 128).\n\nThe models were 3 levels deep with strides of (2, 2, 1). We started with a simple model:\n\n```\n\"channels\": (32, 64, 128, 128)\n\"strides\": (2, 2, 1)\n\"num_res_units\": 1\n```\n\nThis single model with our inference strategy reached 0.744 in public leaderboard score.\n\nLater, we trained more complex U-Nets with configs:\n\n```\n\"channels\": (32, 64, 128, 256),\n\"strides\": (2, 2, 1),\n\"num_res_units\": 2,\n```\nThe model with above config alone achieved a score of 0.759 LB.\n\n```\n\"channels\": (32, 96, 256, 384),\n\"strides\": (2, 2, 1),\n\"num_res_units\": 2,\n```\n\nWe also applied a dropout of 0.2 or 0.3 in the model.\n\n### Training\n\nThe models were pre-trained on six synthetic tomograms denoised with Gaussian denoising—specifically, the 'TS_0', 'TS_1', 'TS_10', 'TS_11', 'TS_12', and 'TS_13' tomograms.\n\n*@sersasj note: Gaussian denoising was applied because it visually improved the WBP tomograms particle visualization and was easy to implement. I hypothesized that better results could be achieved with a more advanced denoiser, but attempting to code a model for denoising used too much of my Kaggle quota, so I gave up.*\n\nPretraining not only reduced the time needed for the models to learn the particles but also increased the LB score by roughly 0.01.\n\nBoth pretraining on synthetic data and fine-tuning uses the following transformations:\n```\nCompose([\n    RandCropByLabelClassesd(\n        keys=[\"image\", \"label\"],\n        label_key=\"label\",\n        spatial_size=[128, 128, 128],\n        num_classes=7,  # background, and all 6 classes\n        num_samples=16,\n    ),\n    RandFlipd(keys=[\"image\", \"label\"], prob=0.5, spatial_axis=0),\n    RandFlipd(keys=[\"image\", \"label\"], prob=0.5, spatial_axis=1),\n    RandFlipd(keys=[\"image\", \"label\"], prob=0.5, spatial_axis=2),\n])\n```\n\nWe used MONAI's DiceCELoss and optimized the models with AdamW, employing a learning rate reduction on plateau and an initial learning rate of 1e-3.\n\nWe also experimented with Exponential Moving Average (EMA), which showed good results. In the last 3 days of the competition, we trained models using other tomo types: \"denoised\", \"ctfdeconvolved\", \"isonetcorrected\". For almost the entire competition, we hadn't found any increase in LB strategies other than geometric data augmentation in training (flip and rotate); nevertheless, training with other tomo types yielded good results.\n\n### Validation strategy\n\nWe primarily relied on out-of-fold predictions from our k-fold models. Specifically, we implemented a 7-fold cross-validation approach where we trained on all tomographies except one, which was used as a validation set. This process was repeated seven times, with each fold serving as the validation set once. This helped us obtain a score that correlated well with the public leaderboard score, typically 0.01 to 0.02 points higher.\n\nOne limitation we encountered was our inability to apply this validation strategy when testing ensembles of our model soups. In these cases, we had to rely on leaked data scores. While this approach didn't correlate as strongly with leaderboard scores, most of the time it allowed us to rank the performance of different ensembles, and compare TTA strategies. \n\n## Details of the submission\n\n### Inference\n\nDuring inference, we use MONAI's sliding_window_inference with 25% overlap and Gaussian reconstruction to handle border predictions more accurately; this boosted the public LB by around 0.002. **One of the greatest findings was making the inference in patches as large as possible; we used a window size of (160, 384, 384), boosting our score by around 0.01 compared to a size of (128, 128, 128).** (Thanks @iamparadox for finding this)\n\nWe also experimented with combining different sizes of predictions, for example (160, 384, 384) + (136, 248, 248), but inference took much longer while yielding inconclusive gains. Therefore, we abandoned this approach in favor of creating more ensembles and employing more TTA.\n\nFor post-processing, we improved particle detection by combining model predictions at the logits level rather than averaging probabilities directly.\n\nWe then refined particle identification using watershed segmentation, which helped separate touching or overlapping particles. The watershed process created binary masks using particle-specific certainty thresholds, computed distance transforms to measure particle separation, identified distinct particles using local maxima as markers, and finally applied watershed segmentation to determine particle boundaries. We switched from using skimage to cucim and cupy to decrease the inference time. \n\nWe also used specific blob threshold sizes for each particle to decrease the number of false positives in our predictions.\n\nMoreover, we used flip along the X and Y axes and transpose TTA.\n   \n### Ensembles\n\nEnsembling was always one of the main ideas in our team. When the team was fully formed 2 weeks before the deadline, we were all around a 0.73 LB score, around 50th place on the leaderboard. At the time, I (@sersasj) had a YOLO scoring 0.682 and a U-Net scoring 0.692 that, when ensembled, reached 0.732 in LB. The rest of the team had only a 3D U-Net that reached 0.739. So it was obvious that we would ensemble everything and achieve amazing results; sadly, it didn't work.\n\nNevertheless, we learned valuable lessons from this experience. My 3D U-Net was ~30 times smaller than the 0.739 scoring U-Net, but it proved quite effective when used with the team's amazing inference techniques: *larger patches, blob thresholds, ensemble at logits level, and watershed processing*. After implementing pretraining and additional training, it achieved a 0.744 leaderboard score. \n\nThanks to our group work (especially @sirapoabchaikunsaeng and @iamparadox), we successfully trained several U-Net variants derived from the 0.744 model by modifying parameters, implementing EMA (Exponential Moving Average), incorporating more data, and applying the techniques mentioned above.\n\nBy doing so, we had various U-Nets that were complementary to each other, but we couldn't use them all. To manage \"using\" all we had, @itsuki9180 presented the idea and code for a technique called model soup, which essentially averages the weights of multiple models from the same pre-trained weights.\nWe applied this to the models from our k-fold training.\n\nOur final ensemble consisted of:\n1. A soup of 7 folds trained with parameters:\n   - Channels: (32, 64, 128, 128)\n   - Number of residual units: 1\n   - Pretrained on synthetic data\n   - Trained on denoised images without EMA\n\n2. A soup of 7 folds trained with parameters:\n   - Channels: (32, 64, 128, 256)\n   - Number of residual units: 2\n   - Pretrained on synthetic data\n   - Trained on all denoised, IsoNet-corrected, and CTF-deconvolved images without EMA\n\n3. A soup of 3 folds (TS_69_2, TS_86_3, TS_99_9) trained with parameters:\n   - Channels: (32, 96, 256, 384)\n   - Number of residual units: 2\n   - Pretrained on synthetic data\n   - Trained on denoised images with EMA\n\n4. A soup of 7 folds trained with parameters:\n   - Channels: (32, 96, 256, 384)\n   - Number of residual units: 2\n   - Pretrained on synthetic data\n   - Trained on all denoised, IsoNet-corrected, and CTF-deconvolved images without EMA\n\n### What didn't work:\n- Increased Augmentations\n- CutMix, MixUp\n- 2D U-Net approaches (reached 0.636 lb only)\n- Ensembling with YOLO (for final solution, works well to reach ~0.73 lb)\n- Bigger and deeper U-Nets\n- Cross entropy only loss functions\n\n\n## Sources section\n\n- [Model Soup paper](https://arxiv.org/abs/2203.05482)\n- [Host's example notebook](https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/3d_unet_monai/train.ipynb)\n- [@fnands notebook](https://www.kaggle.com/code/fnands/baseline-unet-train-submit)\n\n\n## Code\n\n[8th place solution of kaggle czii competition code github\n](https://github.com/IAmPara0x/czii-8th-solution)\n[Submission Notebook](https://www.kaggle.com/code/iamparadox/czii-final-sub-reproduce)\n[Submission Notebook Clean Version](https://www.kaggle.com/code/sirapoabchaikunsaeng/czii-final-sub-reproduce)",
    "3118576": "Congratulations to your team, and very glad to see you @iamparadox winning a gold medal.\n\nI noticed a post on Xiaohongshu regarding accusations against one of your team members. Many Chinese kagglers in the discussion group echoed similar experiences. I don't know the private communications within your team regarding this matter. Here, I am posting the screenshot to remind everyone to be cautious in choosing teammates when teaming up. Of course, if the relevant accusations are proven to be untrue, I am willing to publicly apologize to you @yinhewang .\n\nThis may affect the relationships among your team members, and for that I sincerely apologize.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18538441%2F29191241e615e6c202092157f3203a9f%2FWechatIMG20590.jpg?generation=1739001808068077&alt=media)",
    "3117418": "I am super excited to team up with @sersasj , @iamparadox and @sirapoabchaikunsaeng  to complete this competition! They are all wonderful mates who are very proactive in discussions and didn't hesitate to try out my crazy ideas.\n@sersasj picked out and summarized the ideas that worked for me, but in this post I'll cover the ones that didn't.\n### What I tried that didn't work or didn't used in submission\n1. YOLO approach. I believe many competitors created their solutions based on my baseline, but our team ultimately chose a solution based primarily on 3D-UNet. Therefore, I believe that the solution for YOLO is not worthless, so I will describe it here. In this method, the following preprocessing was applied:\n・For one z, a 3-channel image (z-1, z, z+1) is used instead of a 1-channel image (z, z, z).\n・Scale the image to the range [-1, 1] and then raise the pixel values ​​to the p-th power (approximately 1<p<=3) \n・For min-max scaling, used a 3-channel image with clip values ​​of (0.0001%, 2%), (2%, 98%), and (98%, 99.999%).\nThe second and third ideas take into account the fact that pixels around the target have small values, as considered by @sirapoabchaikunsaeng . They improved LB 0.015~0.03.\nThere were few changes to the training. As a tip, less strong data augmentation tends to produce better scores.\nThere are various ensemble and postprocessing methods for object detection, but I used Weighted Box Fusion. Due to its characteristics, it is necessary to lower the threshold score a little. Applying the techniques above to the publicly available high-scoring notebooks would have been enough to win a bronze medal, but we weren't satisfied with that.\n2. Data generation using ControlNet w/ Stable Diffusion\nAs mentioned in my other discussion posts, I thought the bad scores were due to a lack of training data, so I experimented with generating slice images from the GT of the segmentation map (i.e. the inverse of segmentation). This didn't work at all, and produced terrible results with mode collapse.\n3. Use a larger volume\nI think we all know that a fundamental principle of image processing is that the larger the training image, the better the performance tends to be. So I scaled the (184, 630, 630) volume by a factor of 1.5 and trained the (276, 945, 945) one. This also didn't work well (around LB 0.7), Also, because the training took so long, there was no time to do any in-depth research.\n\n### Things I wanted to try but couldn't\n1. Auxiliary task and aux loss\nThe test dataset only has the denoising method, but the training dataset has other process methods. For example, we can take the denoised image as input and add an auxiliary task of reconstructing the ctfdeconvolved to the segmentation task. Multi-task learning may lead to faster training or better accuracy. I wanted to try this, but I didn't have time.\n2. Contrastive Learning\nI considered contrastive learning with two input images. We suggested an auxiliary task that would give a positive label when the second image was an augmentation or identity of the first image, and a negative label when the first and second images were different, but we didn't have time to do this either.\n\n### Summary\nMany of my suggestions went to waste, but I'm still happy to have contributed even a small amount. I ended up not being of much use to the team, so I would like to take this opportunity to thank the team once again.\n\n### Finally some words\nIn this post, I will first state what I want to say. **Most of our team's experiments (especially mine) have been failures**, Maybe, as have most of the other gold medal winners. **Don't be afraid of failing an experiment or not winning a medal in a competition.** Don't give up just because you lost one competition. Kaggle **NEVER** penalize you for losing. Your competitors are also your best companions to expand your knowledge. Don't stop learning from them, and one day you will become a great Kaggler.\nThere is a Japanese-English word, \"no side\" It is a word used to praise both sides for their good fight after the game is over, regardless of whether they are allies or opponents. This is getting long, but my final word is...\n**NO SIDE!**",
    "3117540": "Congrats all of you!  Very nice job!  Thanks for the great explanation!",
    "3119442": "Congratulations to your team for winning this game, and I am also sorry for yinhe wang's situation.\nIn addition, I would like to know how you use the EMA strategy. When I tried basic EMA, it had almost no effect on improving my score. So I adjusted to update once after each round and tried using the dynamic deck method, which can achieve slight improvements in some cases. I saw that you said EMA has a good effect, so I would like to consult your usage strategy",
    "3118117": "Congratulations!\n\n@iamparadox \nThis is interesting for me.\n>One of the greatest findings was making the inference in patches as large as possible; we used a window size of (160, 384, 384), boosting our score by around 0.01 compared to a size of (128, 128, 128).\n\nDo you train model with same patch size or not ?\nIf not, I cannot understand why this improve score.\nI would appreciate it if you could share any theoretical background.\n",
    "3118096": "brilliant, interesting that augmentations didn't work at all. did you try mirroring?",
    "3117093": "@rezaparaan,\nthank you again for the competition! we're very grateful for all the learnings we've had from this, and also all the friendships we've made (this is the first time we all team up), as well as designing the public and private LB such that there's no massive shakeup!\n\nthis is the first time our team won gold, may I ask you to outline what will happen next? As in what should we provide you with (code / model weights / reports). So that I can make sure that I don't miss anything in the process.\n\nThank you!",
    "3117027": "Congratulations!!! Very insightful ensemble solutions. \nWould you mind sharing the single model performance in the final ensemble models?",
    "3118080": ""
  }
}