{
  "id": 561447,
  "title": "7th place solution [3D-UNet with gaussian heatmaps]",
  "url": "/competitions/czii-cryo-et-object-identification/writeups/kobakos-7th-place-solution-3d-unet-with-gaussian-h",
  "author_name": "",
  "post_date": "2025-02-19T15:21:54.187Z",
  "votes": 28,
  "comment_count": 4,
  "views": 0,
  "content": "<h1>7th place solution</h1>\n<p>First of all, I would like to sincerely thank the competition organizers, dataset providers, and the competitors for this interesting competition. It was very fun experimenting with different models, and I learned a lot in this competition. This is the solution for my 7th place submission.</p>\n<h2>Summary</h2>\n<p>My approach utilized a 3D segmentation models with per class gaussian heatmap prediction. The models were first pre-trained on the simulated dataset (1fold, wbp as val) then fine tuned to the experimental dataset (4 fold). I trained all models using weighted BCE with very high pos_weight, to make the model generate more predictions. Heavy augmentations were performed during training, including Mixup, Cutmix, RandomFlip, Affine (only rotate in the xy plane, used in pre-training only), rot90(only in the xy plane),  and other pixel-value augmentatons. The final submission consisted of an ensemble of three model soups (each combining four folds), applying 4x test-time augmentation (TTA) and a 4x sliding window approach.</p>\n<h2>Pre-processing and data augmentations</h2>\n<p>To reduce the pre-training time, I used the whole simulated dataset for training and used the wbp version of the experimental dataset for validation.</p>\n<p>Preprocessing was minimal, involving percentile clipping (0.1–99.9%) and dataset-specific scaling factors:</p>\n<ul>\n<li>Simulated: ×1.0</li>\n<li>WBP: ×1e4</li>\n<li>Denoised: ×1e5</li>\n</ul>\n<p>I used a sliding window approach with a stride of 64, excluding the last window, to create 1x8x8=64 windows per experiment (0, 64, 128, …, 384, 448 for x and y, 0 for z), then randomly shifted the crop coordinates when generating the training data.</p>\n<p>The augmentations I used were:</p>\n<ul>\n<li>shift (up to 64 from the predetermined crop window)</li>\n<li>cutmix</li>\n<li>mixup</li>\n<li>randomflip (all axis)</li>\n<li>rot90 (only on xy plane)</li>\n<li>affine (scale on every axis, rotate only on xy plane)*only when pre-training</li>\n<li>contrast (0.7 ~ 1.5)</li>\n<li>gamma (0.8 ~ 1.2)</li>\n<li>gaussian noise (0 ~ 0.05)*only when fine tuning</li>\n</ul>\n<h2>Model architecture</h2>\n<p>I used 3 models in the final submission, 2 U-Net based models and 1 DeepLab based model. Among the two U-net based models, one has a backbone of ResNet50d and the other has a backbone of EfficientNetV2-M. The DeepLab based model has a backbone of ResNet50d. I tried training a model with an offset prediction head to improve the localization performance but it did not work in my case, so the final submission only uses a simple segmentation head.  ModelEMA with a decay of 0.995 was for all models except model 207.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2710778%2F65ed0d4fea2275ad4cdc5be27e000385%2Fimage.png?generation=1739977456206695&amp;alt=media\" alt=\"Model description\"></p>\n<p>Out of the 5 feature stages, only using the first 4 stages resulted in comparable or better performance, so some models (resnet50d) in the ensemble do that. I'm guessing it was because the high-resolution feature maps contained better information needed to locate the particles.</p>\n<p>In the final ensemble, backbones of [resnet50d, resnet50d, efficientnetv2-m] were used.</p>\n<p>While it was not used in the final submission, DeepLabV3+ was also able to achieve relatively high scores, while only requiring half the inference time of the U-Net models. It may be able to create a better submission by using DeepLabV3+ and doing more ensembles.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Backbone</th>\n<th>No. skip connections</th>\n<th>Decoder channels</th>\n<th>Initial weights</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Model 207 &amp; 209</td>\n<td>ResNet50d</td>\n<td>4</td>\n<td>[128, 64, 32, 16]</td>\n<td>ra2_in1k</td>\n</tr>\n<tr>\n<td>Model 248 &amp; 249</td>\n<td>EfficientNetV2-M</td>\n<td>5</td>\n<td>[512, 176, 80, 48, 24]</td>\n<td>in21k_ft_in1k</td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Backbone</th>\n<th>No. Highres features</th>\n<th>No. seg features</th>\n<th>Initial weights</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Model 258 &amp; 265</td>\n<td>ResNet50d</td>\n<td>48</td>\n<td>256</td>\n<td>ra2_in1k</td>\n</tr>\n</tbody>\n</table>\n<h2>Loss function</h2>\n<p>The loss function was a weighted BCE loss, where more weights were given to the \"hard\" particles, and regions where the target was positive. The formula for the loss function is given by</p>\n<p>$$<br>\nL_{c, i} = -w_c \\left(p_c \\cdot y_{c, i} \\cdot \\log(t_{c, i}) + (1 - y_{c, i}) \\cdot \\log(1 - t_{c, i}) \\right)<br>\n$$<br>\n$$<br>\nLoss = \\frac{1}{5N} \\sum_{c=1}^{5} \\sum_{i=1}^{N} L_{c, i}<br>\n$$</p>\n<p>where \\(L_{c, i}\\) is the loss for class \\(c\\) and pixel \\(i\\), \\(w_c\\) is the weight for class \\(c\\), \\(p_c\\) is the weight for the positive pixels in class \\(c\\), \\(y_{c, i}\\) is the prediction for class \\(c\\) and pixel \\(i\\), and \\(t_{c, i}\\) is the target for class \\(c\\) and pixel \\(i\\). \\(N\\) is the number of pixels in the image, and the sum is taken over all the pixels and classes.</p>\n<p>The class weights and the positive weights are shown below.<br>\nThe classes are in the order of [apo-ferritin, beta-galactosidase, ribosome, thyroglobulin, virus-like-particle].</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Class Weights</th>\n<th>Positive Weights</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Model 207</td>\n<td>[1, 2, 1, 2, 1]</td>\n<td>[28, 54, 36, 48, 24]</td>\n</tr>\n<tr>\n<td>Model 209</td>\n<td>[1, 2, 1, 2, 1]</td>\n<td>[54, 96, 24, 36, 36]</td>\n</tr>\n<tr>\n<td>Model 248</td>\n<td>[1, 2, 1, 2, 1] × 5/7</td>\n<td>[24, 36, 36, 48, 20]</td>\n</tr>\n<tr>\n<td>Model 249</td>\n<td>[1, 2, 1, 2, 1] × 5/7</td>\n<td>[54, 96, 36, 36, 36]</td>\n</tr>\n<tr>\n<td>Model 258</td>\n<td>[1, 3, 2, 3, 1]</td>\n<td>[42, 48, 72, 85, 54]</td>\n</tr>\n<tr>\n<td>Model 265</td>\n<td>[1, 3, 2, 3, 1]</td>\n<td>[54, 96, 24, 36, 36]</td>\n</tr>\n</tbody>\n</table>\n<h2>Inference</h2>\n<p>Inference was performed on padded volumes of size (192, 128, 128). I created 81 windows out of the [184, 630, 630] tomogram with 64 stride on the x and y axis, resulting in 4x TTA. When combining the results to a single heatmap, averaging the logits was better than averaging the sigmoid-ed heatmaps.<br>\nTo reduce the edge artifact, I also applied a sloped weight function that goes down near the edge.</p>\n<p>Flip augmentations were done as TTA. Depending on the number of models, 8x to 3x (normal, flipx, flipy) TTA was applied.</p>\n<h2>Postprocessing</h2>\n<p>Initially I was using CCL + center of mass to determine the coordinates but I switched to using maxpool to detect the local maxima. To reduce noise, gaussian blur with kernel size 5 and sigma 1.0 was applied before the maxpool, and weighted box fusion with radius=0.5*particle_radius was applied after the local maxima detection. Using the maxpool method significantly improved the inference speed with minimal performance loss (~0.002 reduction), enabling me to do more TTAs and ensembles.</p>\n<p>Due to the very high pos_weight, threshold of around 0.5 was enough to make the model predict many FPs. The final submission used thresholds of [0.3, 0.2, -0.2, -0.2, -0.2] in the <strong>logit space</strong>.</p>\n<h2>Ensemble strategy</h2>\n<p>When ensembling the output of the different models, averaging the logits of the heatmaps were better compared to averaging the sigmoid-ed heatmap or applying WBF on the final predictions. 3 models with 4x Flip TTA, 4x Sliding Window seemed to be the best balance between TTA and ensemble.</p>\n<h2>Probing the LB</h2>\n<p>By making 3 submissions with only the targeted class</p>\n<ul>\n<li>Normal</li>\n<li>Added FP (like appending (0, 0, 0) * 20 to every prediction)</li>\n<li>Doubled prediction<br>\nit is able to calculate the number of TP, FP, FNs by solving the system of equations:</li>\n</ul>\n<pre><code>mat = np.array()\nTP, FP, FN = np.linalg.solve(mat, )\n</code></pre>\n<p>After the total number of particles is obtained, TP, FP, FN can be calculated from 2 submissions. (by changing the bottom row of the coefficient matrix to  <code>1, 0, 1</code> and the target to <code>0, -added_fp*more_fp_score, num_particles</code>)</p>\n<p>At first I was afraid to use this information to tune my model but the probed particle counts (TP + FN) matched the average counts stated in the paper (<a href=\"https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1)\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1)</a>, so I assumed that the private test set had the same particle distribution as the public test set. I was trying to keep my submission count low in the first half of the competition to not overfit to the LB but after this I decided to trust the LB over my CV.</p>\n<h2>Scores</h2>\n<table>\n<thead>\n<tr>\n<th>Description</th>\n<th>Public Score</th>\n<th>Private Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Model 209</td>\n<td>0.77320</td>\n<td>0.76665</td>\n</tr>\n<tr>\n<td>Model 249</td>\n<td>0.77024</td>\n<td>0.76478</td>\n</tr>\n<tr>\n<td>Model 265</td>\n<td>0.76824</td>\n<td>0.76057</td>\n</tr>\n<tr>\n<td>Ensemble of the three</td>\n<td>0.78351</td>\n<td>0.77708</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "3116629",
      "postDate": "02/06/2025 06:26:59",
      "content": "<h1>7th place solution</h1>\n<p>First of all, I would like to sincerely thank the competition organizers, dataset providers, and the competitors for this interesting competition. It was very fun experimenting with different models, and I learned a lot in this competition. This is the solution for my 7th place submission.</p>\n<h2>Summary</h2>\n<p>My approach utilized a 3D segmentation models with per class gaussian heatmap prediction. The models were first pre-trained on the simulated dataset (1fold, wbp as val) then fine tuned to the experimental dataset (4 fold). I trained all models using weighted BCE with very high pos_weight, to make the model generate more predictions. Heavy augmentations were performed during training, including Mixup, Cutmix, RandomFlip, Affine (only rotate in the xy plane, used in pre-training only), rot90(only in the xy plane),  and other pixel-value augmentatons. The final submission consisted of an ensemble of three model soups (each combining four folds), applying 4x test-time augmentation (TTA) and a 4x sliding window approach.</p>\n<h2>Pre-processing and data augmentations</h2>\n<p>To reduce the pre-training time, I used the whole simulated dataset for training and used the wbp version of the experimental dataset for validation.</p>\n<p>Preprocessing was minimal, involving percentile clipping (0.1–99.9%) and dataset-specific scaling factors:</p>\n<ul>\n<li>Simulated: ×1.0</li>\n<li>WBP: ×1e4</li>\n<li>Denoised: ×1e5</li>\n</ul>\n<p>I used a sliding window approach with a stride of 64, excluding the last window, to create 1x8x8=64 windows per experiment (0, 64, 128, …, 384, 448 for x and y, 0 for z), then randomly shifted the crop coordinates when generating the training data.</p>\n<p>The augmentations I used were:</p>\n<ul>\n<li>shift (up to 64 from the predetermined crop window)</li>\n<li>cutmix</li>\n<li>mixup</li>\n<li>randomflip (all axis)</li>\n<li>rot90 (only on xy plane)</li>\n<li>affine (scale on every axis, rotate only on xy plane)*only when pre-training</li>\n<li>contrast (0.7 ~ 1.5)</li>\n<li>gamma (0.8 ~ 1.2)</li>\n<li>gaussian noise (0 ~ 0.05)*only when fine tuning</li>\n</ul>\n<h2>Model architecture</h2>\n<p>I used 3 models in the final submission, 2 U-Net based models and 1 DeepLab based model. Among the two U-net based models, one has a backbone of ResNet50d and the other has a backbone of EfficientNetV2-M. The DeepLab based model has a backbone of ResNet50d. I tried training a model with an offset prediction head to improve the localization performance but it did not work in my case, so the final submission only uses a simple segmentation head.  ModelEMA with a decay of 0.995 was for all models except model 207.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2710778%2F65ed0d4fea2275ad4cdc5be27e000385%2Fimage.png?generation=1739977456206695&amp;alt=media\" alt=\"Model description\"></p>\n<p>Out of the 5 feature stages, only using the first 4 stages resulted in comparable or better performance, so some models (resnet50d) in the ensemble do that. I'm guessing it was because the high-resolution feature maps contained better information needed to locate the particles.</p>\n<p>In the final ensemble, backbones of [resnet50d, resnet50d, efficientnetv2-m] were used.</p>\n<p>While it was not used in the final submission, DeepLabV3+ was also able to achieve relatively high scores, while only requiring half the inference time of the U-Net models. It may be able to create a better submission by using DeepLabV3+ and doing more ensembles.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Backbone</th>\n<th>No. skip connections</th>\n<th>Decoder channels</th>\n<th>Initial weights</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Model 207 &amp; 209</td>\n<td>ResNet50d</td>\n<td>4</td>\n<td>[128, 64, 32, 16]</td>\n<td>ra2_in1k</td>\n</tr>\n<tr>\n<td>Model 248 &amp; 249</td>\n<td>EfficientNetV2-M</td>\n<td>5</td>\n<td>[512, 176, 80, 48, 24]</td>\n<td>in21k_ft_in1k</td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Backbone</th>\n<th>No. Highres features</th>\n<th>No. seg features</th>\n<th>Initial weights</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Model 258 &amp; 265</td>\n<td>ResNet50d</td>\n<td>48</td>\n<td>256</td>\n<td>ra2_in1k</td>\n</tr>\n</tbody>\n</table>\n<h2>Loss function</h2>\n<p>The loss function was a weighted BCE loss, where more weights were given to the \"hard\" particles, and regions where the target was positive. The formula for the loss function is given by</p>\n<p>$$<br>\nL_{c, i} = -w_c \\left(p_c \\cdot y_{c, i} \\cdot \\log(t_{c, i}) + (1 - y_{c, i}) \\cdot \\log(1 - t_{c, i}) \\right)<br>\n$$<br>\n$$<br>\nLoss = \\frac{1}{5N} \\sum_{c=1}^{5} \\sum_{i=1}^{N} L_{c, i}<br>\n$$</p>\n<p>where \\(L_{c, i}\\) is the loss for class \\(c\\) and pixel \\(i\\), \\(w_c\\) is the weight for class \\(c\\), \\(p_c\\) is the weight for the positive pixels in class \\(c\\), \\(y_{c, i}\\) is the prediction for class \\(c\\) and pixel \\(i\\), and \\(t_{c, i}\\) is the target for class \\(c\\) and pixel \\(i\\). \\(N\\) is the number of pixels in the image, and the sum is taken over all the pixels and classes.</p>\n<p>The class weights and the positive weights are shown below.<br>\nThe classes are in the order of [apo-ferritin, beta-galactosidase, ribosome, thyroglobulin, virus-like-particle].</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Class Weights</th>\n<th>Positive Weights</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Model 207</td>\n<td>[1, 2, 1, 2, 1]</td>\n<td>[28, 54, 36, 48, 24]</td>\n</tr>\n<tr>\n<td>Model 209</td>\n<td>[1, 2, 1, 2, 1]</td>\n<td>[54, 96, 24, 36, 36]</td>\n</tr>\n<tr>\n<td>Model 248</td>\n<td>[1, 2, 1, 2, 1] × 5/7</td>\n<td>[24, 36, 36, 48, 20]</td>\n</tr>\n<tr>\n<td>Model 249</td>\n<td>[1, 2, 1, 2, 1] × 5/7</td>\n<td>[54, 96, 36, 36, 36]</td>\n</tr>\n<tr>\n<td>Model 258</td>\n<td>[1, 3, 2, 3, 1]</td>\n<td>[42, 48, 72, 85, 54]</td>\n</tr>\n<tr>\n<td>Model 265</td>\n<td>[1, 3, 2, 3, 1]</td>\n<td>[54, 96, 24, 36, 36]</td>\n</tr>\n</tbody>\n</table>\n<h2>Inference</h2>\n<p>Inference was performed on padded volumes of size (192, 128, 128). I created 81 windows out of the [184, 630, 630] tomogram with 64 stride on the x and y axis, resulting in 4x TTA. When combining the results to a single heatmap, averaging the logits was better than averaging the sigmoid-ed heatmaps.<br>\nTo reduce the edge artifact, I also applied a sloped weight function that goes down near the edge.</p>\n<p>Flip augmentations were done as TTA. Depending on the number of models, 8x to 3x (normal, flipx, flipy) TTA was applied.</p>\n<h2>Postprocessing</h2>\n<p>Initially I was using CCL + center of mass to determine the coordinates but I switched to using maxpool to detect the local maxima. To reduce noise, gaussian blur with kernel size 5 and sigma 1.0 was applied before the maxpool, and weighted box fusion with radius=0.5*particle_radius was applied after the local maxima detection. Using the maxpool method significantly improved the inference speed with minimal performance loss (~0.002 reduction), enabling me to do more TTAs and ensembles.</p>\n<p>Due to the very high pos_weight, threshold of around 0.5 was enough to make the model predict many FPs. The final submission used thresholds of [0.3, 0.2, -0.2, -0.2, -0.2] in the <strong>logit space</strong>.</p>\n<h2>Ensemble strategy</h2>\n<p>When ensembling the output of the different models, averaging the logits of the heatmaps were better compared to averaging the sigmoid-ed heatmap or applying WBF on the final predictions. 3 models with 4x Flip TTA, 4x Sliding Window seemed to be the best balance between TTA and ensemble.</p>\n<h2>Probing the LB</h2>\n<p>By making 3 submissions with only the targeted class</p>\n<ul>\n<li>Normal</li>\n<li>Added FP (like appending (0, 0, 0) * 20 to every prediction)</li>\n<li>Doubled prediction<br>\nit is able to calculate the number of TP, FP, FNs by solving the system of equations:</li>\n</ul>\n<pre><code>mat = np.array()\nTP, FP, FN = np.linalg.solve(mat, )\n</code></pre>\n<p>After the total number of particles is obtained, TP, FP, FN can be calculated from 2 submissions. (by changing the bottom row of the coefficient matrix to  <code>1, 0, 1</code> and the target to <code>0, -added_fp*more_fp_score, num_particles</code>)</p>\n<p>At first I was afraid to use this information to tune my model but the probed particle counts (TP + FN) matched the average counts stated in the paper (<a href=\"https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1)\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1)</a>, so I assumed that the private test set had the same particle distribution as the public test set. I was trying to keep my submission count low in the first half of the competition to not overfit to the LB but after this I decided to trust the LB over my CV.</p>\n<h2>Scores</h2>\n<table>\n<thead>\n<tr>\n<th>Description</th>\n<th>Public Score</th>\n<th>Private Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Model 209</td>\n<td>0.77320</td>\n<td>0.76665</td>\n</tr>\n<tr>\n<td>Model 249</td>\n<td>0.77024</td>\n<td>0.76478</td>\n</tr>\n<tr>\n<td>Model 265</td>\n<td>0.76824</td>\n<td>0.76057</td>\n</tr>\n<tr>\n<td>Ensemble of the three</td>\n<td>0.78351</td>\n<td>0.77708</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "# 7th place solution\n\nFirst of all, I would like to sincerely thank the competition organizers, dataset providers, and the competitors for this interesting competition. It was very fun experimenting with different models, and I learned a lot in this competition. This is the solution for my 7th place submission.\n\n## Summary\n\nMy approach utilized a 3D segmentation models with per class gaussian heatmap prediction. The models were first pre-trained on the simulated dataset (1fold, wbp as val) then fine tuned to the experimental dataset (4 fold). I trained all models using weighted BCE with very high pos_weight, to make the model generate more predictions. Heavy augmentations were performed during training, including Mixup, Cutmix, RandomFlip, Affine (only rotate in the xy plane, used in pre-training only), rot90(only in the xy plane),  and other pixel-value augmentatons. The final submission consisted of an ensemble of three model soups (each combining four folds), applying 4x test-time augmentation (TTA) and a 4x sliding window approach.\n\n## Pre-processing and data augmentations\n\nTo reduce the pre-training time, I used the whole simulated dataset for training and used the wbp version of the experimental dataset for validation.\n\nPreprocessing was minimal, involving percentile clipping (0.1–99.9%) and dataset-specific scaling factors:\n- Simulated: ×1.0\n- WBP: ×1e4\n- Denoised: ×1e5\n\nI used a sliding window approach with a stride of 64, excluding the last window, to create 1x8x8=64 windows per experiment (0, 64, 128, ..., 384, 448 for x and y, 0 for z), then randomly shifted the crop coordinates when generating the training data.\n\nThe augmentations I used were:\n- shift (up to 64 from the predetermined crop window)\n- cutmix\n- mixup\n- randomflip (all axis)\n- rot90 (only on xy plane)\n- affine (scale on every axis, rotate only on xy plane)\\*only when pre-training\n- contrast (0.7 ~ 1.5)\n- gamma (0.8 ~ 1.2)\n- gaussian noise (0 ~ 0.05)\\*only when fine tuning\n\n## Model architecture\n\nI used 3 models in the final submission, 2 U-Net based models and 1 DeepLab based model. Among the two U-net based models, one has a backbone of ResNet50d and the other has a backbone of EfficientNetV2-M. The DeepLab based model has a backbone of ResNet50d. I tried training a model with an offset prediction head to improve the localization performance but it did not work in my case, so the final submission only uses a simple segmentation head.  ModelEMA with a decay of 0.995 was for all models except model 207.\n\n![Model description](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2710778%2F65ed0d4fea2275ad4cdc5be27e000385%2Fimage.png?generation=1739977456206695&alt=media)\n\nOut of the 5 feature stages, only using the first 4 stages resulted in comparable or better performance, so some models (resnet50d) in the ensemble do that. I'm guessing it was because the high-resolution feature maps contained better information needed to locate the particles.\n\nIn the final ensemble, backbones of [resnet50d, resnet50d, efficientnetv2-m] were used.\n\nWhile it was not used in the final submission, DeepLabV3+ was also able to achieve relatively high scores, while only requiring half the inference time of the U-Net models. It may be able to create a better submission by using DeepLabV3+ and doing more ensembles.\n\n| Model            | Backbone         | No. skip connections | Decoder channels         | Initial weights   |\n|------------------|------------------|----------------------|--------------------------|-------------------|\n| Model 207 & 209  | ResNet50d        | 4                    | [128, 64, 32, 16]        | ra2_in1k          |\n| Model 248 & 249  | EfficientNetV2-M | 5                    | [512, 176, 80, 48, 24]    | in21k_ft_in1k     |\n\n| Model           | Backbone  | No. Highres features | No. seg features | Initial weights |\n|-----------------|-----------|----------------------|------------------|-----------------|\n| Model 258 & 265 | ResNet50d | 48                   | 256              | ra2_in1k       |\n\n## Loss function\n\nThe loss function was a weighted BCE loss, where more weights were given to the \"hard\" particles, and regions where the target was positive. The formula for the loss function is given by\n\n$$\nL_{c, i} = -w_c \\left(p_c \\cdot y_{c, i} \\cdot \\log(t_{c, i}) + (1 - y_{c, i}) \\cdot \\log(1 - t_{c, i}) \\right)\n$$\n$$\nLoss = \\frac{1}{5N} \\sum_{c=1}^{5} \\sum_{i=1}^{N} L_{c, i}\n$$\n\nwhere \\\\(L_{c, i}\\\\) is the loss for class \\\\(c\\\\) and pixel \\\\(i\\\\), \\\\(w_c\\\\) is the weight for class \\\\(c\\\\), \\\\(p_c\\\\) is the weight for the positive pixels in class \\\\(c\\\\), \\\\(y_{c, i}\\\\) is the prediction for class \\\\(c\\\\) and pixel \\\\(i\\\\), and \\\\(t_{c, i}\\\\) is the target for class \\\\(c\\\\) and pixel \\\\(i\\\\). \\\\(N\\\\) is the number of pixels in the image, and the sum is taken over all the pixels and classes.\n\nThe class weights and the positive weights are shown below.\nThe classes are in the order of [apo-ferritin, beta-galactosidase, ribosome, thyroglobulin, virus-like-particle].\n| Model      | Class Weights                | Positive Weights        |\n|------------|------------------------------|-------------------------|\n| Model 207  | [1, 2, 1, 2, 1]              | [28, 54, 36, 48, 24]     |\n| Model 209  | [1, 2, 1, 2, 1]              | [54, 96, 24, 36, 36]     |\n| Model 248  | [1, 2, 1, 2, 1] × 5/7         | [24, 36, 36, 48, 20]     |\n| Model 249  | [1, 2, 1, 2, 1] × 5/7         | [54, 96, 36, 36, 36]     |\n| Model 258  | [1, 3, 2, 3, 1]              | [42, 48, 72, 85, 54]     |\n| Model 265  | [1, 3, 2, 3, 1]              | [54, 96, 24, 36, 36]     |\n\n## Inference\n\nInference was performed on padded volumes of size (192, 128, 128). I created 81 windows out of the [184, 630, 630] tomogram with 64 stride on the x and y axis, resulting in 4x TTA. When combining the results to a single heatmap, averaging the logits was better than averaging the sigmoid-ed heatmaps.\nTo reduce the edge artifact, I also applied a sloped weight function that goes down near the edge.\n\nFlip augmentations were done as TTA. Depending on the number of models, 8x to 3x (normal, flipx, flipy) TTA was applied.\n\n## Postprocessing\n\nInitially I was using CCL + center of mass to determine the coordinates but I switched to using maxpool to detect the local maxima. To reduce noise, gaussian blur with kernel size 5 and sigma 1.0 was applied before the maxpool, and weighted box fusion with radius=0.5*particle_radius was applied after the local maxima detection. Using the maxpool method significantly improved the inference speed with minimal performance loss (~0.002 reduction), enabling me to do more TTAs and ensembles.\n\nDue to the very high pos_weight, threshold of around 0.5 was enough to make the model predict many FPs. The final submission used thresholds of [0.3, 0.2, -0.2, -0.2, -0.2] in the **logit space**.\n\n## Ensemble strategy\n\nWhen ensembling the output of the different models, averaging the logits of the heatmaps were better compared to averaging the sigmoid-ed heatmap or applying WBF on the final predictions. 3 models with 4x Flip TTA, 4x Sliding Window seemed to be the best balance between TTA and ensemble.\n\n\n## Probing the LB\n\nBy making 3 submissions with only the targeted class\n- Normal\n- Added FP (like appending (0, 0, 0) * 20 to every prediction)\n- Doubled prediction\nit is able to calculate the number of TP, FP, FNs by solving the system of equations:\n```\nmat = np.array([\n    \t[17*(original_score - 1), original_score, 16*original_score],\n    \t[17*(more_fp_score - 1), more_fp_score, 16*more_fp_score],\n    \t[18*double_pred_score - 17, 2*double_pred_score, 16*double_pred_score]\n\t])\nTP, FP, FN = np.linalg.solve(mat, [0, -added_fp*more_fp_score, 0])\n```\n\nAfter the total number of particles is obtained, TP, FP, FN can be calculated from 2 submissions. (by changing the bottom row of the coefficient matrix to  `1, 0, 1` and the target to `0, -added_fp*more_fp_score, num_particles`)\n\n\nAt first I was afraid to use this information to tune my model but the probed particle counts (TP + FN) matched the average counts stated in the paper (https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1), so I assumed that the private test set had the same particle distribution as the public test set. I was trying to keep my submission count low in the first half of the competition to not overfit to the LB but after this I decided to trust the LB over my CV.\n\n\n## Scores\n| Description                     | Public Score | Private Score |\n|---------------------------------|--------------|---------------|\n| Model 209                             | 0.77320      | 0.76665       |\n| Model 249                             | 0.77024      | 0.76478       |\n| Model 265                             | 0.76824      | 0.76057       |\n|Ensemble of the three| 0.78351 | 0.77708 |",
      "votes": null
    },
    {
      "id": "3116832",
      "postDate": "02/06/2025 11:02:26",
      "content": "<p>Congratulations on your solo gold and the prize! That's a clever technique to probe the LB. </p>",
      "rawMarkdown": "Congratulations on your solo gold and the prize! That's a clever technique to probe the LB.",
      "votes": null
    },
    {
      "id": "3128454",
      "postDate": "02/19/2025 15:16:36",
      "content": "<p>It has been two weeks since the competition has ended, but I realized I was making an unbelievably stupid mistake of mixing up Experiment 265 (DeepLabV3+ architecture) with Experiment 256 (U-Net architecture).  I am really sorry that I did not noticed this sooner. Also, the details of the loss function and the scores of the model has been added.</p>",
      "rawMarkdown": "It has been two weeks since the competition has ended, but I realized I was making an unbelievably stupid mistake of mixing up Experiment 265 (DeepLabV3+ architecture) with Experiment 256 (U-Net architecture).  I am really sorry that I did not noticed this sooner. Also, the details of the loss function and the scores of the model has been added.",
      "votes": null
    },
    {
      "id": "3129915",
      "postDate": "02/21/2025 05:57:17",
      "content": "<p>Your writeup is impressively detailed—I truly appreciate the depth of insights you’ve shared.</p>",
      "rawMarkdown": "Your writeup is impressively detailed—I truly appreciate the depth of insights you’ve shared.",
      "votes": null
    },
    {
      "id": "3365829",
      "postDate": "12/07/2025 11:49:30",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kokikoba\" target=\"_blank\">@kokikoba</a>, Nice Writeup!! How were the class_weights and positive_weights were determined for each model?? Do you use any specific algorithm?</p>",
      "rawMarkdown": "Hi @kokikoba, Nice Writeup!! How were the class_weights and positive_weights were determined for each model?? Do you use any specific algorithm?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3116832,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "02/06/2025 11:02:26",
      "content": "<p>Congratulations on your solo gold and the prize! That's a clever technique to probe the LB. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3128454,
      "author_name": "kokikoba",
      "author_url": "",
      "post_date": "02/19/2025 15:16:36",
      "content": "<p>It has been two weeks since the competition has ended, but I realized I was making an unbelievably stupid mistake of mixing up Experiment 265 (DeepLabV3+ architecture) with Experiment 256 (U-Net architecture).  I am really sorry that I did not noticed this sooner. Also, the details of the loss function and the scores of the model has been added.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3129915,
      "author_name": "linhanwang2",
      "author_url": "",
      "post_date": "02/21/2025 05:57:17",
      "content": "<p>Your writeup is impressively detailed—I truly appreciate the depth of insights you’ve shared.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3365829,
      "author_name": "gowrishankarp",
      "author_url": "",
      "post_date": "12/07/2025 11:49:30",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kokikoba\" target=\"_blank\">@kokikoba</a>, Nice Writeup!! How were the class_weights and positive_weights were determined for each model?? Do you use any specific algorithm?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3116629": "# 7th place solution\n\nFirst of all, I would like to sincerely thank the competition organizers, dataset providers, and the competitors for this interesting competition. It was very fun experimenting with different models, and I learned a lot in this competition. This is the solution for my 7th place submission.\n\n## Summary\n\nMy approach utilized a 3D segmentation models with per class gaussian heatmap prediction. The models were first pre-trained on the simulated dataset (1fold, wbp as val) then fine tuned to the experimental dataset (4 fold). I trained all models using weighted BCE with very high pos_weight, to make the model generate more predictions. Heavy augmentations were performed during training, including Mixup, Cutmix, RandomFlip, Affine (only rotate in the xy plane, used in pre-training only), rot90(only in the xy plane),  and other pixel-value augmentatons. The final submission consisted of an ensemble of three model soups (each combining four folds), applying 4x test-time augmentation (TTA) and a 4x sliding window approach.\n\n## Pre-processing and data augmentations\n\nTo reduce the pre-training time, I used the whole simulated dataset for training and used the wbp version of the experimental dataset for validation.\n\nPreprocessing was minimal, involving percentile clipping (0.1–99.9%) and dataset-specific scaling factors:\n- Simulated: ×1.0\n- WBP: ×1e4\n- Denoised: ×1e5\n\nI used a sliding window approach with a stride of 64, excluding the last window, to create 1x8x8=64 windows per experiment (0, 64, 128, ..., 384, 448 for x and y, 0 for z), then randomly shifted the crop coordinates when generating the training data.\n\nThe augmentations I used were:\n- shift (up to 64 from the predetermined crop window)\n- cutmix\n- mixup\n- randomflip (all axis)\n- rot90 (only on xy plane)\n- affine (scale on every axis, rotate only on xy plane)\\*only when pre-training\n- contrast (0.7 ~ 1.5)\n- gamma (0.8 ~ 1.2)\n- gaussian noise (0 ~ 0.05)\\*only when fine tuning\n\n## Model architecture\n\nI used 3 models in the final submission, 2 U-Net based models and 1 DeepLab based model. Among the two U-net based models, one has a backbone of ResNet50d and the other has a backbone of EfficientNetV2-M. The DeepLab based model has a backbone of ResNet50d. I tried training a model with an offset prediction head to improve the localization performance but it did not work in my case, so the final submission only uses a simple segmentation head.  ModelEMA with a decay of 0.995 was for all models except model 207.\n\n![Model description](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2710778%2F65ed0d4fea2275ad4cdc5be27e000385%2Fimage.png?generation=1739977456206695&alt=media)\n\nOut of the 5 feature stages, only using the first 4 stages resulted in comparable or better performance, so some models (resnet50d) in the ensemble do that. I'm guessing it was because the high-resolution feature maps contained better information needed to locate the particles.\n\nIn the final ensemble, backbones of [resnet50d, resnet50d, efficientnetv2-m] were used.\n\nWhile it was not used in the final submission, DeepLabV3+ was also able to achieve relatively high scores, while only requiring half the inference time of the U-Net models. It may be able to create a better submission by using DeepLabV3+ and doing more ensembles.\n\n| Model            | Backbone         | No. skip connections | Decoder channels         | Initial weights   |\n|------------------|------------------|----------------------|--------------------------|-------------------|\n| Model 207 & 209  | ResNet50d        | 4                    | [128, 64, 32, 16]        | ra2_in1k          |\n| Model 248 & 249  | EfficientNetV2-M | 5                    | [512, 176, 80, 48, 24]    | in21k_ft_in1k     |\n\n| Model           | Backbone  | No. Highres features | No. seg features | Initial weights |\n|-----------------|-----------|----------------------|------------------|-----------------|\n| Model 258 & 265 | ResNet50d | 48                   | 256              | ra2_in1k       |\n\n## Loss function\n\nThe loss function was a weighted BCE loss, where more weights were given to the \"hard\" particles, and regions where the target was positive. The formula for the loss function is given by\n\n$$\nL_{c, i} = -w_c \\left(p_c \\cdot y_{c, i} \\cdot \\log(t_{c, i}) + (1 - y_{c, i}) \\cdot \\log(1 - t_{c, i}) \\right)\n$$\n$$\nLoss = \\frac{1}{5N} \\sum_{c=1}^{5} \\sum_{i=1}^{N} L_{c, i}\n$$\n\nwhere \\\\(L_{c, i}\\\\) is the loss for class \\\\(c\\\\) and pixel \\\\(i\\\\), \\\\(w_c\\\\) is the weight for class \\\\(c\\\\), \\\\(p_c\\\\) is the weight for the positive pixels in class \\\\(c\\\\), \\\\(y_{c, i}\\\\) is the prediction for class \\\\(c\\\\) and pixel \\\\(i\\\\), and \\\\(t_{c, i}\\\\) is the target for class \\\\(c\\\\) and pixel \\\\(i\\\\). \\\\(N\\\\) is the number of pixels in the image, and the sum is taken over all the pixels and classes.\n\nThe class weights and the positive weights are shown below.\nThe classes are in the order of [apo-ferritin, beta-galactosidase, ribosome, thyroglobulin, virus-like-particle].\n| Model      | Class Weights                | Positive Weights        |\n|------------|------------------------------|-------------------------|\n| Model 207  | [1, 2, 1, 2, 1]              | [28, 54, 36, 48, 24]     |\n| Model 209  | [1, 2, 1, 2, 1]              | [54, 96, 24, 36, 36]     |\n| Model 248  | [1, 2, 1, 2, 1] × 5/7         | [24, 36, 36, 48, 20]     |\n| Model 249  | [1, 2, 1, 2, 1] × 5/7         | [54, 96, 36, 36, 36]     |\n| Model 258  | [1, 3, 2, 3, 1]              | [42, 48, 72, 85, 54]     |\n| Model 265  | [1, 3, 2, 3, 1]              | [54, 96, 24, 36, 36]     |\n\n## Inference\n\nInference was performed on padded volumes of size (192, 128, 128). I created 81 windows out of the [184, 630, 630] tomogram with 64 stride on the x and y axis, resulting in 4x TTA. When combining the results to a single heatmap, averaging the logits was better than averaging the sigmoid-ed heatmaps.\nTo reduce the edge artifact, I also applied a sloped weight function that goes down near the edge.\n\nFlip augmentations were done as TTA. Depending on the number of models, 8x to 3x (normal, flipx, flipy) TTA was applied.\n\n## Postprocessing\n\nInitially I was using CCL + center of mass to determine the coordinates but I switched to using maxpool to detect the local maxima. To reduce noise, gaussian blur with kernel size 5 and sigma 1.0 was applied before the maxpool, and weighted box fusion with radius=0.5*particle_radius was applied after the local maxima detection. Using the maxpool method significantly improved the inference speed with minimal performance loss (~0.002 reduction), enabling me to do more TTAs and ensembles.\n\nDue to the very high pos_weight, threshold of around 0.5 was enough to make the model predict many FPs. The final submission used thresholds of [0.3, 0.2, -0.2, -0.2, -0.2] in the **logit space**.\n\n## Ensemble strategy\n\nWhen ensembling the output of the different models, averaging the logits of the heatmaps were better compared to averaging the sigmoid-ed heatmap or applying WBF on the final predictions. 3 models with 4x Flip TTA, 4x Sliding Window seemed to be the best balance between TTA and ensemble.\n\n\n## Probing the LB\n\nBy making 3 submissions with only the targeted class\n- Normal\n- Added FP (like appending (0, 0, 0) * 20 to every prediction)\n- Doubled prediction\nit is able to calculate the number of TP, FP, FNs by solving the system of equations:\n```\nmat = np.array([\n    \t[17*(original_score - 1), original_score, 16*original_score],\n    \t[17*(more_fp_score - 1), more_fp_score, 16*more_fp_score],\n    \t[18*double_pred_score - 17, 2*double_pred_score, 16*double_pred_score]\n\t])\nTP, FP, FN = np.linalg.solve(mat, [0, -added_fp*more_fp_score, 0])\n```\n\nAfter the total number of particles is obtained, TP, FP, FN can be calculated from 2 submissions. (by changing the bottom row of the coefficient matrix to  `1, 0, 1` and the target to `0, -added_fp*more_fp_score, num_particles`)\n\n\nAt first I was afraid to use this information to tune my model but the probed particle counts (TP + FN) matched the average counts stated in the paper (https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1), so I assumed that the private test set had the same particle distribution as the public test set. I was trying to keep my submission count low in the first half of the competition to not overfit to the LB but after this I decided to trust the LB over my CV.\n\n\n## Scores\n| Description                     | Public Score | Private Score |\n|---------------------------------|--------------|---------------|\n| Model 209                             | 0.77320      | 0.76665       |\n| Model 249                             | 0.77024      | 0.76478       |\n| Model 265                             | 0.76824      | 0.76057       |\n|Ensemble of the three| 0.78351 | 0.77708 |",
    "3116832": "Congratulations on your solo gold and the prize! That's a clever technique to probe the LB.",
    "3128454": "It has been two weeks since the competition has ended, but I realized I was making an unbelievably stupid mistake of mixing up Experiment 265 (DeepLabV3+ architecture) with Experiment 256 (U-Net architecture).  I am really sorry that I did not noticed this sooner. Also, the details of the loss function and the scores of the model has been added.",
    "3129915": "Your writeup is impressively detailed—I truly appreciate the depth of insights you’ve shared.",
    "3365829": "Hi @kokikoba, Nice Writeup!! How were the class_weights and positive_weights were determined for each model?? Do you use any specific algorithm?"
  },
  "source": "meta"
}