{
  "id": 561568,
  "title": "2nd Place Solution",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/561568",
  "author_name": "AnnieGo",
  "post_date": "2025-02-06T18:53:56.557000",
  "votes": 44,
  "comment_count": 11,
  "views": 0,
  "content": "<h2>Acknowledgments</h2>\n<p>We sincerely appreciate Kaggle and the competition organizers for offering this invaluable opportunity. We also extend our gratitude to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and <a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a> for their significant contributions. <a href=\"https://www.kaggle.com/code/fnands/baseline-unet-train-submit/notebook\" target=\"_blank\"><strong>fnands's notebook</strong></a> and <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/547350#3056648\" target=\"_blank\"><strong>hengck23's discussion</strong></a> provided a solid foundation for our approach. Finally, I am grateful to my teammate <a href=\"https://www.kaggle.com/luoziqian\" target=\"_blank\">@luoziqian</a> for the excellent collaboration.</p>\n<h2>Summary</h2>\n<p>Our approach is based on an ensemble of multiple lightweight segmentation models, with parameter sizes ranging from 873K to 14.2M. Following segmentation, we computed particle centroids using CC3D and filtered small clusters based on voxel count statistics.</p>\n<h2>Overall Pipeline</h2>\n<p>The overall pipeline is illustrated below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5483160%2F1b63c4ad09d3ec66e33450cb95263adb%2Foverall_pipeline.png?generation=1739272382909281&amp;alt=media\"></p>\n<h2>Model Architecture</h2>\n<p>We initiated our experiments using the MONAI UNet baseline model. Throughout our trials, we observed that models with a large number of parameters were prone to overfitting and often performed worse than lighter models. Consequently, we opted for lightweight architectures such as UNet3D, VoxResNet, VoxHRNet, SegResNet, DynUNet, and DenseVNet.</p>\n<p>However, while experimenting with VoxResNet, VoxHRNet, and DenseVNet, we encountered stability issues, including performance fluctuations and difficulties in convergence. Upon further analysis, we identified that MONAI UNet employs <code>InstanceNorm3d</code> and <code>PReLU</code>. By modifying the normalization and activation layers accordingly, we achieved more stable model performance.</p>\n<p>The model architecture of MONAI UNet is shown below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5483160%2F35bb85a0777083a30ffc41461aa2c154%2Fmonai_unet.png?generation=1738867977514788&amp;alt=media\" alt=\"\"></p>\n<p>The following figure compares the training performance of MONAI UNet with InstanceNorm3d and PReLU versus BatchNorm3d and ReLU across 5 experiments, demonstrating that the use of InstanceNorm3d and PReLU results in more stable training.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5483160%2Fcb7065ced216887833f5f653bea7df8d%2Ftraining_stability_comparison.png?generation=1739272596177762&amp;alt=media\"></p>\n<p>For the final ensemble, we selected models based on their public leaderboard scores:</p>\n<ul>\n<li>UNet3D</li>\n<li>VoxResNet</li>\n<li>VoxHRNet</li>\n<li>SegResNet</li>\n<li>DenseVNet</li>\n<li>UNet2E3D</li>\n</ul>\n<h2>Training Strategy</h2>\n<p>Our training configuration was designed to ensure stability and optimal performance. We found that the <strong>segmentation mask radius</strong>, <strong>loss function</strong> and <strong>data augmentation strategies</strong> played a crucial role in achieving reliable results.</p>\n<h4>Segmentation Mask:</h4>\n<p>We applied ground truth masks with a customized radius for each particle.</p>\n<table>\n<thead>\n<tr>\n<th><strong>Particle Types</strong></th>\n<th><strong>Default Radius</strong></th>\n<th><strong>Luoziqian's Radius</strong></th>\n<th><strong>Lion's Radius</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>apo-ferritin</td>\n<td>60</td>\n<td>60 x 0.5</td>\n<td>80 x 0.4</td>\n</tr>\n<tr>\n<td>beta-galactosidase</td>\n<td>90</td>\n<td>90 x 0.5</td>\n<td>90 x 0.4</td>\n</tr>\n<tr>\n<td>ribosome</td>\n<td>150</td>\n<td>150 x 0.5</td>\n<td>150 x 0.4</td>\n</tr>\n<tr>\n<td>thyroglobulin</td>\n<td>130</td>\n<td>130 x 0.5</td>\n<td>120 x 0.4</td>\n</tr>\n<tr>\n<td>virus-like-particle</td>\n<td>135</td>\n<td>135 x 0.5</td>\n<td>150 x 0.4</td>\n</tr>\n</tbody>\n</table>\n<h4>Training Settings:</h4>\n<p><strong>Lion's</strong> models:</p>\n<ul>\n<li>patch size <code>[128, 256, 256]</code> or <code>[128, 384, 384]</code></li>\n<li><code>200</code> epochs</li>\n<li>7 classes</li>\n<li>Optimizer: AdamW with a learning rate of <code>0.001</code> (no learning rate scheduler)</li>\n<li>Batch size: <code>2</code> or <code>4</code></li>\n<li>Drop path rate: <code>0.1</code> or <code>0.3</code></li>\n<li>Model selection based on the F-beta metric, save top-5 best model</li>\n<li>Loss functions:<ul>\n<li>Tversky Loss</li>\n<li>Cross-Entropy Loss with weights <code>[1.0, 1.0, 0.0, 2.0, 1.0, 2.0, 1.0]</code></li></ul></li>\n</ul>\n<p><strong>Luoziqian's</strong> models</p>\n<ul>\n<li>patch size <code>[128, 200, 200]</code> or <code>[128, 256, 256]</code></li>\n<li><code>100</code> or <code>300</code> epochs</li>\n<li>6 classes</li>\n<li>Early stopping with <code>patience=20</code></li>\n<li>Batch size: <code>1</code> or <code>2</code></li>\n<li>Model selection is based on the evaluation loss. The last 100 models are saved and their F-beta scores are verified</li>\n<li>Loss functions:<ul>\n<li>Dice Loss</li>\n<li>Tversky Loss</li>\n<li>Cross-Entropy Loss, scaled to 0.05 or 0.1</li></ul></li>\n</ul>\n<h4>Augmentations:</h4>\n<p><strong>Lion's</strong> models:</p>\n<ul>\n<li>RandCropByLabelClassesd with ratios <code>[1, 1, 1, 1, 2, 1, 2]</code> for balanced sampling</li>\n<li>RandRotate90d</li>\n<li>RandFlipd</li>\n<li>RandAffined</li>\n<li>RandGaussianNoised</li>\n</ul>\n<p><strong>Luoziqian's</strong> models</p>\n<ul>\n<li>RandCropByLabelClassesd</li>\n<li>RandRotate90d</li>\n<li>RandFlipd</li>\n<li>RandShiftIntensityd</li>\n</ul>\n<h4>Model Performance with 7 TTA Summary (Partial Selection)</h4>\n<table>\n<thead>\n<tr>\n<th><strong>No</strong></th>\n<th><strong>Model</strong></th>\n<th><strong>Developer</strong></th>\n<th><strong>Architecture</strong></th>\n<th><strong>Parameters</strong></th>\n<th><strong>Valid ID</strong></th>\n<th><strong>Normalization</strong></th>\n<th><strong>Activation</strong></th>\n<th><strong>Public LB</strong></th>\n<th><strong>Private LB</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>epoch122-step2952-valid_loss0.3625-val_metric0.8367.ckpt</td>\n<td>Lion</td>\n<td>UNet3D</td>\n<td>1.1M</td>\n<td>TS_86_3</td>\n<td>InstanceNorm3d</td>\n<td>PReLU</td>\n<td>0.77379</td>\n<td>0.76582</td>\n</tr>\n<tr>\n<td>2</td>\n<td>epoch148-step3576-valid_loss1.1154-val_metric0.7722.ckpt</td>\n<td>Lion</td>\n<td>UNet3D</td>\n<td>1.1M</td>\n<td>TS_6_4</td>\n<td>InstanceNorm3d</td>\n<td>PReLU</td>\n<td>0.77021</td>\n<td>0.76725</td>\n</tr>\n<tr>\n<td>3</td>\n<td>epoch153-step3696-valid_loss0.3021-val_metric0.8900.ckpt</td>\n<td>Lion</td>\n<td>UNet3D</td>\n<td>1.6M</td>\n<td>TS_69_2</td>\n<td>InstanceNorm3d</td>\n<td>PReLU</td>\n<td>0.77205</td>\n<td>0.76676</td>\n</tr>\n<tr>\n<td>4</td>\n<td>epoch194-step4680-valid_loss1.0213-val_metric0.8788.ckpt</td>\n<td>Lion</td>\n<td>UNet3D</td>\n<td>1.1M</td>\n<td>TS_69_2</td>\n<td>InstanceNorm3d</td>\n<td>PReLU</td>\n<td>0.77390</td>\n<td>0.76737</td>\n</tr>\n<tr>\n<td>5</td>\n<td>epoch138-step3336-valid_loss0.3690-val_metric0.8476.ckpt</td>\n<td>Lion</td>\n<td>UNet3D</td>\n<td>1.1M</td>\n<td>TS_73_6</td>\n<td>InstanceNorm3d</td>\n<td>PReLU</td>\n<td>0.76543</td>\n<td>0.76025</td>\n</tr>\n<tr>\n<td>6</td>\n<td>epoch152-step3672-valid_loss0.4333-val_metric0.7929.ckpt</td>\n<td>Lion</td>\n<td>DenseVNet</td>\n<td>873K</td>\n<td>TS_6_6</td>\n<td>InstanceNorm3d</td>\n<td>PReLU</td>\n<td>0.76528</td>\n<td>0.75417</td>\n</tr>\n<tr>\n<td>7</td>\n<td>epoch195-step4704-valid_loss0.4258-val_metric0.7914.ckpt</td>\n<td>Lion</td>\n<td>VoxResNet</td>\n<td>7.0M</td>\n<td>TS_6_6</td>\n<td>InstanceNorm3d</td>\n<td>PReLU</td>\n<td>0.77457</td>\n<td>0.76593</td>\n</tr>\n<tr>\n<td>8</td>\n<td>epoch188-step4536-valid_loss0.4231-val_metric0.8659.ckpt</td>\n<td>Lion</td>\n<td>VoxHRNet</td>\n<td>1.4M</td>\n<td>TS_73_6</td>\n<td>InstanceNorm3d</td>\n<td>PReLU</td>\n<td>0.76738</td>\n<td>0.75995</td>\n</tr>\n<tr>\n<td>9</td>\n<td>epoch198-step4776-valid_loss0.3471-val_metric0.8730.ckpt</td>\n<td>Lion</td>\n<td>VoxHRNet</td>\n<td>1.4M</td>\n<td>TS_73_6</td>\n<td>InstanceNorm3d</td>\n<td>PReLU</td>\n<td>0.76135</td>\n<td>0.75848</td>\n</tr>\n<tr>\n<td>10</td>\n<td>epoch133-val_loss0.52-val_metric0.56-step3216.ckpt</td>\n<td>Luoziqian</td>\n<td>UNet3D</td>\n<td>1.1M</td>\n<td>TS_6_4</td>\n<td>BatchNorm3d</td>\n<td>PReLU</td>\n<td>0.76844</td>\n<td>0.76320</td>\n</tr>\n<tr>\n<td>11</td>\n<td>epoch314-val_loss0.54-val_metric0.54-step7560.ckpt</td>\n<td>Luoziqian</td>\n<td>SegResNet</td>\n<td>1.2M</td>\n<td>TS_6_4</td>\n<td>GroupNorm</td>\n<td>ReLU</td>\n<td>0.75521</td>\n<td>0.74647</td>\n</tr>\n<tr>\n<td>12</td>\n<td>epoch114-val_loss0.55-val_metric0.53-Step2760.ckpt</td>\n<td>Luoziqian</td>\n<td>UNet2E3D</td>\n<td>14.2M</td>\n<td>TS_6_4</td>\n<td>BatchNorm3d</td>\n<td>ReLU</td>\n<td>0.73758</td>\n<td>0.72966</td>\n</tr>\n</tbody>\n</table>\n<h2>Ensemble Strategy (Partial Selection)</h2>\n<p>We select the models with the best public leaderboard scores for the final submission.</p>\n<table>\n<thead>\n<tr>\n<th><strong>Models</strong></th>\n<th><strong>Ensemble Strategy</strong></th>\n<th><strong>Certainty Threshold</strong></th>\n<th><strong>Public LB</strong></th>\n<th><strong>Private LB</strong></th>\n<th><strong>Selected</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>[1, 2, 3, 4, 7, 8, 12]</td>\n<td>Average</td>\n<td>0.23</td>\n<td>0.79094</td>\n<td>0.78641</td>\n<td>❌</td>\n</tr>\n<tr>\n<td>[1, 2, 3, 4, 7, 8, 12]</td>\n<td>Average</td>\n<td>0.18</td>\n<td>0.79213</td>\n<td>0.78630</td>\n<td>❌</td>\n</tr>\n<tr>\n<td>[1, 2, 3, 4, 7, 8, 12]</td>\n<td>Average</td>\n<td>0.15</td>\n<td>0.79307</td>\n<td>0.78457</td>\n<td>❌</td>\n</tr>\n<tr>\n<td>[1, 2, 3, 4, 7, 8, 12]</td>\n<td>weighted</td>\n<td>0.15</td>\n<td>0.79247</td>\n<td>0.78417</td>\n<td>❌</td>\n</tr>\n<tr>\n<td>[1, 2, 3, 4, 6, 12]</td>\n<td>Average</td>\n<td>0.15</td>\n<td>0.78701</td>\n<td>0.78389</td>\n<td>❌</td>\n</tr>\n<tr>\n<td>[1, 2, 3, 4, 5, 6, 7]</td>\n<td>Average</td>\n<td>0.15</td>\n<td>0.79104</td>\n<td>0.78277</td>\n<td>❌</td>\n</tr>\n<tr>\n<td>[1, 2, 3, 4, 7, 9, 12]</td>\n<td>Average</td>\n<td>0.15</td>\n<td>0.79391</td>\n<td>0.78283</td>\n<td>✅</td>\n</tr>\n<tr>\n<td>[1, 2, 3, 4, 6, 7, 9, 10, 11, 12]</td>\n<td>Average</td>\n<td>0.15</td>\n<td>0.79355</td>\n<td>0.78381</td>\n<td>✅</td>\n</tr>\n</tbody>\n</table>\n<h2>Post-processing (Cluster Removal)</h2>\n<p>We used <code>CC3D</code> to convert the segmentation results into particle clusters and applied the following strategies for cluster selection:</p>\n<ul>\n<li>Threshold-based clustering (the threshold is set to 0.15 for all classses to improve recall)</li>\n<li>Cluster size filtering</li>\n</ul>\n<h2>Methods that Did Not Yield Improvements:</h2>\n<ul>\n<li>Synthetic data</li>\n<li>Second-stage classification</li>\n<li>Heavy-weight models (transformer-based models)</li>\n</ul>\n<h2>Code</h2>\n<p><a href=\"https://www.kaggle.com/code/luoziqian/7-model-6tta-t4x2-patch-norm-v2?scriptVersionId=220452693\" target=\"_blank\">Inference Notebook</a><br>\n<a href=\"https://github.com/GWwangshuo/Kaggle-2024-CZII-Pub\" target=\"_blank\">Lion's Training Code</a><br>\n<a href=\"https://github.com/luoziqianX/CZII-CryoET-Object-Identification-2nd-luoziqian\" target=\"_blank\">Luoziqian's Training Code</a></p>\n<h2>Reference</h2>\n<p>Gubins, Ilja, et al. \"SHREC 2020: Classification in cryo-electron tomograms.\" <em>Computers &amp; Graphics</em> 91 (2020): 279-289.</p>",
  "messages": [
    {
      "id": 3117241,
      "postDate": "2025-02-06T18:53:56.557Z",
      "content": "<h2>Acknowledgments</h2>\n<p>We sincerely appreciate Kaggle and the competition organizers for offering this invaluable opportunity. We also extend our gratitude to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and <a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a> for their significant contributions. <a href=\"https://www.kaggle.com/code/fnands/baseline-unet-train-submit/notebook\" target=\"_blank\"><strong>fnands's notebook</strong></a> and <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/547350#3056648\" target=\"_blank\"><strong>hengck23's discussion</strong></a> provided a solid foundation for our approach. Finally, I am grateful to my teammate <a href=\"https://www.kaggle.com/luoziqian\" target=\"_blank\">@luoziqian</a> for the excellent collaboration.</p>\n<h2>Summary</h2>\n<p>Our approach is based on an ensemble of multiple lightweight segmentation models, with parameter sizes ranging from 873K to 14.2M. Following segmentation, we computed particle centroids using CC3D and filtered small clusters based on voxel count statistics.</p>\n<h2>Overall Pipeline</h2>\n<p>The overall pipeline is illustrated below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5483160%2F1b63c4ad09d3ec66e33450cb95263adb%2Foverall_pipeline.png?generation=1739272382909281&amp;alt=media\"></p>\n<h2>Model Architecture</h2>\n<p>We initiated our experiments using the MONAI UNet baseline model. Throughout our trials, we observed that models with a large number of parameters were prone to overfitting and often performed worse than lighter models. Consequently, we opted for lightweight architectures such as UNet3D, VoxResNet, VoxHRNet, SegResNet, DynUNet, and DenseVNet.</p>\n<p>However, while experimenting with VoxResNet, VoxHRNet, and DenseVNet, we encountered stability issues, including performance fluctuations and difficulties in convergence. Upon further analysis, we identified that MONAI UNet employs <code>InstanceNorm3d</code> and <code>PReLU</code>. By modifying the normalization and activation layers accordingly, we achieved more stable model performance.</p>\n<p>The model architecture of MONAI UNet is shown below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5483160%2F35bb85a0777083a30ffc41461aa2c154%2Fmonai_unet.png?generation=1738867977514788&amp;alt=media\" alt=\"\"></p>\n<p>The following figure compares the training performance of MONAI UNet with InstanceNorm3d and PReLU versus BatchNorm3d and ReLU across 5 experiments, demonstrating that the use of InstanceNorm3d and PReLU results in more stable training.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5483160%2Fcb7065ced216887833f5f653bea7df8d%2Ftraining_stability_comparison.png?generation=1739272596177762&amp;alt=media\"></p>\n<p>For the final ensemble, we selected models based on their public leaderboard scores:</p>\n<ul>\n<li>UNet3D</li>\n<li>VoxResNet</li>\n<li>VoxHRNet</li>\n<li>SegResNet</li>\n<li>DenseVNet</li>\n<li>UNet2E3D</li>\n</ul>\n<h2>Training Strategy</h2>\n<p>Our training configuration was designed to ensure stability and optimal performance. We found that the <strong>segmentation mask radius</strong>, <strong>loss function</strong> and <strong>data augmentation strategies</strong> played a crucial role in achieving reliable results.</p>\n<h4>Segmentation Mask:</h4>\n<p>We applied ground truth masks with a customized radius for each particle.</p>\n<table>\n<thead>\n<tr>\n<th><strong>Particle Types</strong></th>\n<th><strong>Default Radius</strong></th>\n<th><strong>Luoziqian's Radius</strong></th>\n<th><strong>Lion's Radius</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>apo-ferritin</td>\n<td>60</td>\n<td>60 x 0.5</td>\n<td>80 x 0.4</td>\n</tr>\n<tr>\n<td>beta-galactosidase</td>\n<td>90</td>\n<td>90 x 0.5</td>\n<td>90 x 0.4</td>\n</tr>\n<tr>\n<td>ribosome</td>\n<td>150</td>\n<td>150 x 0.5</td>\n<td>150 x 0.4</td>\n</tr>\n<tr>\n<td>thyroglobulin</td>\n<td>130</td>\n<td>130 x 0.5</td>\n<td>120 x 0.4</td>\n</tr>\n<tr>\n<td>virus-like-particle</td>\n<td>135</td>\n<td>135 x 0.5</td>\n<td>150 x 0.4</td>\n</tr>\n</tbody>\n</table>\n<h4>Training Settings:</h4>\n<p><strong>Lion's</strong> models:</p>\n<ul>\n<li>patch size <code>[128, 256, 256]</code> or <code>[128, 384, 384]</code></li>\n<li><code>200</code> epochs</li>\n<li>7 classes</li>\n<li>Optimizer: AdamW with a learning rate of <code>0.001</code> (no learning rate scheduler)</li>\n<li>Batch size: <code>2</code> or <code>4</code></li>\n<li>Drop path rate: <code>0.1</code> or <code>0.3</code></li>\n<li>Model selection based on the F-beta metric, save top-5 best model</li>\n<li>Loss functions:<ul>\n<li>Tversky Loss</li>\n<li>Cross-Entropy Loss with weights <code>[1.0, 1.0, 0.0, 2.0, 1.0, 2.0, 1.0]</code></li></ul></li>\n</ul>\n<p><strong>Luoziqian's</strong> models</p>\n<ul>\n<li>patch size <code>[128, 200, 200]</code> or <code>[128, 256, 256]</code></li>\n<li><code>100</code> or <code>300</code> epochs</li>\n<li>6 classes</li>\n<li>Early stopping with <code>patience=20</code></li>\n<li>Batch size: <code>1</code> or <code>2</code></li>\n<li>Model selection is based on the evaluation loss. The last 100 models are saved and their F-beta scores are verified</li>\n<li>Loss functions:<ul>\n<li>Dice Loss</li>\n<li>Tversky Loss</li>\n<li>Cross-Entropy Loss, scaled to 0.05 or 0.1</li></ul></li>\n</ul>\n<h4>Augmentations:</h4>\n<p><strong>Lion's</strong> models:</p>\n<ul>\n<li>RandCropByLabelClassesd with ratios <code>[1, 1, 1, 1, 2, 1, 2]</code> for balanced sampling</li>\n<li>RandRotate90d</li>\n<li>RandFlipd</li>\n<li>RandAffined</li>\n<li>RandGaussianNoised</li>\n</ul>\n<p><strong>Luoziqian's</strong> models</p>\n<ul>\n<li>RandCropByLabelClassesd</li>\n<li>RandRotate90d</li>\n<li>RandFlipd</li>\n<li>RandShiftIntensityd</li>\n</ul>\n<h4>Model Performance with 7 TTA Summary (Partial Selection)</h4>\n<table>\n<thead>\n<tr>\n<th><strong>No</strong></th>\n<th><strong>Model</strong></th>\n<th><strong>Developer</strong></th>\n<th><strong>Architecture</strong></th>\n<th><strong>Parameters</strong></th>\n<th><strong>Valid ID</strong></th>\n<th><strong>Normalization</strong></th>\n<th><strong>Activation</strong></th>\n<th><strong>Public LB</strong></th>\n<th><strong>Private LB</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>epoch122-step2952-valid_loss0.3625-val_metric0.8367.ckpt</td>\n<td>Lion</td>\n<td>UNet3D</td>\n<td>1.1M</td>\n<td>TS_86_3</td>\n<td>InstanceNorm3d</td>\n<td>PReLU</td>\n<td>0.77379</td>\n<td>0.76582</td>\n</tr>\n<tr>\n<td>2</td>\n<td>epoch148-step3576-valid_loss1.1154-val_metric0.7722.ckpt</td>\n<td>Lion</td>\n<td>UNet3D</td>\n<td>1.1M</td>\n<td>TS_6_4</td>\n<td>InstanceNorm3d</td>\n<td>PReLU</td>\n<td>0.77021</td>\n<td>0.76725</td>\n</tr>\n<tr>\n<td>3</td>\n<td>epoch153-step3696-valid_loss0.3021-val_metric0.8900.ckpt</td>\n<td>Lion</td>\n<td>UNet3D</td>\n<td>1.6M</td>\n<td>TS_69_2</td>\n<td>InstanceNorm3d</td>\n<td>PReLU</td>\n<td>0.77205</td>\n<td>0.76676</td>\n</tr>\n<tr>\n<td>4</td>\n<td>epoch194-step4680-valid_loss1.0213-val_metric0.8788.ckpt</td>\n<td>Lion</td>\n<td>UNet3D</td>\n<td>1.1M</td>\n<td>TS_69_2</td>\n<td>InstanceNorm3d</td>\n<td>PReLU</td>\n<td>0.77390</td>\n<td>0.76737</td>\n</tr>\n<tr>\n<td>5</td>\n<td>epoch138-step3336-valid_loss0.3690-val_metric0.8476.ckpt</td>\n<td>Lion</td>\n<td>UNet3D</td>\n<td>1.1M</td>\n<td>TS_73_6</td>\n<td>InstanceNorm3d</td>\n<td>PReLU</td>\n<td>0.76543</td>\n<td>0.76025</td>\n</tr>\n<tr>\n<td>6</td>\n<td>epoch152-step3672-valid_loss0.4333-val_metric0.7929.ckpt</td>\n<td>Lion</td>\n<td>DenseVNet</td>\n<td>873K</td>\n<td>TS_6_6</td>\n<td>InstanceNorm3d</td>\n<td>PReLU</td>\n<td>0.76528</td>\n<td>0.75417</td>\n</tr>\n<tr>\n<td>7</td>\n<td>epoch195-step4704-valid_loss0.4258-val_metric0.7914.ckpt</td>\n<td>Lion</td>\n<td>VoxResNet</td>\n<td>7.0M</td>\n<td>TS_6_6</td>\n<td>InstanceNorm3d</td>\n<td>PReLU</td>\n<td>0.77457</td>\n<td>0.76593</td>\n</tr>\n<tr>\n<td>8</td>\n<td>epoch188-step4536-valid_loss0.4231-val_metric0.8659.ckpt</td>\n<td>Lion</td>\n<td>VoxHRNet</td>\n<td>1.4M</td>\n<td>TS_73_6</td>\n<td>InstanceNorm3d</td>\n<td>PReLU</td>\n<td>0.76738</td>\n<td>0.75995</td>\n</tr>\n<tr>\n<td>9</td>\n<td>epoch198-step4776-valid_loss0.3471-val_metric0.8730.ckpt</td>\n<td>Lion</td>\n<td>VoxHRNet</td>\n<td>1.4M</td>\n<td>TS_73_6</td>\n<td>InstanceNorm3d</td>\n<td>PReLU</td>\n<td>0.76135</td>\n<td>0.75848</td>\n</tr>\n<tr>\n<td>10</td>\n<td>epoch133-val_loss0.52-val_metric0.56-step3216.ckpt</td>\n<td>Luoziqian</td>\n<td>UNet3D</td>\n<td>1.1M</td>\n<td>TS_6_4</td>\n<td>BatchNorm3d</td>\n<td>PReLU</td>\n<td>0.76844</td>\n<td>0.76320</td>\n</tr>\n<tr>\n<td>11</td>\n<td>epoch314-val_loss0.54-val_metric0.54-step7560.ckpt</td>\n<td>Luoziqian</td>\n<td>SegResNet</td>\n<td>1.2M</td>\n<td>TS_6_4</td>\n<td>GroupNorm</td>\n<td>ReLU</td>\n<td>0.75521</td>\n<td>0.74647</td>\n</tr>\n<tr>\n<td>12</td>\n<td>epoch114-val_loss0.55-val_metric0.53-Step2760.ckpt</td>\n<td>Luoziqian</td>\n<td>UNet2E3D</td>\n<td>14.2M</td>\n<td>TS_6_4</td>\n<td>BatchNorm3d</td>\n<td>ReLU</td>\n<td>0.73758</td>\n<td>0.72966</td>\n</tr>\n</tbody>\n</table>\n<h2>Ensemble Strategy (Partial Selection)</h2>\n<p>We select the models with the best public leaderboard scores for the final submission.</p>\n<table>\n<thead>\n<tr>\n<th><strong>Models</strong></th>\n<th><strong>Ensemble Strategy</strong></th>\n<th><strong>Certainty Threshold</strong></th>\n<th><strong>Public LB</strong></th>\n<th><strong>Private LB</strong></th>\n<th><strong>Selected</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>[1, 2, 3, 4, 7, 8, 12]</td>\n<td>Average</td>\n<td>0.23</td>\n<td>0.79094</td>\n<td>0.78641</td>\n<td>❌</td>\n</tr>\n<tr>\n<td>[1, 2, 3, 4, 7, 8, 12]</td>\n<td>Average</td>\n<td>0.18</td>\n<td>0.79213</td>\n<td>0.78630</td>\n<td>❌</td>\n</tr>\n<tr>\n<td>[1, 2, 3, 4, 7, 8, 12]</td>\n<td>Average</td>\n<td>0.15</td>\n<td>0.79307</td>\n<td>0.78457</td>\n<td>❌</td>\n</tr>\n<tr>\n<td>[1, 2, 3, 4, 7, 8, 12]</td>\n<td>weighted</td>\n<td>0.15</td>\n<td>0.79247</td>\n<td>0.78417</td>\n<td>❌</td>\n</tr>\n<tr>\n<td>[1, 2, 3, 4, 6, 12]</td>\n<td>Average</td>\n<td>0.15</td>\n<td>0.78701</td>\n<td>0.78389</td>\n<td>❌</td>\n</tr>\n<tr>\n<td>[1, 2, 3, 4, 5, 6, 7]</td>\n<td>Average</td>\n<td>0.15</td>\n<td>0.79104</td>\n<td>0.78277</td>\n<td>❌</td>\n</tr>\n<tr>\n<td>[1, 2, 3, 4, 7, 9, 12]</td>\n<td>Average</td>\n<td>0.15</td>\n<td>0.79391</td>\n<td>0.78283</td>\n<td>✅</td>\n</tr>\n<tr>\n<td>[1, 2, 3, 4, 6, 7, 9, 10, 11, 12]</td>\n<td>Average</td>\n<td>0.15</td>\n<td>0.79355</td>\n<td>0.78381</td>\n<td>✅</td>\n</tr>\n</tbody>\n</table>\n<h2>Post-processing (Cluster Removal)</h2>\n<p>We used <code>CC3D</code> to convert the segmentation results into particle clusters and applied the following strategies for cluster selection:</p>\n<ul>\n<li>Threshold-based clustering (the threshold is set to 0.15 for all classses to improve recall)</li>\n<li>Cluster size filtering</li>\n</ul>\n<h2>Methods that Did Not Yield Improvements:</h2>\n<ul>\n<li>Synthetic data</li>\n<li>Second-stage classification</li>\n<li>Heavy-weight models (transformer-based models)</li>\n</ul>\n<h2>Code</h2>\n<p><a href=\"https://www.kaggle.com/code/luoziqian/7-model-6tta-t4x2-patch-norm-v2?scriptVersionId=220452693\" target=\"_blank\">Inference Notebook</a><br>\n<a href=\"https://github.com/GWwangshuo/Kaggle-2024-CZII-Pub\" target=\"_blank\">Lion's Training Code</a><br>\n<a href=\"https://github.com/luoziqianX/CZII-CryoET-Object-Identification-2nd-luoziqian\" target=\"_blank\">Luoziqian's Training Code</a></p>\n<h2>Reference</h2>\n<p>Gubins, Ilja, et al. \"SHREC 2020: Classification in cryo-electron tomograms.\" <em>Computers &amp; Graphics</em> 91 (2020): 279-289.</p>",
      "rawMarkdown": "## Acknowledgments\n\nWe sincerely appreciate Kaggle and the competition organizers for offering this invaluable opportunity. We also extend our gratitude to @hengck23 and @fnands for their significant contributions. [**fnands's notebook**](https://www.kaggle.com/code/fnands/baseline-unet-train-submit/notebook) and [**hengck23's discussion**](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/547350#3056648) provided a solid foundation for our approach. Finally, I am grateful to my teammate @luoziqian for the excellent collaboration.\n\n## Summary\n\nOur approach is based on an ensemble of multiple lightweight segmentation models, with parameter sizes ranging from 873K to 14.2M. Following segmentation, we computed particle centroids using CC3D and filtered small clusters based on voxel count statistics.\n\n## Overall Pipeline\n\nThe overall pipeline is illustrated below:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5483160%2F1b63c4ad09d3ec66e33450cb95263adb%2Foverall_pipeline.png?generation=1739272382909281&alt=media\" width=\"700\" height=\"auto\">\n\n\n## Model Architecture\n\nWe initiated our experiments using the MONAI UNet baseline model. Throughout our trials, we observed that models with a large number of parameters were prone to overfitting and often performed worse than lighter models. Consequently, we opted for lightweight architectures such as UNet3D, VoxResNet, VoxHRNet, SegResNet, DynUNet, and DenseVNet.\n\nHowever, while experimenting with VoxResNet, VoxHRNet, and DenseVNet, we encountered stability issues, including performance fluctuations and difficulties in convergence. Upon further analysis, we identified that MONAI UNet employs `InstanceNorm3d` and `PReLU`. By modifying the normalization and activation layers accordingly, we achieved more stable model performance.\n\nThe model architecture of MONAI UNet is shown below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5483160%2F35bb85a0777083a30ffc41461aa2c154%2Fmonai_unet.png?generation=1738867977514788&alt=media)\n\nThe following figure compares the training performance of MONAI UNet with InstanceNorm3d and PReLU versus BatchNorm3d and ReLU across 5 experiments, demonstrating that the use of InstanceNorm3d and PReLU results in more stable training.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5483160%2Fcb7065ced216887833f5f653bea7df8d%2Ftraining_stability_comparison.png?generation=1739272596177762&alt=media\" width=\"700\" height=\"auto\">\n\nFor the final ensemble, we selected models based on their public leaderboard scores:\n\n- UNet3D\n- VoxResNet\n- VoxHRNet\n- SegResNet\n- DenseVNet\n- UNet2E3D\n\n## Training Strategy\n\nOur training configuration was designed to ensure stability and optimal performance. We found that the **segmentation mask radius**, **loss function** and **data augmentation strategies** played a crucial role in achieving reliable results.\n\n#### Segmentation Mask:\n\nWe applied ground truth masks with a customized radius for each particle.\n\n| **Particle Types**  | **Default Radius** |**Luoziqian's Radius** | **Lion's Radius** |\n| -------------------   | ------------------    |  -----------------------  |   ----------------   |\n| apo-ferritin               | 60                             | 60 x 0.5                          |  80 x 0.4                | \n| beta-galactosidase  | 90                             | 90 x 0.5                          | 90 x 0.4                 |\n| ribosome                  | 150                            | 150 x 0.5                         | 150 x 0.4                |\n| thyroglobulin            | 130                            | 130 x 0.5                         |120 x 0.4                |\n| virus-like-particle    | 135                            | 135 x 0.5                         | 150 x 0.4               |     \n\n#### Training Settings:\n\n**Lion's** models:\n- patch size `[128, 256, 256]` or `[128, 384, 384]`\n- `200` epochs\n- 7 classes\n- Optimizer: AdamW with a learning rate of `0.001` (no learning rate scheduler)\n- Batch size: `2` or `4`\n- Drop path rate: `0.1` or `0.3`\n- Model selection based on the F-beta metric, save top-5 best model\n- Loss functions:\n  - Tversky Loss\n  - Cross-Entropy Loss with weights `[1.0, 1.0, 0.0, 2.0, 1.0, 2.0, 1.0]`\n\n**Luoziqian's** models\n- patch size `[128, 200, 200]` or `[128, 256, 256]`\n- `100` or `300` epochs\n- 6 classes\n- Early stopping with `patience=20`\n- Batch size: `1` or `2`\n- Model selection is based on the evaluation loss. The last 100 models are saved and their F-beta scores are verified\n- Loss functions:\n  - Dice Loss\n  - Tversky Loss\n  - Cross-Entropy Loss, scaled to 0.05 or 0.1\n\n\n#### Augmentations:\n\n**Lion's** models:\n\n- RandCropByLabelClassesd with ratios `[1, 1, 1, 1, 2, 1, 2]` for balanced sampling\n- RandRotate90d\n- RandFlipd\n- RandAffined\n- RandGaussianNoised\n\n**Luoziqian's** models\n- RandCropByLabelClassesd\n- RandRotate90d\n- RandFlipd\n- RandShiftIntensityd\n\n\n#### Model Performance with 7 TTA Summary (Partial Selection)\n\n| **No** | **Model**                                                |  **Developer**  | **Architecture** | **Parameters** |    **Valid ID**   | **Normalization** | **Activation** | **Public LB** | **Private LB** |\n| ------ | -------------------------------------------------------- | ----------------| ---------------- | -------------- | ----------------- | ----------------- | -------------- | ------------- | -------------- |\n| 1      | epoch122-step2952-valid_loss0.3625-val_metric0.8367.ckpt | Lion            | UNet3D           | 1.1M           | TS_86_3           | InstanceNorm3d    | PReLU          | 0.77379       | 0.76582        |\n| 2      | epoch148-step3576-valid_loss1.1154-val_metric0.7722.ckpt | Lion            | UNet3D           | 1.1M           | TS_6_4            | InstanceNorm3d    | PReLU          | 0.77021       | 0.76725        |\n| 3      | epoch153-step3696-valid_loss0.3021-val_metric0.8900.ckpt | Lion            | UNet3D           | 1.6M           | TS_69_2           | InstanceNorm3d    | PReLU          | 0.77205       | 0.76676        |\n| 4      | epoch194-step4680-valid_loss1.0213-val_metric0.8788.ckpt | Lion            | UNet3D           | 1.1M           | TS_69_2           | InstanceNorm3d    | PReLU          | 0.77390       | 0.76737        |\n| 5      | epoch138-step3336-valid_loss0.3690-val_metric0.8476.ckpt | Lion            | UNet3D           | 1.1M           | TS_73_6           | InstanceNorm3d    | PReLU          | 0.76543       | 0.76025        |\n| 6      | epoch152-step3672-valid_loss0.4333-val_metric0.7929.ckpt | Lion            | DenseVNet        | 873K           | TS_6_6            | InstanceNorm3d    | PReLU          | 0.76528       | 0.75417        |\n| 7      | epoch195-step4704-valid_loss0.4258-val_metric0.7914.ckpt | Lion            | VoxResNet        | 7.0M           | TS_6_6            | InstanceNorm3d    | PReLU          | 0.77457       | 0.76593        |\n| 8      | epoch188-step4536-valid_loss0.4231-val_metric0.8659.ckpt | Lion            | VoxHRNet         | 1.4M           | TS_73_6           | InstanceNorm3d    | PReLU          | 0.76738       | 0.75995        |\n| 9      | epoch198-step4776-valid_loss0.3471-val_metric0.8730.ckpt | Lion            | VoxHRNet         | 1.4M           | TS_73_6           | InstanceNorm3d    | PReLU          | 0.76135       | 0.75848        |\n| 10     | epoch133-val_loss0.52-val_metric0.56-step3216.ckpt       | Luoziqian       | UNet3D           | 1.1M           | TS_6_4            | BatchNorm3d       | PReLU          | 0.76844       | 0.76320        |\n| 11     | epoch314-val_loss0.54-val_metric0.54-step7560.ckpt       | Luoziqian       | SegResNet        | 1.2M           | TS_6_4            | GroupNorm         | ReLU           | 0.75521       | 0.74647        |\n| 12     | epoch114-val_loss0.55-val_metric0.53-Step2760.ckpt       | Luoziqian       | UNet2E3D         | 14.2M          | TS_6_4            | BatchNorm3d       | ReLU           | 0.73758       | 0.72966        |\n\n\n## Ensemble Strategy (Partial Selection)\nWe select the models with the best public leaderboard scores for the final submission.\n\n| **Models**                        | **Ensemble Strategy** | **Certainty Threshold** | **Public LB** | **Private LB** |    **Selected**   |\n| --------------------------------- | --------------------- |  ---------------------  | ------------- | -------------- | ----------------- |\n| [1, 2, 3, 4, 7, 8, 12]            | Average               |          0.23           | 0.79094       | 0.78641        |        ❌         |\n| [1, 2, 3, 4, 7, 8, 12]            | Average               |          0.18           | 0.79213       | 0.78630        |        ❌         |\n| [1, 2, 3, 4, 7, 8, 12]            | Average               |          0.15           | 0.79307       | 0.78457        |        ❌         |\n| [1, 2, 3, 4, 7, 8, 12]            | weighted              |          0.15           | 0.79247       | 0.78417        |        ❌         |\n| [1, 2, 3, 4, 6, 12]               | Average               |          0.15           | 0.78701       | 0.78389        |        ❌         |\n| [1, 2, 3, 4, 5, 6, 7]             | Average               |          0.15           | 0.79104       | 0.78277        |        ❌         |\n| [1, 2, 3, 4, 7, 9, 12]            | Average               |          0.15           | 0.79391       | 0.78283        |        ✅         |\n| [1, 2, 3, 4, 6, 7, 9, 10, 11, 12] | Average               |          0.15           | 0.79355       | 0.78381        |        ✅         |\n\n## Post-processing (Cluster Removal)\n\nWe used `CC3D` to convert the segmentation results into particle clusters and applied the following strategies for cluster selection:\n\n- Threshold-based clustering (the threshold is set to 0.15 for all classses to improve recall)\n- Cluster size filtering\n\n## Methods that Did Not Yield Improvements:\n\n- Synthetic data\n- Second-stage classification\n- Heavy-weight models (transformer-based models)\n\n## Code\n\n[Inference Notebook](https://www.kaggle.com/code/luoziqian/7-model-6tta-t4x2-patch-norm-v2?scriptVersionId=220452693)\n[Lion's Training Code](https://github.com/GWwangshuo/Kaggle-2024-CZII-Pub)\n[Luoziqian's Training Code](https://github.com/luoziqianX/CZII-CryoET-Object-Identification-2nd-luoziqian)\n\n## Reference\n\nGubins, Ilja, et al. \"SHREC 2020: Classification in cryo-electron tomograms.\" *Computers & Graphics* 91 (2020): 279-289.",
      "votes": 44
    },
    {
      "id": 3117309,
      "postDate": "2025-02-06T20:36:24.863Z",
      "content": "<p>Wow, you have very strong single models. Focus on stabilizing the training might be the trick(s) we missed. Congratulations on the gold and the prize!</p>",
      "rawMarkdown": "Wow, you have very strong single models. Focus on stabilizing the training might be the trick(s) we missed. Congratulations on the gold and the prize!",
      "votes": 2
    },
    {
      "id": 3120885,
      "postDate": "2025-02-11T03:15:04.657Z",
      "content": "<p>Congratulations !<br>\nI've realized that there are a lot of room for improvement even if segmentation + cc3d approach and lack of my skills to design experiments.<br>\nIn Luoziqian's Training Code, I cannot find output comp metrics in training code. Did you choose best model by other way ?</p>",
      "rawMarkdown": "Congratulations !\nI've realized that there are a lot of room for improvement even if segmentation + cc3d approach and lack of my skills to design experiments.\nIn Luoziqian's Training Code, I cannot find output comp metrics in training code. Did you choose best model by other way ?",
      "replies": [
        {
          "id": 3120907,
          "postDate": "2025-02-11T03:47:15.863Z",
          "content": "<p>You can change the path_dir in <a href=\"https://github.com/luoziqianX/CZII-CryoET-Object-Identification-2nd-luoziqian/blob/main/compute-cv-7TTA.ipynb\" target=\"_blank\">compute-cv-7TTA.ipynb</a> and run it to calculate the F-beta4 score of ckpts after training. I think the Tversky Loss for each class is enough for me to initially rule out less successful experiments. Since measuring F-beta4 during training would significantly increase the experiment time and my limited computing resources, I did not calculate it during training.</p>",
          "rawMarkdown": "You can change the path_dir in [compute-cv-7TTA.ipynb](https://github.com/luoziqianX/CZII-CryoET-Object-Identification-2nd-luoziqian/blob/main/compute-cv-7TTA.ipynb) and run it to calculate the F-beta4 score of ckpts after training. I think the Tversky Loss for each class is enough for me to initially rule out less successful experiments. Since measuring F-beta4 during training would significantly increase the experiment time and my limited computing resources, I did not calculate it during training.",
          "votes": 4
        }
      ]
    },
    {
      "id": 3119199,
      "postDate": "2025-02-09T02:21:27.833Z",
      "content": "<p>Congratulations! <a href=\"https://www.kaggle.com/sjtuwangshuo\" target=\"_blank\">@sjtuwangshuo</a> .</p>\n<p>Which sample do you use to validate? </p>",
      "rawMarkdown": "Congratulations! @sjtuwangshuo .\n\nWhich sample do you use to validate? ",
      "replies": [
        {
          "id": 3119239,
          "postDate": "2025-02-09T04:05:20.560Z",
          "content": "<p>Initial Strategy:</p>\n<ul>\n<li>Split tomograms into seven folds (e.g., TS_5_4, TS_6_4, TS_69_2, etc.).</li>\n<li>Train each fold and save the models with the top five highest F-beta scores.</li>\n<li>Evaluate each of the top five models using 7× Test-Time Augmentation (TTA) on the same validation sample and compute the F-beta score.</li>\n<li>Select the model with the highest F-beta score after 7× TTA for submission.</li>\n</ul>\n<p>Refinement:<br>\nSubsequently, <code>Luoziqian</code> observed that models validated on ['TS_5_4', 'TS_6_4', 'TS_69_2'] tend to achieve better public leaderboard performance, though not consistently. Based on this observation, I evaluate model performance using 7× TTA on these specific folds and submit the model with the highest F-beta score.</p>\n<p>For details on the validation samples of the selected models used in the final ensemble, please refer to the <code>\"Model Performance with 7× TTA Summary (Partial Selection)\"</code> section of this post.</p>",
          "rawMarkdown": "Initial Strategy:\n\n- Split tomograms into seven folds (e.g., TS_5_4, TS_6_4, TS_69_2, etc.).\n- Train each fold and save the models with the top five highest F-beta scores.\n- Evaluate each of the top five models using 7× Test-Time Augmentation (TTA) on the same validation sample and compute the F-beta score.\n- Select the model with the highest F-beta score after 7× TTA for submission.\n\nRefinement:\nSubsequently, `Luoziqian` observed that models validated on ['TS_5_4', 'TS_6_4', 'TS_69_2'] tend to achieve better public leaderboard performance, though not consistently. Based on this observation, I evaluate model performance using 7× TTA on these specific folds and submit the model with the highest F-beta score.\n\nFor details on the validation samples of the selected models used in the final ensemble, please refer to the `\"Model Performance with 7× TTA Summary (Partial Selection)\"` section of this post."
        }
      ]
    },
    {
      "id": 3118525,
      "postDate": "2025-02-08T06:52:59.977Z",
      "content": "<p>Congratulations! thank you for sharing your solution</p>\n<p>I have a question :why  experiment no.2 and no.10/11/12 have such a huge different valid score(0.7x vs 0.5x,but lb very close)？</p>",
      "rawMarkdown": "Congratulations! thank you for sharing your solution\n\nI have a question :why  experiment no.2 and no.10/11/12 have such a huge different valid score(0.7x vs 0.5x,but lb very close)？\n\n",
      "replies": [
        {
          "id": 3118546,
          "postDate": "2025-02-08T07:30:02.710Z",
          "content": "<p>Could you please clarify what is meant by “0.7x vs. 0.5x, but lb very close”? Please note that the experiments presented thus far represent only a subset of the complete experimental results. Additionally, my colleague, Luoziqian, has developed several models based on the MONAI UNet architecture that achieve a public leaderboard score exceeding 0.760; however, these models do not exhibit improved performance when ensembled.</p>",
          "rawMarkdown": "Could you please clarify what is meant by “0.7x vs. 0.5x, but lb very close”? Please note that the experiments presented thus far represent only a subset of the complete experimental results. Additionally, my colleague, Luoziqian, has developed several models based on the MONAI UNet architecture that achieve a public leaderboard score exceeding 0.760; however, these models do not exhibit improved performance when ensembled.",
          "replies": [
            {
              "id": 3118558,
              "postDate": "2025-02-08T07:42:32.960Z",
              "content": "<p>噢是这样的，我看你实验的model列里面有val_metric这一项，no.2和no.10,no.11,no.12都是用TS_6_4作为验证集，好像这几个的验证集指标差异挺大的，想问问看是不是有什么特殊处理我错过了</p>",
              "rawMarkdown": "噢是这样的，我看你实验的model列里面有val_metric这一项，no.2和no.10,no.11,no.12都是用TS_6_4作为验证集，好像这几个的验证集指标差异挺大的，想问问看是不是有什么特殊处理我错过了"
            },
            {
              "id": 3118572,
              "postDate": "2025-02-08T08:01:55.547Z",
              "content": "<p>Our experiments indicated that several factors significantly affect performance, including spatial size, mask radius, data augmentation strategies, loss functions, normalization layers, and activation layers.</p>\n<p>In the comparison between Experiment <code>No.2</code> and Experiment <code>No.10</code>, the following differences were observed matter:</p>\n<ul>\n<li>Patch spatial size</li>\n<li>Mask radius</li>\n<li>Data augmentation strategies</li>\n<li>Loss functions</li>\n<li>Number of classes</li>\n</ul>\n<p>Furthermore, when comparing Experiment <code>No.2</code> with Experiments <code>No.11</code> and <code>No.12</code>, we found that the MONAI UNet outperformed both SegResNet and UNet2E3D. This improved performance can be attributed to several architectural features of the MONAI UNet:</p>\n<ul>\n<li>The use of <code>InstanceNorm3d</code> in conjunction with <code>PReLU activation</code>.</li>\n<li>An encoder downsampling factor of <code>1/4</code> as opposed to <code>1/8</code>, <code>1/16</code>, or <code>1/32</code>.</li>\n<li>The implementation of a deconvolution layer using <code>ConvTransposed3D</code> instead of conventional upsampling.</li>\n<li>The incorporation of <code>residual blocks</code>, which enhances parameter efficiency and reduce t the likelihood of overfitting.</li>\n</ul>\n<p>For both SegResNet and UNet2E3D, modifications were implemented based on our experimental investigations to facilitate training convergence. Specifically, we:</p>\n<ul>\n<li>Revised the normalization and activation layers.</li>\n<li>Replaced the decoder with a deconvolution layer.</li>\n<li>Reduced the number of decoder channels to create a more lightweight model.</li>\n</ul>\n<p>Despite these modifications, neither SegResNet nor UNet2E3D outperformed the MONAI UNet.</p>",
              "rawMarkdown": "Our experiments indicated that several factors significantly affect performance, including spatial size, mask radius, data augmentation strategies, loss functions, normalization layers, and activation layers.\n\nIn the comparison between Experiment `No.2` and Experiment `No.10`, the following differences were observed matter:\n- Patch spatial size\n- Mask radius\n- Data augmentation strategies\n- Loss functions\n- Number of classes\n\nFurthermore, when comparing Experiment `No.2` with Experiments `No.11` and `No.12`, we found that the MONAI UNet outperformed both SegResNet and UNet2E3D. This improved performance can be attributed to several architectural features of the MONAI UNet:\n\n- The use of `InstanceNorm3d` in conjunction with `PReLU activation`.\n- An encoder downsampling factor of `1/4` as opposed to `1/8`, `1/16`, or `1/32`.\n- The implementation of a deconvolution layer using `ConvTransposed3D` instead of conventional upsampling.\n- The incorporation of `residual blocks`, which enhances parameter efficiency and reduce t the likelihood of overfitting.\n\nFor both SegResNet and UNet2E3D, modifications were implemented based on our experimental investigations to facilitate training convergence. Specifically, we:\n- Revised the normalization and activation layers.\n- Replaced the decoder with a deconvolution layer.\n- Reduced the number of decoder channels to create a more lightweight model.\n\nDespite these modifications, neither SegResNet nor UNet2E3D outperformed the MONAI UNet.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3117336,
      "postDate": "2025-02-06T21:33:08.730Z",
      "content": "<p>Congrats! Glad the notebook helped you get started!</p>",
      "rawMarkdown": "Congrats! Glad the notebook helped you get started!"
    },
    {
      "id": 3118220,
      "postDate": "2025-02-07T18:57:19.297Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3117309,
      "author_name": "Sinan Calisir",
      "author_url": "",
      "post_date": "2025-02-06T20:36:24.863000",
      "content": "<p>Wow, you have very strong single models. Focus on stabilizing the training might be the trick(s) we missed. Congratulations on the gold and the prize!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3120885,
      "author_name": "Aurora_blue",
      "author_url": "",
      "post_date": "2025-02-11T03:15:04.657000",
      "content": "<p>Congratulations !<br>\nI've realized that there are a lot of room for improvement even if segmentation + cc3d approach and lack of my skills to design experiments.<br>\nIn Luoziqian's Training Code, I cannot find output comp metrics in training code. Did you choose best model by other way ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3120907,
          "author_name": "LuoZiqian",
          "author_url": "",
          "post_date": "2025-02-11T03:47:15.863000",
          "content": "<p>You can change the path_dir in <a href=\"https://github.com/luoziqianX/CZII-CryoET-Object-Identification-2nd-luoziqian/blob/main/compute-cv-7TTA.ipynb\" target=\"_blank\">compute-cv-7TTA.ipynb</a> and run it to calculate the F-beta4 score of ckpts after training. I think the Tversky Loss for each class is enough for me to initially rule out less successful experiments. Since measuring F-beta4 during training would significantly increase the experiment time and my limited computing resources, I did not calculate it during training.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 3119199,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2025-02-09T02:21:27.833000",
      "content": "<p>Congratulations! <a href=\"https://www.kaggle.com/sjtuwangshuo\" target=\"_blank\">@sjtuwangshuo</a> .</p>\n<p>Which sample do you use to validate? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3119239,
          "author_name": "AnnieGo",
          "author_url": "",
          "post_date": "2025-02-09T04:05:20.560000",
          "content": "<p>Initial Strategy:</p>\n<ul>\n<li>Split tomograms into seven folds (e.g., TS_5_4, TS_6_4, TS_69_2, etc.).</li>\n<li>Train each fold and save the models with the top five highest F-beta scores.</li>\n<li>Evaluate each of the top five models using 7× Test-Time Augmentation (TTA) on the same validation sample and compute the F-beta score.</li>\n<li>Select the model with the highest F-beta score after 7× TTA for submission.</li>\n</ul>\n<p>Refinement:<br>\nSubsequently, <code>Luoziqian</code> observed that models validated on ['TS_5_4', 'TS_6_4', 'TS_69_2'] tend to achieve better public leaderboard performance, though not consistently. Based on this observation, I evaluate model performance using 7× TTA on these specific folds and submit the model with the highest F-beta score.</p>\n<p>For details on the validation samples of the selected models used in the final ensemble, please refer to the <code>\"Model Performance with 7× TTA Summary (Partial Selection)\"</code> section of this post.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3118525,
      "author_name": "FatLiuyun",
      "author_url": "",
      "post_date": "2025-02-08T06:52:59.977000",
      "content": "<p>Congratulations! thank you for sharing your solution</p>\n<p>I have a question :why  experiment no.2 and no.10/11/12 have such a huge different valid score(0.7x vs 0.5x,but lb very close)？</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3118546,
          "author_name": "AnnieGo",
          "author_url": "",
          "post_date": "2025-02-08T07:30:02.710000",
          "content": "<p>Could you please clarify what is meant by “0.7x vs. 0.5x, but lb very close”? Please note that the experiments presented thus far represent only a subset of the complete experimental results. Additionally, my colleague, Luoziqian, has developed several models based on the MONAI UNet architecture that achieve a public leaderboard score exceeding 0.760; however, these models do not exhibit improved performance when ensembled.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3118558,
              "author_name": "FatLiuyun",
              "author_url": "",
              "post_date": "2025-02-08T07:42:32.960000",
              "content": "<p>噢是这样的，我看你实验的model列里面有val_metric这一项，no.2和no.10,no.11,no.12都是用TS_6_4作为验证集，好像这几个的验证集指标差异挺大的，想问问看是不是有什么特殊处理我错过了</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3118572,
              "author_name": "AnnieGo",
              "author_url": "",
              "post_date": "2025-02-08T08:01:55.547000",
              "content": "<p>Our experiments indicated that several factors significantly affect performance, including spatial size, mask radius, data augmentation strategies, loss functions, normalization layers, and activation layers.</p>\n<p>In the comparison between Experiment <code>No.2</code> and Experiment <code>No.10</code>, the following differences were observed matter:</p>\n<ul>\n<li>Patch spatial size</li>\n<li>Mask radius</li>\n<li>Data augmentation strategies</li>\n<li>Loss functions</li>\n<li>Number of classes</li>\n</ul>\n<p>Furthermore, when comparing Experiment <code>No.2</code> with Experiments <code>No.11</code> and <code>No.12</code>, we found that the MONAI UNet outperformed both SegResNet and UNet2E3D. This improved performance can be attributed to several architectural features of the MONAI UNet:</p>\n<ul>\n<li>The use of <code>InstanceNorm3d</code> in conjunction with <code>PReLU activation</code>.</li>\n<li>An encoder downsampling factor of <code>1/4</code> as opposed to <code>1/8</code>, <code>1/16</code>, or <code>1/32</code>.</li>\n<li>The implementation of a deconvolution layer using <code>ConvTransposed3D</code> instead of conventional upsampling.</li>\n<li>The incorporation of <code>residual blocks</code>, which enhances parameter efficiency and reduce t the likelihood of overfitting.</li>\n</ul>\n<p>For both SegResNet and UNet2E3D, modifications were implemented based on our experimental investigations to facilitate training convergence. Specifically, we:</p>\n<ul>\n<li>Revised the normalization and activation layers.</li>\n<li>Replaced the decoder with a deconvolution layer.</li>\n<li>Reduced the number of decoder channels to create a more lightweight model.</li>\n</ul>\n<p>Despite these modifications, neither SegResNet nor UNet2E3D outperformed the MONAI UNet.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3117336,
      "author_name": "fnands",
      "author_url": "",
      "post_date": "2025-02-06T21:33:08.730000",
      "content": "<p>Congrats! Glad the notebook helped you get started!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3118220,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-02-07T18:57:19.297000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3117241": "## Acknowledgments\n\nWe sincerely appreciate Kaggle and the competition organizers for offering this invaluable opportunity. We also extend our gratitude to @hengck23 and @fnands for their significant contributions. [**fnands's notebook**](https://www.kaggle.com/code/fnands/baseline-unet-train-submit/notebook) and [**hengck23's discussion**](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/547350#3056648) provided a solid foundation for our approach. Finally, I am grateful to my teammate @luoziqian for the excellent collaboration.\n\n## Summary\n\nOur approach is based on an ensemble of multiple lightweight segmentation models, with parameter sizes ranging from 873K to 14.2M. Following segmentation, we computed particle centroids using CC3D and filtered small clusters based on voxel count statistics.\n\n## Overall Pipeline\n\nThe overall pipeline is illustrated below:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5483160%2F1b63c4ad09d3ec66e33450cb95263adb%2Foverall_pipeline.png?generation=1739272382909281&alt=media\" width=\"700\" height=\"auto\">\n\n\n## Model Architecture\n\nWe initiated our experiments using the MONAI UNet baseline model. Throughout our trials, we observed that models with a large number of parameters were prone to overfitting and often performed worse than lighter models. Consequently, we opted for lightweight architectures such as UNet3D, VoxResNet, VoxHRNet, SegResNet, DynUNet, and DenseVNet.\n\nHowever, while experimenting with VoxResNet, VoxHRNet, and DenseVNet, we encountered stability issues, including performance fluctuations and difficulties in convergence. Upon further analysis, we identified that MONAI UNet employs `InstanceNorm3d` and `PReLU`. By modifying the normalization and activation layers accordingly, we achieved more stable model performance.\n\nThe model architecture of MONAI UNet is shown below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5483160%2F35bb85a0777083a30ffc41461aa2c154%2Fmonai_unet.png?generation=1738867977514788&alt=media)\n\nThe following figure compares the training performance of MONAI UNet with InstanceNorm3d and PReLU versus BatchNorm3d and ReLU across 5 experiments, demonstrating that the use of InstanceNorm3d and PReLU results in more stable training.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5483160%2Fcb7065ced216887833f5f653bea7df8d%2Ftraining_stability_comparison.png?generation=1739272596177762&alt=media\" width=\"700\" height=\"auto\">\n\nFor the final ensemble, we selected models based on their public leaderboard scores:\n\n- UNet3D\n- VoxResNet\n- VoxHRNet\n- SegResNet\n- DenseVNet\n- UNet2E3D\n\n## Training Strategy\n\nOur training configuration was designed to ensure stability and optimal performance. We found that the **segmentation mask radius**, **loss function** and **data augmentation strategies** played a crucial role in achieving reliable results.\n\n#### Segmentation Mask:\n\nWe applied ground truth masks with a customized radius for each particle.\n\n| **Particle Types**  | **Default Radius** |**Luoziqian's Radius** | **Lion's Radius** |\n| -------------------   | ------------------    |  -----------------------  |   ----------------   |\n| apo-ferritin               | 60                             | 60 x 0.5                          |  80 x 0.4                | \n| beta-galactosidase  | 90                             | 90 x 0.5                          | 90 x 0.4                 |\n| ribosome                  | 150                            | 150 x 0.5                         | 150 x 0.4                |\n| thyroglobulin            | 130                            | 130 x 0.5                         |120 x 0.4                |\n| virus-like-particle    | 135                            | 135 x 0.5                         | 150 x 0.4               |     \n\n#### Training Settings:\n\n**Lion's** models:\n- patch size `[128, 256, 256]` or `[128, 384, 384]`\n- `200` epochs\n- 7 classes\n- Optimizer: AdamW with a learning rate of `0.001` (no learning rate scheduler)\n- Batch size: `2` or `4`\n- Drop path rate: `0.1` or `0.3`\n- Model selection based on the F-beta metric, save top-5 best model\n- Loss functions:\n  - Tversky Loss\n  - Cross-Entropy Loss with weights `[1.0, 1.0, 0.0, 2.0, 1.0, 2.0, 1.0]`\n\n**Luoziqian's** models\n- patch size `[128, 200, 200]` or `[128, 256, 256]`\n- `100` or `300` epochs\n- 6 classes\n- Early stopping with `patience=20`\n- Batch size: `1` or `2`\n- Model selection is based on the evaluation loss. The last 100 models are saved and their F-beta scores are verified\n- Loss functions:\n  - Dice Loss\n  - Tversky Loss\n  - Cross-Entropy Loss, scaled to 0.05 or 0.1\n\n\n#### Augmentations:\n\n**Lion's** models:\n\n- RandCropByLabelClassesd with ratios `[1, 1, 1, 1, 2, 1, 2]` for balanced sampling\n- RandRotate90d\n- RandFlipd\n- RandAffined\n- RandGaussianNoised\n\n**Luoziqian's** models\n- RandCropByLabelClassesd\n- RandRotate90d\n- RandFlipd\n- RandShiftIntensityd\n\n\n#### Model Performance with 7 TTA Summary (Partial Selection)\n\n| **No** | **Model**                                                |  **Developer**  | **Architecture** | **Parameters** |    **Valid ID**   | **Normalization** | **Activation** | **Public LB** | **Private LB** |\n| ------ | -------------------------------------------------------- | ----------------| ---------------- | -------------- | ----------------- | ----------------- | -------------- | ------------- | -------------- |\n| 1      | epoch122-step2952-valid_loss0.3625-val_metric0.8367.ckpt | Lion            | UNet3D           | 1.1M           | TS_86_3           | InstanceNorm3d    | PReLU          | 0.77379       | 0.76582        |\n| 2      | epoch148-step3576-valid_loss1.1154-val_metric0.7722.ckpt | Lion            | UNet3D           | 1.1M           | TS_6_4            | InstanceNorm3d    | PReLU          | 0.77021       | 0.76725        |\n| 3      | epoch153-step3696-valid_loss0.3021-val_metric0.8900.ckpt | Lion            | UNet3D           | 1.6M           | TS_69_2           | InstanceNorm3d    | PReLU          | 0.77205       | 0.76676        |\n| 4      | epoch194-step4680-valid_loss1.0213-val_metric0.8788.ckpt | Lion            | UNet3D           | 1.1M           | TS_69_2           | InstanceNorm3d    | PReLU          | 0.77390       | 0.76737        |\n| 5      | epoch138-step3336-valid_loss0.3690-val_metric0.8476.ckpt | Lion            | UNet3D           | 1.1M           | TS_73_6           | InstanceNorm3d    | PReLU          | 0.76543       | 0.76025        |\n| 6      | epoch152-step3672-valid_loss0.4333-val_metric0.7929.ckpt | Lion            | DenseVNet        | 873K           | TS_6_6            | InstanceNorm3d    | PReLU          | 0.76528       | 0.75417        |\n| 7      | epoch195-step4704-valid_loss0.4258-val_metric0.7914.ckpt | Lion            | VoxResNet        | 7.0M           | TS_6_6            | InstanceNorm3d    | PReLU          | 0.77457       | 0.76593        |\n| 8      | epoch188-step4536-valid_loss0.4231-val_metric0.8659.ckpt | Lion            | VoxHRNet         | 1.4M           | TS_73_6           | InstanceNorm3d    | PReLU          | 0.76738       | 0.75995        |\n| 9      | epoch198-step4776-valid_loss0.3471-val_metric0.8730.ckpt | Lion            | VoxHRNet         | 1.4M           | TS_73_6           | InstanceNorm3d    | PReLU          | 0.76135       | 0.75848        |\n| 10     | epoch133-val_loss0.52-val_metric0.56-step3216.ckpt       | Luoziqian       | UNet3D           | 1.1M           | TS_6_4            | BatchNorm3d       | PReLU          | 0.76844       | 0.76320        |\n| 11     | epoch314-val_loss0.54-val_metric0.54-step7560.ckpt       | Luoziqian       | SegResNet        | 1.2M           | TS_6_4            | GroupNorm         | ReLU           | 0.75521       | 0.74647        |\n| 12     | epoch114-val_loss0.55-val_metric0.53-Step2760.ckpt       | Luoziqian       | UNet2E3D         | 14.2M          | TS_6_4            | BatchNorm3d       | ReLU           | 0.73758       | 0.72966        |\n\n\n## Ensemble Strategy (Partial Selection)\nWe select the models with the best public leaderboard scores for the final submission.\n\n| **Models**                        | **Ensemble Strategy** | **Certainty Threshold** | **Public LB** | **Private LB** |    **Selected**   |\n| --------------------------------- | --------------------- |  ---------------------  | ------------- | -------------- | ----------------- |\n| [1, 2, 3, 4, 7, 8, 12]            | Average               |          0.23           | 0.79094       | 0.78641        |        ❌         |\n| [1, 2, 3, 4, 7, 8, 12]            | Average               |          0.18           | 0.79213       | 0.78630        |        ❌         |\n| [1, 2, 3, 4, 7, 8, 12]            | Average               |          0.15           | 0.79307       | 0.78457        |        ❌         |\n| [1, 2, 3, 4, 7, 8, 12]            | weighted              |          0.15           | 0.79247       | 0.78417        |        ❌         |\n| [1, 2, 3, 4, 6, 12]               | Average               |          0.15           | 0.78701       | 0.78389        |        ❌         |\n| [1, 2, 3, 4, 5, 6, 7]             | Average               |          0.15           | 0.79104       | 0.78277        |        ❌         |\n| [1, 2, 3, 4, 7, 9, 12]            | Average               |          0.15           | 0.79391       | 0.78283        |        ✅         |\n| [1, 2, 3, 4, 6, 7, 9, 10, 11, 12] | Average               |          0.15           | 0.79355       | 0.78381        |        ✅         |\n\n## Post-processing (Cluster Removal)\n\nWe used `CC3D` to convert the segmentation results into particle clusters and applied the following strategies for cluster selection:\n\n- Threshold-based clustering (the threshold is set to 0.15 for all classses to improve recall)\n- Cluster size filtering\n\n## Methods that Did Not Yield Improvements:\n\n- Synthetic data\n- Second-stage classification\n- Heavy-weight models (transformer-based models)\n\n## Code\n\n[Inference Notebook](https://www.kaggle.com/code/luoziqian/7-model-6tta-t4x2-patch-norm-v2?scriptVersionId=220452693)\n[Lion's Training Code](https://github.com/GWwangshuo/Kaggle-2024-CZII-Pub)\n[Luoziqian's Training Code](https://github.com/luoziqianX/CZII-CryoET-Object-Identification-2nd-luoziqian)\n\n## Reference\n\nGubins, Ilja, et al. \"SHREC 2020: Classification in cryo-electron tomograms.\" *Computers & Graphics* 91 (2020): 279-289.",
    "3117309": "Wow, you have very strong single models. Focus on stabilizing the training might be the trick(s) we missed. Congratulations on the gold and the prize!",
    "3120885": "Congratulations !\nI've realized that there are a lot of room for improvement even if segmentation + cc3d approach and lack of my skills to design experiments.\nIn Luoziqian's Training Code, I cannot find output comp metrics in training code. Did you choose best model by other way ?",
    "3119199": "Congratulations! @sjtuwangshuo .\n\nWhich sample do you use to validate? ",
    "3118525": "Congratulations! thank you for sharing your solution\n\nI have a question :why  experiment no.2 and no.10/11/12 have such a huge different valid score(0.7x vs 0.5x,but lb very close)？\n\n",
    "3117336": "Congrats! Glad the notebook helped you get started!",
    "3118220": ""
  }
}