{
  "id": 561518,
  "title": "6th Place Solution",
  "url": "/competitions/czii-cryo-et-object-identification/writeups/tomoon33-6th-place-solution",
  "author_name": "",
  "post_date": "2025-02-12T12:43:10.877Z",
  "votes": 24,
  "comment_count": 5,
  "views": 0,
  "content": "<p>First of all, I would like to express my sincere gratitude to the competition organizers for hosting this exciting challenge. Thanks to their efforts, I had the opportunity to learn a lot and improve my skills. I also appreciate the Kaggle community for their insightful discussions and valuable resources.<br>\nIn particular, I found <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s discussions and notebooks extremely helpful. The insights and contributions were instrumental in shaping my approach to this competition.</p>\n<h1>Pipeline Overview</h1>\n<ol>\n<li>Use MONAI’s <code>sliding_window_inference</code> to divide the input volume into fixed-size subvolumes and perform inference with 10 different models. Each model predicts a map for each particle type, indicating the locations of particles.</li>\n<li>Average 10 prediction maps for each particle.</li>\n<li>Binarize the averaged map using a threshold specific to each particle.</li>\n<li>Apply connected component analysis to the binarized volume and calculate the centroid of each connected component as the particle’s location.</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15217057%2F3a44dec0b4a6a9662196d5f34279162c%2Foverview.png?generation=1738853660680589&amp;alt=media\" alt=\"\"></p>\n<h1>Models</h1>\n<p>I utilized a 2.5D UNet based on <a href=\"https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder\" target=\"_blank\">this notebook</a>, as well as MONAI’s 3D UNet and SegResNet. For the 2.5D UNet, timm's EfficientNet-B2, EfficientNetV2-B2, ConvNeXt-Nano, and ResNet34d were used as encoders. While the 3D UNet and SegResNet initially struggled to achieve good performance using only the training dataset, they reached a similar level of accuracy as the 2.5D UNet after pretraining on simulated data, which I will describe later.<br>\nTo improve model diversity, I trained a total of 10 models with slight variations in pretraining strategies, data augmentation techniques, and other hyperparameters. For the final submission, I used models trained on the entire training dataset.</p>\n<h1>Training</h1>\n<h2>Preprocessing</h2>\n<p>As a preprocessing step, I first removed outliers and then applied min-max normalization. The input volumes were cropped to a fixed size of 64×128×128.</p>\n<h2>Particle Mask</h2>\n<p>For the training labels, I generated binary masks centered around each particle’s location. Specifically, I set the region within radius × 0.5 to 1 and all other areas to 0. I created these masks separately for each particle.</p>\n<h2>Loss Function</h2>\n<p>Initially, I experimented with BCE and Focal Loss, but these loss functions resulted in extremely slow convergence, likely due to the severe class imbalance. I then tried Dice Loss and Tversky Loss, which significantly accelerated training. However, these loss functions caused the model’s predictions to become overly binary, outputting only extreme values of 0 or 1. <br>\nSince my pipeline ensembles prediction maps by averaging them and then applies a threshold for binarization, such prediction values were undesirable. To address this, I used FocalTversky++ loss as proposed in <a href=\"https://arxiv.org/pdf/2111.00528\" target=\"_blank\">this paper</a>. By using this loss, the model outputs prediction maps that reflect its confidence.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15217057%2Fb1f14e94d7758fd033ee36f9904fad9f%2Fmaps.png?generation=1738851373469946&amp;alt=media\" alt=\"\"></p>\n<h2>Pretraining with Simulated Data</h2>\n<p>I generated my own simulated data using <a href=\"https://github.com/anmartinezs/polnet\" target=\"_blank\">polnet</a> and used it to pretrain some of the models. To enhance the model’s ability to distinguish particles, I designed the simulated data to include particles with shapes similar to those of the target particles in this competition.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15217057%2F8937037775d2e7fec1d786f3d43e25de%2Fsim_row.png?generation=1738852645400856&amp;alt=media\" alt=\"\"></p>\n<h1>Inference</h1>\n<p>For inference, I used MONAI’s <code>sliding_window_inference</code> with an overlap ratio of 0.25. Since predictions near the edges of subvolumes tend to be less stable, I discarded the outermost 8% of predictions to improve reliability.<br>\nI felt that one of the key challenges in this competition was achieving fast inference. To optimize speed, I first converted the models to TensorRT in a separate notebook before submission. During inference, I loaded the pre-converted TensorRT models and leveraged parallel processing with two T4 GPUs.</p>\n<h1>Post Processing</h1>\n<p>Once the ensembled prediction maps were obtained, I binarized them using particle-specific thresholds. These thresholds were initially estimated through cross-validation and later fine-tuned based on leaderboard.<br>\nAfter binarization, I performed connected component analysis on the binary map and extracted the centroid of each connected component as the particle location.</p>\n<h1>What Did Not Work</h1>\n<ul>\n<li>Two-stage model</li>\n<li>Use of tomogram data other than denoised</li>\n</ul>\n<h1>Code</h1>\n<ul>\n<li><a href=\"https://github.com/uchiyama33/czii-6th-place\" target=\"_blank\">https://github.com/uchiyama33/czii-6th-place</a></li>\n<li><a href=\"https://www.kaggle.com/code/tomoon33/czii-submission-6th-place\" target=\"_blank\">https://www.kaggle.com/code/tomoon33/czii-submission-6th-place</a></li>\n</ul>",
  "messages": [
    {
      "id": "3117030",
      "postDate": "02/06/2025 14:59:52",
      "content": "<p>First of all, I would like to express my sincere gratitude to the competition organizers for hosting this exciting challenge. Thanks to their efforts, I had the opportunity to learn a lot and improve my skills. I also appreciate the Kaggle community for their insightful discussions and valuable resources.<br>\nIn particular, I found <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s discussions and notebooks extremely helpful. The insights and contributions were instrumental in shaping my approach to this competition.</p>\n<h1>Pipeline Overview</h1>\n<ol>\n<li>Use MONAI’s <code>sliding_window_inference</code> to divide the input volume into fixed-size subvolumes and perform inference with 10 different models. Each model predicts a map for each particle type, indicating the locations of particles.</li>\n<li>Average 10 prediction maps for each particle.</li>\n<li>Binarize the averaged map using a threshold specific to each particle.</li>\n<li>Apply connected component analysis to the binarized volume and calculate the centroid of each connected component as the particle’s location.</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15217057%2F3a44dec0b4a6a9662196d5f34279162c%2Foverview.png?generation=1738853660680589&amp;alt=media\" alt=\"\"></p>\n<h1>Models</h1>\n<p>I utilized a 2.5D UNet based on <a href=\"https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder\" target=\"_blank\">this notebook</a>, as well as MONAI’s 3D UNet and SegResNet. For the 2.5D UNet, timm's EfficientNet-B2, EfficientNetV2-B2, ConvNeXt-Nano, and ResNet34d were used as encoders. While the 3D UNet and SegResNet initially struggled to achieve good performance using only the training dataset, they reached a similar level of accuracy as the 2.5D UNet after pretraining on simulated data, which I will describe later.<br>\nTo improve model diversity, I trained a total of 10 models with slight variations in pretraining strategies, data augmentation techniques, and other hyperparameters. For the final submission, I used models trained on the entire training dataset.</p>\n<h1>Training</h1>\n<h2>Preprocessing</h2>\n<p>As a preprocessing step, I first removed outliers and then applied min-max normalization. The input volumes were cropped to a fixed size of 64×128×128.</p>\n<h2>Particle Mask</h2>\n<p>For the training labels, I generated binary masks centered around each particle’s location. Specifically, I set the region within radius × 0.5 to 1 and all other areas to 0. I created these masks separately for each particle.</p>\n<h2>Loss Function</h2>\n<p>Initially, I experimented with BCE and Focal Loss, but these loss functions resulted in extremely slow convergence, likely due to the severe class imbalance. I then tried Dice Loss and Tversky Loss, which significantly accelerated training. However, these loss functions caused the model’s predictions to become overly binary, outputting only extreme values of 0 or 1. <br>\nSince my pipeline ensembles prediction maps by averaging them and then applies a threshold for binarization, such prediction values were undesirable. To address this, I used FocalTversky++ loss as proposed in <a href=\"https://arxiv.org/pdf/2111.00528\" target=\"_blank\">this paper</a>. By using this loss, the model outputs prediction maps that reflect its confidence.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15217057%2Fb1f14e94d7758fd033ee36f9904fad9f%2Fmaps.png?generation=1738851373469946&amp;alt=media\" alt=\"\"></p>\n<h2>Pretraining with Simulated Data</h2>\n<p>I generated my own simulated data using <a href=\"https://github.com/anmartinezs/polnet\" target=\"_blank\">polnet</a> and used it to pretrain some of the models. To enhance the model’s ability to distinguish particles, I designed the simulated data to include particles with shapes similar to those of the target particles in this competition.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15217057%2F8937037775d2e7fec1d786f3d43e25de%2Fsim_row.png?generation=1738852645400856&amp;alt=media\" alt=\"\"></p>\n<h1>Inference</h1>\n<p>For inference, I used MONAI’s <code>sliding_window_inference</code> with an overlap ratio of 0.25. Since predictions near the edges of subvolumes tend to be less stable, I discarded the outermost 8% of predictions to improve reliability.<br>\nI felt that one of the key challenges in this competition was achieving fast inference. To optimize speed, I first converted the models to TensorRT in a separate notebook before submission. During inference, I loaded the pre-converted TensorRT models and leveraged parallel processing with two T4 GPUs.</p>\n<h1>Post Processing</h1>\n<p>Once the ensembled prediction maps were obtained, I binarized them using particle-specific thresholds. These thresholds were initially estimated through cross-validation and later fine-tuned based on leaderboard.<br>\nAfter binarization, I performed connected component analysis on the binary map and extracted the centroid of each connected component as the particle location.</p>\n<h1>What Did Not Work</h1>\n<ul>\n<li>Two-stage model</li>\n<li>Use of tomogram data other than denoised</li>\n</ul>\n<h1>Code</h1>\n<ul>\n<li><a href=\"https://github.com/uchiyama33/czii-6th-place\" target=\"_blank\">https://github.com/uchiyama33/czii-6th-place</a></li>\n<li><a href=\"https://www.kaggle.com/code/tomoon33/czii-submission-6th-place\" target=\"_blank\">https://www.kaggle.com/code/tomoon33/czii-submission-6th-place</a></li>\n</ul>",
      "rawMarkdown": "First of all, I would like to express my sincere gratitude to the competition organizers for hosting this exciting challenge. Thanks to their efforts, I had the opportunity to learn a lot and improve my skills. I also appreciate the Kaggle community for their insightful discussions and valuable resources.\nIn particular, I found @hengck23's discussions and notebooks extremely helpful. The insights and contributions were instrumental in shaping my approach to this competition.\n\n# Pipeline Overview\n1.\tUse MONAI’s `sliding_window_inference` to divide the input volume into fixed-size subvolumes and perform inference with 10 different models. Each model predicts a map for each particle type, indicating the locations of particles.\n2.\tAverage 10 prediction maps for each particle.\n3.\tBinarize the averaged map using a threshold specific to each particle.\n4.\tApply connected component analysis to the binarized volume and calculate the centroid of each connected component as the particle’s location.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15217057%2F3a44dec0b4a6a9662196d5f34279162c%2Foverview.png?generation=1738853660680589&alt=media)\n\n# Models\nI utilized a 2.5D UNet based on [this notebook](https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder), as well as MONAI’s 3D UNet and SegResNet. For the 2.5D UNet, timm's EfficientNet-B2, EfficientNetV2-B2, ConvNeXt-Nano, and ResNet34d were used as encoders. While the 3D UNet and SegResNet initially struggled to achieve good performance using only the training dataset, they reached a similar level of accuracy as the 2.5D UNet after pretraining on simulated data, which I will describe later.\nTo improve model diversity, I trained a total of 10 models with slight variations in pretraining strategies, data augmentation techniques, and other hyperparameters. For the final submission, I used models trained on the entire training dataset.\n\n# Training\n## Preprocessing\nAs a preprocessing step, I first removed outliers and then applied min-max normalization. The input volumes were cropped to a fixed size of 64×128×128.\n\n## Particle Mask\nFor the training labels, I generated binary masks centered around each particle’s location. Specifically, I set the region within radius × 0.5 to 1 and all other areas to 0. I created these masks separately for each particle.\n\n## Loss Function\nInitially, I experimented with BCE and Focal Loss, but these loss functions resulted in extremely slow convergence, likely due to the severe class imbalance. I then tried Dice Loss and Tversky Loss, which significantly accelerated training. However, these loss functions caused the model’s predictions to become overly binary, outputting only extreme values of 0 or 1. \nSince my pipeline ensembles prediction maps by averaging them and then applies a threshold for binarization, such prediction values were undesirable. To address this, I used FocalTversky++ loss as proposed in [this paper](https://arxiv.org/pdf/2111.00528). By using this loss, the model outputs prediction maps that reflect its confidence.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15217057%2Fb1f14e94d7758fd033ee36f9904fad9f%2Fmaps.png?generation=1738851373469946&alt=media)\n\n## Pretraining with Simulated Data\nI generated my own simulated data using [polnet](https://github.com/anmartinezs/polnet) and used it to pretrain some of the models. To enhance the model’s ability to distinguish particles, I designed the simulated data to include particles with shapes similar to those of the target particles in this competition.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15217057%2F8937037775d2e7fec1d786f3d43e25de%2Fsim_row.png?generation=1738852645400856&alt=media)\n\n# Inference\nFor inference, I used MONAI’s `sliding_window_inference` with an overlap ratio of 0.25. Since predictions near the edges of subvolumes tend to be less stable, I discarded the outermost 8% of predictions to improve reliability.\nI felt that one of the key challenges in this competition was achieving fast inference. To optimize speed, I first converted the models to TensorRT in a separate notebook before submission. During inference, I loaded the pre-converted TensorRT models and leveraged parallel processing with two T4 GPUs.\n\n# Post Processing\nOnce the ensembled prediction maps were obtained, I binarized them using particle-specific thresholds. These thresholds were initially estimated through cross-validation and later fine-tuned based on leaderboard.\nAfter binarization, I performed connected component analysis on the binary map and extracted the centroid of each connected component as the particle location.\n\n# What Did Not Work\n- Two-stage model\n- Use of tomogram data other than denoised\n\n# Code\n- https://github.com/uchiyama33/czii-6th-place\n- https://www.kaggle.com/code/tomoon33/czii-submission-6th-place",
      "votes": null
    },
    {
      "id": "3117038",
      "postDate": "02/06/2025 15:12:27",
      "content": "<p>Congratulations!!! <br>\nWould it be possible for you to share the single model score in the final solution?</p>",
      "rawMarkdown": "Congratulations!!! \nWould it be possible for you to share the single model score in the final solution?",
      "votes": null
    },
    {
      "id": "3117043",
      "postDate": "02/06/2025 15:19:29",
      "content": "<p>Public LBs of the single model ranged from 0.757 to 0.766. <br>\n(However, the thresholds in the single model were not fine-tuned, so the scores could be a bit higher if optimal thresholds were used.)</p>",
      "rawMarkdown": "Public LBs of the single model ranged from 0.757 to 0.766. \n(However, the thresholds in the single model were not fine-tuned, so the scores could be a bit higher if optimal thresholds were used.)",
      "votes": null
    },
    {
      "id": "3117180",
      "postDate": "02/06/2025 17:44:39",
      "content": "<p>Interesting, thanks for sharing. </p>\n<p>Were you able to measure how much your CV/LB scores changed when using FocalTversky++ over Tversky?</p>",
      "rawMarkdown": "Interesting, thanks for sharing. \n\nWere you able to measure how much your CV/LB scores changed when using FocalTversky++ over Tversky?",
      "votes": null
    },
    {
      "id": "3117868",
      "postDate": "02/07/2025 10:26:24",
      "content": "<p>Single model scores increased by about 0.02 for both CV and LB. While I can't provide exact values, the score improvement with ensembling was significantly larger when using FocalTversky++.</p>",
      "rawMarkdown": "Single model scores increased by about 0.02 for both CV and LB. While I can't provide exact values, the score improvement with ensembling was significantly larger when using FocalTversky++.",
      "votes": null
    },
    {
      "id": "3122272",
      "postDate": "02/12/2025 12:46:04",
      "content": "<p>I’ve updated my solution code.</p>",
      "rawMarkdown": "I’ve updated my solution code.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3117038,
      "author_name": "maedward",
      "author_url": "",
      "post_date": "02/06/2025 15:12:27",
      "content": "<p>Congratulations!!! <br>\nWould it be possible for you to share the single model score in the final solution?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3117043,
          "author_name": "tomoon33",
          "author_url": "",
          "post_date": "02/06/2025 15:19:29",
          "content": "<p>Public LBs of the single model ranged from 0.757 to 0.766. <br>\n(However, the thresholds in the single model were not fine-tuned, so the scores could be a bit higher if optimal thresholds were used.)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3117180,
      "author_name": "brendanartley",
      "author_url": "",
      "post_date": "02/06/2025 17:44:39",
      "content": "<p>Interesting, thanks for sharing. </p>\n<p>Were you able to measure how much your CV/LB scores changed when using FocalTversky++ over Tversky?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3117868,
          "author_name": "tomoon33",
          "author_url": "",
          "post_date": "02/07/2025 10:26:24",
          "content": "<p>Single model scores increased by about 0.02 for both CV and LB. While I can't provide exact values, the score improvement with ensembling was significantly larger when using FocalTversky++.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3122272,
      "author_name": "tomoon33",
      "author_url": "",
      "post_date": "02/12/2025 12:46:04",
      "content": "<p>I’ve updated my solution code.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3117030": "First of all, I would like to express my sincere gratitude to the competition organizers for hosting this exciting challenge. Thanks to their efforts, I had the opportunity to learn a lot and improve my skills. I also appreciate the Kaggle community for their insightful discussions and valuable resources.\nIn particular, I found @hengck23's discussions and notebooks extremely helpful. The insights and contributions were instrumental in shaping my approach to this competition.\n\n# Pipeline Overview\n1.\tUse MONAI’s `sliding_window_inference` to divide the input volume into fixed-size subvolumes and perform inference with 10 different models. Each model predicts a map for each particle type, indicating the locations of particles.\n2.\tAverage 10 prediction maps for each particle.\n3.\tBinarize the averaged map using a threshold specific to each particle.\n4.\tApply connected component analysis to the binarized volume and calculate the centroid of each connected component as the particle’s location.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15217057%2F3a44dec0b4a6a9662196d5f34279162c%2Foverview.png?generation=1738853660680589&alt=media)\n\n# Models\nI utilized a 2.5D UNet based on [this notebook](https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder), as well as MONAI’s 3D UNet and SegResNet. For the 2.5D UNet, timm's EfficientNet-B2, EfficientNetV2-B2, ConvNeXt-Nano, and ResNet34d were used as encoders. While the 3D UNet and SegResNet initially struggled to achieve good performance using only the training dataset, they reached a similar level of accuracy as the 2.5D UNet after pretraining on simulated data, which I will describe later.\nTo improve model diversity, I trained a total of 10 models with slight variations in pretraining strategies, data augmentation techniques, and other hyperparameters. For the final submission, I used models trained on the entire training dataset.\n\n# Training\n## Preprocessing\nAs a preprocessing step, I first removed outliers and then applied min-max normalization. The input volumes were cropped to a fixed size of 64×128×128.\n\n## Particle Mask\nFor the training labels, I generated binary masks centered around each particle’s location. Specifically, I set the region within radius × 0.5 to 1 and all other areas to 0. I created these masks separately for each particle.\n\n## Loss Function\nInitially, I experimented with BCE and Focal Loss, but these loss functions resulted in extremely slow convergence, likely due to the severe class imbalance. I then tried Dice Loss and Tversky Loss, which significantly accelerated training. However, these loss functions caused the model’s predictions to become overly binary, outputting only extreme values of 0 or 1. \nSince my pipeline ensembles prediction maps by averaging them and then applies a threshold for binarization, such prediction values were undesirable. To address this, I used FocalTversky++ loss as proposed in [this paper](https://arxiv.org/pdf/2111.00528). By using this loss, the model outputs prediction maps that reflect its confidence.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15217057%2Fb1f14e94d7758fd033ee36f9904fad9f%2Fmaps.png?generation=1738851373469946&alt=media)\n\n## Pretraining with Simulated Data\nI generated my own simulated data using [polnet](https://github.com/anmartinezs/polnet) and used it to pretrain some of the models. To enhance the model’s ability to distinguish particles, I designed the simulated data to include particles with shapes similar to those of the target particles in this competition.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15217057%2F8937037775d2e7fec1d786f3d43e25de%2Fsim_row.png?generation=1738852645400856&alt=media)\n\n# Inference\nFor inference, I used MONAI’s `sliding_window_inference` with an overlap ratio of 0.25. Since predictions near the edges of subvolumes tend to be less stable, I discarded the outermost 8% of predictions to improve reliability.\nI felt that one of the key challenges in this competition was achieving fast inference. To optimize speed, I first converted the models to TensorRT in a separate notebook before submission. During inference, I loaded the pre-converted TensorRT models and leveraged parallel processing with two T4 GPUs.\n\n# Post Processing\nOnce the ensembled prediction maps were obtained, I binarized them using particle-specific thresholds. These thresholds were initially estimated through cross-validation and later fine-tuned based on leaderboard.\nAfter binarization, I performed connected component analysis on the binary map and extracted the centroid of each connected component as the particle location.\n\n# What Did Not Work\n- Two-stage model\n- Use of tomogram data other than denoised\n\n# Code\n- https://github.com/uchiyama33/czii-6th-place\n- https://www.kaggle.com/code/tomoon33/czii-submission-6th-place",
    "3117038": "Congratulations!!! \nWould it be possible for you to share the single model score in the final solution?",
    "3117043": "Public LBs of the single model ranged from 0.757 to 0.766. \n(However, the thresholds in the single model were not fine-tuned, so the scores could be a bit higher if optimal thresholds were used.)",
    "3117180": "Interesting, thanks for sharing. \n\nWere you able to measure how much your CV/LB scores changed when using FocalTversky++ over Tversky?",
    "3117868": "Single model scores increased by about 0.02 for both CV and LB. While I can't provide exact values, the score improvement with ensembling was significantly larger when using FocalTversky++.",
    "3122272": "I’ve updated my solution code."
  },
  "source": "meta"
}