{
  "id": 561431,
  "title": "9th place solution",
  "url": "/competitions/czii-cryo-et-object-identification/writeups/avengers-9th-place-solution",
  "author_name": "",
  "post_date": "2025-03-03T23:32:00.743Z",
  "votes": 51,
  "comment_count": 3,
  "views": 0,
  "content": "<p>First of all, I would like to express my sincere gratitude to the competition host and the Kaggle staff for organizing such a fascinating competition. I thoroughly enjoyed this competition and learned a great deal in the process!</p>\n<p>Furthermore, I'd like to thank <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a> . <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's discussion and notebook were the starting point for my solution, and Many of the discussions and comments from <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a> were extremely helpful.</p>\n<h2>Summary</h2>\n<p>My solution is simple and straightforward. I performed segmentation using a 3D ConvNeXt-like model and further employed an ensemble of as many models as possible. Subsequently, I computed the centroids of the particles using cc3d and curated the predicted centroids with DBSCAN. </p>\n<h2>Segmentation Mask</h2>\n<p>I used ground truth masks with an adjusted radius for each particle. As a result, the public leaderboard score jumped 0.02~0.04. Concletely, I adjust mask size as shown in the below table; </p>\n<table>\n<thead>\n<tr>\n<th>apo-ferritin(r=60)</th>\n<th>beta-galactosidase(r=90)</th>\n<th>ribosome(r=150)</th>\n<th>thyroglobulin(r=130)</th>\n<th>virus-like-particle(r=135)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>60/2</td>\n<td>90/2</td>\n<td>150/3</td>\n<td>130/3</td>\n<td>135/3</td>\n</tr>\n</tbody>\n</table>\n<p>Since the competition metric requires that the predicted centroid falls within a radius of r×0.5 from the ground truth centroid, I believe that multiplying the radius of each particle by a factor of 0.5 or less is reasonable. </p>\n<h2>Model Architecture</h2>\n<p>In this section, I explain my model architecture both encoder and decoder. </p>\n<h3>Encoder</h3>\n<p>To implement the encoder, I started with the ConvNeXt, since ConvNeXt is very powerful and fast. Then, I customized the model to suit the task that is to predict small particles. These changes from the original ConvNeXt are below;</p>\n<ul>\n<li>2D to 3D for all convolutional layer</li>\n<li>4x4 stem -&gt; 2x2 stem. Because the particles are very small, I believed that a large stem would adversely affect the model's predictions, especially apo-ferritin and beta-galactosidase. This modification slowed down the model's inference, but it improved its cross-validation performance. </li>\n<li>7x7 kernel size in conv block -&gt;3x3. The reason why I modified like this is the same as the stem. This modification improved both inference speed and cv. </li>\n<li>(3, 3, 9, 3) number of block -&gt; (3, 3, 3, 3). this modification is to reduce a number of parameters and get more inference speed. cv did not decrease. </li>\n</ul>\n<h3>Decoder</h3>\n<p>The flow of the decoder is based on U-Net. The implementation of conv block is shown in the following code; </p>\n<pre><code>\n (nn.Module):\n     ():\n        ().__init__()\n        .conv1 = nn.Conv3d(in_channels, in_channels, kernel_size=, padding=, bias=, groups=in_channels)\n        .norm1 = nn.GroupNorm(num_groups=, num_channels=in_channels)\n        .act1 = nn.GELU()\n        .conv2 = nn.Conv3d(in_channels, in_channels*, kernel_size=, bias=)\n        .conv3 = nn.Conv3d(in_channels*, out_channels, kernel_size=, bias=)\n\n     ():\n        \n        x = .conv1(x)\n        x = .norm1(x)\n        x = .conv2(x)\n        x = .act1(x)\n        x = .conv3(x)\n         x\n</code></pre>\n<p>And the implementation of the decoder is shown in the following code; </p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        .up3 = nn.Upsample(scale_factor=(,,), mode=, align_corners=)\n        \n        .dec3 = ConvBlock3D(in_channels=encoder_dims[] + encoder_dims[], out_channels=decoder_dims[])\n       ...\n\n     ():\n        x, f0, f1, f2, f3 = features\n\n        \n        d3 = .up3(f3)\n        d3 = torch.cat([d3, f2], dim=)\n        d3 = .dec3(d3)\n        ...\n</code></pre>\n<p>Since the ground truth for segmentation is simply a sphere, I used a basic <code>nn.Upsample</code> for upsampling instead of <code>nn.ConvTranspose3d</code> to reduce parameters. </p>\n<h2>Other Details</h2>\n<ul>\n<li>loss dunction is bce</li>\n<li>My inference code is based on <a href=\"https://www.kaggle.com/hengck\" target=\"_blank\">@hengck</a> 's <a href=\"https://www.kaggle.com/code/hengck23/1-hr-fast-2d-3d-unet-resnet34d-scanner-tta\" target=\"_blank\">this notebook</a> and the idea of DBScan came from <a href=\"https://www.kaggle.com/linheshen\" target=\"_blank\">@linheshen</a> 's <a href=\"https://www.kaggle.com/code/hengck23/1-hr-fast-2d-3d-unet-resnet34d-scanner-tta\" target=\"_blank\">notebook</a>. </li>\n<li>The window size for inference is (32, 320, 320) (overwrap of z axis is 6). The window size for training is (32, 256, 256). </li>\n<li>TTAs are rot90, 180 and 270. </li>\n<li>preprocessing is only normalization</li>\n</ul>\n<pre><code>     ():\n        lower, upper = np.percentile(x, (, ))\n        x = np.clip(x, lower, upper)\n        x = x - np.(x)\n        x = x / np.(x)\n         x\n</code></pre>\n<ul>\n<li>Augmentations for training are rot90, 180, 270, and xyz flip. </li>\n</ul>\n<h2>What did not work</h2>\n<ul>\n<li>pretraining with synthetic data provided by host</li>\n<li>more and more smaller mask(e.g. factor=0.25 for large particles and factor=0.4 for small particles)</li>\n<li>flip tta instead of rot90</li>\n</ul>\n<h2>Code &amp; Model Checkpoint</h2>\n<p>inference code: <a href=\"https://www.kaggle.com/code/wadakoki/czii-9th-inference\" target=\"_blank\">https://www.kaggle.com/code/wadakoki/czii-9th-inference</a><br>\ntraining code: <a href=\"https://colab.research.google.com/drive/1OG5Rap9zUkKG6fnGZ8kP0JwwFZfSTgCc?usp=sharing\" target=\"_blank\">https://colab.research.google.com/drive/1OG5Rap9zUkKG6fnGZ8kP0JwwFZfSTgCc?usp=sharing</a><br>\nmodel checkpoint: <a href=\"https://www.kaggle.com/datasets/wadakoki/czii-9th-final-models/data\" target=\"_blank\">https://www.kaggle.com/datasets/wadakoki/czii-9th-final-models/data</a><br>\ngenerated mask: <a href=\"https://www.kaggle.com/datasets/wadakoki/czii-mask-small\" target=\"_blank\">https://www.kaggle.com/datasets/wadakoki/czii-mask-small</a><br>\ncode for mask generation: <a href=\"https://www.kaggle.com/code/wadakoki/czii-gen-small-masks\" target=\"_blank\">https://www.kaggle.com/code/wadakoki/czii-gen-small-masks</a></p>\n<p><strong>Requirements for training code on google colab</strong></p>\n<ul>\n<li>GPU: L4</li>\n<li>RAM: High Memory</li>\n</ul>",
  "messages": [
    {
      "id": "3116563",
      "postDate": "02/06/2025 04:29:32",
      "content": "<p>First of all, I would like to express my sincere gratitude to the competition host and the Kaggle staff for organizing such a fascinating competition. I thoroughly enjoyed this competition and learned a great deal in the process!</p>\n<p>Furthermore, I'd like to thank <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a> . <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's discussion and notebook were the starting point for my solution, and Many of the discussions and comments from <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a> were extremely helpful.</p>\n<h2>Summary</h2>\n<p>My solution is simple and straightforward. I performed segmentation using a 3D ConvNeXt-like model and further employed an ensemble of as many models as possible. Subsequently, I computed the centroids of the particles using cc3d and curated the predicted centroids with DBSCAN. </p>\n<h2>Segmentation Mask</h2>\n<p>I used ground truth masks with an adjusted radius for each particle. As a result, the public leaderboard score jumped 0.02~0.04. Concletely, I adjust mask size as shown in the below table; </p>\n<table>\n<thead>\n<tr>\n<th>apo-ferritin(r=60)</th>\n<th>beta-galactosidase(r=90)</th>\n<th>ribosome(r=150)</th>\n<th>thyroglobulin(r=130)</th>\n<th>virus-like-particle(r=135)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>60/2</td>\n<td>90/2</td>\n<td>150/3</td>\n<td>130/3</td>\n<td>135/3</td>\n</tr>\n</tbody>\n</table>\n<p>Since the competition metric requires that the predicted centroid falls within a radius of r×0.5 from the ground truth centroid, I believe that multiplying the radius of each particle by a factor of 0.5 or less is reasonable. </p>\n<h2>Model Architecture</h2>\n<p>In this section, I explain my model architecture both encoder and decoder. </p>\n<h3>Encoder</h3>\n<p>To implement the encoder, I started with the ConvNeXt, since ConvNeXt is very powerful and fast. Then, I customized the model to suit the task that is to predict small particles. These changes from the original ConvNeXt are below;</p>\n<ul>\n<li>2D to 3D for all convolutional layer</li>\n<li>4x4 stem -&gt; 2x2 stem. Because the particles are very small, I believed that a large stem would adversely affect the model's predictions, especially apo-ferritin and beta-galactosidase. This modification slowed down the model's inference, but it improved its cross-validation performance. </li>\n<li>7x7 kernel size in conv block -&gt;3x3. The reason why I modified like this is the same as the stem. This modification improved both inference speed and cv. </li>\n<li>(3, 3, 9, 3) number of block -&gt; (3, 3, 3, 3). this modification is to reduce a number of parameters and get more inference speed. cv did not decrease. </li>\n</ul>\n<h3>Decoder</h3>\n<p>The flow of the decoder is based on U-Net. The implementation of conv block is shown in the following code; </p>\n<pre><code>\n (nn.Module):\n     ():\n        ().__init__()\n        .conv1 = nn.Conv3d(in_channels, in_channels, kernel_size=, padding=, bias=, groups=in_channels)\n        .norm1 = nn.GroupNorm(num_groups=, num_channels=in_channels)\n        .act1 = nn.GELU()\n        .conv2 = nn.Conv3d(in_channels, in_channels*, kernel_size=, bias=)\n        .conv3 = nn.Conv3d(in_channels*, out_channels, kernel_size=, bias=)\n\n     ():\n        \n        x = .conv1(x)\n        x = .norm1(x)\n        x = .conv2(x)\n        x = .act1(x)\n        x = .conv3(x)\n         x\n</code></pre>\n<p>And the implementation of the decoder is shown in the following code; </p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        .up3 = nn.Upsample(scale_factor=(,,), mode=, align_corners=)\n        \n        .dec3 = ConvBlock3D(in_channels=encoder_dims[] + encoder_dims[], out_channels=decoder_dims[])\n       ...\n\n     ():\n        x, f0, f1, f2, f3 = features\n\n        \n        d3 = .up3(f3)\n        d3 = torch.cat([d3, f2], dim=)\n        d3 = .dec3(d3)\n        ...\n</code></pre>\n<p>Since the ground truth for segmentation is simply a sphere, I used a basic <code>nn.Upsample</code> for upsampling instead of <code>nn.ConvTranspose3d</code> to reduce parameters. </p>\n<h2>Other Details</h2>\n<ul>\n<li>loss dunction is bce</li>\n<li>My inference code is based on <a href=\"https://www.kaggle.com/hengck\" target=\"_blank\">@hengck</a> 's <a href=\"https://www.kaggle.com/code/hengck23/1-hr-fast-2d-3d-unet-resnet34d-scanner-tta\" target=\"_blank\">this notebook</a> and the idea of DBScan came from <a href=\"https://www.kaggle.com/linheshen\" target=\"_blank\">@linheshen</a> 's <a href=\"https://www.kaggle.com/code/hengck23/1-hr-fast-2d-3d-unet-resnet34d-scanner-tta\" target=\"_blank\">notebook</a>. </li>\n<li>The window size for inference is (32, 320, 320) (overwrap of z axis is 6). The window size for training is (32, 256, 256). </li>\n<li>TTAs are rot90, 180 and 270. </li>\n<li>preprocessing is only normalization</li>\n</ul>\n<pre><code>     ():\n        lower, upper = np.percentile(x, (, ))\n        x = np.clip(x, lower, upper)\n        x = x - np.(x)\n        x = x / np.(x)\n         x\n</code></pre>\n<ul>\n<li>Augmentations for training are rot90, 180, 270, and xyz flip. </li>\n</ul>\n<h2>What did not work</h2>\n<ul>\n<li>pretraining with synthetic data provided by host</li>\n<li>more and more smaller mask(e.g. factor=0.25 for large particles and factor=0.4 for small particles)</li>\n<li>flip tta instead of rot90</li>\n</ul>\n<h2>Code &amp; Model Checkpoint</h2>\n<p>inference code: <a href=\"https://www.kaggle.com/code/wadakoki/czii-9th-inference\" target=\"_blank\">https://www.kaggle.com/code/wadakoki/czii-9th-inference</a><br>\ntraining code: <a href=\"https://colab.research.google.com/drive/1OG5Rap9zUkKG6fnGZ8kP0JwwFZfSTgCc?usp=sharing\" target=\"_blank\">https://colab.research.google.com/drive/1OG5Rap9zUkKG6fnGZ8kP0JwwFZfSTgCc?usp=sharing</a><br>\nmodel checkpoint: <a href=\"https://www.kaggle.com/datasets/wadakoki/czii-9th-final-models/data\" target=\"_blank\">https://www.kaggle.com/datasets/wadakoki/czii-9th-final-models/data</a><br>\ngenerated mask: <a href=\"https://www.kaggle.com/datasets/wadakoki/czii-mask-small\" target=\"_blank\">https://www.kaggle.com/datasets/wadakoki/czii-mask-small</a><br>\ncode for mask generation: <a href=\"https://www.kaggle.com/code/wadakoki/czii-gen-small-masks\" target=\"_blank\">https://www.kaggle.com/code/wadakoki/czii-gen-small-masks</a></p>\n<p><strong>Requirements for training code on google colab</strong></p>\n<ul>\n<li>GPU: L4</li>\n<li>RAM: High Memory</li>\n</ul>",
      "rawMarkdown": "First of all, I would like to express my sincere gratitude to the competition host and the Kaggle staff for organizing such a fascinating competition. I thoroughly enjoyed this competition and learned a great deal in the process!\n\nFurthermore, I'd like to thank @hengck23 and @davidlist . @hengck23 's discussion and notebook were the starting point for my solution, and Many of the discussions and comments from @davidlist were extremely helpful.\n\n## Summary\n\nMy solution is simple and straightforward. I performed segmentation using a 3D ConvNeXt-like model and further employed an ensemble of as many models as possible. Subsequently, I computed the centroids of the particles using cc3d and curated the predicted centroids with DBSCAN. \n\n## Segmentation Mask\n\nI used ground truth masks with an adjusted radius for each particle. As a result, the public leaderboard score jumped 0.02~0.04. Concletely, I adjust mask size as shown in the below table; \n\n| apo-ferritin(r=60) | beta-galactosidase(r=90) | ribosome(r=150) | thyroglobulin(r=130) | virus-like-particle(r=135) |\n| --- | --- | --- | --- | --- |\n| 60/2 | 90/2 |  150/3 |  130/3 | 135/3 |\n\nSince the competition metric requires that the predicted centroid falls within a radius of r×0.5 from the ground truth centroid, I believe that multiplying the radius of each particle by a factor of 0.5 or less is reasonable. \n\n## Model Architecture\n\nIn this section, I explain my model architecture both encoder and decoder. \n\n### Encoder\n\nTo implement the encoder, I started with the ConvNeXt, since ConvNeXt is very powerful and fast. Then, I customized the model to suit the task that is to predict small particles. These changes from the original ConvNeXt are below;\n\n- 2D to 3D for all convolutional layer\n- 4x4 stem -> 2x2 stem. Because the particles are very small, I believed that a large stem would adversely affect the model's predictions, especially apo-ferritin and beta-galactosidase. This modification slowed down the model's inference, but it improved its cross-validation performance. \n- 7x7 kernel size in conv block ->3x3. The reason why I modified like this is the same as the stem. This modification improved both inference speed and cv. \n- (3, 3, 9, 3) number of block -> (3, 3, 3, 3). this modification is to reduce a number of parameters and get more inference speed. cv did not decrease. \n\n### Decoder\n\nThe flow of the decoder is based on U-Net. The implementation of conv block is shown in the following code; \n\n```python\n# conv block for decoder\nclass ConvBlock3D(nn.Module):\n    def __init__(self, in_channels, out_channels):\n        super().__init__()\n        self.conv1 = nn.Conv3d(in_channels, in_channels, kernel_size=3, padding=1, bias=False, groups=in_channels)\n        self.norm1 = nn.GroupNorm(num_groups=1, num_channels=in_channels)\n        self.act1 = nn.GELU()\n        self.conv2 = nn.Conv3d(in_channels, in_channels*3, kernel_size=1, bias=False)\n        self.conv3 = nn.Conv3d(in_channels*3, out_channels, kernel_size=1, bias=False)\n\n    def forward(self, x):\n        # x: (B, C, D, H, W)\n        x = self.conv1(x)\n        x = self.norm1(x)\n        x = self.conv2(x)\n        x = self.act1(x)\n        x = self.conv3(x)\n        return x\n```\nAnd the implementation of the decoder is shown in the following code; \n\n```python\nclass Decoder3D(nn.Module):\n    def __init__(self,\n                 encoder_dims=[64, 128, 256, 512],\n                 decoder_dims=[32, 64, 128, 256],\n                 out_channels=5):\n        super().__init__()\n        self.up3 = nn.Upsample(scale_factor=(2,2,2), mode='trilinear', align_corners=True)\n        #self.up3 = nn.ConvTranspose3d(encoder_dims[3], encoder_dims[3], kernel_size=2, stride=2)\n        self.dec3 = ConvBlock3D(in_channels=encoder_dims[3] + encoder_dims[2], out_channels=decoder_dims[2])\n       ...\n\n    def forward(self, features):\n        x, f0, f1, f2, f3 = features\n\n        # --- 1) f3 -> f2 ---\n        d3 = self.up3(f3)\n        d3 = torch.cat([d3, f2], dim=1)\n        d3 = self.dec3(d3)\n        ...\n```\n\nSince the ground truth for segmentation is simply a sphere, I used a basic `nn.Upsample` for upsampling instead of `nn.ConvTranspose3d` to reduce parameters. \n\n## Other Details\n\n- loss dunction is bce\n- My inference code is based on @hengck 's [this notebook](https://www.kaggle.com/code/hengck23/1-hr-fast-2d-3d-unet-resnet34d-scanner-tta) and the idea of DBScan came from @linheshen 's [notebook](https://www.kaggle.com/code/hengck23/1-hr-fast-2d-3d-unet-resnet34d-scanner-tta). \n- The window size for inference is (32, 320, 320) (overwrap of z axis is 6). The window size for training is (32, 256, 256). \n- TTAs are rot90, 180 and 270. \n- preprocessing is only normalization\n```python\n    def normalize_numpy(self, x):\n        lower, upper = np.percentile(x, (1, 99))\n        x = np.clip(x, lower, upper)\n        x = x - np.min(x)\n        x = x / np.max(x)\n        return x\n```\n- Augmentations for training are rot90, 180, 270, and xyz flip. \n\n## What did not work\n\n- pretraining with synthetic data provided by host\n- more and more smaller mask(e.g. factor=0.25 for large particles and factor=0.4 for small particles)\n- flip tta instead of rot90\n\n## Code & Model Checkpoint\ninference code: https://www.kaggle.com/code/wadakoki/czii-9th-inference\ntraining code: https://colab.research.google.com/drive/1OG5Rap9zUkKG6fnGZ8kP0JwwFZfSTgCc?usp=sharing\nmodel checkpoint: https://www.kaggle.com/datasets/wadakoki/czii-9th-final-models/data\ngenerated mask: https://www.kaggle.com/datasets/wadakoki/czii-mask-small\ncode for mask generation: https://www.kaggle.com/code/wadakoki/czii-gen-small-masks\n\n**Requirements for training code on google colab**\n- GPU: L4\n- RAM: High Memory",
      "votes": null
    },
    {
      "id": "3116652",
      "postDate": "02/06/2025 06:55:40",
      "content": "<p><a href=\"https://www.kaggle.com/wadakoki\" target=\"_blank\">@wadakoki</a> Thanks for the shout out!  Congratulations!!!</p>",
      "rawMarkdown": "wadakoki Thanks for the shout out!  Congratulations!!!",
      "votes": null
    },
    {
      "id": "3116771",
      "postDate": "02/06/2025 09:42:39",
      "content": "<p>Wow, a great implementation with such important details! Congratulations on your solo gold and prize!</p>",
      "rawMarkdown": "Wow, a great implementation with such important details! Congratulations on your solo gold and prize!",
      "votes": null
    },
    {
      "id": "3118807",
      "postDate": "02/08/2025 14:07:31",
      "content": "<p>[EDIT] I have added the training code, inference code, and the datasets used in this competition to the solution.</p>",
      "rawMarkdown": "[EDIT] I have added the training code, inference code, and the datasets used in this competition to the solution.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3116652,
      "author_name": "davidlist",
      "author_url": "",
      "post_date": "02/06/2025 06:55:40",
      "content": "<p><a href=\"https://www.kaggle.com/wadakoki\" target=\"_blank\">@wadakoki</a> Thanks for the shout out!  Congratulations!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3116771,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "02/06/2025 09:42:39",
      "content": "<p>Wow, a great implementation with such important details! Congratulations on your solo gold and prize!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3118807,
      "author_name": "wadakoki",
      "author_url": "",
      "post_date": "02/08/2025 14:07:31",
      "content": "<p>[EDIT] I have added the training code, inference code, and the datasets used in this competition to the solution.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3116563": "First of all, I would like to express my sincere gratitude to the competition host and the Kaggle staff for organizing such a fascinating competition. I thoroughly enjoyed this competition and learned a great deal in the process!\n\nFurthermore, I'd like to thank @hengck23 and @davidlist . @hengck23 's discussion and notebook were the starting point for my solution, and Many of the discussions and comments from @davidlist were extremely helpful.\n\n## Summary\n\nMy solution is simple and straightforward. I performed segmentation using a 3D ConvNeXt-like model and further employed an ensemble of as many models as possible. Subsequently, I computed the centroids of the particles using cc3d and curated the predicted centroids with DBSCAN. \n\n## Segmentation Mask\n\nI used ground truth masks with an adjusted radius for each particle. As a result, the public leaderboard score jumped 0.02~0.04. Concletely, I adjust mask size as shown in the below table; \n\n| apo-ferritin(r=60) | beta-galactosidase(r=90) | ribosome(r=150) | thyroglobulin(r=130) | virus-like-particle(r=135) |\n| --- | --- | --- | --- | --- |\n| 60/2 | 90/2 |  150/3 |  130/3 | 135/3 |\n\nSince the competition metric requires that the predicted centroid falls within a radius of r×0.5 from the ground truth centroid, I believe that multiplying the radius of each particle by a factor of 0.5 or less is reasonable. \n\n## Model Architecture\n\nIn this section, I explain my model architecture both encoder and decoder. \n\n### Encoder\n\nTo implement the encoder, I started with the ConvNeXt, since ConvNeXt is very powerful and fast. Then, I customized the model to suit the task that is to predict small particles. These changes from the original ConvNeXt are below;\n\n- 2D to 3D for all convolutional layer\n- 4x4 stem -> 2x2 stem. Because the particles are very small, I believed that a large stem would adversely affect the model's predictions, especially apo-ferritin and beta-galactosidase. This modification slowed down the model's inference, but it improved its cross-validation performance. \n- 7x7 kernel size in conv block ->3x3. The reason why I modified like this is the same as the stem. This modification improved both inference speed and cv. \n- (3, 3, 9, 3) number of block -> (3, 3, 3, 3). this modification is to reduce a number of parameters and get more inference speed. cv did not decrease. \n\n### Decoder\n\nThe flow of the decoder is based on U-Net. The implementation of conv block is shown in the following code; \n\n```python\n# conv block for decoder\nclass ConvBlock3D(nn.Module):\n    def __init__(self, in_channels, out_channels):\n        super().__init__()\n        self.conv1 = nn.Conv3d(in_channels, in_channels, kernel_size=3, padding=1, bias=False, groups=in_channels)\n        self.norm1 = nn.GroupNorm(num_groups=1, num_channels=in_channels)\n        self.act1 = nn.GELU()\n        self.conv2 = nn.Conv3d(in_channels, in_channels*3, kernel_size=1, bias=False)\n        self.conv3 = nn.Conv3d(in_channels*3, out_channels, kernel_size=1, bias=False)\n\n    def forward(self, x):\n        # x: (B, C, D, H, W)\n        x = self.conv1(x)\n        x = self.norm1(x)\n        x = self.conv2(x)\n        x = self.act1(x)\n        x = self.conv3(x)\n        return x\n```\nAnd the implementation of the decoder is shown in the following code; \n\n```python\nclass Decoder3D(nn.Module):\n    def __init__(self,\n                 encoder_dims=[64, 128, 256, 512],\n                 decoder_dims=[32, 64, 128, 256],\n                 out_channels=5):\n        super().__init__()\n        self.up3 = nn.Upsample(scale_factor=(2,2,2), mode='trilinear', align_corners=True)\n        #self.up3 = nn.ConvTranspose3d(encoder_dims[3], encoder_dims[3], kernel_size=2, stride=2)\n        self.dec3 = ConvBlock3D(in_channels=encoder_dims[3] + encoder_dims[2], out_channels=decoder_dims[2])\n       ...\n\n    def forward(self, features):\n        x, f0, f1, f2, f3 = features\n\n        # --- 1) f3 -> f2 ---\n        d3 = self.up3(f3)\n        d3 = torch.cat([d3, f2], dim=1)\n        d3 = self.dec3(d3)\n        ...\n```\n\nSince the ground truth for segmentation is simply a sphere, I used a basic `nn.Upsample` for upsampling instead of `nn.ConvTranspose3d` to reduce parameters. \n\n## Other Details\n\n- loss dunction is bce\n- My inference code is based on @hengck 's [this notebook](https://www.kaggle.com/code/hengck23/1-hr-fast-2d-3d-unet-resnet34d-scanner-tta) and the idea of DBScan came from @linheshen 's [notebook](https://www.kaggle.com/code/hengck23/1-hr-fast-2d-3d-unet-resnet34d-scanner-tta). \n- The window size for inference is (32, 320, 320) (overwrap of z axis is 6). The window size for training is (32, 256, 256). \n- TTAs are rot90, 180 and 270. \n- preprocessing is only normalization\n```python\n    def normalize_numpy(self, x):\n        lower, upper = np.percentile(x, (1, 99))\n        x = np.clip(x, lower, upper)\n        x = x - np.min(x)\n        x = x / np.max(x)\n        return x\n```\n- Augmentations for training are rot90, 180, 270, and xyz flip. \n\n## What did not work\n\n- pretraining with synthetic data provided by host\n- more and more smaller mask(e.g. factor=0.25 for large particles and factor=0.4 for small particles)\n- flip tta instead of rot90\n\n## Code & Model Checkpoint\ninference code: https://www.kaggle.com/code/wadakoki/czii-9th-inference\ntraining code: https://colab.research.google.com/drive/1OG5Rap9zUkKG6fnGZ8kP0JwwFZfSTgCc?usp=sharing\nmodel checkpoint: https://www.kaggle.com/datasets/wadakoki/czii-9th-final-models/data\ngenerated mask: https://www.kaggle.com/datasets/wadakoki/czii-mask-small\ncode for mask generation: https://www.kaggle.com/code/wadakoki/czii-gen-small-masks\n\n**Requirements for training code on google colab**\n- GPU: L4\n- RAM: High Memory",
    "3116652": "wadakoki Thanks for the shout out!  Congratulations!!!",
    "3116771": "Wow, a great implementation with such important details! Congratulations on your solo gold and prize!",
    "3118807": "[EDIT] I have added the training code, inference code, and the datasets used in this competition to the solution."
  },
  "source": "meta"
}