{
  "id": 561677,
  "title": "32nd Solution",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/561677",
  "author_name": "wpl",
  "post_date": "2025-02-07T08:04:51.010000",
  "votes": 20,
  "comment_count": 0,
  "views": 0,
  "content": "<h1><strong>Acknowledgements</strong></h1>\n<p>We sincerely appreciate Kaggle and the competition organizers for offering this invaluable opportunity. We also extend our gratitude to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , <a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a> and <a href=\"https://www.kaggle.com/sjtuwangshuo\" target=\"_blank\">@sjtuwangshuo</a> for their significant contributions. Lastly, I would like to sincerely thank my teammates for their dedication and hard work during this time! <a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a> , <a href=\"https://www.kaggle.com/miyamotodaiya\" target=\"_blank\">@miyamotodaiya</a> and <a href=\"https://www.kaggle.com/yingpengchen\" target=\"_blank\">@yingpengchen</a> </p>\n<h1><strong>Model</strong></h1>\n<p>We used the basic UNet3D model provided by MONAI as the primary model, and we implemented DLinkNet3D as an auxiliary model using Torch. Given the limited number of training samples, the model may face issues with generalization. To address this challenge, we designed a memory module to enhance the model's generalization ability. The specific design of the module is as follows:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14279047%2F9aa4594fd24c73286676967216fbe2ba%2Fmemorymodule.png?generation=1740544995457660&amp;alt=media\" alt=\"Mem Block\"><br>\nThis module was added to the bottom layer of the UNet model to enhance the generalization ability of high-dimensional vectors. </p>\n<p>The final models used are as follows:</p>\n<table>\n<thead>\n<tr>\n<th>model name</th>\n<th>nums</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>MemUNet</td>\n<td>3</td>\n</tr>\n<tr>\n<td>UNet</td>\n<td>4</td>\n</tr>\n<tr>\n<td>DLinkNet</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>\n<p>Parameter Settings:</p>\n<pre><code>    ：(, , , )\n    ：(, , )\n    ：\n    ：.\n     Vector Size：(,)\n</code></pre>\n<h1><strong>Training</strong></h1>\n<p>We scaled the original data's radius to 0.48 or 0.5 (thanks to <a href=\"https://www.kaggle.com/miyamotodaiya\" target=\"_blank\">@miyamotodaiya</a> for the experiment). After testing various training sizes (such as 96, 128, 136, 144, 164, and 176), we found that sizes 128 and 164 yielded the best results. Additionally, <a href=\"https://www.kaggle.com/yingpengchen\" target=\"_blank\">@yingpengchen</a> tested different xyz size combinations, and the (48, 256, 256) size performed best. For the output channels, we tried 6, 7, and 8 output channels, with 6 channels providing the best performance.</p>\n<h2><strong>Data Augmentation:</strong></h2>\n<p>We used the following data augmentation methods:</p>\n<ul>\n<li>RanRandCropByLabelClassesd</li>\n<li>RandFlipd</li>\n<li>RandRotated</li>\n<li>RandAffined</li>\n<li>RandGridDistortiond</li>\n<li>RandCoarseDropoutd</li>\n<li>RandScaleIntensityd</li>\n<li>RandShiftIntensityd<br>\nWe set the probability of rotation and flipping to 1 to ensure data diversity.</li>\n</ul>\n<h2><strong>Optimizer:</strong></h2>\n<p>We used the schedulefree.AdamWScheduleFree optimizer introduced by <a href=\"https://www.kaggle.com/miyamotodaiya\" target=\"_blank\">@miyamotodaiya</a> .</p>\n<h2><strong>Loss Function:</strong></h2>\n<ul>\n<li><strong>Weighted Tversky Loss</strong></li>\n<li><strong>Distance Loss</strong> : This loss function performs MSE loss after applying a distance transformation to the ground truth labels, making the model focus more on the central region of the labels. This loss function performs especially well when label overlap occurs.</li>\n</ul>\n<h2><strong>EMA:</strong></h2>\n<p>We employed a method of dynamically adjusting the decay parameter based on the comparison between the current model's score and the best score. This approach yielded an improvement of approximately 0.001 in local tests.</p>\n<h1><strong>Inference</strong></h1>\n<p>For inference, we used a sliding window strategy. Models trained at a size of 96 were used for inference at a size of 128 (since the DLinkNet model is too large to infer at a larger size), while other models were used for inference at sizes 176 or 180. The overlap was set to either 0.15 or 0.5.</p>\n<h2><strong>Inference Time Optimization:</strong></h2>\n<p>We adopted two inference strategies:</p>\n<ol>\n<li><p>Multi-Model Sliding Window Inference:<br>\nThis approach combines multiple models with a sliding window for inference. The final score for this strategy was lb 763. We distributed the models across two GPUs, used multi-processing and TensorRT acceleration, and ran multiple models in parallel. With 7 models, the inference was completed in about 4 hours.</p></li>\n<li><p>Fewer Models with Sliding Window and Extensive TTA:<br>\nThis strategy used fewer models combined with a sliding window and extensive test-time augmentation (TTA), such as flipping and rotating. The final score for this strategy was lb 756. We loaded all models onto two GPUs, split the data into two parts, and used multi-processing and TensorRT acceleration to infer two datasets simultaneously. With 1 model and 7 TTA methods, the inference was completed in about 5 hours.</p></li>\n</ol>\n<p>Finally, using the DataFrame fusion strategy provided by <a href=\"https://www.kaggle.com/miyamotodaiya\" target=\"_blank\">@miyamotodaiya</a> , we merged the results from the two strategies, achieving a final score of lb 768.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14279047%2F6547b0b959fe25ffb55e39f2d11b9294%2Fpipline.png?generation=1740545064775410&amp;alt=media\" alt=\"pipline\"></p>\n<h1><strong>Code</strong></h1>\n<p>Train Code: <a href=\"https://github.com/wplll/Memory-Enhanced-3D-Segmentation\" target=\"_blank\">https://github.com/wplll/Memory-Enhanced-3D-Segmentation</a></p>\n<p>Multi-Model: <a href=\"https://www.kaggle.com/code/peilwang/czii-infer-multi-model\" target=\"_blank\">https://www.kaggle.com/code/peilwang/czii-infer-multi-model</a></p>\n<p>Fewer Models and Extensive TTA: <a href=\"https://www.kaggle.com/code/peilwang/czii-infer-fewer-models-and-extensive-tta\" target=\"_blank\">https://www.kaggle.com/code/peilwang/czii-infer-fewer-models-and-extensive-tta</a></p>\n<p>A big thank you to my teammates for their hard work and support!</p>",
  "messages": [
    {
      "id": 3117750,
      "postDate": "2025-02-07T08:04:51.010Z",
      "content": "<h1><strong>Acknowledgements</strong></h1>\n<p>We sincerely appreciate Kaggle and the competition organizers for offering this invaluable opportunity. We also extend our gratitude to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , <a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a> and <a href=\"https://www.kaggle.com/sjtuwangshuo\" target=\"_blank\">@sjtuwangshuo</a> for their significant contributions. Lastly, I would like to sincerely thank my teammates for their dedication and hard work during this time! <a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a> , <a href=\"https://www.kaggle.com/miyamotodaiya\" target=\"_blank\">@miyamotodaiya</a> and <a href=\"https://www.kaggle.com/yingpengchen\" target=\"_blank\">@yingpengchen</a> </p>\n<h1><strong>Model</strong></h1>\n<p>We used the basic UNet3D model provided by MONAI as the primary model, and we implemented DLinkNet3D as an auxiliary model using Torch. Given the limited number of training samples, the model may face issues with generalization. To address this challenge, we designed a memory module to enhance the model's generalization ability. The specific design of the module is as follows:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14279047%2F9aa4594fd24c73286676967216fbe2ba%2Fmemorymodule.png?generation=1740544995457660&amp;alt=media\" alt=\"Mem Block\"><br>\nThis module was added to the bottom layer of the UNet model to enhance the generalization ability of high-dimensional vectors. </p>\n<p>The final models used are as follows:</p>\n<table>\n<thead>\n<tr>\n<th>model name</th>\n<th>nums</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>MemUNet</td>\n<td>3</td>\n</tr>\n<tr>\n<td>UNet</td>\n<td>4</td>\n</tr>\n<tr>\n<td>DLinkNet</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>\n<p>Parameter Settings:</p>\n<pre><code>    ：(, , , )\n    ：(, , )\n    ：\n    ：.\n     Vector Size：(,)\n</code></pre>\n<h1><strong>Training</strong></h1>\n<p>We scaled the original data's radius to 0.48 or 0.5 (thanks to <a href=\"https://www.kaggle.com/miyamotodaiya\" target=\"_blank\">@miyamotodaiya</a> for the experiment). After testing various training sizes (such as 96, 128, 136, 144, 164, and 176), we found that sizes 128 and 164 yielded the best results. Additionally, <a href=\"https://www.kaggle.com/yingpengchen\" target=\"_blank\">@yingpengchen</a> tested different xyz size combinations, and the (48, 256, 256) size performed best. For the output channels, we tried 6, 7, and 8 output channels, with 6 channels providing the best performance.</p>\n<h2><strong>Data Augmentation:</strong></h2>\n<p>We used the following data augmentation methods:</p>\n<ul>\n<li>RanRandCropByLabelClassesd</li>\n<li>RandFlipd</li>\n<li>RandRotated</li>\n<li>RandAffined</li>\n<li>RandGridDistortiond</li>\n<li>RandCoarseDropoutd</li>\n<li>RandScaleIntensityd</li>\n<li>RandShiftIntensityd<br>\nWe set the probability of rotation and flipping to 1 to ensure data diversity.</li>\n</ul>\n<h2><strong>Optimizer:</strong></h2>\n<p>We used the schedulefree.AdamWScheduleFree optimizer introduced by <a href=\"https://www.kaggle.com/miyamotodaiya\" target=\"_blank\">@miyamotodaiya</a> .</p>\n<h2><strong>Loss Function:</strong></h2>\n<ul>\n<li><strong>Weighted Tversky Loss</strong></li>\n<li><strong>Distance Loss</strong> : This loss function performs MSE loss after applying a distance transformation to the ground truth labels, making the model focus more on the central region of the labels. This loss function performs especially well when label overlap occurs.</li>\n</ul>\n<h2><strong>EMA:</strong></h2>\n<p>We employed a method of dynamically adjusting the decay parameter based on the comparison between the current model's score and the best score. This approach yielded an improvement of approximately 0.001 in local tests.</p>\n<h1><strong>Inference</strong></h1>\n<p>For inference, we used a sliding window strategy. Models trained at a size of 96 were used for inference at a size of 128 (since the DLinkNet model is too large to infer at a larger size), while other models were used for inference at sizes 176 or 180. The overlap was set to either 0.15 or 0.5.</p>\n<h2><strong>Inference Time Optimization:</strong></h2>\n<p>We adopted two inference strategies:</p>\n<ol>\n<li><p>Multi-Model Sliding Window Inference:<br>\nThis approach combines multiple models with a sliding window for inference. The final score for this strategy was lb 763. We distributed the models across two GPUs, used multi-processing and TensorRT acceleration, and ran multiple models in parallel. With 7 models, the inference was completed in about 4 hours.</p></li>\n<li><p>Fewer Models with Sliding Window and Extensive TTA:<br>\nThis strategy used fewer models combined with a sliding window and extensive test-time augmentation (TTA), such as flipping and rotating. The final score for this strategy was lb 756. We loaded all models onto two GPUs, split the data into two parts, and used multi-processing and TensorRT acceleration to infer two datasets simultaneously. With 1 model and 7 TTA methods, the inference was completed in about 5 hours.</p></li>\n</ol>\n<p>Finally, using the DataFrame fusion strategy provided by <a href=\"https://www.kaggle.com/miyamotodaiya\" target=\"_blank\">@miyamotodaiya</a> , we merged the results from the two strategies, achieving a final score of lb 768.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14279047%2F6547b0b959fe25ffb55e39f2d11b9294%2Fpipline.png?generation=1740545064775410&amp;alt=media\" alt=\"pipline\"></p>\n<h1><strong>Code</strong></h1>\n<p>Train Code: <a href=\"https://github.com/wplll/Memory-Enhanced-3D-Segmentation\" target=\"_blank\">https://github.com/wplll/Memory-Enhanced-3D-Segmentation</a></p>\n<p>Multi-Model: <a href=\"https://www.kaggle.com/code/peilwang/czii-infer-multi-model\" target=\"_blank\">https://www.kaggle.com/code/peilwang/czii-infer-multi-model</a></p>\n<p>Fewer Models and Extensive TTA: <a href=\"https://www.kaggle.com/code/peilwang/czii-infer-fewer-models-and-extensive-tta\" target=\"_blank\">https://www.kaggle.com/code/peilwang/czii-infer-fewer-models-and-extensive-tta</a></p>\n<p>A big thank you to my teammates for their hard work and support!</p>",
      "rawMarkdown": "# **Acknowledgements**\nWe sincerely appreciate Kaggle and the competition organizers for offering this invaluable opportunity. We also extend our gratitude to @hengck23 , @fnands and @sjtuwangshuo for their significant contributions. Lastly, I would like to sincerely thank my teammates for their dedication and hard work during this time! @snnclsr , @miyamotodaiya and @yingpengchen \n\n\n# **Model**\nWe used the basic UNet3D model provided by MONAI as the primary model, and we implemented DLinkNet3D as an auxiliary model using Torch. Given the limited number of training samples, the model may face issues with generalization. To address this challenge, we designed a memory module to enhance the model's generalization ability. The specific design of the module is as follows:\n![Mem Block](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14279047%2F9aa4594fd24c73286676967216fbe2ba%2Fmemorymodule.png?generation=1740544995457660&alt=media)\nThis module was added to the bottom layer of the UNet model to enhance the generalization ability of high-dimensional vectors. \n\nThe final models used are as follows:\n|model name|nums|\n|:--------:|:--------:|\n|MemUNet|3|\n|UNet|4|\n|DLinkNet|1|\n\nParameter Settings:\n\n        Channels：(48, 64, 80, 80)\n        Strides：(2, 2, 1)\n        num_res_units：2\n        Dropout：0.3\n        Mem Vector Size：(256,256)\n\n\n# **Training**\nWe scaled the original data's radius to 0.48 or 0.5 (thanks to @miyamotodaiya for the experiment). After testing various training sizes (such as 96, 128, 136, 144, 164, and 176), we found that sizes 128 and 164 yielded the best results. Additionally, @yingpengchen tested different xyz size combinations, and the (48, 256, 256) size performed best. For the output channels, we tried 6, 7, and 8 output channels, with 6 channels providing the best performance.\n\n## **Data Augmentation:**\nWe used the following data augmentation methods:\n- RanRandCropByLabelClassesd\n- RandFlipd\n- RandRotated\n- RandAffined\n- RandGridDistortiond\n- RandCoarseDropoutd\n- RandScaleIntensityd\n- RandShiftIntensityd\nWe set the probability of rotation and flipping to 1 to ensure data diversity.\n\n## **Optimizer:**\nWe used the schedulefree.AdamWScheduleFree optimizer introduced by @miyamotodaiya .\n\n## **Loss Function:**\n- **Weighted Tversky Loss**\n- **Distance Loss** : This loss function performs MSE loss after applying a distance transformation to the ground truth labels, making the model focus more on the central region of the labels. This loss function performs especially well when label overlap occurs.\n\n## **EMA:**\nWe employed a method of dynamically adjusting the decay parameter based on the comparison between the current model's score and the best score. This approach yielded an improvement of approximately 0.001 in local tests.\n\n# **Inference**\nFor inference, we used a sliding window strategy. Models trained at a size of 96 were used for inference at a size of 128 (since the DLinkNet model is too large to infer at a larger size), while other models were used for inference at sizes 176 or 180. The overlap was set to either 0.15 or 0.5.\n\n## **Inference Time Optimization:**\nWe adopted two inference strategies:\n1. Multi-Model Sliding Window Inference:\nThis approach combines multiple models with a sliding window for inference. The final score for this strategy was lb 763. We distributed the models across two GPUs, used multi-processing and TensorRT acceleration, and ran multiple models in parallel. With 7 models, the inference was completed in about 4 hours.\n\n2. Fewer Models with Sliding Window and Extensive TTA:\nThis strategy used fewer models combined with a sliding window and extensive test-time augmentation (TTA), such as flipping and rotating. The final score for this strategy was lb 756. We loaded all models onto two GPUs, split the data into two parts, and used multi-processing and TensorRT acceleration to infer two datasets simultaneously. With 1 model and 7 TTA methods, the inference was completed in about 5 hours.\n\nFinally, using the DataFrame fusion strategy provided by @miyamotodaiya , we merged the results from the two strategies, achieving a final score of lb 768.\n![pipline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14279047%2F6547b0b959fe25ffb55e39f2d11b9294%2Fpipline.png?generation=1740545064775410&alt=media)\n# **Code**\n\nTrain Code: [https://github.com/wplll/Memory-Enhanced-3D-Segmentation](https://github.com/wplll/Memory-Enhanced-3D-Segmentation)\n\nMulti-Model: [https://www.kaggle.com/code/peilwang/czii-infer-multi-model](https://www.kaggle.com/code/peilwang/czii-infer-multi-model)\n\nFewer Models and Extensive TTA: [https://www.kaggle.com/code/peilwang/czii-infer-fewer-models-and-extensive-tta](https://www.kaggle.com/code/peilwang/czii-infer-fewer-models-and-extensive-tta)\n\nA big thank you to my teammates for their hard work and support!\n",
      "votes": 20
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3117750": "# **Acknowledgements**\nWe sincerely appreciate Kaggle and the competition organizers for offering this invaluable opportunity. We also extend our gratitude to @hengck23 , @fnands and @sjtuwangshuo for their significant contributions. Lastly, I would like to sincerely thank my teammates for their dedication and hard work during this time! @snnclsr , @miyamotodaiya and @yingpengchen \n\n\n# **Model**\nWe used the basic UNet3D model provided by MONAI as the primary model, and we implemented DLinkNet3D as an auxiliary model using Torch. Given the limited number of training samples, the model may face issues with generalization. To address this challenge, we designed a memory module to enhance the model's generalization ability. The specific design of the module is as follows:\n![Mem Block](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14279047%2F9aa4594fd24c73286676967216fbe2ba%2Fmemorymodule.png?generation=1740544995457660&alt=media)\nThis module was added to the bottom layer of the UNet model to enhance the generalization ability of high-dimensional vectors. \n\nThe final models used are as follows:\n|model name|nums|\n|:--------:|:--------:|\n|MemUNet|3|\n|UNet|4|\n|DLinkNet|1|\n\nParameter Settings:\n\n        Channels：(48, 64, 80, 80)\n        Strides：(2, 2, 1)\n        num_res_units：2\n        Dropout：0.3\n        Mem Vector Size：(256,256)\n\n\n# **Training**\nWe scaled the original data's radius to 0.48 or 0.5 (thanks to @miyamotodaiya for the experiment). After testing various training sizes (such as 96, 128, 136, 144, 164, and 176), we found that sizes 128 and 164 yielded the best results. Additionally, @yingpengchen tested different xyz size combinations, and the (48, 256, 256) size performed best. For the output channels, we tried 6, 7, and 8 output channels, with 6 channels providing the best performance.\n\n## **Data Augmentation:**\nWe used the following data augmentation methods:\n- RanRandCropByLabelClassesd\n- RandFlipd\n- RandRotated\n- RandAffined\n- RandGridDistortiond\n- RandCoarseDropoutd\n- RandScaleIntensityd\n- RandShiftIntensityd\nWe set the probability of rotation and flipping to 1 to ensure data diversity.\n\n## **Optimizer:**\nWe used the schedulefree.AdamWScheduleFree optimizer introduced by @miyamotodaiya .\n\n## **Loss Function:**\n- **Weighted Tversky Loss**\n- **Distance Loss** : This loss function performs MSE loss after applying a distance transformation to the ground truth labels, making the model focus more on the central region of the labels. This loss function performs especially well when label overlap occurs.\n\n## **EMA:**\nWe employed a method of dynamically adjusting the decay parameter based on the comparison between the current model's score and the best score. This approach yielded an improvement of approximately 0.001 in local tests.\n\n# **Inference**\nFor inference, we used a sliding window strategy. Models trained at a size of 96 were used for inference at a size of 128 (since the DLinkNet model is too large to infer at a larger size), while other models were used for inference at sizes 176 or 180. The overlap was set to either 0.15 or 0.5.\n\n## **Inference Time Optimization:**\nWe adopted two inference strategies:\n1. Multi-Model Sliding Window Inference:\nThis approach combines multiple models with a sliding window for inference. The final score for this strategy was lb 763. We distributed the models across two GPUs, used multi-processing and TensorRT acceleration, and ran multiple models in parallel. With 7 models, the inference was completed in about 4 hours.\n\n2. Fewer Models with Sliding Window and Extensive TTA:\nThis strategy used fewer models combined with a sliding window and extensive test-time augmentation (TTA), such as flipping and rotating. The final score for this strategy was lb 756. We loaded all models onto two GPUs, split the data into two parts, and used multi-processing and TensorRT acceleration to infer two datasets simultaneously. With 1 model and 7 TTA methods, the inference was completed in about 5 hours.\n\nFinally, using the DataFrame fusion strategy provided by @miyamotodaiya , we merged the results from the two strategies, achieving a final score of lb 768.\n![pipline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14279047%2F6547b0b959fe25ffb55e39f2d11b9294%2Fpipline.png?generation=1740545064775410&alt=media)\n# **Code**\n\nTrain Code: [https://github.com/wplll/Memory-Enhanced-3D-Segmentation](https://github.com/wplll/Memory-Enhanced-3D-Segmentation)\n\nMulti-Model: [https://www.kaggle.com/code/peilwang/czii-infer-multi-model](https://www.kaggle.com/code/peilwang/czii-infer-multi-model)\n\nFewer Models and Extensive TTA: [https://www.kaggle.com/code/peilwang/czii-infer-fewer-models-and-extensive-tta](https://www.kaggle.com/code/peilwang/czii-infer-fewer-models-and-extensive-tta)\n\nA big thank you to my teammates for their hard work and support!\n"
  }
}