{
  "id": 561422,
  "title": "13th Place Solution - w/ Cutpaste",
  "url": "/competitions/czii-cryo-et-object-identification/writeups/bartley-13th-place-solution-w-cutpaste",
  "author_name": "",
  "post_date": "2025-02-07T00:00:23.323Z",
  "votes": 23,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Thanks to CZII and Kaggle for hosting this competition, it was a fun one and I learned a lot. Also thanks to the hosts for actively contributing to the forums, this was very helpful for participants.</p>\n<p>TLDR; My solution uses 3D unet models only. Each model uses a pre-trained resnet3d encoder from <a href=\"https://github.com/kenshohara/3D-ResNets-PyTorch\" target=\"_blank\">3D-ResNets-PyTorch</a>, and a decoder with pixel shuffle upsample blocks. I apply EMA, heavy augmentations, and other regularization strategies to improve generalization.</p>\n<h2>Model</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fe214dfda38e641841d6c9b6b2655837c%2Fmodel.jpg?generation=1738811467298781&amp;alt=media\" alt=\"ModelArch\" style=\"max-width: 75%\"></p>\n<p>My final submission uses 2x r3d50, 2x r3d34 and 1x r3d18 backbones. The deeper encoders perform better on CV and LB, but an ensemble of all three works best. The encoders are improved by adding <a href=\"https://arxiv.org/abs/1603.09382\" target=\"_blank\">Stochastic Depth</a> and <a href=\"https://arxiv.org/abs/1810.12890\" target=\"_blank\">DropBlock</a>.</p>\n<p>In the decoder, I use standard pixel shuffle in all layers except the deepest layer. In the deepest layer, I replace the conv3d in pixel shuffle with a depth-wise separable 3D convolution. This significantly reduces the parameter count of the decoder with no drop in performance. Applying this modification to shallower blocks did reduce performance.</p>\n<p>I use 3 auxiliary segmentation heads to regularize the intermediate feature maps. The heads are trained on the max pooled segmentation masks rather than interpolated masks. This was done to force more “confidence” into the intermediate features.</p>\n<p>Each model is trained in 2 stages. In the first stage, I freeze the encoder and train the decoder on the simulated data. In the second stage, I unfreeze all parameters and train on the competition data.</p>\n<h2>Augs</h2>\n<p>Heavy augmentations are very important in this pipeline. I apply cut mix, rotation, flips, pixel intensity, and pixel shift to each batch with 100% probability. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fc674a9687cf0bf3d365cce1769ff432f%2Fcutpaste.jpg?generation=1738813304410522&amp;alt=media\" alt=\"ModelArch\" style=\"max-width: 75%\"></p>\n<p>Another augmentation I use is cut paste. For this augmentation, I crop around each picked particle center and randomly insert them into patches with no picked particles. It works well to apply this augmentation within the same volume, but it was not beneficial to cut paste between volumes.</p>\n<h2>Inference</h2>\n<p>During inference, I use Monai’s <a href=\"https://docs.monai.io/en/latest/inferers.html#sliding-window-inference-function\" target=\"_blank\">sliding_window_inference</a> function to iterate patches. Based on the visualization of OOF predictions locally, I found that a high degree of overlap (eg. 0.5+) helped reduce uncertainty on the edge of patches. Moreover, on a (32, 128, 128) patch I only made predictions on the central (16, 64, 64) pixels. This increased the inference time significantly but improved CV and LB by ~0.005.</p>\n<h2>Final Note</h2>\n<p>On the final day of the competition, I realized that my segmentation masks and coordinate conversion code were incorrect! With no time to retrain any models, my final submissions attempted to correct this with postprocessing. These submissions were my highest-scoring, and I learned a valuable lesson for next time 🙂</p>\n<p>Thanks to everyone for sharing throughout. Happy Kaggling!</p>",
  "messages": [
    {
      "id": "3116522",
      "postDate": "02/06/2025 03:43:00",
      "content": "<p>Thanks to CZII and Kaggle for hosting this competition, it was a fun one and I learned a lot. Also thanks to the hosts for actively contributing to the forums, this was very helpful for participants.</p>\n<p>TLDR; My solution uses 3D unet models only. Each model uses a pre-trained resnet3d encoder from <a href=\"https://github.com/kenshohara/3D-ResNets-PyTorch\" target=\"_blank\">3D-ResNets-PyTorch</a>, and a decoder with pixel shuffle upsample blocks. I apply EMA, heavy augmentations, and other regularization strategies to improve generalization.</p>\n<h2>Model</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fe214dfda38e641841d6c9b6b2655837c%2Fmodel.jpg?generation=1738811467298781&amp;alt=media\" alt=\"ModelArch\" style=\"max-width: 75%\"></p>\n<p>My final submission uses 2x r3d50, 2x r3d34 and 1x r3d18 backbones. The deeper encoders perform better on CV and LB, but an ensemble of all three works best. The encoders are improved by adding <a href=\"https://arxiv.org/abs/1603.09382\" target=\"_blank\">Stochastic Depth</a> and <a href=\"https://arxiv.org/abs/1810.12890\" target=\"_blank\">DropBlock</a>.</p>\n<p>In the decoder, I use standard pixel shuffle in all layers except the deepest layer. In the deepest layer, I replace the conv3d in pixel shuffle with a depth-wise separable 3D convolution. This significantly reduces the parameter count of the decoder with no drop in performance. Applying this modification to shallower blocks did reduce performance.</p>\n<p>I use 3 auxiliary segmentation heads to regularize the intermediate feature maps. The heads are trained on the max pooled segmentation masks rather than interpolated masks. This was done to force more “confidence” into the intermediate features.</p>\n<p>Each model is trained in 2 stages. In the first stage, I freeze the encoder and train the decoder on the simulated data. In the second stage, I unfreeze all parameters and train on the competition data.</p>\n<h2>Augs</h2>\n<p>Heavy augmentations are very important in this pipeline. I apply cut mix, rotation, flips, pixel intensity, and pixel shift to each batch with 100% probability. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fc674a9687cf0bf3d365cce1769ff432f%2Fcutpaste.jpg?generation=1738813304410522&amp;alt=media\" alt=\"ModelArch\" style=\"max-width: 75%\"></p>\n<p>Another augmentation I use is cut paste. For this augmentation, I crop around each picked particle center and randomly insert them into patches with no picked particles. It works well to apply this augmentation within the same volume, but it was not beneficial to cut paste between volumes.</p>\n<h2>Inference</h2>\n<p>During inference, I use Monai’s <a href=\"https://docs.monai.io/en/latest/inferers.html#sliding-window-inference-function\" target=\"_blank\">sliding_window_inference</a> function to iterate patches. Based on the visualization of OOF predictions locally, I found that a high degree of overlap (eg. 0.5+) helped reduce uncertainty on the edge of patches. Moreover, on a (32, 128, 128) patch I only made predictions on the central (16, 64, 64) pixels. This increased the inference time significantly but improved CV and LB by ~0.005.</p>\n<h2>Final Note</h2>\n<p>On the final day of the competition, I realized that my segmentation masks and coordinate conversion code were incorrect! With no time to retrain any models, my final submissions attempted to correct this with postprocessing. These submissions were my highest-scoring, and I learned a valuable lesson for next time 🙂</p>\n<p>Thanks to everyone for sharing throughout. Happy Kaggling!</p>",
      "rawMarkdown": "Thanks to CZII and Kaggle for hosting this competition, it was a fun one and I learned a lot. Also thanks to the hosts for actively contributing to the forums, this was very helpful for participants.\n\nTLDR; My solution uses 3D unet models only. Each model uses a pre-trained resnet3d encoder from [3D-ResNets-PyTorch](https://github.com/kenshohara/3D-ResNets-PyTorch), and a decoder with pixel shuffle upsample blocks. I apply EMA, heavy augmentations, and other regularization strategies to improve generalization.\n\n## Model\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fe214dfda38e641841d6c9b6b2655837c%2Fmodel.jpg?generation=1738811467298781&alt=media\" alt=\"ModelArch\" style=\"max-width: 75%;\">\n\nMy final submission uses 2x r3d50, 2x r3d34 and 1x r3d18 backbones. The deeper encoders perform better on CV and LB, but an ensemble of all three works best. The encoders are improved by adding [Stochastic Depth](https://arxiv.org/abs/1603.09382) and [DropBlock](https://arxiv.org/abs/1810.12890).\n\nIn the decoder, I use standard pixel shuffle in all layers except the deepest layer. In the deepest layer, I replace the conv3d in pixel shuffle with a depth-wise separable 3D convolution. This significantly reduces the parameter count of the decoder with no drop in performance. Applying this modification to shallower blocks did reduce performance.\n\nI use 3 auxiliary segmentation heads to regularize the intermediate feature maps. The heads are trained on the max pooled segmentation masks rather than interpolated masks. This was done to force more “confidence” into the intermediate features.\n\nEach model is trained in 2 stages. In the first stage, I freeze the encoder and train the decoder on the simulated data. In the second stage, I unfreeze all parameters and train on the competition data.\n\n## Augs\n\nHeavy augmentations are very important in this pipeline. I apply cut mix, rotation, flips, pixel intensity, and pixel shift to each batch with 100% probability. \n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fc674a9687cf0bf3d365cce1769ff432f%2Fcutpaste.jpg?generation=1738813304410522&alt=media\" alt=\"ModelArch\" style=\"max-width: 75%;\">\n\nAnother augmentation I use is cut paste. For this augmentation, I crop around each picked particle center and randomly insert them into patches with no picked particles. It works well to apply this augmentation within the same volume, but it was not beneficial to cut paste between volumes.\n\n## Inference\n\nDuring inference, I use Monai’s [sliding_window_inference](https://docs.monai.io/en/latest/inferers.html#sliding-window-inference-function) function to iterate patches. Based on the visualization of OOF predictions locally, I found that a high degree of overlap (eg. 0.5+) helped reduce uncertainty on the edge of patches. Moreover, on a (32, 128, 128) patch I only made predictions on the central (16, 64, 64) pixels. This increased the inference time significantly but improved CV and LB by ~0.005.\n\n## Final Note\n\nOn the final day of the competition, I realized that my segmentation masks and coordinate conversion code were incorrect! With no time to retrain any models, my final submissions attempted to correct this with postprocessing. These submissions were my highest-scoring, and I learned a valuable lesson for next time 🙂\n\nThanks to everyone for sharing throughout. Happy Kaggling!",
      "votes": null
    },
    {
      "id": "3116753",
      "postDate": "02/06/2025 09:15:59",
      "content": "<p>That final note was harsh :D Congratulations on your strong finish and for the detailed write-up! Are you planning to release your code as well?</p>",
      "rawMarkdown": "That final note was harsh :D Congratulations on your strong finish and for the detailed write-up! Are you planning to release your code as well?",
      "votes": null
    },
    {
      "id": "3116940",
      "postDate": "02/06/2025 12:56:54",
      "content": "<p>Congratulations! Thanks for share.</p>\n<p>About cutpaste I've tried almost the same but with spherical cuts to freely rotate each particle. In my case this directly pointed particles so the model learned to localize this kind of processing instead of the particle. What do you think has been my biggest mistake with that? May be was more about differences with freely rotated pixels than some kind of water mark at the edges (that's what I've though).</p>\n<p>\"I use 3 auxiliary segmentation heads to regularize the intermediate feature maps. The heads are trained on the max pooled segmentation masks rather than interpolated masks. This was done to force more “confidence” into the intermediate features.\" Nice one!</p>\n<p>\"I use standard pixel shuffle\" I'll search about that but could you please elaborate it? Is <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.PixelShuffle.html\" target=\"_blank\">this</a>? Thanks. </p>",
      "rawMarkdown": "Congratulations! Thanks for share.\n\nAbout cutpaste I've tried almost the same but with spherical cuts to freely rotate each particle. In my case this directly pointed particles so the model learned to localize this kind of processing instead of the particle. What do you think has been my biggest mistake with that? May be was more about differences with freely rotated pixels than some kind of water mark at the edges (that's what I've though).\n\n\"I use 3 auxiliary segmentation heads to regularize the intermediate feature maps. The heads are trained on the max pooled segmentation masks rather than interpolated masks. This was done to force more “confidence” into the intermediate features.\" Nice one!\n\n\"I use standard pixel shuffle\" I'll search about that but could you please elaborate it? Is [this](https://pytorch.org/docs/stable/generated/torch.nn.PixelShuffle.html)? Thanks.",
      "votes": null
    },
    {
      "id": "3116989",
      "postDate": "02/06/2025 14:05:34",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sacuscreed\" target=\"_blank\">@sacuscreed</a>, thanks for the comment.</p>\n<p>Here are a few things that were important for cutpaste. a) patches with no other picked particles. b) within the same volume. c) heavy cutmix. To add on to your point, I think these changes made it less obvious where the inserted particles were. Interestingly, by training on 100% background-only images with particles inserted using cutpaste, I was able to achieve 0.770/0.768 CV/LB.</p>\n<p>For pixel shuffle, I used Monai's class from <a href=\"https://docs.monai.io/en/stable/networks.html#monai.networks.blocks.SubpixelUpsample\" target=\"_blank\">here</a>. </p>\n<p>Hope this helps.</p>",
      "rawMarkdown": "Hi @sacuscreed, thanks for the comment.\n\nHere are a few things that were important for cutpaste. a) patches with no other picked particles. b) within the same volume. c) heavy cutmix. To add on to your point, I think these changes made it less obvious where the inserted particles were. Interestingly, by training on 100% background-only images with particles inserted using cutpaste, I was able to achieve 0.770/0.768 CV/LB.\n\nFor pixel shuffle, I used Monai's class from [here](https://docs.monai.io/en/stable/networks.html#monai.networks.blocks.SubpixelUpsample). \n\nHope this helps.",
      "votes": null
    },
    {
      "id": "3116991",
      "postDate": "02/06/2025 14:11:53",
      "content": "<p></p>\n<p>Thanks <a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a>. Code is <a href=\"https://github.com/brendanartley/CZII-Competition\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "~~Sure, I will update this comment when complete!~~\n\nThanks @snnclsr. Code is [here](https://github.com/brendanartley/CZII-Competition).",
      "votes": null
    },
    {
      "id": "3117101",
      "postDate": "02/06/2025 16:26:34",
      "content": "<p>Awesome!! Thank you so much!</p>",
      "rawMarkdown": "Awesome!! Thank you so much!",
      "votes": null
    },
    {
      "id": "3185318",
      "postDate": "04/23/2025 06:39:18",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>!</p>",
      "rawMarkdown": "Thanks @brendanartley!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3116753,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "02/06/2025 09:15:59",
      "content": "<p>That final note was harsh :D Congratulations on your strong finish and for the detailed write-up! Are you planning to release your code as well?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3116991,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "02/06/2025 14:11:53",
          "content": "<p></p>\n<p>Thanks <a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a>. Code is <a href=\"https://github.com/brendanartley/CZII-Competition\" target=\"_blank\">here</a>.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3117101,
              "author_name": "snnclsr",
              "author_url": "",
              "post_date": "02/06/2025 16:26:34",
              "content": "<p>Awesome!! Thank you so much!</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3185318,
              "author_name": "davidlist",
              "author_url": "",
              "post_date": "04/23/2025 06:39:18",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3116940,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "02/06/2025 12:56:54",
      "content": "<p>Congratulations! Thanks for share.</p>\n<p>About cutpaste I've tried almost the same but with spherical cuts to freely rotate each particle. In my case this directly pointed particles so the model learned to localize this kind of processing instead of the particle. What do you think has been my biggest mistake with that? May be was more about differences with freely rotated pixels than some kind of water mark at the edges (that's what I've though).</p>\n<p>\"I use 3 auxiliary segmentation heads to regularize the intermediate feature maps. The heads are trained on the max pooled segmentation masks rather than interpolated masks. This was done to force more “confidence” into the intermediate features.\" Nice one!</p>\n<p>\"I use standard pixel shuffle\" I'll search about that but could you please elaborate it? Is <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.PixelShuffle.html\" target=\"_blank\">this</a>? Thanks. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3116989,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "02/06/2025 14:05:34",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sacuscreed\" target=\"_blank\">@sacuscreed</a>, thanks for the comment.</p>\n<p>Here are a few things that were important for cutpaste. a) patches with no other picked particles. b) within the same volume. c) heavy cutmix. To add on to your point, I think these changes made it less obvious where the inserted particles were. Interestingly, by training on 100% background-only images with particles inserted using cutpaste, I was able to achieve 0.770/0.768 CV/LB.</p>\n<p>For pixel shuffle, I used Monai's class from <a href=\"https://docs.monai.io/en/stable/networks.html#monai.networks.blocks.SubpixelUpsample\" target=\"_blank\">here</a>. </p>\n<p>Hope this helps.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3116522": "Thanks to CZII and Kaggle for hosting this competition, it was a fun one and I learned a lot. Also thanks to the hosts for actively contributing to the forums, this was very helpful for participants.\n\nTLDR; My solution uses 3D unet models only. Each model uses a pre-trained resnet3d encoder from [3D-ResNets-PyTorch](https://github.com/kenshohara/3D-ResNets-PyTorch), and a decoder with pixel shuffle upsample blocks. I apply EMA, heavy augmentations, and other regularization strategies to improve generalization.\n\n## Model\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fe214dfda38e641841d6c9b6b2655837c%2Fmodel.jpg?generation=1738811467298781&alt=media\" alt=\"ModelArch\" style=\"max-width: 75%;\">\n\nMy final submission uses 2x r3d50, 2x r3d34 and 1x r3d18 backbones. The deeper encoders perform better on CV and LB, but an ensemble of all three works best. The encoders are improved by adding [Stochastic Depth](https://arxiv.org/abs/1603.09382) and [DropBlock](https://arxiv.org/abs/1810.12890).\n\nIn the decoder, I use standard pixel shuffle in all layers except the deepest layer. In the deepest layer, I replace the conv3d in pixel shuffle with a depth-wise separable 3D convolution. This significantly reduces the parameter count of the decoder with no drop in performance. Applying this modification to shallower blocks did reduce performance.\n\nI use 3 auxiliary segmentation heads to regularize the intermediate feature maps. The heads are trained on the max pooled segmentation masks rather than interpolated masks. This was done to force more “confidence” into the intermediate features.\n\nEach model is trained in 2 stages. In the first stage, I freeze the encoder and train the decoder on the simulated data. In the second stage, I unfreeze all parameters and train on the competition data.\n\n## Augs\n\nHeavy augmentations are very important in this pipeline. I apply cut mix, rotation, flips, pixel intensity, and pixel shift to each batch with 100% probability. \n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fc674a9687cf0bf3d365cce1769ff432f%2Fcutpaste.jpg?generation=1738813304410522&alt=media\" alt=\"ModelArch\" style=\"max-width: 75%;\">\n\nAnother augmentation I use is cut paste. For this augmentation, I crop around each picked particle center and randomly insert them into patches with no picked particles. It works well to apply this augmentation within the same volume, but it was not beneficial to cut paste between volumes.\n\n## Inference\n\nDuring inference, I use Monai’s [sliding_window_inference](https://docs.monai.io/en/latest/inferers.html#sliding-window-inference-function) function to iterate patches. Based on the visualization of OOF predictions locally, I found that a high degree of overlap (eg. 0.5+) helped reduce uncertainty on the edge of patches. Moreover, on a (32, 128, 128) patch I only made predictions on the central (16, 64, 64) pixels. This increased the inference time significantly but improved CV and LB by ~0.005.\n\n## Final Note\n\nOn the final day of the competition, I realized that my segmentation masks and coordinate conversion code were incorrect! With no time to retrain any models, my final submissions attempted to correct this with postprocessing. These submissions were my highest-scoring, and I learned a valuable lesson for next time 🙂\n\nThanks to everyone for sharing throughout. Happy Kaggling!",
    "3116753": "That final note was harsh :D Congratulations on your strong finish and for the detailed write-up! Are you planning to release your code as well?",
    "3116940": "Congratulations! Thanks for share.\n\nAbout cutpaste I've tried almost the same but with spherical cuts to freely rotate each particle. In my case this directly pointed particles so the model learned to localize this kind of processing instead of the particle. What do you think has been my biggest mistake with that? May be was more about differences with freely rotated pixels than some kind of water mark at the edges (that's what I've though).\n\n\"I use 3 auxiliary segmentation heads to regularize the intermediate feature maps. The heads are trained on the max pooled segmentation masks rather than interpolated masks. This was done to force more “confidence” into the intermediate features.\" Nice one!\n\n\"I use standard pixel shuffle\" I'll search about that but could you please elaborate it? Is [this](https://pytorch.org/docs/stable/generated/torch.nn.PixelShuffle.html)? Thanks.",
    "3116989": "Hi @sacuscreed, thanks for the comment.\n\nHere are a few things that were important for cutpaste. a) patches with no other picked particles. b) within the same volume. c) heavy cutmix. To add on to your point, I think these changes made it less obvious where the inserted particles were. Interestingly, by training on 100% background-only images with particles inserted using cutpaste, I was able to achieve 0.770/0.768 CV/LB.\n\nFor pixel shuffle, I used Monai's class from [here](https://docs.monai.io/en/stable/networks.html#monai.networks.blocks.SubpixelUpsample). \n\nHope this helps.",
    "3116991": "~~Sure, I will update this comment when complete!~~\n\nThanks @snnclsr. Code is [here](https://github.com/brendanartley/CZII-Competition).",
    "3117101": "Awesome!! Thank you so much!",
    "3185318": "Thanks @brendanartley!"
  },
  "source": "meta"
}