{
  "id": 573335,
  "title": "training on 3d unet",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/573335",
  "author_name": "thisArmin",
  "post_date": "2025-04-15T00:57:54.357000",
  "votes": 5,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I understand the yolov8 is training on 2d slices. Has anyone tried 3D unet, where you give the whole data instead of slices to model? </p>",
  "messages": [
    {
      "id": 3179047,
      "postDate": "2025-04-15T00:57:54.357Z",
      "content": "<p>I understand the yolov8 is training on 2d slices. Has anyone tried 3D unet, where you give the whole data instead of slices to model? </p>",
      "rawMarkdown": "I understand the yolov8 is training on 2d slices. Has anyone tried 3D unet, where you give the whole data instead of slices to model? ",
      "votes": 5
    },
    {
      "id": 3180287,
      "postDate": "2025-04-16T11:38:32.413Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/thisarmin\" target=\"_blank\">@thisarmin</a> , I've tried adapting the code I used in the CZII competition, where I trained a 3D U-Net. As others mentioned, feeding the full volume at once isn’t possible because of memory constraints.</p>\n<p>I downsampled the volumes by 0.5 and fed the network with crops of size (128, 128, 128) centered around the particle. I also experimented with (24, 256, 256) crops. My best 3D U-Net model scored 0.436 on the leaderboard.</p>\n<p>Since I noticed many people were using 2D models, I also tried a 2D U-Net, resizing slices to (640, 640) and feeding the model 8 Z-slices around the particle center. It didn’t outperform the 3D U-Net.</p>\n<p>For the targets, I used masks that were simple blobs with a ~6px radius centered on the particles.</p>\n<p>Here’s a the code of the MONAI 3D U-Net I used:</p>\n<pre><code> (pl.LightningModule):\n     ():\n        ().__init__()\n        .save_hyperparameters()\n\n        .model = UNet(\n            spatial_dims=.hparams.spatial_dims,\n            in_channels=.hparams.in_channels,\n            out_channels=.hparams.out_channels,\n            channels=.hparams.channels,\n            strides=.hparams.strides,\n            num_res_units=.hparams.num_res_units,\n            dropout=.hparams.dropout,\n        )\n\n\n</code></pre>\n<p>I'm segmenting, then used connected component analysis and selected the blob with highest confidence score.</p>\n<p>During training, the U-Net learns, the validation plots look good and the model is able to detect the motors with only a few artifacts. But when applying a sliding window over the entire tomogram, the model segmentation has a LOT of artifacts.</p>\n<p>Here are some reference images:</p>\n<p>tomo_00e047 169.0 546.0 603.0</p>\n<p>tomo_00e047 z = 20:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2Fcb78b381622d672831d441163aee3a71%2Ftomo_00e047_slice_20_prediction.png?generation=1744803417908613&amp;alt=media\" alt=\"\"></p>\n<p>tomo_00e047 z = 40:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F27ef58a85da26b7c01c34186b97a9e5a%2Ftomo_00e047_slice_40_prediction.png?generation=1744803228442504&amp;alt=media\" alt=\"\"></p>\n<p>tomo_00e047 z = 170:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F0d3a015a0e111d001b31b5a95605e833%2Ftomo_00e047_slice_170_prediction.png?generation=1744803301669830&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hey @thisarmin , I've tried adapting the code I used in the CZII competition, where I trained a 3D U-Net. As others mentioned, feeding the full volume at once isn’t possible because of memory constraints.\n\nI downsampled the volumes by 0.5 and fed the network with crops of size (128, 128, 128) centered around the particle. I also experimented with (24, 256, 256) crops. My best 3D U-Net model scored 0.436 on the leaderboard.\n\nSince I noticed many people were using 2D models, I also tried a 2D U-Net, resizing slices to (640, 640) and feeding the model 8 Z-slices around the particle center. It didn’t outperform the 3D U-Net.\n\nFor the targets, I used masks that were simple blobs with a ~6px radius centered on the particles.\n\nHere’s a the code of the MONAI 3D U-Net I used:\n\n```\nclass UNetModel(pl.LightningModule):\n    def __init__(\n        self, \n        spatial_dims: int = 3,\n        in_channels: int = 1,\n        out_channels: int = Config.NUM_CLASSES,\n        channels: Union[Tuple[int, ...], List[int]] = (32, 64, 128, 128),\n        strides: Union[Tuple[int, ...], List[int]] = (2, 2, 1),\n        num_res_units: int = 1,\n        lr: float = Config.LEARNING_RATE,\n        dropout: float = 0.3,\n    ):\n        super().__init__()\n        self.save_hyperparameters()\n        \n        self.model = UNet(\n            spatial_dims=self.hparams.spatial_dims,\n            in_channels=self.hparams.in_channels,\n            out_channels=self.hparams.out_channels,\n            channels=self.hparams.channels,\n            strides=self.hparams.strides,\n            num_res_units=self.hparams.num_res_units,\n            dropout=self.hparams.dropout,\n        )\n\n# loss = DiceCeLoss \n```\n\nI'm segmenting, then used connected component analysis and selected the blob with highest confidence score.\n\nDuring training, the U-Net learns, the validation plots look good and the model is able to detect the motors with only a few artifacts. But when applying a sliding window over the entire tomogram, the model segmentation has a LOT of artifacts.\n\nHere are some reference images:\n\ntomo_00e047 169.0 546.0 603.0\n\ntomo_00e047 z = 20:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2Fcb78b381622d672831d441163aee3a71%2Ftomo_00e047_slice_20_prediction.png?generation=1744803417908613&alt=media)\n\ntomo_00e047 z = 40:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F27ef58a85da26b7c01c34186b97a9e5a%2Ftomo_00e047_slice_40_prediction.png?generation=1744803228442504&alt=media)\n\ntomo_00e047 z = 170:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F0d3a015a0e111d001b31b5a95605e833%2Ftomo_00e047_slice_170_prediction.png?generation=1744803301669830&alt=media)\n",
      "votes": 3,
      "replies": [
        {
          "id": 3180329,
          "postDate": "2025-04-16T13:08:38.690Z",
          "content": "<p>Hi, nice work! In CZII competition I have boosted score a lot by inferencing with x2 side size of crops e.g 96 for train and 192 for inference. That drops false positive detections a lot. You can also try to filter blobs by size, I hope that will help you</p>",
          "rawMarkdown": "Hi, nice work! In CZII competition I have boosted score a lot by inferencing with x2 side size of crops e.g 96 for train and 192 for inference. That drops false positive detections a lot. You can also try to filter blobs by size, I hope that will help you",
          "votes": 3
        },
        {
          "id": 3180718,
          "postDate": "2025-04-17T03:24:15.457Z",
          "content": "<p>Thank you for sharing,<br>\nWhen I try a 3D unet on kaggle GPU, the ram memory runs out. Is there any pretrained object detection model you use or do you train from scratch? Do you only use kaggle GPU or train model externally?</p>",
          "rawMarkdown": "Thank you for sharing,\nWhen I try a 3D unet on kaggle GPU, the ram memory runs out. Is there any pretrained object detection model you use or do you train from scratch? Do you only use kaggle GPU or train model externally?",
          "replies": [
            {
              "id": 3180940,
              "postDate": "2025-04-17T10:02:02.273Z",
              "content": "<p>In the CZII competition I used Unet3d by monai. I use my local 3090 to run experiments and debug the code. I optimized the submission time significantly, which allows for larger, higher resolution models. Here is my submission notebook <a href=\"https://www.kaggle.com/code/fautei/byu-yolo-optimized-submission-notebook\" target=\"_blank\">https://www.kaggle.com/code/fautei/byu-yolo-optimized-submission-notebook</a>. There is nothing special in the training code, you can use the kernel provided by the host</p>",
              "rawMarkdown": "In the CZII competition I used Unet3d by monai. I use my local 3090 to run experiments and debug the code. I optimized the submission time significantly, which allows for larger, higher resolution models. Here is my submission notebook https://www.kaggle.com/code/fautei/byu-yolo-optimized-submission-notebook. There is nothing special in the training code, you can use the kernel provided by the host"
            },
            {
              "id": 3181110,
              "postDate": "2025-04-17T13:26:29.773Z",
              "content": "<p>Hey <a href=\"https://www.kaggle.com/thisarmin\" target=\"_blank\">@thisarmin</a>, I trained everything from scratch using MONAI 3D unet, I used Kaggle’s GPU. To avoid memory issues, make sure you're using cropped inputs instead of full volumes, and keep the batch size small.</p>\n<p>I’ll try to clean up my code and will share it in the next few days if it helps, maybe someone can achieve a better result with U-Net.</p>",
              "rawMarkdown": "Hey @thisarmin, I trained everything from scratch using MONAI 3D unet, I used Kaggle’s GPU. To avoid memory issues, make sure you're using cropped inputs instead of full volumes, and keep the batch size small.\n\nI’ll try to clean up my code and will share it in the next few days if it helps, maybe someone can achieve a better result with U-Net.\n\n"
            }
          ]
        }
      ]
    },
    {
      "id": 3181938,
      "postDate": "2025-04-18T14:58:13.263Z",
      "content": "<p>If u want to use 3d segmentation, u should use BOTH pos and neg tomograph in training. But the trick is to use 3 label. Pos voxel. Neg voxel. Ignore voxel which are in slice close to annotation and may or may not still contain motor assembly.</p>",
      "rawMarkdown": "If u want to use 3d segmentation, u should use BOTH pos and neg tomograph in training. But the trick is to use 3 label. Pos voxel. Neg voxel. Ignore voxel which are in slice close to annotation and may or may not still contain motor assembly.",
      "votes": 2
    },
    {
      "id": 3181073,
      "postDate": "2025-04-17T12:31:33.483Z",
      "content": "<p>you have a third option. do encoder and decoder in two stage</p>\n<p>assume input is 500x 960x960.</p>\n<p>stage.1:<br>\nyou make a 3d encoder. like most image encoder, then output is 1/32 scaled. this is 15x30x30 feature map.<br>\nit is said test sample has only one motor, so you do a softmax to predict the block that most likely to contain motor. <br>\n(or you can use sigmoid to predict multiple blocks)</p>\n<p>stage.2:<br>\nfor that block in the feature map, you can crop the corresponding region in original  500x 960x960 input.<br>\nyou can then make another decoder verifer model:</p>\n<p>verifier(crop region) = more precise location, label (got or not)</p>",
      "rawMarkdown": "you have a third option. do encoder and decoder in two stage\n\nassume input is 500x 960x960.\n\nstage.1:\nyou make a 3d encoder. like most image encoder, then output is 1/32 scaled. this is 15x30x30 feature map.\nit is said test sample has only one motor, so you do a softmax to predict the block that most likely to contain motor. \n(or you can use sigmoid to predict multiple blocks)\n\n\nstage.2:\nfor that block in the feature map, you can crop the corresponding region in original  500x 960x960 input.\nyou can then make another decoder verifer model:\n\nverifier(crop region) = more precise location, label (got or not)",
      "votes": 2
    },
    {
      "id": 3179720,
      "postDate": "2025-04-15T16:09:38.150Z",
      "content": "<p>Why don't you start with 2d unet? I think you'll lose less information that way.<br>\nHave you read this interesting <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/572008\" target=\"_blank\">discussion</a> ? </p>",
      "rawMarkdown": "Why don't you start with 2d unet? I think you'll lose less information that way.\nHave you read this interesting [discussion](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/572008) ? "
    },
    {
      "id": 3179090,
      "postDate": "2025-04-15T03:08:53.220Z",
      "content": "<p>A 3d unet is a very viable solution for this competition but I would make the following recommendation, do not try to fit the whole input volume to the unet.  I recommend creating voxels (3d patches) of something like 128x128x128 (for example) from the volume because an input of 950x950x500 (for some of the samples) will require way too much memory and will be difficult to learn small features.</p>",
      "rawMarkdown": "A 3d unet is a very viable solution for this competition but I would make the following recommendation, do not try to fit the whole input volume to the unet.  I recommend creating voxels (3d patches) of something like 128x128x128 (for example) from the volume because an input of 950x950x500 (for some of the samples) will require way too much memory and will be difficult to learn small features.",
      "replies": [
        {
          "id": 3179118,
          "postDate": "2025-04-15T04:15:18.820Z",
          "content": "<p>By averaging? The flagellar is already so small, I’m going to lose a lot information folding it 9 times. </p>",
          "rawMarkdown": "By averaging? The flagellar is already so small, I’m going to lose a lot information folding it 9 times. ",
          "replies": [
            {
              "id": 3179783,
              "postDate": "2025-04-15T17:23:04.233Z",
              "content": "<p>You do a full scale crop from the volume, not a resize</p>",
              "rawMarkdown": "You do a full scale crop from the volume, not a resize"
            },
            {
              "id": 3180714,
              "postDate": "2025-04-17T03:16:06.740Z",
              "content": "<p>dividing the data into crops or just crop around the target?</p>",
              "rawMarkdown": "dividing the data into crops or just crop around the target?"
            }
          ]
        }
      ]
    },
    {
      "id": 3208696,
      "postDate": "2025-05-24T14:11:36.613Z",
      "content": "<p>What arch have u tried with 3d unet, consider trying attention gates in skip connections, that may reduce no of FPs being labelled  and bigger window size with bigger kernels if possible</p>\n<p>I currently want to explore 3d OD , we can collaborate if are ok</p>",
      "rawMarkdown": "What arch have u tried with 3d unet, consider trying attention gates in skip connections, that may reduce no of FPs being labelled  and bigger window size with bigger kernels if possible\n\nI currently want to explore 3d OD , we can collaborate if are ok",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3180287,
      "author_name": "Sergio Alvarez",
      "author_url": "",
      "post_date": "2025-04-16T11:38:32.413000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/thisarmin\" target=\"_blank\">@thisarmin</a> , I've tried adapting the code I used in the CZII competition, where I trained a 3D U-Net. As others mentioned, feeding the full volume at once isn’t possible because of memory constraints.</p>\n<p>I downsampled the volumes by 0.5 and fed the network with crops of size (128, 128, 128) centered around the particle. I also experimented with (24, 256, 256) crops. My best 3D U-Net model scored 0.436 on the leaderboard.</p>\n<p>Since I noticed many people were using 2D models, I also tried a 2D U-Net, resizing slices to (640, 640) and feeding the model 8 Z-slices around the particle center. It didn’t outperform the 3D U-Net.</p>\n<p>For the targets, I used masks that were simple blobs with a ~6px radius centered on the particles.</p>\n<p>Here’s a the code of the MONAI 3D U-Net I used:</p>\n<pre><code> (pl.LightningModule):\n     ():\n        ().__init__()\n        .save_hyperparameters()\n\n        .model = UNet(\n            spatial_dims=.hparams.spatial_dims,\n            in_channels=.hparams.in_channels,\n            out_channels=.hparams.out_channels,\n            channels=.hparams.channels,\n            strides=.hparams.strides,\n            num_res_units=.hparams.num_res_units,\n            dropout=.hparams.dropout,\n        )\n\n\n</code></pre>\n<p>I'm segmenting, then used connected component analysis and selected the blob with highest confidence score.</p>\n<p>During training, the U-Net learns, the validation plots look good and the model is able to detect the motors with only a few artifacts. But when applying a sliding window over the entire tomogram, the model segmentation has a LOT of artifacts.</p>\n<p>Here are some reference images:</p>\n<p>tomo_00e047 169.0 546.0 603.0</p>\n<p>tomo_00e047 z = 20:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2Fcb78b381622d672831d441163aee3a71%2Ftomo_00e047_slice_20_prediction.png?generation=1744803417908613&amp;alt=media\" alt=\"\"></p>\n<p>tomo_00e047 z = 40:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F27ef58a85da26b7c01c34186b97a9e5a%2Ftomo_00e047_slice_40_prediction.png?generation=1744803228442504&amp;alt=media\" alt=\"\"></p>\n<p>tomo_00e047 z = 170:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F0d3a015a0e111d001b31b5a95605e833%2Ftomo_00e047_slice_170_prediction.png?generation=1744803301669830&amp;alt=media\" alt=\"\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 3180329,
          "author_name": "Maxim Ilyin",
          "author_url": "",
          "post_date": "2025-04-16T13:08:38.690000",
          "content": "<p>Hi, nice work! In CZII competition I have boosted score a lot by inferencing with x2 side size of crops e.g 96 for train and 192 for inference. That drops false positive detections a lot. You can also try to filter blobs by size, I hope that will help you</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 3180718,
          "author_name": "thisArmin",
          "author_url": "",
          "post_date": "2025-04-17T03:24:15.457000",
          "content": "<p>Thank you for sharing,<br>\nWhen I try a 3D unet on kaggle GPU, the ram memory runs out. Is there any pretrained object detection model you use or do you train from scratch? Do you only use kaggle GPU or train model externally?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3180940,
              "author_name": "Maxim Ilyin",
              "author_url": "",
              "post_date": "2025-04-17T10:02:02.273000",
              "content": "<p>In the CZII competition I used Unet3d by monai. I use my local 3090 to run experiments and debug the code. I optimized the submission time significantly, which allows for larger, higher resolution models. Here is my submission notebook <a href=\"https://www.kaggle.com/code/fautei/byu-yolo-optimized-submission-notebook\" target=\"_blank\">https://www.kaggle.com/code/fautei/byu-yolo-optimized-submission-notebook</a>. There is nothing special in the training code, you can use the kernel provided by the host</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3181110,
              "author_name": "Sergio Alvarez",
              "author_url": "",
              "post_date": "2025-04-17T13:26:29.773000",
              "content": "<p>Hey <a href=\"https://www.kaggle.com/thisarmin\" target=\"_blank\">@thisarmin</a>, I trained everything from scratch using MONAI 3D unet, I used Kaggle’s GPU. To avoid memory issues, make sure you're using cropped inputs instead of full volumes, and keep the batch size small.</p>\n<p>I’ll try to clean up my code and will share it in the next few days if it helps, maybe someone can achieve a better result with U-Net.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3181938,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-04-18T14:58:13.263000",
      "content": "<p>If u want to use 3d segmentation, u should use BOTH pos and neg tomograph in training. But the trick is to use 3 label. Pos voxel. Neg voxel. Ignore voxel which are in slice close to annotation and may or may not still contain motor assembly.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3181073,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-04-17T12:31:33.483000",
      "content": "<p>you have a third option. do encoder and decoder in two stage</p>\n<p>assume input is 500x 960x960.</p>\n<p>stage.1:<br>\nyou make a 3d encoder. like most image encoder, then output is 1/32 scaled. this is 15x30x30 feature map.<br>\nit is said test sample has only one motor, so you do a softmax to predict the block that most likely to contain motor. <br>\n(or you can use sigmoid to predict multiple blocks)</p>\n<p>stage.2:<br>\nfor that block in the feature map, you can crop the corresponding region in original  500x 960x960 input.<br>\nyou can then make another decoder verifer model:</p>\n<p>verifier(crop region) = more precise location, label (got or not)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3179720,
      "author_name": "MathieuD",
      "author_url": "",
      "post_date": "2025-04-15T16:09:38.150000",
      "content": "<p>Why don't you start with 2d unet? I think you'll lose less information that way.<br>\nHave you read this interesting <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/572008\" target=\"_blank\">discussion</a> ? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3179090,
      "author_name": "Connor",
      "author_url": "",
      "post_date": "2025-04-15T03:08:53.220000",
      "content": "<p>A 3d unet is a very viable solution for this competition but I would make the following recommendation, do not try to fit the whole input volume to the unet.  I recommend creating voxels (3d patches) of something like 128x128x128 (for example) from the volume because an input of 950x950x500 (for some of the samples) will require way too much memory and will be difficult to learn small features.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3179118,
          "author_name": "thisArmin",
          "author_url": "",
          "post_date": "2025-04-15T04:15:18.820000",
          "content": "<p>By averaging? The flagellar is already so small, I’m going to lose a lot information folding it 9 times. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3179783,
              "author_name": "DennisSakva",
              "author_url": "",
              "post_date": "2025-04-15T17:23:04.233000",
              "content": "<p>You do a full scale crop from the volume, not a resize</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3180714,
              "author_name": "thisArmin",
              "author_url": "",
              "post_date": "2025-04-17T03:16:06.740000",
              "content": "<p>dividing the data into crops or just crop around the target?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3208696,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-24T14:11:36.613000",
      "content": "<p>What arch have u tried with 3d unet, consider trying attention gates in skip connections, that may reduce no of FPs being labelled  and bigger window size with bigger kernels if possible</p>\n<p>I currently want to explore 3d OD , we can collaborate if are ok</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3179047": "I understand the yolov8 is training on 2d slices. Has anyone tried 3D unet, where you give the whole data instead of slices to model? ",
    "3180287": "Hey @thisarmin , I've tried adapting the code I used in the CZII competition, where I trained a 3D U-Net. As others mentioned, feeding the full volume at once isn’t possible because of memory constraints.\n\nI downsampled the volumes by 0.5 and fed the network with crops of size (128, 128, 128) centered around the particle. I also experimented with (24, 256, 256) crops. My best 3D U-Net model scored 0.436 on the leaderboard.\n\nSince I noticed many people were using 2D models, I also tried a 2D U-Net, resizing slices to (640, 640) and feeding the model 8 Z-slices around the particle center. It didn’t outperform the 3D U-Net.\n\nFor the targets, I used masks that were simple blobs with a ~6px radius centered on the particles.\n\nHere’s a the code of the MONAI 3D U-Net I used:\n\n```\nclass UNetModel(pl.LightningModule):\n    def __init__(\n        self, \n        spatial_dims: int = 3,\n        in_channels: int = 1,\n        out_channels: int = Config.NUM_CLASSES,\n        channels: Union[Tuple[int, ...], List[int]] = (32, 64, 128, 128),\n        strides: Union[Tuple[int, ...], List[int]] = (2, 2, 1),\n        num_res_units: int = 1,\n        lr: float = Config.LEARNING_RATE,\n        dropout: float = 0.3,\n    ):\n        super().__init__()\n        self.save_hyperparameters()\n        \n        self.model = UNet(\n            spatial_dims=self.hparams.spatial_dims,\n            in_channels=self.hparams.in_channels,\n            out_channels=self.hparams.out_channels,\n            channels=self.hparams.channels,\n            strides=self.hparams.strides,\n            num_res_units=self.hparams.num_res_units,\n            dropout=self.hparams.dropout,\n        )\n\n# loss = DiceCeLoss \n```\n\nI'm segmenting, then used connected component analysis and selected the blob with highest confidence score.\n\nDuring training, the U-Net learns, the validation plots look good and the model is able to detect the motors with only a few artifacts. But when applying a sliding window over the entire tomogram, the model segmentation has a LOT of artifacts.\n\nHere are some reference images:\n\ntomo_00e047 169.0 546.0 603.0\n\ntomo_00e047 z = 20:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2Fcb78b381622d672831d441163aee3a71%2Ftomo_00e047_slice_20_prediction.png?generation=1744803417908613&alt=media)\n\ntomo_00e047 z = 40:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F27ef58a85da26b7c01c34186b97a9e5a%2Ftomo_00e047_slice_40_prediction.png?generation=1744803228442504&alt=media)\n\ntomo_00e047 z = 170:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F0d3a015a0e111d001b31b5a95605e833%2Ftomo_00e047_slice_170_prediction.png?generation=1744803301669830&alt=media)\n",
    "3181938": "If u want to use 3d segmentation, u should use BOTH pos and neg tomograph in training. But the trick is to use 3 label. Pos voxel. Neg voxel. Ignore voxel which are in slice close to annotation and may or may not still contain motor assembly.",
    "3181073": "you have a third option. do encoder and decoder in two stage\n\nassume input is 500x 960x960.\n\nstage.1:\nyou make a 3d encoder. like most image encoder, then output is 1/32 scaled. this is 15x30x30 feature map.\nit is said test sample has only one motor, so you do a softmax to predict the block that most likely to contain motor. \n(or you can use sigmoid to predict multiple blocks)\n\n\nstage.2:\nfor that block in the feature map, you can crop the corresponding region in original  500x 960x960 input.\nyou can then make another decoder verifer model:\n\nverifier(crop region) = more precise location, label (got or not)",
    "3179720": "Why don't you start with 2d unet? I think you'll lose less information that way.\nHave you read this interesting [discussion](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/572008) ? ",
    "3179090": "A 3d unet is a very viable solution for this competition but I would make the following recommendation, do not try to fit the whole input volume to the unet.  I recommend creating voxels (3d patches) of something like 128x128x128 (for example) from the volume because an input of 950x950x500 (for some of the samples) will require way too much memory and will be difficult to learn small features.",
    "3208696": "What arch have u tried with 3d unet, consider trying attention gates in skip connections, that may reduce no of FPs being labelled  and bigger window size with bigger kernels if possible\n\nI currently want to explore 3d OD , we can collaborate if are ok"
  }
}