{
  "id": 579502,
  "title": "Help... approaches other than YOLO + how to deal with artifacts?",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/579502",
  "author_name": "homiecal",
  "post_date": "2025-05-17T23:56:23.595000",
  "votes": 2,
  "comment_count": 9,
  "views": 0,
  "content": "<p>For this competition, I have persisted with a 3d segmentation approach using gaussian heatmaps and/or binary masks, mainly for the reason to explore and learn as much as possible by working within just 1 model technique. YOLO seems to be the main contender here, and I have avoided it because I'm more interested in the modelling technique itself rather than post processing - clearly this will lead to a low competition rank, or no rank at all lol.</p>\n<p>So, here comes my writeup, and subsequent questions, any feedback will be sincerely appreciated.</p>\n<h3>Context</h3>\n<p>For my solution, I have been using SegResNet back bone from the Monai library with a point detection head (that gives dense predictions for a gaussian heatmap). The models I am referring to have been trained on 30,000 crops of 32 x 256 x 256 on all tomograms (regardless of motor count).</p>\n<p><strong>What is working</strong></p>\n<ul>\n<li>Cross entropy loss with a positive weight of 64 combined with dice loss, seems to yield the best predictions. CE with positive weights allows the model to learn and predict false positives, and as the model starts to make too many of them the dice loss reigns back in those predictions.</li>\n</ul>\n<p><strong>What is not working</strong><br>\nI haven't been able to yield a single val/score above zero (and it's close to zero even on train), there are simply far too many false positives. The model almost always is distracted by the many artifacts that appear (see <code>tomo_00e047</code> below).</p>\n<ul>\n<li>I've increased patch size significantly to try and capturing more global information (as many of the artifacts occur close to X, Y edges).</li>\n<li>I considered hard negative sampling, but after observing the patches after the increase in size, anecodetally there was enough artifacts present in the training data.</li>\n</ul>\n<p>The predictions below highlight that the model, is not only ignoring artifacts, it's also not good at taking in global context of the image (which it should be able to do with an X,Y patch size of 256). For example, it often, predicts a non-zero probability for all edges of the bacteria, and often fails to recognise that the motors are usually located near visible tails of the bacteria.</p>\n<ul>\n<li>Maybe there still isn't enough training on artifacts?</li>\n<li>Maybe the model technique is not suited to this problem, there needs to be more reinforcement of global pattern (say using a 2d model with full image size?)</li>\n</ul>\n<p>For example, take <code>tomo_00e047</code> for instance, these are hold out predictions:</p>\n<p>Center:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3784115%2F3804c74ffc37b9c8ccdcaae8e4ff7d9d%2FCapture.PNG?generation=1747525756444115&amp;alt=media\" alt=\"center_predictions\"></p>\n<p>Max prediction (what would be decoded from non maximum supression as the centre):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3784115%2F44a043244a7eec8b66c55f8a898e8af4%2FCapture.PNG?generation=1747525848669371&amp;alt=media\" alt=\"max prediction\"></p>\n<p><strong>What has not worked</strong></p>\n<ul>\n<li>Focal loss</li>\n<li>Smaller patch sizes (duh)</li>\n<li>Other soft losses (soft dice loss, soft tverksy loss)</li>\n</ul>\n<h3>Final remarks</h3>\n<ul>\n<li>Yes, there are post processing techniques I could employ, like watershed to get rid of larger blobs and smaller blobs - but I am not interested in this, because the model isn't good enough in the first place to start employing these techniques to boost the score slightly.</li>\n<li>I know YOLO is better and will work, but sticking with this approach will help me learn why 3d segmentation (atleast with the segresnet backbone) doesn't work, I just haven't quite nailed exactly why… which is where I am hoping the kaggle community can elucidate for me :) </li>\n</ul>\n<p>This is kind of an open ended post and might be difficult to respond to, but I'm hoping to get some feedback before the inevitable solutions come out, because why not have multiple feedback loops.</p>",
  "messages": [
    {
      "id": 3204997,
      "postDate": "2025-05-19T06:45:12.007Z",
      "content": "<p>Try increasing the ball size. Try hard balls instead of the Gaussian.</p>",
      "rawMarkdown": "Try increasing the ball size. Try hard balls instead of the Gaussian.",
      "votes": 1
    },
    {
      "id": 3204866,
      "postDate": "2025-05-19T01:50:40.810Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/homiecal\" target=\"_blank\">@homiecal</a>, I don’t really have a answer either, and to be honest, I still don’t fully understand why 3D U-Nets produce so many artifacts.</p>\n<p>I also experimented with 3D methods (specifically 3D U-Net) and ran into similar issues. I ended up with a 0.527 LB score, but still had lots of artifacts (screenshot below). My main strategy to at least get something out was to apply connected components and select the most confident blob. It didn’t solve the artifact problem, but it helped me score.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F97fd8600706d83351d1f2363f692fdd2%2Ftomo_493bea_confidence_comparison.jpg?generation=1747619380923697&amp;alt=media\" alt=\"\"><br>\nI used 24×256×256 patches and DiceCELoss from MONAI. For each tomo, I extracted 6 \"random\" patches, with a bias towards ones that contained motors</p>",
      "rawMarkdown": "Hey @homiecal, I don’t really have a answer either, and to be honest, I still don’t fully understand why 3D U-Nets produce so many artifacts.\n\nI also experimented with 3D methods (specifically 3D U-Net) and ran into similar issues. I ended up with a 0.527 LB score, but still had lots of artifacts (screenshot below). My main strategy to at least get something out was to apply connected components and select the most confident blob. It didn’t solve the artifact problem, but it helped me score.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F97fd8600706d83351d1f2363f692fdd2%2Ftomo_493bea_confidence_comparison.jpg?generation=1747619380923697&alt=media)\nI used 24×256×256 patches and DiceCELoss from MONAI. For each tomo, I extracted 6 \"random\" patches, with a bias towards ones that contained motors\n",
      "votes": 1,
      "replies": [
        {
          "id": 3204886,
          "postDate": "2025-05-19T02:41:35.923Z",
          "content": "<p><a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> Your result is same as mine I got in last month. I find that segmentation model is really sensitive to peak signal in this type of data. One way you can try is:</p>\n<ol>\n<li>produce pixel or patch predictions first</li>\n<li>blurring the patch vector in the feature map corresponding to the location of positive prediction</li>\n<li>feed the blurred feature map to the segmentation head again</li>\n</ol>\n<p>LB might boosts to 0.7+ I think. But I'm not a fan of reading raw voxel in this comp. </p>\n<p>Another way you can do some preprocessing first:</p>\n<ol>\n<li>capture the regions (superpixel or region prop or GraphCut) with low content complexity score (entropy, local contrast, gradient magnitude mean, LBP, Gabor response, haralick feature, edge density, corner density, fractal dimension, fourier spectrum, perimeter-to-area ratio).</li>\n<li>blur that region with large kernel size</li>\n</ol>",
          "rawMarkdown": "@sersasj Your result is same as mine I got in last month. I find that segmentation model is really sensitive to peak signal in this type of data. One way you can try is:\n1. produce pixel or patch predictions first\n2. blurring the patch vector in the feature map corresponding to the location of positive prediction\n3. feed the blurred feature map to the segmentation head again\n\nLB might boosts to 0.7+ I think. But I'm not a fan of reading raw voxel in this comp. \n\nAnother way you can do some preprocessing first:\n1. capture the regions (superpixel or region prop or GraphCut) with low content complexity score (entropy, local contrast, gradient magnitude mean, LBP, Gabor response, haralick feature, edge density, corner density, fractal dimension, fourier spectrum, perimeter-to-area ratio).\n2. blur that region with large kernel size",
          "votes": 1
        }
      ]
    },
    {
      "id": 3204776,
      "postDate": "2025-05-18T19:29:49.377Z",
      "content": "<p>I think context necessary to avoid those artifacts is bigger than 256, out of range for reasonable 3D patches.</p>",
      "rawMarkdown": "I think context necessary to avoid those artifacts is bigger than 256, out of range for reasonable 3D patches.",
      "votes": 1
    },
    {
      "id": 3204437,
      "postDate": "2025-05-18T09:45:15.590Z",
      "content": "<p>I'm using a similar model, I had problems early on with detecting artifacts or the square edges but now I'm achieving good scores on my validation set (LB is another matter however). I can't say exactly where your problem lies because there are many different choices, but I didn't need to focus specifically on the artifacts, I just tinkered a lot with the parameters (including how many negative crops to feed and loss weights).</p>",
      "rawMarkdown": "I'm using a similar model, I had problems early on with detecting artifacts or the square edges but now I'm achieving good scores on my validation set (LB is another matter however). I can't say exactly where your problem lies because there are many different choices, but I didn't need to focus specifically on the artifacts, I just tinkered a lot with the parameters (including how many negative crops to feed and loss weights).",
      "votes": 1
    },
    {
      "id": 3204213,
      "postDate": "2025-05-17T23:56:23.597Z",
      "content": "<p>For this competition, I have persisted with a 3d segmentation approach using gaussian heatmaps and/or binary masks, mainly for the reason to explore and learn as much as possible by working within just 1 model technique. YOLO seems to be the main contender here, and I have avoided it because I'm more interested in the modelling technique itself rather than post processing - clearly this will lead to a low competition rank, or no rank at all lol.</p>\n<p>So, here comes my writeup, and subsequent questions, any feedback will be sincerely appreciated.</p>\n<h3>Context</h3>\n<p>For my solution, I have been using SegResNet back bone from the Monai library with a point detection head (that gives dense predictions for a gaussian heatmap). The models I am referring to have been trained on 30,000 crops of 32 x 256 x 256 on all tomograms (regardless of motor count).</p>\n<p><strong>What is working</strong></p>\n<ul>\n<li>Cross entropy loss with a positive weight of 64 combined with dice loss, seems to yield the best predictions. CE with positive weights allows the model to learn and predict false positives, and as the model starts to make too many of them the dice loss reigns back in those predictions.</li>\n</ul>\n<p><strong>What is not working</strong><br>\nI haven't been able to yield a single val/score above zero (and it's close to zero even on train), there are simply far too many false positives. The model almost always is distracted by the many artifacts that appear (see <code>tomo_00e047</code> below).</p>\n<ul>\n<li>I've increased patch size significantly to try and capturing more global information (as many of the artifacts occur close to X, Y edges).</li>\n<li>I considered hard negative sampling, but after observing the patches after the increase in size, anecodetally there was enough artifacts present in the training data.</li>\n</ul>\n<p>The predictions below highlight that the model, is not only ignoring artifacts, it's also not good at taking in global context of the image (which it should be able to do with an X,Y patch size of 256). For example, it often, predicts a non-zero probability for all edges of the bacteria, and often fails to recognise that the motors are usually located near visible tails of the bacteria.</p>\n<ul>\n<li>Maybe there still isn't enough training on artifacts?</li>\n<li>Maybe the model technique is not suited to this problem, there needs to be more reinforcement of global pattern (say using a 2d model with full image size?)</li>\n</ul>\n<p>For example, take <code>tomo_00e047</code> for instance, these are hold out predictions:</p>\n<p>Center:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3784115%2F3804c74ffc37b9c8ccdcaae8e4ff7d9d%2FCapture.PNG?generation=1747525756444115&amp;alt=media\" alt=\"center_predictions\"></p>\n<p>Max prediction (what would be decoded from non maximum supression as the centre):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3784115%2F44a043244a7eec8b66c55f8a898e8af4%2FCapture.PNG?generation=1747525848669371&amp;alt=media\" alt=\"max prediction\"></p>\n<p><strong>What has not worked</strong></p>\n<ul>\n<li>Focal loss</li>\n<li>Smaller patch sizes (duh)</li>\n<li>Other soft losses (soft dice loss, soft tverksy loss)</li>\n</ul>\n<h3>Final remarks</h3>\n<ul>\n<li>Yes, there are post processing techniques I could employ, like watershed to get rid of larger blobs and smaller blobs - but I am not interested in this, because the model isn't good enough in the first place to start employing these techniques to boost the score slightly.</li>\n<li>I know YOLO is better and will work, but sticking with this approach will help me learn why 3d segmentation (atleast with the segresnet backbone) doesn't work, I just haven't quite nailed exactly why… which is where I am hoping the kaggle community can elucidate for me :) </li>\n</ul>\n<p>This is kind of an open ended post and might be difficult to respond to, but I'm hoping to get some feedback before the inevitable solutions come out, because why not have multiple feedback loops.</p>",
      "rawMarkdown": "For this competition, I have persisted with a 3d segmentation approach using gaussian heatmaps and/or binary masks, mainly for the reason to explore and learn as much as possible by working within just 1 model technique. YOLO seems to be the main contender here, and I have avoided it because I'm more interested in the modelling technique itself rather than post processing - clearly this will lead to a low competition rank, or no rank at all lol.\n\nSo, here comes my writeup, and subsequent questions, any feedback will be sincerely appreciated.\n\n### Context\nFor my solution, I have been using SegResNet back bone from the Monai library with a point detection head (that gives dense predictions for a gaussian heatmap). The models I am referring to have been trained on 30,000 crops of 32 x 256 x 256 on all tomograms (regardless of motor count).\n\n**What is working**\n- Cross entropy loss with a positive weight of 64 combined with dice loss, seems to yield the best predictions. CE with positive weights allows the model to learn and predict false positives, and as the model starts to make too many of them the dice loss reigns back in those predictions.\n\n**What is not working**\nI haven't been able to yield a single val/score above zero (and it's close to zero even on train), there are simply far too many false positives. The model almost always is distracted by the many artifacts that appear (see `tomo_00e047` below).\n- I've increased patch size significantly to try and capturing more global information (as many of the artifacts occur close to X, Y edges).\n- I considered hard negative sampling, but after observing the patches after the increase in size, anecodetally there was enough artifacts present in the training data.\n\nThe predictions below highlight that the model, is not only ignoring artifacts, it's also not good at taking in global context of the image (which it should be able to do with an X,Y patch size of 256). For example, it often, predicts a non-zero probability for all edges of the bacteria, and often fails to recognise that the motors are usually located near visible tails of the bacteria.\n- Maybe there still isn't enough training on artifacts?\n- Maybe the model technique is not suited to this problem, there needs to be more reinforcement of global pattern (say using a 2d model with full image size?)\n\nFor example, take `tomo_00e047` for instance, these are hold out predictions:\n\nCenter:\n![center_predictions](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3784115%2F3804c74ffc37b9c8ccdcaae8e4ff7d9d%2FCapture.PNG?generation=1747525756444115&alt=media)\n\nMax prediction (what would be decoded from non maximum supression as the centre):\n![max prediction](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3784115%2F44a043244a7eec8b66c55f8a898e8af4%2FCapture.PNG?generation=1747525848669371&alt=media)\n\n**What has not worked**\n- Focal loss\n- Smaller patch sizes (duh)\n- Other soft losses (soft dice loss, soft tverksy loss)\n\n### Final remarks\n\n- Yes, there are post processing techniques I could employ, like watershed to get rid of larger blobs and smaller blobs - but I am not interested in this, because the model isn't good enough in the first place to start employing these techniques to boost the score slightly.\n- I know YOLO is better and will work, but sticking with this approach will help me learn why 3d segmentation (atleast with the segresnet backbone) doesn't work, I just haven't quite nailed exactly why... which is where I am hoping the kaggle community can elucidate for me :) \n\nThis is kind of an open ended post and might be difficult to respond to, but I'm hoping to get some feedback before the inevitable solutions come out, because why not have multiple feedback loops.\n\n",
      "votes": 2
    },
    {
      "id": 3205095,
      "postDate": "2025-05-19T10:18:13.310Z",
      "content": "<p>Thanks everyone for some great discussion - <a href=\"https://www.kaggle.com/tom99763\" target=\"_blank\">@tom99763</a> you raise some particularly interesting points… I asked GPT to further elaborate and clarify some things, you are probably right that the noise to signal ratio here is too high and perhaps dense predictions is not the way to go.</p>\n<p>Would you say that if we perform patch predictions on a grid basis say, splitting the image into a 4 x 4 grid, then taking the grid with the highest probability we can refine the offset for the centre of our object - will lead to a similar effect of ignoring the peak intensity signals? </p>\n<ul>\n<li>Using pre-processing techniques doesn't \"feel\" right to me, because I would expect that if the model and loss are good enough then it should figure out these processing techniques itself, however, it seems that the model clearly <em>isn't good enough</em> and so it could be sensible to pre-process and make the models job easier (I guess after all a model will not do well with noise).</li>\n</ul>",
      "rawMarkdown": "Thanks everyone for some great discussion - @tom99763 you raise some particularly interesting points... I asked GPT to further elaborate and clarify some things, you are probably right that the noise to signal ratio here is too high and perhaps dense predictions is not the way to go.\n\nWould you say that if we perform patch predictions on a grid basis say, splitting the image into a 4 x 4 grid, then taking the grid with the highest probability we can refine the offset for the centre of our object - will lead to a similar effect of ignoring the peak intensity signals? \n- Using pre-processing techniques doesn't \"feel\" right to me, because I would expect that if the model and loss are good enough then it should figure out these processing techniques itself, however, it seems that the model clearly *isn't good enough* and so it could be sensible to pre-process and make the models job easier (I guess after all a model will not do well with noise).",
      "replies": [
        {
          "id": 3205207,
          "postDate": "2025-05-19T13:16:30.880Z",
          "content": "<p>Combine all of ingredients:</p>\n<ul>\n<li>produce predictions from segmentation model</li>\n<li>preprocessing the image based on segment results</li>\n<li>feed the processed image to another segmentation model which takes the refined input</li>\n</ul>\n<p>Actually this is similar to my current solution, so I confirm it would work.</p>",
          "rawMarkdown": "Combine all of ingredients:\n* produce predictions from segmentation model\n* preprocessing the image based on segment results\n* feed the processed image to another segmentation model which takes the refined input\n\nActually this is similar to my current solution, so I confirm it would work.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3208849,
      "postDate": "2025-05-24T19:44:46.413Z",
      "content": "<p>I think you should try Unet with attention gates in skip connections , this may reduce the those irrelevant artifacts catching up , also are you just using randomcrop ?? try randomcropwithposneglabel in MONAI</p>",
      "rawMarkdown": "I think you should try Unet with attention gates in skip connections , this may reduce the those irrelevant artifacts catching up , also are you just using randomcrop ?? try randomcropwithposneglabel in MONAI",
      "isDeleted": true
    },
    {
      "id": 3208847,
      "postDate": "2025-05-24T19:44:21.947Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3204997,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2025-05-19T06:45:12.007000",
      "content": "<p>Try increasing the ball size. Try hard balls instead of the Gaussian.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3204866,
      "author_name": "Sergio Alvarez",
      "author_url": "",
      "post_date": "2025-05-19T01:50:40.810000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/homiecal\" target=\"_blank\">@homiecal</a>, I don’t really have a answer either, and to be honest, I still don’t fully understand why 3D U-Nets produce so many artifacts.</p>\n<p>I also experimented with 3D methods (specifically 3D U-Net) and ran into similar issues. I ended up with a 0.527 LB score, but still had lots of artifacts (screenshot below). My main strategy to at least get something out was to apply connected components and select the most confident blob. It didn’t solve the artifact problem, but it helped me score.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F97fd8600706d83351d1f2363f692fdd2%2Ftomo_493bea_confidence_comparison.jpg?generation=1747619380923697&amp;alt=media\" alt=\"\"><br>\nI used 24×256×256 patches and DiceCELoss from MONAI. For each tomo, I extracted 6 \"random\" patches, with a bias towards ones that contained motors</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3204886,
          "author_name": "Tom",
          "author_url": "",
          "post_date": "2025-05-19T02:41:35.923000",
          "content": "<p><a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> Your result is same as mine I got in last month. I find that segmentation model is really sensitive to peak signal in this type of data. One way you can try is:</p>\n<ol>\n<li>produce pixel or patch predictions first</li>\n<li>blurring the patch vector in the feature map corresponding to the location of positive prediction</li>\n<li>feed the blurred feature map to the segmentation head again</li>\n</ol>\n<p>LB might boosts to 0.7+ I think. But I'm not a fan of reading raw voxel in this comp. </p>\n<p>Another way you can do some preprocessing first:</p>\n<ol>\n<li>capture the regions (superpixel or region prop or GraphCut) with low content complexity score (entropy, local contrast, gradient magnitude mean, LBP, Gabor response, haralick feature, edge density, corner density, fractal dimension, fourier spectrum, perimeter-to-area ratio).</li>\n<li>blur that region with large kernel size</li>\n</ol>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3204776,
      "author_name": "Ángel Jacinto Sánchez Ruiz",
      "author_url": "",
      "post_date": "2025-05-18T19:29:49.377000",
      "content": "<p>I think context necessary to avoid those artifacts is bigger than 256, out of range for reasonable 3D patches.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3204437,
      "author_name": "tennogh",
      "author_url": "",
      "post_date": "2025-05-18T09:45:15.590000",
      "content": "<p>I'm using a similar model, I had problems early on with detecting artifacts or the square edges but now I'm achieving good scores on my validation set (LB is another matter however). I can't say exactly where your problem lies because there are many different choices, but I didn't need to focus specifically on the artifacts, I just tinkered a lot with the parameters (including how many negative crops to feed and loss weights).</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3205095,
      "author_name": "homiecal",
      "author_url": "",
      "post_date": "2025-05-19T10:18:13.310000",
      "content": "<p>Thanks everyone for some great discussion - <a href=\"https://www.kaggle.com/tom99763\" target=\"_blank\">@tom99763</a> you raise some particularly interesting points… I asked GPT to further elaborate and clarify some things, you are probably right that the noise to signal ratio here is too high and perhaps dense predictions is not the way to go.</p>\n<p>Would you say that if we perform patch predictions on a grid basis say, splitting the image into a 4 x 4 grid, then taking the grid with the highest probability we can refine the offset for the centre of our object - will lead to a similar effect of ignoring the peak intensity signals? </p>\n<ul>\n<li>Using pre-processing techniques doesn't \"feel\" right to me, because I would expect that if the model and loss are good enough then it should figure out these processing techniques itself, however, it seems that the model clearly <em>isn't good enough</em> and so it could be sensible to pre-process and make the models job easier (I guess after all a model will not do well with noise).</li>\n</ul>",
      "votes": 0,
      "replies": [
        {
          "id": 3205207,
          "author_name": "Tom",
          "author_url": "",
          "post_date": "2025-05-19T13:16:30.880000",
          "content": "<p>Combine all of ingredients:</p>\n<ul>\n<li>produce predictions from segmentation model</li>\n<li>preprocessing the image based on segment results</li>\n<li>feed the processed image to another segmentation model which takes the refined input</li>\n</ul>\n<p>Actually this is similar to my current solution, so I confirm it would work.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3208849,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-24T19:44:46.413000",
      "content": "<p>I think you should try Unet with attention gates in skip connections , this may reduce the those irrelevant artifacts catching up , also are you just using randomcrop ?? try randomcropwithposneglabel in MONAI</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3208847,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-24T19:44:21.947000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3204997": "Try increasing the ball size. Try hard balls instead of the Gaussian.",
    "3204866": "Hey @homiecal, I don’t really have a answer either, and to be honest, I still don’t fully understand why 3D U-Nets produce so many artifacts.\n\nI also experimented with 3D methods (specifically 3D U-Net) and ran into similar issues. I ended up with a 0.527 LB score, but still had lots of artifacts (screenshot below). My main strategy to at least get something out was to apply connected components and select the most confident blob. It didn’t solve the artifact problem, but it helped me score.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F97fd8600706d83351d1f2363f692fdd2%2Ftomo_493bea_confidence_comparison.jpg?generation=1747619380923697&alt=media)\nI used 24×256×256 patches and DiceCELoss from MONAI. For each tomo, I extracted 6 \"random\" patches, with a bias towards ones that contained motors\n",
    "3204776": "I think context necessary to avoid those artifacts is bigger than 256, out of range for reasonable 3D patches.",
    "3204437": "I'm using a similar model, I had problems early on with detecting artifacts or the square edges but now I'm achieving good scores on my validation set (LB is another matter however). I can't say exactly where your problem lies because there are many different choices, but I didn't need to focus specifically on the artifacts, I just tinkered a lot with the parameters (including how many negative crops to feed and loss weights).",
    "3204213": "For this competition, I have persisted with a 3d segmentation approach using gaussian heatmaps and/or binary masks, mainly for the reason to explore and learn as much as possible by working within just 1 model technique. YOLO seems to be the main contender here, and I have avoided it because I'm more interested in the modelling technique itself rather than post processing - clearly this will lead to a low competition rank, or no rank at all lol.\n\nSo, here comes my writeup, and subsequent questions, any feedback will be sincerely appreciated.\n\n### Context\nFor my solution, I have been using SegResNet back bone from the Monai library with a point detection head (that gives dense predictions for a gaussian heatmap). The models I am referring to have been trained on 30,000 crops of 32 x 256 x 256 on all tomograms (regardless of motor count).\n\n**What is working**\n- Cross entropy loss with a positive weight of 64 combined with dice loss, seems to yield the best predictions. CE with positive weights allows the model to learn and predict false positives, and as the model starts to make too many of them the dice loss reigns back in those predictions.\n\n**What is not working**\nI haven't been able to yield a single val/score above zero (and it's close to zero even on train), there are simply far too many false positives. The model almost always is distracted by the many artifacts that appear (see `tomo_00e047` below).\n- I've increased patch size significantly to try and capturing more global information (as many of the artifacts occur close to X, Y edges).\n- I considered hard negative sampling, but after observing the patches after the increase in size, anecodetally there was enough artifacts present in the training data.\n\nThe predictions below highlight that the model, is not only ignoring artifacts, it's also not good at taking in global context of the image (which it should be able to do with an X,Y patch size of 256). For example, it often, predicts a non-zero probability for all edges of the bacteria, and often fails to recognise that the motors are usually located near visible tails of the bacteria.\n- Maybe there still isn't enough training on artifacts?\n- Maybe the model technique is not suited to this problem, there needs to be more reinforcement of global pattern (say using a 2d model with full image size?)\n\nFor example, take `tomo_00e047` for instance, these are hold out predictions:\n\nCenter:\n![center_predictions](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3784115%2F3804c74ffc37b9c8ccdcaae8e4ff7d9d%2FCapture.PNG?generation=1747525756444115&alt=media)\n\nMax prediction (what would be decoded from non maximum supression as the centre):\n![max prediction](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3784115%2F44a043244a7eec8b66c55f8a898e8af4%2FCapture.PNG?generation=1747525848669371&alt=media)\n\n**What has not worked**\n- Focal loss\n- Smaller patch sizes (duh)\n- Other soft losses (soft dice loss, soft tverksy loss)\n\n### Final remarks\n\n- Yes, there are post processing techniques I could employ, like watershed to get rid of larger blobs and smaller blobs - but I am not interested in this, because the model isn't good enough in the first place to start employing these techniques to boost the score slightly.\n- I know YOLO is better and will work, but sticking with this approach will help me learn why 3d segmentation (atleast with the segresnet backbone) doesn't work, I just haven't quite nailed exactly why... which is where I am hoping the kaggle community can elucidate for me :) \n\nThis is kind of an open ended post and might be difficult to respond to, but I'm hoping to get some feedback before the inevitable solutions come out, because why not have multiple feedback loops.\n\n",
    "3205095": "Thanks everyone for some great discussion - @tom99763 you raise some particularly interesting points... I asked GPT to further elaborate and clarify some things, you are probably right that the noise to signal ratio here is too high and perhaps dense predictions is not the way to go.\n\nWould you say that if we perform patch predictions on a grid basis say, splitting the image into a 4 x 4 grid, then taking the grid with the highest probability we can refine the offset for the centre of our object - will lead to a similar effect of ignoring the peak intensity signals? \n- Using pre-processing techniques doesn't \"feel\" right to me, because I would expect that if the model and loss are good enough then it should figure out these processing techniques itself, however, it seems that the model clearly *isn't good enough* and so it could be sensible to pre-process and make the models job easier (I guess after all a model will not do well with noise).",
    "3208849": "I think you should try Unet with attention gates in skip connections , this may reduce the those irrelevant artifacts catching up , also are you just using randomcrop ?? try randomcropwithposneglabel in MONAI",
    "3208847": ""
  }
}