{
  "id": 561709,
  "title": "95th Place Solution for the CZII - CryoET Object Identification Competition",
  "url": "/competitions/czii-cryo-et-object-identification/writeups/homiecal-95th-place-solution-for-the-czii-cryoet-o",
  "author_name": "",
  "post_date": "2025-02-07T11:07:29.305590200Z",
  "votes": 9,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thank you to the organisers of the competition - this is my first competition and I am impressed with the responsiveness of the competition hosts, the quality of the data and research problem. </p>\n<p>Also many thanks to <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a>, the CryoET personal tutor, and congratulations on a gold medal, well deserved.</p>\n<h2>Solution Overview</h2>\n<p>I tried many things, but my final solution is very straightforward, and similar to many other solution posts. I employed a 3D U-Net structure through the Monai library, and incrementally improved the training pipeline, augmentations, and post-processing pipeline to yield better results. My initial 3D U-Net scored 0.58 (0.57) public leaderboard, with essentially the same model structure I managed to push this to 0.729 (0.719) public leaderboard.</p>\n<h3>Training Pipeline</h3>\n<ul>\n<li>Segmentation masks were created from particle centres to be 0.5*radius provided. </li>\n<li><code>TverksyLoss</code> function was used with <code>alpha=0.5</code> and <code>beta=1</code>.</li>\n<li>Augmentations: I used a patch_size of <code>[48, 196, 196]</code>, the larger the patch size for training the better the results - I also found that larger dims in the x,y plane were favoured over the z plane (i.e. better scores with <code>[48, 196, 196]</code> opposed to <code>[128, 128, 128]</code>).</li>\n</ul>\n<pre><code>Compose([\n        RandCropByLabelClassesd(\n            keys=[, ],\n            label_key=,\n            spatial_size=[, , ],\n            num_classes=,\n            num_samples=,\n            ratios=[,,,,,,]\n        ),\n        RandRotate90d(keys=[, ], prob=, spatial_axes=[, ]),\n        RandFlipd(keys=[, ], prob=, spatial_axis=[,,]),    \n        RandRotated(keys=[,],range_x=(-, ), range_y=(-, ), prob=, padding_mode=),\n        RandGibbsNoised(keys=[],prob=),\n        RandAdjustContrastd(keys=[],prob=,gamma=(,),retain_stats=),\n])\n</code></pre>\n<ul>\n<li>I ensembled 7 models from k-fold validation, with the following model config:</li>\n</ul>\n<pre><code>model = UNet(\n            spatial_dims=,\n            in_channels=,\n            out_channels=+, \n            channels=[, , , ],\n            strides=[, , ],\n            num_res_units=,\n        )\n</code></pre>\n<ul>\n<li>Optimizer: <code>Adam</code> with <code>CosineAnnealingLR</code> learning rate = 1e-3 and scheduled to 1e-4 after 100 iterations.</li>\n</ul>\n<h3>Inference Pipeline</h3>\n<p>A sliding window inference was used over patches of dimensions <code>[96,288,288]</code> with an <code>overlap=0.4</code> and guassian weighting with <code>sigma=0.1</code>. A larger patch, less border artifacts, additionally the larger overlap also less border artifacts (but at the cost of an increase in runtime), the config mentioned was the best balance between time and score, and <code>overlap=0.5</code> increased the run time enough that I could only use 6/7 k-models.</p>\n<p>A lot of improvements in my score came from tweaking the post processing. I used watershed algorithm provided by competition hosts, but moved all operations to the GPU using the <code>cupyx</code> library (this decreased computation time by ~70-80%). I applied <code>min_particle_size = 0.3</code> and <code>max_particle_size=1.0</code> filters on the masks created from watershed algorithm. The filter of <code>min_particle_size = 0.3</code> was key to filtering out many false positives, and noise created by the interpretation of artefacts as particles (see below TS_99_9, z-slice=7):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3784115%2F46d5a3957fdbcdffacb6439c44d6ce4d%2Fexample.JPG?generation=1738924826529887&amp;alt=media\" alt=\"TS_99_9_zslice_7\"></p>\n<h2>What I tried to implement</h2>\n<p>I spent sometime implementing the focal loss from scratch from the CentreNet paper and use guassian weigthed heatmaps, but was unable to a 3D FCNN model to get any meaningful results. I saw the first solution took this approach initially with impressive results, so I know it is possible.</p>\n<p>As a last minute stretch and inspired by other solutions, in the last 3 days, I attempted to bring a 2D object detection (ResNet18 backbone, R-FasterCNN RPN) model into the pipeline to filter out false positives in 2D plane after the watershed algorithm selected particle centers from 3D U-Net masks. I had time only to run 3 training iterations, and thus the performance was far too poor to provide any meaningful results. To improve the metric it's 4 false positives removed for every 1 false negative introduced, and the 2D object det model did not reach this benchmark.</p>\n<h3>Other things that did not work</h3>\n<ul>\n<li>Struggled to get anything meaningful out of the synthetic data</li>\n<li>Heatmaps with Focal/CE Loss</li>\n<li>Deeper 3D U-Nets</li>\n<li>Attention U-Net: I thought introducing spatial attention, might benefit this task by filtering out irrelevant feature maps that introduce false positives, but I got slightly worse results than the regular 3D U-Net implemented by Monai. My rationale for this was, it might be due to difference in decoder upsample, where Monai's imeplmentation of the Attention U-Net differs from the original paper where the encoder is concatenated to the decoder before a transposed convolution (instead of after like the standard 3D U-Net implemented by Monai).</li>\n</ul>\n<h3>Things I wanted to try</h3>\n<ul>\n<li>Copy Paste augmentation suggested by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>.</li>\n<li>More effort into 2D object det model</li>\n</ul>\n<h2>Conclusion</h2>\n<p>This was a lot of fun, and I have a lot to learn from some fantastic solutions. It might be time to invest in my own GPU setup, flicking between Google Colab and Kaggle Kernels got a bit tiring 😄 </p>",
  "messages": [
    {
      "id": "3117886",
      "postDate": "02/07/2025 11:07:29",
      "content": "<p>Thank you to the organisers of the competition - this is my first competition and I am impressed with the responsiveness of the competition hosts, the quality of the data and research problem. </p>\n<p>Also many thanks to <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a>, the CryoET personal tutor, and congratulations on a gold medal, well deserved.</p>\n<h2>Solution Overview</h2>\n<p>I tried many things, but my final solution is very straightforward, and similar to many other solution posts. I employed a 3D U-Net structure through the Monai library, and incrementally improved the training pipeline, augmentations, and post-processing pipeline to yield better results. My initial 3D U-Net scored 0.58 (0.57) public leaderboard, with essentially the same model structure I managed to push this to 0.729 (0.719) public leaderboard.</p>\n<h3>Training Pipeline</h3>\n<ul>\n<li>Segmentation masks were created from particle centres to be 0.5*radius provided. </li>\n<li><code>TverksyLoss</code> function was used with <code>alpha=0.5</code> and <code>beta=1</code>.</li>\n<li>Augmentations: I used a patch_size of <code>[48, 196, 196]</code>, the larger the patch size for training the better the results - I also found that larger dims in the x,y plane were favoured over the z plane (i.e. better scores with <code>[48, 196, 196]</code> opposed to <code>[128, 128, 128]</code>).</li>\n</ul>\n<pre><code>Compose([\n        RandCropByLabelClassesd(\n            keys=[, ],\n            label_key=,\n            spatial_size=[, , ],\n            num_classes=,\n            num_samples=,\n            ratios=[,,,,,,]\n        ),\n        RandRotate90d(keys=[, ], prob=, spatial_axes=[, ]),\n        RandFlipd(keys=[, ], prob=, spatial_axis=[,,]),    \n        RandRotated(keys=[,],range_x=(-, ), range_y=(-, ), prob=, padding_mode=),\n        RandGibbsNoised(keys=[],prob=),\n        RandAdjustContrastd(keys=[],prob=,gamma=(,),retain_stats=),\n])\n</code></pre>\n<ul>\n<li>I ensembled 7 models from k-fold validation, with the following model config:</li>\n</ul>\n<pre><code>model = UNet(\n            spatial_dims=,\n            in_channels=,\n            out_channels=+, \n            channels=[, , , ],\n            strides=[, , ],\n            num_res_units=,\n        )\n</code></pre>\n<ul>\n<li>Optimizer: <code>Adam</code> with <code>CosineAnnealingLR</code> learning rate = 1e-3 and scheduled to 1e-4 after 100 iterations.</li>\n</ul>\n<h3>Inference Pipeline</h3>\n<p>A sliding window inference was used over patches of dimensions <code>[96,288,288]</code> with an <code>overlap=0.4</code> and guassian weighting with <code>sigma=0.1</code>. A larger patch, less border artifacts, additionally the larger overlap also less border artifacts (but at the cost of an increase in runtime), the config mentioned was the best balance between time and score, and <code>overlap=0.5</code> increased the run time enough that I could only use 6/7 k-models.</p>\n<p>A lot of improvements in my score came from tweaking the post processing. I used watershed algorithm provided by competition hosts, but moved all operations to the GPU using the <code>cupyx</code> library (this decreased computation time by ~70-80%). I applied <code>min_particle_size = 0.3</code> and <code>max_particle_size=1.0</code> filters on the masks created from watershed algorithm. The filter of <code>min_particle_size = 0.3</code> was key to filtering out many false positives, and noise created by the interpretation of artefacts as particles (see below TS_99_9, z-slice=7):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3784115%2F46d5a3957fdbcdffacb6439c44d6ce4d%2Fexample.JPG?generation=1738924826529887&amp;alt=media\" alt=\"TS_99_9_zslice_7\"></p>\n<h2>What I tried to implement</h2>\n<p>I spent sometime implementing the focal loss from scratch from the CentreNet paper and use guassian weigthed heatmaps, but was unable to a 3D FCNN model to get any meaningful results. I saw the first solution took this approach initially with impressive results, so I know it is possible.</p>\n<p>As a last minute stretch and inspired by other solutions, in the last 3 days, I attempted to bring a 2D object detection (ResNet18 backbone, R-FasterCNN RPN) model into the pipeline to filter out false positives in 2D plane after the watershed algorithm selected particle centers from 3D U-Net masks. I had time only to run 3 training iterations, and thus the performance was far too poor to provide any meaningful results. To improve the metric it's 4 false positives removed for every 1 false negative introduced, and the 2D object det model did not reach this benchmark.</p>\n<h3>Other things that did not work</h3>\n<ul>\n<li>Struggled to get anything meaningful out of the synthetic data</li>\n<li>Heatmaps with Focal/CE Loss</li>\n<li>Deeper 3D U-Nets</li>\n<li>Attention U-Net: I thought introducing spatial attention, might benefit this task by filtering out irrelevant feature maps that introduce false positives, but I got slightly worse results than the regular 3D U-Net implemented by Monai. My rationale for this was, it might be due to difference in decoder upsample, where Monai's imeplmentation of the Attention U-Net differs from the original paper where the encoder is concatenated to the decoder before a transposed convolution (instead of after like the standard 3D U-Net implemented by Monai).</li>\n</ul>\n<h3>Things I wanted to try</h3>\n<ul>\n<li>Copy Paste augmentation suggested by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>.</li>\n<li>More effort into 2D object det model</li>\n</ul>\n<h2>Conclusion</h2>\n<p>This was a lot of fun, and I have a lot to learn from some fantastic solutions. It might be time to invest in my own GPU setup, flicking between Google Colab and Kaggle Kernels got a bit tiring 😄 </p>",
      "rawMarkdown": "Thank you to the organisers of the competition - this is my first competition and I am impressed with the responsiveness of the competition hosts, the quality of the data and research problem. \n\nAlso many thanks to @davidlist, the CryoET personal tutor, and congratulations on a gold medal, well deserved.\n\n## Solution Overview\n\nI tried many things, but my final solution is very straightforward, and similar to many other solution posts. I employed a 3D U-Net structure through the Monai library, and incrementally improved the training pipeline, augmentations, and post-processing pipeline to yield better results. My initial 3D U-Net scored 0.58 (0.57) public leaderboard, with essentially the same model structure I managed to push this to 0.729 (0.719) public leaderboard.\n\n\n### Training Pipeline\n\n- Segmentation masks were created from particle centres to be 0.5*radius provided. \n- `TverksyLoss` function was used with `alpha=0.5` and `beta=1`.\n- Augmentations: I used a patch_size of `[48, 196, 196]`, the larger the patch size for training the better the results - I also found that larger dims in the x,y plane were favoured over the z plane (i.e. better scores with `[48, 196, 196]` opposed to `[128, 128, 128]`).\n\n```python\nCompose([\n        RandCropByLabelClassesd(\n            keys=[\"image\", \"label\"],\n            label_key=\"label\",\n            spatial_size=[48, 196, 196],\n            num_classes=7,\n            num_samples=8,\n            ratios=[2,1,1,2,1,1,2]\n        ),\n        RandRotate90d(keys=[\"image\", \"label\"], prob=0.5, spatial_axes=[1, 2]),\n        RandFlipd(keys=[\"image\", \"label\"], prob=0.5, spatial_axis=[0,1,2]),    \n        RandRotated(keys=['image','label'],range_x=(-1, 1), range_y=(-1, 1), prob=0.25, padding_mode='zeros'),\n        RandGibbsNoised(keys=['image'],prob=0.2),\n        RandAdjustContrastd(keys=['image'],prob=0.2,gamma=(0.5,2),retain_stats=True),\n])\n```\n\n- I ensembled 7 models from k-fold validation, with the following model config:\n\n```python\nmodel = UNet(\n            spatial_dims=3,\n            in_channels=1,\n            out_channels=6+1, \n            channels=[32, 64, 80, 80],\n            strides=[2, 2, 1],\n            num_res_units=2,\n        )\n```\n\n- Optimizer: `Adam` with `CosineAnnealingLR` learning rate = 1e-3 and scheduled to 1e-4 after 100 iterations.\n\n### Inference Pipeline\n\nA sliding window inference was used over patches of dimensions `[96,288,288]` with an `overlap=0.4` and guassian weighting with `sigma=0.1`. A larger patch, less border artifacts, additionally the larger overlap also less border artifacts (but at the cost of an increase in runtime), the config mentioned was the best balance between time and score, and `overlap=0.5` increased the run time enough that I could only use 6/7 k-models.\n\nA lot of improvements in my score came from tweaking the post processing. I used watershed algorithm provided by competition hosts, but moved all operations to the GPU using the `cupyx` library (this decreased computation time by ~70-80%). I applied `min_particle_size = 0.3` and `max_particle_size=1.0` filters on the masks created from watershed algorithm. The filter of `min_particle_size = 0.3` was key to filtering out many false positives, and noise created by the interpretation of artefacts as particles (see below TS_99_9, z-slice=7):\n![TS_99_9_zslice_7](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3784115%2F46d5a3957fdbcdffacb6439c44d6ce4d%2Fexample.JPG?generation=1738924826529887&alt=media)\n\n\n\n\n## What I tried to implement\n\nI spent sometime implementing the focal loss from scratch from the CentreNet paper and use guassian weigthed heatmaps, but was unable to a 3D FCNN model to get any meaningful results. I saw the first solution took this approach initially with impressive results, so I know it is possible.\n\nAs a last minute stretch and inspired by other solutions, in the last 3 days, I attempted to bring a 2D object detection (ResNet18 backbone, R-FasterCNN RPN) model into the pipeline to filter out false positives in 2D plane after the watershed algorithm selected particle centers from 3D U-Net masks. I had time only to run 3 training iterations, and thus the performance was far too poor to provide any meaningful results. To improve the metric it's 4 false positives removed for every 1 false negative introduced, and the 2D object det model did not reach this benchmark.\n\n### Other things that did not work\n- Struggled to get anything meaningful out of the synthetic data\n- Heatmaps with Focal/CE Loss\n- Deeper 3D U-Nets\n- Attention U-Net: I thought introducing spatial attention, might benefit this task by filtering out irrelevant feature maps that introduce false positives, but I got slightly worse results than the regular 3D U-Net implemented by Monai. My rationale for this was, it might be due to difference in decoder upsample, where Monai's imeplmentation of the Attention U-Net differs from the original paper where the encoder is concatenated to the decoder before a transposed convolution (instead of after like the standard 3D U-Net implemented by Monai).\n\n### Things I wanted to try\n- Copy Paste augmentation suggested by @hengck23.\n- More effort into 2D object det model\n\n## Conclusion\nThis was a lot of fun, and I have a lot to learn from some fantastic solutions. It might be time to invest in my own GPU setup, flicking between Google Colab and Kaggle Kernels got a bit tiring 😄",
      "votes": null
    },
    {
      "id": "3117915",
      "postDate": "02/07/2025 12:12:04",
      "content": "<p><a href=\"https://www.kaggle.com/homiecal\" target=\"_blank\">@homiecal</a> Hello and congratulations on your medal! May I ask you about the watershed? I remember attempting it myself but having memory issues with it? I also don’t recall seeing it in the competition provided starter solution. Thanks!</p>",
      "rawMarkdown": "homiecal Hello and congratulations on your medal! May I ask you about the watershed? I remember attempting it myself but having memory issues with it? I also don’t recall seeing it in the competition provided starter solution. Thanks!",
      "votes": null
    },
    {
      "id": "3119410",
      "postDate": "02/09/2025 08:36:14",
      "content": "<blockquote>\n  <p>Also many thanks to <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a>, the CryoET personal tutor, and congratulations on a gold medal, well deserved.  </p>\n</blockquote>\n<p>Funny.  😀  And thank you!  Congrats on your first medal!!!</p>",
      "rawMarkdown": ">Also many thanks to @davidlist, the CryoET personal tutor, and congratulations on a gold medal, well deserved.  \n\nFunny.  😀  And thank you!  Congrats on your first medal!!!",
      "votes": null
    },
    {
      "id": "3120286",
      "postDate": "02/10/2025 11:09:35",
      "content": "<p>Hey Andrei, here's a snippet of the function below. I inference using the full image size (184, 630, 630) and manage the GPU contraints by moving operations to the CPU where possible, and deleting data throughout the pipeline. You can clog up memory quite quickly, because each step of the process creates a new tensor (binary dilation, erosion etc.).</p>\n<p>The core of this function can be found in the <a href=\"https://github.com/copick/copick-utils/blob/main/src/copick_utils/segmentation/picks_from_segmentation.py\" target=\"_blank\">copick-utils</a> repo <code>picks_from_segmentation()</code>.</p>\n<pre><code> torch\n cupy  cp\n cupyx.scipy.ndimage  distance_transform_edt, label, maximum_filter, binary_erosion, binary_dilation\n skimage.segmentation  watershed\n skimage.measure  regionprops\n skimage.morphology  ball\n\n ():\n    \n    log = []  \n\n     ():\n        end_time = time.time()\n        log.append({: step_name, : np.(end_time - start_time, )})\n\n    start_time = time.time()\n    log_step(, start_time)\n    segmentation_gpu = cp.asarray(segmentation)\n\n    \n    step_start = time.time()\n    \n    binary_mask = (segmentation_gpu == segmentation_idx).astype(cp.int32)\n\n    \n     np.(binary_mask) == :\n        log_step(, step_start)\n        ()\n         log\n\n    log_step(, step_start)\n\n    \n    step_start = time.time()\n    struct_elem = ball()\n    struct_elem_gpu = cp.asarray(struct_elem)\n    eroded = cp.asarray(binary_erosion(binary_mask, struct_elem_gpu))\n    dilated = cp.asarray(binary_dilation(eroded, struct_elem_gpu))\n    log_step(, step_start)\n\n    \n     eroded\n\n    \n    step_start = time.time()\n    \n    distance = distance_transform_edt(dilated)\n    log_step(, step_start)\n\n    step_start = time.time()\n    local_max = distance == maximum_filter(distance, size=maxima_filter_size)\n    log_step(, step_start)\n\n    \n    step_start = time.time()\n    markers, _ = label(local_max)\n    \n     local_max\n    watershed_labels = watershed(-distance.get(), markers.get(), mask=dilated.get())\n    watershed_labels_cpu = watershed_labels.astype()\n    \n     watershed_labels, distance, markers, dilated\n    log_step(, step_start)\n\n    \n    step_start = time.time()\n    all_centroids = []\n     region  regionprops(watershed_labels_cpu):\n         min_particle_size &lt;= region.area &lt;= max_particle_size:\n            all_centroids.append(region.centroid)\n    log_step(, step_start)\n     watershed_labels_cpu\n\n    \n    step_start = time.time()\n     all_centroids:\n         submission:\n             [{\n                    : run.name,\n                    : pickable_object,\n                    : c[] * voxel_spacing,\n                    : c[] * voxel_spacing,\n                    : c[] * voxel_spacing,\n                }\n                 c  all_centroids\n            ]   \n        :\n            pick_set = run.new_picks(pickable_object, session_id, user_id)\n\n            pick_set.points = [\n                CopickPoint(\n                    location=CopickLocation(\n                        **{\n                            : c[] * voxel_spacing,\n                            : c[] * voxel_spacing,\n                            : c[] * voxel_spacing,\n                        }\n                    )\n                )\n                 c  all_centroids\n            ]\n            pick_set.store()\n            log_step(, step_start)\n            ()\n            ()\n             log\n\n\n    :\n        log_step(, step_start)\n        ()\n         log\n</code></pre>",
      "rawMarkdown": "Hey Andrei, here's a snippet of the function below. I inference using the full image size (184, 630, 630) and manage the GPU contraints by moving operations to the CPU where possible, and deleting data throughout the pipeline. You can clog up memory quite quickly, because each step of the process creates a new tensor (binary dilation, erosion etc.).\n\nThe core of this function can be found in the [copick-utils](https://github.com/copick/copick-utils/blob/main/src/copick_utils/segmentation/picks_from_segmentation.py) repo `picks_from_segmentation()`.\n\n```python\nimport torch\nimport cupy as cp\nfrom cupyx.scipy.ndimage import distance_transform_edt, label, maximum_filter, binary_erosion, binary_dilation\nfrom skimage.segmentation import watershed\nfrom skimage.measure import regionprops\nfrom skimage.morphology import ball\n\ndef picks_from_segmentation_time_gpu(\n    segmentation,\n    segmentation_idx,\n    maxima_filter_size,\n    min_particle_size,\n    max_particle_size,\n    session_id,\n    user_id,\n    pickable_object,\n    run,\n    voxel_spacing=1,\n    submission: bool = False,\n):\n    \"\"\"\n    Process a specific label in the segmentation, extract centroids, and save them as picks.\n\n    Args:\n        segmentation (np.ndarray): Multilabel segmentation array.\n        segmentation_idx (int): The specific label from the segmentation to process.\n        maxima_filter_size (int): Size of the maximum detection filter.\n        min_particle_size (int): Minimum size threshold for particles.\n        max_particle_size (int): Maximum size threshold for particles.\n        session_id (str): Session ID for pick saving.\n        user_id (str): User ID for pick saving.\n        pickable_object (str): The name of the object to save picks for.\n        run: A Copick run object that manages pick saving.\n        voxel_spacing (int): The voxel spacing used to scale pick locations (default 1).\n    \"\"\"\n    log = []  # Logging list for timestamps\n\n    def log_step(step_name, start_time):\n        end_time = time.time()\n        log.append({\"step\": step_name, \"duration\": np.round(end_time - start_time, 5)})\n\n    start_time = time.time()\n    log_step(\"Start function\", start_time)\n    segmentation_gpu = cp.asarray(segmentation)\n\n    # Create a binary mask for the specific segmentation label\n    step_start = time.time()\n    # Create a binary mask for the specific segmentation label\n    binary_mask = (segmentation_gpu == segmentation_idx).astype(cp.int32)\n\n    # Skip if the segmentation label is not present\n    if np.sum(binary_mask) == 0:\n        log_step(\"No segmentation label found\", step_start)\n        print(f\"No segmentation with label {segmentation_idx} found.\")\n        return log\n\n    log_step(\"Binary mask created\", step_start)\n\n    # Structuring element for erosion and dilation\n    step_start = time.time()\n    struct_elem = ball(1)\n    struct_elem_gpu = cp.asarray(struct_elem)\n    eroded = cp.asarray(binary_erosion(binary_mask, struct_elem_gpu))\n    dilated = cp.asarray(binary_dilation(eroded, struct_elem_gpu))\n    log_step(\"Erosion and dilation\", step_start)\n\n    # Release GPU memory\n    del eroded\n\n    # Distance transform and local maxima detection\n    step_start = time.time()\n    # Distance transform and local maxima detection\n    distance = distance_transform_edt(dilated)\n    log_step(\"Distance transform\", step_start)\n\n    step_start = time.time()\n    local_max = distance == maximum_filter(distance, size=maxima_filter_size)\n    log_step(\"Local maxima detection\", step_start)\n\n    # Watershed segmentation\n    step_start = time.time()\n    markers, _ = label(local_max)\n    # Release GPU memory\n    del local_max\n    watershed_labels = watershed(-distance.get(), markers.get(), mask=dilated.get())\n    watershed_labels_cpu = watershed_labels.astype(int)\n    # Release GPU memory\n    del watershed_labels, distance, markers, dilated\n    log_step(\"Watershed segmentation\", step_start)\n\n    # Extract region properties and filter based on particle size\n    step_start = time.time()\n    all_centroids = []\n    for region in regionprops(watershed_labels_cpu):\n        if min_particle_size <= region.area <= max_particle_size:\n            all_centroids.append(region.centroid)\n    log_step(\"Region properties extracted and filtered\", step_start)\n    del watershed_labels_cpu\n\n    # Save centroids as picks\n    step_start = time.time()\n    if all_centroids:\n        if submission:\n            return [{\n                    \"experiment\": run.name,\n                    \"particle_type\": pickable_object,\n                    \"x\": c[2] * voxel_spacing,\n                    \"y\": c[1] * voxel_spacing,\n                    \"z\": c[0] * voxel_spacing,\n                }\n                for c in all_centroids\n            ]   \n        else:\n            pick_set = run.new_picks(pickable_object, session_id, user_id)\n\n            pick_set.points = [\n                CopickPoint(\n                    location=CopickLocation(\n                        **{\n                            \"x\": c[2] * voxel_spacing,\n                            \"y\": c[1] * voxel_spacing,\n                            \"z\": c[0] * voxel_spacing,\n                        }\n                    )\n                )\n                for c in all_centroids\n            ]\n            pick_set.store()\n            log_step(\"Centroids saved successfully\", step_start)\n            print(f\"Centroids for label {segmentation_idx} saved successfully.\")\n            print(f\"Log time stamps for {segmentation_idx}: {log}\")\n            return log\n\n\n    else:\n        log_step(\"No valid centroids found\", step_start)\n        print(f\"No valid centroids found for label {segmentation_idx}.\")\n        return log\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3117915,
      "author_name": "andreizamfir",
      "author_url": "",
      "post_date": "02/07/2025 12:12:04",
      "content": "<p><a href=\"https://www.kaggle.com/homiecal\" target=\"_blank\">@homiecal</a> Hello and congratulations on your medal! May I ask you about the watershed? I remember attempting it myself but having memory issues with it? I also don’t recall seeing it in the competition provided starter solution. Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3120286,
          "author_name": "homiecal",
          "author_url": "",
          "post_date": "02/10/2025 11:09:35",
          "content": "<p>Hey Andrei, here's a snippet of the function below. I inference using the full image size (184, 630, 630) and manage the GPU contraints by moving operations to the CPU where possible, and deleting data throughout the pipeline. You can clog up memory quite quickly, because each step of the process creates a new tensor (binary dilation, erosion etc.).</p>\n<p>The core of this function can be found in the <a href=\"https://github.com/copick/copick-utils/blob/main/src/copick_utils/segmentation/picks_from_segmentation.py\" target=\"_blank\">copick-utils</a> repo <code>picks_from_segmentation()</code>.</p>\n<pre><code> torch\n cupy  cp\n cupyx.scipy.ndimage  distance_transform_edt, label, maximum_filter, binary_erosion, binary_dilation\n skimage.segmentation  watershed\n skimage.measure  regionprops\n skimage.morphology  ball\n\n ():\n    \n    log = []  \n\n     ():\n        end_time = time.time()\n        log.append({: step_name, : np.(end_time - start_time, )})\n\n    start_time = time.time()\n    log_step(, start_time)\n    segmentation_gpu = cp.asarray(segmentation)\n\n    \n    step_start = time.time()\n    \n    binary_mask = (segmentation_gpu == segmentation_idx).astype(cp.int32)\n\n    \n     np.(binary_mask) == :\n        log_step(, step_start)\n        ()\n         log\n\n    log_step(, step_start)\n\n    \n    step_start = time.time()\n    struct_elem = ball()\n    struct_elem_gpu = cp.asarray(struct_elem)\n    eroded = cp.asarray(binary_erosion(binary_mask, struct_elem_gpu))\n    dilated = cp.asarray(binary_dilation(eroded, struct_elem_gpu))\n    log_step(, step_start)\n\n    \n     eroded\n\n    \n    step_start = time.time()\n    \n    distance = distance_transform_edt(dilated)\n    log_step(, step_start)\n\n    step_start = time.time()\n    local_max = distance == maximum_filter(distance, size=maxima_filter_size)\n    log_step(, step_start)\n\n    \n    step_start = time.time()\n    markers, _ = label(local_max)\n    \n     local_max\n    watershed_labels = watershed(-distance.get(), markers.get(), mask=dilated.get())\n    watershed_labels_cpu = watershed_labels.astype()\n    \n     watershed_labels, distance, markers, dilated\n    log_step(, step_start)\n\n    \n    step_start = time.time()\n    all_centroids = []\n     region  regionprops(watershed_labels_cpu):\n         min_particle_size &lt;= region.area &lt;= max_particle_size:\n            all_centroids.append(region.centroid)\n    log_step(, step_start)\n     watershed_labels_cpu\n\n    \n    step_start = time.time()\n     all_centroids:\n         submission:\n             [{\n                    : run.name,\n                    : pickable_object,\n                    : c[] * voxel_spacing,\n                    : c[] * voxel_spacing,\n                    : c[] * voxel_spacing,\n                }\n                 c  all_centroids\n            ]   \n        :\n            pick_set = run.new_picks(pickable_object, session_id, user_id)\n\n            pick_set.points = [\n                CopickPoint(\n                    location=CopickLocation(\n                        **{\n                            : c[] * voxel_spacing,\n                            : c[] * voxel_spacing,\n                            : c[] * voxel_spacing,\n                        }\n                    )\n                )\n                 c  all_centroids\n            ]\n            pick_set.store()\n            log_step(, step_start)\n            ()\n            ()\n             log\n\n\n    :\n        log_step(, step_start)\n        ()\n         log\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3119410,
      "author_name": "davidlist",
      "author_url": "",
      "post_date": "02/09/2025 08:36:14",
      "content": "<blockquote>\n  <p>Also many thanks to <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a>, the CryoET personal tutor, and congratulations on a gold medal, well deserved.  </p>\n</blockquote>\n<p>Funny.  😀  And thank you!  Congrats on your first medal!!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3117886": "Thank you to the organisers of the competition - this is my first competition and I am impressed with the responsiveness of the competition hosts, the quality of the data and research problem. \n\nAlso many thanks to @davidlist, the CryoET personal tutor, and congratulations on a gold medal, well deserved.\n\n## Solution Overview\n\nI tried many things, but my final solution is very straightforward, and similar to many other solution posts. I employed a 3D U-Net structure through the Monai library, and incrementally improved the training pipeline, augmentations, and post-processing pipeline to yield better results. My initial 3D U-Net scored 0.58 (0.57) public leaderboard, with essentially the same model structure I managed to push this to 0.729 (0.719) public leaderboard.\n\n\n### Training Pipeline\n\n- Segmentation masks were created from particle centres to be 0.5*radius provided. \n- `TverksyLoss` function was used with `alpha=0.5` and `beta=1`.\n- Augmentations: I used a patch_size of `[48, 196, 196]`, the larger the patch size for training the better the results - I also found that larger dims in the x,y plane were favoured over the z plane (i.e. better scores with `[48, 196, 196]` opposed to `[128, 128, 128]`).\n\n```python\nCompose([\n        RandCropByLabelClassesd(\n            keys=[\"image\", \"label\"],\n            label_key=\"label\",\n            spatial_size=[48, 196, 196],\n            num_classes=7,\n            num_samples=8,\n            ratios=[2,1,1,2,1,1,2]\n        ),\n        RandRotate90d(keys=[\"image\", \"label\"], prob=0.5, spatial_axes=[1, 2]),\n        RandFlipd(keys=[\"image\", \"label\"], prob=0.5, spatial_axis=[0,1,2]),    \n        RandRotated(keys=['image','label'],range_x=(-1, 1), range_y=(-1, 1), prob=0.25, padding_mode='zeros'),\n        RandGibbsNoised(keys=['image'],prob=0.2),\n        RandAdjustContrastd(keys=['image'],prob=0.2,gamma=(0.5,2),retain_stats=True),\n])\n```\n\n- I ensembled 7 models from k-fold validation, with the following model config:\n\n```python\nmodel = UNet(\n            spatial_dims=3,\n            in_channels=1,\n            out_channels=6+1, \n            channels=[32, 64, 80, 80],\n            strides=[2, 2, 1],\n            num_res_units=2,\n        )\n```\n\n- Optimizer: `Adam` with `CosineAnnealingLR` learning rate = 1e-3 and scheduled to 1e-4 after 100 iterations.\n\n### Inference Pipeline\n\nA sliding window inference was used over patches of dimensions `[96,288,288]` with an `overlap=0.4` and guassian weighting with `sigma=0.1`. A larger patch, less border artifacts, additionally the larger overlap also less border artifacts (but at the cost of an increase in runtime), the config mentioned was the best balance between time and score, and `overlap=0.5` increased the run time enough that I could only use 6/7 k-models.\n\nA lot of improvements in my score came from tweaking the post processing. I used watershed algorithm provided by competition hosts, but moved all operations to the GPU using the `cupyx` library (this decreased computation time by ~70-80%). I applied `min_particle_size = 0.3` and `max_particle_size=1.0` filters on the masks created from watershed algorithm. The filter of `min_particle_size = 0.3` was key to filtering out many false positives, and noise created by the interpretation of artefacts as particles (see below TS_99_9, z-slice=7):\n![TS_99_9_zslice_7](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3784115%2F46d5a3957fdbcdffacb6439c44d6ce4d%2Fexample.JPG?generation=1738924826529887&alt=media)\n\n\n\n\n## What I tried to implement\n\nI spent sometime implementing the focal loss from scratch from the CentreNet paper and use guassian weigthed heatmaps, but was unable to a 3D FCNN model to get any meaningful results. I saw the first solution took this approach initially with impressive results, so I know it is possible.\n\nAs a last minute stretch and inspired by other solutions, in the last 3 days, I attempted to bring a 2D object detection (ResNet18 backbone, R-FasterCNN RPN) model into the pipeline to filter out false positives in 2D plane after the watershed algorithm selected particle centers from 3D U-Net masks. I had time only to run 3 training iterations, and thus the performance was far too poor to provide any meaningful results. To improve the metric it's 4 false positives removed for every 1 false negative introduced, and the 2D object det model did not reach this benchmark.\n\n### Other things that did not work\n- Struggled to get anything meaningful out of the synthetic data\n- Heatmaps with Focal/CE Loss\n- Deeper 3D U-Nets\n- Attention U-Net: I thought introducing spatial attention, might benefit this task by filtering out irrelevant feature maps that introduce false positives, but I got slightly worse results than the regular 3D U-Net implemented by Monai. My rationale for this was, it might be due to difference in decoder upsample, where Monai's imeplmentation of the Attention U-Net differs from the original paper where the encoder is concatenated to the decoder before a transposed convolution (instead of after like the standard 3D U-Net implemented by Monai).\n\n### Things I wanted to try\n- Copy Paste augmentation suggested by @hengck23.\n- More effort into 2D object det model\n\n## Conclusion\nThis was a lot of fun, and I have a lot to learn from some fantastic solutions. It might be time to invest in my own GPU setup, flicking between Google Colab and Kaggle Kernels got a bit tiring 😄",
    "3117915": "homiecal Hello and congratulations on your medal! May I ask you about the watershed? I remember attempting it myself but having memory issues with it? I also don’t recall seeing it in the competition provided starter solution. Thanks!",
    "3119410": ">Also many thanks to @davidlist, the CryoET personal tutor, and congratulations on a gold medal, well deserved.  \n\nFunny.  😀  And thank you!  Congrats on your first medal!!!",
    "3120286": "Hey Andrei, here's a snippet of the function below. I inference using the full image size (184, 630, 630) and manage the GPU contraints by moving operations to the CPU where possible, and deleting data throughout the pipeline. You can clog up memory quite quickly, because each step of the process creates a new tensor (binary dilation, erosion etc.).\n\nThe core of this function can be found in the [copick-utils](https://github.com/copick/copick-utils/blob/main/src/copick_utils/segmentation/picks_from_segmentation.py) repo `picks_from_segmentation()`.\n\n```python\nimport torch\nimport cupy as cp\nfrom cupyx.scipy.ndimage import distance_transform_edt, label, maximum_filter, binary_erosion, binary_dilation\nfrom skimage.segmentation import watershed\nfrom skimage.measure import regionprops\nfrom skimage.morphology import ball\n\ndef picks_from_segmentation_time_gpu(\n    segmentation,\n    segmentation_idx,\n    maxima_filter_size,\n    min_particle_size,\n    max_particle_size,\n    session_id,\n    user_id,\n    pickable_object,\n    run,\n    voxel_spacing=1,\n    submission: bool = False,\n):\n    \"\"\"\n    Process a specific label in the segmentation, extract centroids, and save them as picks.\n\n    Args:\n        segmentation (np.ndarray): Multilabel segmentation array.\n        segmentation_idx (int): The specific label from the segmentation to process.\n        maxima_filter_size (int): Size of the maximum detection filter.\n        min_particle_size (int): Minimum size threshold for particles.\n        max_particle_size (int): Maximum size threshold for particles.\n        session_id (str): Session ID for pick saving.\n        user_id (str): User ID for pick saving.\n        pickable_object (str): The name of the object to save picks for.\n        run: A Copick run object that manages pick saving.\n        voxel_spacing (int): The voxel spacing used to scale pick locations (default 1).\n    \"\"\"\n    log = []  # Logging list for timestamps\n\n    def log_step(step_name, start_time):\n        end_time = time.time()\n        log.append({\"step\": step_name, \"duration\": np.round(end_time - start_time, 5)})\n\n    start_time = time.time()\n    log_step(\"Start function\", start_time)\n    segmentation_gpu = cp.asarray(segmentation)\n\n    # Create a binary mask for the specific segmentation label\n    step_start = time.time()\n    # Create a binary mask for the specific segmentation label\n    binary_mask = (segmentation_gpu == segmentation_idx).astype(cp.int32)\n\n    # Skip if the segmentation label is not present\n    if np.sum(binary_mask) == 0:\n        log_step(\"No segmentation label found\", step_start)\n        print(f\"No segmentation with label {segmentation_idx} found.\")\n        return log\n\n    log_step(\"Binary mask created\", step_start)\n\n    # Structuring element for erosion and dilation\n    step_start = time.time()\n    struct_elem = ball(1)\n    struct_elem_gpu = cp.asarray(struct_elem)\n    eroded = cp.asarray(binary_erosion(binary_mask, struct_elem_gpu))\n    dilated = cp.asarray(binary_dilation(eroded, struct_elem_gpu))\n    log_step(\"Erosion and dilation\", step_start)\n\n    # Release GPU memory\n    del eroded\n\n    # Distance transform and local maxima detection\n    step_start = time.time()\n    # Distance transform and local maxima detection\n    distance = distance_transform_edt(dilated)\n    log_step(\"Distance transform\", step_start)\n\n    step_start = time.time()\n    local_max = distance == maximum_filter(distance, size=maxima_filter_size)\n    log_step(\"Local maxima detection\", step_start)\n\n    # Watershed segmentation\n    step_start = time.time()\n    markers, _ = label(local_max)\n    # Release GPU memory\n    del local_max\n    watershed_labels = watershed(-distance.get(), markers.get(), mask=dilated.get())\n    watershed_labels_cpu = watershed_labels.astype(int)\n    # Release GPU memory\n    del watershed_labels, distance, markers, dilated\n    log_step(\"Watershed segmentation\", step_start)\n\n    # Extract region properties and filter based on particle size\n    step_start = time.time()\n    all_centroids = []\n    for region in regionprops(watershed_labels_cpu):\n        if min_particle_size <= region.area <= max_particle_size:\n            all_centroids.append(region.centroid)\n    log_step(\"Region properties extracted and filtered\", step_start)\n    del watershed_labels_cpu\n\n    # Save centroids as picks\n    step_start = time.time()\n    if all_centroids:\n        if submission:\n            return [{\n                    \"experiment\": run.name,\n                    \"particle_type\": pickable_object,\n                    \"x\": c[2] * voxel_spacing,\n                    \"y\": c[1] * voxel_spacing,\n                    \"z\": c[0] * voxel_spacing,\n                }\n                for c in all_centroids\n            ]   \n        else:\n            pick_set = run.new_picks(pickable_object, session_id, user_id)\n\n            pick_set.points = [\n                CopickPoint(\n                    location=CopickLocation(\n                        **{\n                            \"x\": c[2] * voxel_spacing,\n                            \"y\": c[1] * voxel_spacing,\n                            \"z\": c[0] * voxel_spacing,\n                        }\n                    )\n                )\n                for c in all_centroids\n            ]\n            pick_set.store()\n            log_step(\"Centroids saved successfully\", step_start)\n            print(f\"Centroids for label {segmentation_idx} saved successfully.\")\n            print(f\"Log time stamps for {segmentation_idx}: {log}\")\n            return log\n\n\n    else:\n        log_step(\"No valid centroids found\", step_start)\n        print(f\"No valid centroids found for label {segmentation_idx}.\")\n        return log\n```"
  },
  "source": "meta"
}