{
  "id": 561402,
  "title": "33 place solution",
  "url": "/competitions/czii-cryo-et-object-identification/writeups/code-hacker-33-place-solution",
  "author_name": "",
  "post_date": "2025-02-06T00:23:06.090Z",
  "votes": 19,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I use Data Augmentation</p>\n<p><code>random_transforms = Compose([\n    RandCropByLabelClassesd(\n        keys=[\"image\", \"label\"],\n        label_key=\"label\",\n        spatial_size=[128, 128, 128],\n        num_classes=7,\n        num_samples=my_num_samples\n    ),\n    RandRotate90d(\n        keys=[\"image\", \"label\"],\n        prob=0.3,\n        spatial_axes=[0, 1]\n    ),\n    RandRotate90d(\n        keys=[\"image\", \"label\"],\n        prob=0.2,\n        spatial_axes=[1, 2]\n    ),\n    RandFlipd(\n        keys=[\"image\", \"label\"],\n        prob=0.3,\n        spatial_axis=0\n    ),\n    RandFlipd(\n        keys=[\"image\", \"label\"],\n        prob=0.3,\n        spatial_axis=1\n    ),\n    # Optionally, you can also flip along the third axis:\n    # RandFlipd(\n    #     keys=[\"image\", \"label\"],\n    #     prob=0.3,\n    #     spatial_axis=2\n    # ),\n    RandAffined(\n        keys=[\"image\", \"label\"],\n        prob=0.5,\n        rotate_range=(0.17, 0.17, 0.17),\n        scale_range=(0.05, 0.05, 0.05),\n        mode=(\"bilinear\", \"nearest\"),\n        padding_mode=\"zeros\"\n    ),\n    RandScaleIntensityd(\n        keys=\"image\",\n        prob=0.2,\n        factors=0.1\n    ),\n    RandShiftIntensityd(\n        keys=\"image\",\n        prob=0.2,\n        offsets=0.1\n    ),\n    RandAdjustContrastd(\n        keys=\"image\",\n        prob=0.2,\n        gamma=(0.9, 1.1)\n    ),\n    RandHistogramShiftd(\n        keys=\"image\",\n        prob=0.2,\n        num_control_points=10\n    )\n])\n</code></p>\n<p><strong>Model Architecture</strong><br>\nMy model is based on the MONAI UNet and uses the following configurations:</p>\n<p>Channels: (64, 128, 256, 256)<br>\nStrides Pattern: (2, 2, 1)<br>\nNumber of Residual Units: 1<br>\nDue to limited GPU resources, I utilized an L4 GPU.</p>\n<p><strong>Performance Results</strong><br>\nPure 3D UNet: Achieved a public score of 0.722 and a private score of 0.726.<br>\nEnsemble with YOLO: Reached a public score of 0.757 and a private score of 0.755.</p>\n<p><strong>Additional Enhancements</strong><br>\nI also incorporated a filtering mechanism that ignores any cluster where the standard deviation of each coordinate exceeds 40% of the particle radius.</p>\n<p>For further details, please refer to the notebook.<br>\n<a href=\"https://www.kaggle.com/code/junhanzangai/czii-cryo-s\" target=\"_blank\">https://www.kaggle.com/code/junhanzangai/czii-cryo-s</a><br>\n<a href=\"https://www.kaggle.com/code/junhanzangai/submission-test\\\" target=\"_blank\">https://www.kaggle.com/code/junhanzangai/submission-test\\</a></p>\n<p>And i attach my code, too.</p>\n<p>Thank you for giving me this opportunity.</p>",
  "messages": [
    {
      "id": "3116419",
      "postDate": "02/06/2025 00:21:54",
      "content": "<p>I use Data Augmentation</p>\n<p><code>random_transforms = Compose([\n    RandCropByLabelClassesd(\n        keys=[\"image\", \"label\"],\n        label_key=\"label\",\n        spatial_size=[128, 128, 128],\n        num_classes=7,\n        num_samples=my_num_samples\n    ),\n    RandRotate90d(\n        keys=[\"image\", \"label\"],\n        prob=0.3,\n        spatial_axes=[0, 1]\n    ),\n    RandRotate90d(\n        keys=[\"image\", \"label\"],\n        prob=0.2,\n        spatial_axes=[1, 2]\n    ),\n    RandFlipd(\n        keys=[\"image\", \"label\"],\n        prob=0.3,\n        spatial_axis=0\n    ),\n    RandFlipd(\n        keys=[\"image\", \"label\"],\n        prob=0.3,\n        spatial_axis=1\n    ),\n    # Optionally, you can also flip along the third axis:\n    # RandFlipd(\n    #     keys=[\"image\", \"label\"],\n    #     prob=0.3,\n    #     spatial_axis=2\n    # ),\n    RandAffined(\n        keys=[\"image\", \"label\"],\n        prob=0.5,\n        rotate_range=(0.17, 0.17, 0.17),\n        scale_range=(0.05, 0.05, 0.05),\n        mode=(\"bilinear\", \"nearest\"),\n        padding_mode=\"zeros\"\n    ),\n    RandScaleIntensityd(\n        keys=\"image\",\n        prob=0.2,\n        factors=0.1\n    ),\n    RandShiftIntensityd(\n        keys=\"image\",\n        prob=0.2,\n        offsets=0.1\n    ),\n    RandAdjustContrastd(\n        keys=\"image\",\n        prob=0.2,\n        gamma=(0.9, 1.1)\n    ),\n    RandHistogramShiftd(\n        keys=\"image\",\n        prob=0.2,\n        num_control_points=10\n    )\n])\n</code></p>\n<p><strong>Model Architecture</strong><br>\nMy model is based on the MONAI UNet and uses the following configurations:</p>\n<p>Channels: (64, 128, 256, 256)<br>\nStrides Pattern: (2, 2, 1)<br>\nNumber of Residual Units: 1<br>\nDue to limited GPU resources, I utilized an L4 GPU.</p>\n<p><strong>Performance Results</strong><br>\nPure 3D UNet: Achieved a public score of 0.722 and a private score of 0.726.<br>\nEnsemble with YOLO: Reached a public score of 0.757 and a private score of 0.755.</p>\n<p><strong>Additional Enhancements</strong><br>\nI also incorporated a filtering mechanism that ignores any cluster where the standard deviation of each coordinate exceeds 40% of the particle radius.</p>\n<p>For further details, please refer to the notebook.<br>\n<a href=\"https://www.kaggle.com/code/junhanzangai/czii-cryo-s\" target=\"_blank\">https://www.kaggle.com/code/junhanzangai/czii-cryo-s</a><br>\n<a href=\"https://www.kaggle.com/code/junhanzangai/submission-test\\\" target=\"_blank\">https://www.kaggle.com/code/junhanzangai/submission-test\\</a></p>\n<p>And i attach my code, too.</p>\n<p>Thank you for giving me this opportunity.</p>",
      "rawMarkdown": "I use Data Augmentation\n\n`random_transforms = Compose([\n    RandCropByLabelClassesd(\n        keys=[\"image\", \"label\"],\n        label_key=\"label\",\n        spatial_size=[128, 128, 128],\n        num_classes=7,\n        num_samples=my_num_samples\n    ),\n    RandRotate90d(\n        keys=[\"image\", \"label\"],\n        prob=0.3,\n        spatial_axes=[0, 1]\n    ),\n    RandRotate90d(\n        keys=[\"image\", \"label\"],\n        prob=0.2,\n        spatial_axes=[1, 2]\n    ),\n    RandFlipd(\n        keys=[\"image\", \"label\"],\n        prob=0.3,\n        spatial_axis=0\n    ),\n    RandFlipd(\n        keys=[\"image\", \"label\"],\n        prob=0.3,\n        spatial_axis=1\n    ),\n    # Optionally, you can also flip along the third axis:\n    # RandFlipd(\n    #     keys=[\"image\", \"label\"],\n    #     prob=0.3,\n    #     spatial_axis=2\n    # ),\n    RandAffined(\n        keys=[\"image\", \"label\"],\n        prob=0.5,\n        rotate_range=(0.17, 0.17, 0.17),\n        scale_range=(0.05, 0.05, 0.05),\n        mode=(\"bilinear\", \"nearest\"),\n        padding_mode=\"zeros\"\n    ),\n    RandScaleIntensityd(\n        keys=\"image\",\n        prob=0.2,\n        factors=0.1\n    ),\n    RandShiftIntensityd(\n        keys=\"image\",\n        prob=0.2,\n        offsets=0.1\n    ),\n    RandAdjustContrastd(\n        keys=\"image\",\n        prob=0.2,\n        gamma=(0.9, 1.1)\n    ),\n    RandHistogramShiftd(\n        keys=\"image\",\n        prob=0.2,\n        num_control_points=10\n    )\n])\n`\n\n**Model Architecture**\nMy model is based on the MONAI UNet and uses the following configurations:\n\nChannels: (64, 128, 256, 256)\nStrides Pattern: (2, 2, 1)\nNumber of Residual Units: 1\nDue to limited GPU resources, I utilized an L4 GPU.\n\n**Performance Results**\nPure 3D UNet: Achieved a public score of 0.722 and a private score of 0.726.\nEnsemble with YOLO: Reached a public score of 0.757 and a private score of 0.755.\n\n**Additional Enhancements**\nI also incorporated a filtering mechanism that ignores any cluster where the standard deviation of each coordinate exceeds 40% of the particle radius.\n\nFor further details, please refer to the notebook.\nhttps://www.kaggle.com/code/junhanzangai/czii-cryo-s\nhttps://www.kaggle.com/code/junhanzangai/submission-test\\\n\nAnd i attach my code, too.\n\nThank you for giving me this opportunity.",
      "votes": null
    },
    {
      "id": "3116421",
      "postDate": "02/06/2025 00:25:22",
      "content": "<p>Oh, wow!  I really like that filter idea!</p>",
      "rawMarkdown": "Oh, wow!  I really like that filter idea!",
      "votes": null
    },
    {
      "id": "3116425",
      "postDate": "02/06/2025 00:29:45",
      "content": "<p>Congratulations! Wow, I didn't expect YOLO to boost the scores that much. That's amazing! </p>",
      "rawMarkdown": "Congratulations! Wow, I didn't expect YOLO to boost the scores that much. That's amazing!",
      "votes": null
    },
    {
      "id": "3116442",
      "postDate": "02/06/2025 00:56:18",
      "content": "<p>In my case, filtering id is main for scoring up.</p>",
      "rawMarkdown": "In my case, filtering id is main for scoring up.",
      "votes": null
    },
    {
      "id": "3116443",
      "postDate": "02/06/2025 00:58:02",
      "content": "<p>In fact, as I working with 3D, I  realized that 2.5D could have a good impact.</p>",
      "rawMarkdown": "In fact, as I working with 3D, I  realized that 2.5D could have a good impact.",
      "votes": null
    },
    {
      "id": "3116449",
      "postDate": "02/06/2025 01:01:15",
      "content": "<p>so if i understand correctly was that mostly getting rid of non-spherical particles?  or rather particle candidates i suppose.</p>",
      "rawMarkdown": "so if i understand correctly was that mostly getting rid of non-spherical particles?  or rather particle candidates i suppose.",
      "votes": null
    },
    {
      "id": "3116458",
      "postDate": "02/06/2025 01:14:56",
      "content": "<p>Your understanding is partly correct but there's a nuance. The filtering mechanism isn't intended solely to remove non-spherical particles. Instead, it's designed to ignore any candidate cluster where the standard deviation along any coordinate exceeds 40% of the expected particle radius. This serves as a proxy for detecting clusters with overly dispersed predictions—which likely indicates unstable or false-positive detections—rather than a direct assessment of the particle’s inherent shape. In other words, it's more about ensuring the predicted candidate clusters are spatially compact and consistent with what you'd expect for a true particle detection.</p>\n<p>Additionally, I used the following code to filter out as many genuine particles as possible, but the score actually dropped.<br>\n`BLOB_THRESHOLD = 250<br>\nclasses = [1, 2, 3, 4, 5, 6]<br>\ntotal_start = time.time()</p>\n<p>with torch.no_grad():<br>\n    location_df = []</p>\n<pre><code> run  root.runs:\n    run_start = time.time()  \n\n    (run)\n\n    \n    load_start = time.time()\n    tomo = run.get_voxel_spacing()\n    tomo_arr = tomo.get_tomogram(tomo_type).numpy()  \n    load_end = time.time()\n    ()\n\n    \n    prep_start = time.time()\n    data_dict = [{: tomo_arr}]\n    tomo_ds = CacheDataset(data=data_dict, transform=inference_transforms, cache_rate=)\n    volume_tensor = tomo_ds[][].unsqueeze().to()  \n    prep_end = time.time()\n    ()\n\n    \n    \n    infer_start = time.time()\n    out_logits = sliding_window_inference(\n        inputs=volume_tensor,\n        roi_size=(, , ),\n        sw_batch_size=,\n        predictor=ensemble_tta_predictor,  \n        overlap=,\n        mode=\n    )\n    infer_end = time.time()\n    ()\n\n    \n    post_start = time.time()\n    out_probs = torch.softmax(out_logits, dim=)  \n    out_probs_np = out_probs[].cpu().numpy()      \n    reconstructed_mask = np.argmax(out_probs_np, axis=)  \n    post_end = time.time()\n    ()\n\n    \n    aspect_ratio_thresholds = {\n        : ,\n        : ,\n        : ,\n        : ,\n        : \n    }\n\n    cc_start = time.time()\n    location = {}\n\n    \n     c  classes:\n        \n        cc = cc3d.connected_components(reconstructed_mask == c)\n        stats = cc3d.statistics(cc)\n\n        \n        centroids = np.array(stats[][:])  \n        bbox_list = stats[][:]         \n        voxel_counts = np.array(stats[][:])\n\n        valid_indices = []\n        \n        class_name = id_to_name[c]\n        ar_threshold = aspect_ratio_thresholds.get(class_name, )\n\n         i, bbox  (bbox_list):\n            \n            \n            slice_z, slice_y, slice_x = bbox  \n            depth  = slice_z.stop - slice_z.start\n            height = slice_y.stop - slice_y.start\n            width  = slice_x.stop - slice_x.start\n\n            \n            aspect_ratio = (width, height, depth) / (width, height, depth)\n\n            \n             voxel_counts[i] &gt; BLOB_THRESHOLD  aspect_ratio &gt;= ar_threshold:\n                valid_indices.append(i)\n\n         (valid_indices) &gt; :\n            \n            valid_centroids = centroids[valid_indices]\n            \n            valid_centroids = valid_centroids *   \n            \n            valid_xyz = np.ascontiguousarray(valid_centroids[:, ::-])\n            location[class_name] = valid_xyz\n\n    cc_end = time.time()\n    ()\n\n    \n    df = dict_to_df(location, run.name)\n    location_df.append(df)\n\n    run_end = time.time()\n    ()\n\n\nlocation_df = pd.concat(location_df)\n</code></pre>\n<h1>End timing of the entire pipeline</h1>\n<p>total_end = time.time()<br>\n`</p>",
      "rawMarkdown": "Your understanding is partly correct but there's a nuance. The filtering mechanism isn't intended solely to remove non-spherical particles. Instead, it's designed to ignore any candidate cluster where the standard deviation along any coordinate exceeds 40% of the expected particle radius. This serves as a proxy for detecting clusters with overly dispersed predictions—which likely indicates unstable or false-positive detections—rather than a direct assessment of the particle’s inherent shape. In other words, it's more about ensuring the predicted candidate clusters are spatially compact and consistent with what you'd expect for a true particle detection.\n\nAdditionally, I used the following code to filter out as many genuine particles as possible, but the score actually dropped.\n`BLOB_THRESHOLD = 250\nclasses = [1, 2, 3, 4, 5, 6]\ntotal_start = time.time()\n\nwith torch.no_grad():\n    location_df = []\n\n    for run in root.runs:\n        run_start = time.time()  # Start time for a single run\n\n        print(run)\n\n        # 1) Load volume (10Å voxel)\n        load_start = time.time()\n        tomo = run.get_voxel_spacing(10)\n        tomo_arr = tomo.get_tomogram(tomo_type).numpy()  # shape: (X, Y, Z)\n        load_end = time.time()\n        print(f\"[Timer] Volume load time: {load_end - load_start:.3f} sec\")\n\n        # 2) Load dataset (preprocessing)\n        prep_start = time.time()\n        data_dict = [{\"image\": tomo_arr}]\n        tomo_ds = CacheDataset(data=data_dict, transform=inference_transforms, cache_rate=1.0)\n        volume_tensor = tomo_ds[0][\"image\"].unsqueeze(0).to(\"cuda\")  # (1,1,X,Y,Z)\n        prep_end = time.time()\n        print(f\"[Timer] Dataset prep time: {prep_end - prep_start:.3f} sec\")\n\n        # 3) Sliding Window Inference\n        #    -> Replace with predictor=ensemble_tta_predictor\n        infer_start = time.time()\n        out_logits = sliding_window_inference(\n            inputs=volume_tensor,\n            roi_size=(128, 128, 128),\n            sw_batch_size=6,\n            predictor=ensemble_tta_predictor,  # <-- Performing ensemble + TTA here\n            overlap=0.25,\n            mode=\"gaussian\"\n        )\n        infer_end = time.time()\n        print(f\"[Timer] SW Inference (Ensemble+TTA) time: {infer_end - infer_start:.3f} sec\")\n\n        # 4) Softmax then argmax\n        post_start = time.time()\n        out_probs = torch.softmax(out_logits, dim=1)  # (1,7,X,Y,Z)\n        out_probs_np = out_probs[0].cpu().numpy()      # (7, X, Y, Z)\n        reconstructed_mask = np.argmax(out_probs_np, axis=0)  # (X, Y, Z)\n        post_end = time.time()\n        print(f\"[Timer] Postprocess (softmax+argmax) time: {post_end - post_start:.3f} sec\")\n\n        # Set aspect ratio thresholds for each class (reflecting actual shape information)\n        aspect_ratio_thresholds = {\n            \"apo-ferritin\": 0.8,\n            \"beta-galactosidase\": 0.6,\n            \"ribosome\": 0.65,\n            \"thyroglobulin\": 0.6,\n            \"virus-like-particle\": 0.8\n        }\n        \n        cc_start = time.time()\n        location = {}\n        \n        # Assume classes are numbered as [1, 2, 3, 4, 5, 6] and map them to the actual class names using id_to_name.\n        for c in classes:\n            # Perform connected components analysis on regions where reconstructed_mask == c\n            cc = cc3d.connected_components(reconstructed_mask == c)\n            stats = cc3d.statistics(cc)\n            \n            # Exclude the background label (0) and use indices starting from 1 for actual objects\n            centroids = np.array(stats[\"centroids\"][1:])  # Typically returned in [z, y, x] order\n            bbox_list = stats[\"bounding_boxes\"][1:]         # Bounding box for each object, format: (slice_z, slice_y, slice_x)\n            voxel_counts = np.array(stats[\"voxel_counts\"][1:])\n            \n            valid_indices = []\n            # Set the class name and threshold (default 0.7 if not specified)\n            class_name = id_to_name[c]\n            ar_threshold = aspect_ratio_thresholds.get(class_name, 0.7)\n            \n            for i, bbox in enumerate(bbox_list):\n                # Each bounding box is in the form (slice_z, slice_y, slice_x)\n                # Calculate the length using the start and stop values of each slice object.\n                slice_z, slice_y, slice_x = bbox  # Slice object for each axis\n                depth  = slice_z.stop - slice_z.start\n                height = slice_y.stop - slice_y.start\n                width  = slice_x.stop - slice_x.start\n                \n                # Aspect ratio: Divide the minimum length by the maximum length (closer to 1 indicates a more spherical shape)\n                aspect_ratio = min(width, height, depth) / max(width, height, depth)\n                \n                # Condition: Voxel count is greater than BLOB_THRESHOLD and aspect ratio exceeds the class-specific threshold\n                if voxel_counts[i] > BLOB_THRESHOLD and aspect_ratio >= ar_threshold:\n                    valid_indices.append(i)\n            \n            if len(valid_indices) > 0:\n                # Select the centroids of the valid objects\n                valid_centroids = centroids[valid_indices]\n                # Multiply by a factor (e.g., 10.012444) for voxel size correction\n                valid_centroids = valid_centroids * 10.012444  \n                # If the cc3d result is in [z, y, x] order, change it to [x, y, z]\n                valid_xyz = np.ascontiguousarray(valid_centroids[:, ::-1])\n                location[class_name] = valid_xyz\n        \n        cc_end = time.time()\n        print(f\"[Timer] Connected components + centroids + shape filtering time: {cc_end - cc_start:.3f} sec\")\n        \n        # Convert the result to a DataFrame using dict_to_df and store it in location_df.\n        df = dict_to_df(location, run.name)\n        location_df.append(df)\n\n        run_end = time.time()\n        print(f\"[Timer] Single run total time: {run_end - run_start:.3f} sec\\n\")\n\n    # Concatenate results from all runs\n    location_df = pd.concat(location_df)\n\n# End timing of the entire pipeline\ntotal_end = time.time()\n`",
      "votes": null
    },
    {
      "id": "3116461",
      "postDate": "02/06/2025 01:30:16",
      "content": "<p>Yeah, that makes sense.  Thanks!</p>",
      "rawMarkdown": "Yeah, that makes sense.  Thanks!",
      "votes": null
    },
    {
      "id": "3116945",
      "postDate": "02/06/2025 13:03:35",
      "content": "<p>congratulation!</p>",
      "rawMarkdown": "congratulation!",
      "votes": null
    },
    {
      "id": "3117064",
      "postDate": "02/06/2025 15:47:06",
      "content": "<p>Congratulations! I also used 3DUnet model, but it didn`t do so well. Would you mind telling tips for hyperparameter tuning?</p>",
      "rawMarkdown": "Congratulations! I also used 3DUnet model, but it didn`t do so well. Would you mind telling tips for hyperparameter tuning?",
      "votes": null
    },
    {
      "id": "3117080",
      "postDate": "02/06/2025 16:06:58",
      "content": "<p>Similar to #8(<a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561515)\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561515)</a>, i found dropout, batchnorm and channels to be key. The increase in channels affected my model, so I spent a lot of time optimizing them. If you look at my notes, you'll see that there is learning for many different channels.</p>\n<p><a href=\"https://www.kaggle.com/code/junhanzangai/submission-test/\" target=\"_blank\">https://www.kaggle.com/code/junhanzangai/submission-test/</a></p>",
      "rawMarkdown": "Similar to #8(https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561515), i found dropout, batchnorm and channels to be key. The increase in channels affected my model, so I spent a lot of time optimizing them. If you look at my notes, you'll see that there is learning for many different channels.\n\nhttps://www.kaggle.com/code/junhanzangai/submission-test/",
      "votes": null
    },
    {
      "id": "3117387",
      "postDate": "02/07/2025 00:03:21",
      "content": "<p>thank you for great sharing!</p>",
      "rawMarkdown": "thank you for great sharing!",
      "votes": null
    },
    {
      "id": "3179351",
      "postDate": "04/15/2025 10:50:27",
      "content": "<p>Hello, Thank you for very informative sharing!!!<br>\nI'm studying YOLO now. If possible, could you share the train code of YOLO?</p>",
      "rawMarkdown": "Hello, Thank you for very informative sharing!!!\nI'm studying YOLO now. If possible, could you share the train code of YOLO?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3116421,
      "author_name": "davidlist",
      "author_url": "",
      "post_date": "02/06/2025 00:25:22",
      "content": "<p>Oh, wow!  I really like that filter idea!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3116442,
          "author_name": "junhanzangai",
          "author_url": "",
          "post_date": "02/06/2025 00:56:18",
          "content": "<p>In my case, filtering id is main for scoring up.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3116449,
              "author_name": "davidlist",
              "author_url": "",
              "post_date": "02/06/2025 01:01:15",
              "content": "<p>so if i understand correctly was that mostly getting rid of non-spherical particles?  or rather particle candidates i suppose.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3116458,
                  "author_name": "junhanzangai",
                  "author_url": "",
                  "post_date": "02/06/2025 01:14:56",
                  "content": "<p>Your understanding is partly correct but there's a nuance. The filtering mechanism isn't intended solely to remove non-spherical particles. Instead, it's designed to ignore any candidate cluster where the standard deviation along any coordinate exceeds 40% of the expected particle radius. This serves as a proxy for detecting clusters with overly dispersed predictions—which likely indicates unstable or false-positive detections—rather than a direct assessment of the particle’s inherent shape. In other words, it's more about ensuring the predicted candidate clusters are spatially compact and consistent with what you'd expect for a true particle detection.</p>\n<p>Additionally, I used the following code to filter out as many genuine particles as possible, but the score actually dropped.<br>\n`BLOB_THRESHOLD = 250<br>\nclasses = [1, 2, 3, 4, 5, 6]<br>\ntotal_start = time.time()</p>\n<p>with torch.no_grad():<br>\n    location_df = []</p>\n<pre><code> run  root.runs:\n    run_start = time.time()  \n\n    (run)\n\n    \n    load_start = time.time()\n    tomo = run.get_voxel_spacing()\n    tomo_arr = tomo.get_tomogram(tomo_type).numpy()  \n    load_end = time.time()\n    ()\n\n    \n    prep_start = time.time()\n    data_dict = [{: tomo_arr}]\n    tomo_ds = CacheDataset(data=data_dict, transform=inference_transforms, cache_rate=)\n    volume_tensor = tomo_ds[][].unsqueeze().to()  \n    prep_end = time.time()\n    ()\n\n    \n    \n    infer_start = time.time()\n    out_logits = sliding_window_inference(\n        inputs=volume_tensor,\n        roi_size=(, , ),\n        sw_batch_size=,\n        predictor=ensemble_tta_predictor,  \n        overlap=,\n        mode=\n    )\n    infer_end = time.time()\n    ()\n\n    \n    post_start = time.time()\n    out_probs = torch.softmax(out_logits, dim=)  \n    out_probs_np = out_probs[].cpu().numpy()      \n    reconstructed_mask = np.argmax(out_probs_np, axis=)  \n    post_end = time.time()\n    ()\n\n    \n    aspect_ratio_thresholds = {\n        : ,\n        : ,\n        : ,\n        : ,\n        : \n    }\n\n    cc_start = time.time()\n    location = {}\n\n    \n     c  classes:\n        \n        cc = cc3d.connected_components(reconstructed_mask == c)\n        stats = cc3d.statistics(cc)\n\n        \n        centroids = np.array(stats[][:])  \n        bbox_list = stats[][:]         \n        voxel_counts = np.array(stats[][:])\n\n        valid_indices = []\n        \n        class_name = id_to_name[c]\n        ar_threshold = aspect_ratio_thresholds.get(class_name, )\n\n         i, bbox  (bbox_list):\n            \n            \n            slice_z, slice_y, slice_x = bbox  \n            depth  = slice_z.stop - slice_z.start\n            height = slice_y.stop - slice_y.start\n            width  = slice_x.stop - slice_x.start\n\n            \n            aspect_ratio = (width, height, depth) / (width, height, depth)\n\n            \n             voxel_counts[i] &gt; BLOB_THRESHOLD  aspect_ratio &gt;= ar_threshold:\n                valid_indices.append(i)\n\n         (valid_indices) &gt; :\n            \n            valid_centroids = centroids[valid_indices]\n            \n            valid_centroids = valid_centroids *   \n            \n            valid_xyz = np.ascontiguousarray(valid_centroids[:, ::-])\n            location[class_name] = valid_xyz\n\n    cc_end = time.time()\n    ()\n\n    \n    df = dict_to_df(location, run.name)\n    location_df.append(df)\n\n    run_end = time.time()\n    ()\n\n\nlocation_df = pd.concat(location_df)\n</code></pre>\n<h1>End timing of the entire pipeline</h1>\n<p>total_end = time.time()<br>\n`</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3116461,
                      "author_name": "davidlist",
                      "author_url": "",
                      "post_date": "02/06/2025 01:30:16",
                      "content": "<p>Yeah, that makes sense.  Thanks!</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3116425,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "02/06/2025 00:29:45",
      "content": "<p>Congratulations! Wow, I didn't expect YOLO to boost the scores that much. That's amazing! </p>",
      "votes": null,
      "replies": [
        {
          "id": 3116443,
          "author_name": "junhanzangai",
          "author_url": "",
          "post_date": "02/06/2025 00:58:02",
          "content": "<p>In fact, as I working with 3D, I  realized that 2.5D could have a good impact.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3116945,
      "author_name": "",
      "author_url": "",
      "post_date": "02/06/2025 13:03:35",
      "content": "<p>congratulation!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3117064,
      "author_name": "ayaha0619",
      "author_url": "",
      "post_date": "02/06/2025 15:47:06",
      "content": "<p>Congratulations! I also used 3DUnet model, but it didn`t do so well. Would you mind telling tips for hyperparameter tuning?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3117080,
          "author_name": "junhanzangai",
          "author_url": "",
          "post_date": "02/06/2025 16:06:58",
          "content": "<p>Similar to #8(<a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561515)\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561515)</a>, i found dropout, batchnorm and channels to be key. The increase in channels affected my model, so I spent a lot of time optimizing them. If you look at my notes, you'll see that there is learning for many different channels.</p>\n<p><a href=\"https://www.kaggle.com/code/junhanzangai/submission-test/\" target=\"_blank\">https://www.kaggle.com/code/junhanzangai/submission-test/</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 3117387,
              "author_name": "ayaha0619",
              "author_url": "",
              "post_date": "02/07/2025 00:03:21",
              "content": "<p>thank you for great sharing!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3179351,
      "author_name": "taiki2",
      "author_url": "",
      "post_date": "04/15/2025 10:50:27",
      "content": "<p>Hello, Thank you for very informative sharing!!!<br>\nI'm studying YOLO now. If possible, could you share the train code of YOLO?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3116419": "I use Data Augmentation\n\n`random_transforms = Compose([\n    RandCropByLabelClassesd(\n        keys=[\"image\", \"label\"],\n        label_key=\"label\",\n        spatial_size=[128, 128, 128],\n        num_classes=7,\n        num_samples=my_num_samples\n    ),\n    RandRotate90d(\n        keys=[\"image\", \"label\"],\n        prob=0.3,\n        spatial_axes=[0, 1]\n    ),\n    RandRotate90d(\n        keys=[\"image\", \"label\"],\n        prob=0.2,\n        spatial_axes=[1, 2]\n    ),\n    RandFlipd(\n        keys=[\"image\", \"label\"],\n        prob=0.3,\n        spatial_axis=0\n    ),\n    RandFlipd(\n        keys=[\"image\", \"label\"],\n        prob=0.3,\n        spatial_axis=1\n    ),\n    # Optionally, you can also flip along the third axis:\n    # RandFlipd(\n    #     keys=[\"image\", \"label\"],\n    #     prob=0.3,\n    #     spatial_axis=2\n    # ),\n    RandAffined(\n        keys=[\"image\", \"label\"],\n        prob=0.5,\n        rotate_range=(0.17, 0.17, 0.17),\n        scale_range=(0.05, 0.05, 0.05),\n        mode=(\"bilinear\", \"nearest\"),\n        padding_mode=\"zeros\"\n    ),\n    RandScaleIntensityd(\n        keys=\"image\",\n        prob=0.2,\n        factors=0.1\n    ),\n    RandShiftIntensityd(\n        keys=\"image\",\n        prob=0.2,\n        offsets=0.1\n    ),\n    RandAdjustContrastd(\n        keys=\"image\",\n        prob=0.2,\n        gamma=(0.9, 1.1)\n    ),\n    RandHistogramShiftd(\n        keys=\"image\",\n        prob=0.2,\n        num_control_points=10\n    )\n])\n`\n\n**Model Architecture**\nMy model is based on the MONAI UNet and uses the following configurations:\n\nChannels: (64, 128, 256, 256)\nStrides Pattern: (2, 2, 1)\nNumber of Residual Units: 1\nDue to limited GPU resources, I utilized an L4 GPU.\n\n**Performance Results**\nPure 3D UNet: Achieved a public score of 0.722 and a private score of 0.726.\nEnsemble with YOLO: Reached a public score of 0.757 and a private score of 0.755.\n\n**Additional Enhancements**\nI also incorporated a filtering mechanism that ignores any cluster where the standard deviation of each coordinate exceeds 40% of the particle radius.\n\nFor further details, please refer to the notebook.\nhttps://www.kaggle.com/code/junhanzangai/czii-cryo-s\nhttps://www.kaggle.com/code/junhanzangai/submission-test\\\n\nAnd i attach my code, too.\n\nThank you for giving me this opportunity.",
    "3116421": "Oh, wow!  I really like that filter idea!",
    "3116425": "Congratulations! Wow, I didn't expect YOLO to boost the scores that much. That's amazing!",
    "3116442": "In my case, filtering id is main for scoring up.",
    "3116443": "In fact, as I working with 3D, I  realized that 2.5D could have a good impact.",
    "3116449": "so if i understand correctly was that mostly getting rid of non-spherical particles?  or rather particle candidates i suppose.",
    "3116458": "Your understanding is partly correct but there's a nuance. The filtering mechanism isn't intended solely to remove non-spherical particles. Instead, it's designed to ignore any candidate cluster where the standard deviation along any coordinate exceeds 40% of the expected particle radius. This serves as a proxy for detecting clusters with overly dispersed predictions—which likely indicates unstable or false-positive detections—rather than a direct assessment of the particle’s inherent shape. In other words, it's more about ensuring the predicted candidate clusters are spatially compact and consistent with what you'd expect for a true particle detection.\n\nAdditionally, I used the following code to filter out as many genuine particles as possible, but the score actually dropped.\n`BLOB_THRESHOLD = 250\nclasses = [1, 2, 3, 4, 5, 6]\ntotal_start = time.time()\n\nwith torch.no_grad():\n    location_df = []\n\n    for run in root.runs:\n        run_start = time.time()  # Start time for a single run\n\n        print(run)\n\n        # 1) Load volume (10Å voxel)\n        load_start = time.time()\n        tomo = run.get_voxel_spacing(10)\n        tomo_arr = tomo.get_tomogram(tomo_type).numpy()  # shape: (X, Y, Z)\n        load_end = time.time()\n        print(f\"[Timer] Volume load time: {load_end - load_start:.3f} sec\")\n\n        # 2) Load dataset (preprocessing)\n        prep_start = time.time()\n        data_dict = [{\"image\": tomo_arr}]\n        tomo_ds = CacheDataset(data=data_dict, transform=inference_transforms, cache_rate=1.0)\n        volume_tensor = tomo_ds[0][\"image\"].unsqueeze(0).to(\"cuda\")  # (1,1,X,Y,Z)\n        prep_end = time.time()\n        print(f\"[Timer] Dataset prep time: {prep_end - prep_start:.3f} sec\")\n\n        # 3) Sliding Window Inference\n        #    -> Replace with predictor=ensemble_tta_predictor\n        infer_start = time.time()\n        out_logits = sliding_window_inference(\n            inputs=volume_tensor,\n            roi_size=(128, 128, 128),\n            sw_batch_size=6,\n            predictor=ensemble_tta_predictor,  # <-- Performing ensemble + TTA here\n            overlap=0.25,\n            mode=\"gaussian\"\n        )\n        infer_end = time.time()\n        print(f\"[Timer] SW Inference (Ensemble+TTA) time: {infer_end - infer_start:.3f} sec\")\n\n        # 4) Softmax then argmax\n        post_start = time.time()\n        out_probs = torch.softmax(out_logits, dim=1)  # (1,7,X,Y,Z)\n        out_probs_np = out_probs[0].cpu().numpy()      # (7, X, Y, Z)\n        reconstructed_mask = np.argmax(out_probs_np, axis=0)  # (X, Y, Z)\n        post_end = time.time()\n        print(f\"[Timer] Postprocess (softmax+argmax) time: {post_end - post_start:.3f} sec\")\n\n        # Set aspect ratio thresholds for each class (reflecting actual shape information)\n        aspect_ratio_thresholds = {\n            \"apo-ferritin\": 0.8,\n            \"beta-galactosidase\": 0.6,\n            \"ribosome\": 0.65,\n            \"thyroglobulin\": 0.6,\n            \"virus-like-particle\": 0.8\n        }\n        \n        cc_start = time.time()\n        location = {}\n        \n        # Assume classes are numbered as [1, 2, 3, 4, 5, 6] and map them to the actual class names using id_to_name.\n        for c in classes:\n            # Perform connected components analysis on regions where reconstructed_mask == c\n            cc = cc3d.connected_components(reconstructed_mask == c)\n            stats = cc3d.statistics(cc)\n            \n            # Exclude the background label (0) and use indices starting from 1 for actual objects\n            centroids = np.array(stats[\"centroids\"][1:])  # Typically returned in [z, y, x] order\n            bbox_list = stats[\"bounding_boxes\"][1:]         # Bounding box for each object, format: (slice_z, slice_y, slice_x)\n            voxel_counts = np.array(stats[\"voxel_counts\"][1:])\n            \n            valid_indices = []\n            # Set the class name and threshold (default 0.7 if not specified)\n            class_name = id_to_name[c]\n            ar_threshold = aspect_ratio_thresholds.get(class_name, 0.7)\n            \n            for i, bbox in enumerate(bbox_list):\n                # Each bounding box is in the form (slice_z, slice_y, slice_x)\n                # Calculate the length using the start and stop values of each slice object.\n                slice_z, slice_y, slice_x = bbox  # Slice object for each axis\n                depth  = slice_z.stop - slice_z.start\n                height = slice_y.stop - slice_y.start\n                width  = slice_x.stop - slice_x.start\n                \n                # Aspect ratio: Divide the minimum length by the maximum length (closer to 1 indicates a more spherical shape)\n                aspect_ratio = min(width, height, depth) / max(width, height, depth)\n                \n                # Condition: Voxel count is greater than BLOB_THRESHOLD and aspect ratio exceeds the class-specific threshold\n                if voxel_counts[i] > BLOB_THRESHOLD and aspect_ratio >= ar_threshold:\n                    valid_indices.append(i)\n            \n            if len(valid_indices) > 0:\n                # Select the centroids of the valid objects\n                valid_centroids = centroids[valid_indices]\n                # Multiply by a factor (e.g., 10.012444) for voxel size correction\n                valid_centroids = valid_centroids * 10.012444  \n                # If the cc3d result is in [z, y, x] order, change it to [x, y, z]\n                valid_xyz = np.ascontiguousarray(valid_centroids[:, ::-1])\n                location[class_name] = valid_xyz\n        \n        cc_end = time.time()\n        print(f\"[Timer] Connected components + centroids + shape filtering time: {cc_end - cc_start:.3f} sec\")\n        \n        # Convert the result to a DataFrame using dict_to_df and store it in location_df.\n        df = dict_to_df(location, run.name)\n        location_df.append(df)\n\n        run_end = time.time()\n        print(f\"[Timer] Single run total time: {run_end - run_start:.3f} sec\\n\")\n\n    # Concatenate results from all runs\n    location_df = pd.concat(location_df)\n\n# End timing of the entire pipeline\ntotal_end = time.time()\n`",
    "3116461": "Yeah, that makes sense.  Thanks!",
    "3116945": "congratulation!",
    "3117064": "Congratulations! I also used 3DUnet model, but it didn`t do so well. Would you mind telling tips for hyperparameter tuning?",
    "3117080": "Similar to #8(https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/561515), i found dropout, batchnorm and channels to be key. The increase in channels affected my model, so I spent a lot of time optimizing them. If you look at my notes, you'll see that there is learning for many different channels.\n\nhttps://www.kaggle.com/code/junhanzangai/submission-test/",
    "3117387": "thank you for great sharing!",
    "3179351": "Hello, Thank you for very informative sharing!!!\nI'm studying YOLO now. If possible, could you share the train code of YOLO?"
  },
  "source": "meta"
}