{
  "id": 578156,
  "title": "Brief Insights on Ensemble Functions",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/578156",
  "author_name": "Tom",
  "post_date": "2025-05-09T05:12:26.843000",
  "votes": 37,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Currently, I am developing ensemble functions for object localization. Here, I’d like to share some brief insights on the methods I’ve explored so far:</p>\n<h1>Average &amp; Weighted Sum</h1>\n<ul>\n<li>Works well when your models have low diversity.</li>\n<li>Performance may degrade if your models are highly diverse.</li>\n<li>Weight selection is done through brute-force search.</li>\n</ul>\n<h1>Maximum Function</h1>\n<ul>\n<li>Ensures that object presence is not missed.</li>\n<li>Crucially, it allows you to detect false positives from individual models — if the overall score drops after adding another model, it suggests that the new model is predicting false positives.</li>\n</ul>\n<h1>Clustering (HDBSCAN)</h1>\n<p>HDBSCAN is an effective tool for handling clusters with varying shapes and densities. I initially tried it and am now conducting deeper experiments on its properties for object localization.</p>\n<ul>\n<li>If a model predicts multiple keypoints in an area, there is a high possibility that an object is present.</li>\n<li>Suppresses peak signals, reducing false positives. It also helps verify whether your current score improvements rely on peak signals, which can cause instability and shake-up.</li>\n<li>Parameters like <code>threshold</code> from your model, <code>min_cluster_size</code>, and <code>min_samples</code> from HDBSCAN provide flexible control over prediction strength. Particularly effective for models trained with Gaussian balls or other irregularly shaped labels (e.g., keypoints annotated on motors or bacteria). I find it works especially well in 3D models, where the additional z-dimension improves clustering.</li>\n<li>Helps address prediction diversity across multiple models.</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>methods</th>\n<th>worse case LB</th>\n<th>best case LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>average</td>\n<td>0.69</td>\n<td>0.8</td>\n</tr>\n<tr>\n<td>maximum</td>\n<td>0.81</td>\n<td>0.84</td>\n</tr>\n<tr>\n<td>clustering</td>\n<td>0.76</td>\n<td>0.84</td>\n</tr>\n</tbody>\n</table>\n<h1>Notes</h1>\n<p>After reviewing the YOLO predictions using the maximum function and hdbscan, I found that YOLO actually produces many peak signals on the hidden test set (including almost all public notebooks and my best local model). This suggests that many high scores are the result of lucky predictions.</p>",
  "messages": [
    {
      "id": 3198136,
      "postDate": "2025-05-09T05:12:26.843Z",
      "content": "<p>Currently, I am developing ensemble functions for object localization. Here, I’d like to share some brief insights on the methods I’ve explored so far:</p>\n<h1>Average &amp; Weighted Sum</h1>\n<ul>\n<li>Works well when your models have low diversity.</li>\n<li>Performance may degrade if your models are highly diverse.</li>\n<li>Weight selection is done through brute-force search.</li>\n</ul>\n<h1>Maximum Function</h1>\n<ul>\n<li>Ensures that object presence is not missed.</li>\n<li>Crucially, it allows you to detect false positives from individual models — if the overall score drops after adding another model, it suggests that the new model is predicting false positives.</li>\n</ul>\n<h1>Clustering (HDBSCAN)</h1>\n<p>HDBSCAN is an effective tool for handling clusters with varying shapes and densities. I initially tried it and am now conducting deeper experiments on its properties for object localization.</p>\n<ul>\n<li>If a model predicts multiple keypoints in an area, there is a high possibility that an object is present.</li>\n<li>Suppresses peak signals, reducing false positives. It also helps verify whether your current score improvements rely on peak signals, which can cause instability and shake-up.</li>\n<li>Parameters like <code>threshold</code> from your model, <code>min_cluster_size</code>, and <code>min_samples</code> from HDBSCAN provide flexible control over prediction strength. Particularly effective for models trained with Gaussian balls or other irregularly shaped labels (e.g., keypoints annotated on motors or bacteria). I find it works especially well in 3D models, where the additional z-dimension improves clustering.</li>\n<li>Helps address prediction diversity across multiple models.</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>methods</th>\n<th>worse case LB</th>\n<th>best case LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>average</td>\n<td>0.69</td>\n<td>0.8</td>\n</tr>\n<tr>\n<td>maximum</td>\n<td>0.81</td>\n<td>0.84</td>\n</tr>\n<tr>\n<td>clustering</td>\n<td>0.76</td>\n<td>0.84</td>\n</tr>\n</tbody>\n</table>\n<h1>Notes</h1>\n<p>After reviewing the YOLO predictions using the maximum function and hdbscan, I found that YOLO actually produces many peak signals on the hidden test set (including almost all public notebooks and my best local model). This suggests that many high scores are the result of lucky predictions.</p>",
      "rawMarkdown": "Currently, I am developing ensemble functions for object localization. Here, I’d like to share some brief insights on the methods I’ve explored so far:\n\n# Average & Weighted Sum\n* Works well when your models have low diversity.\n* Performance may degrade if your models are highly diverse.\n* Weight selection is done through brute-force search.\n\n# Maximum Function\n* Ensures that object presence is not missed.\n* Crucially, it allows you to detect false positives from individual models — if the overall score drops after adding another model, it suggests that the new model is predicting false positives.\n\n# Clustering (HDBSCAN)\nHDBSCAN is an effective tool for handling clusters with varying shapes and densities. I initially tried it and am now conducting deeper experiments on its properties for object localization.\n* If a model predicts multiple keypoints in an area, there is a high possibility that an object is present.\n* Suppresses peak signals, reducing false positives. It also helps verify whether your current score improvements rely on peak signals, which can cause instability and shake-up.\n* Parameters like `threshold` from your model, `min_cluster_size`, and `min_samples` from HDBSCAN provide flexible control over prediction strength. Particularly effective for models trained with Gaussian balls or other irregularly shaped labels (e.g., keypoints annotated on motors or bacteria). I find it works especially well in 3D models, where the additional z-dimension improves clustering.\n* Helps address prediction diversity across multiple models.\n\n| methods | worse case LB | best case LB |\n| --- | --- |\n| average |0.69|0.8|\n| maximum |0.81|0.84|\n| clustering |0.76|0.84|\n\n# Notes\nAfter reviewing the YOLO predictions using the maximum function and hdbscan, I found that YOLO actually produces many peak signals on the hidden test set (including almost all public notebooks and my best local model). This suggests that many high scores are the result of lucky predictions.",
      "votes": 37
    },
    {
      "id": 3206242,
      "postDate": "2025-05-21T04:34:01.907Z",
      "content": "<p>Additional tips in HDBSCAN:</p>\n<h3>Preserving basic predictions</h3>\n<p>Assume we have two predictions from two models, with each model predicting a single keypoint in 3D volume. Setting <code>min_samples=1</code> and <code>min_cluster_size=2</code> allows HDBSCAN to preserve similar predictions made by both models without needing prediction multiple points at the target.</p>\n<h3>One neighbor as core point with five minimum cluster capacity:</h3>\n<table>\n<thead>\n<tr>\n<th>Setting</th>\n<th>Value</th>\n<th>Behavior</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>min_samples</code></td>\n<td>1</td>\n<td>Almost every point can be a core point → high recall, permissive clustering</td>\n</tr>\n<tr>\n<td><code>min_cluster_size</code></td>\n<td>5</td>\n<td>Clusters must have ≥5 points → filters out small/noisy clusters</td>\n</tr>\n</tbody>\n</table>\n<p>If we can ensemble a large number of models, we gain more flexibility to control the recall rate, making the clustering process more robust.</p>",
      "rawMarkdown": "Additional tips in HDBSCAN:\n\n### Preserving basic predictions\nAssume we have two predictions from two models, with each model predicting a single keypoint in 3D volume. Setting `min_samples=1` and `min_cluster_size=2` allows HDBSCAN to preserve similar predictions made by both models without needing prediction multiple points at the target.\n\n### One neighbor as core point with five minimum cluster capacity:\n| Setting | Value | Behavior |\n| --- | --- |\n| `min_samples`  | 1 | Almost every point can be a core point → high recall, permissive clustering |\n| `min_cluster_size`  | 5 | Clusters must have ≥5 points → filters out small/noisy clusters |\n\nIf we can ensemble a large number of models, we gain more flexibility to control the recall rate, making the clustering process more robust.",
      "votes": 1,
      "replies": [
        {
          "id": 3206403,
          "postDate": "2025-05-21T08:54:20.330Z",
          "content": "<blockquote>\n  <p>Assume we have two predictions from two models, with each model predicting a single keypoint in 3D volume. Setting min_samples=1 and min_cluster_size=2</p>\n</blockquote>\n<p>For both models you take the highest confidence prediction in 3d then ? This config means if both predictions are close it's motor else it's noise. If I understood correctly this will lead to miss true positives </p>\n<blockquote>\n  <p>If we can ensemble a large number of models, we gain more flexibility to control the recall rate, making the clustering process more robust.</p>\n</blockquote>\n<p>It doesn't mean we will be able to catch more motors. Depending on your models, the noise clusters might actually get bigger :/<br>\nI noticed that it's not easy to tune hdbscan/dbscan and I think that there is a high risk of overfitting to the public lb</p>",
          "rawMarkdown": ">Assume we have two predictions from two models, with each model predicting a single keypoint in 3D volume. Setting min_samples=1 and min_cluster_size=2\n\nFor both models you take the highest confidence prediction in 3d then ? This config means if both predictions are close it's motor else it's noise. If I understood correctly this will lead to miss true positives \n\n>If we can ensemble a large number of models, we gain more flexibility to control the recall rate, making the clustering process more robust.\n\nIt doesn't mean we will be able to catch more motors. Depending on your models, the noise clusters might actually get bigger :/\nI noticed that it's not easy to tune hdbscan/dbscan and I think that there is a high risk of overfitting to the public lb\n",
          "replies": [
            {
              "id": 3206409,
              "postDate": "2025-05-21T09:12:07.453Z",
              "content": "<p><a href=\"https://www.kaggle.com/victorvannobel\" target=\"_blank\">@victorvannobel</a> </p>\n<pre><code>For both models you take  highest confidence prediction  d  ? This config means  both predictions are  s noise. If I understood correctly this will lead  miss  positives\n</code></pre>\n<p>You can duplicate your prediction and plus 1e-4. This secure your prediction.</p>\n<pre><code>It doesn\nI noticed that it\n</code></pre>\n<p>Yeah the biggest problem is tunning the hyperparameter. We need to balance FP and FN through the parameter tunning. I prefer accepting more FP since the metric is fbeta score with beta=2.</p>\n<p>Furthermore, developing how to select the cluster is also a challenge, currently I was trying confidence-based selection, but it turns out that cluster-sized-based selection is better. This needs further development.</p>",
              "rawMarkdown": "@victorvannobel \n\n```\nFor both models you take the highest confidence prediction in 3d then ? This config means if both predictions are close it's motor else it's noise. If I understood correctly this will lead to miss true positives\n```\nYou can duplicate your prediction and plus 1e-4. This secure your prediction.\n\n```\nIt doesn't mean we will be able to catch more motors. Depending on your models, the noise clusters might actually get bigger :/\nI noticed that it's not easy to tune hdbscan/dbscan and I think that there is a high risk of overfitting to the public lb\n```\nYeah the biggest problem is tunning the hyperparameter. We need to balance FP and FN through the parameter tunning. I prefer accepting more FP since the metric is fbeta score with beta=2.\n\nFurthermore, developing how to select the cluster is also a challenge, currently I was trying confidence-based selection, but it turns out that cluster-sized-based selection is better. This needs further development."
            },
            {
              "id": 3207061,
              "postDate": "2025-05-22T07:02:49.313Z",
              "content": "<p>What is the image size for your model? I tried using DBSCAN but got only 0.765</p>",
              "rawMarkdown": "What is the image size for your model? I tried using DBSCAN but got only 0.765"
            }
          ]
        }
      ]
    },
    {
      "id": 3200048,
      "postDate": "2025-05-12T00:13:24.490Z",
      "content": "<p>Thanks for your insights! I'm a beginner, so I have some basic questions. <br>\nDo you do the ensembling before or after 3D NMS? <br>\nI was just ensembling by taking all the predictions of all the models and treating that as the effectively single model and then doing 3D NMS. But I have much less control than the methods you mentioned as a result…. Is that the same as the Maximum method?<br>\nAnd how do you predict worst case and best case LB? <br>\nAlso maximum seems is working just as well , if not better than clustering … I would guess that because clustering will also take the Z axis into account better, I'll expect it to have better performance but it isn't the case ? <br>\nAnd what is the Gaussian ball model ? I have been trying 3D models and keypoint detectors, but they seem to be much harder to train and to learn something meaningful compared to the templated yolo model. </p>",
      "rawMarkdown": "Thanks for your insights! I'm a beginner, so I have some basic questions. \nDo you do the ensembling before or after 3D NMS? \nI was just ensembling by taking all the predictions of all the models and treating that as the effectively single model and then doing 3D NMS. But I have much less control than the methods you mentioned as a result.... Is that the same as the Maximum method?\nAnd how do you predict worst case and best case LB? \nAlso maximum seems is working just as well , if not better than clustering ... I would guess that because clustering will also take the Z axis into account better, I'll expect it to have better performance but it isn't the case ? \nAnd what is the Gaussian ball model ? I have been trying 3D models and keypoint detectors, but they seem to be much harder to train and to learn something meaningful compared to the templated yolo model. ",
      "votes": 1,
      "replies": [
        {
          "id": 3200053,
          "postDate": "2025-05-12T00:46:48.887Z",
          "content": "<p><a href=\"https://www.kaggle.com/iamudit\" target=\"_blank\">@iamudit</a>  The 3D NMS implemented in the public notebook is effectively equivalent to simply taking the maximum. The filtering step is redundant because they directly select <code>final_detection = detection[0]</code>.</p>\n<p>If you're doing clustering, you need to ensure that there are enough detected points; otherwise, clustering won’t work. Clustering is actually better than taking the maximum when you have 10+ ensemble models. Selecting the cluster with the largest size is similar to majority voting.</p>\n<p>I understand it's quite challenging to complete predictions with 10+ models within 12 hours. That’s why I developed a two-stage method: the ensemble operates in the second stage, and the input at that stage is much smaller compared to the raw volume.</p>\n<p>Also, you can’t effectively train a 3D model by providing only a single point—there's too much semantic ambiguity. A simple improvement for the labels is to annotate all points within 1000 angstroms of the target point as positive. Using a Gaussian ball follows the same idea but instead of assigning hard labels, it gives a soft, continuous score—like applying a 3D Gaussian function.</p>",
          "rawMarkdown": "@iamudit  The 3D NMS implemented in the public notebook is effectively equivalent to simply taking the maximum. The filtering step is redundant because they directly select `final_detection = detection[0]`.\n\nIf you're doing clustering, you need to ensure that there are enough detected points; otherwise, clustering won’t work. Clustering is actually better than taking the maximum when you have 10+ ensemble models. Selecting the cluster with the largest size is similar to majority voting.\n\nI understand it's quite challenging to complete predictions with 10+ models within 12 hours. That’s why I developed a two-stage method: the ensemble operates in the second stage, and the input at that stage is much smaller compared to the raw volume.\n\nAlso, you can’t effectively train a 3D model by providing only a single point—there's too much semantic ambiguity. A simple improvement for the labels is to annotate all points within 1000 angstroms of the target point as positive. Using a Gaussian ball follows the same idea but instead of assigning hard labels, it gives a soft, continuous score—like applying a 3D Gaussian function.\n",
          "votes": 4,
          "replies": [
            {
              "id": 3201130,
              "postDate": "2025-05-13T12:55:55.673Z",
              "content": "<p><a href=\"https://www.kaggle.com/tom99763\" target=\"_blank\">@tom99763</a>  Thank you for your explanation~! <br>\nI see, 10+ ensemble models is impressive and at that point with the high number of detections,clustering would be much better than a simple maximum.<br>\nI'm doing a similar staging as you and still, currently fitting in more than 2 models within the time limit seems challenging,  I guess lowering the image resolution would increase the speed and therefore the raw number of models, (apart from discarding many slices till multiple models fit) l (but also compromise the prediction accuracies ) so there is a LOT of experimentation to be done to determine the best configuration !!! </p>\n<p>Thanks for explaining the 3D model and the Gaussian ball. Do you think it has the potential to beat the 2 stage model ? where the first stage has a (comparatively)simple task and 2nd stage has a simple task (both in lower dimensionality) . These 3D models suffer from the curse of dimensionality. <br>\nBut I completely agree that if trained well enough and made light enough, it would be a viable candidate in the ensemble pipeline.</p>",
              "rawMarkdown": "@tom99763  Thank you for your explanation~! \nI see, 10+ ensemble models is impressive and at that point with the high number of detections,clustering would be much better than a simple maximum.\nI'm doing a similar staging as you and still, currently fitting in more than 2 models within the time limit seems challenging,  I guess lowering the image resolution would increase the speed and therefore the raw number of models, (apart from discarding many slices till multiple models fit) l (but also compromise the prediction accuracies ) so there is a LOT of experimentation to be done to determine the best configuration !!! \n\nThanks for explaining the 3D model and the Gaussian ball. Do you think it has the potential to beat the 2 stage model ? where the first stage has a (comparatively)simple task and 2nd stage has a simple task (both in lower dimensionality) . These 3D models suffer from the curse of dimensionality. \nBut I completely agree that if trained well enough and made light enough, it would be a viable candidate in the ensemble pipeline.\n"
            },
            {
              "id": 3201582,
              "postDate": "2025-05-14T04:59:28.587Z",
              "content": "<p><a href=\"https://www.kaggle.com/iamudit\" target=\"_blank\">@iamudit</a> My solution is two-stage + 3d (I have published <a href=\"https://www.kaggle.com/code/tom99763/first-idea-to-extract-keypoints-byu\" target=\"_blank\">the initial idea of my solution</a>), but without using traditional 3D models like UNet or SegResNet. I found an effective way to replicate the functionality of those models while keeping the architecture extremely lightweight. Even with over 10 models in the pipeline, the total inference time remains around 7 hours which is the same as running a single model.</p>\n<p>I'm still finding more models to stack, thinking that 20+ models is possible.</p>",
              "rawMarkdown": "@iamudit My solution is two-stage + 3d (I have published [the initial idea of my solution](https://www.kaggle.com/code/tom99763/first-idea-to-extract-keypoints-byu)), but without using traditional 3D models like UNet or SegResNet. I found an effective way to replicate the functionality of those models while keeping the architecture extremely lightweight. Even with over 10 models in the pipeline, the total inference time remains around 7 hours which is the same as running a single model.\n\nI'm still finding more models to stack, thinking that 20+ models is possible.",
              "votes": 5
            },
            {
              "id": 3202231,
              "postDate": "2025-05-15T04:31:52.157Z",
              "content": "<p>Thanks!! It looks like a very interesting solution. Looking forward to your solution writeup! :)</p>",
              "rawMarkdown": "Thanks!! It looks like a very interesting solution. Looking forward to your solution writeup! :)",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3201017,
      "postDate": "2025-05-13T10:55:45.087Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3206242,
      "author_name": "Tom",
      "author_url": "",
      "post_date": "2025-05-21T04:34:01.907000",
      "content": "<p>Additional tips in HDBSCAN:</p>\n<h3>Preserving basic predictions</h3>\n<p>Assume we have two predictions from two models, with each model predicting a single keypoint in 3D volume. Setting <code>min_samples=1</code> and <code>min_cluster_size=2</code> allows HDBSCAN to preserve similar predictions made by both models without needing prediction multiple points at the target.</p>\n<h3>One neighbor as core point with five minimum cluster capacity:</h3>\n<table>\n<thead>\n<tr>\n<th>Setting</th>\n<th>Value</th>\n<th>Behavior</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>min_samples</code></td>\n<td>1</td>\n<td>Almost every point can be a core point → high recall, permissive clustering</td>\n</tr>\n<tr>\n<td><code>min_cluster_size</code></td>\n<td>5</td>\n<td>Clusters must have ≥5 points → filters out small/noisy clusters</td>\n</tr>\n</tbody>\n</table>\n<p>If we can ensemble a large number of models, we gain more flexibility to control the recall rate, making the clustering process more robust.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3206403,
          "author_name": "Victor",
          "author_url": "",
          "post_date": "2025-05-21T08:54:20.330000",
          "content": "<blockquote>\n  <p>Assume we have two predictions from two models, with each model predicting a single keypoint in 3D volume. Setting min_samples=1 and min_cluster_size=2</p>\n</blockquote>\n<p>For both models you take the highest confidence prediction in 3d then ? This config means if both predictions are close it's motor else it's noise. If I understood correctly this will lead to miss true positives </p>\n<blockquote>\n  <p>If we can ensemble a large number of models, we gain more flexibility to control the recall rate, making the clustering process more robust.</p>\n</blockquote>\n<p>It doesn't mean we will be able to catch more motors. Depending on your models, the noise clusters might actually get bigger :/<br>\nI noticed that it's not easy to tune hdbscan/dbscan and I think that there is a high risk of overfitting to the public lb</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3206409,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2025-05-21T09:12:07.453000",
              "content": "<p><a href=\"https://www.kaggle.com/victorvannobel\" target=\"_blank\">@victorvannobel</a> </p>\n<pre><code>For both models you take  highest confidence prediction  d  ? This config means  both predictions are  s noise. If I understood correctly this will lead  miss  positives\n</code></pre>\n<p>You can duplicate your prediction and plus 1e-4. This secure your prediction.</p>\n<pre><code>It doesn\nI noticed that it\n</code></pre>\n<p>Yeah the biggest problem is tunning the hyperparameter. We need to balance FP and FN through the parameter tunning. I prefer accepting more FP since the metric is fbeta score with beta=2.</p>\n<p>Furthermore, developing how to select the cluster is also a challenge, currently I was trying confidence-based selection, but it turns out that cluster-sized-based selection is better. This needs further development.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3207061,
              "author_name": "Rustam Bazarbayev",
              "author_url": "",
              "post_date": "2025-05-22T07:02:49.313000",
              "content": "<p>What is the image size for your model? I tried using DBSCAN but got only 0.765</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3200048,
      "author_name": "Udit Jain",
      "author_url": "",
      "post_date": "2025-05-12T00:13:24.490000",
      "content": "<p>Thanks for your insights! I'm a beginner, so I have some basic questions. <br>\nDo you do the ensembling before or after 3D NMS? <br>\nI was just ensembling by taking all the predictions of all the models and treating that as the effectively single model and then doing 3D NMS. But I have much less control than the methods you mentioned as a result…. Is that the same as the Maximum method?<br>\nAnd how do you predict worst case and best case LB? <br>\nAlso maximum seems is working just as well , if not better than clustering … I would guess that because clustering will also take the Z axis into account better, I'll expect it to have better performance but it isn't the case ? <br>\nAnd what is the Gaussian ball model ? I have been trying 3D models and keypoint detectors, but they seem to be much harder to train and to learn something meaningful compared to the templated yolo model. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3200053,
          "author_name": "Tom",
          "author_url": "",
          "post_date": "2025-05-12T00:46:48.887000",
          "content": "<p><a href=\"https://www.kaggle.com/iamudit\" target=\"_blank\">@iamudit</a>  The 3D NMS implemented in the public notebook is effectively equivalent to simply taking the maximum. The filtering step is redundant because they directly select <code>final_detection = detection[0]</code>.</p>\n<p>If you're doing clustering, you need to ensure that there are enough detected points; otherwise, clustering won’t work. Clustering is actually better than taking the maximum when you have 10+ ensemble models. Selecting the cluster with the largest size is similar to majority voting.</p>\n<p>I understand it's quite challenging to complete predictions with 10+ models within 12 hours. That’s why I developed a two-stage method: the ensemble operates in the second stage, and the input at that stage is much smaller compared to the raw volume.</p>\n<p>Also, you can’t effectively train a 3D model by providing only a single point—there's too much semantic ambiguity. A simple improvement for the labels is to annotate all points within 1000 angstroms of the target point as positive. Using a Gaussian ball follows the same idea but instead of assigning hard labels, it gives a soft, continuous score—like applying a 3D Gaussian function.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 3201130,
              "author_name": "Udit Jain",
              "author_url": "",
              "post_date": "2025-05-13T12:55:55.673000",
              "content": "<p><a href=\"https://www.kaggle.com/tom99763\" target=\"_blank\">@tom99763</a>  Thank you for your explanation~! <br>\nI see, 10+ ensemble models is impressive and at that point with the high number of detections,clustering would be much better than a simple maximum.<br>\nI'm doing a similar staging as you and still, currently fitting in more than 2 models within the time limit seems challenging,  I guess lowering the image resolution would increase the speed and therefore the raw number of models, (apart from discarding many slices till multiple models fit) l (but also compromise the prediction accuracies ) so there is a LOT of experimentation to be done to determine the best configuration !!! </p>\n<p>Thanks for explaining the 3D model and the Gaussian ball. Do you think it has the potential to beat the 2 stage model ? where the first stage has a (comparatively)simple task and 2nd stage has a simple task (both in lower dimensionality) . These 3D models suffer from the curse of dimensionality. <br>\nBut I completely agree that if trained well enough and made light enough, it would be a viable candidate in the ensemble pipeline.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3201582,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2025-05-14T04:59:28.587000",
              "content": "<p><a href=\"https://www.kaggle.com/iamudit\" target=\"_blank\">@iamudit</a> My solution is two-stage + 3d (I have published <a href=\"https://www.kaggle.com/code/tom99763/first-idea-to-extract-keypoints-byu\" target=\"_blank\">the initial idea of my solution</a>), but without using traditional 3D models like UNet or SegResNet. I found an effective way to replicate the functionality of those models while keeping the architecture extremely lightweight. Even with over 10 models in the pipeline, the total inference time remains around 7 hours which is the same as running a single model.</p>\n<p>I'm still finding more models to stack, thinking that 20+ models is possible.</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 3202231,
              "author_name": "Udit Jain",
              "author_url": "",
              "post_date": "2025-05-15T04:31:52.157000",
              "content": "<p>Thanks!! It looks like a very interesting solution. Looking forward to your solution writeup! :)</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3201017,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-13T10:55:45.087000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3198136": "Currently, I am developing ensemble functions for object localization. Here, I’d like to share some brief insights on the methods I’ve explored so far:\n\n# Average & Weighted Sum\n* Works well when your models have low diversity.\n* Performance may degrade if your models are highly diverse.\n* Weight selection is done through brute-force search.\n\n# Maximum Function\n* Ensures that object presence is not missed.\n* Crucially, it allows you to detect false positives from individual models — if the overall score drops after adding another model, it suggests that the new model is predicting false positives.\n\n# Clustering (HDBSCAN)\nHDBSCAN is an effective tool for handling clusters with varying shapes and densities. I initially tried it and am now conducting deeper experiments on its properties for object localization.\n* If a model predicts multiple keypoints in an area, there is a high possibility that an object is present.\n* Suppresses peak signals, reducing false positives. It also helps verify whether your current score improvements rely on peak signals, which can cause instability and shake-up.\n* Parameters like `threshold` from your model, `min_cluster_size`, and `min_samples` from HDBSCAN provide flexible control over prediction strength. Particularly effective for models trained with Gaussian balls or other irregularly shaped labels (e.g., keypoints annotated on motors or bacteria). I find it works especially well in 3D models, where the additional z-dimension improves clustering.\n* Helps address prediction diversity across multiple models.\n\n| methods | worse case LB | best case LB |\n| --- | --- |\n| average |0.69|0.8|\n| maximum |0.81|0.84|\n| clustering |0.76|0.84|\n\n# Notes\nAfter reviewing the YOLO predictions using the maximum function and hdbscan, I found that YOLO actually produces many peak signals on the hidden test set (including almost all public notebooks and my best local model). This suggests that many high scores are the result of lucky predictions.",
    "3206242": "Additional tips in HDBSCAN:\n\n### Preserving basic predictions\nAssume we have two predictions from two models, with each model predicting a single keypoint in 3D volume. Setting `min_samples=1` and `min_cluster_size=2` allows HDBSCAN to preserve similar predictions made by both models without needing prediction multiple points at the target.\n\n### One neighbor as core point with five minimum cluster capacity:\n| Setting | Value | Behavior |\n| --- | --- |\n| `min_samples`  | 1 | Almost every point can be a core point → high recall, permissive clustering |\n| `min_cluster_size`  | 5 | Clusters must have ≥5 points → filters out small/noisy clusters |\n\nIf we can ensemble a large number of models, we gain more flexibility to control the recall rate, making the clustering process more robust.",
    "3200048": "Thanks for your insights! I'm a beginner, so I have some basic questions. \nDo you do the ensembling before or after 3D NMS? \nI was just ensembling by taking all the predictions of all the models and treating that as the effectively single model and then doing 3D NMS. But I have much less control than the methods you mentioned as a result.... Is that the same as the Maximum method?\nAnd how do you predict worst case and best case LB? \nAlso maximum seems is working just as well , if not better than clustering ... I would guess that because clustering will also take the Z axis into account better, I'll expect it to have better performance but it isn't the case ? \nAnd what is the Gaussian ball model ? I have been trying 3D models and keypoint detectors, but they seem to be much harder to train and to learn something meaningful compared to the templated yolo model. ",
    "3201017": ""
  }
}