{
  "id": 583164,
  "title": "13th Place Solution - 2.5D YOLO Ensemble with DBSCAN",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/victor-13th-place-solution-2-5d-yolo-ensemble-with",
  "author_name": "",
  "post_date": "2025-06-05T06:13:56.437382300Z",
  "votes": 18,
  "comment_count": 2,
  "views": 0,
  "content": "<p><strong>First of all, I would like to thank <a href=\"https://www.kaggle.com/andrewjdarley\" target=\"_blank\">@andrewjdarley</a> and <a href=\"https://www.kaggle.com/fautei\" target=\"_blank\">@fautei</a> for their great notebooks which were very useful, and <a href=\"https://www.kaggle.com/yyyy0201\" target=\"_blank\">@yyyy0201</a> for sharing valuable insights and ideas.</strong></p>\n<p>I’m happy to finish in the top 50 for the first competition in which I invested time.</p>\n<p>My solution involves ensembling 6 2.5D YOLO models with DBSCAN. I’ll describe it in four parts: labeling, preprocessing, training, and postprocessing.</p>\n<p>It was quite hard to train multiple models and regularly compute CV scores of my pipeline as I only had access to Kaggle's computing resources.</p>\n<h2>Labeling</h2>\n<p>I used tomograms containing one or more motors. I also corrected some mislabeled data by running my inference pipeline on the training set, followed by visual inspection. I didn't use random slices from tomograms without motors as they would mostly be useless noise. However, using the hard negative slices that my models struggled with could have improved the results.</p>\n<p>The external data from <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> was resized to a low resolution and I didn’t have the technical resources to load the original tomograms and apply the desired preprocessing. Therefore, no external data was used, although it could have improved results since YOLO performs better with a large number of training images.</p>\n<p>I used 3-4 slices below and above the slice containing the center of the motor along the z-axis. The bounding boxes sizes were 24x24 and 30x30.</p>\n<p>I randomly split the data into 80% for training and 20% for validation. After reviewing the slices in both sets, I noticed a good tomogram distribution, thanks to a lucky seed. It helped to get strong solo models with good generalization. I also applied few augmentations to the validation set including gaussian, median, average blurs, CLAHE and RandomBrightnessContrast.</p>\n<h2>Preprocessing</h2>\n<p>The preprocessing only consisted of 2nd and 98th percentile normalization. During inference, the slices were resized to 1024×1024 using letterbox.</p>\n<p>I decided to join this competition to focus exclusively on 2.5D models. During inference, the inputs given to the YOLO models were RGB slices with slice z-2 in the R channel and slice z+2 in the B channel. During training, slice z-1 and z+1 were used in R and B channels. The LB scores were better using 2 slices below and above.</p>\n<h2>Training</h2>\n<p>My final pipeline ensembles 6 YOLO models: 8s, 9s, 10m, 2x 11s, 11m.<br>\nThe training args:</p>\n<pre><code>- epochs: \n- batch:  - \n- imgsz: \n- dropout: \n- lr0:  - \n- lrf: \n- weight_decay:  - \n- scale: \n- mixup:  - \n- copy_paste: \n- mosaic: \n</code></pre>\n<p>I also modified the default YOLO augmentations. They are the same as those I applied to the validation set.</p>\n<p>I noticed that high-resolution tomograms didn’t contain any motors but as I said in <strong>Preprocessing</strong>, I didn't use hard negative. My models were trained on tomograms with resolutions ranging approximately from 920 to 1000.</p>\n<p>My best solo models (8s, 11s) scored both 0.83+ (it could still be improved).</p>\n<h2>Postprocessing</h2>\n<p>For each model, the confidence threshold was set to 0.35 and the top 10 detections were kept.</p>\n<p>I used TTA (h-flip, v-flip) during inference but I didn't merge the TTA detections using NMS or WBF. I decided to ensemble the detections using DBSCAN. I normalized the 3D coordinates of the detections. Selecting the <code>eps</code> parameter was kinda easy, I set it to 0.02. Selecting the <code>min_samples</code> parameter was more challenging. Based on my results on the public LB and in order to avoid missing true positives on the private LB, I set it to 22. </p>\n<p>I noticed that clusters elongated along the x or y axis were more likely to be false positives, whereas valid detections were typically elongated along the z-axis but I didn’t follow up on that idea.</p>\n<p>During inference 3 models were running on even slices and 3 models on odd slices.</p>\n<p>It was an interesting challenge thanks to domain shift and I learned a lot.</p>\n<p>Thanks for reading,<br>\nVictor</p>",
  "messages": [
    {
      "id": "3217556",
      "postDate": "06/05/2025 06:13:56",
      "content": "<p><strong>First of all, I would like to thank <a href=\"https://www.kaggle.com/andrewjdarley\" target=\"_blank\">@andrewjdarley</a> and <a href=\"https://www.kaggle.com/fautei\" target=\"_blank\">@fautei</a> for their great notebooks which were very useful, and <a href=\"https://www.kaggle.com/yyyy0201\" target=\"_blank\">@yyyy0201</a> for sharing valuable insights and ideas.</strong></p>\n<p>I’m happy to finish in the top 50 for the first competition in which I invested time.</p>\n<p>My solution involves ensembling 6 2.5D YOLO models with DBSCAN. I’ll describe it in four parts: labeling, preprocessing, training, and postprocessing.</p>\n<p>It was quite hard to train multiple models and regularly compute CV scores of my pipeline as I only had access to Kaggle's computing resources.</p>\n<h2>Labeling</h2>\n<p>I used tomograms containing one or more motors. I also corrected some mislabeled data by running my inference pipeline on the training set, followed by visual inspection. I didn't use random slices from tomograms without motors as they would mostly be useless noise. However, using the hard negative slices that my models struggled with could have improved the results.</p>\n<p>The external data from <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> was resized to a low resolution and I didn’t have the technical resources to load the original tomograms and apply the desired preprocessing. Therefore, no external data was used, although it could have improved results since YOLO performs better with a large number of training images.</p>\n<p>I used 3-4 slices below and above the slice containing the center of the motor along the z-axis. The bounding boxes sizes were 24x24 and 30x30.</p>\n<p>I randomly split the data into 80% for training and 20% for validation. After reviewing the slices in both sets, I noticed a good tomogram distribution, thanks to a lucky seed. It helped to get strong solo models with good generalization. I also applied few augmentations to the validation set including gaussian, median, average blurs, CLAHE and RandomBrightnessContrast.</p>\n<h2>Preprocessing</h2>\n<p>The preprocessing only consisted of 2nd and 98th percentile normalization. During inference, the slices were resized to 1024×1024 using letterbox.</p>\n<p>I decided to join this competition to focus exclusively on 2.5D models. During inference, the inputs given to the YOLO models were RGB slices with slice z-2 in the R channel and slice z+2 in the B channel. During training, slice z-1 and z+1 were used in R and B channels. The LB scores were better using 2 slices below and above.</p>\n<h2>Training</h2>\n<p>My final pipeline ensembles 6 YOLO models: 8s, 9s, 10m, 2x 11s, 11m.<br>\nThe training args:</p>\n<pre><code>- epochs: \n- batch:  - \n- imgsz: \n- dropout: \n- lr0:  - \n- lrf: \n- weight_decay:  - \n- scale: \n- mixup:  - \n- copy_paste: \n- mosaic: \n</code></pre>\n<p>I also modified the default YOLO augmentations. They are the same as those I applied to the validation set.</p>\n<p>I noticed that high-resolution tomograms didn’t contain any motors but as I said in <strong>Preprocessing</strong>, I didn't use hard negative. My models were trained on tomograms with resolutions ranging approximately from 920 to 1000.</p>\n<p>My best solo models (8s, 11s) scored both 0.83+ (it could still be improved).</p>\n<h2>Postprocessing</h2>\n<p>For each model, the confidence threshold was set to 0.35 and the top 10 detections were kept.</p>\n<p>I used TTA (h-flip, v-flip) during inference but I didn't merge the TTA detections using NMS or WBF. I decided to ensemble the detections using DBSCAN. I normalized the 3D coordinates of the detections. Selecting the <code>eps</code> parameter was kinda easy, I set it to 0.02. Selecting the <code>min_samples</code> parameter was more challenging. Based on my results on the public LB and in order to avoid missing true positives on the private LB, I set it to 22. </p>\n<p>I noticed that clusters elongated along the x or y axis were more likely to be false positives, whereas valid detections were typically elongated along the z-axis but I didn’t follow up on that idea.</p>\n<p>During inference 3 models were running on even slices and 3 models on odd slices.</p>\n<p>It was an interesting challenge thanks to domain shift and I learned a lot.</p>\n<p>Thanks for reading,<br>\nVictor</p>",
      "rawMarkdown": "**First of all, I would like to thank @andrewjdarley and @fautei for their great notebooks which were very useful, and @yyyy0201 for sharing valuable insights and ideas.**\n\nI’m happy to finish in the top 50 for the first competition in which I invested time.\n\nMy solution involves ensembling 6 2.5D YOLO models with DBSCAN. I’ll describe it in four parts: labeling, preprocessing, training, and postprocessing.\n\nIt was quite hard to train multiple models and regularly compute CV scores of my pipeline as I only had access to Kaggle's computing resources.\n\n## Labeling\n\nI used tomograms containing one or more motors. I also corrected some mislabeled data by running my inference pipeline on the training set, followed by visual inspection. I didn't use random slices from tomograms without motors as they would mostly be useless noise. However, using the hard negative slices that my models struggled with could have improved the results.\n\nThe external data from @brendanartley was resized to a low resolution and I didn’t have the technical resources to load the original tomograms and apply the desired preprocessing. Therefore, no external data was used, although it could have improved results since YOLO performs better with a large number of training images.\n\nI used 3-4 slices below and above the slice containing the center of the motor along the z-axis. The bounding boxes sizes were 24x24 and 30x30.\n\nI randomly split the data into 80% for training and 20% for validation. After reviewing the slices in both sets, I noticed a good tomogram distribution, thanks to a lucky seed. It helped to get strong solo models with good generalization. I also applied few augmentations to the validation set including gaussian, median, average blurs, CLAHE and RandomBrightnessContrast.\n\n## Preprocessing\n\nThe preprocessing only consisted of 2nd and 98th percentile normalization. During inference, the slices were resized to 1024×1024 using letterbox.\n\nI decided to join this competition to focus exclusively on 2.5D models. During inference, the inputs given to the YOLO models were RGB slices with slice z-2 in the R channel and slice z+2 in the B channel. During training, slice z-1 and z+1 were used in R and B channels. The LB scores were better using 2 slices below and above.\n\n## Training\n\nMy final pipeline ensembles 6 YOLO models: 8s, 9s, 10m, 2x 11s, 11m.\nThe training args:\n```python\n- epochs: 50\n- batch: 8 - 16\n- imgsz: 960\n- dropout: 0.1\n- lr0: 0.0001 - 0.0005\n- lrf: 0.1\n- weight_decay: 0.0005 - 0.001\n- scale: 0.4\n- mixup: 0.1 - 0.2\n- copy_paste: 0.0\n- mosaic: 1\n```\n\nI also modified the default YOLO augmentations. They are the same as those I applied to the validation set.\n\nI noticed that high-resolution tomograms didn’t contain any motors but as I said in **Preprocessing**, I didn't use hard negative. My models were trained on tomograms with resolutions ranging approximately from 920 to 1000.\n\nMy best solo models (8s, 11s) scored both 0.83+ (it could still be improved).\n\n## Postprocessing\n \nFor each model, the confidence threshold was set to 0.35 and the top 10 detections were kept.\n\nI used TTA (h-flip, v-flip) during inference but I didn't merge the TTA detections using NMS or WBF. I decided to ensemble the detections using DBSCAN. I normalized the 3D coordinates of the detections. Selecting the `eps` parameter was kinda easy, I set it to 0.02. Selecting the `min_samples` parameter was more challenging. Based on my results on the public LB and in order to avoid missing true positives on the private LB, I set it to 22. \n\nI noticed that clusters elongated along the x or y axis were more likely to be false positives, whereas valid detections were typically elongated along the z-axis but I didn’t follow up on that idea.\n\nDuring inference 3 models were running on even slices and 3 models on odd slices.\n\nIt was an interesting challenge thanks to domain shift and I learned a lot.\n\nThanks for reading,\nVictor",
      "votes": null
    },
    {
      "id": "3218183",
      "postDate": "06/06/2025 00:37:58",
      "content": "<blockquote>\n  <p>I decided to join this competition to focus exclusively on 2.5D models. During inference, the inputs given to the YOLO models were RGB slices with slice z-2 in the R channel and slice z+2 in the B channel. During training, slice z-1 and z+1 were used in R and B channels</p>\n</blockquote>\n<p>I had not thought of 2.5D YOLO. It's good!<br>\nI also tried DBScan, but DFS(depth-first search scored better.</p>",
      "rawMarkdown": ">I decided to join this competition to focus exclusively on 2.5D models. During inference, the inputs given to the YOLO models were RGB slices with slice z-2 in the R channel and slice z+2 in the B channel. During training, slice z-1 and z+1 were used in R and B channels\n\nI had not thought of 2.5D YOLO. It's good!\nI also tried DBScan, but DFS(depth-first search scored better.",
      "votes": null
    },
    {
      "id": "3218509",
      "postDate": "06/06/2025 09:30:10",
      "content": "<p>It's a good way I think to do 2.5D without changing the arch of the models. You can still use pretrained weights but you are limited to only 3 channels. Maybe using +2 -2 on test data was better because of high number of slices per tomo, I don't know.</p>\n<p>For dbscan, setting eps to 0.01 (only dividing by 2) did not change my public LB but I got a very few FN on CV and I didn't want to risk it on private LB. Well, it could have led me to top 10 but it was almost impossible to predict 😂 (not choosing the best subs also happened to many teams). Actually I didn't know that the test data was balanced as one mentionned it, maybe I should give data hacking a try next time. Even if it's f2 metric, it's important to reduce the FP to get a top pipeline imo.</p>\n<p>It's not a hard pipeline thus it didn't overfit so much to public LB.</p>\n<p>There is room for improvement in the ensembling strategy, in how thresholds are set (I think about the quantile thresholding from bartley, I would like to try that strategy in another comp one day)…</p>",
      "rawMarkdown": "It's a good way I think to do 2.5D without changing the arch of the models. You can still use pretrained weights but you are limited to only 3 channels. Maybe using +2 -2 on test data was better because of high number of slices per tomo, I don't know.\n\nFor dbscan, setting eps to 0.01 (only dividing by 2) did not change my public LB but I got a very few FN on CV and I didn't want to risk it on private LB. Well, it could have led me to top 10 but it was almost impossible to predict 😂 (not choosing the best subs also happened to many teams). Actually I didn't know that the test data was balanced as one mentionned it, maybe I should give data hacking a try next time. Even if it's f2 metric, it's important to reduce the FP to get a top pipeline imo.\n\nIt's not a hard pipeline thus it didn't overfit so much to public LB.\n\nThere is room for improvement in the ensembling strategy, in how thresholds are set (I think about the quantile thresholding from bartley, I would like to try that strategy in another comp one day)...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3218183,
      "author_name": "minfuka",
      "author_url": "",
      "post_date": "06/06/2025 00:37:58",
      "content": "<blockquote>\n  <p>I decided to join this competition to focus exclusively on 2.5D models. During inference, the inputs given to the YOLO models were RGB slices with slice z-2 in the R channel and slice z+2 in the B channel. During training, slice z-1 and z+1 were used in R and B channels</p>\n</blockquote>\n<p>I had not thought of 2.5D YOLO. It's good!<br>\nI also tried DBScan, but DFS(depth-first search scored better.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3218509,
          "author_name": "victorvannobel",
          "author_url": "",
          "post_date": "06/06/2025 09:30:10",
          "content": "<p>It's a good way I think to do 2.5D without changing the arch of the models. You can still use pretrained weights but you are limited to only 3 channels. Maybe using +2 -2 on test data was better because of high number of slices per tomo, I don't know.</p>\n<p>For dbscan, setting eps to 0.01 (only dividing by 2) did not change my public LB but I got a very few FN on CV and I didn't want to risk it on private LB. Well, it could have led me to top 10 but it was almost impossible to predict 😂 (not choosing the best subs also happened to many teams). Actually I didn't know that the test data was balanced as one mentionned it, maybe I should give data hacking a try next time. Even if it's f2 metric, it's important to reduce the FP to get a top pipeline imo.</p>\n<p>It's not a hard pipeline thus it didn't overfit so much to public LB.</p>\n<p>There is room for improvement in the ensembling strategy, in how thresholds are set (I think about the quantile thresholding from bartley, I would like to try that strategy in another comp one day)…</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3217556": "**First of all, I would like to thank @andrewjdarley and @fautei for their great notebooks which were very useful, and @yyyy0201 for sharing valuable insights and ideas.**\n\nI’m happy to finish in the top 50 for the first competition in which I invested time.\n\nMy solution involves ensembling 6 2.5D YOLO models with DBSCAN. I’ll describe it in four parts: labeling, preprocessing, training, and postprocessing.\n\nIt was quite hard to train multiple models and regularly compute CV scores of my pipeline as I only had access to Kaggle's computing resources.\n\n## Labeling\n\nI used tomograms containing one or more motors. I also corrected some mislabeled data by running my inference pipeline on the training set, followed by visual inspection. I didn't use random slices from tomograms without motors as they would mostly be useless noise. However, using the hard negative slices that my models struggled with could have improved the results.\n\nThe external data from @brendanartley was resized to a low resolution and I didn’t have the technical resources to load the original tomograms and apply the desired preprocessing. Therefore, no external data was used, although it could have improved results since YOLO performs better with a large number of training images.\n\nI used 3-4 slices below and above the slice containing the center of the motor along the z-axis. The bounding boxes sizes were 24x24 and 30x30.\n\nI randomly split the data into 80% for training and 20% for validation. After reviewing the slices in both sets, I noticed a good tomogram distribution, thanks to a lucky seed. It helped to get strong solo models with good generalization. I also applied few augmentations to the validation set including gaussian, median, average blurs, CLAHE and RandomBrightnessContrast.\n\n## Preprocessing\n\nThe preprocessing only consisted of 2nd and 98th percentile normalization. During inference, the slices were resized to 1024×1024 using letterbox.\n\nI decided to join this competition to focus exclusively on 2.5D models. During inference, the inputs given to the YOLO models were RGB slices with slice z-2 in the R channel and slice z+2 in the B channel. During training, slice z-1 and z+1 were used in R and B channels. The LB scores were better using 2 slices below and above.\n\n## Training\n\nMy final pipeline ensembles 6 YOLO models: 8s, 9s, 10m, 2x 11s, 11m.\nThe training args:\n```python\n- epochs: 50\n- batch: 8 - 16\n- imgsz: 960\n- dropout: 0.1\n- lr0: 0.0001 - 0.0005\n- lrf: 0.1\n- weight_decay: 0.0005 - 0.001\n- scale: 0.4\n- mixup: 0.1 - 0.2\n- copy_paste: 0.0\n- mosaic: 1\n```\n\nI also modified the default YOLO augmentations. They are the same as those I applied to the validation set.\n\nI noticed that high-resolution tomograms didn’t contain any motors but as I said in **Preprocessing**, I didn't use hard negative. My models were trained on tomograms with resolutions ranging approximately from 920 to 1000.\n\nMy best solo models (8s, 11s) scored both 0.83+ (it could still be improved).\n\n## Postprocessing\n \nFor each model, the confidence threshold was set to 0.35 and the top 10 detections were kept.\n\nI used TTA (h-flip, v-flip) during inference but I didn't merge the TTA detections using NMS or WBF. I decided to ensemble the detections using DBSCAN. I normalized the 3D coordinates of the detections. Selecting the `eps` parameter was kinda easy, I set it to 0.02. Selecting the `min_samples` parameter was more challenging. Based on my results on the public LB and in order to avoid missing true positives on the private LB, I set it to 22. \n\nI noticed that clusters elongated along the x or y axis were more likely to be false positives, whereas valid detections were typically elongated along the z-axis but I didn’t follow up on that idea.\n\nDuring inference 3 models were running on even slices and 3 models on odd slices.\n\nIt was an interesting challenge thanks to domain shift and I learned a lot.\n\nThanks for reading,\nVictor",
    "3218183": ">I decided to join this competition to focus exclusively on 2.5D models. During inference, the inputs given to the YOLO models were RGB slices with slice z-2 in the R channel and slice z+2 in the B channel. During training, slice z-1 and z+1 were used in R and B channels\n\nI had not thought of 2.5D YOLO. It's good!\nI also tried DBScan, but DFS(depth-first search scored better.",
    "3218509": "It's a good way I think to do 2.5D without changing the arch of the models. You can still use pretrained weights but you are limited to only 3 channels. Maybe using +2 -2 on test data was better because of high number of slices per tomo, I don't know.\n\nFor dbscan, setting eps to 0.01 (only dividing by 2) did not change my public LB but I got a very few FN on CV and I didn't want to risk it on private LB. Well, it could have led me to top 10 but it was almost impossible to predict 😂 (not choosing the best subs also happened to many teams). Actually I didn't know that the test data was balanced as one mentionned it, maybe I should give data hacking a try next time. Even if it's f2 metric, it's important to reduce the FP to get a top pipeline imo.\n\nIt's not a hard pipeline thus it didn't overfit so much to public LB.\n\nThere is room for improvement in the ensembling strategy, in how thresholds are set (I think about the quantile thresholding from bartley, I would like to try that strategy in another comp one day)..."
  },
  "source": "meta"
}