{
  "id": 583144,
  "title": "17th Place Solution - Ultralytics + Timm",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/sergio-alvarez-paradox-17th-place-solution-ultraly",
  "author_name": "",
  "post_date": "2025-06-05T13:57:41.210Z",
  "votes": 30,
  "comment_count": 9,
  "views": 0,
  "content": "<h1>BYU Bacterial Flagellar Motors Competition - 17th Place Solution</h1>\n<p>First, we thank the competition hosts <a href=\"https://www.kaggle.com/andrewjdarleym\" target=\"_blank\">@andrewjdarleym</a>, <a href=\"https://www.kaggle.com/braxtonowens\" target=\"_blank\">@braxtonowens</a> and Kaggle staff for organizing this competition. Below, we introduce the solution of team <strong>Sergio Alvarez + Paradox</strong> -- <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a>, <a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a>.</p>\n<h2>Context</h2>\n<ul>\n<li><strong>Business context</strong>: <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/overview\" target=\"_blank\">Competition overview</a></li>\n<li><strong>Data context</strong>: <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/data\" target=\"_blank\">Competition data</a></li>\n</ul>\n<h2>Background &amp; Evolution</h2>\n<p>After achieving a good position in the CZII competition, our first approach (although not yet teamed up) was to try U-nets. However, this didn't work well, with the best model scoring only ~0.5 on the public leaderboard.</p>\n<p>Next, I experimented with <a href=\"https://www.kaggle.com/code/junkoda/speed-up-inference\" target=\"_blank\">Jun Koda's models</a>, which used Timm backbones with simple classification and segmentation heads. I managed to achieve 0.770 on the public LB but got stuck.</p>\n<p>After seeing discussions and public notebooks with good scores using YOLO, I decided to test it and achieved 0.792 on the public LB with YOLO10X. This got me thinking: if simple techniques with Timm backbones + a basic head got me 0.770, why not use the YOLO ultralytics framework that already has established augmentations, loss functions, SOTA neck for feature combination, and optimized bbox prediction heads? </p>\n<p>Paradox and I explored this approach, and here's our solution!</p>\n<h2>Solution Overview</h2>\n<p>Our <strong>17th place solution</strong> consisted of 3 YOLO-like (if I can call them that) models using <code>convnextv2_base.fcmae_ft_in22k_in1k</code> as the backbone. We extracted features from P4, P3, and P2 for the neck and made predictions with a single head at P3 (stride / 8).</p>\n<h3>Model Architectures</h3>\n<p>I started implementing Timm integration, but someone more intelligent than me had already done it! We just want to thank <a href=\"https://github.com/yjwong1999\" target=\"_blank\">yjwong1999</a>, who created the <a href=\"https://github.com/ultralytics/ultralytics/pull/19609\" target=\"_blank\">PR that we based our approach on</a>. We just added some stuff to it and created cfgs.</p>\n<p>We developed two main model configurations that vary only in the \"neck\" design:</p>\n<h4>Model 1:</h4>\n<ul>\n<li><strong>Backbone</strong>: <code>convnextv2_base.fcmae_ft_in22k_in1k</code></li>\n<li><strong>Neck</strong>: SPPF + C2PSA → upsample P4 and concatenate with P3 → head</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F222be1bb10a5511492902cc79f1a8379%2FCONFIG_1.png?generation=1749095926701628&amp;alt=media\" alt=\"\"></p>\n<h4>Model 2:</h4>\n<ul>\n<li><strong>Backbone</strong>: <code>convnextv2_base.fcmae_ft_in22k_in1k</code></li>\n<li><strong>Neck</strong>: SPPF + C2PSA → upsample P4, adaptive_avg_pool2d at P2, concatenate P3 and upsampled P4 → head</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F942884ec3348873a43b94f2b56bcf3a6%2FCONFIG_2.png?generation=1749095942195298&amp;alt=media\" alt=\"\"></p>\n<p>Paradox had the amazing idea of removing P5 scale and features after seeing <a href=\"https://www.kaggle.com/code/tatamikenn/reverse-engineering-yolo\" target=\"_blank\">@tatamikenn's excellent notebook</a> about YOLO features. This modification improved our scores by ~0.015 on the public LB.<br>\nThe adaptive_avg_pool2d/AVG block in the architecture is from <a href=\"https://github.com/yang-0201/MHAF-YOLO\" target=\"_blank\">@yyyy0201 remarkable mhaf-yolo</a>.</p>\n<h3>Training Configuration</h3>\n<p>We trained with 80% of the images that contained motors with trust = 4. And additionaly used 80% <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/569921\" target=\"_blank\">@bartley external dataset</a> dataset with trust = 0. The rest was used as validation.</p>\n<p>We used these standard training parameters:</p>\n<pre><code>results = model.train(\n    data=(yaml_path),\n    epochs=,\n    batch=,\n    imgsz=,\n    optimizer=,\n    lr0=,\n    lrf=,\n    warmup_epochs=,\n    dropout=,\n    project=(RUNS_DIR),\n    exist_ok=,\n    name=,\n    patience=,   \n    save_period=,\n    val=,   \n    mosaic=,\n    close_mosaic=,\n    mixup=,\n    flipud=,\n    scale=,\n    degrees=,\n    seed=,\n    deterministic=,\n    label_smoothing=,  \n    augment=,\n    device=,\n)\n</code></pre>\n<h3>Validation Strategy Challenges</h3>\n<p>Validation was problematic throughout the competition. We struggled to find correlation between our validation scores and public leaderboard performance, as our models were scoring 0.97+ in CV but showing different performance on the LB.</p>\n<p>We also faced significant challenges in epoch selection, which led to <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/576756\" target=\"_blank\">this discussion</a>.</p>\n<p>To mitigate the epoch selection problem, we used the model soup technique to average weights of multiple epochs.</p>\n<p>Ultimately, we compared models using mAP50, mAP50-95, precision, and recall curves to estimate performance. This didn't work so well.</p>\n<h2>Final Submission Details</h2>\n<h3>Ensemble Strategy</h3>\n<p>Our final ensemble consisted of three main components:</p>\n<ol>\n<li>First model configuration with standard training</li>\n<li>First model configuration with more augmentations:</li>\n</ol>\n<pre><code>   T = [\n       A.Blur(p=),\n       A.MedianBlur(p=),\n       A.ToGray(p=),\n       A.CLAHE(p=),\n       A.RandomBrightnessContrast(p=),\n       A.RandomGamma(p=),\n       A.ImageCompression(quality_upper=, quality_lower=, p=),\n       A.ShiftScaleRotate(p=),\n       A.GaussNoise(p=),\n       A.GaussianBlur(p=),  \n       A.UnsharpMask(p=)\n   ]\n</code></pre>\n<ol>\n<li>Second model configuration with standard training</li>\n</ol>\n<h3>Inference Strategy</h3>\n<p>We used <code>concentration = 0.5</code> for selecting slices, meaning we predicted on half of the available slices to increase speed.</p>\n<p>For aggregating predictions, we experimented with HDBSCAN clustering with the following parameters:</p>\n<ul>\n<li>min_samples: 1</li>\n<li>Cluster size: 4</li>\n<li>EPS: 50  </li>\n<li>YOLO threshold: 0.4</li>\n</ul>\n<p>We also added a confidence adjustment based on cluster size to favor bigger cluster.</p>\n<h2>What Didn't Work</h2>\n<ul>\n<li><strong>Alternative Backbones</strong>: We tested <code>tf_efficientnetv2_l.in21k_ft_in1k</code>, <code>resnext50_32x4d.fb_swsl_ig1b_ft_in1k</code>, and <code>caformer_b36.sail_in22k</code>. They achieved ~0.8 performance, none surpassed ConvNeXt base.</li>\n<li><strong>Larger Models</strong>: ConvNeXt base was already quite huge. ConvNeXt large actually performed worse, so we stopped exploring larger architectures.</li>\n<li><strong>Different Head</strong>: I've tried using ConvNeXt blocks in head, improved map50-95 but not in LB.</li>\n</ul>\n<h2>Bonus PB</h2>\n<p>This probably happened to many teams, but we didn't select our top-scoring notebook for the private leaderboard. In fact, we didn't even select the top 5.</p>\n<p><strong>Our True Best</strong>: Using the second model configuration alone (a different epoch soup combination than what we used in the ensemble) achieved <strong>0.856</strong> on the private LB, which would have been in the gold zone, indicating that P2 features were actually really important.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2Fc1c7ddc0295a36ea7048938f84d88802%2FPB.png?generation=1749095003450830&amp;alt=media\" alt=\"PB\"></p>\n<h2>Resources &amp; Acknowledgments</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/tatamikenn/reverse-engineering-yolo\" target=\"_blank\">@tatamikenn's YOLO features notebook</a></li>\n<li><a href=\"https://www.kaggle.com/code/junkoda/speed-up-inference\" target=\"_blank\">Jun Koda's speed optimization</a> </li>\n<li><a href=\"https://arxiv.org/abs/2203.05482\" target=\"_blank\">Model Soup paper</a></li>\n<li><a href=\"https://github.com/yang-0201/MHAF-YOLO\" target=\"_blank\">MHAF-YOLO</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/569921\" target=\"_blank\">Bartley External Dataset</a></li>\n<li><a href=\"https://github.com/ultralytics/ultralytics/pull/19609\" target=\"_blank\">Timm+Ultralytcs</a></li>\n</ul>",
  "messages": [
    {
      "id": "3217481",
      "postDate": "06/05/2025 04:05:18",
      "content": "<h1>BYU Bacterial Flagellar Motors Competition - 17th Place Solution</h1>\n<p>First, we thank the competition hosts <a href=\"https://www.kaggle.com/andrewjdarleym\" target=\"_blank\">@andrewjdarleym</a>, <a href=\"https://www.kaggle.com/braxtonowens\" target=\"_blank\">@braxtonowens</a> and Kaggle staff for organizing this competition. Below, we introduce the solution of team <strong>Sergio Alvarez + Paradox</strong> -- <a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a>, <a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a>.</p>\n<h2>Context</h2>\n<ul>\n<li><strong>Business context</strong>: <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/overview\" target=\"_blank\">Competition overview</a></li>\n<li><strong>Data context</strong>: <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/data\" target=\"_blank\">Competition data</a></li>\n</ul>\n<h2>Background &amp; Evolution</h2>\n<p>After achieving a good position in the CZII competition, our first approach (although not yet teamed up) was to try U-nets. However, this didn't work well, with the best model scoring only ~0.5 on the public leaderboard.</p>\n<p>Next, I experimented with <a href=\"https://www.kaggle.com/code/junkoda/speed-up-inference\" target=\"_blank\">Jun Koda's models</a>, which used Timm backbones with simple classification and segmentation heads. I managed to achieve 0.770 on the public LB but got stuck.</p>\n<p>After seeing discussions and public notebooks with good scores using YOLO, I decided to test it and achieved 0.792 on the public LB with YOLO10X. This got me thinking: if simple techniques with Timm backbones + a basic head got me 0.770, why not use the YOLO ultralytics framework that already has established augmentations, loss functions, SOTA neck for feature combination, and optimized bbox prediction heads? </p>\n<p>Paradox and I explored this approach, and here's our solution!</p>\n<h2>Solution Overview</h2>\n<p>Our <strong>17th place solution</strong> consisted of 3 YOLO-like (if I can call them that) models using <code>convnextv2_base.fcmae_ft_in22k_in1k</code> as the backbone. We extracted features from P4, P3, and P2 for the neck and made predictions with a single head at P3 (stride / 8).</p>\n<h3>Model Architectures</h3>\n<p>I started implementing Timm integration, but someone more intelligent than me had already done it! We just want to thank <a href=\"https://github.com/yjwong1999\" target=\"_blank\">yjwong1999</a>, who created the <a href=\"https://github.com/ultralytics/ultralytics/pull/19609\" target=\"_blank\">PR that we based our approach on</a>. We just added some stuff to it and created cfgs.</p>\n<p>We developed two main model configurations that vary only in the \"neck\" design:</p>\n<h4>Model 1:</h4>\n<ul>\n<li><strong>Backbone</strong>: <code>convnextv2_base.fcmae_ft_in22k_in1k</code></li>\n<li><strong>Neck</strong>: SPPF + C2PSA → upsample P4 and concatenate with P3 → head</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F222be1bb10a5511492902cc79f1a8379%2FCONFIG_1.png?generation=1749095926701628&amp;alt=media\" alt=\"\"></p>\n<h4>Model 2:</h4>\n<ul>\n<li><strong>Backbone</strong>: <code>convnextv2_base.fcmae_ft_in22k_in1k</code></li>\n<li><strong>Neck</strong>: SPPF + C2PSA → upsample P4, adaptive_avg_pool2d at P2, concatenate P3 and upsampled P4 → head</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F942884ec3348873a43b94f2b56bcf3a6%2FCONFIG_2.png?generation=1749095942195298&amp;alt=media\" alt=\"\"></p>\n<p>Paradox had the amazing idea of removing P5 scale and features after seeing <a href=\"https://www.kaggle.com/code/tatamikenn/reverse-engineering-yolo\" target=\"_blank\">@tatamikenn's excellent notebook</a> about YOLO features. This modification improved our scores by ~0.015 on the public LB.<br>\nThe adaptive_avg_pool2d/AVG block in the architecture is from <a href=\"https://github.com/yang-0201/MHAF-YOLO\" target=\"_blank\">@yyyy0201 remarkable mhaf-yolo</a>.</p>\n<h3>Training Configuration</h3>\n<p>We trained with 80% of the images that contained motors with trust = 4. And additionaly used 80% <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/569921\" target=\"_blank\">@bartley external dataset</a> dataset with trust = 0. The rest was used as validation.</p>\n<p>We used these standard training parameters:</p>\n<pre><code>results = model.train(\n    data=(yaml_path),\n    epochs=,\n    batch=,\n    imgsz=,\n    optimizer=,\n    lr0=,\n    lrf=,\n    warmup_epochs=,\n    dropout=,\n    project=(RUNS_DIR),\n    exist_ok=,\n    name=,\n    patience=,   \n    save_period=,\n    val=,   \n    mosaic=,\n    close_mosaic=,\n    mixup=,\n    flipud=,\n    scale=,\n    degrees=,\n    seed=,\n    deterministic=,\n    label_smoothing=,  \n    augment=,\n    device=,\n)\n</code></pre>\n<h3>Validation Strategy Challenges</h3>\n<p>Validation was problematic throughout the competition. We struggled to find correlation between our validation scores and public leaderboard performance, as our models were scoring 0.97+ in CV but showing different performance on the LB.</p>\n<p>We also faced significant challenges in epoch selection, which led to <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/576756\" target=\"_blank\">this discussion</a>.</p>\n<p>To mitigate the epoch selection problem, we used the model soup technique to average weights of multiple epochs.</p>\n<p>Ultimately, we compared models using mAP50, mAP50-95, precision, and recall curves to estimate performance. This didn't work so well.</p>\n<h2>Final Submission Details</h2>\n<h3>Ensemble Strategy</h3>\n<p>Our final ensemble consisted of three main components:</p>\n<ol>\n<li>First model configuration with standard training</li>\n<li>First model configuration with more augmentations:</li>\n</ol>\n<pre><code>   T = [\n       A.Blur(p=),\n       A.MedianBlur(p=),\n       A.ToGray(p=),\n       A.CLAHE(p=),\n       A.RandomBrightnessContrast(p=),\n       A.RandomGamma(p=),\n       A.ImageCompression(quality_upper=, quality_lower=, p=),\n       A.ShiftScaleRotate(p=),\n       A.GaussNoise(p=),\n       A.GaussianBlur(p=),  \n       A.UnsharpMask(p=)\n   ]\n</code></pre>\n<ol>\n<li>Second model configuration with standard training</li>\n</ol>\n<h3>Inference Strategy</h3>\n<p>We used <code>concentration = 0.5</code> for selecting slices, meaning we predicted on half of the available slices to increase speed.</p>\n<p>For aggregating predictions, we experimented with HDBSCAN clustering with the following parameters:</p>\n<ul>\n<li>min_samples: 1</li>\n<li>Cluster size: 4</li>\n<li>EPS: 50  </li>\n<li>YOLO threshold: 0.4</li>\n</ul>\n<p>We also added a confidence adjustment based on cluster size to favor bigger cluster.</p>\n<h2>What Didn't Work</h2>\n<ul>\n<li><strong>Alternative Backbones</strong>: We tested <code>tf_efficientnetv2_l.in21k_ft_in1k</code>, <code>resnext50_32x4d.fb_swsl_ig1b_ft_in1k</code>, and <code>caformer_b36.sail_in22k</code>. They achieved ~0.8 performance, none surpassed ConvNeXt base.</li>\n<li><strong>Larger Models</strong>: ConvNeXt base was already quite huge. ConvNeXt large actually performed worse, so we stopped exploring larger architectures.</li>\n<li><strong>Different Head</strong>: I've tried using ConvNeXt blocks in head, improved map50-95 but not in LB.</li>\n</ul>\n<h2>Bonus PB</h2>\n<p>This probably happened to many teams, but we didn't select our top-scoring notebook for the private leaderboard. In fact, we didn't even select the top 5.</p>\n<p><strong>Our True Best</strong>: Using the second model configuration alone (a different epoch soup combination than what we used in the ensemble) achieved <strong>0.856</strong> on the private LB, which would have been in the gold zone, indicating that P2 features were actually really important.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2Fc1c7ddc0295a36ea7048938f84d88802%2FPB.png?generation=1749095003450830&amp;alt=media\" alt=\"PB\"></p>\n<h2>Resources &amp; Acknowledgments</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/tatamikenn/reverse-engineering-yolo\" target=\"_blank\">@tatamikenn's YOLO features notebook</a></li>\n<li><a href=\"https://www.kaggle.com/code/junkoda/speed-up-inference\" target=\"_blank\">Jun Koda's speed optimization</a> </li>\n<li><a href=\"https://arxiv.org/abs/2203.05482\" target=\"_blank\">Model Soup paper</a></li>\n<li><a href=\"https://github.com/yang-0201/MHAF-YOLO\" target=\"_blank\">MHAF-YOLO</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/569921\" target=\"_blank\">Bartley External Dataset</a></li>\n<li><a href=\"https://github.com/ultralytics/ultralytics/pull/19609\" target=\"_blank\">Timm+Ultralytcs</a></li>\n</ul>",
      "rawMarkdown": "# BYU Bacterial Flagellar Motors Competition - 17th Place Solution\n\nFirst, we thank the competition hosts @andrewjdarleym, @braxtonowens and Kaggle staff for organizing this competition. Below, we introduce the solution of team **Sergio Alvarez + Paradox** -- @sersasj, @iamparadox.\n\n## Context\n\n- **Business context**: [Competition overview](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/overview)\n- **Data context**: [Competition data](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/data)\n\n## Background & Evolution\n\nAfter achieving a good position in the CZII competition, our first approach (although not yet teamed up) was to try U-nets. However, this didn't work well, with the best model scoring only ~0.5 on the public leaderboard.\n\nNext, I experimented with [Jun Koda's models](https://www.kaggle.com/code/junkoda/speed-up-inference), which used Timm backbones with simple classification and segmentation heads. I managed to achieve 0.770 on the public LB but got stuck.\n\nAfter seeing discussions and public notebooks with good scores using YOLO, I decided to test it and achieved 0.792 on the public LB with YOLO10X. This got me thinking: if simple techniques with Timm backbones + a basic head got me 0.770, why not use the YOLO ultralytics framework that already has established augmentations, loss functions, SOTA neck for feature combination, and optimized bbox prediction heads? \n\nParadox and I explored this approach, and here's our solution!\n\n## Solution Overview\n\nOur **17th place solution** consisted of 3 YOLO-like (if I can call them that) models using `convnextv2_base.fcmae_ft_in22k_in1k` as the backbone. We extracted features from P4, P3, and P2 for the neck and made predictions with a single head at P3 (stride / 8).\n\n### Model Architectures\n\nI started implementing Timm integration, but someone more intelligent than me had already done it! We just want to thank [yjwong1999](https://github.com/yjwong1999), who created the [PR that we based our approach on](https://github.com/ultralytics/ultralytics/pull/19609). We just added some stuff to it and created cfgs.\n\nWe developed two main model configurations that vary only in the \"neck\" design:\n\n#### Model 1: \n- **Backbone**: `convnextv2_base.fcmae_ft_in22k_in1k`\n- **Neck**: SPPF + C2PSA → upsample P4 and concatenate with P3 → head\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F222be1bb10a5511492902cc79f1a8379%2FCONFIG_1.png?generation=1749095926701628&alt=media)\n\n#### Model 2:  \n- **Backbone**: `convnextv2_base.fcmae_ft_in22k_in1k`\n- **Neck**: SPPF + C2PSA → upsample P4, adaptive_avg_pool2d at P2, concatenate P3 and upsampled P4 → head\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F942884ec3348873a43b94f2b56bcf3a6%2FCONFIG_2.png?generation=1749095942195298&alt=media)\n\n\nParadox had the amazing idea of removing P5 scale and features after seeing [@tatamikenn's excellent notebook](https://www.kaggle.com/code/tatamikenn/reverse-engineering-yolo) about YOLO features. This modification improved our scores by ~0.015 on the public LB.\nThe adaptive_avg_pool2d/AVG block in the architecture is from [@yyyy0201 remarkable mhaf-yolo](https://github.com/yang-0201/MHAF-YOLO).\n\n### Training Configuration\n\nWe trained with 80% of the images that contained motors with trust = 4. And additionaly used 80% [@bartley external dataset](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/569921) dataset with trust = 0. The rest was used as validation.\n\nWe used these standard training parameters:\n\n```python\nresults = model.train(\n    data=str(yaml_path),\n    epochs=40,\n    batch=2,\n    imgsz=960,\n    optimizer='AdamW',\n    lr0=1e-4,\n    lrf=0.1,\n    warmup_epochs=0,\n    dropout=0.1,\n    project=str(RUNS_DIR),\n    exist_ok=True,\n    name=f\"fold{fold_idx}\",\n    patience=100,   \n    save_period=1,\n    val=True,   \n    mosaic=0.5,\n    close_mosaic=0,\n    mixup=0.4,\n    flipud=0.5,\n    scale=0.25,\n    degrees=45,\n    seed=42,\n    deterministic=True,\n    label_smoothing=0.1,  # Note: isn't doing anything actually\n    augment=True,\n    device=0,\n)\n```\n\n### Validation Strategy Challenges\n\nValidation was problematic throughout the competition. We struggled to find correlation between our validation scores and public leaderboard performance, as our models were scoring 0.97+ in CV but showing different performance on the LB.\n\nWe also faced significant challenges in epoch selection, which led to [this discussion](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/576756).\n\nTo mitigate the epoch selection problem, we used the model soup technique to average weights of multiple epochs.\n\nUltimately, we compared models using mAP50, mAP50-95, precision, and recall curves to estimate performance. This didn't work so well.\n\n## Final Submission Details\n\n### Ensemble Strategy\n\nOur final ensemble consisted of three main components:\n\n1. First model configuration with standard training\n2. First model configuration with more augmentations:\n   ```python\n   T = [\n       A.Blur(p=0.1),\n       A.MedianBlur(p=0.1),\n       A.ToGray(p=0.1),\n       A.CLAHE(p=0.1),\n       A.RandomBrightnessContrast(p=0.1),\n       A.RandomGamma(p=0.1),\n       A.ImageCompression(quality_upper=100, quality_lower=60, p=0.1),\n       A.ShiftScaleRotate(p=0.1),\n       A.GaussNoise(p=0.1),\n       A.GaussianBlur(p=0.1),  \n       A.UnsharpMask(p=0.1)\n   ]\n   ```\n3. Second model configuration with standard training\n\n### Inference Strategy\n\nWe used `concentration = 0.5` for selecting slices, meaning we predicted on half of the available slices to increase speed.\n\nFor aggregating predictions, we experimented with HDBSCAN clustering with the following parameters:\n- min_samples: 1\n- Cluster size: 4\n- EPS: 50  \n- YOLO threshold: 0.4\n\nWe also added a confidence adjustment based on cluster size to favor bigger cluster.\n\n\n## What Didn't Work\n\n- **Alternative Backbones**: We tested `tf_efficientnetv2_l.in21k_ft_in1k`, `resnext50_32x4d.fb_swsl_ig1b_ft_in1k`, and `caformer_b36.sail_in22k`. They achieved ~0.8 performance, none surpassed ConvNeXt base.\n- **Larger Models**: ConvNeXt base was already quite huge. ConvNeXt large actually performed worse, so we stopped exploring larger architectures.\n- **Different Head**: I've tried using ConvNeXt blocks in head, improved map50-95 but not in LB.\n\n\n## Bonus PB\n\nThis probably happened to many teams, but we didn't select our top-scoring notebook for the private leaderboard. In fact, we didn't even select the top 5.\n\n**Our True Best**: Using the second model configuration alone (a different epoch soup combination than what we used in the ensemble) achieved **0.856** on the private LB, which would have been in the gold zone, indicating that P2 features were actually really important.\n\n![PB](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2Fc1c7ddc0295a36ea7048938f84d88802%2FPB.png?generation=1749095003450830&alt=media)\n\n\n## Resources & Acknowledgments\n\n- [@tatamikenn's YOLO features notebook](https://www.kaggle.com/code/tatamikenn/reverse-engineering-yolo)\n- [Jun Koda's speed optimization](https://www.kaggle.com/code/junkoda/speed-up-inference) \n- [Model Soup paper](https://arxiv.org/abs/2203.05482)\n- [MHAF-YOLO](https://github.com/yang-0201/MHAF-YOLO)\n- [Bartley External Dataset](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/569921)\n- [Timm+Ultralytcs](https://github.com/ultralytics/ultralytics/pull/19609)",
      "votes": null
    },
    {
      "id": "3217485",
      "postDate": "06/05/2025 04:08:27",
      "content": "<p>Nice work! Thanks for sharing.👏</p>",
      "rawMarkdown": "Nice work! Thanks for sharing.👏",
      "votes": null
    },
    {
      "id": "3217518",
      "postDate": "06/05/2025 05:10:03",
      "content": "<p><a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> Congratulations for 17 place. </p>\n<blockquote>\n  <p>Paradox had the amazing idea of removing P5 scale and features after seeing <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a>'s excellent notebook about YOLO features. This modification improved our scores by ~0.015 on the public LB.</p>\n</blockquote>\n<p>Thank you for mentioning. Actually, I also came up with this idea to reduce processing time, but to lazy to implement. Seems interesting it also gains score.</p>",
      "rawMarkdown": "sersasj Congratulations for 17 place. \n\n> Paradox had the amazing idea of removing P5 scale and features after seeing @tatamikenn's excellent notebook about YOLO features. This modification improved our scores by ~0.015 on the public LB.\n\nThank you for mentioning. Actually, I also came up with this idea to reduce processing time, but to lazy to implement. Seems interesting it also gains score.",
      "votes": null
    },
    {
      "id": "3217609",
      "postDate": "06/05/2025 07:47:39",
      "content": "<p>Congratulations! Looking forward to your code.</p>",
      "rawMarkdown": "Congratulations! Looking forward to your code.",
      "votes": null
    },
    {
      "id": "3217628",
      "postDate": "06/05/2025 08:11:33",
      "content": "<p><a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> Is this the definition of model 2?</p>\n<pre><code>\n\n\n\n  \n \n  \n   [, , ]\n\n\n\n  \n   [, , , [, , , , , ]] \n   [, , , [, ]] \n   [, , , [, ]] \n   [, , , [, ]] \n   [, , , [, ]] \n   [, , , []]   \n\n\n   [, , , [, , ]] \n   [, , , [, ]] \n   [[, , ], , , []] \n   [, , , [, ]] \n   [[], , , []] \n</code></pre>",
      "rawMarkdown": "sersasj Is this the definition of model 2?\n\n```yaml\n\n# Ultralytics YOLO 🚀, AGPL-3.0 license\n# 👀 ref: https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/583144#3217518\n\n# Parameters\nnc: 80 # number of classes\nscales: # there is no scaling for this model; the following scales are used only to bypass errors in model parsing\n  # [depth, width, max_channels]\n  n: [1.00, 1.00, 2048]\n\n\nbackbone:\n  # [from, number, module, args]\n  - [-1, 1, Timm, [2048, \"convnextv2_base.fcmae_ft_in22k_in1k\", True, True, 0, True]] # - 0\n  - [0, 1, Index, [128, 0]] # P2 (1, 128, 1/8, 1/8) - 1\n  - [0, 1, Index, [256, 1]] # P3 (1, 256, 1/16, 1/16) - 2\n  - [0, 1, Index, [512, 2]] # P4 (1, 512, 1/32, 1/32) - 3\n  - [-1, 1, SPPF, [512, 5]] # SPFF - 4\n  - [-1, 2, C2PSA, [512]]   # C2PSA - 5\n\nhead:\n  - [-1, 1, nn.Upsample, [None, 2, \"nearest\"]] # Upsample P4 output - 6\n  - [1, 1, nn.AvgPool2d, [2, 2]] # Avg pooling P2 output - 7\n  - [[7, 2, 6], 1, Concat, [1]] # 8\n  - [-1, 2, C3k2, [512, False]] # 9\n  - [[-1], 1, Detect, [nc]] # Detect()\n```",
      "votes": null
    },
    {
      "id": "3217646",
      "postDate": "06/05/2025 08:56:06",
      "content": "<p>Yup that's correct!</p>",
      "rawMarkdown": "Yup that's correct!",
      "votes": null
    },
    {
      "id": "3217665",
      "postDate": "06/05/2025 09:21:10",
      "content": "<p>The interesting thing here is we never used feature maps from last layer i.e. 1024x30x30</p>",
      "rawMarkdown": "The interesting thing here is we never used feature maps from last layer i.e. 1024x30x30",
      "votes": null
    },
    {
      "id": "3217681",
      "postDate": "06/05/2025 09:50:40",
      "content": "<p><a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a> Thank you for the clarification.</p>\n<blockquote>\n  <p>The interesting thing here is we never used feature maps from the last layer, i.e. 1024×30×30</p>\n</blockquote>\n<p>Is this design choice simply based on the fact that YOLO doesn't make much use of lower-resolution feature maps?  <br>\nOr did you observe a performance drop when including that layer?</p>",
      "rawMarkdown": "iamparadox Thank you for the clarification.\n\n> The interesting thing here is we never used feature maps from the last layer, i.e. 1024×30×30\n\nIs this design choice simply based on the fact that YOLO doesn't make much use of lower-resolution feature maps?  \nOr did you observe a performance drop when including that layer?",
      "votes": null
    },
    {
      "id": "3217720",
      "postDate": "06/05/2025 10:58:10",
      "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> We initially had the original standard YOLOv11 neck, upsampling and downsampling P3, P4, and P5 . It got us 0.822 on the public LB. We removed P5 and simplified the neck structure with the idea of reducing model complexity and maybe increasing its speed, but it also improved the LB score because we jumped to 0.839.</p>",
      "rawMarkdown": "tatamikenn We initially had the original standard YOLOv11 neck, upsampling and downsampling P3, P4, and P5 . It got us 0.822 on the public LB. We removed P5 and simplified the neck structure with the idea of reducing model complexity and maybe increasing its speed, but it also improved the LB score because we jumped to 0.839.",
      "votes": null
    },
    {
      "id": "3217739",
      "postDate": "06/05/2025 11:14:40",
      "content": "<p><a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> I see. Thanks.</p>",
      "rawMarkdown": "sersasj I see. Thanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3217485,
      "author_name": "eyh123",
      "author_url": "",
      "post_date": "06/05/2025 04:08:27",
      "content": "<p>Nice work! Thanks for sharing.👏</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3217518,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "06/05/2025 05:10:03",
      "content": "<p><a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> Congratulations for 17 place. </p>\n<blockquote>\n  <p>Paradox had the amazing idea of removing P5 scale and features after seeing <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a>'s excellent notebook about YOLO features. This modification improved our scores by ~0.015 on the public LB.</p>\n</blockquote>\n<p>Thank you for mentioning. Actually, I also came up with this idea to reduce processing time, but to lazy to implement. Seems interesting it also gains score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3217609,
      "author_name": "playwithme",
      "author_url": "",
      "post_date": "06/05/2025 07:47:39",
      "content": "<p>Congratulations! Looking forward to your code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3217628,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "06/05/2025 08:11:33",
      "content": "<p><a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> Is this the definition of model 2?</p>\n<pre><code>\n\n\n\n  \n \n  \n   [, , ]\n\n\n\n  \n   [, , , [, , , , , ]] \n   [, , , [, ]] \n   [, , , [, ]] \n   [, , , [, ]] \n   [, , , [, ]] \n   [, , , []]   \n\n\n   [, , , [, , ]] \n   [, , , [, ]] \n   [[, , ], , , []] \n   [, , , [, ]] \n   [[], , , []] \n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 3217646,
          "author_name": "iamparadox",
          "author_url": "",
          "post_date": "06/05/2025 08:56:06",
          "content": "<p>Yup that's correct!</p>",
          "votes": null,
          "replies": [
            {
              "id": 3217665,
              "author_name": "iamparadox",
              "author_url": "",
              "post_date": "06/05/2025 09:21:10",
              "content": "<p>The interesting thing here is we never used feature maps from last layer i.e. 1024x30x30</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3217681,
                  "author_name": "tatamikenn",
                  "author_url": "",
                  "post_date": "06/05/2025 09:50:40",
                  "content": "<p><a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a> Thank you for the clarification.</p>\n<blockquote>\n  <p>The interesting thing here is we never used feature maps from the last layer, i.e. 1024×30×30</p>\n</blockquote>\n<p>Is this design choice simply based on the fact that YOLO doesn't make much use of lower-resolution feature maps?  <br>\nOr did you observe a performance drop when including that layer?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3217720,
                      "author_name": "sersasj",
                      "author_url": "",
                      "post_date": "06/05/2025 10:58:10",
                      "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> We initially had the original standard YOLOv11 neck, upsampling and downsampling P3, P4, and P5 . It got us 0.822 on the public LB. We removed P5 and simplified the neck structure with the idea of reducing model complexity and maybe increasing its speed, but it also improved the LB score because we jumped to 0.839.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3217739,
                          "author_name": "tatamikenn",
                          "author_url": "",
                          "post_date": "06/05/2025 11:14:40",
                          "content": "<p><a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> I see. Thanks.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3217481": "# BYU Bacterial Flagellar Motors Competition - 17th Place Solution\n\nFirst, we thank the competition hosts @andrewjdarleym, @braxtonowens and Kaggle staff for organizing this competition. Below, we introduce the solution of team **Sergio Alvarez + Paradox** -- @sersasj, @iamparadox.\n\n## Context\n\n- **Business context**: [Competition overview](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/overview)\n- **Data context**: [Competition data](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/data)\n\n## Background & Evolution\n\nAfter achieving a good position in the CZII competition, our first approach (although not yet teamed up) was to try U-nets. However, this didn't work well, with the best model scoring only ~0.5 on the public leaderboard.\n\nNext, I experimented with [Jun Koda's models](https://www.kaggle.com/code/junkoda/speed-up-inference), which used Timm backbones with simple classification and segmentation heads. I managed to achieve 0.770 on the public LB but got stuck.\n\nAfter seeing discussions and public notebooks with good scores using YOLO, I decided to test it and achieved 0.792 on the public LB with YOLO10X. This got me thinking: if simple techniques with Timm backbones + a basic head got me 0.770, why not use the YOLO ultralytics framework that already has established augmentations, loss functions, SOTA neck for feature combination, and optimized bbox prediction heads? \n\nParadox and I explored this approach, and here's our solution!\n\n## Solution Overview\n\nOur **17th place solution** consisted of 3 YOLO-like (if I can call them that) models using `convnextv2_base.fcmae_ft_in22k_in1k` as the backbone. We extracted features from P4, P3, and P2 for the neck and made predictions with a single head at P3 (stride / 8).\n\n### Model Architectures\n\nI started implementing Timm integration, but someone more intelligent than me had already done it! We just want to thank [yjwong1999](https://github.com/yjwong1999), who created the [PR that we based our approach on](https://github.com/ultralytics/ultralytics/pull/19609). We just added some stuff to it and created cfgs.\n\nWe developed two main model configurations that vary only in the \"neck\" design:\n\n#### Model 1: \n- **Backbone**: `convnextv2_base.fcmae_ft_in22k_in1k`\n- **Neck**: SPPF + C2PSA → upsample P4 and concatenate with P3 → head\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F222be1bb10a5511492902cc79f1a8379%2FCONFIG_1.png?generation=1749095926701628&alt=media)\n\n#### Model 2:  \n- **Backbone**: `convnextv2_base.fcmae_ft_in22k_in1k`\n- **Neck**: SPPF + C2PSA → upsample P4, adaptive_avg_pool2d at P2, concatenate P3 and upsampled P4 → head\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F942884ec3348873a43b94f2b56bcf3a6%2FCONFIG_2.png?generation=1749095942195298&alt=media)\n\n\nParadox had the amazing idea of removing P5 scale and features after seeing [@tatamikenn's excellent notebook](https://www.kaggle.com/code/tatamikenn/reverse-engineering-yolo) about YOLO features. This modification improved our scores by ~0.015 on the public LB.\nThe adaptive_avg_pool2d/AVG block in the architecture is from [@yyyy0201 remarkable mhaf-yolo](https://github.com/yang-0201/MHAF-YOLO).\n\n### Training Configuration\n\nWe trained with 80% of the images that contained motors with trust = 4. And additionaly used 80% [@bartley external dataset](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/569921) dataset with trust = 0. The rest was used as validation.\n\nWe used these standard training parameters:\n\n```python\nresults = model.train(\n    data=str(yaml_path),\n    epochs=40,\n    batch=2,\n    imgsz=960,\n    optimizer='AdamW',\n    lr0=1e-4,\n    lrf=0.1,\n    warmup_epochs=0,\n    dropout=0.1,\n    project=str(RUNS_DIR),\n    exist_ok=True,\n    name=f\"fold{fold_idx}\",\n    patience=100,   \n    save_period=1,\n    val=True,   \n    mosaic=0.5,\n    close_mosaic=0,\n    mixup=0.4,\n    flipud=0.5,\n    scale=0.25,\n    degrees=45,\n    seed=42,\n    deterministic=True,\n    label_smoothing=0.1,  # Note: isn't doing anything actually\n    augment=True,\n    device=0,\n)\n```\n\n### Validation Strategy Challenges\n\nValidation was problematic throughout the competition. We struggled to find correlation between our validation scores and public leaderboard performance, as our models were scoring 0.97+ in CV but showing different performance on the LB.\n\nWe also faced significant challenges in epoch selection, which led to [this discussion](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/576756).\n\nTo mitigate the epoch selection problem, we used the model soup technique to average weights of multiple epochs.\n\nUltimately, we compared models using mAP50, mAP50-95, precision, and recall curves to estimate performance. This didn't work so well.\n\n## Final Submission Details\n\n### Ensemble Strategy\n\nOur final ensemble consisted of three main components:\n\n1. First model configuration with standard training\n2. First model configuration with more augmentations:\n   ```python\n   T = [\n       A.Blur(p=0.1),\n       A.MedianBlur(p=0.1),\n       A.ToGray(p=0.1),\n       A.CLAHE(p=0.1),\n       A.RandomBrightnessContrast(p=0.1),\n       A.RandomGamma(p=0.1),\n       A.ImageCompression(quality_upper=100, quality_lower=60, p=0.1),\n       A.ShiftScaleRotate(p=0.1),\n       A.GaussNoise(p=0.1),\n       A.GaussianBlur(p=0.1),  \n       A.UnsharpMask(p=0.1)\n   ]\n   ```\n3. Second model configuration with standard training\n\n### Inference Strategy\n\nWe used `concentration = 0.5` for selecting slices, meaning we predicted on half of the available slices to increase speed.\n\nFor aggregating predictions, we experimented with HDBSCAN clustering with the following parameters:\n- min_samples: 1\n- Cluster size: 4\n- EPS: 50  \n- YOLO threshold: 0.4\n\nWe also added a confidence adjustment based on cluster size to favor bigger cluster.\n\n\n## What Didn't Work\n\n- **Alternative Backbones**: We tested `tf_efficientnetv2_l.in21k_ft_in1k`, `resnext50_32x4d.fb_swsl_ig1b_ft_in1k`, and `caformer_b36.sail_in22k`. They achieved ~0.8 performance, none surpassed ConvNeXt base.\n- **Larger Models**: ConvNeXt base was already quite huge. ConvNeXt large actually performed worse, so we stopped exploring larger architectures.\n- **Different Head**: I've tried using ConvNeXt blocks in head, improved map50-95 but not in LB.\n\n\n## Bonus PB\n\nThis probably happened to many teams, but we didn't select our top-scoring notebook for the private leaderboard. In fact, we didn't even select the top 5.\n\n**Our True Best**: Using the second model configuration alone (a different epoch soup combination than what we used in the ensemble) achieved **0.856** on the private LB, which would have been in the gold zone, indicating that P2 features were actually really important.\n\n![PB](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2Fc1c7ddc0295a36ea7048938f84d88802%2FPB.png?generation=1749095003450830&alt=media)\n\n\n## Resources & Acknowledgments\n\n- [@tatamikenn's YOLO features notebook](https://www.kaggle.com/code/tatamikenn/reverse-engineering-yolo)\n- [Jun Koda's speed optimization](https://www.kaggle.com/code/junkoda/speed-up-inference) \n- [Model Soup paper](https://arxiv.org/abs/2203.05482)\n- [MHAF-YOLO](https://github.com/yang-0201/MHAF-YOLO)\n- [Bartley External Dataset](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/569921)\n- [Timm+Ultralytcs](https://github.com/ultralytics/ultralytics/pull/19609)",
    "3217485": "Nice work! Thanks for sharing.👏",
    "3217518": "sersasj Congratulations for 17 place. \n\n> Paradox had the amazing idea of removing P5 scale and features after seeing @tatamikenn's excellent notebook about YOLO features. This modification improved our scores by ~0.015 on the public LB.\n\nThank you for mentioning. Actually, I also came up with this idea to reduce processing time, but to lazy to implement. Seems interesting it also gains score.",
    "3217609": "Congratulations! Looking forward to your code.",
    "3217628": "sersasj Is this the definition of model 2?\n\n```yaml\n\n# Ultralytics YOLO 🚀, AGPL-3.0 license\n# 👀 ref: https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/583144#3217518\n\n# Parameters\nnc: 80 # number of classes\nscales: # there is no scaling for this model; the following scales are used only to bypass errors in model parsing\n  # [depth, width, max_channels]\n  n: [1.00, 1.00, 2048]\n\n\nbackbone:\n  # [from, number, module, args]\n  - [-1, 1, Timm, [2048, \"convnextv2_base.fcmae_ft_in22k_in1k\", True, True, 0, True]] # - 0\n  - [0, 1, Index, [128, 0]] # P2 (1, 128, 1/8, 1/8) - 1\n  - [0, 1, Index, [256, 1]] # P3 (1, 256, 1/16, 1/16) - 2\n  - [0, 1, Index, [512, 2]] # P4 (1, 512, 1/32, 1/32) - 3\n  - [-1, 1, SPPF, [512, 5]] # SPFF - 4\n  - [-1, 2, C2PSA, [512]]   # C2PSA - 5\n\nhead:\n  - [-1, 1, nn.Upsample, [None, 2, \"nearest\"]] # Upsample P4 output - 6\n  - [1, 1, nn.AvgPool2d, [2, 2]] # Avg pooling P2 output - 7\n  - [[7, 2, 6], 1, Concat, [1]] # 8\n  - [-1, 2, C3k2, [512, False]] # 9\n  - [[-1], 1, Detect, [nc]] # Detect()\n```",
    "3217646": "Yup that's correct!",
    "3217665": "The interesting thing here is we never used feature maps from last layer i.e. 1024x30x30",
    "3217681": "iamparadox Thank you for the clarification.\n\n> The interesting thing here is we never used feature maps from the last layer, i.e. 1024×30×30\n\nIs this design choice simply based on the fact that YOLO doesn't make much use of lower-resolution feature maps?  \nOr did you observe a performance drop when including that layer?",
    "3217720": "tatamikenn We initially had the original standard YOLOv11 neck, upsampling and downsampling P3, P4, and P5 . It got us 0.822 on the public LB. We removed P5 and simplified the neck structure with the idea of reducing model complexity and maybe increasing its speed, but it also improved the LB score because we jumped to 0.839.",
    "3217739": "sersasj I see. Thanks."
  },
  "source": "meta"
}