{
  "id": 601039,
  "title": " ⭐🥈2nd Place Solution 🥈⭐",
  "url": "/competitions/multi-class-object-detection-challenge/writeups/2nd-place-solution",
  "author_name": "",
  "post_date": "2025-08-26T21:49:48.363Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<h1>Solution Overview</h1>\n<p>First off, a big thank you to the organisers for putting on this challenge and for hosting the live sessions. <a href=\"https://www.kaggle.com/rebekahduality\" target=\"_blank\">@rebekahduality</a> <a href=\"https://www.kaggle.com/rishikeshjadhav22\" target=\"_blank\">@rishikeshjadhav22</a> <br>\nIt was a lot of fun to take part and to see how other competitors tackled the problem. ❤️🤩</p>\n<p>I wanted to share how I approached this challenge and what worked for me (and what didnt) . Hopefully it’s helpful for anyone coming back to this problem later or just curious about how the top solutions were built.</p>\n<hr>\n<h2>Data sources</h2>\n<p>The training data went well beyond the starter set. Here’s everything I used:</p>\n<p>1 - <strong>Multi-class Object Detection Challenge</strong> – the official competition dataset.  <br>\nI used only the <strong>train set</strong> for training. For validation, I relied exclusively on the <strong>real portion of the validation set</strong> to ensure that the validation set better reflects the target domain.  <br>\n(+1000 images)</p>\n<p>2 - <a href=\"https://www.kaggle.com/datasets/thelastsmilodon/extra-synthetic-data\" target=\"_blank\"><strong>Extra synthetic data</strong></a> – my main synthetic dataset generated using <strong>FalconCloud</strong>.  <br>\nIt is divided into nine output folders (<code>output (1)</code> through <code>output (9)</code>), each containing its own <code>images</code> and <code>labels</code> subdirectories.  </p>\n<p>My approach was two-staged:  </p>\n<ul>\n<li>In the first 5–6 scenarios, I focused on generating as much data as possible to enrich the dataset.  </li>\n<li>Afterwards, I relied on the <strong>Controlled scenario</strong> to address specific shortcomings observed during model training:  <ul>\n<li><strong>Misclassification of cups, candles, and cylindrical objects as soup cans</strong>  <br>\nImproved by introducing twin objects and capturing multiple views from varied angles and distances.  </li>\n<li><strong>Failure to detect occluded objects</strong>  <br>\nImproved by creating scenarios with similar occlusions and ensuring diverse captures.  </li>\n<li><strong>Failure to detect distant soup cans</strong>  <br>\nImproved by adding distant images under varied lighting and backgrounds.  </li>\n<li><strong>Limited diversity in object placement</strong>  <br>\nImproved by frequently rearranging objects, varying orientation, and including edge-case scenarios.  </li></ul></li>\n</ul>\n<p>(+904 images)</p>\n<p>3 - <a href=\"https://www.kaggle.com/datasets/kadirkrtls/falcon-multiclass-cheerios-soupv2\" target=\"_blank\"><strong>Falcon-Multiclass-Cheerios-Soup V2</strong></a> – a synthetic dataset generated via FalconCloud by <a href=\"https://www.kaggle.com/kadirkrtls\" target=\"_blank\">@kadirkrtls</a> (thanks for sharing).  <br>\n<em>(Didn’t use it in my top notebook though)</em></p>\n<p>4 - <strong>Sample synthetic data generated</strong> – a basic sample of synthetic images I created on FalconCloud and shared publicly at the start of the competition.  <br>\n(+100 images)</p>\n<p>5 - <a href=\"https://www.kaggle.com/datasets/thelastsmilodon/synthetic-soup1\" target=\"_blank\"><strong>Synthetic-soup1</strong></a> – dataset from the previous soup detection competition.  <br>\nI used it in several experiments, but the improvements (if any) were not worth the extra training time, so I excluded it from the final training.  </p>\n<p>In total, the combined dataset finally consisted of <strong>~2,000 images</strong> synthetic images for training,  carefully curated to maximize diversity and domain relevance. As for the validation set it consisted of <strong>77 images</strong> real images.  </p>\n<blockquote>\n<pre><code>\nbase_path       = \nsynthetic_base  = \nsynthetic_base2 = \nsynthetic_base3 = \n\n\nsynthetic_train_dirs = [(p)  p  Path(synthetic_base).rglob()]\nsynthetic_val_dirs   = [(p)  p  Path(synthetic_base).rglob()]\n\n\nsynthetic_train_dirs2 = [(p)  p  Path(synthetic_base2).rglob()]\nsynthetic_val_dirs2   = [(p)  p  Path(synthetic_base2).rglob()]\n\n\nsynthetic_train_dirs3 = [(p)  p  Path(synthetic_base3).rglob()]\nsynthetic_val_dirs3   = [(p)  p  Path(synthetic_base3).rglob()]\n\n\ndata_yaml = {\n    : (\n        [\n            ,\n            ,\n            ,\n        ]\n        + synthetic_train_dirs\n        + synthetic_val_dirs\n        \n        \n        \n        \n    ),\n    :   ,\n    :  ,\n    :    ,\n    : [, ]\n}\n\n\n (, )  f:\n    yaml.safe_dump(data_yaml, f, default_flow_style=)\n</code></pre>\n  <blockquote>\n    <p>rest f the functions can be found in the notebooks attached</p>\n  </blockquote>\n</blockquote>\n<hr>\n<h2>Training pipeline</h2>\n<p>I used the Ultralytics implementation of YOLO to train three separate models with different hyperparameters on the combined dataset. The training notebook installed the <code>ultralytics</code> package and <code>ensemble‑boxes</code> for post‑processing. Key settings were:</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Backbone / Architecture</th>\n<th>Image Size</th>\n<th>Epochs</th>\n<th>Batch Size</th>\n<th>Learning Rate</th>\n<th>weight decay</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Model 1</td>\n<td>yolo11x</td>\n<td>672 (with multiscale)</td>\n<td>12</td>\n<td>4</td>\n<td>1e-3</td>\n<td>3e-3</td>\n</tr>\n<tr>\n<td>Model 2</td>\n<td>YOLO12x</td>\n<td>512</td>\n<td>20</td>\n<td>16</td>\n<td>1e-3</td>\n<td>1e-3</td>\n</tr>\n<tr>\n<td>Model 3</td>\n<td>YOLOv8x</td>\n<td>640</td>\n<td>30</td>\n<td>32</td>\n<td>1e-3</td>\n<td>2e-3</td>\n</tr>\n</tbody>\n</table>\n<h3>Notes on Training and Hyperparameters</h3>\n<ul>\n<li><p><strong>Hyperparameter Selection</strong>  <br>\nMost of the hyperparameters were tuned through extensive trial and error. Surprisingly, I found that <strong>heavy augmentations</strong> actually hurt performance and significantly increased training time. Instead, using <strong>softer augmentations</strong> (especially color/saturation-related) along with <strong>MixUp</strong> and <strong>CutMix</strong> yielded much better results. More importantly, adding <strong>diverse synthetic images</strong> turned out to be the single most effective way to improve generalization. <em>(Exact values can be found in the training notebook.)</em></p></li>\n<li><p><strong>Optimizer Choice (SGD)</strong>  <br>\nI chose <strong>SGD with momentum</strong> instead of the default Adam/AdamW because it provided more stable convergence and better final accuracy in this competition setting. It also converged much faster, allowing the use of merely 15-30 epochs . </p></li>\n<li><p><strong>Image Size</strong>  <br>\nTraining with <strong>larger image sizes</strong> offered noticeable boosts in mAP, especially for smaller objects like soup cans and cups. However, this came at the cost of much longer training times and occasional <strong>out-of-memory (OOM) errors</strong> on Kaggle GPUs. A balance had to be struck between accuracy gains and computational feasibility.</p></li>\n<li><p><strong>Batch Size and Accumulation</strong>  <br>\nDue to memory constraints with larger image sizes, I experimented with <strong>smaller batch sizes combined with gradient accumulation</strong>, which allowed stable training without compromising performance.</p></li>\n<li><p><strong>Learning Rate Scheduling</strong>  <br>\nA <strong>cosine annealing schedule</strong> with warm restarts worked best, enabling the model to escape local minima and generalize better. Step-based schedules were also tested but performed slightly worse.</p></li>\n<li><p><strong>Freezing initial layers</strong>  <br>\nTraining only the last few layers was very helpful as it reduced training time and memory usage allowing the use of much larger image sizes and model architectures. </p></li>\n<li><p><strong>Augmentation Insights</strong>  </p>\n<ul>\n<li><em>Color augmentations</em> (hue, saturation, brightness) were useful in <strong>moderation</strong>, but when applied too aggressively they actually <strong>slowed convergence</strong> and increased training time without delivering meaningful performance gains.  </li>\n<li><em>Synthetic diversity</em> — introducing variations in <strong>object placement, occlusion patterns, and lighting conditions</strong> — proved to be <strong>far more impactful</strong> than heavy augmentations, offering better generalization and robustness.  </li>\n<li><em>MixUp and CutMix</em> were particularly effective for this task, as they <strong>regularized the models</strong> by creating blended training examples and reducing overconfidence on single features. This helped the ensemble handle <strong>class overlap and cluttered backgrounds</strong> more reliably.  </li></ul></li>\n<li><p><strong>Validation Strategy</strong>  <br>\nUsing only the <strong>real portion of the validation set</strong> was crucial, since validating on synthetic images gave overly optimistic results and did not reflect the target domain.</p></li>\n</ul>\n<hr>\n<h2>Ensembling and post‑processing</h2>\n<p>After training the three models (YOLOv11x, YOLOv12x, YOLOv8x), I generated predictions on the test set and applied an ensembling pipeline followed by lightweight post-processing:</p>\n<ol>\n<li><p><strong>Model Ensembling with Weighted Box Fusion (WBF):</strong>  <br>\nThe outputs of the three models were combined using the <code>ensemble-boxes</code> library.  </p>\n<ul>\n<li>WBF merges highly overlapping bounding boxes across models by computing a <strong>confidence-weighted average of their coordinates</strong>.  </li>\n<li>This approach leverages the complementary strengths of each model:  <ul>\n<li>YOLOv11x (672, multiscale) contributed robustness to scale variations.  </li>\n<li>YOLOv12x (512) captured efficiency and faster convergence.  </li>\n<li>YOLOv8x (640) provided strong generalisation and stability.  </li></ul></li>\n<li>The ensemble was consistently more reliable than any single model, improving localisation accuracy and reducing duplicate detections.</li></ul></li>\n<li><p><strong>Confidence Filtering (Top-k per class):</strong>  <br>\nFor each class in every image, I retained only the <strong>top-2 predictions ranked by confidence</strong>.  </p>\n<ul>\n<li>This simple heuristic removed low-confidence false positives without hurting recall.  </li>\n<li>It was especially effective in reducing spurious boxes around small or cluttered objects (cups, soup cans, etc.).  </li>\n<li>After this step, the final ensemble achieved a <strong>mAP50 score of 0.996</strong>, showing that even lightweight post-processing can significantly polish ensemble results.</li></ul></li>\n</ol>\n<blockquote>\n<pre><code> os\n pathlib  Path\n pandas  pd\n csv\n ultralytics  YOLO\n ensemble_boxes  weighted_boxes_fusion\n PIL  Image\n\nmodel_paths = [\n    ,\n    ,\n    ,\n]\n\ntest_images_path = \noutput_dir = \n\nconf = \niou_thr = \nskip_box_thr = conf\nimage_sizes = [,,,,,,,]\n\nmodels = [YOLO(path)  path  model_paths]\npredictions = run_inference(models, image_sizes, test_images_path,conf=conf,iou_thr=iou_thr)\n\nimage_ids = (((((predictions.values())).values())).keys())\n\napply_wbf_and_save_final_submission(predictions, image_ids,iou_thr=,skip_box_thr=,conf_post=)\n</code></pre>\n</blockquote>\n<hr>\n<h2>Reproducibility and tips</h2>\n<ul>\n<li>All data paths and hyper‑parameters are stored in the notebook’s <code>data.yaml</code> and <code>args.yaml</code>, so rerunning on similar hardware should reproduce the results.</li>\n<li>When generating additional synthetic data in FalconCloud, vary the camera angles, lighting and object scales. This diversity helps the model generalise better.</li>\n<li>If you have access to more powerful hardware, increasing the image size (e.g. 800) and batch size may yield further improvements. Larger YOLO models (like <code>yolov9e</code>) and longer training schedules are also worth exploring.</li>\n</ul>\n<hr>\n<h2>🏁 Closing Thoughts</h2>\n<p>Bringing together the <strong>official competition dataset</strong>, <strong>multiple synthetic datasets</strong>, and a <strong>private dataset</strong> resulted in a training set that was both <strong>rich</strong> and <strong>diverse</strong>. This variety played a key role in helping the models generalize better to real-world validation data.  </p>\n<p>Training <strong>three different YOLO models</strong> and combining their predictions through <strong>Weighted Box Fusion (WBF)</strong> — followed by a lightweight <strong>top-k filtering step</strong> — turned out to be a simple yet very effective strategy. This ensemble approach consistently improved localisation, reduced duplicate detections, and ultimately delivered a <strong>high mAP50 score</strong> on the leaderboard.  </p>\n<h3>💬 Feedback &amp; Updates</h3>\n<ul>\n<li>If you have any <strong>comments or questions</strong>, feel free to share them below — I’d be glad to discuss.  </li>\n<li>I will continue to <strong>update this write-up</strong> if necessary.  </li>\n</ul>\n<hr>\n<p>⭐ <strong>If you found this helpful, please consider upvoting — it helps a lot. Thanks!</strong></p>",
  "messages": [
    {
      "id": "3275211",
      "postDate": "08/26/2025 06:08:05",
      "content": "<h1>Solution Overview</h1>\n<p>First off, a big thank you to the organisers for putting on this challenge and for hosting the live sessions. <a href=\"https://www.kaggle.com/rebekahduality\" target=\"_blank\">@rebekahduality</a> <a href=\"https://www.kaggle.com/rishikeshjadhav22\" target=\"_blank\">@rishikeshjadhav22</a> <br>\nIt was a lot of fun to take part and to see how other competitors tackled the problem. ❤️🤩</p>\n<p>I wanted to share how I approached this challenge and what worked for me (and what didnt) . Hopefully it’s helpful for anyone coming back to this problem later or just curious about how the top solutions were built.</p>\n<hr>\n<h2>Data sources</h2>\n<p>The training data went well beyond the starter set. Here’s everything I used:</p>\n<p>1 - <strong>Multi-class Object Detection Challenge</strong> – the official competition dataset.  <br>\nI used only the <strong>train set</strong> for training. For validation, I relied exclusively on the <strong>real portion of the validation set</strong> to ensure that the validation set better reflects the target domain.  <br>\n(+1000 images)</p>\n<p>2 - <a href=\"https://www.kaggle.com/datasets/thelastsmilodon/extra-synthetic-data\" target=\"_blank\"><strong>Extra synthetic data</strong></a> – my main synthetic dataset generated using <strong>FalconCloud</strong>.  <br>\nIt is divided into nine output folders (<code>output (1)</code> through <code>output (9)</code>), each containing its own <code>images</code> and <code>labels</code> subdirectories.  </p>\n<p>My approach was two-staged:  </p>\n<ul>\n<li>In the first 5–6 scenarios, I focused on generating as much data as possible to enrich the dataset.  </li>\n<li>Afterwards, I relied on the <strong>Controlled scenario</strong> to address specific shortcomings observed during model training:  <ul>\n<li><strong>Misclassification of cups, candles, and cylindrical objects as soup cans</strong>  <br>\nImproved by introducing twin objects and capturing multiple views from varied angles and distances.  </li>\n<li><strong>Failure to detect occluded objects</strong>  <br>\nImproved by creating scenarios with similar occlusions and ensuring diverse captures.  </li>\n<li><strong>Failure to detect distant soup cans</strong>  <br>\nImproved by adding distant images under varied lighting and backgrounds.  </li>\n<li><strong>Limited diversity in object placement</strong>  <br>\nImproved by frequently rearranging objects, varying orientation, and including edge-case scenarios.  </li></ul></li>\n</ul>\n<p>(+904 images)</p>\n<p>3 - <a href=\"https://www.kaggle.com/datasets/kadirkrtls/falcon-multiclass-cheerios-soupv2\" target=\"_blank\"><strong>Falcon-Multiclass-Cheerios-Soup V2</strong></a> – a synthetic dataset generated via FalconCloud by <a href=\"https://www.kaggle.com/kadirkrtls\" target=\"_blank\">@kadirkrtls</a> (thanks for sharing).  <br>\n<em>(Didn’t use it in my top notebook though)</em></p>\n<p>4 - <strong>Sample synthetic data generated</strong> – a basic sample of synthetic images I created on FalconCloud and shared publicly at the start of the competition.  <br>\n(+100 images)</p>\n<p>5 - <a href=\"https://www.kaggle.com/datasets/thelastsmilodon/synthetic-soup1\" target=\"_blank\"><strong>Synthetic-soup1</strong></a> – dataset from the previous soup detection competition.  <br>\nI used it in several experiments, but the improvements (if any) were not worth the extra training time, so I excluded it from the final training.  </p>\n<p>In total, the combined dataset finally consisted of <strong>~2,000 images</strong> synthetic images for training,  carefully curated to maximize diversity and domain relevance. As for the validation set it consisted of <strong>77 images</strong> real images.  </p>\n<blockquote>\n<pre><code>\nbase_path       = \nsynthetic_base  = \nsynthetic_base2 = \nsynthetic_base3 = \n\n\nsynthetic_train_dirs = [(p)  p  Path(synthetic_base).rglob()]\nsynthetic_val_dirs   = [(p)  p  Path(synthetic_base).rglob()]\n\n\nsynthetic_train_dirs2 = [(p)  p  Path(synthetic_base2).rglob()]\nsynthetic_val_dirs2   = [(p)  p  Path(synthetic_base2).rglob()]\n\n\nsynthetic_train_dirs3 = [(p)  p  Path(synthetic_base3).rglob()]\nsynthetic_val_dirs3   = [(p)  p  Path(synthetic_base3).rglob()]\n\n\ndata_yaml = {\n    : (\n        [\n            ,\n            ,\n            ,\n        ]\n        + synthetic_train_dirs\n        + synthetic_val_dirs\n        \n        \n        \n        \n    ),\n    :   ,\n    :  ,\n    :    ,\n    : [, ]\n}\n\n\n (, )  f:\n    yaml.safe_dump(data_yaml, f, default_flow_style=)\n</code></pre>\n  <blockquote>\n    <p>rest f the functions can be found in the notebooks attached</p>\n  </blockquote>\n</blockquote>\n<hr>\n<h2>Training pipeline</h2>\n<p>I used the Ultralytics implementation of YOLO to train three separate models with different hyperparameters on the combined dataset. The training notebook installed the <code>ultralytics</code> package and <code>ensemble‑boxes</code> for post‑processing. Key settings were:</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Backbone / Architecture</th>\n<th>Image Size</th>\n<th>Epochs</th>\n<th>Batch Size</th>\n<th>Learning Rate</th>\n<th>weight decay</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Model 1</td>\n<td>yolo11x</td>\n<td>672 (with multiscale)</td>\n<td>12</td>\n<td>4</td>\n<td>1e-3</td>\n<td>3e-3</td>\n</tr>\n<tr>\n<td>Model 2</td>\n<td>YOLO12x</td>\n<td>512</td>\n<td>20</td>\n<td>16</td>\n<td>1e-3</td>\n<td>1e-3</td>\n</tr>\n<tr>\n<td>Model 3</td>\n<td>YOLOv8x</td>\n<td>640</td>\n<td>30</td>\n<td>32</td>\n<td>1e-3</td>\n<td>2e-3</td>\n</tr>\n</tbody>\n</table>\n<h3>Notes on Training and Hyperparameters</h3>\n<ul>\n<li><p><strong>Hyperparameter Selection</strong>  <br>\nMost of the hyperparameters were tuned through extensive trial and error. Surprisingly, I found that <strong>heavy augmentations</strong> actually hurt performance and significantly increased training time. Instead, using <strong>softer augmentations</strong> (especially color/saturation-related) along with <strong>MixUp</strong> and <strong>CutMix</strong> yielded much better results. More importantly, adding <strong>diverse synthetic images</strong> turned out to be the single most effective way to improve generalization. <em>(Exact values can be found in the training notebook.)</em></p></li>\n<li><p><strong>Optimizer Choice (SGD)</strong>  <br>\nI chose <strong>SGD with momentum</strong> instead of the default Adam/AdamW because it provided more stable convergence and better final accuracy in this competition setting. It also converged much faster, allowing the use of merely 15-30 epochs . </p></li>\n<li><p><strong>Image Size</strong>  <br>\nTraining with <strong>larger image sizes</strong> offered noticeable boosts in mAP, especially for smaller objects like soup cans and cups. However, this came at the cost of much longer training times and occasional <strong>out-of-memory (OOM) errors</strong> on Kaggle GPUs. A balance had to be struck between accuracy gains and computational feasibility.</p></li>\n<li><p><strong>Batch Size and Accumulation</strong>  <br>\nDue to memory constraints with larger image sizes, I experimented with <strong>smaller batch sizes combined with gradient accumulation</strong>, which allowed stable training without compromising performance.</p></li>\n<li><p><strong>Learning Rate Scheduling</strong>  <br>\nA <strong>cosine annealing schedule</strong> with warm restarts worked best, enabling the model to escape local minima and generalize better. Step-based schedules were also tested but performed slightly worse.</p></li>\n<li><p><strong>Freezing initial layers</strong>  <br>\nTraining only the last few layers was very helpful as it reduced training time and memory usage allowing the use of much larger image sizes and model architectures. </p></li>\n<li><p><strong>Augmentation Insights</strong>  </p>\n<ul>\n<li><em>Color augmentations</em> (hue, saturation, brightness) were useful in <strong>moderation</strong>, but when applied too aggressively they actually <strong>slowed convergence</strong> and increased training time without delivering meaningful performance gains.  </li>\n<li><em>Synthetic diversity</em> — introducing variations in <strong>object placement, occlusion patterns, and lighting conditions</strong> — proved to be <strong>far more impactful</strong> than heavy augmentations, offering better generalization and robustness.  </li>\n<li><em>MixUp and CutMix</em> were particularly effective for this task, as they <strong>regularized the models</strong> by creating blended training examples and reducing overconfidence on single features. This helped the ensemble handle <strong>class overlap and cluttered backgrounds</strong> more reliably.  </li></ul></li>\n<li><p><strong>Validation Strategy</strong>  <br>\nUsing only the <strong>real portion of the validation set</strong> was crucial, since validating on synthetic images gave overly optimistic results and did not reflect the target domain.</p></li>\n</ul>\n<hr>\n<h2>Ensembling and post‑processing</h2>\n<p>After training the three models (YOLOv11x, YOLOv12x, YOLOv8x), I generated predictions on the test set and applied an ensembling pipeline followed by lightweight post-processing:</p>\n<ol>\n<li><p><strong>Model Ensembling with Weighted Box Fusion (WBF):</strong>  <br>\nThe outputs of the three models were combined using the <code>ensemble-boxes</code> library.  </p>\n<ul>\n<li>WBF merges highly overlapping bounding boxes across models by computing a <strong>confidence-weighted average of their coordinates</strong>.  </li>\n<li>This approach leverages the complementary strengths of each model:  <ul>\n<li>YOLOv11x (672, multiscale) contributed robustness to scale variations.  </li>\n<li>YOLOv12x (512) captured efficiency and faster convergence.  </li>\n<li>YOLOv8x (640) provided strong generalisation and stability.  </li></ul></li>\n<li>The ensemble was consistently more reliable than any single model, improving localisation accuracy and reducing duplicate detections.</li></ul></li>\n<li><p><strong>Confidence Filtering (Top-k per class):</strong>  <br>\nFor each class in every image, I retained only the <strong>top-2 predictions ranked by confidence</strong>.  </p>\n<ul>\n<li>This simple heuristic removed low-confidence false positives without hurting recall.  </li>\n<li>It was especially effective in reducing spurious boxes around small or cluttered objects (cups, soup cans, etc.).  </li>\n<li>After this step, the final ensemble achieved a <strong>mAP50 score of 0.996</strong>, showing that even lightweight post-processing can significantly polish ensemble results.</li></ul></li>\n</ol>\n<blockquote>\n<pre><code> os\n pathlib  Path\n pandas  pd\n csv\n ultralytics  YOLO\n ensemble_boxes  weighted_boxes_fusion\n PIL  Image\n\nmodel_paths = [\n    ,\n    ,\n    ,\n]\n\ntest_images_path = \noutput_dir = \n\nconf = \niou_thr = \nskip_box_thr = conf\nimage_sizes = [,,,,,,,]\n\nmodels = [YOLO(path)  path  model_paths]\npredictions = run_inference(models, image_sizes, test_images_path,conf=conf,iou_thr=iou_thr)\n\nimage_ids = (((((predictions.values())).values())).keys())\n\napply_wbf_and_save_final_submission(predictions, image_ids,iou_thr=,skip_box_thr=,conf_post=)\n</code></pre>\n</blockquote>\n<hr>\n<h2>Reproducibility and tips</h2>\n<ul>\n<li>All data paths and hyper‑parameters are stored in the notebook’s <code>data.yaml</code> and <code>args.yaml</code>, so rerunning on similar hardware should reproduce the results.</li>\n<li>When generating additional synthetic data in FalconCloud, vary the camera angles, lighting and object scales. This diversity helps the model generalise better.</li>\n<li>If you have access to more powerful hardware, increasing the image size (e.g. 800) and batch size may yield further improvements. Larger YOLO models (like <code>yolov9e</code>) and longer training schedules are also worth exploring.</li>\n</ul>\n<hr>\n<h2>🏁 Closing Thoughts</h2>\n<p>Bringing together the <strong>official competition dataset</strong>, <strong>multiple synthetic datasets</strong>, and a <strong>private dataset</strong> resulted in a training set that was both <strong>rich</strong> and <strong>diverse</strong>. This variety played a key role in helping the models generalize better to real-world validation data.  </p>\n<p>Training <strong>three different YOLO models</strong> and combining their predictions through <strong>Weighted Box Fusion (WBF)</strong> — followed by a lightweight <strong>top-k filtering step</strong> — turned out to be a simple yet very effective strategy. This ensemble approach consistently improved localisation, reduced duplicate detections, and ultimately delivered a <strong>high mAP50 score</strong> on the leaderboard.  </p>\n<h3>💬 Feedback &amp; Updates</h3>\n<ul>\n<li>If you have any <strong>comments or questions</strong>, feel free to share them below — I’d be glad to discuss.  </li>\n<li>I will continue to <strong>update this write-up</strong> if necessary.  </li>\n</ul>\n<hr>\n<p>⭐ <strong>If you found this helpful, please consider upvoting — it helps a lot. Thanks!</strong></p>",
      "rawMarkdown": "# Solution Overview\n\nFirst off, a big thank you to the organisers for putting on this challenge and for hosting the live sessions. @rebekahduality @rishikeshjadhav22 \nIt was a lot of fun to take part and to see how other competitors tackled the problem. ❤️🤩\n\nI wanted to share how I approached this challenge and what worked for me (and what didnt) . Hopefully it’s helpful for anyone coming back to this problem later or just curious about how the top solutions were built.\n\n---\n\n## Data sources\n\nThe training data went well beyond the starter set. Here’s everything I used:\n\n1 - **Multi-class Object Detection Challenge** – the official competition dataset.  \nI used only the **train set** for training. For validation, I relied exclusively on the **real portion of the validation set** to ensure that the validation set better reflects the target domain.  \n(+1000 images)\n\n2 - [**Extra synthetic data**](https://www.kaggle.com/datasets/thelastsmilodon/extra-synthetic-data) – my main synthetic dataset generated using **FalconCloud**.  \nIt is divided into nine output folders (`output (1)` through `output (9)`), each containing its own `images` and `labels` subdirectories.  \n\nMy approach was two-staged:  \n- In the first 5–6 scenarios, I focused on generating as much data as possible to enrich the dataset.  \n- Afterwards, I relied on the **Controlled scenario** to address specific shortcomings observed during model training:  \n  * **Misclassification of cups, candles, and cylindrical objects as soup cans**  \n    Improved by introducing twin objects and capturing multiple views from varied angles and distances.  \n  * **Failure to detect occluded objects**  \n    Improved by creating scenarios with similar occlusions and ensuring diverse captures.  \n  * **Failure to detect distant soup cans**  \n    Improved by adding distant images under varied lighting and backgrounds.  \n  * **Limited diversity in object placement**  \n    Improved by frequently rearranging objects, varying orientation, and including edge-case scenarios.  \n\n(+904 images)\n\n3 - [**Falcon-Multiclass-Cheerios-Soup V2**](https://www.kaggle.com/datasets/kadirkrtls/falcon-multiclass-cheerios-soupv2) – a synthetic dataset generated via FalconCloud by @kadirkrtls (thanks for sharing).  \n*(Didn’t use it in my top notebook though)*\n\n4 - **Sample synthetic data generated** – a basic sample of synthetic images I created on FalconCloud and shared publicly at the start of the competition.  \n(+100 images)\n\n5 - [**Synthetic-soup1**](https://www.kaggle.com/datasets/thelastsmilodon/synthetic-soup1) – dataset from the previous soup detection competition.  \nI used it in several experiments, but the improvements (if any) were not worth the extra training time, so I excluded it from the final training.  \n\nIn total, the combined dataset finally consisted of **~2,000 images** synthetic images for training,  carefully curated to maximize diversity and domain relevance. As for the validation set it consisted of **77 images** real images.  \n\n> ```python\n> # Paths\n> base_path       = \"/kaggle/input/multi-class-object-detection-challenge/Starter_Dataset\"\n> synthetic_base  = \"/kaggle/input/extra-synthetic-data\"\n> synthetic_base2 = \"/kaggle/input/falcon-multiclass-cheerios-soupv2/falcon-multiclass-cheerios-soupV2\"\n> synthetic_base3 = \"/kaggle/working/synthetic_soup_extra\"\n>\n> # Gather synthetic train/val image dirs\n> synthetic_train_dirs = [str(p) for p in Path(synthetic_base).rglob(\"train/images\")]\n> synthetic_val_dirs   = [str(p) for p in Path(synthetic_base).rglob(\"val/images\")]\n>\n> # Gather synthetic train/val image dirs (dataset 2)\n> synthetic_train_dirs2 = [str(p) for p in Path(synthetic_base2).rglob(\"train/images\")]\n> synthetic_val_dirs2   = [str(p) for p in Path(synthetic_base2).rglob(\"val/images\")]\n>\n> # Gather synthetic train/val image dirs (dataset 3)\n> synthetic_train_dirs3 = [str(p) for p in Path(synthetic_base3).rglob(\"train/images\")]\n> synthetic_val_dirs3   = [str(p) for p in Path(synthetic_base3).rglob(\"val/images\")]\n>\n> # Build YAML dict\n> data_yaml = {\n>     \"train\": (\n>         [\n>             f\"{base_path}/train/images\",\n>             '/kaggle/input/sample-synthetic-data-generated/home/ubuntu/Output/2025-07-15-06-26-28/train/images',\n>             '/kaggle/input/sample-synthetic-data-generated/home/ubuntu/Output/2025-07-15-06-26-28/val/images',\n>         ]\n>         + synthetic_train_dirs\n>         + synthetic_val_dirs\n>         # + synthetic_train_dirs2\n>         # + synthetic_val_dirs2\n>         # + synthetic_train_dirs3\n>         # + synthetic_val_dirs3\n>     ),\n>     \"val\":   '/kaggle/working/val_real/images',\n>     \"test\":  f\"{base_path}/TestImages\",\n>     \"nc\":    2,\n>     \"names\": [\"cheerios\", \"Soup\"]\n> }\n>\n> # Save to data.yaml\n> with open(\"data.yaml\", \"w\") as f:\n>     yaml.safe_dump(data_yaml, f, default_flow_style=False)\n> ```\n\n>> rest f the functions can be found in the notebooks attached\n\n---\n## Training pipeline\n\nI used the Ultralytics implementation of YOLO to train three separate models with different hyperparameters on the combined dataset. The training notebook installed the `ultralytics` package and `ensemble‑boxes` for post‑processing. Key settings were:\n\n| Model | Backbone / Architecture | Image Size | Epochs | Batch Size | Learning Rate |weight decay|\n|-------|--------------------------|------------|--------|------------|---------------|------------|\n| Model 1 | yolo11x | 672 (with multiscale)| 12 | 4| 1e-3 |3e-3|\n| Model 2 | YOLO12x | 512 | 20 | 16| 1e-3 |1e-3|\n| Model 3 | YOLOv8x | 640| 30 | 32 | 1e-3 |2e-3|\n\n### Notes on Training and Hyperparameters\n\n- **Hyperparameter Selection**  \n  Most of the hyperparameters were tuned through extensive trial and error. Surprisingly, I found that **heavy augmentations** actually hurt performance and significantly increased training time. Instead, using **softer augmentations** (especially color/saturation-related) along with **MixUp** and **CutMix** yielded much better results. More importantly, adding **diverse synthetic images** turned out to be the single most effective way to improve generalization. *(Exact values can be found in the training notebook.)*\n\n- **Optimizer Choice (SGD)**  \n  I chose **SGD with momentum** instead of the default Adam/AdamW because it provided more stable convergence and better final accuracy in this competition setting. It also converged much faster, allowing the use of merely 15-30 epochs . \n\n- **Image Size**  \n  Training with **larger image sizes** offered noticeable boosts in mAP, especially for smaller objects like soup cans and cups. However, this came at the cost of much longer training times and occasional **out-of-memory (OOM) errors** on Kaggle GPUs. A balance had to be struck between accuracy gains and computational feasibility.\n\n- **Batch Size and Accumulation**  \n  Due to memory constraints with larger image sizes, I experimented with **smaller batch sizes combined with gradient accumulation**, which allowed stable training without compromising performance.\n\n- **Learning Rate Scheduling**  \n  A **cosine annealing schedule** with warm restarts worked best, enabling the model to escape local minima and generalize better. Step-based schedules were also tested but performed slightly worse.\n\n- **Freezing initial layers**  \n  Training only the last few layers was very helpful as it reduced training time and memory usage allowing the use of much larger image sizes and model architectures. \n\n- **Augmentation Insights**  \n     - *Color augmentations* (hue, saturation, brightness) were useful in **moderation**, but when applied too aggressively they actually **slowed convergence** and increased training time without delivering meaningful performance gains.  \n     - *Synthetic diversity* — introducing variations in **object placement, occlusion patterns, and lighting conditions** — proved to be **far more impactful** than heavy augmentations, offering better generalization and robustness.  \n     - *MixUp and CutMix* were particularly effective for this task, as they **regularized the models** by creating blended training examples and reducing overconfidence on single features. This helped the ensemble handle **class overlap and cluttered backgrounds** more reliably.  \n\n- **Validation Strategy**  \n  Using only the **real portion of the validation set** was crucial, since validating on synthetic images gave overly optimistic results and did not reflect the target domain.\n\n---\n## Ensembling and post‑processing\n\nAfter training the three models (YOLOv11x, YOLOv12x, YOLOv8x), I generated predictions on the test set and applied an ensembling pipeline followed by lightweight post-processing:\n\n1. **Model Ensembling with Weighted Box Fusion (WBF):**  \n   The outputs of the three models were combined using the `ensemble-boxes` library.  \n   - WBF merges highly overlapping bounding boxes across models by computing a **confidence-weighted average of their coordinates**.  \n   - This approach leverages the complementary strengths of each model:  \n     - YOLOv11x (672, multiscale) contributed robustness to scale variations.  \n     - YOLOv12x (512) captured efficiency and faster convergence.  \n     - YOLOv8x (640) provided strong generalisation and stability.  \n   - The ensemble was consistently more reliable than any single model, improving localisation accuracy and reducing duplicate detections.\n\n2. **Confidence Filtering (Top-k per class):**  \n   For each class in every image, I retained only the **top-2 predictions ranked by confidence**.  \n   - This simple heuristic removed low-confidence false positives without hurting recall.  \n   - It was especially effective in reducing spurious boxes around small or cluttered objects (cups, soup cans, etc.).  \n   - After this step, the final ensemble achieved a **mAP50 score of 0.996**, showing that even lightweight post-processing can significantly polish ensemble results.\n\n> ```python\n> import os\n> from pathlib import Path\n> import pandas as pd\n> import csv\n> from ultralytics import YOLO\n> from ensemble_boxes import weighted_boxes_fusion\n> from PIL import Image\n>\n> model_paths = [\n>     '/kaggle/working/runs1/train/train/weights/best.pt',\n>     '/kaggle/working/runs3/train/train/weights/best.pt',\n>     '/kaggle/working/runs2/train/train/weights/best.pt',\n> ]\n>\n> test_images_path = \"/kaggle/input/multi-class-object-detection-challenge/testImages/images\"\n> output_dir = \"/kaggle/working/predictions/labels\"\n>\n> conf = 0.0001\n> iou_thr = 0.3\n> skip_box_thr = conf\n> image_sizes = [640,800,864,1024,1216,512,1344,2048]\n>\n> models = [YOLO(path) for path in model_paths]\n> predictions = run_inference(models, image_sizes, test_images_path,conf=conf,iou_thr=iou_thr)\n>\n> image_ids = list(next(iter(next(iter(predictions.values())).values())).keys())\n>\n> apply_wbf_and_save_final_submission(predictions, image_ids,iou_thr=0.4,skip_box_thr=0.001,conf_post=0.14)\n> ```\n\n---\n## Reproducibility and tips\n\n- All data paths and hyper‑parameters are stored in the notebook’s `data.yaml` and `args.yaml`, so rerunning on similar hardware should reproduce the results.\n- When generating additional synthetic data in FalconCloud, vary the camera angles, lighting and object scales. This diversity helps the model generalise better.\n- If you have access to more powerful hardware, increasing the image size (e.g. 800) and batch size may yield further improvements. Larger YOLO models (like `yolov9e`) and longer training schedules are also worth exploring.\n\n---\n## 🏁 Closing Thoughts\n\nBringing together the **official competition dataset**, **multiple synthetic datasets**, and a **private dataset** resulted in a training set that was both **rich** and **diverse**. This variety played a key role in helping the models generalize better to real-world validation data.  \n\nTraining **three different YOLO models** and combining their predictions through **Weighted Box Fusion (WBF)** — followed by a lightweight **top-k filtering step** — turned out to be a simple yet very effective strategy. This ensemble approach consistently improved localisation, reduced duplicate detections, and ultimately delivered a **high mAP50 score** on the leaderboard.  \n\n\n### 💬 Feedback & Updates  \n- If you have any **comments or questions**, feel free to share them below — I’d be glad to discuss.  \n- I will continue to **update this write-up** if necessary.  \n\n---\n\n⭐ **If you found this helpful, please consider upvoting — it helps a lot. Thanks!**",
      "votes": null
    },
    {
      "id": "3275274",
      "postDate": "08/26/2025 08:50:34",
      "content": "<p>Congratulations! Thanks for the detailed solution, can you tell me how you performed the validation? Your private model has risen significantly in speed, how did you calculate this?</p>",
      "rawMarkdown": "Congratulations! Thanks for the detailed solution, can you tell me how you performed the validation? Your private model has risen significantly in speed, how did you calculate this?",
      "votes": null
    },
    {
      "id": "3276739",
      "postDate": "08/26/2025 21:49:27",
      "content": "<p>Thanks a lot! 🙏 I actually published an early incomplete draft by mistake — the updated version is now live, so I’d recommend giving it another read.</p>\n<p>For validation, I only used the real portion of the validation set, and I also relied on plotting and visual inspection of predictions to make sure the models generalized well.</p>",
      "rawMarkdown": "Thanks a lot! 🙏 I actually published an early incomplete draft by mistake — the updated version is now live, so I’d recommend giving it another read.\n\nFor validation, I only used the real portion of the validation set, and I also relied on plotting and visual inspection of predictions to make sure the models generalized well.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3275274,
      "author_name": "antonoof",
      "author_url": "",
      "post_date": "08/26/2025 08:50:34",
      "content": "<p>Congratulations! Thanks for the detailed solution, can you tell me how you performed the validation? Your private model has risen significantly in speed, how did you calculate this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3276739,
          "author_name": "thelastsmilodon",
          "author_url": "",
          "post_date": "08/26/2025 21:49:27",
          "content": "<p>Thanks a lot! 🙏 I actually published an early incomplete draft by mistake — the updated version is now live, so I’d recommend giving it another read.</p>\n<p>For validation, I only used the real portion of the validation set, and I also relied on plotting and visual inspection of predictions to make sure the models generalized well.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3275211": "# Solution Overview\n\nFirst off, a big thank you to the organisers for putting on this challenge and for hosting the live sessions. @rebekahduality @rishikeshjadhav22 \nIt was a lot of fun to take part and to see how other competitors tackled the problem. ❤️🤩\n\nI wanted to share how I approached this challenge and what worked for me (and what didnt) . Hopefully it’s helpful for anyone coming back to this problem later or just curious about how the top solutions were built.\n\n---\n\n## Data sources\n\nThe training data went well beyond the starter set. Here’s everything I used:\n\n1 - **Multi-class Object Detection Challenge** – the official competition dataset.  \nI used only the **train set** for training. For validation, I relied exclusively on the **real portion of the validation set** to ensure that the validation set better reflects the target domain.  \n(+1000 images)\n\n2 - [**Extra synthetic data**](https://www.kaggle.com/datasets/thelastsmilodon/extra-synthetic-data) – my main synthetic dataset generated using **FalconCloud**.  \nIt is divided into nine output folders (`output (1)` through `output (9)`), each containing its own `images` and `labels` subdirectories.  \n\nMy approach was two-staged:  \n- In the first 5–6 scenarios, I focused on generating as much data as possible to enrich the dataset.  \n- Afterwards, I relied on the **Controlled scenario** to address specific shortcomings observed during model training:  \n  * **Misclassification of cups, candles, and cylindrical objects as soup cans**  \n    Improved by introducing twin objects and capturing multiple views from varied angles and distances.  \n  * **Failure to detect occluded objects**  \n    Improved by creating scenarios with similar occlusions and ensuring diverse captures.  \n  * **Failure to detect distant soup cans**  \n    Improved by adding distant images under varied lighting and backgrounds.  \n  * **Limited diversity in object placement**  \n    Improved by frequently rearranging objects, varying orientation, and including edge-case scenarios.  \n\n(+904 images)\n\n3 - [**Falcon-Multiclass-Cheerios-Soup V2**](https://www.kaggle.com/datasets/kadirkrtls/falcon-multiclass-cheerios-soupv2) – a synthetic dataset generated via FalconCloud by @kadirkrtls (thanks for sharing).  \n*(Didn’t use it in my top notebook though)*\n\n4 - **Sample synthetic data generated** – a basic sample of synthetic images I created on FalconCloud and shared publicly at the start of the competition.  \n(+100 images)\n\n5 - [**Synthetic-soup1**](https://www.kaggle.com/datasets/thelastsmilodon/synthetic-soup1) – dataset from the previous soup detection competition.  \nI used it in several experiments, but the improvements (if any) were not worth the extra training time, so I excluded it from the final training.  \n\nIn total, the combined dataset finally consisted of **~2,000 images** synthetic images for training,  carefully curated to maximize diversity and domain relevance. As for the validation set it consisted of **77 images** real images.  \n\n> ```python\n> # Paths\n> base_path       = \"/kaggle/input/multi-class-object-detection-challenge/Starter_Dataset\"\n> synthetic_base  = \"/kaggle/input/extra-synthetic-data\"\n> synthetic_base2 = \"/kaggle/input/falcon-multiclass-cheerios-soupv2/falcon-multiclass-cheerios-soupV2\"\n> synthetic_base3 = \"/kaggle/working/synthetic_soup_extra\"\n>\n> # Gather synthetic train/val image dirs\n> synthetic_train_dirs = [str(p) for p in Path(synthetic_base).rglob(\"train/images\")]\n> synthetic_val_dirs   = [str(p) for p in Path(synthetic_base).rglob(\"val/images\")]\n>\n> # Gather synthetic train/val image dirs (dataset 2)\n> synthetic_train_dirs2 = [str(p) for p in Path(synthetic_base2).rglob(\"train/images\")]\n> synthetic_val_dirs2   = [str(p) for p in Path(synthetic_base2).rglob(\"val/images\")]\n>\n> # Gather synthetic train/val image dirs (dataset 3)\n> synthetic_train_dirs3 = [str(p) for p in Path(synthetic_base3).rglob(\"train/images\")]\n> synthetic_val_dirs3   = [str(p) for p in Path(synthetic_base3).rglob(\"val/images\")]\n>\n> # Build YAML dict\n> data_yaml = {\n>     \"train\": (\n>         [\n>             f\"{base_path}/train/images\",\n>             '/kaggle/input/sample-synthetic-data-generated/home/ubuntu/Output/2025-07-15-06-26-28/train/images',\n>             '/kaggle/input/sample-synthetic-data-generated/home/ubuntu/Output/2025-07-15-06-26-28/val/images',\n>         ]\n>         + synthetic_train_dirs\n>         + synthetic_val_dirs\n>         # + synthetic_train_dirs2\n>         # + synthetic_val_dirs2\n>         # + synthetic_train_dirs3\n>         # + synthetic_val_dirs3\n>     ),\n>     \"val\":   '/kaggle/working/val_real/images',\n>     \"test\":  f\"{base_path}/TestImages\",\n>     \"nc\":    2,\n>     \"names\": [\"cheerios\", \"Soup\"]\n> }\n>\n> # Save to data.yaml\n> with open(\"data.yaml\", \"w\") as f:\n>     yaml.safe_dump(data_yaml, f, default_flow_style=False)\n> ```\n\n>> rest f the functions can be found in the notebooks attached\n\n---\n## Training pipeline\n\nI used the Ultralytics implementation of YOLO to train three separate models with different hyperparameters on the combined dataset. The training notebook installed the `ultralytics` package and `ensemble‑boxes` for post‑processing. Key settings were:\n\n| Model | Backbone / Architecture | Image Size | Epochs | Batch Size | Learning Rate |weight decay|\n|-------|--------------------------|------------|--------|------------|---------------|------------|\n| Model 1 | yolo11x | 672 (with multiscale)| 12 | 4| 1e-3 |3e-3|\n| Model 2 | YOLO12x | 512 | 20 | 16| 1e-3 |1e-3|\n| Model 3 | YOLOv8x | 640| 30 | 32 | 1e-3 |2e-3|\n\n### Notes on Training and Hyperparameters\n\n- **Hyperparameter Selection**  \n  Most of the hyperparameters were tuned through extensive trial and error. Surprisingly, I found that **heavy augmentations** actually hurt performance and significantly increased training time. Instead, using **softer augmentations** (especially color/saturation-related) along with **MixUp** and **CutMix** yielded much better results. More importantly, adding **diverse synthetic images** turned out to be the single most effective way to improve generalization. *(Exact values can be found in the training notebook.)*\n\n- **Optimizer Choice (SGD)**  \n  I chose **SGD with momentum** instead of the default Adam/AdamW because it provided more stable convergence and better final accuracy in this competition setting. It also converged much faster, allowing the use of merely 15-30 epochs . \n\n- **Image Size**  \n  Training with **larger image sizes** offered noticeable boosts in mAP, especially for smaller objects like soup cans and cups. However, this came at the cost of much longer training times and occasional **out-of-memory (OOM) errors** on Kaggle GPUs. A balance had to be struck between accuracy gains and computational feasibility.\n\n- **Batch Size and Accumulation**  \n  Due to memory constraints with larger image sizes, I experimented with **smaller batch sizes combined with gradient accumulation**, which allowed stable training without compromising performance.\n\n- **Learning Rate Scheduling**  \n  A **cosine annealing schedule** with warm restarts worked best, enabling the model to escape local minima and generalize better. Step-based schedules were also tested but performed slightly worse.\n\n- **Freezing initial layers**  \n  Training only the last few layers was very helpful as it reduced training time and memory usage allowing the use of much larger image sizes and model architectures. \n\n- **Augmentation Insights**  \n     - *Color augmentations* (hue, saturation, brightness) were useful in **moderation**, but when applied too aggressively they actually **slowed convergence** and increased training time without delivering meaningful performance gains.  \n     - *Synthetic diversity* — introducing variations in **object placement, occlusion patterns, and lighting conditions** — proved to be **far more impactful** than heavy augmentations, offering better generalization and robustness.  \n     - *MixUp and CutMix* were particularly effective for this task, as they **regularized the models** by creating blended training examples and reducing overconfidence on single features. This helped the ensemble handle **class overlap and cluttered backgrounds** more reliably.  \n\n- **Validation Strategy**  \n  Using only the **real portion of the validation set** was crucial, since validating on synthetic images gave overly optimistic results and did not reflect the target domain.\n\n---\n## Ensembling and post‑processing\n\nAfter training the three models (YOLOv11x, YOLOv12x, YOLOv8x), I generated predictions on the test set and applied an ensembling pipeline followed by lightweight post-processing:\n\n1. **Model Ensembling with Weighted Box Fusion (WBF):**  \n   The outputs of the three models were combined using the `ensemble-boxes` library.  \n   - WBF merges highly overlapping bounding boxes across models by computing a **confidence-weighted average of their coordinates**.  \n   - This approach leverages the complementary strengths of each model:  \n     - YOLOv11x (672, multiscale) contributed robustness to scale variations.  \n     - YOLOv12x (512) captured efficiency and faster convergence.  \n     - YOLOv8x (640) provided strong generalisation and stability.  \n   - The ensemble was consistently more reliable than any single model, improving localisation accuracy and reducing duplicate detections.\n\n2. **Confidence Filtering (Top-k per class):**  \n   For each class in every image, I retained only the **top-2 predictions ranked by confidence**.  \n   - This simple heuristic removed low-confidence false positives without hurting recall.  \n   - It was especially effective in reducing spurious boxes around small or cluttered objects (cups, soup cans, etc.).  \n   - After this step, the final ensemble achieved a **mAP50 score of 0.996**, showing that even lightweight post-processing can significantly polish ensemble results.\n\n> ```python\n> import os\n> from pathlib import Path\n> import pandas as pd\n> import csv\n> from ultralytics import YOLO\n> from ensemble_boxes import weighted_boxes_fusion\n> from PIL import Image\n>\n> model_paths = [\n>     '/kaggle/working/runs1/train/train/weights/best.pt',\n>     '/kaggle/working/runs3/train/train/weights/best.pt',\n>     '/kaggle/working/runs2/train/train/weights/best.pt',\n> ]\n>\n> test_images_path = \"/kaggle/input/multi-class-object-detection-challenge/testImages/images\"\n> output_dir = \"/kaggle/working/predictions/labels\"\n>\n> conf = 0.0001\n> iou_thr = 0.3\n> skip_box_thr = conf\n> image_sizes = [640,800,864,1024,1216,512,1344,2048]\n>\n> models = [YOLO(path) for path in model_paths]\n> predictions = run_inference(models, image_sizes, test_images_path,conf=conf,iou_thr=iou_thr)\n>\n> image_ids = list(next(iter(next(iter(predictions.values())).values())).keys())\n>\n> apply_wbf_and_save_final_submission(predictions, image_ids,iou_thr=0.4,skip_box_thr=0.001,conf_post=0.14)\n> ```\n\n---\n## Reproducibility and tips\n\n- All data paths and hyper‑parameters are stored in the notebook’s `data.yaml` and `args.yaml`, so rerunning on similar hardware should reproduce the results.\n- When generating additional synthetic data in FalconCloud, vary the camera angles, lighting and object scales. This diversity helps the model generalise better.\n- If you have access to more powerful hardware, increasing the image size (e.g. 800) and batch size may yield further improvements. Larger YOLO models (like `yolov9e`) and longer training schedules are also worth exploring.\n\n---\n## 🏁 Closing Thoughts\n\nBringing together the **official competition dataset**, **multiple synthetic datasets**, and a **private dataset** resulted in a training set that was both **rich** and **diverse**. This variety played a key role in helping the models generalize better to real-world validation data.  \n\nTraining **three different YOLO models** and combining their predictions through **Weighted Box Fusion (WBF)** — followed by a lightweight **top-k filtering step** — turned out to be a simple yet very effective strategy. This ensemble approach consistently improved localisation, reduced duplicate detections, and ultimately delivered a **high mAP50 score** on the leaderboard.  \n\n\n### 💬 Feedback & Updates  \n- If you have any **comments or questions**, feel free to share them below — I’d be glad to discuss.  \n- I will continue to **update this write-up** if necessary.  \n\n---\n\n⭐ **If you found this helpful, please consider upvoting — it helps a lot. Thanks!**",
    "3275274": "Congratulations! Thanks for the detailed solution, can you tell me how you performed the validation? Your private model has risen significantly in speed, how did you calculate this?",
    "3276739": "Thanks a lot! 🙏 I actually published an early incomplete draft by mistake — the updated version is now live, so I’d recommend giving it another read.\n\nFor validation, I only used the real portion of the validation set, and I also relied on plotting and visual inspection of predictions to make sure the models generalized well."
  },
  "source": "meta"
}