{
  "id": 667479,
  "title": "4th Place Solution",
  "url": "/competitions/duality-ai-lunate-ai-geospatial-object-detection/writeups/4th-place-solution",
  "author_name": "",
  "post_date": "2026-01-12T21:04:26.783Z",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p>👋 Hello everyone, in this writeup we describe the approach used by our team in this competition.</p>\n<h2>Preprocessing</h2>\n<p>We used several independently trained YOLO models.<br>\nImages were converted to RGB and processed at multiple input resolutions.  Degenerate bounding boxes were removed before ensembling.</p>\n<h2>Models</h2>\n<p>Our final solution is based on an ensemble of different YOLO architectures:</p>\n<ul>\n<li>YOLOv8x  </li>\n<li>YOLO11x  </li>\n<li>YOLO12x  </li>\n</ul>\n<p>Each model was trained separately with its own configuration and then used together at inference time.</p>\n<h2>Multi-Scale Inference</h2>\n<p>For each model, predictions were generated at multiple image resolutions:</p>\n<p><code>image_sizes = [1024, 1280, 1440, 1600]</code></p>\n<p>This allows the models to better detect objects of different scales.<br>\nAlthough resizing images can reduce fine-grained details, we found this approach to be more stable than sliding-window methods, especially for objects that appear at the global image level.</p>\n<h2>Ensembling with Weighted Boxes Fusion</h2>\n<p>All predictions from different models and image sizes were merged using <strong>Weighted Boxes Fusion (WBF)</strong>.</p>\n<p>The idea is that each model and resolution acts as a weak detector.<br>\nWhen multiple predictions agree spatially, WBF aligns them and assigns a higher confidence to consistent detections.</p>\n<p>This strategy helps:</p>\n<ul>\n<li>suppress isolated false positives  </li>\n<li>stabilize localization across scales  </li>\n<li>improve robustness without complex post-processing  </li>\n</ul>\n<h2>Reproducibility</h2>\n<p>Full inference and ensembling pipeline:<br>\n<a href=\"https://www.kaggle.com/code/tmdofi/preds-ensemble?scriptVersionId=289277492\" target=\"_blank\">Preds ensemble</a></p>\n<h2>Conclusion</h2>\n<p>In short, the solution is based on model diversity, multi-scale inference, and box-level ensembling.</p>\n<p>We would be happy to see how the top-3 teams approached this task and what strategies they found most effective</p>\n<p>Thanks to the organizers for an interesting and well-designed competition!</p>",
  "messages": [
    {
      "id": "3390251",
      "postDate": "01/12/2026 21:00:30",
      "content": "<p>👋 Hello everyone, in this writeup we describe the approach used by our team in this competition.</p>\n<h2>Preprocessing</h2>\n<p>We used several independently trained YOLO models.<br>\nImages were converted to RGB and processed at multiple input resolutions.  Degenerate bounding boxes were removed before ensembling.</p>\n<h2>Models</h2>\n<p>Our final solution is based on an ensemble of different YOLO architectures:</p>\n<ul>\n<li>YOLOv8x  </li>\n<li>YOLO11x  </li>\n<li>YOLO12x  </li>\n</ul>\n<p>Each model was trained separately with its own configuration and then used together at inference time.</p>\n<h2>Multi-Scale Inference</h2>\n<p>For each model, predictions were generated at multiple image resolutions:</p>\n<p><code>image_sizes = [1024, 1280, 1440, 1600]</code></p>\n<p>This allows the models to better detect objects of different scales.<br>\nAlthough resizing images can reduce fine-grained details, we found this approach to be more stable than sliding-window methods, especially for objects that appear at the global image level.</p>\n<h2>Ensembling with Weighted Boxes Fusion</h2>\n<p>All predictions from different models and image sizes were merged using <strong>Weighted Boxes Fusion (WBF)</strong>.</p>\n<p>The idea is that each model and resolution acts as a weak detector.<br>\nWhen multiple predictions agree spatially, WBF aligns them and assigns a higher confidence to consistent detections.</p>\n<p>This strategy helps:</p>\n<ul>\n<li>suppress isolated false positives  </li>\n<li>stabilize localization across scales  </li>\n<li>improve robustness without complex post-processing  </li>\n</ul>\n<h2>Reproducibility</h2>\n<p>Full inference and ensembling pipeline:<br>\n<a href=\"https://www.kaggle.com/code/tmdofi/preds-ensemble?scriptVersionId=289277492\" target=\"_blank\">Preds ensemble</a></p>\n<h2>Conclusion</h2>\n<p>In short, the solution is based on model diversity, multi-scale inference, and box-level ensembling.</p>\n<p>We would be happy to see how the top-3 teams approached this task and what strategies they found most effective</p>\n<p>Thanks to the organizers for an interesting and well-designed competition!</p>",
      "rawMarkdown": "👋 Hello everyone, in this writeup we describe the approach used by our team in this competition.\n\n\n## Preprocessing\nWe used several independently trained YOLO models.  \nImages were converted to RGB and processed at multiple input resolutions.  Degenerate bounding boxes were removed before ensembling.\n\n\n## Models\nOur final solution is based on an ensemble of different YOLO architectures:\n\n- YOLOv8x  \n- YOLO11x  \n- YOLO12x  \n\nEach model was trained separately with its own configuration and then used together at inference time.\n\n## Multi-Scale Inference\nFor each model, predictions were generated at multiple image resolutions:\n\n`image_sizes = [1024, 1280, 1440, 1600]`\n\nThis allows the models to better detect objects of different scales.  \nAlthough resizing images can reduce fine-grained details, we found this approach to be more stable than sliding-window methods, especially for objects that appear at the global image level.\n\n\n## Ensembling with Weighted Boxes Fusion\nAll predictions from different models and image sizes were merged using **Weighted Boxes Fusion (WBF)**.\n\nThe idea is that each model and resolution acts as a weak detector.  \nWhen multiple predictions agree spatially, WBF aligns them and assigns a higher confidence to consistent detections.\n\nThis strategy helps:\n- suppress isolated false positives  \n- stabilize localization across scales  \n- improve robustness without complex post-processing  \n\n\n## Reproducibility\nFull inference and ensembling pipeline:  \n[Preds ensemble](https://www.kaggle.com/code/tmdofi/preds-ensemble?scriptVersionId=289277492)\n\n## Conclusion\nIn short, the solution is based on model diversity, multi-scale inference, and box-level ensembling.\n\nWe would be happy to see how the top-3 teams approached this task and what strategies they found most effective\n\nThanks to the organizers for an interesting and well-designed competition!",
      "votes": null
    },
    {
      "id": "3390266",
      "postDate": "01/12/2026 21:46:23",
      "content": "<p>Great breakdown, thank you for sharing!</p>",
      "rawMarkdown": "Great breakdown, thank you for sharing!",
      "votes": null
    },
    {
      "id": "3390407",
      "postDate": "01/13/2026 05:16:40",
      "content": "<p>Actually, the code pipeline is almost the same for everyone. The final score mainly depends on who prepared a better dataset. Since Falcon is a generative model, each participant ends up with different data quality, which directly affects the results.</p>",
      "rawMarkdown": "Actually, the code pipeline is almost the same for everyone. The final score mainly depends on who prepared a better dataset. Since Falcon is a generative model, each participant ends up with different data quality, which directly affects the results.",
      "votes": null
    },
    {
      "id": "3390630",
      "postDate": "01/13/2026 15:02:28",
      "content": "<p>yeah, but i think SAHI in this task is better</p>",
      "rawMarkdown": "yeah, but i think SAHI in this task is better",
      "votes": null
    },
    {
      "id": "3390704",
      "postDate": "01/13/2026 19:38:33",
      "content": "<p>We do ensure that the scenarios we release are able to produce the needed quality of data before launching the competition. Would you be interested in live calls for each competition where go over strategies for generating effective data? We are very down for hosting those if people think it would be useful!\nSynthetic data use does require the ability to align the synthetic data characteristics with the real-world needs. This is a major service we provide our commercial customers, so we include it as a crucial part of our competitions. </p>",
      "rawMarkdown": "We do ensure that the scenarios we release are able to produce the needed quality of data before launching the competition. Would you be interested in live calls for each competition where go over strategies for generating effective data? We are very down for hosting those if people think it would be useful!\nSynthetic data use does require the ability to align the synthetic data characteristics with the real-world needs. This is a major service we provide our commercial customers, so we include it as a crucial part of our competitions.",
      "votes": null
    },
    {
      "id": "3390709",
      "postDate": "01/13/2026 20:11:08",
      "content": "<p>Are there any measurements or assumptions about how SAHI works? will it be better in this competition?</p>",
      "rawMarkdown": "Are there any measurements or assumptions about how SAHI works? will it be better in this competition?",
      "votes": null
    },
    {
      "id": "3390734",
      "postDate": "01/13/2026 21:38:38",
      "content": "<p>We don't have any measurements, but this would be a great competition to use SAHI. One of the downsides of traditional training is that objects with fewer pixels represented could get lost if images get scaled down. Since this competition has the challenge of detecting smaller objects, that's a huge problem.\nSAHI would help bypass that problem, so it would be a great strategy!</p>",
      "rawMarkdown": "We don't have any measurements, but this would be a great competition to use SAHI. One of the downsides of traditional training is that objects with fewer pixels represented could get lost if images get scaled down. Since this competition has the challenge of detecting smaller objects, that's a huge problem.\nSAHI would help bypass that problem, so it would be a great strategy!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3390266,
      "author_name": "rebekahduality",
      "author_url": "",
      "post_date": "01/12/2026 21:46:23",
      "content": "<p>Great breakdown, thank you for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3390407,
      "author_name": "siamarefin",
      "author_url": "",
      "post_date": "01/13/2026 05:16:40",
      "content": "<p>Actually, the code pipeline is almost the same for everyone. The final score mainly depends on who prepared a better dataset. Since Falcon is a generative model, each participant ends up with different data quality, which directly affects the results.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3390630,
          "author_name": "antonoof",
          "author_url": "",
          "post_date": "01/13/2026 15:02:28",
          "content": "<p>yeah, but i think SAHI in this task is better</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3390704,
          "author_name": "rebekahduality",
          "author_url": "",
          "post_date": "01/13/2026 19:38:33",
          "content": "<p>We do ensure that the scenarios we release are able to produce the needed quality of data before launching the competition. Would you be interested in live calls for each competition where go over strategies for generating effective data? We are very down for hosting those if people think it would be useful!\nSynthetic data use does require the ability to align the synthetic data characteristics with the real-world needs. This is a major service we provide our commercial customers, so we include it as a crucial part of our competitions. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3390709,
              "author_name": "antonoof",
              "author_url": "",
              "post_date": "01/13/2026 20:11:08",
              "content": "<p>Are there any measurements or assumptions about how SAHI works? will it be better in this competition?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3390734,
                  "author_name": "rebekahduality",
                  "author_url": "",
                  "post_date": "01/13/2026 21:38:38",
                  "content": "<p>We don't have any measurements, but this would be a great competition to use SAHI. One of the downsides of traditional training is that objects with fewer pixels represented could get lost if images get scaled down. Since this competition has the challenge of detecting smaller objects, that's a huge problem.\nSAHI would help bypass that problem, so it would be a great strategy!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3390251": "👋 Hello everyone, in this writeup we describe the approach used by our team in this competition.\n\n\n## Preprocessing\nWe used several independently trained YOLO models.  \nImages were converted to RGB and processed at multiple input resolutions.  Degenerate bounding boxes were removed before ensembling.\n\n\n## Models\nOur final solution is based on an ensemble of different YOLO architectures:\n\n- YOLOv8x  \n- YOLO11x  \n- YOLO12x  \n\nEach model was trained separately with its own configuration and then used together at inference time.\n\n## Multi-Scale Inference\nFor each model, predictions were generated at multiple image resolutions:\n\n`image_sizes = [1024, 1280, 1440, 1600]`\n\nThis allows the models to better detect objects of different scales.  \nAlthough resizing images can reduce fine-grained details, we found this approach to be more stable than sliding-window methods, especially for objects that appear at the global image level.\n\n\n## Ensembling with Weighted Boxes Fusion\nAll predictions from different models and image sizes were merged using **Weighted Boxes Fusion (WBF)**.\n\nThe idea is that each model and resolution acts as a weak detector.  \nWhen multiple predictions agree spatially, WBF aligns them and assigns a higher confidence to consistent detections.\n\nThis strategy helps:\n- suppress isolated false positives  \n- stabilize localization across scales  \n- improve robustness without complex post-processing  \n\n\n## Reproducibility\nFull inference and ensembling pipeline:  \n[Preds ensemble](https://www.kaggle.com/code/tmdofi/preds-ensemble?scriptVersionId=289277492)\n\n## Conclusion\nIn short, the solution is based on model diversity, multi-scale inference, and box-level ensembling.\n\nWe would be happy to see how the top-3 teams approached this task and what strategies they found most effective\n\nThanks to the organizers for an interesting and well-designed competition!",
    "3390266": "Great breakdown, thank you for sharing!",
    "3390407": "Actually, the code pipeline is almost the same for everyone. The final score mainly depends on who prepared a better dataset. Since Falcon is a generative model, each participant ends up with different data quality, which directly affects the results.",
    "3390630": "yeah, but i think SAHI in this task is better",
    "3390704": "We do ensure that the scenarios we release are able to produce the needed quality of data before launching the competition. Would you be interested in live calls for each competition where go over strategies for generating effective data? We are very down for hosting those if people think it would be useful!\nSynthetic data use does require the ability to align the synthetic data characteristics with the real-world needs. This is a major service we provide our commercial customers, so we include it as a crucial part of our competitions.",
    "3390709": "Are there any measurements or assumptions about how SAHI works? will it be better in this competition?",
    "3390734": "We don't have any measurements, but this would be a great competition to use SAHI. One of the downsides of traditional training is that objects with fewer pixels represented could get lost if images get scaled down. Since this competition has the challenge of detecting smaller objects, that's a huge problem.\nSAHI would help bypass that problem, so it would be a great strategy!"
  },
  "source": "meta"
}